Hongshuo Tian

dblp:276/3242 · DBLP profile ↗
← Back
28ranked-venue papers
4as first author
27since 2021 · last 2027
0000-0001-7635-0961ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 4 first-author · 21 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2027 Beyond content: A dual-channel approach for social bot detection via unmasking behavioral sequence camouflage
Hongshuo Tian, Jinlin Guo, Xianzhu Liu, Ning Xu 0003, Lanjun Wang
Expert Syst. Appl.2
2026 Towards social-aware image captioning via chain-of-thought prompting
Shenyuan Zhang, Ning Xu 0003, Quanhan Wu, Jinlin Guo, Hongshuo Tian, Anan Liu
Expert Syst. Appl.6
2026 Exploring fine-grained multimodal prompts for visual relation detection with adaptation of VLMs
Hongshuo Tian
Multim. Syst.2
2026 Medical VLP Model Is Vulnerable: Toward Multimodal Adversarial Attack on Large Medical Vision-Language Models
abstract
Medical Visual Question Answering (Medical VQA) is an essential task that facilitates the automated interpretation of complex clinical imagery with corresponding textual questions, thereby supporting both clinicians and patients in making informed medical decisions. With the rapid progress of Vision-Language Pretraining (VLP) in general domains, the development of medical VLP models has emerged as a rapidly growing interdisciplinary area at the intersection of artificial intelligence (AI) and healthcare. However, few works have been proposed to evaluate the adversarial robustness of medical VLP models, which faces two primary challenges: (1) the complexity of medical texts, stemming from the presence of terminologies, poses significant challenges for models in comprehending the text for adversarial attack; (2) the diversity of medical images arises from the variety of anatomical regions depicted, which requires models to determine critical anatomical regions for attack. In this paper, we propose a novel multimodal adversarial attack generator for evaluating the robustness of medical VLP models. Specifically, for the complexity of medical texts, we integrate medical knowledge when crafting text adversarial samples, which can facilitate the terminologies understanding and adversarial strength; for the diversity of medical images, we divide the anatomical regions into either global or local regions in medical images, which are determined by learned balance weights for perturbations. Our experimental study not only provides a quantitative understanding in medical VLP models, but also underscores the critical need for thorough safety evaluations before implementing them in real-world medical applications.
Zimu Lu, Ning Xu 0003, Hongshuo Tian, Lanjun Wang, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.3
2026 MEF-GD: Multimodal Enhancement and Fusion Network for Garment Designer
abstract
In recent years, with advancements in generative models, an increasing number of garment design methods have been proposed. A generative model capable of generating garment images from text and sketches can provide designers with valuable visual references and creative inspiration to aid in the design process. Existing multimodal garment design methods face the challenge of lacking precise control over the generated results in relation to both sketches and text. In this paper, we propose Multimodal Enhancement and Fusion Network for Garment Design (MEF-GD). Our model inputs image conditions into Stable Diffusion based on ControlNet. On one hand, directly inputting image conditions can lead to feature forgetting, defined as the phenomenon in deep neural networks where previously learned feature representations are lost. To address this issue, we propose a multiple feature injection module to more effectively enhance image condition features. On the other hand, ControlNet fuses control features into Stable Diffusion through pointwise addition, which ignores the interaction between multimodal features and results in the fused features being biased towards the control features, overlooking Stable Diffusion features. To address this limitation, we introduce content-guided attention for more effective feature fusion and improve the expression of text features. Additionally, existing datasets often contain vague textual descriptions of garments. It is difficult to train the model on such a dataset to learn accurate alignment between generated image and the textual descriptions. To address this issue, we have designed a multimodal large model text optimization module to improve the quality and clarity of text generation. Compared to existing multimodal garment design methods, MEF-GD achieves more effective alignment with both textual and sketch-based inputs in generating garment images. Compared to MGD, MEF-GD achieves a decrease of 2.44 in FID and an increase of 0.83 in CLIP Score on Multi-VITON-HD dataset. The code will be available at https://github.com/fengyun691340/MEF-GD.
Dan Song 0006, Jianhao Zeng, Hongshuo Tian, Bolun Zheng, Rongbao Kang, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.4
2026 Constituency-Tree-Induced Vision-Language Alignment for Multimodal Large Language Models
abstract
Multimodal large language models (MLLMs) integrate sophisticated large vision models (LVMs) to empower large language models (LLMs) with vision ability to perceive, reason, and interact in vision-language (V-L) tasks, while the modality bridge between two specialists becomes the bottleneck that translates visual signals into linguistic representations. However, most of the existing methods train the modality bridge with coarse-grained image-text pairs, neglecting the structural mapping between V-L semantics that facilitates modality translation from LVMs to LLMs. To mitigate this, we propose a Constituency-Tree-Induced Multimodal Bridging mechanism (CTIMB) that learns the fine-grained connection from LVMs to LLMs by the structural guidance from multi-modal constituency tree. Our approach consists of: 1) the multi-modal constituency-tree parser that jointly exploits the semantic structure of vision and language; 2) the lightweight connector that translates visual signals into linguistic representation and re-arranges them according to the constituency-tree structure; 3) the dynamic construction loss that aids in aligning the semantic structures derived from the tree parser and the connector. The CTIMB can learn the fine-grained mapping between visual and linguistic semantics, seamlessly bridge the LVMs and LLMs to enhance V-L tasks, and is more cost-efficient compared with current methods. Extensive experiments have demonstrated that our method more accurately interprets the visual features, enabling LLMs to conduct downstream tasks more effectively, and achieve superior performance with less training cost.
Yingchen Zhai, Ning Xu 0003, Hongshuo Tian, Bolun Zheng, Chenggang Yan 0001, Jinbo Cao, Rongbao Kang, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.3
2026 MMToT: Multi-Modal Token-of-Thought Reasoning for Large Models
abstract
With the development of Large models (LMs), recent methods tend to leverage them to complete various downstream tasks, like VQA and image caption. A typical method is Chain-of-Thought (CoT) prompting, which improves the reasoning abilities of LMs by providing intermediate steps. However, existing CoT prompts have two main drawbacks: 1) They typically represent CoT through multiple sentences, which often introduce irrelevant textual or visual context that may confuse LMs. 2) The current CoT methods fail to consider how the contributions of different tokens vary for answer inference. For this, we propose the Multi-Modal Token-of-Thought (MMToT), a novel token-level prompt method to improve LMs' multi-modal reasoning capabilities. Furthermore, MMToT stress on two strengths against CoT: 1) To prevent LMs from being affected by irrelevant contexts, we propose to extract explicit multi-modal tokens rather than sentences to construct MMToT, enhancing the reliability of generated answers. 2) To ensure that LMs prioritize tokens with high contribution scores during answer generation, we propose a confident decision-making module to evaluate and integrate each token's contribution in MMToT. Compared to existing methods, the proposed MMToT demonstrates superior performance on Science-QA, MATH, OKVQA, and VQA-introspect datasets. Furthermore, ablation studies and visualization results validate the effectiveness and interpretability of MMToT.
Ning Xu 0003, Zimu Lu, Hongshuo Tian, Bolun Zheng, Jinbo Cao, Anan Liu
IEEE Trans. Multim.3
2026 Knowledge and multi-detail enhanced GAN for human-driven text-to-image synthesis
abstract
Human-driven text-to-image synthesis aims to create controllable images, which not only adhere to the semantic of given text but also incorporate the visual characteristics of given human. For example, given “a man on the beach” (text) along with a photo of human, the model aims to generate an image depicting the human on the beach. Although current diffusion-based methods have shown promise in this task, they face two major limitations: (1) The generated images appear to be a bit stiff and unnatural, almost like collages of human and backgrounds; (2) The details of human in the generated image are inconsistent with those in the input, losing the original identity. To address these issues, we present the Knowledge and Multi-Detail Enhanced GAN for the task of human-driven text-to-image synthesis. It employs external knowledge as references to improve the harmony between human and backgrounds, and uses CLIP’s multi-layer features to intensify human details. First, we search the database to retrieve external images that are similar to the given text, serving as our knowledge. Second, to preserve the human details, we present the Multi-Detail Enhancer, which uses the image encoder of CLIP to extract human representation at multiple levels. Third, to enhance the human-background naturalness, we present the Knowledge Attention Enhancer, which can seamlessly blend human, text, and knowledge by attentively retain useful information and filter out noise from knowledge. Finally, we introduce the dual discriminators to guide the entire network, which can facilitate the accurate capture of human details and generation of images. Extensive experiments demonstrate the superiority of our method with its efficiency and lower computational demands. It is about 300 times faster than diffusion-based models, uses only 5% of the parameters, and completes training in just two days on three V100 GPUs.
Ning Xu 0003, Zhewen Shen, Hongshuo Tian, Bolun Zheng, Chenggang Yan 0001, Jinbo Cao, Rongbao Kang, Anan Liu
Vis. Informatics3
2025 When Headlines Meet Minds: Empowering News Recommendations with Social Simulator
abstract
Personalized news recommendation aims to deliver content aligned with user interests. However, most existing methods rely on the objective textual content of news, overlooking the subjective social review that reflects how the news is socially perceived. Inspired by social constructionism, we propose Social Review-aware Recommendation (SRec), a novel framework that integrates both objective content and the social review. The latter is constructed through group deliberation modeled by an agent-based social simulator, providing structured representations of collective understandings toward news. In addition, SRec incorporates a reasoning-guided explanation module that produces interpretable rationales by aligning user preferences with the social review of news. Experimental results on the MIND-small and MIND-large datasets demonstrate that SRec improves AUC by at least 2.45% over competitive baselines. Further analysis confirms the value of the social review generated by the simulator, and shows the flexibility of SRec as a lightweight enhancement to existing recommendation systems.
Yanwei Xie, Weizhi Nie, Lanjun Wang, Hongshuo Tian, Changtai Shi, Anan Liu
ACM Multimedia4
2025 Mixture of causal experts: A causal perspective to build dual-level mixture-of-experts models
Ning Xu 0003, Hongshuo Tian, Bolun Zheng, Jinbo Cao, Anan Liu
Expert Syst. Appl.3
2025 Bidirectional Mask Selection for Zero-Shot Referring Image Segmentation
abstract
Zero-shot referring image segmentation (RIS) aims to segment a referent mask via a natural language expression, without any training. Although existing research has made some progress, the lack of a training process in zero-shot learning results in insufficient information, leading to poor zero-shot segmentation performance. We propose a Bidirectional Mask Selection (BMS) framework, which is the first work to incorporate the negative masks into zero-shot RIS. Our idea is based on leveraging the negative masks’ semantic context information around target semantic to enhance the understanding of cross-modal fine-grained correlation. Further, we propose a novel mask adaptive fusion strategy to combine the complementary information from positive and negative masks without additional training. In the experiments, BMS has demonstrated outstanding performance on three prominent RIS datasets, and it has surpassed even the most advanced weakly supervised methods on the RefCOCOg datasets. Code will be available athttps://github.com/pcc-99/BMS.
Wenhui Li 0001, Weizhi Nie, Hongshuo Tian, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.4
2025 ClipMix for Domain Generalization
abstract
Domain Generalization (DG) is a growing field in machine learning that aims to train the model across multiple source domains, thereby enabling effective generalization to new, unseen target domains. Recent studies suggest that data augmentation, which enhances the diversity of the source domain, might be a promising solution to address this task. Current data augmentation methods use random fusion coefficients or local regional fusion, which cannot adaptively design the weights based on data, or preserve the integrity of original semantics. Inspired by the pre-trained model CLIP, which contains extensive multimodal knowledge, we propose ClipMix to address these limitations. Firstly, we use the CLIP model as the external knowledge to adaptively evaluate the alignment between images and their labels, using this alignment to assess the complexity of learning each image and guide adaptive augmentation. Secondly, we implement a label shift mechanism to dynamically assign soft labels to fused images, helping the model focus on hard-to-learn patterns and also gather domain-agnostic representation. Furthermore, we enhance the diversity of fused images at both the pixel and feature levels. Experimental results across sixteen domains from four databases verify the effectiveness of our method.
Anan Liu, Hao-Chen Li, Wenhui Li 0001, Dan Song 0006, Hongshuo Tian, Lanjun Wang
IEEE Trans. Circuits Syst. Video Technol.5
2025 Enriched Image Captioning Based on Knowledge Divergence and Focus
abstract
Image captioning is a fundamental task in computer vision that aims to generate precise and comprehensive descriptions of images automatically. Intuitively, humans initially rely on the image content, e.g., “cake on a plate”, to gradually gather relevant knowledge facts e.g., “birthday party”, “candles”, which is a process referred to as divergence. Then, we perform step-by-step reasoning based on the images to refine, and rearrange these knowledge facts for explicit sentence generation, a process referred to as focus. However, existing image captioning methods mainly rely on the encode-decode framework that does not well fit the “divergence-focus” nature of the task. To this end, we propose the knowledge “divergence-focus” method for Image Captioning (K-DFIC) to gather and polish knowledge facts for image understanding, which consists of two components: (a) Knowledge Divergence Module aims to leverage the divergence capability of large-scale pre-trained model to acquire knowledge facts relevant to the image content. To achieve this, we design a scene-graph-aware prompt that serves as a “trigger” for GPT-3.5, encouraging it to “diverge” and generate more sophisticated, human-like knowledge. (b) Knowledge Focus Module aims to refine acquired knowledge facts and further rearrange them in a coherent manner. We design the interactive refining network to encode knowledge, which is refined with the visual features to remove irrelevant words. Then, to generate fluent image descriptions, we design the large-scale pre-trained model-based rearrangement method to estimate the importance of each knowledge word for an image. Finally, we fuse the refined knowledge and visual features to assist the decoder in generating captions. We demonstrate the superiority of our approach through extensive experiments on the MSCOCO dataset. Our approach surpasses state-of-the-art performance across all metrics in the Karpathy split. For example, our model obtains the best CIDEr-D score of 148.4%. Additional ablation studies and visualization further validate our effectiveness.
Anan Liu, Quanhan Wu, Ning Xu 0003, Hongshuo Tian, Lanjun Wang
IEEE Trans. Circuits Syst. Video Technol.4
2025 Few-Shot In-Context Learning for Implicit Semantic Multimodal Content Detection and Interpretation
abstract
In recent years, the field of explicit semantic multimodal content research makes significant progress. However, research on content with implicit semantics, such as online memes, remains insufficient. Memes often convey implicit semantics through metaphors and may sometimes contain hateful information. To address this issue, researchers propose a task for detecting hateful memes, opening up new avenues for exploring implicit semantics. The hateful meme detection currently faces two main problems: 1) the rapid emergence of meme content makes continuous tracking and detection difficult; 2) current methods often lack interpretability, which limits the understanding and trust in the detection results. To make a better understanding of memes, we analyze the definition of metaphor from social science and identify the three key factors of metaphor: socio-cultural knowledge, metaphorical tenor, and metaphorical representation pattern. According to these key factors, we guide a multimodal large language model (MLLM) to infer the metaphors expressed in memes step by step. Particularly, we propose a hateful meme detection and interpretation framework, which has four modules. We first leverage a multimodal generative search method to obtain socio-cultural knowledge relevant to visual objects of memes. Then, we use socio-cultural knowledge to instruct the MLLM to assess the social-cultural relevance scores between visual objects and textual information, and identify the metaphorical tenor of memes. Meanwhile, we apply a representative interpretation method to provide representative cases of memes and analyze these cases to explore metaphorical representation pattern. Finally, a chain-of-thought prompt is constructed to integrate the output of the above modules, guiding the MLLM to accurately detect and interpret hateful memes. Our method achieves state-of-the-art performance on three hateful meme detection benchmarks and performs better than supervised training models on the hateful meme interpretation benchmark.
Xiuxian Wang, Lanjun Wang, Yuting Su 0001, Hongshuo Tian, Guoqing Jin, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.4
2024 CAT-DM: Controllable Accelerated Virtual Try-On with Diffusion Model
abstract
Generative Adversarial Networks (GANs) dominate the research field in image-based virtual try-on, but have not resolved problems such as unnatural deformation of garments and the blurry generation quality. While the generative quality of diffusion models is impressive, achieving controllability poses a significant challenge when applying it to virtual try-on and multiple denoising iterations limit its potential for real-time applications. In this paper, we propose Controllable Accelerated virtual Try-on with Diffusion Model (CAT-DM). To enhance the controllability, a basic diffusion-based virtual try-on network is designed, which utilizes ControlNet to introduce additional control conditions and improves the feature extraction of garment images. In terms of acceleration, CAT-DM initiates a reverse denoising process with an implicit distribution generated by a pre-trained GAN-based model. Compared with previous try-on methods based on diffusion models, CAT-DM not only retains the pattern and texture details of the in-shop garment but also reduces the sampling steps without compromising generation quality. Extensive experiments demonstrate the superiority of CAT-DM against both GAN-based and diffusion-based methods in producing more real-istic images and accurately reproducing garment patterns.
Jianhao Zeng, Dan Song 0006, Weizhi Nie, Hongshuo Tian, Anan Liu
CVPR4
2024 Cross-Modal Coherence-Enhanced Feedback Prompting for News Captioning
abstract
News Captioning involves generating the descriptions for news images based on the detailed content of related news articles. Given that these articles often contain extensive information not directly related to the image, captions may end up misaligned with the visual content. To mitigate this issue, we propose the novel cross-modal coherence-enhanced feedback prompting method to clarify the crucial elements that align closely with the visual content for news captioning. Specifically, we first adapt CLIP to develop a news-specific image-text matching module, enriched with insights from language model MPNet using a matching-score comparative loss, which facilitates effective cross-modal knowledge distillation. This module enhance the coherence between images and each news sentences via rating confidence. Then, we design confidence-aware prompts to fine-tune LLaVA model with by LoRa strategy, focusing on essential details in extensive articles. Lastly, we evaluate the generated news caption with refined CLIP, constructing confidence-feedback prompts to further enhance LLaVA through feedback learning, which iteratively refine captions to improve its accuracy. Extensive experiments conduct on two public datasets, GoodNews and NYTimes800k, have validated the effectiveness of our method.
Ning Xu 0003, Hongshuo Tian, Anan Liu
ACM Multimedia4
2024 Gaussian Distribution-Aware Commonsense Knowledge Learning for Scene Graph Generation
abstract
Knowledge-based Scene Graph Generation (SGG) requires external commonsense knowledge beyond the visual scene to infer the relation between objects. Such knowledge can be obtained in a variety of forms, such as vision, text, and graph. However, there are two drawbacks as follows: 1) commonsense knowledge essentially has uncertainty, but current works usually represent knowledge in a deterministic manner, which is not well matched to its nature, 2) using commonsense knowledge without denoising will introduce irrelevant information. This can increase the burden on the relation classifier and only obtain marginal gains over a large amount of data. In this paper, we propose a novel Gaussian distribution-aware commonsense knowledge learning method for SGG. First, we associate each object pair with a Gaussian distribution, which parametrizes visual context and commonsense as mean and variance, respectively. We prove that Gaussian modeling can provide a probabilistic soft space to measure the uncertainty of external knowledge, which allows diverse predictions. Second, to reduce semantic noise in commonsense, we sample multiple variables from the Gaussian distribution and train multi-expert classifiers, which can be dynamically examined for the ensemble softmax classification. Extensive comparative experiments on two benchmarks confirm that our method can achieve competitive performance against the state-of-the-art. Ablation studies verify the essential roles of individual components. Moreover, the visualization of multi-expert classifiers confirms our ability to integrate commonsense for relation inference.
Hongshuo Tian, Ning Xu 0003, Mohan Kankanhalli, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.1
2024 Rule-Driven News Captioning
abstract
News captioning task aims to generate sentences by describing named entities or concrete events for an image with its news article. Existing methods have achieved remarkable results by relying on the large-scale pre-trained models, which primarily focus on the correlations between the input news content and the output predictions. However, the news captioning requires adhering to some fundamental rules of news reporting, such as accurately describing the individuals and actions associated with the event. In this paper, we propose the rule-driven news captioning method, which can generate image descriptions following designated rule signal. Specifically, we first design the news-aware semantic rule for the descriptions. This rule incorporates the primary action depicted in the image (e.g., “performing”) and the roles played by named entities involved in the action (e.g., “Agent” and “Place”). Second, we inject this semantic rule into the large-scale pre-trained model, BART, with the prefix-tuning strategy, where multiple encoder layers are embedded with news-aware semantic rule. Finally, we can effectively guide BART to generate news sentences that comply with the designated rule. Extensive experiments on two widely used datasets (i.e., GoodNews and NYTimes800k) demonstrate the effectiveness of our method.
Ning Xu 0003, Hongshuo Tian, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.3
2024 Multi-Modal Validation and Domain Interaction Learning for Knowledge-Based Visual Question Answering
abstract
Knowledge-based Visual Question Answering (KB-VQA) aims to answer the image-aware question via the external knowledge, which requires an agent to not only understand images but also explicitly retrieve and integrate knowledge facts. Intuitively, to accurately answer the question, we humans can validate the retrieved knowledge based on our memory, and then align the knowledge facts with the image regions to infer answers. However, most existing methods ignore the process of knowledge validation and alignment. In this paper, we propose the Multi-Modal Validation and Domain Interaction Learning method, which consists of two components: 1) Multi-modal validation for knowledge retrieval. We propose the multi-modal validation module (MMV) to evaluate the confidence of each retrieved knowledge fact via images and questions, which preserves knowledge candidates effective for inferring answers. 2) Domain interaction for knowledge integration. We propose the Domain Interaction TRansformer module (DI-TR) to align visual regions with knowledge facts by the interaction learning in the improved transformer. Specifically, the inter-domain and intra-domain masks are injected into each self-attention layer to control the integration scope. The proposed method outperforms several strong baselines on three widely-used knowledge-based datasets: KRVQA, OK-VQA and VQA2.0. Extensive experiments and ablation studies demonstrate the effectiveness of multi-modal knowledge validation and domain interaction learning.
Ning Xu 0003, Anan Liu, Hongshuo Tian, Yongdong Zhang 0001
IEEE Trans. Knowl. Data Eng.4
2024 Counterfactual Visual Dialog: Robust Commonsense Knowledge Learning From Unbiased Training
abstract
Visual Dialog (VD) requires an agent to answer the current question by engaging in a conversation with humans referring to an image. Despite the recent progress, it is beneficial to introduce external commonsense knowledge to fully understand the given image and dialog history. However, the existing knowledge-based VD models are inclined to rely on severe learning bias brought by commonsense, e.g., the retrieved$< {\mathtt{{bus}}}, {\mathtt{capable\;of}}, {\mathtt{transport\;people}}>$,$< {\mathtt{{bus}}}, {\mathtt{is\;a}}, {\mathtt{public\;transport}}>$, and$< {\mathtt{{bus}}}, {\mathtt{is\;a}}, {\mathtt{car}}>$can induce a spurious correlation between the question “What is the bus used for?” and the false answer “City bus”. There are two challenges to make commonsense learning more robust against spurious correlations: 1) how to disentangle the true effect of “good” commonsense knowledge from the whole, and 2) how to estimate and remove the effect of “bad” commonsense bias on answers. In this article, we propose a novel CounterFactual Commonsense learning scheme for the Visual Dialog task (CFC-VD). First, comparing with the causal graph of existing VD models, we add one new commonsense node and one new link to multi-modal information from history, question, and image. Since the retrieved knowledge prior is subtle and uncontrollable, we consider it as an unobserved confounder in the commonsense node, which leads to spurious correlations for the answer inference. Then, to remove the effect of the confounder, we formulate it as the direct causal effect of commonsense on answers and remove the direct language effect by subtracting it from the total causal effect via counterfactual reasoning. Experimental results certify the effectiveness of our method on the prevailing Visdial v0.9 and Visdial v1.0 datasets.
Anan Liu, Ning Xu 0003, Hongshuo Tian, Jing Liu 0002, Yongdong Zhang 0001
IEEE Trans. Multim.4
2024 Event-Aware Retrospective Learning for Knowledge-Based Image Captioning
abstract
External knowledge has been widely applied in image captioning tasks to enrich the generated sentences. However, existing methods retrieve knowledge by considering only semantic relevance while ignoring whether they are useful for captioning. For example, when querying “person” in external knowledge, the most relevant concepts may be “wearing shirt” or “riding horse” statistically, which are not consistent with image contents and introduce noise to generated sentences. Intuitively, we humans can iteratively correlate visual clues with corresponding knowledge to distinguish useful clues from noise. Therefore, we propose an event-aware retrospective learning network for knowledge-based image captioning, which employs a retrospective validation mechanism on captioning models to align the retrieved knowledge with visual contents. This approach is an event-aware perspective and helps select useful knowledge that corresponds to visual facts. To better align images and knowledge, 1) we design an event-aware retrieval algorithm that clusters word-centered knowledge into triplet-centered knowledge (i.e., from “” to “- edge -”, which provides an event context to facilitate knowledge retrieval and validation. 2) We revisit image contents to retrospectively validate retrieved knowledge by aligning the visual representation between knowledge and image. We summarize the visual characteristics of each knowledge event from the visual genome dataset to help learn which knowledge does not exist in the visual scene and should be discarded. 3) We adopt a dynamic knowledge fusion module that calibrates image and knowledge representations for sentence generation, which includes a knowledge-controlled gate unit that jointly calculates visual and semantic features in event-aware patterns. Compared to current knowledge-based captioning methods, the proposed network retrospectively learns the visual facts by event-aware retrieval and knowledge-image visual alignment, which regularizes the knowledge-incorporated captioning with visual evidence. Extensive experiments on the MS-COCO dataset demonstrate the effectiveness of our method. Ablation studies and visualization demonstrate the advantages of each component of the proposed model.
Anan Liu, Yingchen Zhai, Ning Xu 0003, Hongshuo Tian, Weizhi Nie, Yongdong Zhang 0001
IEEE Trans. Multim.4
2024 Learning to Supervise Knowledge Retrieval Over a Tree Structure for Visual Question Answering
abstract
Knowledge-based visual question answering (KBVQA) aims to retrieve the external knowledge out of images to answer questions. However, current methods always introduce various irrelevant knowledge due to two drawbacks: (1) Synonymy issue. Existing methods heavily rely on words from questions or object labels in images to match knowledge from databases, which disregards the same word may hold multiple meanings within different contexts. (2) Knowledge uncertainty issue. Due to the absence of supervisory signals, recent methods can not determine which knowledge is applicable for answer inference, which can mislead to admit useless knowledge. To address these two problems, we propose to supervise the process of knowledge retrieval over a tree structure for KB-VQA task. For the synonymy issue, we construct a hierarchical knowledge tree to capture the subordination information between knowledge facts, mitigating the impact of synonyms on knowledge retrieval. For the knowledge uncertainty issue, we use the retrieval history as the ground truth to supervise the knowledge retrieval, which facilitates the QA model to form an explicit path of knowledge facts for answer understanding. Finally, we integrate the image, question, and retrieved knowledge into a variant of transformer to predict answers. Experimental results validate the effectiveness of the proposed method on KR-VQA, OK-VQA and VQA v2 datasets.
Ning Xu 0003, Zimu Lu, Hongshuo Tian, Rongbao Kang, Jinbo Cao, Yongdong Zhang 0001, Anan Liu
IEEE Trans. Multim.3
2023 Towards Confidence-Aware Commonsense Knowledge Integration for Scene Graph Generation
abstract
Commonsense knowledge has been widely explored to improve Scene Graph Generation (SGG). Existing methods simply incorporate the described relations of knowledge bases into each part of the scene for a concrete understanding. However, they ignore the discussion about whether a visual scene needs to associate commonsense knowledge for making inferences. Specifically, the difficulty of relation recognition varies from its type. Some frequent spatial relations (e.g. on) usually produce less perception error even without any prior information, while others involved many rules and patterns (e.g. throwing) possess few samples and require to combine with some commonsense knowledge as supplementary. In this paper, we propose a novel confidence-aware commonsense knowledge integration for SGG. Firstly, we depend on mutual information maximization to design a hybrid-attention module, which decreases the uncertainty in representation learning given external knowledge. Second, we introduce an extra branch for SGG network to perform confidence estimation independent of any ground truth labels, in which the output scalar explicitly reflects the difficulty of visual recognition. This value is equipped with the ability to balance the demand for commonsense knowledge in a given scene. Experiments are conducted with the backbone of MOTIFS on Visual Genome (VG) and our method effectively promotes the metric of mRecall with little performance hit for metric Recall, especially for predicting unseen relations.
Hongshuo Tian, Ning Xu 0003, Yanhui Wang 0001, Chenggang Yan 0001, Bolun Zheng, Xuanya Li, Anan Liu
ICME1
2023 Knowledge Prompt Makes Composed Pre-Trained Models Zero-Shot News Captioner
abstract
News image captioning aims to generate descriptions containing concrete named entities for news images by leveraging relevant news articles. However, existing approaches suffer from two shortcomings: 1) lack of commonsense knowledge required to understand named entities, and 2) limited multimodal context modeling capabilities. In this paper, we propose to migrate the ability of large-scale pre-trained models for news image captioning. To acquire factual knowledge for describing named entities, we induce a pre-trained language model for commonsense knowledge reasoning using context-aware knowledge prompts. To compose a new multimodal context modeling capability, we coordinate pre-trained models by a unified language representation and constrain joint multimodal context reasoning with cross-modal consistency objective. Experimental results on GoodNews and NYTimes datasets show that our proposed method exhibits considerable captioning capabilities even without training on news data.
Yanhui Wang 0001, Ning Xu 0003, Hongshuo Tian, Yulong Duan, Xuanya Li, Anan Liu
ICME3
2022 Toward Region-Aware Attention Learning for Scene Graph Generation
abstract
Scene graph generation (SGGen) is a challenging task due to a complex visual context of an image. Intuitively, the human visual system can volitionally focus on attended regions by salient stimuli associated with visual cues. For example, to infer the relationship between man and horse, the interaction between human leg and horseback can provide strong visual evidence to predict the predicate ride. Besides, the attended region face can also help to determine the object man. Till now, most of the existing works studied the SGGen by extracting coarse-grained bounding box features while understanding fine-grained visual regions received limited attention. To mitigate the drawback, this article proposes a region-aware attention learning method. The key idea is to explicitly construct the attention space to explore salient regions with the object and predicate inferences. First, we extract a set of regions in an image with the standard detection pipeline. Each region regresses to an object. Second, we propose the object-wise attention graph neural network (GNN), which incorporates attention modules into the graph structure to discover attended regions for object inference. Third, we build the predicate-wise co-attention GNN to jointly highlight subject's and object's attended regions for predicate inference. Particularly, each subject-object pair is connected with one of the latent predicates to construct one triplet. The proposed intra-triplet and inter-triplet learning mechanism can help discover the pair-wise attended regions to infer predicates. Extensive experiments on two popular benchmarks demonstrate the superiority of the proposed method. Additional ablation studies and visualization further validate its effectiveness.
Anan Liu, Hongshuo Tian, Ning Xu 0003, Weizhi Nie, Yongdong Zhang 0001, Mohan Kankanhalli
IEEE Trans. Neural Networks Learn. Syst.2
2021 Mask and Predict: Multi-step Reasoning for Scene Graph Generation
abstract
Scene Graph Generation (SGG) aims to parse the image as a set of semantics, containing objects and their relations. Currently, the SGG methods only stay at presenting the intuitive detection in the image, such as the triplet "logo on board". Intuitively, we humans can further refine these intuitive detections as rational descriptions like "flower painted on surfboard". However, most of existing methods always formulate SGG as a straightforward task, only limited by the manner of one-time prediction, which focuses on a single-pass pipeline and predicts all the semantic. Therefore, to handle this problem, we propose a novel multi-step reasoning manner for SGG. Concretely, we break SGG into two explicit learning stages, including intuitive training stage (ITS) and rational training stage (RTS). In the first stage, we follow the traditional SGG processing to detect objects and relationships, yielding an intuitive scene graph. In the second stage, we perform multi-step reasoning to refine the intuitive scene graph. For each step of reasoning, it consists of two kinds of operations: mask and predict. According to primary predictions and their confidences, we constantly select and mask the low-confidence predictions, which features are optimized and predicted again. After several iterations, all of intuitive semantics will gradually tend to be revised with high confidences, yielding a rational scene graph. Extensive experiments on Visual Genome prove the superiority of the proposed method. Additional ablation studies and visualization cases further validate its effectiveness.
Hongshuo Tian, Ning Xu 0003, Anan Liu, Chenggang Yan 0001, Zhendong Mao 0001, Yongdong Zhang 0001
ACM Multimedia1
2021 Coupled-dynamic learning for vision and language: Exploring Interaction between different tasks
Ning Xu 0003, Hongshuo Tian, Yanhui Wang 0001, Weizhi Nie, Dan Song 0006, Anan Liu, Wu Liu 0005
Pattern Recognit.2
2020 Part-Aware Interactive Learning for Scene Graph Generation
abstract
Generating scene graph to describe the whereabouts and interactions of objects in an image has attracted increasing attention of researchers. Most existing methods explore object-level visual context or bodypart-object cooperation with the message passing structure, which can not meet the part-aware interaction nature of scene graph. Normally, a subject interacts with an object through crucial parts in each other. Besides, the correlation among parts within an identical object can also help predicting objects and their relationships. Hence, both of subject and object parts and their intra- and inter-object correlations should be fully considered for scene graph generation. In this paper, we propose a part-aware interactive learning method, which are divided into the intra-object and inter-object scenarios. First, we detect objects from an image and further decompose each one into a set of parts. Second, the part-aware graph attention module is proposed to refine part features via the intra-object message passing, and the refined features are incorporated for object inference. Third, the visual mutual attention module is designed to discover part-aware correlated visual cues precisely for predicate inference. It can highlight the subject-related object parts and the object-related subject parts during inter-object interactive learning. We demonstrate the superiority of our method against the state of the arts on Visual Genome. Ablation studies and visualization further validate its effectiveness.
Hongshuo Tian, Ning Xu 0003, Anan Liu, Yongdong Zhang 0001
ACM Multimedia1