Qingbao Huang

dblp:62/3489 · DBLP profile ↗
← Back
50ranked-venue papers
9as first author
43since 2021 · last 2026
0000-0001-7691-347XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 6 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 5 first-author · 23 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Hybrid-DMKG: A Hybrid Reasoning Framework over Dynamic Multimodal Knowledge Graphs for Multimodal Multihop QA with Knowledge Editing
abstract
Multimodal Knowledge Editing (MKE) extends traditional knowledge editing to settings involving both textual and visual modalities. However, existing MKE benchmarks primarily assess final answer correctness, neglecting the quality of intermediate reasoning and robustness to visually rephrased inputs. To address this limitation, we introduce MMQAKE, the first benchmark for multimodal multihop question answering with knowledge editing. MMQAKE evaluates: (1) a model’s ability to reason over 2–5-hop factual chains that span both text and images, including performance at each intermediate step; (2) robustness to visually rephrased inputs in multihop questions. Our evaluation shows that current MKE methods often struggle to consistently update and reason over multimodal reasoning chains following knowledge edits. To overcome these challenges, we propose Hybrid-DMKG, a hybrid reasoning framework built on a dynamic multimodal knowledge graph (DMKG) to enable accurate multihop reasoning over updated multimodal knowledge. Hybrid-DMKG first uses a large language model to decompose multimodal multihop questions into sequential sub-questions, then applies a multimodal retrieval model to locate updated facts by jointly encoding each sub-question with candidate entities and their associated images. For answer inference, a hybrid reasoning module operates over the DMKG via two parallel paths: (1) relation-linking prediction; (2) RAG Reasoning with large vision-language models. A background-reflective decision module then aggregates evidence from both paths to select the most credible answer. Experimental results on MMQAKE show that Hybrid-DMKG significantly outperforms existing MKE approaches, achieving higher accuracy and improved robustness to knowledge updates.
Qingfei Huang, Bingshan Zhu, Yi Cai 0001, Qingbao Huang, Changmeng Zheng, Zikun Deng, Tao Wang 0036
AAAI5
2026 Knowledge-enhanced Chinese multimodal hate speech detection
Qingbao Huang, Pijian Li, Xingmao Zhang, Shizhen Chen, Haonan Cheng, Zhiyue Liu
Expert Syst. Appl.1
2026 SAM-MPA: A SAM-Based Motion Perception and Aggregation Framework for Referring Video Segmentation
abstract
Referring video object segmentation relies on natural language descriptions to identify and segment target objects in videos, and has achieved substantial progress in recent years. However, most prior studies process videos in a frame-by-frame manner, failing to fully exploit temporal information. Recently, the large-scale segmentation model Segment Anything Model (SAM) has attracted considerable attention due to its strong segmentation capability and impressive zero-shot generalization. Nevertheless, SAM still exhibits limitations when handling complex action-oriented descriptions. Motivated by these observations, we propose a novel SAM-based Motion-Perception Aggregation framework for referring video object segmentation, termed SAM-MPA, which consists of four modules. DINO-SAM leverages the powerful segmentation ability of SAM to perform initial video segmentation guided by textual prompts, generating object-level masks. The Kalman Filtering Motion Modeling module injects explicit object motion modeling into DINO-SAM, improving segmentation robustness under occlusion and fast-motion scenarios. The motion-aware aggregation module effectively captures object action cues at multiple temporal scales, thereby enhancing global video understanding. The text-token matching module further enforces semantic consistency between the segmentation results and the referring expressions. Extensive experiments on challenging RVOS benchmarks demonstrate that SAM-MPA provides a competitive and efficient SAM-based solution for motion-centric referring video object segmentation, while offering a favorable trade-off between performance and computational cost compared with conventional non-MLLM baselines. The code is available at https://github.com/GXU-LIPE/SAM-MPA.
Fang Gao 0001, Ao Lu, Qingbao Huang, Jun Yu 0001
IEEE Internet Things J.4
2026 ExCap: Entity-aware zero-shot image captioning via faithful synthetic image-text alignment
Qipeng Jiang, Zhiyue Liu, Wenkai Zhou, Qingbao Huang, Jiahai Wang
Knowl. Based Syst.5
2026 Anchor-Based Multimodal Verification: A Dynamic Query Framework for Fake News Forensics in Short Videos
abstract
The proliferation of maliciously altered short videos on social media platforms poses a significant threat to information security ecosystems, eroding public trust in digital media. Despite recent advancements in detecting fake video news, significant challenges remain in the forensic analysis of short videos, leading to issues of bias. First, as technology rapidly advances, fake videos are becoming increasingly semantically convincing, undermining the effectiveness of current classification methods. Second, the heterogeneous nature of video modalities (visual, textual, audio) creates critical challenges for models to learn discriminative feature representations. To address these challenges, we propose a dynamic query framework for fake news forensics in short videos, termed the Semantic Guided Adaptive Network (SGAN). Our approach is motivated by the need to utilize superficial alignment to identify suspicious manipulations through anchor-based verification and to leverage the adaptive capability of learnable queries to learn the heterogeneous boundary in each modality. Specifically, SGAN comprises a verification module and a flexible query learning module. The verification module employs text as the anchor to verify detailed context, mining fine-grained information while emphasizing key features, thereby providing candidate manipulations for downstream modules. The query learning module leverages learnable queries to map heterogeneous forensic features and integrates them through multi-level fusion for decision-making. Extensive experiments conducted on two widely used datasets demonstrate the effectiveness and generalization of the proposed method.
Pijian Li, Qingbao Huang, Feng Shuang 0002, Yi Cai 0001, Haonan Cheng, Qing Li 0001
IEEE Trans. Inf. Forensics Secur.2
2026 DESSM: Dual Encoder-Based State Space Model for Image Inpainting
abstract
Image inpainting represents a fundamental and challenging problem in computer vision, requiring the synthesis of visually plausible content for missing regions while preserving both textural details and structural coherence. While current approaches employ auxiliary networks and attention mechanisms to capture structural priors and expand receptive fields, they remain constrained by two fundamental limitations: (1) insufficient interaction between textural and structural priors, and (2) the quadratic computational complexity inherent in attention operations. To overcome these challenges, we present DESSM, an innovative Dual Encoder-based State space Model for image inpainting that achieves efficient global context modeling with linear computational complexity. Our DESSM integrates three synergistically designed modules: (1) a dual-branch encoder for complementary learning of textural patterns and structural priors, (2) a Feature Cross Fusion Block (FCFB) enabling dynamic feature interaction while adaptively suppressing redundant information, and (3) a Spatial-Channel joint Selective scan Block (SCSB) for efficient long-range dependency modeling. Comprehensive evaluations across four standard benchmarks (i.e., CelebA, CelebA-HQ, Places2, and Paris StreetView) demonstrate that our DESSM achieves state-of-the-art performance in both visual fidelity and computational efficiency.
Rongrong Zhou, Peizhou Cai, Pijian Li, Qingbao Huang
IEEE Trans. Multim.5
2026 Metaphorical Visual Question Answering: Benchmark and Knowledge-Enhanced Metaphor Understanding Method
abstract
Fact and common-sense reasoning grounded in metaphorical imagery constitute a more challenging form of visual question answering (VQA). Under this form, models typically cannot obtain answers directly from images. Models first need to identify and comprehend the metaphorical components within the image, subsequently integrating prior knowledge to establish the mappings between the target and source domains depicted in the image. To evaluate this capability, we propose a VQA benchmark based on metaphorical images (METAVQA), which measures the understanding of the model of metaphorical images in the form of VQA. Experimental results indicate that interpreting metaphorical images requires robust prior knowledge and a strong ability to understand abstract components, which remains a challenge for LLMs. To improve the metaphorical understanding capability of LLMs, we propose a knowledge-enhanced method (KEMU). This method utilizes LLMs to extract image triplets, retrieve implicit metaphorical knowledge, and reason through questions step-by-step using chain of thought prompting, and pretrained classifier to match the refined reasoning output to the provided answer choices. Validated on the METAVQA dataset, KEMU outperforms LLMs in both one-hop and multi-hop questions. Our code and benchmark can be seen inhttps://github.com/VILAN-Lab/METAVQA.
Qingbao Huang, Peihang He, Pijian Li, Yang Tian 0008, Yi Cai 0001, Qing Li 0001
IEEE Trans. Multim.1
2026 Implement Referring Expression Comprehension by Extending Auto-focus Lens to Locked Vision Model
abstract
Referring Expression Comprehension (REC) aims to achieve fine-grained cross-modal content alignment. The traditional two-stage approaches, by decomposing REC into localization (region proposal) and comprehension (expression-based ranking), lead to the isolation of continuous image information and heavily rely on the quality of the proposals. In this article, we propose a point-based two-stage framework for REC to quickly achieve localization by inserting a language-modulated auto-focus module into the locked vision model. Specifically, we redefine REC as two processes: point-based cross-modal comprehension and point-based instance localization. For the comprehension stage, we reconstruct the raw annotations into soft masks at the feature point level as a metric of cross-modal correlation. With this indirect metric, REC can be approximated as a binary classification problem, which fundamentally avoids the impact of isolated regions. Remarkably, soft masks are shape-independent, which means our method is extremely general. By switching different vision models, different types of predictions (e.g., localization and segmentation) can be obtained. Experiments on multiple benchmarks demonstrate the feasibility and potential of our point-based paradigm. Our code will be public at https://github.com/VILAN-Lab/PBREC-AF .
Shiyi Zheng, Peizhi Zhao, Qingbao Huang, Yi Cai 0001, Haonan Cheng, Qi Wu 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2025 Look Around Before Locating: Considering Content and Structure Information for Visual Grounding
abstract
As a long-term challenge and fundamental requirement in vision and language tasks, visual grounding aims to localize a target referred by a natural language query. The regional annotations form a superficial correlation between the subject of expression and some common visual entities, which hinder models from comprehending the linguistic content and structure. However, current one-stage methods struggle to uniformly model the visual and linguistic structure due to the structural gap between continuous image patches and discrete text tokens. In this paper, we propose a semi-structured reasoning framework for visual grounding to gradually comprehend the linguistic content and structure. Specifically, we devise a cross-modal content alignment module to effectively align unlabeled contextual information into a stable semantic space corrected by token-level prior knowledge obtained with CLIP. A multi-branch modulated localization module is also established to obtain modulation grounding by linguistic structure. Through a soft split mechanism, our method can destructure the expression into a fixed semi-structure (i.e., subject and context) while ensuring the completeness of linguistic content. Our method is thus capable of building a semi-structured reasoning system to effectively comprehend the linguistic content and structure by content alignment and structure modulated grounding. Experimental results on five widely-used datasets validate the performance improvements of our proposed method.
Shiyi Zheng, Peizhi Zhao, Zhilong Zheng, Peihang He, Haonan Cheng, Yi Cai 0001, Qingbao Huang
AAAI7
2025 Collaborative Multi-LoRA Experts with Achievement-based Multi-Tasks Loss for Unified Multimodal Information Extraction
abstract
Multimodal Information Extraction (MIE) has gained attention for extracting structured information from multimedia sources. Traditional methods tackle MIE tasks separately, missing opportunities to share knowledge across tasks. Recent approaches unify these tasks into a generation problem using instruction-based T5 models with visual adaptors, optimized through full-parameter fine-tuning. However, this method is computationally intensive, and multi-task fine-tuning often faces gradient conflicts, limiting performance. To address these challenges, we propose collaborative multi-LoRA experts with achievement-based multi-task loss (C-LoRAE) for MIE tasks. C-LoRAE extends the low-rank adaptation (LoRA) method by incorporating a universal expert to learn shared multimodal knowledge from cross-MIE tasks and task-specific experts to learn specialized instructional task features. This configuration enhances the model’s generalization ability across multiple tasks while maintaining the independence of various instruction tasks and mitigating gradient conflicts. Additionally, we propose an achievement-based multi-task loss to balance training progress across tasks, addressing the imbalance caused by varying numbers of training samples in MIE tasks. Experimental results on seven benchmark datasets across three key MIE tasks demonstrate that C-LoRAE achieves superior overall performance compared to traditional fine-tuning methods and LoRA methods while utilizing a comparable number of training parameters to LoRA.
Yi Cai 0001, Qing Li 0001, Qingbao Huang, Zikun Deng, Tao Wang 0036
IJCAI5
2025 Deep Learning-Based Knowledge Injection for Metaphor Detection: A Comprehensive Review
Zhiyue Liu, Xingmao Zhang, Qingbao Huang
NLPCC (3)4
2025 Multi-modal Metaphor Explanation: A Dataset and Benchmark
Wenye Zhao, Pijian Li, Qingbao Huang
PRCV (12)5
2025 Decoding before aligning: Scale-Adaptive Early-Decoding Transformer for visual grounding
Liuwu Li, Yi Cai 0001, Jiexin Wang 0002, Cantao Wu, Qingbao Huang, Qing Li 0001
Neurocomputing5
2025 Grouped top-down reasoning with hierarchical window transformer for visual grounding
Liuwu Li, Zhuoming Zheng, Yuqi Bu, Cantao Wu, Shubin Huang, Qingbao Huang, Yi Cai 0001
Inf. Process. Manag.6
2025 Synthesize then align: Modality alignment augmentation for zero-shot image captioning with synthetic data
Zhiyue Liu, Xin Ling, Qingbao Huang, Jiahai Wang
Knowl. Based Syst.4
2025 Visual primitives as words: Alignment and interaction for compositional zero-shot learning
Feng Shuang 0002, Jiahuan Li, Qingbao Huang, Wenye Zhao, Dongsheng Xu 0001, Haonan Cheng
Pattern Recognit.3
2025 DEVICE: Depth and Visual Concepts Aware Transformer for OCR-based image captioning
Dongsheng Xu 0001, Qingbao Huang, Xingmao Zhang, Haonan Cheng, Yi Cai 0001
Pattern Recognit.2
2025 Error-Aware Generative Reasoning for Zero-Shot Visual Grounding
abstract
Zero-shot visual grounding is the task of identifying and localizing an object in an image based on a referring expression without task-specific training. Existing methods employ heuristic rules to step-by-step perform visual perception for visual grounding. Despite their remarkable performance, there are still two limitations. First, such a rule-based manner struggles with expressions that are not covered by predefined rules. Second, existing methods lack a mechanism for identifying and correcting visual perceptual errors of incomplete information, resulting in cascading errors caused by reasoning based on incomplete visual perception results. In this article, we propose an Error-Aware Generative Reasoning (EAGR) method for zero-shot visual grounding. To address the limited adaptability of existing methods, a reasoning chain generator is presented, which prompts LLMs to dynamically generate reasoning chains for specific referring expressions. This generative manner eliminates the reliance on human-written heuristic rules. To mitigate visual perceptual errors of incomplete information, an error-aware mechanism is presented to elicit LLMs to identify these errors and explore correction strategies. Experimental results on four benchmarks show that EAGR outperforms state-of-the-art zero-shot methods by up to 10% and an average of 7%.
Yuqi Bu, Xin Wu 0003, Yi Cai 0001, Qiong Liu 0006, Tao Wang 0036, Qingbao Huang
IEEE Trans. Multim.6
2024 Rethinking Two-Stage Referring Expression Comprehension: A Novel Grounding and Segmentation Method Modulated by Point
abstract
As a fundamental and challenging task in the vision and language domain, Referring Expression Comprehension (REC) has shown impressive improvements recently. However, for a complex task that couples the comprehension of abstract concepts and the localization of concrete instances, one-stage approaches are bottlenecked by computing and data resources. To obtain a low-cost solution, the prevailing two-stage approaches decouple REC into localization (region proposal) and comprehension (region-expression matching) at region-level, but the solution based on isolated regions cannot sufficiently utilize the context and is usually limited by the quality of proposals. Therefore, it is necessary to rebuild an efficient two-stage solution system. In this paper, we propose a point-based two-stage framework for REC, in which the two stages are redefined as point-based cross-modal comprehension and point-based instance localization. Specifically, we reconstruct the raw bounding box and segmentation mask into center and mass scores as soft ground-truth for measuring point-level cross-modal correlations. With the soft ground-truth, REC can be approximated as a binary classification problem, which fundamentally avoids the impact of isolated regions on the optimization process. Remarkably, the consistent metrics between center and mass scores allow our system to directly optimize grounding and segmentation by utilizing the same architecture. Experiments on multiple benchmarks show the feasibility and potential of our point-based paradigm. Our code available at https://github.com/VILAN-Lab/PBREC-MT.
Peizhi Zhao, Shiyi Zheng, Wenye Zhao, Dongsheng Xu 0001, Pijian Li, Yi Cai 0001, Qingbao Huang
AAAI7
2024 Can ChatGPT's Performance be Improved on Verb Metaphor Detection Tasks? Bootstrapping and Combining Tacit Knowledge
abstract
Metaphors detection, as an important task in the field of NLP, has been receiving sustained academic attention in recent years.Current researches focus supervised metaphors detection systems, which usually require large-scale, high-quality labeled data support.The emerge of large language models (e.g., ChatGPT) has made many NLP tasks (e.g., automatic summarization and dialogue systems) a qualitative leap.However, it is worth noting that the use of ChatGPT for unsupervised metaphors detection is often challenged with less-thanexpected performance.Therefore, the aim of our work is to explore how to bootstrap and combine ChatGPT by detecting the most prevalent verb metaphors among metaphors.Our approach first utilizes ChatGPT to obtain literal collocations of target verbs and subjectobject pairs of verbs in the text to be detected.Subsequently, these literal collocations and subject-object pairs are mapped to the same set of topics, and finally the verb metaphors are detected through the analysis of entailment relations.The experimental results show that our method achieves the best performance on the unsupervised verb metaphors detection task compared to existing unsupervised methods or direct prediction using ChatGPT.Our code is available at https://github.com/VILAN- Lab/
Puli Chen, Qingbao Huang
ACL (1)3
2024 Hate Speech Detection for the Power Domain
Qingbao Huang, Zehua Deng, Shizhen Chen, Feng Shuang 0002
NLPCC (4)1
2024 A Knowledge-Enhanced and Topic-Guided Domain Adaptation Model for Aspect-Based Sentiment Analysis
abstract
Cross-domain aspect-based sentiment analysis has recently attracted significant attention, which can effectively alleviate the problem of lacking large-scale labeled data for supervised learning based methods. Most of current methods mainly focus on extracting domain-shared syntactic features to conduct the domain adaptation. Due to the language and syntax are diverse between domains, these methods lack generalization and even lead to syntactic transfer errors. External knowledge graphs have rich domain commonsense and share the relational structures between source and target domains. The domain-shared relational structure can effectively bridge the gap across domains and solve the problem of syntactic transfer errors. Moreover, not all the introduced external knowledge is equally important for the cross-domain aspect-based sentiment analysis. Motivated by these, we propose a knowledge-enhanced and topic-guided cross domain aspect-based sentiment analysis model with the domain-shared commonsense relational structure learning module and the topic-guided knowledge attention module. Extensive experiments are conducted and the experimental results evaluate the effectiveness of our proposed model.
Yushi Zeng, Guohua Wang 0003, Haopeng Ren, Yi Cai 0001, Ho-fung Leung, Qing Li 0001, Qingbao Huang
IEEE Trans. Affect. Comput.7
2024 Multi-Granularity Feature Fusion for Image-Guided Story Ending Generation
abstract
Image-guided Story Ending Generation aims at generating a reasonable and logical ending given a story context and an ending-related image. The existing models have achieved some success by fusing global image features with story context through an attention mechanism. However, they ignore the logical relationship between the story context and the image regions, and have not considered the high-level semantic features of the image such as visual sentiment. This may cause the generated ending inconsistent with the logic or sentiment of the given information. In this paper, we propose aMulti-Granularity featureFusion (MGF) model to solve this problem. Concretely, we first employ an image sentiment extractor to grasp the sentiment features of the image as part of the global image features. We then design a scene subgraph selector to capture the image features of the key region by picking the scene subgraph most relevant to the context. Finally, we fuse the textual and visual features from object level, region level, and global level, respectively. Our model is thereby capable of effectively capturing the key region features and visual sentiment of the image, so as to generate a more logical and sentimental ending. Experimental results show that our MGF model outperforms the state-of-the-art models on most metrics.
Pijian Li, Qingbao Huang, Yi Cai 0001, Feng Shuang 0002, Qing Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 Video Question Generation for Dynamic Changes
abstract
Video question generation task aims to generate meaningful questions about a video targeting an answer. Existing methods merely focus on the static appearance features in the image frames or simply identify a motion in the video to ask general questions. However, a video contains dynamically changing visual content that deserves to be questioned, e.g., changes in object motions, object states and relationships among objects, which is more practical and closer to the dynamic world we live in. In this paper, we propose a difference-aware video question generation model that aims to generate questions about temporal differences in the video, i.e., capturing the dynamic changes between image frames of a video to ask questions. To capture the dynamic changes between image frames, we utilize a temporal difference extractor to localize the differences for each frame pair of a video through an attention mechanism. Then, we introduce an answer-aware module to capture the answer-related image frame pair containing their differences for question generation, which aims to guide our model to focus on answer-related content for questioning. Finally, the output of the answer-aware module is sent to a decoder module to generate questions. Extensive experiments on SVQA and MSVD-QA datasets show that the proposed model outperforms state-of-the-art models, e.g., our model achieves at least 17.1% improvement over existing models in the SVQA dataset. This is because our model can generate questions similar to ground truths that involve changes between image frames in videos. Our code is available at https://github.com/Gary-code/D-VQG.
Jiayuan Xie, Yi Cai 0001, Qingbao Huang, Qing Li 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 Region-Focused Network for Dense Captioning
abstract
Dense captioning is a very critical but under-explored task, which aims to densely detect localized regions-of-interest (RoIs) and describe them with natural language in a given image. Although recent studies tried to fuse multi-scale features from different visual instances to generate more accurate descriptions, their methods still suffer from the lack of exploration of relation semantic information in images, leading to less informative descriptions. Furthermore, indiscriminately fusing all visual instance features will introduce redundant information, resulting in poor matching between descriptions and corresponding regions. In this work, we propose a Region-Focused Network (RFN) to address these issues. Specifically, to fully comprehend the images, we first extract the object-level features, and encode the interaction and position relations between objects to enhance the object representations. Then, to decrease the interference from redundant information about the target region, we extract the most relevant information to the region. Finally, a region-based Transformer is employed to compose and align the previous mined information and generate the corresponding descriptions. Extensive experiments on Visual Genome V1.0 and V1.2 datasets show that our RFN model outperforms the state-of-the-art methods, thus verifying its effectiveness. Our code is available at https://github.com/VILAN-Lab/DesCap .
Qingbao Huang, Pijian Li, Youji Huang, Feng Shuang 0002, Yi Cai 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2023 Linking People across Text and Images Based on Social Relation Reasoning
Peizhi Zhao, Pijian Li, Yi Cai 0001, Qingbao Huang
AAAI5
2023 Scene-text Oriented Visual Entailment: Task, Dataset and Solution
abstract
Visual Entailment (VE) is a fine-grained reasoning task aiming to predict whether the image semantically entails a hypothesis in textual form.Existing studies of VE only focus on basic visual attributes but largely overlook the importance of scene text, which usually entails rich semantic information and crucial clues (e.g., time, place, affiliation, and topic), leading to superficial design of hypothesis or incorrect entailment prediction. To fill this gap, we propose a new task called scene-text oriented Visual Entailment (STOVE), which requires models to predict whether an image semantically entails the corresponding hypothesis designed based on the scene text-centered visual information.STOVE task challenges a model to deeply understand the interplay between language and images containing scene text, requiring aligning hypotheses tokens, scene text, and visual contents.To support the researches on STOVE, we further collect a dataset termed TextVE, consisting of 23,864 images and 47,728 hypotheses related to scene text, which is constructed with the strategy of minimizing biases.Additionally, we present a baseline named MMTVE applying a multimodal transformer to model the spatial, semantic, and visual reasoning relations between multiple scene text tokens, hypotheses, and visual features.Experimental results illustrate that our model is effective in comprehending STOVE and achieves outstanding performance.Our codes are available at https://github.com/VISLANG-Lab/TextVE.
Nan Li 0055, Pijian Li, Dongsheng Xu 0001, Wenye Zhao, Yi Cai 0001, Qingbao Huang
ACM Multimedia6
2023 Zero-TextCap: Zero-shot Framework for Text-based Image Captioning
abstract
Text-based image captioning is a vital but under-explored task, which aims to describe images by captions containing scene text automatically. Recent studies have made encouraging progress, but they are still suffering from two issues. Firstly, current models cannot capture and generate scene text in non-Latin script languages, which severely limits the objectivity and the information completeness of generated captions. Secondly, current models tend to describe images with monotonous and templated style, which greatly limits the diversity of the generated captions. Although the above-mentioned issues can be alleviated through carefully designed annotations, this process is undoubtedly laborious and time-consuming. To address the above issues, we propose a Zero-shot Framework for Text-based Image Captioning (Zero-TextCap). Concretely, to generate candidate sentences starting from the prompt 'Image of' and iteratively refine them to improve the quality and diversity of captions, we introduce a Hybrid-sampling masked language model (H-MLM). To read multi-lingual scene text and model the relationships between them, we introduce a robust OCR system. To ensure that the captions generated by H-MLM contain scene text and are highly relevant to the image, we propose a CLIP-based generation guidance module to insert OCR tokens and filter candidate sentences. Our Zero-TextCap is capable of generalizing captions containing multi-lingual scene text and boosting the diversity of captions. Sufficient experiments demonstrate the effectiveness of our proposed Zero-TextCap. Our codes are available at https://github.com/Gemhuang79/Zero_TextCap.
Dongsheng Xu 0001, Wenye Zhao, Yi Cai 0001, Qingbao Huang
ACM Multimedia4
2023 A Two-Stage Chinese Medical Video Retrieval Framework with LLM
Ningjie Lei, Jinxiang Cai, Yixin Qian, Zhilong Zheng, Zhiyue Liu, Qingbao Huang
NLPCC (3)7
2023 Enhancing Paraphrase Question Generation With Prior Knowledge
abstract
Paraphrase question generation (PQG) aims to rewrite a given original question to a new paraphrase question, where the paraphrase question needs to have the same expressed meaning as the original question, but have a difference in expression form. Existing methods on PQG mainly focus on synonym substitution or word order adjustment based on the original question. However, rewriting based on the word-level may not guarantee the difference between paraphrase questions and original questions. In this paper, we propose a knowledge-aware paraphrase question generation model. Our model first employs a knowledge extractor to extract the prior knowledge related to the original question from the knowledge base. Then an attention mechanism and a gate mechanism are introduced in our model to selectively utilize the extracted prior knowledge for rewriting, which helps to expand the content of the generated question to maximize the difference. Additionally, we use a discriminator module to promote the generated paraphrase to be semantically close to the original question and the ground truth. Specifically, the loss function of the discriminator penalizes the excessive distance between the representation of the paraphrase question and the ground truth. Extensive experiments on the Quora dataset show that the proposed model outperforms the baselines. Further, our model is applied to the SQuAD dataset, which proves the generalization ability of our model in the existing QA dataset.
Jiayuan Xie, Wenhao Fang, Qingbao Huang, Yi Cai 0001, Tao Wang 0036
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Scene-Text Oriented Referring Expression Comprehension
abstract
Referring expression comprehension (REC) aims to identify and locate a specific object in visual scenes referred to by a natural language expression. Existing studies of REC only focus on basic visual attributes and neglect scene text. Since scene text has the functions of object identification and disambiguation, it is naturally and frequently used to refer to objects. However, existing methods do not explicitly recognize text in images and fail to align scene text mentioned in expressions with the text shown in images, resulting in object localization errors. This article takes the first step toward addressing these limitations. First, we introduce a new task called scene-text oriented referring expression comprehension, which aims to align visual cues and textual semantics of scene text with referring expressions and visual contents. Second, we propose a scene text awareness network that can bridge the gap between texts from two modalities by grounding visual representations of expression-correlated scene texts. Specifically, we propose a correlated text extraction module to solve the problem of lacking semantic understanding, and a correlated region activation module to address the fixed alignment problem and absent alignment problem. These modules ensure that the proposed method focuses on local regions that are most relevant to scene text, thus mitigating the misalignment of scene text with irrelevant regions. Third, to conduct quantitative evaluations, we establish a new benchmark dataset called RefText. Experimental results demonstrate that the proposed method can effectively comprehend scene-text oriented referring expressions and achieves excellent performance.
Yuqi Bu, Liuwu Li, Jiayuan Xie, Qiong Liu 0006, Yi Cai 0001, Qingbao Huang, Qing Li 0001
IEEE Trans. Multim.6
2023 Visual Paraphrase Generation with Key Information Retained
abstract
Visual paraphrase generation task aims to rewrite a given image-related original sentence into a new paraphrase, where the paraphrase needs to have the same expressed meaning as the original sentence but have a difference in expression form. Existing studies mainly extract two semantic vectors to represent the entire image and the entire original sentence, respectively, for paraphrase generation. However, these semantic vectors for an image or a sentence may lead to the model failing to focus on some key objects in the original sentence, which may generate semantically inconsistent sentences by changing key object information. In this article, we propose an object-level paraphrase generation model, which generates paraphrases by adjusting the permutation of key objects and modifying their associated descriptions. To adjust the permutation of key objects, an object-sorting module aims to obtain new object sequences based on the key object information and original sentences. Then, a sequence generation module sequentially generates paraphrases based on the permutation of the newly object sequences. Each generation step focuses on different image features associated with different key objects to generate descriptions with differences. Furthermore, we use a semantic discriminator module to promote the generated paraphrase to be semantically close to the original sentence. Specifically, the loss function of the discriminator penalizes the excessive distance between the paraphrase and the original sentence. Extensive experiments on the MS COCO dataset show that the proposed model outperforms the baselines.
Jiayuan Xie, Yi Cai 0001, Qingbao Huang, Qing Li 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2022 SSAP: Storylines and Sentiment Aware Pre-Trained Model for Story Ending Generation
abstract
As an interesting but under-explored task, story ending generation aims at generating an appropriate ending for an incomplete story. The challenges of the task are to deeply understand the story context, mine the storylines hidden in the story, and generate rational endings in logic and sentiment. Although existing pre-trained approaches have been proven effective to this task, how to learn to generate endings with appropriate plots and sufficient sentimental information still remains a major challenge. One possible reason is that an over reliance on external commonsense knowledge beyond the storylines and sentimental trends information hidden in the story context could lead to generation deviating from the main theme. To address this issue, we propose a two-stageStroylines andSentimentAwarePre-trained model (SSAP) for generating sentimentally relevant story endings. We apply a classifier for discriminating the sentiment of the story, and then employ a pre-trained language model, combining with storylines information, to conditionally generate sentences that match both the logic and sentiment of the story. Automatic and manual evaluations show that, without integrating external knowledge, our model can produce more consistent and diverse story endings than state-of-the-art baselines.
Qingbao Huang, Jing Li 0049, Linzhang Mo, Yi Cai 0001, Qing Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Diverse Distractor Generation for Constructing High-Quality Multiple Choice Questions
abstract
Distractor generation task aims to generate incorrect options (i.e., distractors) for multiple choice questions from an article.Existing methods for this task often utilize a standard encoder-decoder framework. However, these methods often tend to generate semantically similar distractors, since the same article representations are used to generate different distractors. Multiple generated distractors with similar semantics are considered equivalent. Because the correct answer is unique, students can eliminate these distractors even without reading the article. In this paper, we propose a multi-selector generation network (MSG-Net) that generates distractors with rich semantics based on different sentences in an article. MSG-Net adopts a multi-selector mechanism to select multiple different sentences in an article that are useful to generate diverse distractors. Specifically, a question-aware and answer-aware mechanism are introduced to assist in selecting useful key sentences, where each key sentence is coherent with the question and not equivalent to the answer. MSG-Net can generate diverse distractors based on each selected key sentence with different semantics. Extensive experiments on the RACE dataset and Cosmos QA dataset show that the proposed model outperforms the state-of-the-art models in generating diverse distractors.
Jiayuan Xie, Ningxin Peng, Yi Cai 0001, Tao Wang 0036, Qingbao Huang
IEEE ACM Trans. Audio Speech Lang. Process.5
2022 Knowledge-Based Visual Question Generation
abstract
Visual question generation task aims to generate meaningful questions about an image targeting an answer. Existing methods focus on the visual concepts in the image for question generation. However, humans inevitably use their knowledge related to visual objects in images to construct questions. In this paper, we propose a knowledge-based visual question generation model that can integrate visual concepts and non-visual knowledge to generate questions. To obtain visual concepts, we utilize a pre-trained object detection model to obtain object-level features of each object in the image. To obtain useful non-visual knowledge, we first retrieve the knowledge from the knowledge-base related to the visual objects in the image. Considering that not all retrieved knowledge is helpful for this task, we introduce an answer-aware module to capture the candidate knowledge related to the answer from the retrieved knowledge, which ensures that the generated content can be targeted at the answer. Finally, object-level representations containing visual concepts and non-visual knowledge are sent to a decoder module to generate questions. Extensive experiments on the FVQA and KBVQA datasets show that the proposed model outperforms the state-of-the-art models.
Jiayuan Xie, Wenhao Fang, Yi Cai 0001, Qingbao Huang, Qing Li 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 Image Difference Captioning With Instance-Level Fine-Grained Feature Representation
abstract
The task of image difference captioning aims at locating changed objects in similar image pairs and describing the difference with natural language. The key challenges of this task are to comprehend the context of image pairs sufficiently and locate the changed objects accurately in the presence of viewpoint change. Previous studies focus on pixel-level image features, neglecting rich explicit features of objects in an image pair which are beneficial to generate a fine-grained difference caption. Additionally, existing generative models suffer from accurately locate the differences in the interference of viewpoint change. To address these issues, we propose an Instance-Level Fine-Grained Difference Captioning (IFDC) model, which consists of a fine-grained feature extraction module, a multi-round feature fusion module, a similarity-based difference finding module, and a difference captioning module. To describe the changed objects comprehensively, we extract the fine-grained features, i.e., visual features, semantic features, and positional features at instance-level, as the objects’ representation. To enhance the model’s immunity to viewpoint change, we design a similarity-based difference finding module to locate the changed objects accurately. Extensive experiments show that our IFDC model achieves comparable performance with the state-of-the-art models on the datasets of CLEVR-Change and Spot-the-Diff, thus verifying the effectiveness of our proposed model. Our source code is available athttps://github.com/VISLANG-Lab/IFDC.
Qingbao Huang, Jielong Wei, Yi Cai 0001, Hanyu Liang, Ho-fung Leung, Qing Li 0001
IEEE Trans. Multim.1
2022 Suppressing Biased Samples for Robust VQA
abstract
Most existing visual question answering (VQA) models strongly rely on language bias to answer questions, i.e., they always tend to fit question-answer pairs on the train split and perform poorly on the test spilt when the answer distributions are different. This behavior makes them hard to be applied in real scenarios. To reduce the language biases, previous studies mainly integrate modules to overcome language priors (ensemble-based methods) or generate additional training data to balance dataset biases (data-balanced methods). However, all the existing ensemble-based methods drop their accuracies on the VQA v2 dataset, while data-balanced methods may introduce new biases and cannot guarantee the quality of the generated data. In this paper, we propose a model-agnostic training scheme called Suppressing Biased Samples (SBS) to overcome language priors. SBS consists of two collaborative parts, i.e., a Data Classifier Module to divide the dataset into biased samples and unbiased samples by utilizing the similarity in the semantic space, and a Bias Penalty Module to suppress the biased samples to weaken their influence. As a new way of balancing data to address language bias, SBS overcomes the shortcomings of previous data-balanced methods. Experimental results show that our method can be merged into other bias-reduction methods and achieves a new state-of-the-art performance on the commonly used VQA-CP v2 dataset.
Ninglin Ouyang, Qingbao Huang, Pijian Li, Yi Cai 0001, Bin Liu 0053, Ho-fung Leung, Qing Li 0001
IEEE Trans. Multim.2
2021 Entity Guided Question Generation with Contextual Structure and Sequence Information Capturing
abstract
Question generation is a challenging task and has attracted widespread attention in recent years. Although previous studies have made great progress, there are still two main shortcomings: First, previous work did not simultaneously capture the sequence information and structure information hidden in the context, which results in poor results of the generated questions. Second, the generated questions cannot be answered by the given context. To tackle these issues, we propose an entity guided question generation model with contextual structure information and sequence information capturing. We use a Graph Convolutional Network and a Bidirectional Long Short Term Memory Network to capture the structure information and sequence information of the context, simultaneously. In addition, to improve the answerability of the generated questions, we use an entity-guided approach to obtain question type from the answer, and jointly encode the answer and question type. Both automatic and manual metrics show that our model can generate comparable questions with state-of-the-art models. Our code is available at https://github.com/VISLANG-Lab/EGSS.
Qingbao Huang, Mingyi Fu, Linzhang Mo, Yi Cai 0001, Pijian Li, Qing Li 0001, Ho-fung Leung
AAAI1
2021 Story Ending Generation with Multi-Level Graph Convolutional Networks over Dependency Trees
abstract
As an interesting and challenging task, story ending generation aims at generating a reasonable and coherent ending for a given story context. The key challenge of the task is to comprehend the context sufficiently and capture the hidden logic information effectively, which has not been well explored by most existing generative models. To tackle this issue, we propose a context-aware Multi-level Graph Convolutional Networks over Dependency Parse (MGCN-DP) trees to capture dependency relations and context clues more effectively. We utilize dependency parse trees to facilitate capturing relations and events in the context implicitly, and Multi-level Graph Convolutional Networks to update and deliver the representation crossing levels to obtain richer contextual information. Both automatic and manual evaluations show that our MGCN-DP can achieve comparable performance with state-of-the-art models. Our source code is available at https://github.com/VISLANG-Lab/MLGCN-DP.
Qingbao Huang, Linzhang Mo, Pijian Li, Yi Cai 0001, Qingguang Liu, Jielong Wei, Qing Li 0001, Ho-fung Leung
AAAI1
2021 Scene Graph with 3D Information for Change Captioning
abstract
Change captioning aims to describe the differences in image pairs with natural language. It is an interesting task under-explored with two main challenges: describing the relative position relationship between objects correctly and overcoming the disturbances from viewpoint changes. To address these issues, we propose a three-dimensional (3D) information aware Scene Graph based Change Captioning (SGCC) model. We extract the semantic attributes of objects and the 3D information of images (i.e., depths of objects, relative two-dimensional image plane distances, and relative angles between objects) to construct the scene graphs for image pairs, then aggregate the nodes representations with a graph convolutional network. Owing to the relative position relationships between objects and the scene graphs, our model thereby is capable of assisting observers to locate the changed objects quickly and being immune to the viewpoint change to some extent. Extensive experiments show that our SGCC model achieves competitive performance with the state-of-the-art models on the CLEVR-Change and Spot-the-Diff datasets, thus verifying the effectiveness of our proposed model. Codes are available at https://github.com/VISLANG-Lab/SGCC.
Zeming Liao, Qingbao Huang, Mingyi Fu, Yi Cai 0001, Qing Li 0001
ACM Multimedia2
2021 Multiple Objects-Aware Visual Question Generation
abstract
Visual question generation task aims to generate meaningful questions about an image according to a target answer. Existing studies mainly focus on merely one object related to the target answer in an image to generate a question. However, a target answer is often related to multiple key objects in an image, which focuses on only one object may mislead its model to generate questions that are only related to partial fragments of the answer. To address this problem, we propose a multi-objects aware generation model to capture all key objects related to an answer and generate the corresponding question. We first introduce a co-attention network to capture the relationship between each object in an image and the answer, and then extract the key objects that are related to the answer. Then, a graph network is introduced to capture the relationships between the key objects and other objects in the image that are not related to the answer, which helps generate questions that involve more visual content. Finally, the learned information from the graph network is fed into a standard decoder module to produce questions. Extensive experiments on the VQA v2.0 dataset show that the proposed model outperforms the state-of-the-art models.
Jiayuan Xie, Yi Cai 0001, Qingbao Huang, Tao Wang 0036
ACM Multimedia3
2021 Incorporating sentimental trend into gated mechanism based transformer network for story ending generation
Linzhang Mo, Jielong Wei, Qingbao Huang, Yi Cai 0001, Qingguang Liu, Xingmao Zhang, Qing Li 0001
Neurocomputing3
2021 Topic-level knowledge sub-graphs for multi-turn dialogue generation
Jing Li 0049, Qingbao Huang, Yi Cai 0001, Mingyi Fu, Qing Li 0001
Knowl. Based Syst.2
2020 Aligned Dual Channel Graph Convolutional Network for Visual Question Answering
abstract
Visual question answering aims to answer the natural language question about a given image. Existing graph-based methods only focus on the relations between objects in an image and neglect the importance of the syntactic dependency relations between words in a question. To simultaneously capture the relations between objects in an image and the syntactic dependency relations between words in a question, we propose a novel dual channel graph convolutional network (DC-GCN) for better combining visual and textual advantages. The DC-GCN model consists of three parts: an I-GCN module to capture the relations between objects in an image, a Q-GCN module to capture the syntactic dependency relations between words in a question, and an attention alignment module to align image representations and question representations. Experimental results show that our model achieves comparable performance with the state-of-theart approaches.
Qingbao Huang, Jielong Wei, Yi Cai 0001, Changmeng Zheng, Ho-fung Leung, Qing Li 0001
ACL1
2020 Incorporating context-relevant concepts into convolutional neural networks for short text classification
Yi Cai 0001, Xin Wu 0003, Xue Lei, Qingbao Huang, Ho-fung Leung, Qing Li 0001
Neurocomputing5
2020 Recurrent neural network with pooling operation and attention mechanism for sentiment analysis: A multi-task learning approach
Yi Cai 0001, Qingbao Huang, Zejun Lin, Qing Li 0001
Knowl. Based Syst.2
2014 Adaptive dynamic programming-based optimal tracking control for nonlinear systems using general value iteration
abstract
For the optimal tracking control problem of affine nonlinear systems, a general value iteration algorithm based on adaptive dynamic programming is proposed in this paper. By system transformation, the optimal tracking problem is converted into the optimal regulating problem for the tracking error dynamics. Then, general value iteration algorithm is developed to obtain the optimal control with convergence analysis. Considering the advantages of echo state network, we use three echo state networks with levenberg-Marquardt (LM) adjusting algorithm to approximate the system, the cost function and the control law. A simulation example is given to demonstrate the effectiveness of the presented scheme.
Xiaofeng Lin 0002, Weikai Kong, Chunning Song, Qingbao Huang
ADPRL5
2011 Research on Water Level Optimal Control of Boiler Drum Based on Dual Heuristic Dynamic Programming
Qingbao Huang, Shaojian Song, Xiaofeng Lin 0002, Kui Peng
ISNN (1)1
2009 Temperature Control in Cement Rotary Kiln with Neural Network-Based Heuristic Dynamic Programming
Xiaofeng Lin 0002, Tangbo Liu, Deguang Cao, Qingbao Huang
ISNN (2)4
2007 Project-Based Artificial Neural Networks Development Software and Applications
Xiaofeng Lin 0002, Shaojian Song, Chunning Song, Qingbao Huang, Xiao xiao Song
ISNN (3)4