Changmeng Zheng

dblp:254/8208 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0002-2945-8248ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Multi-Agent Undercover Gaming: Hallucination Removal Through Counterfactual Test for Multimodal Reasoning
abstract
Hallucination continues to pose a major obstacle in the reasoning capabilities of large language models (LLMs). Although the Multi-Agent Debate (MAD) paradigm offers a promising solution by promoting consensus among multiple agents to enhance reliability, it relies on the unrealistic assumption that all debaters are rational and reflective, which is a condition that may not hold when agents themselves are prone to hallucinations. To address this gap, we introduce the Multi-agent Undercover Gaming (MUG) protocol, inspired by social deduction games like ''Who is Undercover?''. MUG reframes MAD as a process of detecting ''undercover'' agents (those suffering from hallucinations) by employing multimodal counterfactual tests. Specifically, we modify reference images to introduce counterfactual evidence and observe whether agents can accurately identify these changes, providing ground-truth for identifying hallucinating agents and enabling robust, crowd-powered multimodal reasoning. MUG advances MAD protocols along three key dimensions: (1) enabling factual verification beyond statistical consensus through counterfactual testing; (2) introducing cross-evidence reasoning via dynamically modified evidence sources instead of relying on static inputs; and (3) fostering active reasoning, where agents engage in probing discussions rather than passively answering questions. Collectively, these innovations offer a more reliable and effective framework for multimodal reasoning in LLMs.
Dayong Liang, Xiaoyong Wei, Changmeng Zheng
AAAI3
2026 Hybrid-DMKG: A Hybrid Reasoning Framework over Dynamic Multimodal Knowledge Graphs for Multimodal Multihop QA with Knowledge Editing
abstract
Multimodal Knowledge Editing (MKE) extends traditional knowledge editing to settings involving both textual and visual modalities. However, existing MKE benchmarks primarily assess final answer correctness, neglecting the quality of intermediate reasoning and robustness to visually rephrased inputs. To address this limitation, we introduce MMQAKE, the first benchmark for multimodal multihop question answering with knowledge editing. MMQAKE evaluates: (1) a model’s ability to reason over 2–5-hop factual chains that span both text and images, including performance at each intermediate step; (2) robustness to visually rephrased inputs in multihop questions. Our evaluation shows that current MKE methods often struggle to consistently update and reason over multimodal reasoning chains following knowledge edits. To overcome these challenges, we propose Hybrid-DMKG, a hybrid reasoning framework built on a dynamic multimodal knowledge graph (DMKG) to enable accurate multihop reasoning over updated multimodal knowledge. Hybrid-DMKG first uses a large language model to decompose multimodal multihop questions into sequential sub-questions, then applies a multimodal retrieval model to locate updated facts by jointly encoding each sub-question with candidate entities and their associated images. For answer inference, a hybrid reasoning module operates over the DMKG via two parallel paths: (1) relation-linking prediction; (2) RAG Reasoning with large vision-language models. A background-reflective decision module then aggregates evidence from both paths to select the most credible answer. Experimental results on MMQAKE show that Hybrid-DMKG significantly outperforms existing MKE approaches, achieving higher accuracy and improved robustness to knowledge updates.
Qingfei Huang, Bingshan Zhu, Yi Cai 0001, Qingbao Huang, Changmeng Zheng, Zikun Deng, Tao Wang 0036
AAAI6
2024 A Picture Is Worth a Graph: A Blueprint Debate Paradigm for Multimodal Reasoning
abstract
This paper presents a pilot study aimed at introducing multi-agent debate into multimodal reasoning. The study addresses two key challenges: the trivialization of opinions resulting from excessive summarization and the diversion of focus caused by distractor concepts introduced from images. These challenges stem from the inductive (bottom-up) nature of existing debating schemes. To address the issue, we propose a deductive (top-down) debating approach called Blueprint Debate on Graphs (BDoG). In BDoG, debates are confined to a blueprint graph to prevent opinion trivialization through world-level summarization. Moreover, by storing evidence in branches within the graph, BDoG mitigates distractions caused by frequent but irrelevant concepts. Extensive experiments validate that BDoG is able to achieve state-of-the-art results in ScienceQA and MMBench with significant improvements over previous methods. The source code can be accessed at https://github.com/thecharm/BDoG.
Changmeng Zheng, Dayong Liang, Wengyu Zhang, Xiaoyong Wei, Tat-Seng Chua, Qing Li 0001
ACM Multimedia1
2024 Towards Bridged Vision and Language: Learning Cross-Modal Knowledge Representation for Relation Extraction
abstract
In natural language processing, relation extraction (RE) is to detect and classify the semantic relationship of two given entities within a sentence. Previous RE methods consider only the textual contents and suffer performance decline in social media when texts lack contexts. Incorporating text-related visual information can supplement the missing semantics for relation extraction in social media posts. However, textual relations are usually abstract and of high-level semantics, which causes the semantic gap between visual contents and textual expressions. In this paper, we propose RECK - a neural network for relation extraction with cross-modal knowledge representations. Different from previous multimodal methods training a common subspace for all modalities, we bridge the semantic gaps by explicitly selecting knowledge paths from external knowledge through the cross-modal object-entity pairs. We further extend the paths into a knowledge graph, and adopt a graph attention network to capture the multi-grained relevant concepts which can provide higher level and key semantics information from external knowledge. Besides, we employ a cross-modal attention mechanism to align and fuse the multimodal information. Experimental results on a multimodal RE dataset show that our model achieves new state-of-the-art performance with knowledge evidence.
Guohua Wang 0003, Changmeng Zheng, Yi Cai 0001, Ze Fu, Yaowei Wang 0001, Xiaoyong Wei, Qing Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 Rethinking Multimodal Entity and Relation Extraction from a Translation Point of View
abstract
We revisit the multimodal entity and relation extraction from a translation point of view.Special attention is paid on the misalignment issue in text-image datasets which may mislead the learning.We are motivated by the fact that the cross-modal misalignment is a similar problem of cross-lingual divergence issue in machine translation.The problem can then be transformed and existing solutions can be borrowed by treating a text and its paired image as the translation to each other.We implement a multimodal back-translation using diffusionbased generative models for pseudo-paralleled pairs and a divergence estimator by constructing a high-resource corpora as a bridge for low-resource learners.Fine-grained confidence scores are generated to indicate both types and degrees of alignments with which better representations are obtained.The method has been validated in the experiments by outperforming 14 state-of-the-art methods in both entity and relation extraction tasks.The source code is available at https://github.com/thecharm/TMR.
Changmeng Zheng, Yi Cai 0001, Xiaoyong Wei, Qing Li 0001
ACL (1)1
2023 DRAKE: Deep Pair-Wise Relation Alignment for Knowledge-Enhanced Multimodal Scene Graph Generation in Social Media Posts
abstract
Scene Graph Generation (SGG) is a typical computer vision task that detects objects and corresponding predicates in an image. Existing SGG methods focus on modeling visual contexts to generate scene graphs and are conducted on well-annotated datasets with high-quality images. However, the quality is unguaranteed for images in social media posts, so that some images may be incomplete or occluded by some obstacles, hence might not provide sufficient visual context for SGG. Therefore, previous methods might result in missing or false visual relationship detection due to lacking visual contexts. To effectively generate the scene graphs in social media, we study multimodal scene graph generation (MSG) in this paper. MSG aims to develop visual scene graphs from images in social media posts with the support of text sentences. However, leveraging textual contents by simple multimodal alignment such as object-level alignment neglects the inherent pair-wise mapping between multimodal object pairs. To address the limitations, we propose a method named Deep pair-wise Relation Alignment for Knowledge-Enhanced (DRAKE) multimodal scene graph generation. The model supplements the missing visual contexts with well-aligned textual knowledge. It first represents the textual information into object-aware knowledge representation with the help of vision data. Furthermore, our proposed DRAKE facilitates the interaction of the info between multimodal pair-wise representations. A multimodal context enhancement layer can be devised to help the model generate the scene graph. To evaluate the model performance of SGG on social media images, we propose a social media SGG dataset called MSG. We comprehensively analyze the effectiveness of our proposed method on the MSG dataset. The experimental results on the MSG dataset indicate that our model outperforms the previous methods. To fairly compare our method with other SGG models, we also conduct experiments on the Visual Genome dataset for more analysis The MSG dataset is released onhttps://github.com/FuZe4ever/MSG.
Ze Fu, Changmeng Zheng, Yi Cai 0001, Xiaoyong Wei, Yaowei Wang 0001, Qing Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 Knowledge-Enhanced Scene Graph Generation with Multimodal Relation Alignment (Student Abstract)
abstract
Existing scene graph generation methods suffer the limitations when the image lacks of sufficient visual contexts. To address this limitation, we propose a knowledge-enhanced scene graph generation model with multimodal relation alignment, which supplements the missing visual contexts by well-aligned textual knowledge. First, we represent the textual information into contextualized knowledge which is guided by the visual objects to enhance the contexts. Furthermore, we align the multimodal relation triplets by co-attention module for better semantics fusion. The experimental results show the effectiveness of our method.
Ze Fu, Changmeng Zheng, Yi Cai 0001
AAAI3
2021 An Entity-Aware Adversarial Domain Adaptation Network for Cross-Domain Named Entity Recognition (Student Abstract)
abstract
Existing methods for named entity recognition (NER) are critically relied on the amount of labeled data. However, these methods suffer from performance decline in a new domain which is fully-unlabeled. To handle the situation, we propose an entity-aware adversarial domain adaptation network, which utilizes the labeled data from source domain and then adapts to unlabeled target domain. We first apply adversarial training to reduce the distribution gap between different domains. Furthermore, we introduce an entity-aware attention to guide adversarial to achieve the alignment of entity features. The experimental results show that our model outperforms the state-of-the-art approaches.
Qi Peng 0002, Changmeng Zheng, Yi Cai 0001, Tao Wang 0036, Haoran Xie 0001, Qing Li 0001
AAAI2
2021 MNRE: A Challenge Multimodal Dataset for Neural Relation Extraction with Visual Evidence in Social Media Posts
abstract
Extracting relations in social media posts is challenging when sentences lack of contexts. However, images related to these sentences can supplement such missing contexts and help to identify relations precisely. To this end, we present a multimodal neural relation extraction dataset (MNRE), consisting of 10000+ sentences on 31 relations derived from Twitter and annotated by crowdworkers. The subject and object entities are recognized by a pretrained NER tool and then filtered by crowdworkers. All the relations are identified manually. One sentence is tagged with one related image. We develop a multimodal relation extraction baseline model and the experimental results show that introducing multimodal information improves relation extraction performance in social media texts. Still, our detailed analysis points out the difficulties of aligning relations in texts and images, which can be addressed for future research. All details and resources about the dataset and baselines are released on https://github.com/thecharm/MNRE.
Changmeng Zheng, Ze Fu, Yi Cai 0001
ICME1
2021 Multimodal Relation Extraction with Efficient Graph Alignment
abstract
Relation extraction (RE) is a fundamental process in constructing knowledge graphs. However, previous methods on relation extraction suffer sharp performance decline in short and noisy social media texts due to a lack of contexts. Fortunately, the related visual contents (objects and their relations) in social media posts can supplement the missing semantics and help to extract relations precisely. We introduce the multimodal relation extraction (MRE), a task that identifies textual relations with visual clues. To tackle this problem, we present a large-scale dataset which contains 15000+ sentences with 23 pre-defined relation categories. Considering that the visual relations among objects are corresponding to textual relations, we develop a dual graph alignment method to capture this correlation for better performance. Experimental results demonstrate that visual contents help to identify relations more precisely against the text-only baselines. Besides, our alignment method can find the correlations between vision and language, resulting in better performance. Our dataset and code are available at https://github.com/thecharm/Mega.
Changmeng Zheng, Ze Fu, Yi Cai 0001, Qing Li 0001, Tao Wang 0036
ACM Multimedia1
2021 Unsupervised cross-domain named entity recognition using entity-aware adversarial training
Qi Peng 0002, Changmeng Zheng, Yi Cai 0001, Tao Wang 0036, Haoran Xie 0001, Qing Li 0001
Neural Networks2
2021 Object-Aware Multimodal Named Entity Recognition in Social Media Posts With Adversarial Learning
abstract
Named Entity Recognition (NER) in social media posts is challenging since texts are usually short and contexts are lacking. Most recent works show that visual information can boost the NER performance since images can provide complementary contextual information for texts. However, the image-level features ignore the mapping relations between fine-grained visual objects and textual entities, which results in error detection in entities with different types. To better exploit visual and textual information in NER, we propose an adversarial gated bilinear attention neural network (AGBAN). The model jointly extracts entity-related features from both visual objects and texts, and leverages an adversarial training to map two different representations into a shared representation. As a result, domain information contained in an image can be transferred and applied for extracting named entities in the text associated with the image. Experimental results on Tweets dataset demonstrate that our model outperforms the state-of-the-art methods. Moreover, we systematically evaluate the effectiveness of the proposed gated bilinear attention network in capturing the interactions of mutimodal features visual objects and textual words. Our results indicate that the adversarial training can effectively exploit commonalities across heterogeneous data sources, which leads to improved performance in NER when compared to models purely exploiting text data or combining the image-level visual features.
Changmeng Zheng, Tao Wang 0036, Yi Cai 0001, Qing Li 0001
IEEE Trans. Multim.1
2020 Aligned Dual Channel Graph Convolutional Network for Visual Question Answering
abstract
Visual question answering aims to answer the natural language question about a given image. Existing graph-based methods only focus on the relations between objects in an image and neglect the importance of the syntactic dependency relations between words in a question. To simultaneously capture the relations between objects in an image and the syntactic dependency relations between words in a question, we propose a novel dual channel graph convolutional network (DC-GCN) for better combining visual and textual advantages. The DC-GCN model consists of three parts: an I-GCN module to capture the relations between objects in an image, a Q-GCN module to capture the syntactic dependency relations between words in a question, and an attention alignment module to align image representations and question representations. Experimental results show that our model achieves comparable performance with the state-of-theart approaches.
Qingbao Huang, Jielong Wei, Yi Cai 0001, Changmeng Zheng, Ho-fung Leung, Qing Li 0001
ACL4
2020 Controllable Abstractive Sentence Summarization with Guiding Entities
abstract
Entities are the major proportion and build up the topic of text summaries.Although existing text summarization models can produce promising results of automatic metrics, for example, ROUGE, it is difficult to guarantee that an entity is contained in generated summaries.In this paper, we propose a controllable abstractive sentence summarization model which generates summaries with guiding entities.Instead of generating summaries from left to right, we start with a selected entity, generate the left part first, then the right part of a complete summary.Compared to previous entity-based text summarization models, our method can ensure that entities appear in final output summaries rather than generating the complete sentence with implicit entity and article representations.Our model can also generate more novel entities with them incorporated into outputs directly.To evaluate the informativeness of the proposed model, we develop a fine-grained informativeness metrics in the relevance, extraness and omission perspectives.We conduct experiments in two widely-used sentence summarization datasets and experimental results show that our model outperforms the state-of-the-art methods in both automatic evaluation scores and informativeness metrics.
Changmeng Zheng, Yi Cai 0001, Guanjie Zhang, Qing Li 0001
COLING1
2020 Incorporating Concept Information into Term Weighting Schemes for Topic Models
Huakui Zhang, Yi Cai 0001, Bingshan Zhu, Changmeng Zheng, Kai Yang 0007, Raymond Chi-Wing Wong, Qing Li 0001
DASFAA (2)4
2020 Multimodal Representation with Embedded Visual Guiding Objects for Named Entity Recognition in Social Media Posts
abstract
Visual contexts often help to recognize named entities more precisely in short texts such as tweets or snapchat. For example, one can identify "Charlie'' as a name of a dog according to the user posts. Previous works on multimodal named entity recognition ignore the corresponding relations of visual objects and entities. Visual objects are considered as fine-grained image representations. For a sentence with multiple entity types, objects of the relevant image can be utilized to capture different entity information. In this paper, we propose a neural network which combines object-level image information and character-level text information to predict entities. Vision and language are bridged by leveraging object labels as embeddings, and a dense co-attention mechanism is introduced for fine-grained interactions. Experimental results in Twitter dataset demonstrate that our method outperforms the state-of-the-art methods.
Changmeng Zheng, Yi Cai 0001, Ho-fung Leung, Qing Li 0001
ACM Multimedia2
2019 A Boundary-aware Neural Model for Nested Named Entity Recognition
abstract
Changmeng Zheng, Yi Cai, Jingyun Xu, Ho-fung Leung, Guandong Xu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Changmeng Zheng, Yi Cai 0001, Ho-fung Leung, Guandong Xu
EMNLP/IJCNLP (1)1