VLDB 2026 Research / reviewers in the wild / expert
Chaoya Jiang
dblp:270/4680
· DBLP profile ↗
15ranked-venue papers
11as first author
13since 2021 · last 2025
0009-0009-7282-159XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 7 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference OptimizationabstractAs language models continue to scale, Large Language Models (LLMs) have exhibited emerging capabilities in In-Context Learning (ICL), enabling them to solve language tasks by prefixing a few in-context demonstrations (ICDs) as context. Inspired by these advancements, researchers have extended these techniques to develop Large Multimodal Models (LMMs) with ICL capabilities. However, existing LMMs face a critical issue: they often fail to effectively leverage the visual context in multimodal demonstrations and instead simply follow textual patterns. This indicates that LMMs do not achieve effective alignment between multimodal demonstrations and model outputs. To address this problem, we propose Symbol Demonstration Direct Preference Optimization (SymDPO). Specifically, SymDPO aims to break the traditional paradigm of constructing multimodal demonstrations by using random symbols to replace text answers within instances. This forces the model to carefully understand the demonstration images and establish a relationship between the images and the symbols to answer questions correctly. We validate the effectiveness of this method on multiple benchmarks, demonstrating that with SymDPO, LMMs can more effectively understand the multimodal context within examples and utilize this knowledge to answer questions better. Code is available at https://github.com/APiaoG/SymDPO. Hongrui Jia, Chaoya Jiang, Haiyang Xu 0001, Wei Ye 0004, Mengfan Dong, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Shikun Zhang |
CVPR | 2 |
| 2025 | VLM-R³: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-ThoughtabstractRecently, reasoning-based MLLMs have achieved a degree of success in generating long-form textual reasoning chains. However, they still struggle with complex tasks that necessitate dynamic and iterative focusing on and revisiting of visual regions to achieve precise grounding of textual reasoning in visual evidence. We introduce VLM-R³ (Visual Language Model with Region Recognition, Reasoning, and Refinement ), a framework that equips an MLLM with the ability to (i) decide when additional visual evidence is needed, (ii) determine where to ground within the image, and (iii) seamlessly weave the relevant sub-image content back into an interleaved chain-of-thought. The core of our method is \textbf{Region-Conditioned Reinforcement Policy Optimization (R-GRPO)}, a training paradigm that rewards the model for selecting informative regions, formulating appropriate transformations (e.g. crop, zoom), and integrating the resulting visual context into subsequent reasoning steps. To bootstrap this policy, we compile a modest but carefully curated Visuo-Lingual Interleaved Rationale (VLIR) corpus that provides step-level supervision on region selection and textual justification. Extensive experiments on MathVista, ScienceQA, and other benchmarks show that VLM-R$^3$ sets a new state of the art in zero-shot and few-shot settings, with the largest gains appearing on questions demanding subtle spatial reasoning or fine-grained visual cue extraction. Chaoya Jiang, Yongrui Heng, Wei Ye 0004, Haiyang Xu 0001, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Shikun Zhang |
NeurIPS | 1 |
| 2024 | TiMix: Text-Aware Image Mixing for Effective Vision-Language Pre-trainingabstractSelf-supervised Multi-modal Contrastive Learning (SMCL) remarkably advances modern Vision-Language Pre-training (VLP) models by aligning visual and linguistic modalities. Due to noises in web-harvested text-image pairs, however, scaling up training data volume in SMCL presents considerable obstacles in terms of computational cost and data inefficiency. To improve data efficiency in VLP, we propose Text-aware Image Mixing (TiMix), which integrates mix-based data augmentation techniques into SMCL, yielding significant performance improvements without significantly increasing computational overhead. We provide a theoretical analysis of TiMix from a mutual information (MI) perspective, showing that mixed data samples for cross-modal contrastive learning implicitly serve as a regularizer for the contrastive loss. The experimental results demonstrate that TiMix exhibits a comparable performance on downstream tasks, even with a reduced amount of training data and shorter training time, when benchmarked against existing methods. This work empirically and theoretically demonstrates the potential of data mixing for data-efficient and computationally viable VLP, benefiting broader VLP model adoption in practical scenarios. Our code is available on https://github.com/chaoyajiang/TiMiX/tree/main. Chaoya Jiang, Wei Ye 0004, Haiyang Xu 0001, Qinghao Ye, Ming Yan 0008, Ji Zhang 0011, Shikun Zhang |
AAAI | 1 |
| 2024 | Enhancing In-Context Learning via Implicit Demonstration AugmentationabstractXiaoling Zhou, Wei Ye, Yidong Wang, Chaoya Jiang, Zhemg Lee, Rui Xie, Shikun Zhang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Xiaoling Zhou, Wei Ye 0004, Yidong Wang 0003, Chaoya Jiang, Zhemg Lee, Rui Xie 0003, Shikun Zhang |
ACL (1) | 4 |
| 2024 | Hallucination Augmented Contrastive Learning for Multimodal Large Language ModelabstractMulti-modal large language models (MLLMs) have been shown to efficiently integrate natural language with visual information to handle multi-modal tasks. However, MLLMs still face a fundamental limitation of hallucinations, where they tend to generate erroneous or fabricated information. In this paper, we address hallucinations in MLLMs from a novel perspective of representation learning. We first analyzed the representation distribution of textual and visual tokens in MLLM, revealing two important findings: 1) there is a significant gap between textual and visual representations, indicating unsatisfactory cross-modal representation alignment; 2) representations of texts that contain and do not contain hallucinations are entangled, making it challenging to distinguish them. These two observations inspire us with a simple yet effective method to mitigate hallucinations. Specifically, we introduce contrastive learning into MLLMs and use text with hallucination as hard negative examples, naturally bringing representations of non-hallucinative text and visual samples closer while pushing way representations of non-hallucinating and hallucinative text. We evaluate our method quantitatively and qualitatively, showing its effectiveness in reducing hallucination occurrences and improving performance across multiple benchmarks. On the MMhal-Bench benchmark, our method obtains a 34.66% /29.5% improvement over the baseline MiniGPT-4/LLaVA. Our code is available on https://github.com/X-PLUG/mPLUG-HalOwl/tree/main/hacl. Chaoya Jiang, Haiyang Xu 0001, Mengfan Dong, Wei Ye 0004, Ming Yan 0008, Qinghao Ye, Ji Zhang 0011, Fei Huang 0002, Shikun Zhang |
CVPR | 1 |
| 2024 | MIBench: Evaluating Multimodal Large Language Models over Multiple ImagesabstractHaowei Liu, Xi Zhang, Haiyang Xu, Yaya Shi, Chaoya Jiang, Ming Yan, Ji Zhang, Fei Huang, Chunfeng Yuan, Bing Li, Weiming Hu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Haiyang Xu 0001, Yaya Shi, Chaoya Jiang, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004 |
EMNLP | 5 |
| 2024 | PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning OptimizationabstractInstruction tuning large language models (LLMs) remains a challenging task, owing to the complexity of hyperparameter selection and the difficulty involved in evaluating the tuned models. To determine the optimal hyperparameters, an automatic, robust, and reliable evaluation benchmark is essential. However, establishing such a benchmark is not a trivial task due to the challenges associated with evaluation accuracy and privacy protection. In response to these challenges, we introduce a judge large language model, named PandaLM, which is trained to distinguish the superior model given several LLMs. PandaLM's focus extends beyond just the objective correctness of responses, which is the main focus of traditional evaluation datasets. It addresses vital subjective factors such as relative conciseness, clarity, adherence to instructions, comprehensiveness, and formality. To ensure the reliability of PandaLM, we collect a diverse human-annotated test dataset, where all contexts are generated by humans and labels are aligned with human preferences. Our findings reveal that PandaLM-7B offers a performance comparable to both GPT-3.5 and GPT-4. Impressively, PandaLM-70B surpasses their performance. PandaLM enables the evaluation of LLM to be fairer but with less cost, evidenced by significant improvements achieved by models tuned through PandaLM compared to their counterparts trained with default Alpaca's hyperparameters. In addition, PandaLM does not depend on API-based evaluations, thus avoiding potential data leakage. Yidong Wang 0003, Zhuohao Yu 0001, Wenjin Yao, Zhengran Zeng, Linyi Yang, Cunxiang Wang, Hao Chen 0102, Chaoya Jiang, Rui Xie 0003, Jindong Wang 0001, Xing Xie 0001, Wei Ye 0004, Shikun Zhang, Yue Zhang 0004 |
ICLR | 8 |
| 2024 | Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language ModelsabstractLarge Vision-Language Models (LVLMs) exhibit remarkable capabilities but struggle with ''hallucinations''-inconsistencies between images and their descriptions. Previous hallucination evaluation studies on LVLMs have identified hallucinations in terms of objects, attributes, and relations but overlooked complex hallucinations that create an entire narrative around a fictional entity. In this paper, we introduce a refined taxonomy of hallucinations, featuring a new category: Event Hallucination. We then utilize advanced LLMs to generate and filter fine-grained hallucinatory data consisting of various types of hallucinations, with a particular focus on event hallucinations, laying the groundwork for integrating discriminative and generative evaluation methods within our universal evaluation framework. The proposed benchmark distinctively assesses LVLMs' ability to tackle a broad spectrum of hallucinations, making it a reliable and comprehensive tool for gauging LVLMs' efficacy in handling hallucinations. We will release our code and data. Chaoya Jiang, Hongrui Jia, Mengfan Dong, Wei Ye 0004, Haiyang Xu 0001, Ming Yan 0008, Ji Zhang 0011, Shikun Zhang |
ACM Multimedia | 1 |
| 2024 | MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language ModelabstractThis paper presents MaVEn, an innovative Multi-granularity Visual Encoding framework designed to enhance the capabilities of Multimodal Large Language Models (MLLMs) in multi-image reasoning. Current MLLMs primarily focus on single-image visual understanding, limiting their ability to interpret and integrate information across multiple images. MaVEn addresses this limitation by combining discrete visual symbol sequences, which abstract coarse-grained semantic concepts, with traditional continuous representation sequences that model fine-grained features. This dual approach bridges the semantic gap between visual and textual data, thereby improving the model's ability to process and interpret information from multiple images effectively. Additionally, we design a dynamic reduction mechanism by for long-sequence continuous features to enhance multi-image processing efficiency. Experimental results demonstrate that MaVEn significantly enhances MLLMs' understanding in complex multi-image scenarios, while also improving performance in single-image contexts. Chaoya Jiang, Hongrui Jia, Haiyang Xu 0001, Wei Ye 0004, Mengfan Dong, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Shikun Zhang |
NeurIPS | 1 |
| 2023 | Vision Language Pre-training by Contrastive Learning with Cross-Modal Similarity RegulationabstractCross-modal contrastive learning in vision language pretraining (VLP) faces the challenge of (partial) false negatives.In this paper, we study this problem from the perspective of Mutual Information (MI) optimization.It is common sense that InfoNCE loss used in contrastive learning will maximize the lower bound of MI between anchors and their positives, while we theoretically prove that MI involving negatives also matters when noises commonly exist.Guided by a more general lower bound form for optimization, we propose a contrastive learning strategy regulated by progressively refined cross-modal similarity, to more accurately optimize MI between an image/text anchor and its negative texts/images instead of improperly minimizing it.Our method performs competitively on four downstream cross-modal tasks and systematically balances the beneficial and harmful effects of (partial) false negative samples under theoretical guidance. Chaoya Jiang, Wei Ye 0004, Haiyang Xu 0001, Songfang Huang, Fei Huang 0002, Shikun Zhang |
ACL (1) | 1 |
| 2023 | BUS : Efficient and Effective Vision-language Pre-training with Bottom-Up Patch SummarizationabstractVision Transformer (ViT) based Vision-Language Pre-training (VLP) models have demonstrated impressive performance in various tasks. However, the lengthy visual token sequences fed into ViT can lead to training inefficiency and ineffectiveness. Existing efforts address the challenge by either bottom-level patch extraction in the ViT backbone or top-level patch abstraction outside, not balancing training efficiency and effectiveness well. Inspired by text summarization in natural language processing, we propose a Bottom-Up Patch Summarization approach named BUS, coordinating bottom-level extraction and top-level abstraction to learn a concise summary of lengthy visual token sequences efficiently. Specifically, We incorporate a Text-Semantics-Aware Patch Selector (TSPS) into the ViT backbone to perform a coarse-grained visual token extraction and then attach a flexible Transformer-based Patch Abstraction Decoder (PAD) upon the backbone for top-level visual abstraction. This bottom-up collaboration enables our BUS to yield high training efficiency while maintaining or even improving effectiveness. We evaluate our approach on various visual-language understanding and generation tasks and show competitive downstream task performance while boosting the training efficiency by 50%. Additionally, our model achieves state-of-the-art performance on many downstream tasks by increasing input image resolution without increasing computational costs over baselines. Chaoya Jiang, Haiyang Xu 0001, Wei Ye 0004, Qinghao Ye, Chenliang Li 0003, Ming Yan 0008, Bin Bi, Shikun Zhang, Fei Huang 0002, Songfang Huang |
ICCV | 1 |
| 2023 | COPA : Efficient Vision-Language Pre-training through Collaborative Object- and Patch-Text AlignmentabstractVision-Language Pre-training (VLP) methods based on object detection enjoy the rich knowledge of fine-grained object-text alignment but at the cost of computationally expensive inference. Recent Visual-Transformer (ViT)-based approaches circumvent this issue while struggling with long visual sequences without detailed cross-modal alignment information. This paper introduces a ViT-based VLP technique that efficiently incorporates object information through a novel patch-text alignment mechanism. Specifically, we convert object-level signals into patch-level ones and devise a Patch-Text Alignment pre-training task (PTA) to learn a text-aware patch detector. By using off-the-shelf delicate object annotations in 5% training images, we jointly train PTA with other conventional VLP objectives in an end-to-end manner, bypassing the high computational cost of object detection and yielding an effective patch detector that accurately detects text-relevant patches, thus considerably reducing patch sequences and accelerating computation within the ViT backbone. Our experiments on a variety of widely-used benchmarks reveal that our method achieves a speedup of nearly 88% compared to prior VLP models while maintaining competitive or superior performance on downstream tasks with similar model size and data scale. Chaoya Jiang, Haiyang Xu 0001, Wei Ye 0004, Qinghao Ye, Chenliang Li 0003, Ming Yan 0008, Bin Bi, Shikun Zhang, Fei Huang 0002, Ji Zhang 0011 |
ACM Multimedia | 1 |
| 2022 | TRIPS: Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch SelectionabstractVision Transformers (ViTs) have been widely used in large-scale Vision and Language Pretraining (VLP) models.Though previous VLP works have proved the effectiveness of ViTs, they still suffer from computational efficiency brought by the long visual sequence.To tackle this problem, in this paper, we propose an efficient vision-and-language pre-training model with Text-Relevant Image Patch Selection, namely TRIPS, which reduces the visual sequence progressively with a text-guided patchselection layer in the visual backbone for efficient training and inference.The patchselection layer can dynamically compute textdependent visual attention to identify the attentive image tokens with text guidance and fuse inattentive ones in an end-to-end manner.Meanwhile, TRIPS does not introduce extra parameters to ViTs.Experimental results on a variety of popular benchmark datasets demonstrate that TRIPS gain a speedup of 40% over previous similar VLP models, yet with competitive or better downstream task performance. Chaoya Jiang, Haiyang Xu 0001, Chenliang Li 0003, Ming Yan 0008, Wei Ye 0004, Shikun Zhang, Bin Bi, Songfang Huang |
EMNLP | 1 |
| 2020 | Similarity Learning For Cover Song Identification Using Cross-Similarity Matrices of Multi-Level Deep SequencesabstractIn recent years, several deep learning models have been proposed for cover song identification and they have been designed to learn fixed-length feature vectors for music tracks. However, the aspect of temporal progression of music, which is important for measuring the melody similarity between two tracks, is not well represented by fixed-length vectors. In this paper, we propose a new Siamese network architecture for music melody similarity metric learning. The architecture consists of two parts. One part is a network for learning the deep sequence representation of music tracks, and the other is a similarity estimation network which takes as input the cross-similarity matrices calculated from the deep sequences of a pair of tracks. The two networks are jointly trained and optimized to achieve high melody similarity prediction accuracy. Experiments conducted on several public datasets demonstrate the superiority of the proposed architecture. Chaoya Jiang, Deshun Yang, Xiaoou Chen |
ICASSP | 1 |
| 2020 | Learn A Robust Representation For Cover Song Identification Via Aggregating Local And Global Music Temporal ContextabstractRecently, deep learning models have been proposed for cover song identification and designed to learn fixed-length feature vectors for music recordings. However, the aspect of the temporal progression of music, which is important for measuring the melody similarity between two recordings, is not well exploited in those models. In this paper, we propose a new Siamese architecture to learn deep representations for cover song identification where Dilated Temporal Pyramid Convolution is used to exploit the local temporal context and Temporal Self-Attention to exploit the global temporal context in music recordings. In addition to the traditional block which calculates the similarity between a pair of recordings, we add a classification block to classify the recordings to their respective cliques. By combining the regression loss and the classification loss, our model can leam more robust and discriminative latent representations. The representations extracted by our model show substantial superiority to existing hand-crafted features and learned deep features. Experimental results show that our approach far outperforms the state-of the-art methods on several public datasets. Chaoya Jiang, Deshun Yang, Xiaoou Chen |
ICME | 1 |