VLDB 2026 Research / reviewers in the wild / expert
Haoyu Cao 0001
dblp:319/9928-1
· DBLP profile ↗
21ranked-venue papers
5as first author
21since 2021 · last 2026
0000-0002-3789-9705ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 12 · 2 first-author · 12 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multimodal Table Understanding with Difficulty-aware Reinforcement LearningabstractMultimodal table understanding, which aims for a comprehensive grasp of table content by integrating cellular text, tabular structure, and visual presentation, remains a core yet challenging area of research. We identify that the structural complexity of a table, quantifiable by intrinsic properties such as the ratio of merged cells and the total number of cells, presents a significant obstacle for existing models. Our empirical analysis reveals that the performance of leading Multimodal Large Language Models (MLLMs) deteriorates markedly as table complexity increases, exposing a critical vulnerability in their ability to perceive and reason over intricate tabular data. To address this challenge, we propose MM-Table-R1, a model enhanced through difficulty-aware reinforcement learning (RL) post-training strategy. Specifically, we introduce both task-level and data-level curriculum learning. The task-level curriculum is designed to establish a capability ladder, where the model first learns basic perceptual and semantic alignment of table data, and then progresses to acquiring multi-step reasoning capabilities. The data-level curriculum ensures that the model is not exposed to difficult samples prematurely, facilitating a more gradual and effective learning process. Furthermore, we invest considerable effort in constructing a high-quality, large-scale training corpus by curating and processing data from diverse open-source table datasets, ensuring that each instance is paired with an objectively verifiable reward signal. Demonstrating exceptional parameter efficiency, our 3B-parameter model sets a new benchmark by surpassing both established 3B and 7B models, including those specifically designed for table reasoning. Chaohu Liu, Haoyu Cao 0001, YongXiang Hua, Linli Xu 0002 |
AAAI | 2 |
| 2026 | Evaluating the Adversarial Robustness of Vision-Language Models via Internal Feature PerturbationsabstractVision-language models (VLMs), such as BLIP-2 and LLaVA, have significantly advanced multimodal understanding but exhibit critical vulnerabilities to visual adversarial perturbations. The high efficacy of untargeted attacks, in particular, poses significant concerns for their operational robustness. Conventional attack methods generate adversarial examples by backpropagating the language modeling loss from the final output to the input image. However, the deep architecture of the integrated large language models (LLMs) often diminishes this gradient flow, limiting the attack’s effectiveness in perturbing the visual domain. To address this limitation, we introduce a novel untargeted attack method based on Maximizing Information Entropy (MIE). Our approach enhances attack efficacy not only by maximizing the information entropy of the model’s final output but also by directly inducing uncertainty within its internal feature representations. This layer-wise perturbation strategy disrupts the model’s cognitive process more comprehensively than relying on the final output layer alone. We provide a theoretical analysis demonstrating that the uncertainty induced by MIE is greater than or equal to that of conventional output-only attacks. Comprehensive quantitative evaluations across multiple VLM architectures and datasets confirm that our method consistently outperforms existing techniques, thereby establishing a new, more rigorous benchmark for assessing adversarial robustness in vision-language models. Chaohu Liu, Yubo Wang 0010, Haoyu Cao 0001, Deqiang Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | DREAM: Document Reconstruction via End-to-end Autoregressive Model
Xin Li 0118, Mingming Gong, Jianxin Dai, Antai Guo, Xinghua Jiang, Haoyu Cao 0001, Yinsong Liu, Deqiang Jiang, Xing Sun 0001 |
ACM Multimedia | 7 |
| 2025 | Input Domain Aware MoE: Decoupling Routing Decisions from Task Optimization in Mixture of ExpertsabstractSparse Mixture of Experts (sMoE) has become a pivotal approach for scaling large vision-language models, offering substantial capacity while maintaining computational efficiency through dynamic, sparse activation of experts. However, existing routing mechanisms, typically based on similarity scoring, struggle to effectively capture the underlying input structure. This limitation leads to a trade-off between expert specialization and balanced computation, hindering both scalability and performance. We propose Input Domain Aware MoE, a novel routing framework that leverages a probabilistic mixture model to better partition the input space. By modeling routing probabilities as a mixture of distributions, our method enables experts to develop clear specialization boundaries while achieving balanced utilization. Unlike conventional approaches, our routing mechanism is trained independently of task-specific objectives, allowing for stable optimization and decisive expert assignments. Empirical results on vision-language tasks demonstrate that our method consistently outperforms existing sMoE approaches, achieving higher task performance and improved expert utilization balance. YongXiang Hua, Haoyu Cao 0001, Zhou Tao, Bocheng Li, Chaohu Liu, Linli Xu 0002 |
ACM Multimedia | 2 |
| 2025 | CROP: Integrating Topological and Spatial Structures via Cross-View Prefixes for Molecular LLMsabstractRecent advances in molecular science have been propelled significantly by large language models (LLMs). However, their effectiveness is limited when relying solely on molecular sequences, which fail to capture the complex structures of molecules. Beyond sequence representation, molecules exhibit two complementary structural views: the first focuses on the topological relationships between atoms, as exemplified by the graph view; and the second emphasizes the spatial configuration of molecules, as represented by the image view. The two types of views provide unique insights into molecular structures. To leverage these views collaboratively, we propose the CROss-view Prefixes (CROP) to enhance LLMs' molecular understanding through efficient multi-view integration. CROP possesses two advantages: (i) efficiency: by jointly resampling multiple structural views into fixed-length prefixes, it avoids excessive consumption of the LLM's limited context length and allows easy expansion to more views; (ii) effectiveness: by utilizing the LLM's self-encoded molecular sequences to guide the resampling process, it boosts the quality of the generated prefixes. Specifically, our framework features a carefully designed SMILES Guided Resampler for view resampling, and a Structural Embedding Gate for converting the resulting embeddings into LLM's prefixes. Extensive experiments demonstrate the superiority of CROP in tasks including molecule captioning, IUPAC name prediction and molecule property prediction. Jianting Tang, Yubo Wang 0010, Haoyu Cao 0001, Linli Xu 0002 |
ACM Multimedia | 3 |
| 2025 | VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech InteractionabstractRecent Multimodal Large Language Models (MLLMs) have typically focused on integrating visual and textual modalities, with less emphasis placed on the role of speech in enhancing interaction. However, speech plays a crucial role in multimodal dialogue systems, and implementing high-performance in both vision and speech tasks remains a challenge due to the fundamental modality differences. In this paper, we propose a carefully designed multi-stage training methodology that progressively trains LLM to understand both visual and speech information, ultimately enabling fluent vision and speech interaction. Our approach not only preserves strong vision-language capacity, but also enables efficient speech-to-speech dialogue capabilities without separate ASR and TTS modules, significantly accelerating multimodal end-to-end response speed. By comparing against state-of-the-art counterparts across benchmarks for image, video, and speech, we demonstrate that our omni model is equipped with both strong visual and speech capabilities, making omni understanding and interaction. Chaoyou Fu, Haojia Lin, Yifan Zhang 0004, Yunhang Shen, Haoyu Cao 0001, Zuwei Long, Heting Gao, Ke Li 0015, Xiawu Zheng, Rongrong Ji, Xing Sun 0001, Caifeng Shan, Ran He 0001 |
NeurIPS | 7 |
| 2025 | VITA-Audio: Fast Interleaved Audio-Text Token Generation for Efficient Large Speech-Language ModelabstractWith the growing requirement for natural human-computer interaction, speech-based systems receive increasing attention as speech is one of the most common forms of daily communication. However, the existing speech models still experience high latency when generating the first audio token during streaming, which poses a significant bottleneck for deployment. To address this issue, we propose VITA-Audio, an end-to-end large speech model with fast audio-text token generation. Specifically, we introduce a lightweight Multiple Cross-modal Token Prediction (MCTP) module that efficiently generates multiple audio tokens within a single model forward pass, which not only accelerates the inference but also significantly reduces the latency for generating the first audio in streaming scenarios. In addition, a four-stage progressive training strategy is explored to achieve model acceleration with minimal loss of speech quality. To our knowledge, VITA-Audio is the first multi-modal large language model capable of generating audio output during the first forward pass, enabling real-time conversational capabilities with minimal latency. VITA-Audio is fully reproducible and is trained on open-source data only. Experimental results demonstrate that our model achieves an inference speedup of 3~5x at the 7B parameter scale, but also significantly outperforms open-source models of similar model size on multiple benchmarks for automatic speech recognition (ASR), text-to-speech (TTS), and spoken question answering (SQA) tasks. Zuwei Long, Yunhang Shen, Chaoyou Fu, Heting Gao, Lijiang Li, Peixian Chen, Mengdan Zhang, Jian Li 0062, Jinlong Peng, Haoyu Cao 0001, Ke Li 0015, Rongrong Ji, Xing Sun 0001 |
NeurIPS | 11 |
| 2025 | DiffCTE: Consistent Visual Text Editing with High Style Fidelity via Diffusion Model
Haoyu Cao 0001, Anqi Gou, Haobin Cao |
PRCV (5) | 1 |
| 2025 | Towards Trustworthy Document Information Extraction with Self-evaluation Mechanism
Haoyu Cao 0001, Anqi Gou, Haobin Cao |
PRCV (18) | 1 |
| 2024 | Talk With Human-like Agents: Empathetic Dialogue Through Perceptible Acoustic Reception and ReactionabstractLarge Language Model (LLM)-enhanced agents become increasingly prevalent in Human-AI communication, offering vast potential from entertainment to professional domains.However, current multi-modal dialogue systems overlook the acoustic information present in speech, which is crucial for understanding human communication nuances.This oversight can lead to misinterpretations of speakers' intentions, resulting in inconsistent or even contradictory responses within dialogues.To bridge this gap, in this paper, we propose PerceptiveAgent, an empathetic multi-modal dialogue system designed to discern deeper or more subtle meanings beyond the literal interpretations of words through the integration of speech modality perception.Employing LLMs as a cognitive core, Per-ceptiveAgent perceives acoustic information from input speech and generates empathetic responses based on speaking styles described in natural language.Experimental results indicate that PerceptiveAgent excels in contextual understanding by accurately discerning the speakers' true intentions in scenarios where the linguistic meaning is either contrary to or inconsistent with the speaker's true feelings, producing more nuanced and expressive spoken dialogues.Code is publicly available at: https://github.com/Haoqiu-Yan/ PerceptiveAgent. Haoqiu Yan, Yongxin Zhu 0003, Haoyu Cao 0001, Deqiang Jiang, Linli Xu 0002 |
ACL (1) | 5 |
| 2024 | Few-shot Temporal Pruning Accelerates Diffusion Models for Text GenerationabstractDiffusion models have achieved significant success in computer vision and shown immense potential in natural language processing applications, particularly for text generation tasks. However, generating high-quality text using these models often necessitates thousands of iterations, leading to slow sampling rates. Existing acceleration methods either neglect the importance of the distribution of sampling steps, resulting in compromised performance with smaller number of iterations, or require additional training, introducing considerable computational overheads. In this paper, we present Few-shot Temporal Pruning, a novel technique designed to accelerate diffusion models for text generation without supplementary training while effectively leveraging limited data. Employing a Bayesian optimization approach, our method effectively eliminates redundant sampling steps during the sampling process, thereby enhancing the generation speed. A comprehensive evaluation of discrete and continuous diffusion models across various tasks, including machine translation, question generation, and paraphrasing, reveals that our approach achieves competitive performance even with minimal sampling steps after down to less than 1 minute of optimization, yielding a significant acceleration of up to 400x in text generation tasks. Bocheng Li, Zhujin Gao, Yongxin Zhu 0003, Kun Yin, Haoyu Cao 0001, Deqiang Jiang, Linli Xu 0002 |
LREC/COLING | 5 |
| 2024 | Enhancing Visual Document Understanding with Contrastive Learning in Large Visual-Language ModelsabstractRecently, the advent of Large Visual-Language Models (LVLMs) has received increasing attention across various domains, particularly in the field of visual document understanding (VDU). Different from conventional vision-language tasks, VDU is specifically concerned with text-rich scenarios containing abundant document elements. Nevertheless, the importance of fine-grained features remains largely unexplored within the community of LVLMs, leading to suboptimal performance in text-rich scenarios. In this paper, we abbreviate it as the fine-grained feature collapse issue. With the aim of filling this gap, we propose a contrastive learning framework, termed Document Object COntrastive learning (DoCo), specifically tailored for the downstream tasks of VDU. DoCo leverages an auxiliary multimodal encoder to obtain the features of document objects and align them to the visual features generated by the vision encoder of LVLM, which enhances visual representation in text-rich scenarios. It can represent that the contrastive learning between the visual holistic representations and the multimodal fine-grained features of document objects can assist the vision encoder in acquiring more effective visual cues, thereby enhancing the comprehension of text-rich documents in LVLMs. We also demonstrate that the proposed DoCo serves as a plug-and-play pre-training method, which can be employed in the pre-training of various LVLMs without inducing any increase in computational complexity during the inference process. Extensive experimental results on multiple benchmarks of VDU reveal that LVLMs equipped with our proposed DoCo can achieve superior performance and mitigate the gap between VDU and generic vision-language tasks. Xin Li 0118, Xinghua Jiang, Mingming Gong, Haoyu Cao 0001, Yinsong Liu, Deqiang Jiang, Xing Sun 0001 |
CVPR | 6 |
| 2024 | HRVDA: High-Resolution Visual Document AssistantabstractLeveraging vast training data, multimodal large language models (MLLMs) have demonstrated formidable general visual comprehension capabilities and achieved remarkable performance across various tasks. However, their performance in visual document understanding still leaves much room for improvement. This discrepancy is primarily attributed to the fact that visual document understanding is a fine-grained prediction task. In natural scenes, MLLMs typically use low-resolution images, leading to a substantial loss of visual information. Furthermore, general-purpose MLLMs do not excel in handling document-oriented instructions. In this paper, we propose a High-Resolution Visual Document Assistant (HRVDA), which bridges the gap between MLLMs and visual document understanding. This model employs a content filtering mechanism and an instruction filtering module to separately filter out the content-agnostic visual tokens and instruction-agnostic visual tokens, thereby achieving efficient model training and inference for high-resolution images. In addition, we construct a document-oriented visual instruction tuning dataset and apply a multi-stage training strategy to enhance the model's document modeling capabilities. Extensive experiments demonstrate that our model achieves state-of-the-art performance across multiple document understanding datasets, while maintaining training efficiency and inference speed comparable to low-resolution models. Chaohu Liu, Kun Yin, Haoyu Cao 0001, Xinghua Jiang, Xin Li 0118, Yinsong Liu, Deqiang Jiang, Xing Sun 0001, Linli Xu 0002 |
CVPR | 3 |
| 2024 | Break the Visual Perception: Adversarial Attacks Targeting Encoded Visual Tokens of Large Vision-Language ModelsabstractLarge vision-language models (LVLMs) integrate visual information into large language models, showcasing remarkable multi-modal conversational capabilities. However, the visual modules introduces new challenges in terms of robustness for LVLMs, as attackers can craft adversarial images that are visually clean but may mislead the model to generate incorrect answers. In general, LVLMs rely on vision encoders to transform images into visual tokens, which are crucial for the language models to perceive image contents effectively. Therefore, we are curious about one question: Can LVLMs still generate correct responses when the encoded visual tokens are attacked and disrupting the visual information? To this end, we propose a non-targeted attack method referred to as VT-Attack (Visual Tokens Attack), which constructs adversarial examples from multiple perspectives, with the goal of comprehensively disrupting feature representations and inherent relationships as well as the semantic properties of visual tokens output by image encoders. Using only access to the image encoder in the proposed attack, the generated adversarial examples exhibit transferability across diverse LVLMs utilizing the same image encoder and generality across different tasks. Extensive experiments validate the superior attack performance of the VT-Attack over baseline methods, demonstrating its effectiveness in attacking LVLMs with image encoders, which in turn can provide guidance on the robustness of LVLMs, particularly in terms of the stability of the visual feature space. Yubo Wang 0010, Chaohu Liu, Yanqiu Qu, Haoyu Cao 0001, Deqiang Jiang, Linli Xu 0002 |
ACM Multimedia | 4 |
| 2024 | Communication-efficient clustered federated learning via model distance
Mao Zhang 0002, Yifei Cheng 0002, Changcun Bao, Haoyu Cao 0001, Deqiang Jiang, Linli Xu 0002 |
Mach. Learn. | 5 |
| 2024 | Turning a CLIP Model Into a Scene Text SpotterabstractWe exploit the potential of the large-scale Contrastive Language-Image Pretraining (CLIP) model to enhance scene text detection and spotting tasks, transforming it into a robust backbone, FastTCM-CR50. This backbone utilizes visual prompt learning and cross-attention in CLIP to extract image and text-based prior knowledge. Using predefined and learnable prompts, FastTCM-CR50 introduces an instance-language matching process to enhance the synergy between image and text embeddings, thereby refining text regions. Our Bimodal Similarity Matching (BSM) module facilitates dynamic language prompt generation, enabling offline computations and improving performance. FastTCM-CR50 offers several advantages: 1) It can enhance existing text detectors and spotters, improving performance by an average of 1.6% and 1.5%, respectively. 2) It outperforms the previous TCM-CR50 backbone, yielding an average improvement of 0.2% and 0.55% in text detection and spotting tasks, along with a 47.1% increase in inference speed. 3) It showcases robust few-shot training capabilities. Utilizing only 10% of the supervised data, FastTCM-CR50 improves performance by an average of 26.5% and 4.7% for text detection and spotting tasks, respectively. 4) It consistently enhances performance on out-of-distribution text detection and spotting datasets, particularly the NightTime-ArT subset from ICDAR2019-ArT and the DOTA dataset for oriented object detection. Wenwen Yu, Xingkui Zhu, Haoyu Cao 0001, Xing Sun 0001, Xiang Bai |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Attention Where It Matters: Rethinking Visual Document Understanding with Selective Region ConcentrationabstractWe propose a novel end-to-end document understanding model called SeRum (SElective Region Understanding Model) for extracting meaningful information from document images, including document analysis, retrieval, and office automation. Unlike state-of-the-art approaches that rely on multi-stage technical schemes and are computationally expensive, SeRum converts document image understanding and recognition tasks into a local decoding process of the visual tokens of interest, using a content-aware token merge module. This mechanism enables the model to pay more attention to regions of interest generated by the query decoder, improving the model’s effectiveness and speeding up the decoding speed of the generative scheme. We also designed several pre-training tasks to enhance the understanding and local awareness of the model. Experimental results demonstrate that SeRum achieves state-of-the-art performance on document understanding tasks and competitive results on text spotting tasks. SeRum represents a substantial advancement towards enabling efficient and effective end-to-end document understanding. Haoyu Cao 0001, Changcun Bao, Chaohu Liu, Kun Yin, Hao Liu 0003, Yinsong Liu, Deqiang Jiang, Xing Sun 0001 |
ICCV | 1 |
| 2023 | ICDAR 2023 Competition on Structured Text Extraction from Visually-Rich Document Images
Wenwen Yu, Chengquan Zhang, Haoyu Cao 0001, Wei Hua 0005, Bohan Li 0010, Mingrui Chen 0001, Jianfeng Kuang, Mengjun Cheng, Yuning Du, Shikun Feng, Xiaoguang Hu, Pengyuan Lv, Yuechen Yu, Wanxiang Che, Errui Ding, Cheng-Lin Liu 0001, Jiebo Luo 0001, Shuicheng Yan, Min Zhang 0005, Dimosthenis Karatzas, Xing Sun 0001, Jingdong Wang 0001, Xiang Bai |
ICDAR (2) | 3 |
| 2022 | Query-driven Generative Network for Document Information Extraction in the WildabstractThis paper focuses on solving Document Information Extraction (DIE) in the wild problem, which is rarely explored before. In contrast to existing studies mainly tailored for document cases in known templates with predefined layouts and keys under the ideal input without OCR errors involved, we aim to build up a more practical DIE paradigm for real-world scenarios where input document images may contain unknown layouts and keys in the scenes of the problematic OCR results. To achieve this goal, we propose a novel architecture, termed Query-driven Generative Network (QGN), which is equipped with two consecutive modules, i.e., Layout Context-aware Module (LCM) and Structured Generation Module (SGM). Given a document image with unseen layouts and fields, the former LCM yields the value prefix candidates serving as the query prompts for the SGM to generate the final key-value pairs even with OCR noise. To further investigate the potential of our method, we create a new large-scale dataset, named LArge-scale STructured Documents (LastDoc4000), containing 4,000 documents with 1,511 layouts and 3,500 different keys. In experiments, we demonstrate that our QGN consistently achieves the best F1-score on the new LastDoc4000 dataset by at most 30.32% absolute improvement. A more comprehensive experimental analysis and experiments on other public benchmarks also verify the effectiveness and robustness of our proposed method for the wild DIE task. Haoyu Cao 0001, Xin Li 0118, Jiefeng Ma, Deqiang Jiang, Antai Guo, Yiqing Hu, Hao Liu 0003, Yinsong Liu, Bo Ren 0002 |
ACM Multimedia | 1 |
| 2022 | Relational Representation Learning in Visually-Rich DocumentsabstractRelational understanding is critical for a number of visually-rich documents (VRDs) understanding tasks. Through multi-modal pre-training, recent studies provide comprehensive contextual representations and exploit them as prior knowledge for downstream tasks. In spite of their impressive results, we observe that the widespread relational hints (e.g., relation of key/value fields on receipts) built upon contextual knowledge are not excavated yet. To mitigate this gap, we propose DocReL, a Document Relational Representation Learning framework. The major challenge of DocReL roots in the variety of relations. From the simplest pairwise relation to the complex global structure, it is infeasible to conduct supervised training due to the definition of relation varies and even conflicts in different tasks. To deal with the unpredictable definition of relations, we propose a novel contrastive learning task named Relational Consistency Modeling (RCM), which harnesses the fact that existing relations should be consistent in differently augmented positive views. RCM provides relational representations which are more compatible to the urgent need of downstream tasks, even without any knowledge about the exact definition of relation. DocReL achieves better performance on a wide variety of VRD relational understanding tasks, including table structure recognition, key information extraction and reading order detection. Xin Li 0118, Yiqing Hu, Haoyu Cao 0001, Deqiang Jiang, Yinsong Liu, Bo Ren 0002 |
ACM Multimedia | 4 |
| 2022 | GMN: Generative Multi-modal Network for Practical Document Information ExtractionabstractHaoyu Cao, Jiefeng Ma, Antai Guo, Yiqing Hu, Hao Liu, Deqiang Jiang, Yinsong Liu, Bo Ren. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Haoyu Cao 0001, Jiefeng Ma, Antai Guo, Yiqing Hu, Hao Liu 0003, Deqiang Jiang, Yinsong Liu, Bo Ren 0002 |
NAACL-HLT | 1 |