VLDB 2026 Research / reviewers in the wild / expert
Chaohu Liu
dblp:356/2510
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0001-7588-4264ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multimodal Table Understanding with Difficulty-aware Reinforcement LearningabstractMultimodal table understanding, which aims for a comprehensive grasp of table content by integrating cellular text, tabular structure, and visual presentation, remains a core yet challenging area of research. We identify that the structural complexity of a table, quantifiable by intrinsic properties such as the ratio of merged cells and the total number of cells, presents a significant obstacle for existing models. Our empirical analysis reveals that the performance of leading Multimodal Large Language Models (MLLMs) deteriorates markedly as table complexity increases, exposing a critical vulnerability in their ability to perceive and reason over intricate tabular data. To address this challenge, we propose MM-Table-R1, a model enhanced through difficulty-aware reinforcement learning (RL) post-training strategy. Specifically, we introduce both task-level and data-level curriculum learning. The task-level curriculum is designed to establish a capability ladder, where the model first learns basic perceptual and semantic alignment of table data, and then progresses to acquiring multi-step reasoning capabilities. The data-level curriculum ensures that the model is not exposed to difficult samples prematurely, facilitating a more gradual and effective learning process. Furthermore, we invest considerable effort in constructing a high-quality, large-scale training corpus by curating and processing data from diverse open-source table datasets, ensuring that each instance is paired with an objectively verifiable reward signal. Demonstrating exceptional parameter efficiency, our 3B-parameter model sets a new benchmark by surpassing both established 3B and 7B models, including those specifically designed for table reasoning. Chaohu Liu, Haoyu Cao 0001, YongXiang Hua, Linli Xu 0002 |
AAAI | 1 |
| 2026 | Evaluating the Adversarial Robustness of Vision-Language Models via Internal Feature PerturbationsabstractVision-language models (VLMs), such as BLIP-2 and LLaVA, have significantly advanced multimodal understanding but exhibit critical vulnerabilities to visual adversarial perturbations. The high efficacy of untargeted attacks, in particular, poses significant concerns for their operational robustness. Conventional attack methods generate adversarial examples by backpropagating the language modeling loss from the final output to the input image. However, the deep architecture of the integrated large language models (LLMs) often diminishes this gradient flow, limiting the attack’s effectiveness in perturbing the visual domain. To address this limitation, we introduce a novel untargeted attack method based on Maximizing Information Entropy (MIE). Our approach enhances attack efficacy not only by maximizing the information entropy of the model’s final output but also by directly inducing uncertainty within its internal feature representations. This layer-wise perturbation strategy disrupts the model’s cognitive process more comprehensively than relying on the final output layer alone. We provide a theoretical analysis demonstrating that the uncertainty induced by MIE is greater than or equal to that of conventional output-only attacks. Comprehensive quantitative evaluations across multiple VLM architectures and datasets confirm that our method consistently outperforms existing techniques, thereby establishing a new, more rigorous benchmark for assessing adversarial robustness in vision-language models. Chaohu Liu, Yubo Wang 0010, Haoyu Cao 0001, Deqiang Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Tracking the Copyright of Large Vision-Language Models through Parameter Learning Adversarial ImagesabstractLarge vision-language models (LVLMs) have demonstrated remarkable image understanding and dialogue capabilities, allowing them to handle a variety of visual question answering tasks. However, their widespread availability raises concerns about unauthorized usage and copyright infringement, where users or individuals can develop their own LVLMs by fine-tuning published models. In this paper, we propose a novel method called Parameter Learning Attack (PLA) for tracking the copyright of LVLMs without modifying the original model. Specifically, we construct adversarial images through targeted attacks against the original model, enabling it to generate specific outputs. To ensure these attacks remain effective on potential fine-tuned models to trigger copyright tracking, we allow the original model to learn the trigger images by updating parameters in the opposite direction during the adversarial attack process. Notably, the proposed method can be applied after the release of the original model, thus not affecting the model’s performance and behavior. To simulate real-world applications, we fine-tune the original model using various strategies across diverse datasets, creating a range of models for copyright verification. Extensive experiments demonstrate that our method can more effectively identify the original copyright of fine-tuned models compared to baseline methods. Therefore, this work provides a powerful tool for tracking copyrights and detecting unlicensed usage of LVLMs. Yubo Wang 0010, Jianting Tang, Chaohu Liu, Linli Xu 0002 |
ICLR | 3 |
| 2025 | Input Domain Aware MoE: Decoupling Routing Decisions from Task Optimization in Mixture of ExpertsabstractSparse Mixture of Experts (sMoE) has become a pivotal approach for scaling large vision-language models, offering substantial capacity while maintaining computational efficiency through dynamic, sparse activation of experts. However, existing routing mechanisms, typically based on similarity scoring, struggle to effectively capture the underlying input structure. This limitation leads to a trade-off between expert specialization and balanced computation, hindering both scalability and performance. We propose Input Domain Aware MoE, a novel routing framework that leverages a probabilistic mixture model to better partition the input space. By modeling routing probabilities as a mixture of distributions, our method enables experts to develop clear specialization boundaries while achieving balanced utilization. Unlike conventional approaches, our routing mechanism is trained independently of task-specific objectives, allowing for stable optimization and decisive expert assignments. Empirical results on vision-language tasks demonstrate that our method consistently outperforms existing sMoE approaches, achieving higher task performance and improved expert utilization balance. YongXiang Hua, Haoyu Cao 0001, Zhou Tao, Bocheng Li, Chaohu Liu, Linli Xu 0002 |
ACM Multimedia | 6 |
| 2024 | HRVDA: High-Resolution Visual Document AssistantabstractLeveraging vast training data, multimodal large language models (MLLMs) have demonstrated formidable general visual comprehension capabilities and achieved remarkable performance across various tasks. However, their performance in visual document understanding still leaves much room for improvement. This discrepancy is primarily attributed to the fact that visual document understanding is a fine-grained prediction task. In natural scenes, MLLMs typically use low-resolution images, leading to a substantial loss of visual information. Furthermore, general-purpose MLLMs do not excel in handling document-oriented instructions. In this paper, we propose a High-Resolution Visual Document Assistant (HRVDA), which bridges the gap between MLLMs and visual document understanding. This model employs a content filtering mechanism and an instruction filtering module to separately filter out the content-agnostic visual tokens and instruction-agnostic visual tokens, thereby achieving efficient model training and inference for high-resolution images. In addition, we construct a document-oriented visual instruction tuning dataset and apply a multi-stage training strategy to enhance the model's document modeling capabilities. Extensive experiments demonstrate that our model achieves state-of-the-art performance across multiple document understanding datasets, while maintaining training efficiency and inference speed comparable to low-resolution models. Chaohu Liu, Kun Yin, Haoyu Cao 0001, Xinghua Jiang, Xin Li 0118, Yinsong Liu, Deqiang Jiang, Xing Sun 0001, Linli Xu 0002 |
CVPR | 1 |
| 2024 | Break the Visual Perception: Adversarial Attacks Targeting Encoded Visual Tokens of Large Vision-Language ModelsabstractLarge vision-language models (LVLMs) integrate visual information into large language models, showcasing remarkable multi-modal conversational capabilities. However, the visual modules introduces new challenges in terms of robustness for LVLMs, as attackers can craft adversarial images that are visually clean but may mislead the model to generate incorrect answers. In general, LVLMs rely on vision encoders to transform images into visual tokens, which are crucial for the language models to perceive image contents effectively. Therefore, we are curious about one question: Can LVLMs still generate correct responses when the encoded visual tokens are attacked and disrupting the visual information? To this end, we propose a non-targeted attack method referred to as VT-Attack (Visual Tokens Attack), which constructs adversarial examples from multiple perspectives, with the goal of comprehensively disrupting feature representations and inherent relationships as well as the semantic properties of visual tokens output by image encoders. Using only access to the image encoder in the proposed attack, the generated adversarial examples exhibit transferability across diverse LVLMs utilizing the same image encoder and generality across different tasks. Extensive experiments validate the superior attack performance of the VT-Attack over baseline methods, demonstrating its effectiveness in attacking LVLMs with image encoders, which in turn can provide guidance on the robustness of LVLMs, particularly in terms of the stability of the visual feature space. Yubo Wang 0010, Chaohu Liu, Yanqiu Qu, Haoyu Cao 0001, Deqiang Jiang, Linli Xu 0002 |
ACM Multimedia | 2 |
| 2023 | Attention Where It Matters: Rethinking Visual Document Understanding with Selective Region ConcentrationabstractWe propose a novel end-to-end document understanding model called SeRum (SElective Region Understanding Model) for extracting meaningful information from document images, including document analysis, retrieval, and office automation. Unlike state-of-the-art approaches that rely on multi-stage technical schemes and are computationally expensive, SeRum converts document image understanding and recognition tasks into a local decoding process of the visual tokens of interest, using a content-aware token merge module. This mechanism enables the model to pay more attention to regions of interest generated by the query decoder, improving the model’s effectiveness and speeding up the decoding speed of the generative scheme. We also designed several pre-training tasks to enhance the understanding and local awareness of the model. Experimental results demonstrate that SeRum achieves state-of-the-art performance on document understanding tasks and competitive results on text spotting tasks. SeRum represents a substantial advancement towards enabling efficient and effective end-to-end document understanding. Haoyu Cao 0001, Changcun Bao, Chaohu Liu, Kun Yin, Hao Liu 0003, Yinsong Liu, Deqiang Jiang, Xing Sun 0001 |
ICCV | 3 |