EDBT 2026 Demo / reviewers in the wild / expert
Zichen Wen
dblp:366/3991
· DBLP profile ↗
16ranked-venue papers
4as first author
16since 2021 · last 2026
0009-0002-6157-5898ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 3 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | D²Pruner: Debiased Importance and Structural Diversity for MLLM Token PruningabstractProcessing long visual token sequences poses a significant computational burden on Multimodal Large Language Models (MLLMs). While token pruning offers a path to acceleration, we find that current methods, while adequate for general understanding, catastrophically fail on fine-grained localization tasks. We attribute this failure to the inherent flaws of the two prevailing strategies: importance-based methods suffer from a strong positional bias, an inherent model artifact that distracts from semantic content, while diversity-based methods exhibit structural blindness, disregarding the user's prompt and spatial redundancy. To address this, we introduce D²Pruner, a framework that rectifies these issues by uniquely combining debiased importance with a structural pruning mechanism. Our method first secures a core set of the most critical tokens as pivots based on a debiased attention score. It then performs a Maximal Independent Set (MIS) selection on the remaining tokens, which are modeled on a hybrid graph where edges signify spatial proximity and semantic similarity. This process iteratively preserves the most important and available token while removing its neighbors, ensuring that the supplementary tokens are chosen to maximize importance and diversity. Extensive experiments demonstrate that D²Pruner achieves exceptional efficiency and fidelity. Evelyn Zhang, Fufu Yu, Aoqi Wu, Zichen Wen, Shouhong Ding, Biqing Qi, Linfeng Zhang 0001 |
AAAI | 4 |
| 2026 | AgentSlimming: Towards Efficient and Cost-Aware Multi-Agent SystemsabstractLarge Language Model-based Multi-Agent Systems (MAS) have demonstrated remarkable capabilities in complex tasks.However, manually designing optimal communication topologies is labor-intensive, while automated expansion methods often result in bloated structures with redundant agents, leading to excessive token consumption.To address this problem, we introduce AgentSlimming, a plug-and-play compression framework for graph-structured multiagent workflows.Motivated by pruning and quantization in neural networks, AgentSlimming compresses workflows by first estimating the importance score of each agent with a hybrid mechanism, and then removes redundant agents or replaces them with low-cost ones, where each operation is validated using a baseline-anchored acceptance rule to prevent performance collapse.Experiments show that AgentSlimming reduces average token cost by up to 78.9% with negligible performance degradation, and sometimes even improves accuracy, achieving a strong Pareto-optimal tradeoff between cost and quality.Our code is publicly available at https://github.com/ CitrusYL/AgentSlimming. Yulang Chen, Haoxuan Peng, Zichen Wen, Dongrui Liu |
ACL (1) | 4 |
| 2026 | Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression MethodsabstractChenfei Liao, Wensong Wang, Zichen Wen, Xu Zheng, Yiyu Wang, Haocong He, Yuanhuiyi Lyu, Lutao Jiang, Xin Zou, Yuqian Fu, Bin Ren, Linfeng Zhang, Xuming Hu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Chenfei Liao, Wensong Wang, Zichen Wen, Xu Zheng 0002, Haocong He, Yuanhuiyi Lyu, Lutao Jiang, Xin Zou 0001, Yuqian Fu, Bin Ren 0005, Linfeng Zhang 0001, Xuming Hu |
ACL (1) | 3 |
| 2026 | Enhancing Multi-View Clustering: A Sufficient Information-Theoretic Approach for Consistency Acquisition and Redundancy EliminationabstractMulti-view clustering (MVC) has gained widespread recognition as a valuable technique for enhancing clustering performance by harnessing diverse data sources. Nonetheless, current methods mainly concentrate on obtaining consistent information, often ignoring the risk of redundant information across different views. In this study, we propose a novel methodology, called Sufficient Multi-View Clustering (STMVC), which evaluates the multi-view clustering framework through an information-theoretic lens, intending to learn inter-view consistency information while removing redundant information among views. Specifically, we first utilize variational analysis to extract inter-view consistency information, and to further enhance the consistency information and minimize the redundant information between different views, we propose a sufficient representation lower bound. Furthermore, in order to improve the adaptability and generalizability of our proposed approach, we expand the application of STMVC to single-view scenarios and incomplete multi-view scenarios. The STMVC method provides a promising solution to the challenge of multi-view clustering and introduces a fresh perspective for analyzing multi-view data. To validate our model, we conducted a theoretical analysis based on the Bayesian error rate, and experiments on several multi-view datasets and single-view datasets show the outstanding performance of STMVC. Yazhou Ren 0001, Zichen Wen, Junlong Ke, Chenhang Cui, Yonghao Huang, Xinyue Chen 0004, Philip S. Yu, Lifang He 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Data Whisperer: Efficient Data Selection for Task-Specific LLM Fine-Tuning via Few-Shot In-Context LearningabstractFine-tuning large language models (LLMs) on task-specific data is essential for their effective deployment. As dataset sizes grow, efficiently selecting optimal subsets for training becomes crucial to balancing performance and computational costs. Traditional data selection methods often require fine-tuning a scoring model on the target dataset, which is time-consuming and resource-intensive, or rely on heuristics that fail to fully leverage the model’s predictive capabilities. To address these challenges, we propose Data Whisperer, an efficient, training-free, attention-based method that leverages few-shot in-context learning with the model to be fine-tuned. Comprehensive evaluations were conducted on both raw and synthetic datasets across diverse tasks and models. Notably, Data Whisperer achieves superior performance compared to the full GSM8K dataset on the Llama-3-8B-Instruct model, using just 10% of the data, and outperforms existing methods with a 3.1-point improvement and a 7.4× speedup. Shaobo Wang 0001, Xiangqi Jin, Jize Wang, Zichen Wen, Conghui He, Xuming Hu, Linfeng Zhang 0001 |
ACL (1) | 7 |
| 2025 | Stop Looking for "Important Tokens" in Multimodal Language Models: Duplication Matters MoreabstractVision tokens in multimodal large language models often dominate huge computational overhead due to their excessive length compared to linguistic modality. Abundant recent methods aim to solve this problem with token pruning, which first defines an importance criterion for tokens and then prunes the unimportant vision tokens during inference. However, in this paper, we show that the importance is not an ideal indicator to decide whether a token should be pruned. Surprisingly, it usually results in inferior performance than random token pruning and leading to incompatibility to efficient attention computation operators. Instead, we propose DART (Duplication-Aware Reduction of Tokens), which prunes tokens based on its duplication with other tokens, leading to significant and training-free acceleration. Concretely, DART selects a small subset of pivot tokens and then retains the tokens with low duplication to the pivots, ensuring minimal information loss during token pruning. Experiments demonstrate that DART can prune 88.9% vision tokens while maintaining comparable performance, leading to a 1.99\times and 2.99\times speed-up in total time and prefilling stage, respectively, with good compatibility to efficient attention operators. Zichen Wen, Yifeng Gao 0003, Shaobo Wang 0001, Junyuan Zhang, Qintong Zhang, Conghui He, Linfeng Zhang 0001 |
EMNLP | 1 |
| 2025 | LEGION: Learning to Ground and Explain for Synthetic Image DetectionabstractThe rapid advancements in generative technology have emerged as a double-edged sword. While offering powerful tools that enhance convenience, they also pose significant social concerns. As defenders, current synthetic image detection methods often lack artifact-level textual interpretability and are overly focused on image manipulation detection, and current datasets usually suffer from outdated generators and a lack of fine-grained annotations. In this paper, we introduce SynthScars, a high-quality and diverse dataset consisting of 12,236 fully synthetic images with human-expert annotations. It features 4 distinct image content types, 3 categories of artifacts, and fine-grained annotations covering pixel-level segmentation, detailed textual explanations, and artifact category labels. Furthermore, we propose LEGION (LEarning to Ground and explain for Synthetic Image detectiON), a multimodal large language model (MLLM)-based image forgery analysis framework that integrates artifact detection, segmentation, and explanation. Building upon this capability, we further explore LEGION as a controller, integrating it into image refinement pipelines to guide the generation of higher-quality and more realistic images. Extensive experiments show that LEGION outperforms existing methods across multiple benchmarks, particularly surpassing the second-best traditional expert on SynthScars by 3.31% in mIoU and 7.75% in F1 score. Moreover, the refined images generated under its guidance exhibit stronger alignment with human preferences. The code, model, and dataset will be released. Hengrui Kang, Siwei Wen, Zichen Wen, Junyan Ye, Peilin Feng, Baichuan Zhou, Bin Wang 0065, Dahua Lin, Linfeng Zhang 0001, Conghui He |
ICCV | 3 |
| 2025 | OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented GenerationabstractRetrieval-augmented Generation (RAG) enhances Large Language Models (LLMs) by integrating external knowledge to reduce hallucinations and incorporate up-to-date information without retraining. As an essential part of RAG, external knowledge bases are commonly built by extracting structured data from unstructured PDF documents using Optical Character Recognition (OCR). However, given the imperfect prediction of OCR and the inherent non-uniform representation of structured data, knowledge bases inevitably contain various OCR noises. In this paper, we introduce OHRBench, the first benchmark for understanding the cascading impact of OCR on RAG systems. OHRBench includes 8,561 carefully selected unstructured document images from seven real-world RAG application domains, along with 8,498 Q&A pairs derived from multimodal elements in documents, challenging existing OCR solutions used for RAG. To better understand OCR's impact on RAG systems, we identify two primary types of OCR noise: Semantic Noise and Formatting Noise and apply perturbation to generate a set of structured data with varying degrees of each OCR noise. Using OHRBench, we first conduct a comprehensive evaluation of current OCR solutions and reveal that none is competent for constructing high-quality knowledge bases for RAG systems. We then systematically evaluate the impact of these two noise types and demonstrate the trend relationship between the degree of OCR noise and RAG performance. Our OHRBench, including PDF documents, Q&As, and the ground truth structured data are released at: https://github.com/opendatalab/OHR-Bench Junyuan Zhang, Qintong Zhang, Bin Wang 0065, Linke Ouyang, Zichen Wen, Ka-Ho Chow 0001, Conghui He, Wentao Zhang 0001 |
ICCV | 5 |
| 2025 | Multi-View Graph Clustering via Node-Guided Contrastive EncodingabstractMulti-view clustering has gained significant attention for integrating multi-view information in multimedia applications. With the growing complexity of graph data, multi-view graph clustering (MVGC) has become increasingly important. Existing methods primarily use Graph Neural Networks (GNNs) to encode structural and feature information, but applying GNNs within contrastive learning poses specific challenges, such as integrating graph data with node features and handling both homophilic and heterophilic graphs. To address these challenges, this paper introduces Node-Guided Contrastive Encoding (NGCE), a novel MVGC approach that leverages node features to guide embedding generation. NGCE enhances compatibility with GNN filtering, effectively integrates homophilic and heterophilic information, and strengthens contrastive learning across views. Extensive experiments demonstrate its robust performance on six homophilic and heterophilic multi-view benchmark datasets. Yazhou Ren 0001, Junlong Ke, Zichen Wen, Yang Yang 0002, Xiaorong Pu, Lifang He 0001 |
ICML | 3 |
| 2025 | MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?abstractWhile text-to-image models like GPT-4o-Image and FLUX are rapidly proliferating, they often encounter challenges such as hallucination, bias, and the production of unsafe, low-quality output. To effectively address these issues, it is crucial to align these models with desired behaviors based on feedback from a multimodal judge. Despite their significance, current multimodal judges frequently undergo inadequate evaluation of their capabilities and limitations, potentially leading to misalignment and unsafe fine-tuning outcomes. To address this issue, we introduce MJ-Bench, a novel benchmark which incorporates a comprehensive preference dataset to evaluate multimodal judges in providing feedback for image generation models across six key perspectives: alignment, safety, image quality, bias, composition, and visualization. Specifically, we evaluate a large variety of multimodal judges including smaller-sized CLIP-based scoring models, open-source VLMs, and close-source VLMs on each decomposed subcategory of our preference dataset. Experiments reveal that close-source VLMs generally provide better feedback, with GPT-4o outperforming other judges in average. Compared with open-source VLMs, smaller-sized scoring models can provide better feedback regarding text-image alignment and image quality, while VLMs provide more accurate feedback regarding safety and generation bias due to their stronger reasoning capabilities. Further studies in feedback scale reveal that VLM judges can generally provide more accurate and stable feedback in natural language than numerical scales. Notably, human evaluations on end-to-end and fine-tuned models using separate feedback from these multimodal judges provide similar conclusions, further confirming the effectiveness of MJ-Bench. Zhaorun Chen, Zichen Wen, Yichao Du, Yiyang Zhou, Chenhang Cui, Siwei Han, Jen Weng, Chaoqi Wang, Zhengwei Tong, Leria Huang, Canyu Chen, Haoqin Tu, Qinghao Ye, Zhihong Zhu 0001, Zhuokai Zhao, Rafael Rafailov, Chelsea Finn, Huaxiu Yao |
NeurIPS | 2 |
| 2025 | Efficient Multi-modal Large Language Models via Progressive Consistency DistillationabstractVisual tokens consume substantial computational resources in multi-modal large models (MLLMs), significantly compromising their efficiency. Recent works have attempted to improve efficiency by compressing visual tokens during training, either through modifications to model components or by introducing additional parameters. However, they often overlook the increased learning difficulty caused by such compression, as the model’s parameter space struggles to quickly adapt to the substantial perturbations in the feature space induced by token compression. In this work, we propose to develop Efficient MLLMs via Progressive Consistency Distillation (EPIC), a progressive learning framework. Specifically, by decomposing the feature space perturbations introduced by token compression along the token-wise and layer-wise dimensions, we introduce token consistency distillation and layer consistency distillation, respectively, aiming to reduce the training difficulty by leveraging guidance from a teacher model and following a progressive learning trajectory. Extensive experiments demonstrate the superior effectiveness, robustness, and generalization capabilities of our proposed framework. Zichen Wen, Shaobo Wang 0001, Yufa Zhou 0001, Junyuan Zhang, Qintong Zhang, Yifeng Gao 0003, Zhaorun Chen, Bin Wang 0065, Conghui He, Linfeng Zhang 0001 |
NeurIPS | 1 |
| 2025 | Spot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact ExplanationabstractWith the rapid advancement of Artificial Intelligence Generated Content (AIGC) technologies, synthetic images have become increasingly prevalent in everyday life, posing new challenges for authenticity assessment and detection. Despite the effectiveness of existing methods in evaluating image authenticity and locating forgeries, these approaches often lack human interpretability and do not fully address the growing complexity of synthetic data. To tackle these challenges, we introduce FakeVLM, a specialized large multimodal model designed for both general synthetic image and DeepFake detection tasks. FakeVLM not only excels in distinguishing real from fake images but also provides clear, natural language explanations for image artifacts, enhancing interpretability. Additionally, we present FakeClue, a comprehensive dataset containing over 100,000 images across seven categories, annotated with fine-grained artifact clues in natural language. FakeVLM demonstrates performance comparable to expert models while eliminating the need for additional classifiers, making it a robust solution for synthetic data detection. Extensive evaluations across multiple datasets confirm the superiority of FakeVLM in both authenticity classification and artifact explanation tasks, setting a new benchmark for synthetic image detection. The code, model weights, and dataset can be found here: https://github.com/opendatalab/FakeVLM. Siwei Wen, Junyan Ye, Peilin Feng, Hengrui Kang, Zichen Wen, Yize Chen, Jiang Wu 0003, Conghui He |
NeurIPS | 5 |
| 2025 | EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action ModelsabstractVision-Language-Action (VLA) models, particularly diffusion-based architectures, demonstrate transformative potential for embodied intelligence but are severely hampered by high computational and memory demands stemming from extensive inherent and inference-time redundancies. While existing acceleration efforts often target isolated inefficiencies, such piecemeal solutions typically fail to holistically address the varied computational and memory bottlenecks across the entire VLA pipeline, thereby limiting practical deployability. We introduce EfficientVLA, a structured and training-free inference acceleration framework that systematically eliminates these barriers by cohesively exploiting multifaceted redundancies. EfficientVLA synergistically integrates three targeted strategies: (1) pruning of functionally inconsequential layers from the language module, guided by an analysis of inter-layer redundancies; (2) optimizing the visual processing pathway through a task-aware strategy that selects a compact, diverse set of visual tokens, balancing task-criticality with informational coverage; and (3) alleviating temporal computational redundancy within the iterative diffusion-based action head by strategically caching and reusing key intermediate features.
We apply our method to a standard VLA model CogACT, yielding a $1.93\times$ inference speedup and reduces FLOPs to 28.9%, with only a 0.6% success rate drop in the SIMPLER benchmark. Yantai Yang, Zichen Wen, Luo Zhongwei, Chang Zou, Chuan Wen, Linfeng Zhang 0001 |
NeurIPS | 3 |
| 2024 | Homophily-Related: Adaptive Hybrid Graph Filter for Multi-View Graph ClusteringabstractRecently there is a growing focus on graph data, and multi-view graph clustering has become a popular area of research interest. Most of the existing methods are only applicable to homophilous graphs, yet the extensive real-world graph data can hardly fulfill the homophily assumption, where the connected nodes tend to belong to the same class. Several studies have pointed out that the poor performance on heterophilous graphs is actually due to the fact that conventional graph neural networks (GNNs), which are essentially low-pass filters, discard information other than the low-frequency information on the graph. Nevertheless, on certain graphs, particularly heterophilous ones, neglecting high-frequency information and focusing solely on low-frequency information impedes the learning of node representations. To break this limitation, our motivation is to perform graph filtering that is closely related to the homophily degree of the given graph, with the aim of fully leveraging both low-frequency and high-frequency signals to learn distinguishable node embedding. In this work, we propose Adaptive Hybrid Graph Filter for Multi-View Graph Clustering (AHGFC). Specifically, a graph joint process and graph joint aggregation matrix are first designed by using the intrinsic node features and adjacency relationship, which makes the low and high-frequency signals on the graph more distinguishable. Then we design an adaptive hybrid graph filter that is related to the homophily degree, which learns the node embedding based on the graph joint aggregation matrix. After that, the node embedding of each view is weighted and fused into a consensus embedding for the downstream task. Experimental results show that our proposed model performs well on six datasets containing homophilous and heterophilous graphs. Zichen Wen, Yawen Ling, Yazhou Ren 0001, Jianpeng Chen, Xiaorong Pu, Lifang He 0001 |
AAAI | 1 |
| 2024 | Integrating Vision-Language Semantic Graphs in Multi-View Clustering
Junlong Ke, Zichen Wen, Yechenhao Yang, Chenhang Cui, Yazhou Ren 0001, Xiaorong Pu, Lifang He 0001 |
IJCAI | 2 |
| 2024 | Dual-Optimized Adaptive Graph Reconstruction for Multi-View Graph ClusteringabstractMulti-view clustering is an important machine learning task for multi-media data, encompassing various domains such as images, videos, and texts. Moreover, with the growing abundance of graph data, the significance of multi-view graph clustering (MVGC) has become evident. Most existing methods focus on graph neural networks (GNNs) to extract information from both graph structure and feature data to learn distinguishable node representations. However, traditional GNNs are designed with the assumption of homophilous graphs, making them unsuitable for widely prevalent heterophilous graphs. Several techniques have been introduced to enhance GNNs for heterophilous graphs. While these methods partially mitigate the heterophilous graph issue, they often neglect the advantages of traditional GNNs, such as their simplicity, interpretability, and efficiency. In this paper, we propose a novel multi-view graph clustering method based on dual-optimized adaptive graph reconstruction, named DOAGC. It mainly aims to reconstruct the graph structure adapted to traditional GNNs to deal with heterophilous graph issues while maintaining the advantages of traditional GNNs. Specifically, we first develop an adaptive graph reconstruction mechanism that accounts for node correlation and original structural information. To further optimize the reconstruction graph, we design a dual optimization strategy and demonstrate the feasibility of our optimization strategy through mutual information theory. Numerous experiments demonstrate that DOAGC effectively mitigates the heterophilous graph problem. Zichen Wen, Yazhou Ren 0001, Yawen Ling, Chenhang Cui, Xiaorong Pu, Lifang He 0001 |
ACM Multimedia | 1 |