EDBT 2026 Demo / reviewers in the wild / expert
Lichang Chen
dblp:151/6212
· DBLP profile ↗
19ranked-venue papers
6as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 6 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General IntelligenceabstractAudio comprehension-including speech, non-speech sounds, and music-is essential for achieving human-level intelligence. Consequently, AI agents must demonstrate holistic audio understanding to qualify as generally intelligent. However, evaluating auditory intelligence comprehensively remains challenging. To address this gap, we introduce MMAU-Pro, the most comprehensive and rigorously curated benchmark for assessing audio intelligence in AI systems. MMAU-Pro contains 5,305 instances, where each instance has one or more audios paired with human expert-generated question-answer pairs, spanning speech, sound, music, and their combinations. Unlike existing benchmarks, MMAU-Pro evaluates auditory intelligence across 49 unique skills and multiple complex dimensions, including long-form audio comprehension, spatial audio reasoning, multi-audio understanding, among others. All questions are meticulously designed to require deliberate multi-hop reasoning, including both multiple-choice and open-ended response formats. Importantly, audio data is sourced directly ``from the wild" rather than from existing datasets with known distributions. We evaluate 22 leading open-source and proprietary multimodal AI models, revealing significant limitations: even state-of-the-art models such as Gemini 2.5 Flash and Audio Flamingo 3 achieve only 59.2% and 51.7% accuracy, respectively, approaching random performance in multiple categories. Our extensive analysis highlights specific shortcomings and provides novel insights, offering actionable perspectives for the community to enhance future AI systems' progression toward audio general intelligence. The benchmark and code is available at https://sonalkum.github.io/mmau-pro. Sonal Kumar, Simon Sedlácek, Vaibhavi Lokegaonkar, Fernando López, Wenyi Yu, Nishit Anand, Hyeonggon Ryu, Lichang Chen, Maxim Plicka, Miroslav Hlavácek, William Fineas Ellingwood, Sathvik Udupa, Siyuan Hou, Allison Ferner, Sara Barahona, Cecilia Bolaños, Satish Rahi, Laura Herrera-Alarcón, Satvik Dixit, Rupali S. Patil, Soham Deshmukh, Lasha Koroshinadze, L. Paola García-Perera, Eleni Zanou, Themos Stafylakis, Joon Son Chung, David F. Harwath, Dinesh Manocha, Alicia Lozano-Diez, Santosh Kesiraju, Sreyan Ghosh, Ramani Duraiswami |
AAAI | 8 |
| 2025 | Unveiling the Impact of Coding Data Instruction Fine-Tuning on Large Language Models ReasoningabstractInstruction Fine-Tuning (IFT) significantly enhances the zero-shot capabilities of pretrained Large Language Models (LLMs). While coding data is known to boost LLM reasoning abilities during pretraining, its role in activating internal reasoning capacities during IFT remains understudied. This paper investigates a key question: How does coding data impact LLMs' reasoning capacities during IFT stage? To explore this, we thoroughly examine the impact of coding data across different coding data proportions, model families, sizes, and reasoning domains, from various perspectives. Specifically, we create three IFT datasets with increasing coding data proportions, fine-tune six LLM backbones across different families and scales on these datasets, evaluate the tuned models' performance across twelve tasks in three reasoning domains, and analyze the outcomes from three broad-to-granular perspectives: overall, domain-level, and task-specific. Our holistic analysis provides valuable insights into each perspective. First, coding data tuning enhances the overall reasoning capabilities of LLMs across different model families and scales. Moreover, while the impact of coding data varies by domain, it shows consistent trends within each domain across different model families and scales. Additionally, coding data generally provides comparable task-specific benefits across model families, with optimal proportions in IFT datasets being task-dependent. Xinlu Zhang, Zhiyu Chen 0002, Xi Ye 0003, Xianjun Yang, Lichang Chen, William Yang Wang, Linda R. Petzold |
AAAI | 5 |
| 2025 | From Lists to Emojis: How Format Bias Affects Model AlignmentabstractIn this paper, we study format biases in reinforcement learning from human feedback (RLHF).We observe that many widely-used preference models-including human evaluators, GPT-4, and top-ranking models on the RewardBench benchmark-exhibit strong biases towards specific format patterns, such as lists, links, bold text, and emojis.Furthermore, large language models (LLMs) can exploit these biases to achieve higher rankings on popular benchmarks like AlpacaEval and LMSYS Chatbot Arena.One notable example is verbosity bias, where current preference models favor longer responses that appear more comprehensive, even when their quality is equal to or lower than shorter responses.However, format biases beyond verbosity remain largely underexplored.In this work, we extend the study of biases in preference learning beyond the commonly recognized length bias, offering a comprehensive analysis of a wider range of format biases.Additionally, we show that with a small amount of biased data (less than 1%), we can inject significant bias into the reward model.Moreover, these format biases can also be easily exploited by downstream alignment algorithms, such as best-of-n sampling and online iterative DPO, as it is usually easier to manipulate the format than to improve the quality of responses.Our findings emphasize the need to disentangle format and content both for designing alignment algorithms and evaluating models. Xuanchang Zhang, Wei Xiong 0015, Lichang Chen, Tianyi Zhou 0001, Heng Huang 0001, Tong Zhang 0001 |
ACL (1) | 3 |
| 2025 | RRM: Robust Reward Model Training Mitigates Reward HackingabstractReward models (RMs) play a pivotal role in aligning large language models (LLMs) with human preferences. However, traditional RM training, which relies on response pairs tied to specific prompts, struggles to disentangle prompt-driven preferences from prompt-independent artifacts, such as response length and format. In this work, we expose a fundamental limitation of current RM training methods, where RMs fail to effectively distinguish between contextual signals and irrelevant artifacts when determining preferences. To address this, we introduce a causal framework that learns preferences independent of these artifacts and propose a novel data augmentation technique designed to eliminate them. Extensive experiments show that our approach successfully filters out undesirable artifacts, yielding a more robust reward model (RRM). Our RRM improves the performance of a pairwise reward model trained on Gemma-2-9b-it, on Reward-Bench, increasing accuracy from 80.61% to 84.15%. Additionally, we train two DPO policies using both the RM and RRM, demonstrating that the RRM significantly enhances DPO-aligned policies, improving MT-Bench scores from 7.27 to 8.31 and length-controlled win-rates in AlpacaEval-2 from 33.46% to 52.49%. Tianqi Liu 0002, Wei Xiong 0015, Jie Ren 0006, Lichang Chen, Rishabh Joshi, Zhen Qin 0001, Tianhe Yu, Daniel Sohn, Anastasia Makarova, Jeremiah Z. Liu, Bilal Piot, Abraham Ittycheriah, Aviral Kumar, Mohammad Saleh |
ICLR | 4 |
| 2025 | OmnixR: Evaluating Omni-modality Language Models on Reasoning across ModalitiesabstractWe introduce \textbf{OmnixR}, an evaluation suite designed to benchmark state-of-the-art Omni-modality Language Models (OLMs), such as GPT-4o and Gemini.
Evaluating OLMs, which integrate multiple modalities such as text, vision, and audio, presents unique challenges.
Particularly, the user message might often consist of multiple modalities, such that OLMs have to establish holistic understanding and reasoning across modalities to accomplish the task.
Existing benchmarks are limited to single-modality or dual-modality tasks (e.g., image+text or video+text), overlooking comprehensive multi-modal assessments of model reasoning.
To address this, OmnixR offers two evaluation variants: (1) OmnixR-synth: a synthetic dataset generated automatically by translating text into multiple modalities—audio, images, video, and hybrids Omnify!. (2) OmnixR-real: a real-world dataset, manually curated and annotated by experts, for evaluating cross-modal reasoning in natural settings.
OmnixR presents a unique evaluation towards assessing OLMs over a diverse mix of modalities, such as a question that involves video, audio, and text, providing a rigorous cross-modal reasoning testbed than any existing benchmarks.
Our experiments find that all state-of-the-art OLMs struggles with OmnixR questions that require integrating information from multiple modalities to answer.
Further analysis highlight differences in reasoning behavior and underscoring the challenges of omni-modal AI alignment. Lichang Chen, Hexiang Hu, Yandong Li, Pranav Shyam, Tianyi Zhou 0001, Heng Huang 0001, Ming-Hsuan Yang 0001, Boqing Gong |
ICLR | 1 |
| 2025 | Creativity or Brute Force? Using Brainteasers as a Window into the Problem-Solving Abilities of Large Language ModelsabstractAccuracy remains a standard metric for evaluating AI systems, but it offers limited insight into how models arrive at their solutions.
In this work, we introduce a benchmark based on brainteasers written in long narrative form to probe more deeply into the types of reasoning strategies that models use. Brainteasers are well-suited for this goal because they can be solved with multiple approaches, such as a few-step solution that uses a creative insight or a longer solution that uses more brute force.
We investigate large language models (LLMs) across multiple layers of reasoning, focusing not only on correctness but also on the quality and creativity of their solutions.
We investigate many aspects of the reasoning process: (1) semantic parsing of the brainteasers into precise mathematical competition style formats; (2) self-correcting solutions based on gold solutions; (3) producing step-by-step sketches of solutions; and (4) making use of hints.
We find that LLMs are in many cases able to find creative, insightful solutions to brainteasers, suggesting that they capture some of the capacities needed to solve novel problems in creative ways. Nonetheless, there also remain situations where they rely on brute force despite the availability of more efficient, creative solutions, highlighting a potential direction for improvement in the reasoning abilities of LLMs. Simeng Han, Howard Dai, Stephen Xia, Grant Zhang, Chen Liu 0020, Lichang Chen, Hongyuan Mei, Jiayuan Mao, Tom McCoy 0001 |
NeurIPS | 6 |
| 2024 | Hallusionbench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language ModelsabstractWe introduce “HALLUSIONBENCH11“Hallusion” is a portmanteau of “hallucination” and “illusion.”,” a comprehensive benchmark designed for the evaluation of image-context rea-soning. This benchmark presents significant challenges to advanced large visual-language models (LVLMs), such as GPT-4V(ision), Gemini Pro Vision, Claude 3, and LLaVA-1.5, by emphasizing nuanced understanding and interpre-tation of visual data. The benchmark comprises 346 images paired with 1129 questions, all meticulously crafted by human experts. We introduce a novel structure for these visual questions designed to establish control groups. This structure enables us to conduct a quantitative analysis of the models' response tendencies, logical consistency, and various failure modes. In our evaluation on Hallusion-bench, we benchmarked 15 different models, highlighting a 31.42% question-pair accuracy achieved by the state-of-the-art GPT-4V. Notably, all other evaluated models achieve accuracy below 16%. Moreover, our analysis not only high-lights the observed failure modes, including language hal-lucination and visual illusion but also deepens an under-standing of these pitfalls. Our comprehensive case studies within Hallusionbench shed light on the challenges of hallucination and illusion in LVLMs. Based on these in-sights, we suggest potential pathways for their future im-provement. The benchmark and codebase can be accessed at https://github.com/tianyi-labIHallusionBench. Tianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian, Zongxia Li, Xiaoyu Liu 0003, Xijun Wang 0002, Lichang Chen, Furong Huang, Yaser Yacoob, Dinesh Manocha, Tianyi Zhou 0001 |
CVPR | 8 |
| 2024 | Spectrum AUC Difference (SAUCD): Human-Aligned 3D Shape EvaluationabstractExisting 3D mesh shape evaluation metrics mainly focus on the overall shape but are usually less sensitive to local details. This makes them inconsistent with human evaluation, as human perception cares about both overall and detailed shape. In this paper, we propose an analytic metric named Spectrum Area Under the Curve Difference (SAUCD) that demonstrates better consistency with human evaluation. To compare the difference between two shapes, we first transform the 3D mesh to the spectrum domain using the discrete Laplace-Beltrami operator and Fourier transform. Then, we calculate the Area Under the Curve (AUC) difference between the two spectrums, so that each frequency band that captures either the overall or detailed shape is equitably considered. Taking human sensitivity across frequency bands into account, we further extend our metric by learning suitable weights for each frequency band which better aligns with human perception. To measure the performance of SAUCD, we build a 3D mesh evaluation dataset called Shape Grading, along with manual annotations from more than 800 subjects. By measuring the correlation between our metric and human evaluation, we demonstrate that SAUCD is well aligned with human evaluation, and outperforms previous 3D mesh metrics. Our project page: https://bit.ly/saucd. Tianyu Luan, Zhong Li 0007, Lichang Chen, Yi Xu 0002, Junsong Yuan 0001 |
CVPR | 5 |
| 2024 | Prompting Language-Informed Distribution for Compositional Zero-Shot Learning
Wentao Bao, Lichang Chen, Heng Huang 0001, Yu Kong 0001 |
ECCV (14) | 2 |
| 2024 | AlpaGasus: Training a Better Alpaca with Fewer DataabstractLarge language models~(LLMs) strengthen instruction-following capability through instruction-finetuning (IFT) on supervised instruction/response data. However, widely used IFT datasets (e.g., Alpaca's 52k data) surprisingly contain many low-quality instances with incorrect or irrelevant responses, which are misleading and detrimental to IFT. In this paper, we propose a simple and effective data selection strategy that automatically identifies and removes low-quality data using a strong LLM (e.g., ChatGPT). To this end, we introduce Alpagasus, which is finetuned on only 9k high-quality data filtered from the 52k Alpaca data. Alpagasus significantly outperforms the original Alpaca as evaluated by GPT-4 on multiple test sets and the controlled human study. Its 13B variant matches $>90\%$ performance of its teacher LLM (i.e., Text-Davinci-003) on test tasks. It also provides 5.7x faster training, reducing the training time for a 7B variant from 80 minutes (for Alpaca) to 14 minutes \footnote{We apply IFT for the same number of epochs as Alpaca(7B) but on fewer data, using 4$\times$NVIDIA A100 (80GB) GPUs and following the original Alpaca setting and hyperparameters.}. In the experiment, we also demonstrate that our method can work not only for machine-generated datasets but also for human-written datasets. Overall, Alpagasus demonstrates a novel data-centric IFT paradigm that can be generally applied to instruction-tuning data, leading to faster training and better instruction-following models. Lichang Chen, Kalpa Gunaratna, Vikas Yadav, Vijay Srinivasan, Tianyi Zhou 0001, Heng Huang 0001, Hongxia Jin |
ICLR | 1 |
| 2024 | Unbiased Watermark for Large Language ModelsabstractThe recent advancements in large language models (LLMs) have sparked a growing apprehension regarding the potential misuse. One approach to mitigating this risk is to incorporate watermarking techniques into LLMs, allowing for the tracking and attribution of model outputs. This study examines a crucial aspect of watermarking: how significantly watermarks impact the quality of model-generated outputs. Previous studies have suggested a trade-off between watermark strength and output quality. However, our research demonstrates that it is possible to integrate watermarks without affecting the output probability distribution with appropriate implementation. We refer to this type of watermark as an unbiased watermark. This has significant implications for the use of LLMs, as it becomes impossible for users to discern whether a service provider has incorporated watermarks or not. Furthermore, the presence of watermarks does not compromise the performance of the model in downstream tasks, ensuring that the overall utility of the language model is preserved. Our findings contribute to the ongoing discussion around responsible AI development, suggesting that unbiased watermarks can serve as an effective means of tracking and attributing model outputs without sacrificing output quality. Zhengmian Hu, Lichang Chen, Xidong Wu, Hongyang Zhang 0001, Heng Huang 0001 |
ICLR | 2 |
| 2024 | ODIN: Disentangled Reward Mitigates Hacking in RLHFabstractIn this work, we study the issue of reward hacking on the response length, a challenge emerging in Reinforcement Learning from Human Feedback (RLHF) on LLMs. A well-formatted, verbose but less helpful response from the LLMs can often deceive LLMs or even human evaluators and achieve high scores. The same issue also holds for some reward models in RL. To address the challenges in both training and evaluation, we establish a more reliable evaluation protocol for comparing different training configurations, which inspects the trade-off between LLM evaluation score and response length obtained by varying training hyperparameters. Based on this evaluation, we conduct large-scale studies, where the results shed insights into the efficacy of hyperparameters and tricks used in RL on mitigating length bias. We further propose to improve the reward model by jointly training two linear heads to predict the preference, one trained to correlate with length and the other trained to decorrelate with length and therefore focusing more on the actual content. We then discard the length head in RL to ignore the spurious length reward. Experiments demonstrate that our approach eliminates the reward correlation with length, and improves the obtained policy by a significant margin. Lichang Chen, Chen Zhu 0001, Jiuhai Chen, Davit Soselia, Tianyi Zhou 0001, Tom Goldstein, Heng Huang 0001, Mohammad Shoeybi, Bryan Catanzaro |
ICML | 1 |
| 2024 | InstructZero: Efficient Instruction Optimization for Black-Box Large Language ModelsabstractLarge language models (LLMs) are instruction followers but the performance varies under different instructions. It is challenging to create the best instruction, especially for black-box LLMs on which backpropagation is forbidden. Instead of directly optimizing the discrete instruction, we optimize a low-dimensional soft prompt applied to an open-source LLM to generate the instruction for the black-box LLM. In each optimization step of the proposed method InstructZero, a soft prompt is converted into an instruction by the open-source LLM, which is then submitted to the black-box LLM for zero-shot evaluation, whose result is sent to Bayesian optimization to produce new soft prompts improving the zero-shot performance. We evaluate InstructZero on different combinations of open-source LLMs and APIs including Vicuna and ChatGPT. InstructZero outperforms SOTA auto-instruction methods across a variety of downstream tasks. Lichang Chen, Jiuhai Chen, Tom Goldstein, Heng Huang 0001, Tianyi Zhou 0001 |
ICML | 1 |
| 2024 | From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction TuningabstractMing Li, Yong Zhang, Zhitao Li, Jiuhai Chen, Lichang Chen, Ning Cheng, Jianzong Wang, Tianyi Zhou, Jing Xiao. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Ming Li 0010, Yong Zhang 0058, Zhitao Li 0002, Jiuhai Chen, Lichang Chen, Ning Cheng 0001, Jianzong Wang, Tianyi Zhou 0001, Jing Xiao 0006 |
NAACL-HLT | 5 |
| 2024 | Backdooring Instruction-Tuned Large Language Models with Virtual Prompt InjectionabstractJun Yan, Vikas Yadav, Shiyang Li, Lichang Chen, Zheng Tang, Hai Wang, Vijay Srinivasan, Xiang Ren, Hongxia Jin. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Vikas Yadav, Lichang Chen, Vijay Srinivasan, Hongxia Jin |
NAACL-HLT | 4 |
| 2024 | Task-Aware Sampling Layer for Point-Wise AnalysisabstractSampling, grouping, and aggregation are three important components in the multi-scale analysis of point clouds. In this paper, we present a novel data-driven sampler learning strategy for point-wise analysis tasks. Unlike the widely used sampling technique, Farthest Point Sampling (FPS), we propose to learn sampling and downstream applications jointly. Our key insight is that uniform sampling methods like FPS are not always optimal for different tasks: sampling more points around boundary areas can make the point-wise classification easier for segmentation. Towards this end, we propose a novel sampler learning strategy that learns sampling point displacement supervised by task-related ground truth information and can be trained jointly with the underlying tasks. We further demonstrate our methods in various point-wise analysis tasks, including semantic part segmentation, point cloud completion, and keypoint detection. Our experiments show that jointly learning of the sampler and task brings better performance than using FPS in various point-based networks. Yiqun Lin, Lichang Chen, Chongyang Ma, Xiaoguang Han 0001, Shuguang Cui |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | PTP: Boosting Stability and Performance of Prompt Tuning with Perturbation-Based RegularizerabstractRecent studies show that prompt tuning can better leverage the power of large language models than fine-tuning on downstream natural language understanding tasks.Nonetheless, current prompt tuning methods encounter instability during training, marked by a high variance in scores given different random seeds.In addressing this crucial issue, we uncover that the loss landscape of standard prompt tuning, when visualized, is remarkably steep, i.e., minor alterations in the input data can trigger substantial fluctuations in the loss landscape, which is an essential factor that leads to the training instability.In light of this finding, we incorporate perturbation-based regularizers to temper the loss landscape within the prompt tuning process.We thus present a novel algorithm, called Prompt Tuning with Perturbation-based regularizer (PTP), that can significantly reduce training instability and concurrently enhance the performance of prompt tuning.Specifically, we design two variants of perturbation-based regularizers: one that employs random noise, and another that uses an adversarial approach.Importantly, our proposed perturbations display flexibility in both the text and embedding spaces.Extensive experiments show the effectiveness of our proposed methods in stabilizing the training.Our new algorithms improve the state-of-the-art prompt tuning methods by 1.94% and 2.34% on SuperGLUE and FewGLUE benchmarks, respectively. Lichang Chen, Jiuhai Chen, Heng Huang 0001, Minhao Cheng |
EMNLP | 1 |
| 2020 | Graph Edit Distance Reward: Learning to Edit Scene Graph
Lichang Chen, Guosheng Lin, Qingyao Wu |
ECCV (19) | 1 |
| 2014 | Design and implementation of gaze tracking system with iPadabstractGaze tracking is the process of measuring gaze point or the motion of an eye relative to the head. Gaze tracking technique provides us a brand new way of human computer interaction. In addition, eye gaze tracking can be also applied to support seriously disabled people in using computer. In this paper, we explored the use of eye gaze tracking technology on a tablet device, designed and implemented an eye tracking system on an iPad device. In our system, gaze estimation is based on analyzing the appearances of eyes which are retrieved by the built-in camera on the iPad. Artificial neural networks are employed to estimate the location of the user gaze from the eye region image. The results indicate that it is possible to obtain an accuracy of 82.5% with the proposed system. Jiajin Zhang, Liu Di, Lichang Chen |
RTCSA | 3 |