Chi Chen 0005

dblp:21/1794-5 · DBLP profile ↗
← Back
19ranked-venue papers
8as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 6 first-author · 15 since 2021Databases, data management, data science and information retrieval · 4 · 4 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021
YearPublicationVenuePosition
2026 LLaVA-UHD v2: Exploiting Hierarchical Vision Granularity in MLLMs via Inverse Semantic Pyramid
abstract
Vision transformers (ViTs) are widely employed in multimodal large language models (MLLMs) for visual encoding. However, they exhibit inferior performance on tasks regarding fine-grained visual perception. We attribute this to the inner limitations of ViTs in capturing diverse visual semantic levels. To address this, we present Hierarchical window (Hiwin) transformer as a plug-and-play solution for MLLMs, centered around our inverse semantic pyramid (ISP). Hiwin transformer comprises two key modules: (i) a visual detail injection module, which progressively injects low-level visual details into high-level language-aligned semantics features, thereby constructing an ISP, and (ii) a hierarchical window attention module, which leverages cross-scale windows to condense multi-level semantics from the ISP. Notably, our design achieves an average boost of 3.7% across 14 benchmarks compared with the baseline method, 9.3% on DocVQA for instance.
Zonghao Guo, Xuesong Yang, Chi Chen 0005, Yuan Yao 0013, Tat-Seng Chua, Maosong Sun 0001
AAAI7
2026 MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding
abstract
Fuwen Luo, Shengfeng Lou, Chi Chen, Ziyue Wang, Chenliang Li, Weizhou Shen, Jiyue Guo, Peng Li, Ming Yan, Ji Zhang, Fei Huang, Yang Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Fuwen Luo, Shengfeng Lou, Chi Chen 0005, Ziyue Wang 0002, Chenliang Li 0003, Weizhou Shen, Jiyue Guo, Peng Li 0030, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Yang Liu 0005
ACL (1)3
2026 RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic Evaluation
abstract
Bingxian Wu, Yu Zhang, Zonghao Guo, Tang Liu, Chen Qian, Yuxiang Lu, Xingbo Du, Yanghao Li, Yidan Zhang, Chi Chen, Ling Yao, Chenghu Zhou, Maosong Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Bingxian Wu, Yu Zhang 0186, Zonghao Guo, Xingbo Du, Yanghao Li, Chi Chen 0005, Ling Yao, Chenghu Zhou, Maosong Sun 0001
ACL (1)10
2025 ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models
abstract
Ziyue Wang, Chi Chen, Fuwen Luo, Yurui Dong, Yuanchi Zhang, Yuzhuang Xu, Xiaolong Wang, Peng Li, Yang Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Ziyue Wang 0002, Chi Chen 0005, Fuwen Luo, Yurui Dong 0001, Yuanchi Zhang, Yuzhuang Xu, Xiaolong Wang 0014, Peng Li 0030, Yang Liu 0005
ACL (1)2
2025 ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation
abstract
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in chart understanding tasks.However, interpreting charts with textual descriptions often leads to information loss, as it fails to fully capture the dense information embedded in charts.In contrast, parsing charts into code provides lossless representations that can effectively contain all critical details.Although existing open-source MLLMs have achieved success in chart understanding tasks, they still face two major challenges when applied to chart-to-code tasks: (1) Low executability and poor restoration of chart details in the generated code and (2) Lack of large-scale and diverse training data.To address these challenges, we propose ChartCoder, the first dedicated chart-to-code MLLM, which leverages Code LLMs as the language backbone to enhance the executability of the generated code.Furthermore, we introduce Chart2Code-160k, the first large-scale and diverse dataset for chartto-code generation, and propose the Snippetof-Thought (SoT) method, which transforms direct chart-to-code generation data into stepby-step generation.Experiments demonstrate that ChartCoder, with only 7B parameters, surpasses existing open-source MLLMs on chartto-code benchmarks, achieving superior chart restoration and code excitability.Our code is available at https://github.com/thunlp/ ChartCoder.89 seed code with 27 chart types Available functions and parameters
Xuanle Zhao, Xianzhen Luo, Qi Shi 0002, Chi Chen 0005, Shuo Wang 0013, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)4
2025 AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization
abstract
Recently, model merging methods have demonstrated powerful strengths in combining abilities on various tasks from multiple Large Language Models (LLMs). While previous model merging methods mainly focus on merging homogeneous models with identical architecture, they meet challenges when dealing with Multimodal Large Language Models (MLLMs) with inherent heterogeneous property, including differences in model architecture and the asymmetry in the parameter space. In this work, we propose AdaMMS1, a novel model merging method tailored for heterogeneous MLLMs. Our method tackles the challenges in three steps: mapping, merging and searching. Specifically, we first design mapping function between models to apply model merging on MLLMs with different architecture. Then we apply linear interpolation on model weights to actively adapt the asymmetry in the heterogeneous MLLMs. Finally in the hyper-parameter searching step, we propose an unsupervised hyper-parameter selection method for model merging. As the first model merging method capable of merging heterogeneous MLLMs without labeled data, extensive experiments on various model combinations demonstrated that AdaMMS outperforms previous model merging methods on various vision-language benchmarks.2
Yiyang Du, Xiaochen Wang 0002, Chi Chen 0005, Jiabo Ye, Peng Li 0030, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Zhifang Sui, Maosong Sun 0001, Yang Liu 0005
CVPR3
2025 Think in Safety: Unveiling and Mitigating Safety Alignment Collapse in Multimodal Large Reasoning Model
abstract
The rapid development of Multimodal Large Reasoning Models (MLRMs) has demonstrated broad application potential, yet their safety and reliability remain critical concerns that require systematic exploration.To address this gap, we conduct a comprehensive and systematic safety evaluation of 13 MLRMs across 5 benchmarks and unveil prevalent safety degradation phenomena in most advanced models.Moreover, our analysis reveals distinct safety patterns across different benchmarks: significant safety degradation is observed across jailbreak robustness benchmarks, whereas safetyawareness benchmarks demonstrate less pronounced degradation.In particular, the long thought process in some scenarios even enhances safety performance.Therefore, it is a potential approach to address safety issues in MLRMs by leveraging the intrinsic reasoning capabilities of the model to detect unsafe intent.To operationalize this insight, we construct a multimodal tuning dataset that incorporates a safety-oriented thought process.Experimental results from fine-tuning existing MLRMs with this dataset effectively enhance the safety on both jailbreak robustness and safety-awareness benchmarks.This study provides a new perspective for developing safe MLRMs. 1 Warning: this paper contains example data that may be offensive or harmful.
Xinyue Lou, You Li 0010, Jin An Xu, Chi Chen 0005
EMNLP5
2025 How Do Multimodal Large Language Models Handle Complex Multimodal Reasoning? Placing Them in an Extensible Escape Game
Ziyue Wang 0002, Yurui Dong 0001, Fuwen Luo, Minyuan Ruan, Zhili Cheng, Chi Chen 0005, Peng Li 0030, Yang Liu 0005
ICCV6
2025 CITR: Efficient Long Video Understanding Needs Causal Importance
abstract
Long video understanding is essential for various practical applications including surveillance and film analysis. While recent Vision-Language Models (VLMs) have advanced performance in this domain, efficiency remains a key challenge, especially for hour-long videos. Existing methods commonly reduce visual tokens via compression in the vision encoder, but token count still grows linearly with video length. Alternative approaches apply importance-based token reduction in the language model, yet their non-causal design limits efficiency gains to offline, single-query settings. In this work, we emphasize the need for causal importance estimation-where a token's relevance is determined only from prior context-to enable efficient, real-time long video understanding. We propose ØurMethod, a Causal Importance-based Token Reduction framework to reduce visual token redundancy in long video understanding tasks, enabling practical memory control and enhanced computational efficiency. Experiments on both offline and streaming benchmarks show that ØurMethod reduces latency by 49% in offline multi-query scenarios and effectively controls chunked prefilling time in streaming, all within a 24GB memory footprint and with less than 1% performance drop. The code and appendix are available at https://github.com/Columbine21/CITR.
Yanghao Li, Yuxiang Huang 0001, Chi Chen 0005, Shuo Wang 0013, Zhinan Gou
ACM Multimedia5
2024 Model Composition for Multimodal Large Language Models
abstract
Chi Chen, Yiyang Du, Zheng Fang, Ziyue Wang, Fuwen Luo, Peng Li, Ming Yan, Ji Zhang, Fei Huang, Maosong Sun, Yang Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Chi Chen 0005, Yiyang Du, Ziyue Wang 0002, Fuwen Luo, Peng Li 0030, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Maosong Sun 0001, Yang Liu 0005
ACL (1)1
2024 CODIS: Benchmarking Context-dependent Visual Comprehension for Multimodal Large Language Models
abstract
Fuwen Luo, Chi Chen, Zihao Wan, Zhaolu Kang, Qidong Yan, Yingjie Li, Xiaolong Wang, Siyu Wang, Ziyue Wang, Xiaoyue Mi, Peng Li, Ning Ma, Maosong Sun, Yang Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Fuwen Luo, Chi Chen 0005, Zihao Wan, Zhaolu Kang, Qidong Yan, Yingjie Li 0009, Xiaolong Wang 0014, Ziyue Wang 0002, Xiaoyue Mi, Peng Li 0030, Maosong Sun 0001, Yang Liu 0005
ACL (1)2
2024 Browse and Concentrate: Comprehending Multimodal Content via Prior-LLM Context Fusion
abstract
Ziyue Wang, Chi Chen, Yiqi Zhu, Fuwen Luo, Peng Li, Ming Yan, Ji Zhang, Fei Huang, Maosong Sun, Yang Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Ziyue Wang 0002, Chi Chen 0005, Yiqi Zhu, Fuwen Luo, Peng Li 0030, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Maosong Sun 0001, Yang Liu 0005
ACL (1)2
2023 Weakly Supervised Vision-and-Language Pre-training with Relative Representations
abstract
Weakly supervised vision-and-language pretraining (WVLP), which learns cross-modal representations with limited cross-modal supervision, has been shown to effectively reduce the data cost of pre-training while maintaining decent performance on downstream tasks.However, current WVLP methods use only local descriptions of images, i.e., object tags, as cross-modal anchors to construct weaklyaligned image-text pairs for pre-training.This affects the data quality and thus the effectiveness of pre-training.In this paper, we propose to directly take a small number of aligned image-text pairs as anchors, and represent each unaligned image and text by its similarities to these anchors, i.e., relative representations.We build a WVLP framework based on the relative representations, namely RELIT 1 , which collects high-quality weakly-aligned imagetext pairs from large-scale image-only and text-only data for pre-training through relative representation-based retrieval and generation.Experiments on four downstream tasks show that RELIT achieves new state-of-the-art results under the weakly supervised setting 2 .
Chi Chen 0005, Peng Li 0030, Maosong Sun 0001, Yang Liu 0005
ACL (1)1
2022 End-to-End Unsupervised Vision-and-Language Pre-training with Referring Expression Matching
abstract
Recently there has been an emerging interest in unsupervised vision-and-language pre-training (VLP) that learns multimodal representations without parallel image-caption data.These pioneering works significantly reduce the cost of VLP on data collection and achieve promising results compared to supervised VLP.However, existing unsupervised VLP methods take as input pre-extracted region-based visual features from external object detectors, which both limits flexibility and reduces computational efficiency.In this paper, we explore end-to-end unsupervised VLP with a vision encoder to directly encode images.The vision encoder is pre-trained on image-only data and jointly optimized during multimodal pre-training.To further enhance the learned cross-modal features, we propose a novel pre-training task that predicts which patches contain an object referred to in natural language from the encoded visual features.Extensive experiments on four visionand-language tasks show that our approach outperforms previous unsupervised VLP methods and obtains new state-of-the-art results 1 .
Chi Chen 0005, Peng Li 0030, Maosong Sun 0001, Yang Liu 0005
EMNLP1
2022 EXTR: Click-Through Rate Prediction with Externalities in E-Commerce Sponsored Search
abstract
Click-Through Rate (CTR) prediction, estimating the probability of a user clicking on items, plays a key fundamental role in sponsored search. E-commerce platforms display organic search results and advertisements (ads), collectively called items, together as a mixed list. The items displayed around the predicted ad, i.e. external items, may affect the user clicking on the predicted. Previous CTR models assume the user click only relies on the ad itself, which overlooks the effects of external items, referred to as external effects, or externalities. During the advertising prediction, the organic results have been generated by the organic system, while the final displayed ads on multiple ad slots have not been figured out, which leads to two challenges: 1) the predicted (target) ad may win any ad slot, bringing about diverse externalities. 2) external ads are undetermined, resulting in incomplete externalities. Facing the above challenges, inspired by the Transformer, we propose EXternality TRansformer (EXTR) which regards target ad with all slots as query and external items as key&value to model externalities in all exposure situations in parallel. Furthermore, we design a Potential Allocation Generator (PAG) for EXTR, to learn the allocation of potential external ads to complete the externalities. Extensive experimental results on Alibaba datasets demonstrate the effectiveness of externalities in the task of CTR prediction and illustrate that our proposed approach can bring significant profits to the real-world e-commerce platform. EXTR now has been successfully deployed in the online search advertising system in Alibaba, serving the main traffic.
Chi Chen 0005, Kangzhi Zhao, Junsheng Zhou, Hongbo Deng, Jian Xu 0015, Bo Zheng 0007, Yong Zhang 0002, Chunxiao Xing
KDD1
2021 Mask-Align: Self-Supervised Neural Word Alignment
abstract
Chi Chen, Maosong Sun, Yang Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Chi Chen 0005, Maosong Sun 0001, Yang Liu 0005
ACL/IJCNLP (1)1
2019 Investment Behaviors Can Tell What Inside: Exploring Stock Intrinsic Properties for Stock Trend Prediction
abstract
Stock trend prediction, aiming at predicting future price trend of stocks, plays a key role in seeking maximized profit from the stock investment. Recent years have witnessed increasing efforts in applying machine learning techniques, especially deep learning, to pursue more promising stock prediction. While deep learning has given rise to significant improvement, human investors still retain the leading position due to their understanding on stock intrinsic properties, which can imply invaluable principles for stock prediction. In this paper, we propose to extract and explore stock intrinsic properties to enhance stock trend prediction. Fortunately, we discover that the repositories of investment behaviors within mutual fund portfolio data form up a gold mine to extract latent representations of stock properties, since such collective investment behaviors can reflect the professional fund managers' common beliefs on stock intrinsic properties. Powered by extracted stock properties, we further propose to model the dynamic market state and trend using stock representations so as to generate the dynamic correlation between the stock and the market, and then we aggregate such correlation with dynamic stock indicators to achieve more accurate stock prediction. Extensive experiments on real-world stock market data demonstrate the effectiveness of stock properties extracted from collective investment behaviors in the task of stock prediction.
Chi Chen 0005, Li Zhao 0007, Jiang Bian 0002, Chunxiao Xing, Tie-Yan Liu
KDD1
2017 A System for Recognizing Entities and Extracting Relations from Electronic Medical Records
abstract
Digging rich knowledge from clinical texts becomes a popular topic today. Knowledge graph has been widely used to integrate and manage abundant knowledge. Entity recognition and relation extraction play important roles in constructing knowledge graphs. In this paper, we develop a system to recognize entities and extract their relations from clinical texts in Electronic Medical Records. Our system implements four major functions: manual entity annotation, automatic entity recognition, manual relation annotation and automatic relation extraction. Tools of entity annotation and relation annotation are designed for professionals to help them manually annotate objects given original clinical texts. Moreover, entity recognition and relation recognition, which CRF and CNN are applied in, are accessible for professionals before manual annotation in order to increase the efficiency. Our system has been used in several applications, such as medical knowledge graph construction and health QA system.
Chi Chen 0005, Chunxiao Xing
WISA1
2017 When Will a Repost Cascade Settle Down?
Chi Chen 0005, Hongliang Tian, Jie Tang 0001, Chunxiao Xing
WISE (1)1