Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Longze Chen

dblp:378/2223 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
0009-0009-9177-8152ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Vision and language · 27% Language models and text generation · 26% Efficient and distributed learning · 20%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%
Databases, data mining, and information retrieval
1 paper
Web and social media mining · 100%

Topics — the 26 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
vision-language model
1.722025
VCM: Vision Concept Modeling with Adaptive Vision Token Compression via Instruction Fine-Tuning · NeurIPS 2025
DEEM: Diffusion models serve as the eyes of large language models for image perception · ICLR 2025
Computer vision › Vision and language
multimodal reasoning
1.012026
EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection · SIGIR 2026
Web and social media mining
content moderation
1.012026
EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection · SIGIR 2026
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization › long-context modeling
long-context language model
1.022024
Long Context is Not Long at All: A Prospector of Long-Dependency Data for Large Language Models · ACL (1) 2024
Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA · EMNLP 2024
Natural language and speech › Language models and text generation
code language models
0.912025
Hierarchical Context Pruning: Optimizing Real-World Code Completion with Repository-Level Pretrained Code LLMs · AAAI 2025
Computer vision › Vision and language
cross-modal alignment
0.912025
OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-time Emotional Speech Synthesis · NeurIPS 2025
Machine learning › Generative modeling
diffusion model
0.912025
DEEM: Diffusion models serve as the eyes of large language models for image perception · ICLR 2025
Natural language and speech › Speech recognition and synthesis › speech synthesis
emotional speech synthesis
0.912025
OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-time Emotional Speech Synthesis · NeurIPS 2025
Machine learning › Efficient and distributed learning
inference acceleration
0.912025
CLaSp: In-Context Layer Skip for Self-Speculative Decoding · ACL (1) 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
VCM: Vision Concept Modeling with Adaptive Vision Token Compression via Instruction Fine-Tuning · NeurIPS 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
DEEM: Diffusion models serve as the eyes of large language models for image perception · ICLR 2025
Natural language and speech › Language models and text generation › multimodal language model
omni-modal large language model
0.912025
OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-time Emotional Speech Synthesis · NeurIPS 2025
Machine learning › Efficient and distributed learning › inference acceleration
speculative decoding
0.912025
CLaSp: In-Context Layer Skip for Self-Speculative Decoding · ACL (1) 2025
Machine learning › Efficient and distributed learning › token reduction
vision token compression
0.912025
VCM: Vision Concept Modeling with Adaptive Vision Token Compression via Instruction Fine-Tuning · NeurIPS 2025
Computer vision › Image recognition and object detection › visual concept understanding
visual concept modeling
0.912025
VCM: Vision Concept Modeling with Adaptive Vision Token Compression via Instruction Fine-Tuning · NeurIPS 2025
Program synthesis and code generation
code completion
0.912025
Hierarchical Context Pruning: Optimizing Real-World Code Completion with Repository-Level Pretrained Code LLMs · AAAI 2025
Program synthesis and code generation › code completion
repository-level code completion
0.912025
Hierarchical Context Pruning: Optimizing Real-World Code Completion with Repository-Level Pretrained Code LLMs · AAAI 2025
Natural language and speech › Language models and text generation
long context
0.812024
Marathon: A Race Through the Realm of Long Context with Large Language Models · ACL (1) 2024
Natural language and speech › Language models and text generation › large language model evaluation › capability evaluation
long-context evaluation
0.812024
Marathon: A Race Through the Realm of Long Context with Large Language Models · ACL (1) 2024
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
long-context question answering
0.812024
Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA · EMNLP 2024
Natural language and speech › Question answering and dialogue systems › knowledge-intensive question answering
multi-document question answering
0.812024
Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA · EMNLP 2024
Computer vision › Vision and language › vision-language model
vision-language model evaluation
0.312026
EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection · SIGIR 2026
Natural language and speech › Language models and text generation
hallucination mitigation
0.312025
DEEM: Diffusion models serve as the eyes of large language models for image perception · ICLR 2025
Machine learning › Efficient and distributed learning
inference efficiency
0.312025
VCM: Vision Concept Modeling with Adaptive Vision Token Compression via Instruction Fine-Tuning · NeurIPS 2025
Natural language and speech › Language models and text generation
large language model
0.312025
CLaSp: In-Context Layer Skip for Self-Speculative Decoding · ACL (1) 2025
Computer vision › Vision and language
vision-language pretraining
0.312025
OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-time Emotional Speech Synthesis · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

multimodal reasoning · 2.0multi-agent decomposition · 2.0prompt construction · 1.7context pruning · 1.7layer skipping · 0.9in-context learning · 0.9generative feedback · 0.9direct preference optimization · 0.9diffusion model · 0.9CLIP-ViT · 0.9
YearPublicationVenuePosition
2026 EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection
abstract
E-commerce platforms increasingly rely on Large Language Models (LLMs) and Vision Language Models (VLMs) to detect illicit or misleading product content. However, these models remain vulnerable to evasive content, which refers to inputs that have been deliberately modified through techniques such as word splitting, euphemistic language, or image cropping to conceal policy violations while still conveying prohibited claims. Crucially, detecting such content requires a model to simultaneously master two capabilities: accurately comprehending complex rules, and correctly inferring the true intent behind deliberately obfuscated multimodal inputs. While prior work has separately explored LLM reasoning over complex rules and LLM-based detection of evasive content, no existing benchmark combines both within a unified evaluation framework. This gap is particularly consequential in e-commerce, where accurate moderation demands that both capabilities operate in concert. To address this gap, we introduce EVADE-Bench, the first expert-curated Chinese multimodal benchmark specifically designed to evaluate LLMs and VLMs on evasive content detection in real-world e-commerce scenarios. The dataset contains 2,833 annotated text samples and 13,961 annotated images spanning six violation categories. Our comprehensive evaluation of 26 open- and closed-source LLMs and VLMs reveals that even state-of-the-art models frequently misclassify evasive samples. We further demonstrate that clearer rule categorization significantly improves model prediction consistency and reduces false predictions, highlighting the critical role of benchmark design in enabling reliable evaluation. We analyze common error patterns across these models and identify systematic limitations in their ability to reason over metaphorical expressions and complex regulatory rules. To explore paths for performance improvement, we investigate the feasibility of multi-agent decomposition for multimodal reasoning, wherein visual description and logical inference are decoupled into separate agents, and find that this strategy yields notable accuracy gains. By releasing EVADE-Bench, we provide the first rigorous standard for evaluating evasive content detection and aim to support the development of safer and more trustworthy content moderation systems. The dataset is publicly available at https://huggingface.co/datasets/koenshen/EVADE-Bench.
Ancheng Xu, Guanghu Yuan, Longze Chen, Jiehui Zhou, Hengyu Chang, Hamid Alinejad-Rokny, Min Yang 0007
SIGIR5
2025 Hierarchical Context Pruning: Optimizing Real-World Code Completion with Repository-Level Pretrained Code LLMs
abstract
Some of the latest released Code Large Language Models (Code LLMs) have been trained on repository-level code data, enabling them to perceive repository structures and utilize cross-file code information. This capability allows us to directly concatenate the content of repository code files in prompts to achieve repository-level code completion. However, in real development scenarios, directly concatenating all code repository files in a prompt can easily exceed the context window of Code LLMs, leading to a significant decline in completion performance. Additionally, overly long prompts can increase completion latency, negatively impacting the user experience. In this study, we conducted extensive experiments, including completion error analysis, topology dependency analysis, and cross-file content analysis, to investigate the factors affecting repository-level code completion. Based on the conclusions drawn from these preliminary experiments, we proposed a strategy called **Hierarchical Context Pruning (HCP)** to construct high-quality completion prompts. We applied the **HCP** to six Code LLMs and evaluated them on the CrossCodeEval dataset. The experimental results showed that, compared to previous methods, the prompts constructed using our **HCP** strategy achieved higher completion accuracy on five out of six Code LLMs. Additionally, the **HCP** managed to keep the prompt length around 8k tokens (whereas the full repository code is approximately 50k tokens), significantly improving completion throughput. Our code and data will be publicly available.
Lei Zhang 0201, Yunshui Li, Jiaming Li 0004, Xiaobo Xia, Jiaxi Yang 0004, Run Luo, Minzheng Wang 0001, Longze Chen, Junhao Liu 0001, Qiang Qu 0001, Min Yang 0007
AAAI8
2025 CLaSp: In-Context Layer Skip for Self-Speculative Decoding
abstract
Longze Chen, Renke Shan, Huiming Wang, Lu Wang, Ziqiang Liu, Run Luo, Jiawei Wang, Hamid Alinejad-Rokny, Min Yang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Longze Chen, Renke Shan, Run Luo, Hamid Alinejad-Rokny, Min Yang 0007
ACL (1)1
2025 DEEM: Diffusion models serve as the eyes of large language models for image perception
abstract
The development of large language models (LLMs) has significantly advanced the emergence of large multimodal models (LMMs). While LMMs have achieved tremendous success by promoting the synergy between multimodal comprehension and creation, they often face challenges when confronted with out-of-distribution data, such as which can hardly distinguish orientation, quantity, color, structure, etc. This is primarily due to their reliance on image encoders trained to encode images into task-relevant features, which may lead them to disregard irrelevant details. Delving into the modeling capabilities of diffusion models for images naturally prompts the question: Can diffusion models serve as the eyes of large language models for image perception? In this paper, we propose DEEM, a simple but effective approach that utilizes the generative feedback of diffusion models to align the semantic distributions of the image encoder. This addresses the drawbacks of previous methods that solely relied on image encoders like CLIP-ViT, thereby enhancing the model's resilience against out-of-distribution samples and reducing visual hallucinations. Importantly, this is achieved without requiring additional training modules and with fewer training parameters. We extensively evaluated DEEM on both our newly constructed RobustVQA benchmark and other well-known benchmarks, POPE and MMVP, for visual hallucination and perception. In particular, DEEM improves LMM's visual perception performance to a large extent (e.g., 4\% ↑ on RobustVQA, 6.5\% ↑ on MMVP and 12.8 \% ↑ on POPE ). Compared to the state-of-the-art interleaved content generation models, DEEM exhibits enhanced robustness and a superior capacity to alleviate model hallucinations while utilizing fewer trainable parameters, less pre-training data (10\%), and a smaller base model size. Extensive experiments demonstrate that DEEM enhances the performance of LMMs on various downstream tasks without inferior performance in the long term, including visual question answering, image captioning, and text-conditioned image synthesis.
Run Luo, Yunshui Li, Longze Chen, Wanwei He, Ting-En Lin, Lei Zhang 0201, Zikai Song, Hamid Alinejad-Rokny, Xiaobo Xia, Tongliang Liu, Binyuan Hui, Min Yang 0007
ICLR3
2025 OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-time Emotional Speech Synthesis
abstract
Recent advancements in omnimodal learning have significantly improved understanding and generation across images, text, and speech, yet these developments remain predominantly confined to proprietary models. The lack of high-quality omnimodal datasets and the challenges of real-time emotional speech synthesis have notably hindered progress in open-source research. To address these limitations, we introduce OpenOmni, a two-stage training framework that integrates omnimodal alignment and speech generation to develop a state-of-the-art omnimodal large language model. In the alignment phase, a pretrained speech model undergoes further training on image-text tasks, enabling (near) zero-shot generalization from vision to speech, outperforming models trained on tri-modal datasets. In the speech generation phase, a lightweight decoder is trained on speech tasks with direct preference optimization, which enables real-time emotional speech synthesis with high fidelity. Extensive experiments demonstrate that OpenOmni surpasses state-of-the-art models across omnimodal, vision-language, and speech-language benchmarks. It achieves a 4-point absolute improvement on OmniBench over the leading open-source model VITA, despite using 5$\times$ fewer training examples and a smaller model size (7B vs. 7$\times$8B). Besides, OpenOmni achieves real-time speech generation with less than 1 second latency at non-autoregressive mode, reducing inference time by 5$\times$ compared to autoregressive methods, and improves emotion classification accuracy by 7.7\%. The codebase is available at https://github.com/RainBowLuoCS/OpenOmni.
Run Luo, Ting-En Lin, Haonan Zhang 0003, Yuchuan Wu, Yongbin Li 0001, Longze Chen, Jiaming Li 0004, Lei Zhang 0201, Xiaobo Xia, Hamid Alinejad-Rokny, Fei Huang 0002, Min Yang 0007
NeurIPS7
2025 VCM: Vision Concept Modeling with Adaptive Vision Token Compression via Instruction Fine-Tuning
abstract
Large vision-language models (LVLMs) have emerged as foundational tools for real-world AI applications. Despite their remarkable capabilities, current LVLMs process entire images at the token level, leading to significant inefficiencies compared to human cognition, which selectively focuses on high-level vision concepts. This token-level redundancy becomes increasingly problematic for high-resolution images and long video sequences, resulting in large computational costs and limited scalability in practical applications. To address this limitation, we introduce the concept of a vision concept model, a novel paradigm that enables LVLMs to dynamically extract the most relevant vision concepts from complex inputs, based on task-specific instructions. To optimize this vision concept modeling process, we propose VCM, a self-supervised framework that leverages vision-language correlations across diverse instances. VCM is designed to learn meaningful vision concepts without the need for expensive concept-level annotations. At its core, it employs a forward-backward optimization algorithm that supports LVLMs to adjust concept granularity and spatial alignment dynamically. Experiments demonstrate that VCM remarkably reduces computational costs (e.g., achieving up to 85\% fewer FLOPs for LLaVA-1.5-7B), while maintaining strong performance across a series of vision-language tasks. The codebase is available at https://github.com/RainBowLuoCS/VCM.
Run Luo, Renke Shan, Longze Chen, Min Yang 0007, Xiaobo Xia
NeurIPS3
2024 Long Context is Not Long at All: A Prospector of Long-Dependency Data for Large Language Models
abstract
Longze Chen, Ziqiang Liu, Wanwei He, Yinhe Zheng, Hao Sun, Yunshui Li, Run Luo, Min Yang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Longze Chen, Wanwei He, Yinhe Zheng, Yunshui Li, Run Luo, Min Yang 0007
ACL (1)1
2024 Marathon: A Race Through the Realm of Long Context with Large Language Models
abstract
Lei Zhang, Yunshui Li, Ziqiang Liu, Jiaxi Yang, Junhao Liu, Longze Chen, Run Luo, Min Yang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Lei Zhang 0201, Yunshui Li, Jiaxi Yang 0004, Junhao Liu 0001, Longze Chen, Run Luo, Min Yang 0007
ACL (1)6
2024 Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA
abstract
Minzheng Wang, Longze Chen, Fu Cheng, Shengyi Liao, Xinghua Zhang, Bingli Wu, Haiyang Yu, Nan Xu, Lei Zhang, Run Luo, Yunshui Li, Min Yang, Fei Huang, Yongbin Li. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Minzheng Wang 0001, Longze Chen, Fu Cheng, Shengyi Liao, Xinghua Zhang 0001, Bingli Wu, Haiyang Yu 0003, Nan Xu 0004, Lei Zhang 0201, Run Luo, Yunshui Li, Min Yang 0007, Fei Huang 0002, Yongbin Li 0001
EMNLP2