Lei Zhang 0201

dblp:97/8704-201 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0003-0053-1840ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Language models and text generation · 52% Vision and language · 20% Question answering and dialogue systems · 10%
Software engineering, system software, and programming languages
3 papers
Program synthesis and code generation · 82% Software testing · 18%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 27 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
code language models
1.722025
CodeArena: Evaluating and Aligning CodeLLMs on Human Preference · EMNLP 2025
Hierarchical Context Pruning: Optimizing Real-World Code Completion with Repository-Level Pretrained Code LLMs · AAAI 2025
Program synthesis and code generation › code completion
fill-in-the-middle
1.012026
From Completion to Editing: Unlocking Context-Aware Code Infilling via Search-and-Replace Instruction Tuning · ACL (1) 2026
Natural language and speech › Language models and text generation
code generation
0.912025
CodeArena: Evaluating and Aligning CodeLLMs on Human Preference · EMNLP 2025
Computer vision › Vision and language
cross-modal alignment
0.912025
OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-time Emotional Speech Synthesis · NeurIPS 2025
Machine learning › Generative modeling
diffusion model
0.912025
DEEM: Diffusion models serve as the eyes of large language models for image perception · ICLR 2025
Natural language and speech › Speech recognition and synthesis › speech synthesis
emotional speech synthesis
0.912025
OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-time Emotional Speech Synthesis · NeurIPS 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
DEEM: Diffusion models serve as the eyes of large language models for image perception · ICLR 2025
Natural language and speech › Language models and text generation › multimodal language model
omni-modal large language model
0.912025
OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-time Emotional Speech Synthesis · NeurIPS 2025
Computer vision › Vision and language
vision-language model
0.912025
DEEM: Diffusion models serve as the eyes of large language models for image perception · ICLR 2025
Program synthesis and code generation
code completion
0.912025
Hierarchical Context Pruning: Optimizing Real-World Code Completion with Repository-Level Pretrained Code LLMs · AAAI 2025
Program synthesis and code generation › code completion
repository-level code completion
0.912025
Hierarchical Context Pruning: Optimizing Real-World Code Completion with Repository-Level Pretrained Code LLMs · AAAI 2025
Software testing
test-driven development
0.912025
Synthesizing Software Engineering Data in a Test-Driven Manner · ICML 2025
Natural language and speech › Language models and text generation › instruction tuning
instruction data selection
0.812024
One-Shot Learning as Instruction Data Prospector for Large Language Models · ACL (1) 2024
Natural language and speech › Language models and text generation
instruction tuning
0.812024
One-Shot Learning as Instruction Data Prospector for Large Language Models · ACL (1) 2024
Natural language and speech › Language models and text generation
long context
0.812024
Marathon: A Race Through the Realm of Long Context with Large Language Models · ACL (1) 2024
Natural language and speech › Language models and text generation › large language model evaluation › capability evaluation
long-context evaluation
0.812024
Marathon: A Race Through the Realm of Long Context with Large Language Models · ACL (1) 2024
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
long-context question answering
0.812024
Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA · EMNLP 2024
Natural language and speech › Question answering and dialogue systems › knowledge-intensive question answering
multi-document question answering
0.812024
Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA · EMNLP 2024
Machine learning › Representation and self-supervised learning
contrastive learning
0.612022
Image-Text Retrieval via Contrastive Learning with Auxiliary Generative Features and Support-set Regularization · SIGIR 2022
Multimedia analysis and retrieval
cross-modal retrieval
0.612022
Image-Text Retrieval via Contrastive Learning with Auxiliary Generative Features and Support-set Regularization · SIGIR 2022
Multimedia analysis and retrieval › cross-modal retrieval
image-text retrieval
0.612022
Image-Text Retrieval via Contrastive Learning with Auxiliary Generative Features and Support-set Regularization · SIGIR 2022
Natural language and speech › Language models and text generation
alignment
0.312025
CodeArena: Evaluating and Aligning CodeLLMs on Human Preference · EMNLP 2025
Natural language and speech › Language models and text generation
hallucination mitigation
0.312025
DEEM: Diffusion models serve as the eyes of large language models for image perception · ICLR 2025
Natural language and speech › Language models and text generation › alignment
preference alignment
0.312025
CodeArena: Evaluating and Aligning CodeLLMs on Human Preference · EMNLP 2025
Computer vision › Vision and language
vision-language pretraining
0.312025
OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-time Emotional Speech Synthesis · NeurIPS 2025
Program synthesis and code generation
code generation with language models
0.312025
Synthesizing Software Engineering Data in a Test-Driven Manner · ICML 2025
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization › long-context modeling
long-context language model
0.212024
Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA · EMNLP 2024

Methods — techniques the papers use, named apart from their topics

prompt construction · 1.7context pruning · 1.7search-and-replace instruction tuning · 1.0runtime dependency graph · 0.9reinforcement learning from human feedback · 0.9progressive multimodal alignment · 0.9generative feedback · 0.9fine-tuning · 0.9direct preference optimization · 0.9diffusion model · 0.9benchmark evaluation · 0.9CLIP-ViT · 0.9data valuation · 0.8support-set regularization · 0.6generative features · 0.6contrastive learning · 0.6
YearPublicationVenuePosition
2026 From Completion to Editing: Unlocking Context-Aware Code Infilling via Search-and-Replace Instruction Tuning
abstract
Jiajun Zhang, Zeyu Cui, Jiaxi Yang, Lei Zhang, Yuheng Jing, Zeyao Ma, Tianyi Bai, Zilei Wang, Qiang Liu, Liang Wang, Binyuan Hui, Junyang Lin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiajun Zhang 0012, Zeyu Cui, Jiaxi Yang 0004, Lei Zhang 0201, Yuheng Jing, Zeyao Ma, Tianyi Bai, Zilei Wang, Qiang Liu 0006, Liang Wang 0001, Binyuan Hui, Junyang Lin
ACL (1)4
2025 Hierarchical Context Pruning: Optimizing Real-World Code Completion with Repository-Level Pretrained Code LLMs
abstract
Some of the latest released Code Large Language Models (Code LLMs) have been trained on repository-level code data, enabling them to perceive repository structures and utilize cross-file code information. This capability allows us to directly concatenate the content of repository code files in prompts to achieve repository-level code completion. However, in real development scenarios, directly concatenating all code repository files in a prompt can easily exceed the context window of Code LLMs, leading to a significant decline in completion performance. Additionally, overly long prompts can increase completion latency, negatively impacting the user experience. In this study, we conducted extensive experiments, including completion error analysis, topology dependency analysis, and cross-file content analysis, to investigate the factors affecting repository-level code completion. Based on the conclusions drawn from these preliminary experiments, we proposed a strategy called **Hierarchical Context Pruning (HCP)** to construct high-quality completion prompts. We applied the **HCP** to six Code LLMs and evaluated them on the CrossCodeEval dataset. The experimental results showed that, compared to previous methods, the prompts constructed using our **HCP** strategy achieved higher completion accuracy on five out of six Code LLMs. Additionally, the **HCP** managed to keep the prompt length around 8k tokens (whereas the full repository code is approximately 50k tokens), significantly improving completion throughput. Our code and data will be publicly available.
Lei Zhang 0201, Yunshui Li, Jiaming Li 0004, Xiaobo Xia, Jiaxi Yang 0004, Run Luo, Minzheng Wang 0001, Longze Chen, Junhao Liu 0001, Qiang Qu 0001, Min Yang 0007
AAAI1
2025 CodeArena: Evaluating and Aligning CodeLLMs on Human Preference
abstract
Jian Yang, Jiaxi Yang, Wei Zhang, Jin Ke, Yibo Miao, Lei Zhang, Liqun Yang, Zeyu Cui, Yichang Zhang, Zhoujun Li, Binyuan Hui, Junyang Lin. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Jian Yang 0003, Jiaxi Yang 0004, Wei Zhang 0021, Yibo Miao, Lei Zhang 0201, Liqun Yang, Zeyu Cui, Yichang Zhang, Zhoujun Li 0001, Binyuan Hui, Junyang Lin
EMNLP6
2025 DEEM: Diffusion models serve as the eyes of large language models for image perception
abstract
The development of large language models (LLMs) has significantly advanced the emergence of large multimodal models (LMMs). While LMMs have achieved tremendous success by promoting the synergy between multimodal comprehension and creation, they often face challenges when confronted with out-of-distribution data, such as which can hardly distinguish orientation, quantity, color, structure, etc. This is primarily due to their reliance on image encoders trained to encode images into task-relevant features, which may lead them to disregard irrelevant details. Delving into the modeling capabilities of diffusion models for images naturally prompts the question: Can diffusion models serve as the eyes of large language models for image perception? In this paper, we propose DEEM, a simple but effective approach that utilizes the generative feedback of diffusion models to align the semantic distributions of the image encoder. This addresses the drawbacks of previous methods that solely relied on image encoders like CLIP-ViT, thereby enhancing the model's resilience against out-of-distribution samples and reducing visual hallucinations. Importantly, this is achieved without requiring additional training modules and with fewer training parameters. We extensively evaluated DEEM on both our newly constructed RobustVQA benchmark and other well-known benchmarks, POPE and MMVP, for visual hallucination and perception. In particular, DEEM improves LMM's visual perception performance to a large extent (e.g., 4\% ↑ on RobustVQA, 6.5\% ↑ on MMVP and 12.8 \% ↑ on POPE ). Compared to the state-of-the-art interleaved content generation models, DEEM exhibits enhanced robustness and a superior capacity to alleviate model hallucinations while utilizing fewer trainable parameters, less pre-training data (10\%), and a smaller base model size. Extensive experiments demonstrate that DEEM enhances the performance of LMMs on various downstream tasks without inferior performance in the long term, including visual question answering, image captioning, and text-conditioned image synthesis.
Run Luo, Yunshui Li, Longze Chen, Wanwei He, Ting-En Lin, Lei Zhang 0201, Zikai Song, Hamid Alinejad-Rokny, Xiaobo Xia, Tongliang Liu, Binyuan Hui, Min Yang 0007
ICLR7
2025 Synthesizing Software Engineering Data in a Test-Driven Manner
abstract
We introduce **SWE-Flow**, a novel data synthesis framework grounded in Test-Driven Development (TDD). Unlike existing software engineering data that rely on human-submitted issues, **SWE-Flow** automatically infers incremental development steps directly from unit tests, which inherently encapsulate high-level requirements. The core of **SWE-Flow** is the construction of a Runtime Dependency Graph (RDG), which precisely captures function interactions, enabling the generation of a structured, step-by-step *development schedule*. At each step, **SWE-Flow** produces a partial codebase, the corresponding unit tests, and the necessary code modifications, resulting in fully verifiable TDD tasks. With this approach, we generated 16,061 training instances and 2,020 test instances from real-world GitHub projects, creating the **SWE-Flow-Eval** benchmark. Our experiments show that fine-tuning open model on this dataset significantly improves performance in TDD-based coding. To facilitate further research, we release all code, datasets, models, and Docker images at [Github](https://github.com/Hambaobao/SWE-Flow).
Lei Zhang 0201, Jiaxi Yang 0004, Min Yang 0007, Jian Yang 0003, Mouxiang Chen, Jiajun Zhang 0012, Zeyu Cui, Binyuan Hui, Junyang Lin
ICML1
2025 OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-time Emotional Speech Synthesis
abstract
Recent advancements in omnimodal learning have significantly improved understanding and generation across images, text, and speech, yet these developments remain predominantly confined to proprietary models. The lack of high-quality omnimodal datasets and the challenges of real-time emotional speech synthesis have notably hindered progress in open-source research. To address these limitations, we introduce OpenOmni, a two-stage training framework that integrates omnimodal alignment and speech generation to develop a state-of-the-art omnimodal large language model. In the alignment phase, a pretrained speech model undergoes further training on image-text tasks, enabling (near) zero-shot generalization from vision to speech, outperforming models trained on tri-modal datasets. In the speech generation phase, a lightweight decoder is trained on speech tasks with direct preference optimization, which enables real-time emotional speech synthesis with high fidelity. Extensive experiments demonstrate that OpenOmni surpasses state-of-the-art models across omnimodal, vision-language, and speech-language benchmarks. It achieves a 4-point absolute improvement on OmniBench over the leading open-source model VITA, despite using 5$\times$ fewer training examples and a smaller model size (7B vs. 7$\times$8B). Besides, OpenOmni achieves real-time speech generation with less than 1 second latency at non-autoregressive mode, reducing inference time by 5$\times$ compared to autoregressive methods, and improves emotion classification accuracy by 7.7\%. The codebase is available at https://github.com/RainBowLuoCS/OpenOmni.
Run Luo, Ting-En Lin, Haonan Zhang 0003, Yuchuan Wu, Yongbin Li 0001, Longze Chen, Jiaming Li 0004, Lei Zhang 0201, Xiaobo Xia, Hamid Alinejad-Rokny, Fei Huang 0002, Min Yang 0007
NeurIPS9
2025 MSCFF-Net: multi-scale context feature fusion network for polyp segmentation
Lei Zhang 0201, Songlin Yin, Ge Zhang 0009
Multim. Syst.2
2024 One-Shot Learning as Instruction Data Prospector for Large Language Models
abstract
Yunshui Li, Binyuan Hui, Xiaobo Xia, Jiaxi Yang, Min Yang, Lei Zhang, Shuzheng Si, Ling-Hao Chen, Junhao Liu, Tongliang Liu, Fei Huang, Yongbin Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yunshui Li, Binyuan Hui, Xiaobo Xia, Jiaxi Yang 0004, Min Yang 0007, Lei Zhang 0201, Shuzheng Si, Junhao Liu 0001, Tongliang Liu, Fei Huang 0002, Yongbin Li 0001
ACL (1)6
2024 Marathon: A Race Through the Realm of Long Context with Large Language Models
abstract
Lei Zhang, Yunshui Li, Ziqiang Liu, Jiaxi Yang, Junhao Liu, Longze Chen, Run Luo, Min Yang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Lei Zhang 0201, Yunshui Li, Jiaxi Yang 0004, Junhao Liu 0001, Longze Chen, Run Luo, Min Yang 0007
ACL (1)1
2024 Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA
abstract
Minzheng Wang, Longze Chen, Fu Cheng, Shengyi Liao, Xinghua Zhang, Bingli Wu, Haiyang Yu, Nan Xu, Lei Zhang, Run Luo, Yunshui Li, Min Yang, Fei Huang, Yongbin Li. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Minzheng Wang 0001, Longze Chen, Fu Cheng, Shengyi Liao, Xinghua Zhang 0001, Bingli Wu, Haiyang Yu 0003, Nan Xu 0004, Lei Zhang 0201, Run Luo, Yunshui Li, Min Yang 0007, Fei Huang 0002, Yongbin Li 0001
EMNLP9
2023 Lifelong language learning with adaptive uncertainty regularization
Lei Zhang 0201, Fajie Yuan, Binzong Geng, Min Yang 0007
Inf. Sci.1
2022 Image-Text Retrieval via Contrastive Learning with Auxiliary Generative Features and Support-set Regularization
abstract
In this paper, we bridge the heterogeneity gap between different modalities and improve image-text retrieval by taking advantage of auxiliary image-to-text and text-to-image generative features with contrastive learning. Concretely, contrastive learning is devised to narrow the distance between the aligned image-text pairs and push apart the distance between the unaligned pairs from both inter- and intra-modality perspectives with the help of cross-modal retrieval features and auxiliary generative features. In addition, we devise a support-set regularization term to further improve contrastive learning by constraining the distance between each image/text and its corresponding cross-modal support-set information contained in the same semantic category. To evaluate the effectiveness of the proposed method, we conduct experiments on three benchmark datasets (i.e., MIRFLICKR-25K, NUS-WIDE, MS COCO). Experimental results show that our model significantly outperforms the strong baselines for cross-modal image-text retrieval. For reproducibility, we submit the code and data publicly at: \urlhttps://github.com/Hambaobao/CRCGS.
Lei Zhang 0201, Min Yang 0007, Chengming Li 0004, Ruifeng Xu 0001
SIGIR1