VLDB 2026 Research / reviewers in the wild / expert
Chengqi Lyu
dblp:319/5244
· DBLP profile ↗
10ranked-venue papers
0as first author
10since 2021 · last 2025
0000-0002-4438-1314ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Language models and text generation · 40% Image recognition and object detection · 27% Reinforcement learning · 16% | |
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 100% |
Topics — the 23 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
hallucination |
1.6 | 2 | 2025 | Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs · ICLR 2025 ANAH: Analytical Annotation of Hallucinations in Large Language Models · ACL (1) 2024 |
Natural language and speech › Language models and text generation
hallucination detection |
1.5 | 2 | 2024 | ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language Models · NeurIPS 2024 ANAH: Analytical Annotation of Hallucinations in Large Language Models · ACL (1) 2024 |
Machine learning › Reinforcement learning › reward learning
reward modeling |
1.1 | 2 | 2025 | Pre-Trained Policy Discriminators are General Reward Models · NeurIPS 2025 CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward · EMNLP 2025 |
Natural language and speech › Language models and text generation
alignment |
0.9 | 1 | 2025 | Pre-Trained Policy Discriminators are General Reward Models · NeurIPS 2025 |
Natural language and speech › Language models and text generation › trustworthy language model › large language model reliability
factuality |
0.9 | 1 | 2025 | Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs · ICLR 2025 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.9 | 1 | 2025 | CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward · EMNLP 2025 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
0.9 | 1 | 2025 | Pre-Trained Policy Discriminators are General Reward Models · NeurIPS 2025 |
Natural language and speech › Language models and text generation
hallucination mitigation |
0.8 | 1 | 2024 | ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language Models · NeurIPS 2024 |
Program synthesis and code generation
code generation with language models |
0.8 | 1 | 2024 | AlchemistCoder: Harmonizing and Eliciting Code Capability by Hindsight Tuning on Multi-source Data · NeurIPS 2024 |
Computer vision › Image recognition and object detection › object detection
end-to-end object detection |
0.7 | 1 | 2023 | Dense Distinct Query for End-to-End Object Detection · CVPR 2023 |
Computer vision › Image recognition and object detection › object detection › detector training
label assignment |
0.7 | 1 | 2023 | Dense Distinct Query for End-to-End Object Detection · CVPR 2023 |
Computer vision › Image recognition and object detection
object detection |
0.7 | 1 | 2023 | Consistent-Teacher: Towards Reducing Inconsistent Pseudo-Targets in Semi-Supervised Object Detection · CVPR 2023 |
Machine learning › Trustworthy machine learning › learning from noisy data
pseudo-label noise |
0.7 | 1 | 2023 | Consistent-Teacher: Towards Reducing Inconsistent Pseudo-Targets in Semi-Supervised Object Detection · CVPR 2023 |
Computer vision › Image recognition and object detection › object detection
semi-supervised object detection |
0.7 | 1 | 2023 | Consistent-Teacher: Towards Reducing Inconsistent Pseudo-Targets in Semi-Supervised Object Detection · CVPR 2023 |
Computer vision › Image recognition and object detection › object detection › object detection evaluation
object detection benchmark |
0.6 | 1 | 2022 | MMRotate: A Rotated Object Detection Benchmark using PyTorch · ACM Multimedia 2022 |
Computer vision › Image recognition and object detection › object detection
oriented object detection |
0.6 | 1 | 2022 | MMRotate: A Rotated Object Detection Benchmark using PyTorch · ACM Multimedia 2022 |
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization |
0.3 | 1 | 2025 | Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs · ICLR 2025 |
Machine learning › Reinforcement learning
policy optimization |
0.3 | 1 | 2025 | Pre-Trained Policy Discriminators are General Reward Models · NeurIPS 2025 |
Natural language and speech › Language models and text generation
preference optimization |
0.3 | 1 | 2025 | Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs · ICLR 2025 |
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement fine-tuning |
0.3 | 1 | 2025 | Pre-Trained Policy Discriminators are General Reward Models · NeurIPS 2025 |
Natural language and speech › Information extraction and text analysis
data annotation |
0.2 | 1 | 2024 | ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language Models · NeurIPS 2024 |
Natural language and speech › Question answering and dialogue systems › answer generation
generative question answering |
0.2 | 1 | 2024 | ANAH: Analytical Annotation of Hallucinations in Large Language Models · ACL (1) 2024 |
Natural language and speech › Language models and text generation
instruction tuning |
0.2 | 1 | 2024 | AlchemistCoder: Harmonizing and Eliciting Code Capability by Hindsight Tuning on Multi-source Data · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
verifier training · 0.9sentence-level masking · 0.9scaling law analysis · 0.9reward modeling · 0.9pre-training · 0.9policy discriminator · 0.9direct preference optimization · 0.9instruction evolution · 0.8human-in-the-loop annotation · 0.8hindsight relabeling · 0.8generative annotator · 0.8discriminative annotator · 0.8data-specific prompts · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome RewardabstractShudong Liu, Hongwei Liu, Junnan Liu, Linchen Xiao, Songyang Gao, Chengqi Lyu, Yuzhe Gu, Wenwei Zhang, Derek F. Wong, Songyang Zhang, Kai Chen. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Shudong Liu 0007, Linchen Xiao, Songyang Gao, Chengqi Lyu, Yuzhe Gu, Derek F. Wong, Songyang Zhang 0001, Kai Chen 0026 |
EMNLP | 6 |
| 2025 | Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMsabstractLarge language models (LLMs) exhibit hallucinations (i.e., unfaithful or nonsensical information) when serving as AI assistants in various domains. Since hallucinations always come with truthful content in the LLM responses, previous factuality alignment methods that conduct response-level preference learning inevitably introduced noises during training. Therefore, this paper proposes a fine-grained factuality alignment method based on Direct Preference Optimization (DPO), called Mask-DPO. Incorporating sentence-level factuality as mask signals, Mask-DPO only learns from factually correct sentences in the preferred samples and prevents the penalty on factual contents in the not preferred samples, which resolves the ambiguity in the preference learning. Extensive experimental results demonstrate that Mask-DPO can significantly improve the factuality of LLMs responses to questions from both in-domain and out-of-domain datasets, although these questions and their corresponding topics are unseen during training. Only trained on the ANAH train set, the score of Llama3.1-8B-Instruct on the ANAH test set is improved from 49.19% to 77.53%, even surpassing the score of Llama3.1-70B-Instruct (53.44%), while its FactScore on the out-of-domain Biography dataset is also improved from 30.29% to 39.39%. We further study the generalization property of Mask-DPO using different training sample scaling strategies and find that scaling the number of topics in the dataset is more effective than the number of questions. We provide a hypothesis of what factual alignment is doing with LLMs, on the implication of this phenomenon, and conduct proof-of-concept experiments to verify it. We hope the method and the findings pave the way for future research on scaling factuality alignment. Yuzhe Gu, Chengqi Lyu, Dahua Lin, Kai Chen 0026 |
ICLR | 3 |
| 2025 | Pre-Trained Policy Discriminators are General Reward ModelsabstractWe offer a novel perspective on reward modeling by formulating it as a policy discriminator, which quantifies the difference between two policies to generate a reward signal, guiding the training policy towards a target policy with desired behaviors. Based on this conceptual insight, we propose a scalable pre-training method named POLicy DiscriminAtive LeaRning (POLAR), which trains a reward model (RM) to discern identical policies and discriminate different ones. Unlike traditional reward modeling methods relying on absolute preferences, POLAR captures the relative difference between one policy and an arbitrary target policy, which is a scalable, high-level optimization objective suitable for modeling generic ranking relationships. Leveraging the POLAR pre-training paradigm, we present a series of RMs with parameter scales from 1.8B to 7B. Empirical results show that POLAR substantially outperforms traditional non-pre-trained methods, significantly enhancing RM performance.
For instance, POLAR-7B could improve preference accuracy from 54.8% to 81.0% on STEM tasks and from 57.9% to 85.5% on creative writing tasks compared to SOTA baselines.
POLAR also shows robust generalization capabilities in RLHF using Reinforcement Fine-tuning (RFT), providing reliable reward signals and markedly enhancing policy performance—improving LLaMa3.1-8B from an average of 47.36% to 56.33% and Qwen2.5-32B from 64.49% to 70.47% on 20 benchmarks.
Moreover, scaling experiments reveal a clear power-law relationship between computation and performance, supported by linear correlation coefficients approaching 0.99.
The impressive performance, strong generalization, and scaling properties suggest that POLAR is a promising direction for developing general and strong reward models. Shihan Dou, Shichun Liu, Yuming Yang 0001, Yicheng Zou, Yunhua Zhou, Shuhao Xing, Chenhao Huang, Qiming Ge, Haijun Lv, Demin Song, Songyang Gao, Chengqi Lyu, Enyu Zhou, Honglin Guo, Zhiheng Xi, Qipeng Guo, Tao Gui, Qi Zhang 0001, Xipeng Qiu, Xuanjing Huang 0001, Kai Chen 0026 |
NeurIPS | 12 |
| 2024 | ANAH: Analytical Annotation of Hallucinations in Large Language ModelsabstractReducing the 'hallucination' problem of Large Language Models (LLMs) is crucial for their wide applications.A comprehensive and finegrained measurement of the hallucination is the first key step for the governance of this issue but is under-explored in the community.Thus, we present ANAH, a bilingual dataset that offers ANalytical Annotation of Hallucinations in LLMs within Generative Question Answering.Each answer sentence in our dataset undergoes rigorous annotation, involving the retrieval of a reference fragment, the judgment of the hallucination type, and the correction of hallucinated content.ANAH consists of ∼12k sentence-level annotations for ∼4.3kLLM responses covering over 700 topics, constructed by a human-in-the-loop pipeline.Thanks to the fine granularity of the hallucination annotations, we can quantitatively confirm that the hallucinations of LLMs progressively accumulate in the answer and use ANAH to train and evaluate hallucination annotators.We conduct extensive experiments on studying generative and discriminative annotators and show that, although current open-source LLMs have difficulties in fine-grained hallucination annotation, the generative annotator trained with ANAH can surpass all open-source LLMs and GPT-3.5, obtain performance competitive with GPT-4, and exhibits better generalization ability on unseen questions. 1 Ziwei Ji 0001, Yuzhe Gu, Chengqi Lyu, Dahua Lin, Kai Chen 0026 |
ACL (1) | 4 |
| 2024 | Fake Alignment: Are LLMs Really Aligned Well?abstractYixu Wang, Yan Teng, Kexin Huang, Chengqi Lyu, Songyang Zhang, Wenwei Zhang, Xingjun Ma, Yu-Gang Jiang, Yu Qiao, Yingchun Wang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Yixu Wang, Yan Teng 0002, Chengqi Lyu, Songyang Zhang 0001, Xingjun Ma, Yu-Gang Jiang 0001, Yu Qiao 0001, Yingchun Wang 0004 |
NAACL-HLT | 4 |
| 2024 | ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language ModelsabstractLarge language models (LLMs) exhibit hallucinations in long-form question-answering tasks across various domains and wide applications. Current hallucination detection and mitigation datasets are limited in domain and size, which struggle to scale due to prohibitive labor costs and insufficient reliability of existing hallucination annotators. To facilitate the scalable oversight of LLM hallucinations, this paper introduces an iterative self-training framework that simultaneously and progressively scales up the annotation dataset and improves the accuracy of the annotator. Based on the Expectation Maximization algorithm, in each iteration, the framework first applies an automatic hallucination annotation pipeline for a scaled dataset and then trains a more accurate annotator on the dataset. This new annotator is adopted in the annotation pipeline for the next iteration. Extensive experimental results demonstrate that the finally obtained hallucination annotator with only 7B parameters surpasses GPT-4 and obtains new state-of-the-art hallucination detection results on HaluEval and HalluQA by zero-shot inference. Such an annotator can not only evaluate the hallucination levels of various LLMs on the large-scale dataset but also help to mitigate the hallucination of LLMs generations, with the Natural Language Inference metric increasing from 25% to 37% on HaluEval. Yuzhe Gu, Ziwei Ji 0001, Chengqi Lyu, Dahua Lin, Kai Chen 0026 |
NeurIPS | 4 |
| 2024 | AlchemistCoder: Harmonizing and Eliciting Code Capability by Hindsight Tuning on Multi-source DataabstractOpen-source Large Language Models (LLMs) and their specialized variants, particularly Code LLMs, have recently delivered impressive performance. However, previous Code LLMs are typically fine-tuned on single-source data with limited quality and diversity, which may insufficiently elicit the potential of pre-trained Code LLMs. In this paper, we present AlchemistCoder, a series of Code LLMs with enhanced code generation and generalization capabilities fine-tuned on multi-source data. To achieve this, we pioneer to unveil inherent conflicts among the various styles and qualities in multi-source code corpora and introduce data-specific prompts with hindsight relabeling, termed AlchemistPrompts, to harmonize different data sources and instruction-response pairs. Additionally, we propose incorporating the data construction process into the fine-tuning data as code comprehension tasks, including instruction evolution, data filtering, and code review. Extensive experiments demonstrate that AlchemistCoder holds a clear lead among all models of the same size (6.7B/7B) and rivals or even surpasses larger models (15B/33B/70B), showcasing the efficacy of our method in refining instruction-following capabilities and advancing the boundaries of code intelligence. Source code and models are available at https://github.com/InternLM/AlchemistCoder. Zifan Song, Yudong Wang 0002, Kuikun Liu, Chengqi Lyu, Demin Song, Qipeng Guo, Hang Yan 0001, Dahua Lin, Kai Chen 0026, Cairong Zhao |
NeurIPS | 5 |
| 2023 | Consistent-Teacher: Towards Reducing Inconsistent Pseudo-Targets in Semi-Supervised Object DetectionabstractIn this study, we dive deep into the inconsistency of pseudo targets in semi-supervised object detection (SSOD). Our core observation is that the oscillating pseudo-targets undermine the training of an accurate detector. It injects noise into the student's training, leading to severe overfitting problems. Therefore, we propose a systematic solution, termed Consistent-Teacher, to reduce the inconsistency. First, adaptive anchor assignment (ASA) substitutes the static IoU-based strategy, which enables the student network to be resistant to noisy pseudo-bounding boxes. Then we calibrate the subtask predictions by designing a 3D feature alignment module (FAM-3D). It allows each classification feature to adaptively query the optimal feature vector for the regression task at arbitrary scales and locations. Lastly, a Gaussian Mixture Model (GMM) dynamically revises the score threshold of pseudo-bboxes, which stabilizes the number of ground truths at an early stage and remedies the unreliable supervision signal during training. Consistent-Teacher provides strong results on a large range of SSOD evaluations. It achieves 40.0 mAP with ResNet-50 backbone given only 10% of annotated MS-COCO data, which surpasses previous base-lines using pseudo labels by around 3 mAP. When trained on fully annotated MS-COCO with additional unlabeled data, the performance further increases to 47.7 mAP. Our code is available at https://github.com/Adamdad/ConsistentTeacher. Xinjiang Wang, Xingyi Yang, Yijiang Li, Litong Feng, Shijie Fang, Chengqi Lyu, Kai Chen 0002, Wayne Zhang 0001 |
CVPR | 7 |
| 2023 | Dense Distinct Query for End-to-End Object DetectionabstractOne-to-one label assignment in object detection has successfully obviated the need for non-maximum suppression (NMS) as postprocessing and makes the pipeline end-to-end. However, it triggers a new dilemma as the widely used sparse queries cannot guarantee a high recall, while dense queries inevitably bring more similar queries and encounter optimization difficulties. As both sparse and dense queries are problematic, then what are the expected queries in end-to-end object detection? This paper shows that the solution should be Dense Distinct Queries (DDQ). Concretely, we first lay dense queries like traditional detectors and then select distinct ones for one-to-one assignments. DDQ blends the advantages of traditional and recent end-to-end detectors and significantly improves the performance of various detectors including FCN, R-CNN, and DETRs. Most impressively, DDQ-DETR achieves 52.1 AP on MS-COCO dataset within 12 epochs using a ResNet-50 backbone, outperforming all existing detectors in the same setting. DDQ also shares the benefit of end-to-end detectors in crowded scenes and achieves 93.8 AP on Crowd-Human. We hope DDQ can inspire researchers to consider the complementarity between traditional methods and end-to-end detectors. The source code can be found at https://github.com/jshilong/DDQ. Xinjiang Wang, Jiaqi Wang 0003, Jiangmiao Pang, Chengqi Lyu, Ping Luo 0002, Kai Chen 0026 |
CVPR | 5 |
| 2022 | MMRotate: A Rotated Object Detection Benchmark using PyTorchabstractWe present an open-source toolbox, named MMRotate, which provides a coherent algorithm framework of training, inferring, and evaluation for the popular rotated object detection algorithm based on deep learning. MMRotate implements 18 state-of-the-art algorithms and supports the three most frequently used angle definition methods. To facilitate future research and industrial applications of rotated object detection-related problems, we also provide a large number of trained models and detailed benchmarks to give insights into the performance of rotated object detection. MMRotate is publicly released at https://github.com/open-mmlab/mmrotate. Yue Zhou 0005, Xue Yang 0005, Gefan Zhang, Yanyi Liu, Liping Hou, Xue Jiang 0001, Xingzhao Liu, Junchi Yan, Chengqi Lyu, Kai Chen 0026 |
ACM Multimedia | 10 |