EDBT 2026 Demo / reviewers in the wild / expert
Huijie Lv
dblp:371/4393
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 67% Reinforcement learning · 17% Efficient and distributed learning · 15% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › large language model evaluation
benchmark contamination |
1.0 | 1 | 2026 | Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination · AAAI 2026 |
Natural language and speech › Language models and text generation › large language model evaluation
data contamination |
1.0 | 1 | 2026 | Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination · AAAI 2026 |
Natural language and speech › Language models and text generation
mathematical reasoning |
1.0 | 1 | 2026 | Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination · AAAI 2026 |
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement learning for language models |
1.0 | 1 | 2026 | Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination · AAAI 2026 |
Machine learning › Efficient and distributed learning › data selection
data selection for fine-tuning |
0.9 | 1 | 2025 | Measuring Data Diversity for Instruction Tuning: A Systematic Analysis and A Reliable Metric · ACL (1) 2025 |
Natural language and speech › Language models and text generation
instruction tuning |
0.9 | 1 | 2025 | Measuring Data Diversity for Instruction Tuning: A Systematic Analysis and A Reliable Metric · ACL (1) 2025 |
Methods — techniques the papers use, named apart from their topics
synthetic benchmark generation · 1.0reinforcement learning · 1.0novelty-based diversity metric · 0.9greedy data selection · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data ContaminationabstractReasoning in large language models has long been a central research focus, and recent studies employing reinforcement learning (RL) have introduced diverse methods that yield substantial performance gains with minimal or even no external supervision. Surprisingly, some studies even suggest that random or incorrect reward signals can enhance performance. However, these breakthroughs are predominantly observed for the mathematically strong Qwen2.5 series on benchmarks such as MATH-500, AMC, and AIME, and seldom transfer to models like Llama, which warrants a more in-depth investigation. In this work, our empirical analysis reveals that pre-training on massive web-scale corpora leaves Qwen2.5 susceptible to data contamination in widely used benchmarks. Consequently, conclusions derived from contaminated benchmarks on Qwen2.5 series may be unreliable. To obtain trustworthy evaluation results, we introduce a generator that creates fully clean arithmetic problems of arbitrary length and difficulty, dubbed RandomCalculation. Using this leakage-free dataset, we show that only accurate reward signals yield steady improvements that surpass the base model’s performance boundary in mathematical reasoning, whereas random or incorrect rewards do not. Moreover, we conduct more fine-grained analyses to elucidate the factors underlying the different performance observed on the MATH-500 and RandomCalculation benchmarks. Consequently, we recommend that future studies evaluate models on uncontaminated benchmarks and, when feasible, test various model series to ensure trustworthy conclusions about RL and related methods. Mingqi Wu, Zhihao Zhang 0002, Qiaole Dong, Zhiheng Xi, Jun Zhao 0019, Senjie Jin, Xiaoran Fan, Yuhao Zhou 0005, Huijie Lv, Ming Zhang 0030, Yanwei Fu 0001, Qin Liu 0010, Songyang Zhang 0001, Qi Zhang 0001 |
AAAI | 9 |
| 2025 | Measuring Data Diversity for Instruction Tuning: A Systematic Analysis and A Reliable MetricabstractData diversity is crucial for the instruction tuning of large language models. Existing studies have explored various diversity-aware data selection methods to construct high-quality datasets and enhance model performance. However, the fundamental problem of precisely defining and measuring data diversity remains underexplored, limiting clear guidance for data engineering. To address this, we systematically analyze 11 existing diversity measurement methods by evaluating their correlation with model performance through extensive fine-tuning experiments. Our results indicate that a reliable diversity measure should properly account for both inter-sample differences and the information density in the sample space. Building on this, we propose NovelSum, a new diversity metric based on sample-level “novelty.” Experiments on both simulated and real-world data show that NovelSum accurately captures diversity variations and achieves a 0.97 correlation with instruction-tuned model performance, highlighting its value in guiding data engineering practices. With NovelSum as an optimization objective, we further develop a greedy, diversity-oriented data selection strategy that outperforms existing approaches, validating both the effectiveness and practical significance of our metric. Yuming Yang 0001, Junjie Ye 0005, Shihan Dou, Xiao Wang 0042, Huijie Lv, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 7 |