Yuming Yang 0001

dblp:222/1970-1 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0003-9518-7372ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Language models and text generation · 53% Reinforcement learning · 16% Efficient and distributed learning · 13%

Topics — the 18 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
agent benchmarking
1.012026
AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments · ACL (1) 2026
Natural language and speech › Language models and text generation
chain-of-thought reasoning
1.012026
Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Alignment · ACL (1) 2026
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
1.012026
Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Alignment · ACL (1) 2026
Natural language and speech › Language models and text generation
LLM agents
1.012026
AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments · ACL (1) 2026
Natural language and speech › Language models and text generation
alignment
0.912025
Pre-Trained Policy Discriminators are General Reward Models · NeurIPS 2025
Machine learning › Efficient and distributed learning › data selection
data selection for fine-tuning
0.912025
Measuring Data Diversity for Instruction Tuning: A Systematic Analysis and A Reliable Metric · ACL (1) 2025
Natural language and speech › Language models and text generation
hallucination mitigation
0.912025
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs · ICLR 2025
Natural language and speech › Language models and text generation
instruction tuning
0.912025
Measuring Data Diversity for Instruction Tuning: A Systematic Analysis and A Reliable Metric · ACL (1) 2025
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.912025
Pre-Trained Policy Discriminators are General Reward Models · NeurIPS 2025
Machine learning › Reinforcement learning › reward learning
reward modeling
0.912025
Pre-Trained Policy Discriminators are General Reward Models · NeurIPS 2025
Machine learning › Trustworthy machine learning
robustness
0.912025
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs · ICLR 2025
Natural language and speech › Language models and text generation › large language model › large language model adaptation
supervised fine-tuning
0.912025
Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levels · EMNLP 2025
Natural language and speech › Language models and text generation › alignment
sycophancy
0.912025
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs · ICLR 2025
Computer vision › Vision and language
vision-language model
0.912025
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs · ICLR 2025
Machine learning › Trustworthy machine learning › hallucination
vision-language model hallucination
0.912025
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs · ICLR 2025
Natural language and speech › Language models and text generation › large language model › knowledge in language models
knowledge retention
0.312025
Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levels · EMNLP 2025
Machine learning › Reinforcement learning
policy optimization
0.312025
Pre-Trained Policy Discriminators are General Reward Models · NeurIPS 2025
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement fine-tuning
0.312025
Pre-Trained Policy Discriminators are General Reward Models · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

alignment metric · 1.0token-level analysis · 0.9supervised fine-tuning · 0.9prompting · 0.9pre-training · 0.9policy discriminator · 0.9parameter-level analysis · 0.9novelty-based diversity metric · 0.9greedy data selection · 0.9direct preference optimization · 0.9
YearPublicationVenuePosition
2026 AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments
abstract
Zhiheng Xi, Dingwen Yang, Jiaqi Liu, Jixuan Huang, Honglin Guo, Baodai Huang, Tinggang Chen, Qi Zhang, Zhonghang Lu, Chenyu Liu, Jiajun Sun, Jiazheng Zhang, Dingwei Zhu, Xin Guo, Junzhe Wang, Zhihao Zhang, Yuming Yang, Junjie Ye, Minghe Gao, Dongrui Liu, Jiaming Ji, Guohao Li, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhiheng Xi, Dingwen Yang, Jixuan Huang, Honglin Guo, Baodai Huang, Tinggang Chen, Qi Zhang 0001, Zhonghang Lu, Jiazheng Zhang, Dingwei Zhu, Junzhe Wang 0001, Zhihao Zhang 0002, Yuming Yang 0001, Junjie Ye 0005, Minghe Gao, Dongrui Liu, Jiaming Ji, Tao Gui, Xuanjing Huang 0001
ACL (1)17
2026 Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Alignment
abstract
Yuming Yang, Mingyoung Lai, Wanxu Zhao, Xiaoran Fan, Zhiheng Xi, Mingqi Wu, Chiyue Huang, Jun Zhao, Haijun Lv, Jian Tong, Yunhua Zhou, Yicheng Zou, Qipeng Guo, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yuming Yang 0001, Mingyoung Lai, Wanxu Zhao, Xiaoran Fan, Zhiheng Xi, Mingqi Wu, Chiyue Huang, Jun Zhao 0019, Haijun Lv, Jian Tong, Yunhua Zhou, Yicheng Zou, Qipeng Guo, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ACL (1)1
2025 Measuring Data Diversity for Instruction Tuning: A Systematic Analysis and A Reliable Metric
abstract
Data diversity is crucial for the instruction tuning of large language models. Existing studies have explored various diversity-aware data selection methods to construct high-quality datasets and enhance model performance. However, the fundamental problem of precisely defining and measuring data diversity remains underexplored, limiting clear guidance for data engineering. To address this, we systematically analyze 11 existing diversity measurement methods by evaluating their correlation with model performance through extensive fine-tuning experiments. Our results indicate that a reliable diversity measure should properly account for both inter-sample differences and the information density in the sample space. Building on this, we propose NovelSum, a new diversity metric based on sample-level “novelty.” Experiments on both simulated and real-world data show that NovelSum accurately captures diversity variations and achieves a 0.97 correlation with instruction-tuned model performance, highlighting its value in guiding data engineering practices. With NovelSum as an optimization objective, we further develop a greedy, diversity-oriented data selection strategy that outperforms existing approaches, validating both the effectiveness and practical significance of our metric.
Yuming Yang 0001, Junjie Ye 0005, Shihan Dou, Xiao Wang 0042, Huijie Lv, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ACL (1)1
2025 Beyond Boundaries: Learning a Universal Entity Taxonomy across Datasets and Languages for Open Named Entity Recognition
abstract
Open Named Entity Recognition (NER), which involves identifying arbitrary types of entities from arbitrary domains, remains challenging for Large Language Models (LLMs). Recent studies suggest that fine-tuning LLMs on extensive NER data can boost their performance. However, training directly on existing datasets neglects their inconsistent entity definitions and redundant data, limiting LLMs to dataset-specific learning and hindering out-of-domain adaptation. To address this, we present B2NERD, a compact dataset designed to guide LLMs’ generalization in Open NER under a universal entity taxonomy. B2NERD is refined from 54 existing English and Chinese datasets using a two-step process. First, we detect inconsistent entity definitions across datasets and clarify them by distinguishable label names to construct a universal taxonomy of 400+ entity types. Second, we address redundancy using a data pruning strategy that selects fewer samples with greater category and semantic diversity. Comprehensive evaluation shows that B2NERD significantly enhances LLMs’ Open NER capabilities. Our B2NER models, trained on B2NERD, outperform GPT-4 by 6.8-12.0 F1 points and surpass previous methods in 3 out-of-domain benchmarks across 15 datasets and 6 languages. The data, models, and code are publicly available at https://github.com/UmeanNever/B2NER.
Yuming Yang 0001, Wantong Zhao, Caishuang Huang, Junjie Ye 0005, Xiao Wang 0042, Huiyuan Zheng, Xueying Xu, Kaixin Huang, Yunke Zhang, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
COLING1
2025 Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levels
abstract
Junjie Ye, Yuming Yang, Yang Nan, Shuo Li, Qi Zhang, Tao Gui, Xuanjing Huang, Peng Wang, Zhongchao Shi, Jianping Fan. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Junjie Ye 0005, Yuming Yang 0001, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001, Peng Wang 0095, Zhongchao Shi, Jianping Fan 0007
EMNLP2
2025 Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
abstract
In the study of LLMs, sycophancy represents a prevalent hallucination that poses significant challenges to these models. Specifically, LLMs often fail to adhere to original correct responses, instead blindly agreeing with users' opinions, even when those opinions are incorrect or malicious. However, research on sycophancy in visual language models (VLMs) has been scarce. In this work, we extend the exploration of sycophancy from LLMs to VLMs, introducing the MM-SY benchmark to evaluate this phenomenon. We present evaluation results from multiple representative models, addressing the gap in sycophancy research for VLMs. To mitigate sycophancy, we propose a synthetic dataset for training and employ methods based on prompts, supervised fine-tuning, and DPO. Our experiments demonstrate that these methods effectively alleviate sycophancy in VLMs. Additionally, we probe VLMs to assess the semantic impact of sycophancy and analyze the attention distribution of visual tokens. Our findings indicate that the ability to prevent sycophancy is predominantly observed in higher layers of the model. The lack of attention to image knowledge in these higher layers may contribute to sycophancy, and enhancing image attention at high layers proves beneficial in mitigating this issue.
Xiaoran Fan, Linsheng Lu, Leyi Yang, Yuming Yang 0001, Zhiheng Xi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ICLR6
2025 Pre-Trained Policy Discriminators are General Reward Models
abstract
We offer a novel perspective on reward modeling by formulating it as a policy discriminator, which quantifies the difference between two policies to generate a reward signal, guiding the training policy towards a target policy with desired behaviors. Based on this conceptual insight, we propose a scalable pre-training method named POLicy DiscriminAtive LeaRning (POLAR), which trains a reward model (RM) to discern identical policies and discriminate different ones. Unlike traditional reward modeling methods relying on absolute preferences, POLAR captures the relative difference between one policy and an arbitrary target policy, which is a scalable, high-level optimization objective suitable for modeling generic ranking relationships. Leveraging the POLAR pre-training paradigm, we present a series of RMs with parameter scales from 1.8B to 7B. Empirical results show that POLAR substantially outperforms traditional non-pre-trained methods, significantly enhancing RM performance. For instance, POLAR-7B could improve preference accuracy from 54.8% to 81.0% on STEM tasks and from 57.9% to 85.5% on creative writing tasks compared to SOTA baselines. POLAR also shows robust generalization capabilities in RLHF using Reinforcement Fine-tuning (RFT), providing reliable reward signals and markedly enhancing policy performance—improving LLaMa3.1-8B from an average of 47.36% to 56.33% and Qwen2.5-32B from 64.49% to 70.47% on 20 benchmarks. Moreover, scaling experiments reveal a clear power-law relationship between computation and performance, supported by linear correlation coefficients approaching 0.99. The impressive performance, strong generalization, and scaling properties suggest that POLAR is a promising direction for developing general and strong reward models.
Shihan Dou, Shichun Liu, Yuming Yang 0001, Yicheng Zou, Yunhua Zhou, Shuhao Xing, Chenhao Huang, Qiming Ge, Haijun Lv, Demin Song, Songyang Gao, Chengqi Lyu, Enyu Zhou, Honglin Guo, Zhiheng Xi, Qipeng Guo, Tao Gui, Qi Zhang 0001, Xipeng Qiu, Xuanjing Huang 0001, Kai Chen 0026
NeurIPS3
2019 Modeling and Experimental Design for MOOC Dropout Prediction: A Replication Perspective
Josh Gardner 0001, Yuming Yang 0001, Ryan Baker 0001, Christopher Brooks 0001
EDM2