Junjie Ye 0005

dblp:19/8588-5 · DBLP profile ↗
← Back
20ranked-venue papers
6as first author
20since 2021 · last 2026
0009-0004-0921-6323ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 6 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
YearPublicationVenuePosition
2026 What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study
abstract
Speech-language models (SLMs) offer a promising path toward unifying speech and text understanding and generation. However, challenges remain in achieving effective cross-modal alignment and high-quality speech generation. In this work, we systematically investigate the role of speech tokenizer designs in LLM-centric SLMs, augmented by speech heads and speaker modeling. We compare coupled, semi-decoupled, and fully decoupled speech tokenizers under a fair SLM framework and find that decoupled tokenization significantly improves alignment and synthesis quality. To address the information density mismatch between speech and text, we introduce multi-token prediction (MTP) into SLMs, enabling each hidden state to decode multiple speech tokens. This leads to up to 12× faster decoding and a substantial drop in word error rate (from 6.07 to 3.01). Furthermore, we propose a speaker-aware generation paradigm and introduce RoleTriviaQA, a large-scale role-playing knowledge QA benchmark with diverse speaker identities. Experiments demonstrate that our methods enhance both knowledge understanding and speaker consistency.
Xiaoran Fan, Yangfan Gao, Jingfei Xiong, Hang Yan 0001, Yifei Cao, Zhihao Zhang 0002, Zhiheng Xi, Yuhao Zhou 0005, Senjie Jin, Changhao Jiang, Junjie Ye 0005, Ming Zhang 0030, Zhenhua Han, Yunke Zhang, Demei Yan, Shaokang Dong, Tao Gui
AAAI14
2026 MetaAct-RL: Training Language Models for Reasoning Through Meta-Action-Based Reinforcement Learning
abstract
Outcome-based reinforcement learning has made notable advances in training language models (LMs) for reasoning. However, without explicit incentives and controls, this paradigm has limitations and instability in eliciting high-quality reasoning trajectories with diverse actions—particularly for models whose pretraining lacked extensive reasoning-related data. To this end, we introduce MetaAct-RL, a new RL framework that frames LMs’ thinking as sequential decision making over meta-actions. In this framework, the model chooses and executes a high-level action at each step—such as forward reasoning, critique, or refinement—to gradually reach the correct answer. To encourage deeper exploration, richer action diversity, and to improve sampling efficiency in the RL optimization process, MetaAct-RL incorporates appropriate length-based reward and regularization, and a key-state restart mechanism. Extensive experiments across six benchmarks show that MetaAct-RL improves reasoning performance by 7.99 on Llama3.2-1B and 7.17 on Llama3.1-8B relative to vanilla RL method. Moreover, on the challenging AIME-2024, our method outperforms the vanilla RL by 7.5 with Qwen2.5-1.5B.
Zhiheng Xi, Yiwen Ding, Senjie Jin, Shichun Liu, Jixuan Huang, Dingwen Yang, Jiafu Tang, Boyang Hong, Junjie Ye 0005, Shihan Dou, Ming Zhang 0030, Jian Guan 0002, Wei Wu 0014, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
AAAI11
2026 Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-Training
abstract
Changhao Jiang, Ming Zhang, Yifei Cao, Junjie Ye, Xiaoran Fan, Shihan Dou, Zhiheng Xi, Jiajun Sun, Yi Dong, Yujiong Shen, Jingqi Tong, Baoyu Fan, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Changhao Jiang, Ming Zhang 0030, Yifei Cao, Junjie Ye 0005, Xiaoran Fan, Shihan Dou, Zhiheng Xi, Yujiong Shen, Jingqi Tong, Baoyu Fan, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ACL (1)4
2026 DARM: Distribution-Aware Reward Modeling by Alleviating Biases from Low Preference-Context Dependency Data
abstract
Shaofan Liu, Guoqiang Zhang, Shihan Dou, Huiyuan Zheng, Yiming Zhou, Junjie Ye, Shaowen Wang, Shichun Liu, Jiazheng Zhang, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Shaofan Liu, Shihan Dou, Huiyuan Zheng, Junjie Ye 0005, Shichun Liu, Jiazheng Zhang, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ACL (1)6
2026 AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments
abstract
Zhiheng Xi, Dingwen Yang, Jiaqi Liu, Jixuan Huang, Honglin Guo, Baodai Huang, Tinggang Chen, Qi Zhang, Zhonghang Lu, Chenyu Liu, Jiajun Sun, Jiazheng Zhang, Dingwei Zhu, Xin Guo, Junzhe Wang, Zhihao Zhang, Yuming Yang, Junjie Ye, Minghe Gao, Dongrui Liu, Jiaming Ji, Guohao Li, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhiheng Xi, Dingwen Yang, Jixuan Huang, Honglin Guo, Baodai Huang, Tinggang Chen, Qi Zhang 0001, Zhonghang Lu, Jiazheng Zhang, Dingwei Zhu, Junzhe Wang 0001, Zhihao Zhang 0002, Yuming Yang 0001, Junjie Ye 0005, Minghe Gao, Dongrui Liu, Jiaming Ji, Tao Gui, Xuanjing Huang 0001
ACL (1)18
2026 VRPO: Rethinking Value Modeling for Robust RL under Noisy Supervision in LLM Post-Training
abstract
Dingwei Zhu, Shihan Dou, Zhiheng Xi, Senjie Jin, Guoqiang Zhang, Jiazheng Zhang, Junjie Ye, Mingxu Chai, Enyu Zhou, Ming Zhang, Yuhui Wang, Caishuang Huang, Chenhao Huang, Yunke Zhang, Yuran Wang, Tao Gui, Qi Zhang, Xipeng Qiu, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Dingwei Zhu, Shihan Dou, Zhiheng Xi, Senjie Jin, Jiazheng Zhang, Junjie Ye 0005, Mingxu Chai, Enyu Zhou, Ming Zhang 0030, Caishuang Huang, Chenhao Huang, Yunke Zhang, Tao Gui, Qi Zhang 0001, Xipeng Qiu, Xuanjing Huang 0001
ACL (1)7
2025 Alleviating Shifted Distribution in Human Preference Alignment through Meta-Learning
abstract
The capability of the reward model (RM) is crucial for the success of Reinforcement Learning from Human Feedback (RLHF) in aligning with human preferences. However, as training progresses, the output space distribution of the policy model shifts. The RM, initially trained on responses sampled from the output distribution of the early policy model, gradually loses its ability to distinguish between responses from the newly shifted distribution. This issue is further compounded when the RM, trained on a specific data distribution, struggles to generalize to examples outside of that distribution. These two issues can be united as a challenge posed by the shifted distribution of the environment. To surmount this challenge, we introduce MetaRM, a novel method leveraging meta-learning to adapt the RM to the shifted environment distribution. MetaRM optimizes the RM in an alternating way, by preserving both the preferences of the original preference pairs, as well as maximizing discrimination power over new examples of the shifted distribution. Extensive experiments demonstrate that MetaRM can iteratively enhance the performance of human preference alignment by improving the RM's capacity to identify subtle differences in samples of shifted distributions.
Shihan Dou, Yan Liu 0002, Enyu Zhou, Songyang Gao, Tianlong Li, Limao Xiong, Haoxiang Jia, Junjie Ye 0005, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
AAAI9
2025 Measuring Data Diversity for Instruction Tuning: A Systematic Analysis and A Reliable Metric
abstract
Data diversity is crucial for the instruction tuning of large language models. Existing studies have explored various diversity-aware data selection methods to construct high-quality datasets and enhance model performance. However, the fundamental problem of precisely defining and measuring data diversity remains underexplored, limiting clear guidance for data engineering. To address this, we systematically analyze 11 existing diversity measurement methods by evaluating their correlation with model performance through extensive fine-tuning experiments. Our results indicate that a reliable diversity measure should properly account for both inter-sample differences and the information density in the sample space. Building on this, we propose NovelSum, a new diversity metric based on sample-level “novelty.” Experiments on both simulated and real-world data show that NovelSum accurately captures diversity variations and achieves a 0.97 correlation with instruction-tuned model performance, highlighting its value in guiding data engineering practices. With NovelSum as an optimization objective, we further develop a greedy, diversity-oriented data selection strategy that outperforms existing approaches, validating both the effectiveness and practical significance of our metric.
Yuming Yang 0001, Junjie Ye 0005, Shihan Dou, Xiao Wang 0042, Huijie Lv, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ACL (1)3
2025 ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool Use
abstract
Junjie Ye, Zhengyin Du, Xuesong Yao, Weijian Lin, Yufei Xu, Zehui Chen, Zaiyuan Wang, Sining Zhu, Zhiheng Xi, Siyu Yuan, Tao Gui, Qi Zhang, Xuanjing Huang, Jiecao Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Junjie Ye 0005, Zhengyin Du, Xuesong Yao, Weijian Lin, Yufei Xu, Zaiyuan Wang, Sining Zhu, Zhiheng Xi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001, Jiecao Chen
ACL (1)1
2025 Beyond Boundaries: Learning a Universal Entity Taxonomy across Datasets and Languages for Open Named Entity Recognition
abstract
Open Named Entity Recognition (NER), which involves identifying arbitrary types of entities from arbitrary domains, remains challenging for Large Language Models (LLMs). Recent studies suggest that fine-tuning LLMs on extensive NER data can boost their performance. However, training directly on existing datasets neglects their inconsistent entity definitions and redundant data, limiting LLMs to dataset-specific learning and hindering out-of-domain adaptation. To address this, we present B2NERD, a compact dataset designed to guide LLMs’ generalization in Open NER under a universal entity taxonomy. B2NERD is refined from 54 existing English and Chinese datasets using a two-step process. First, we detect inconsistent entity definitions across datasets and clarify them by distinguishable label names to construct a universal taxonomy of 400+ entity types. Second, we address redundancy using a data pruning strategy that selects fewer samples with greater category and semantic diversity. Comprehensive evaluation shows that B2NERD significantly enhances LLMs’ Open NER capabilities. Our B2NER models, trained on B2NERD, outperform GPT-4 by 6.8-12.0 F1 points and surpass previous methods in 3 out-of-domain benchmarks across 15 datasets and 6 languages. The data, models, and code are publicly available at https://github.com/UmeanNever/B2NER.
Yuming Yang 0001, Wantong Zhao, Caishuang Huang, Junjie Ye 0005, Xiao Wang 0042, Huiyuan Zheng, Xueying Xu, Kaixin Huang, Yunke Zhang, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
COLING4
2025 ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios
abstract
Existing evaluations of tool learning primarily focus on validating the alignment of selected tools for large language models (LLMs) with expected outcomes. However, these approaches rely on a limited set of scenarios where answers can be pre-determined. Furthermore, a sole emphasis on outcomes disregards the complex capabilities required for LLMs to effectively use tools. To tackle this issue, we propose ToolEyes, a fine-grained system tailored for the evaluation of the LLMs’ tool learning capabilities in authentic scenarios. The system meticulously examines seven real-world scenarios, analyzing five dimensions crucial to LLMs in tool learning: format alignment, intent comprehension, behavior planning, tool selection, and answer organization. Additionally, ToolEyes incorporates a tool library boasting approximately 600 tools, serving as an intermediary between LLMs and the physical world. Evaluations involving ten LLMs across three categories reveal a preference for specific scenarios and limited cognitive abilities in tool learning. Intriguingly, expanding the model size even exacerbates the hindrance to tool learning. The code and data are available at https://github.com/Junjie-Ye/ToolEyes.
Junjie Ye 0005, Songyang Gao, Caishuang Huang, Yilong Wu, Sixian Li, Xiaoran Fan, Shihan Dou, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001
COLING1
2025 CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios
abstract
Pretrained language models (LMs) are prone to arithmetic errors.Existing work showed limited success in probing numeric values from models' representations, indicating that these errors can be attributed to the inherent unreliability of distributionally learned embeddings in representing exact quantities.However, we observe that previous probing methods are inadequate for the emergent structure of learned number embeddings with sinusoidal patterns.In response, we propose a novel probing technique that decodes numeric values from input embeddings with near-perfect accuracy across a range of open-source LMs.This proves that after the sole pre-training, LMs represent numbers with remarkable precision.Finally, we find that the embeddings' precision, judged by our probe's accuracy, explains a large portion of LM's errors in elementary arithmetic, and show that aligning the embeddings with the pattern our probes discover can mitigate these errors.
Shiting Huang, Junjie Ye 0005, Lin Chen 0019, Feng Zhao 0004
EMNLP5
2025 Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levels
abstract
Junjie Ye, Yuming Yang, Yang Nan, Shuo Li, Qi Zhang, Tao Gui, Xuanjing Huang, Peng Wang, Zhongchao Shi, Jianping Fan. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Junjie Ye 0005, Yuming Yang 0001, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001, Peng Wang 0095, Zhongchao Shi, Jianping Fan 0007
EMNLP1
2024 ToolSword: Unveiling Safety Issues of Large Language Models in Tool Learning Across Three Stages
abstract
Junjie Ye, Sixian Li, Guanyu Li, Caishuang Huang, Songyang Gao, Yilong Wu, Qi Zhang, Tao Gui, Xuanjing Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Junjie Ye 0005, Sixian Li, Caishuang Huang, Songyang Gao, Yilong Wu, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001
ACL (1)1
2024 Improving Discriminative Capability of Reward Models in RLHF Using Contrastive Learning
abstract
Lu Chen, Rui Zheng, Binghai Wang, Senjie Jin, Caishuang Huang, Junjie Ye, Zhihao Zhang, Yuhao Zhou, Zhiheng Xi, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Lu Chen 0001, Binghai Wang, Senjie Jin, Caishuang Huang, Junjie Ye 0005, Zhihao Zhang 0002, Yuhao Zhou 0005, Zhiheng Xi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
EMNLP6
2024 RoTBench: A Multi-Level Benchmark for Evaluating the Robustness of Large Language Models in Tool Learning
abstract
Junjie Ye, Yilong Wu, Songyang Gao, Caishuang Huang, Sixian Li, Guanyu Li, Xiaoran Fan, Qi Zhang, Tao Gui, Xuanjing Huang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Junjie Ye 0005, Yilong Wu, Songyang Gao, Caishuang Huang, Sixian Li, Xiaoran Fan, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001
EMNLP1
2024 TransferTOD: A Generalizable Chinese Multi-Domain Task-Oriented Dialogue System with Transfer Capabilities
abstract
Ming Zhang, Caishuang Huang, Yilong Wu, Shichun Liu, Huiyuan Zheng, Yurui Dong, Yujiong Shen, Shihan Dou, Jun Zhao, Junjie Ye, Qi Zhang, Tao Gui, Xuanjing Huang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Ming Zhang 0030, Caishuang Huang, Yilong Wu, Shichun Liu, Huiyuan Zheng, Yurui Dong 0001, Yujiong Shen, Shihan Dou, Jun Zhao 0019, Junjie Ye 0005, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001
EMNLP10
2024 Linear Alignment: A Closed-form Solution for Aligning Human Preferences without Tuning and Feedback
abstract
The success of AI assistants based on Language Models (LLMs) hinges on Reinforcement Learning from Human Feedback (RLHF) to comprehend and align with user intentions. However, traditional alignment algorithms, such as PPO, are hampered by complex annotation and training requirements. This reliance limits the applicability of RLHF and hinders the development of professional assistants tailored to diverse human preferences. In this work, we introduce Linear Alignment, a novel algorithm that aligns language models with human preferences in one single inference step, eliminating the reliance on data annotation and model training. Linear alignment incorporates a new parameterization for policy optimization under divergence constraints, which enables the extraction of optimal policy in a closed-form manner and facilitates the direct estimation of the aligned response. Extensive experiments on both general and personalized preference datasets demonstrate that linear alignment significantly enhances the performance and efficiency of LLM alignment across diverse scenarios.
Songyang Gao, Qiming Ge, Shihan Dou, Junjie Ye 0005, Xiao Wang 0001, Yicheng Zou, Zhi Chen 0006, Hang Yan 0001, Qi Zhang 0001, Dahua Lin
ICML5
2022 Causal Intervention Improves Implicit Sentiment Analysis
abstract
Despite having achieved great success for sentiment analysis, existing neural models struggle with implicit sentiment analysis. It is because they may latch onto spurious correlations (“shortcuts”, e.g., focusing only on explicit sentiment words), resulting in undermining the effectiveness and robustness of the learned model. In this work, we propose a CausaL intervention model for implicit sEntiment ANalysis using instrumental variable (CLEAN). We first review sentiment analysis from a causal perspective and analyze the confounders existing in this task. Then, we introduce instrumental variable to eliminate the confounding causal effects, thus extracting the pure causal effect between sentence and sentiment. We compare the proposed CLEAN with several strong baselines on both the general implicit sentiment analysis and aspect-based implicit sentiment analysis tasks. The results indicate the great advantages of our model and the efficacy of implicit sentiment reasoning.
Siyin Wang, Jie Zhou 0015, Changzhi Sun, Junjie Ye 0005, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
COLING4
2022 Sentiment-aware multimodal pre-training for multimodal sentiment analysis
Junjie Ye 0005, Jie Zhou 0015, Rui Wang 0005, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
Knowl. Based Syst.1