Yuyang Ding

dblp:253/6501 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Escaping the Echo Trap: On Credit Assignment Failure in Multi-turn LLM Self-Reflection
abstract
Linxuan Du, Guangquan Xue, Xiaobo Liang, Qipeng Huang, Yuyang Ding, Xinyu Shi, Zhang Yijun, Ji Qi, Wenpeng Zhu, Juntao Li, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Linxuan Du, Guangquan Xue, Xiaobo Liang, Qipeng Huang, Yuyang Ding, Xinyu Shi 0005, Zhang Yijun, Wenpeng Zhu, Juntao Li 0005, Min Zhang 0005
ACL (1)5
2026 DUAL RM: Beyond Rule-based Preference Reward Modeling via Meta-Reward
abstract
Xiaobo Liang, Wanfu Wang, Qipeng Huang, Yuyang Ding, Zecheng Tang, Yixin Ji, Qianben Chen, Zhe Zhao, Kehai Chen, Juntao Li, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xiaobo Liang, Wanfu Wang, Qipeng Huang, Yuyang Ding, Zecheng Tang, Yixin Ji, Qianben Chen, Kehai Chen, Juntao Li 0005, Min Zhang 0005
ACL (1)4
2026 A survey of slow thinking-based reasoning LLMs using reinforcement learning and test-time scaling law
Qianjun Pan, Wenkai Ji, Yuyang Ding, Junsong Li, Shilian Chen, Jie Zhou 0015, Qin Chen 0001, Min Zhang 0068, Yulan Wu, Liang He 0001
Inf. Process. Manag.3
2025 Unleashing LLM Reasoning Capability via Scalable Question Synthesis from Scratch
abstract
Improving the mathematical reasoning capabilities of Large Language Models (LLMs) is critical for advancing artificial intelligence. However, access to extensive, diverse, and high-quality reasoning datasets remains a significant challenge, particularly for the open-source community. In this paper, we propose ScaleQuest, a novel, scalable, and cost-effective data synthesis method that enables the generation of large-scale mathematical reasoning datasets using lightweight 7B-scale models. ScaleQuest introduces a two-stage question-tuning process comprising Question Fine-Tuning (QFT) and Question Preference Optimization (QPO) to unlock the question generation capabilities of problem-solving models. By generating diverse questions from scratch – without relying on powerful proprietary models or seed data – we produce a dataset of 1 million problem-solution pairs. Our experiments demonstrate that models trained on our data outperform existing open-source datasets in both in-domain and out-of-domain evaluations. Furthermore, our approach shows continued performance improvement as the volume of training data increases, highlighting its potential for ongoing data scaling. The extensive improvements observed in code reasoning tasks demonstrate the generalization capabilities of our proposed method. Our work provides the open-source community with a practical solution to enhance the mathematical reasoning abilities of LLMs.
Yuyang Ding, Xinyu Shi 0005, Xiaobo Liang, Juntao Li 0005, Zhaopeng Tu, Qiaoming Zhu, Min Zhang 0005
ACL (1)1
2025 Watch out for Bears: Do People Behave Differently in Perceptual and Financial Decisions?
Yuyang Ding, Duarte Gonçalves, Maarten Speekenbrink
CogSci1
2025 Teaching LLMs for Step-Level Automatic Math Correction via Reinforcement Learning
abstract
Automatic math correction aims to check students’ solutions to mathematical problems via artificial intelligence technologies. Most existing studies focus on judging the final answer at the problem level, while they ignore detailed feedback on each step in a math problem-solving process, which requires abilities of semantic understanding and reasoning. In this paper, we propose a reinforcement learning (RL)-based method to boost large language model (LLM) for step-level automatic math correction, named StepAMC. Particularly, we convert the step-level automatic math correction within the text classification task into an RL problem to enhance the reasoning capabilities of LLMs. Then, we design a space-constrained policy network to improve the stability of RL. Then, we introduce a fine-grained reward network to convert the binary human feedback into a continuous value. We conduct extensive experiments over two benchmark datasets and the results show that our model outperforms the eleven strong baselines.
Junsong Li, Jie Zhou 0015, Yutao Yang, Bihao Zhan, Qianjun Pan, Yuyang Ding, Qin Chen 0001, Jiang Bo, Xin Lin 0001, Liang He 0001
ICME6
2025 SCAN: Self-Denoising Monte Carlo Annotation for Robust Process Reward Learning
abstract
Process reward models (PRMs) offer fine-grained, step-level evaluations that facilitate deeper reasoning processes in large language models (LLMs), proving effective in complex tasks like mathematical reasoning. However, developing PRMs is challenging due to the high cost and limited scalability of human-annotated data. Synthetic data from Monte Carlo (MC) estimation is a promising alternative but suffers from a high noise ratio, which can cause overfitting and hinder large-scale training. In this work, we conduct a preliminary study on the noise distribution in synthetic data from MC estimation, identifying that annotation models tend to both underestimate and overestimate step correctness due to limitations in their annotation capabilities. Building on these insights, we propose {\bf S}elf-Denoising Monte {\bf C}arlo {\bf An}notation (\textsc{Scan}), an efficient data synthesis and noise-tolerant learning framework. Our key findings indicate that: (1) Even lightweight models (e.g., 1.5B parameters) can produce high-quality annotations through self-denoising strategy, enabling PRMs to achieve superior performance with only 6\% the inference cost required by vanilla MC estimation. (2) With our robust learning strategy, PRMs can effectively learn from this weak supervision, achieving a 39.2 F1 score improvement (from 19.9 to 59.1) in ProcessBench. Despite using only a compact synthetic dataset, our models surpass strong baselines, including those trained on large-scale human-annotated datasets such as PRM800K. Furthermore, performance continues to improve as we scale up the synthetic data, highlighting the potential of \textsc{Scan} for scalable, cost-efficient, and robust PRM training.
Yuyang Ding, Xinyu Shi 0005, Juntao Li 0005, Xiaobo Liang, Zhaopeng Tu, Min Zhang 0005
NeurIPS1
2025 OpenBA: an open-sourced 15B bilingual asymmetric Seq2Seq model pre-trained from scratch
Juntao Li 0005, Zecheng Tang, Yuyang Ding, Pinzheng Wang, Pei Guo, Wangjie You, Wenliang Chen, Guohong Fu, Qiaoming Zhu, Guodong Zhou 0001, Min Zhang 0005
Sci. China Inf. Sci.3
2025 Towards DS-NER: Unveiling and Addressing Latent Noise in Distant Annotations
abstract
Distantly supervised named entity recognition (DS-NER) has emerged as a cheap and convenient alternative to traditional human annotation methods, enabling the automatic generation of training data by aligning text with external resources. Despite the many efforts in noise measurement methods, few works focus on the latent noise distribution between different distant annotation methods. In this work, we explore the effectiveness and robustness of DS-NER by two aspects: (1) distant annotation techniques, which encompasses both traditional rule-based methods and the innovative large language model supervision approach, and (2) noise assessment, for which we introduce a novel framework. This framework addresses the challenges by distinctly categorizing them into theunlabeled-entity problem (UEP)and thenoisy-entity problem (NEP), subsequently providing specialized solutions for each. Our proposed method achieves significant improvements on eight real-world distant supervision datasets originating from three different data sources and involving four distinct annotation techniques, confirming its superiority over current state-of-the-art methods.
Yuyang Ding, Juntao Li 0005, Jiajie Xu 0001, Pingfu Chao, Xiaofang Zhou 0001, Min Zhang 0005
IEEE Trans. Knowl. Data Eng.1
2024 Boosting Large Language Models with Socratic Method for Conversational Mathematics Teaching
abstract
With the introduction of large language models (LLMs), automatic math reasoning has seen tremendous success. However, current methods primarily focus on providing solutions or using techniques like Chain-of-Thought to enhance problem-solving accuracy. In this paper, we focus on improving the capability of mathematics teaching via a Socratic teaching-based LLM (SocraticLLM), which guides learners toward profound thinking with clarity and self-discovery via conversation. We collect and release a high-quality mathematical teaching dataset, named SocraticMATH, which provides Socratic-style conversations of problems with extra knowledge. Also, we propose a knowledge-enhanced LLM as a strong baseline to generate reliable responses with review, guidance/heuristic, rectification, and summarization. Experimental results show the great advantages of SocraticLLM by comparing it with several strong generative models. The codes and datasets are available on https://github.com/ECNU-ICALK/SocraticMath.
Yuyang Ding, Hanglei Hu, Jie Zhou 0015, Qin Chen 0001, Bo Jiang 0016, Liang He 0001
CIKM1
2024 CMD: a framework for Context-aware Model self-Detoxification
abstract
Text detoxification aims to minimize the risk of language models producing toxic content.However, existing detoxification methods fail to balance the detoxification effectiveness and generation quality.This issue arises from neglecting the constraints imposed by the context: language models are designed to generate output that closely matches the given context, while detoxification methods strive to ensure the safety of the output, even if it deviates semantically from the context.Given this, we introduce a Context-aware Model self-Detoxification (CMD) framework that pays attention to both the context and the detoxification process, i.e., first detoxifying the context and then making the language model generate along the safe context.Specifically, CMD framework involves two phases: utilizing language models to synthesize data and applying these data for training.We also introduce a toxic contrastive loss that encourages the model generation away from the negative toxic samples.Experiments on various LLMs have verified the effectiveness of our MSD framework, which can yield the best performance compared to baselines. 1 Warning: cases in this paper may contain offensive content.
Zecheng Tang, Keyan Zhou, Juntao Li 0005, Yuyang Ding, Pinzheng Wang, Yan Bowen, Renjie Hua, Min Zhang 0005
EMNLP4
2022 SelfMix: Robust Learning against Textual Label Noise with Self-Mixup Training
abstract
The conventional success of textual classification relies on annotated data, and the new paradigm of pre-trained language models (PLMs) still requires a few labeled data for downstream tasks. However, in real-world applications, label noise inevitably exists in training data, damaging the effectiveness, robustness, and generalization of the models constructed on such data. Recently, remarkable achievements have been made to mitigate this dilemma in visual data, while only a few explore textual data. To fill this gap, we present SelfMix, a simple yet effective method, to handle label noise in text classification tasks. SelfMix uses the Gaussian Mixture Model to separate samples and leverages semi-supervised learning. Unlike previous works requiring multiple models, our method utilizes the dropout mechanism on a single model to reduce the confirmation bias in self-training and introduces a textual level mixup training strategy. Experimental results on three text classification benchmarks with different types of text show that the performance of our proposed method outperforms these strong baselines designed for both textual and visual data under different noise ratios and noise types. Our anonymous code is available at https://github.com/noise-learning/SelfMix.
Chenchen Dai, Yuyang Ding, Juntao Li 0005, Wenliang Chen, Min Zhang 0005
COLING3
2022 Rain Streak Removal From Light Field Images
abstract
Raining is a common weather condition, and may seriously degrade the performances of outdoor computer vision systems, such as surveillance and autonomous navigation. Rain streaks may exhibit diverse appearances in the captured images, depending on their distances from the camera. For example, sparse rain streaks near the camera lens may appear as continuous and translucent strips, while distant densely accumulated rain streaks are more like fog and mist. Existing rain removal methods are mainly based on a single input image. However, on a single image, it is difficult to estimate a reliable depth map for rain removal. A light field image (LFI) records abundant structural and texture information of the target scene by capturing multi-perspective sub-aperture views with a single exposure. With a LFI, it is easier to estimate the depth maps, and rain streak locations across sub-aperture views are highly correlated. We observe that rain streaks usually have different slops and/or chromaic values, compared with the background scene, along the epipolar plane images (EPIs) of an LFI. Thus, we propose to make use of 3D EPIs to detect rain streaks and restore the background. To this end, we propose a novel GAN architecture to remove rain streaks from an LFI. Our method takes as input a 3D EPI, i.e., a stacked of sub-aperture views along the same row of a rainy LFI. It first estimates the disparity maps for the 3D EPI by utilizing an auto-encoder based depth estimation sub-network. The disparity maps concatenated with the input sub-aperture views are then fed into a non-local residual block, and two branched autoencoder sub-networks are used to extract rain-streaks and recover rain-free sub-aperture views. Extensive experiments conducted on both synthetic real-world-like LFIs and real-world LFIs demonstrate the effectiveness of our method.
Yuyang Ding, Tao Yan 0001, Fan Zhang 0063, Yuan Liu 0021, Rynson W. H. Lau
IEEE Trans. Circuits Syst. Video Technol.1