Qianyu He

dblp:317/0033 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
16since 2021 · last 2026
0000-0001-9810-2408ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 What Makes an Ideal Quote? Recommending "Unexpected yet Rational" Quotations via Novelty
abstract
Powei Chang, Jin Xiao, Guanglei Yue, Qianyu He, Yanghua Xiao, Deqing Yang, Jiaqing Liang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Powei Chang, Guanglei Yue, Qianyu He, Yanghua Xiao, Deqing Yang, Jiaqing Liang
ACL (1)4
2026 Metaphor Reasoning is Meta-reasoning
abstract
Qianyu He, Junting Lu, Yikai Zhang, Siyu Yuan, Xiaojun Meng, Jiansheng Wei, Jiaqing Liang, Yanghua Xiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Qianyu He, Junting Lu, Yikai Zhang 0004, Xiaojun Meng, Jiansheng Wei, Jiaqing Liang, Yanghua Xiao
ACL (1)1
2026 Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following
abstract
Qingyu Ren, Qianyu He, Powei Chang, Jie Zeng, Zeye Sun, Fei Yu, Jiaqing Liang, Yanghua Xiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Qingyu Ren, Qianyu He, Powei Chang, Jie Zeng 0003, Zeye Sun, Jiaqing Liang, Yanghua Xiao
ACL (1)2
2025 Order Doesn't Matter, But Reasoning Does: Training LLMs with Order-Centric Augmentation
abstract
Logical reasoning is essential for large language models (LLMs) to ensure accurate and coherent inference.However, LLMs struggle with reasoning order variations and fail to generalize across logically equivalent transformations.LLMs often rely on fixed sequential patterns rather than true logical understanding.To address this issue, we introduce an order-centric data augmentation framework based on commutativity in logical reasoning.We first randomly shuffle independent premises to introduce condition order augmentation.For reasoning steps, we construct a directed acyclic graph (DAG) to model dependencies between steps, which allows us to identify valid reorderings of steps while preserving logical correctness.By leveraging order-centric augmentations, models can develop a more flexible and generalized reasoning process.Finally, we conduct extensive experiments across multiple logical reasoning benchmarks, demonstrating that our method significantly enhances LLMs' reasoning performance and adaptability to diverse logical structures.We release our codes and augmented data in https://github.com/qianxiHe147/ Order-
Qianxi He, Qianyu He, Jiaqing Liang, Weikang Zhou, Zeye Sun, Yanghua Xiao
EMNLP2
2025 Think Thrice Before You Act: Progressive Thought Refinement in Large Language Models
abstract
Recent advancements in large language models (LLMs) have demonstrated that progressive refinement, rather than providing a single answer, results in more accurate and thoughtful outputs. However, existing methods often rely heavily on supervision signals to evaluate previous responses, making it difficult to effectively assess output quality in more open-ended scenarios. Additionally, these methods are typically designed for specific tasks, which limits their generalization to new domains. To address these limitations, we propose Progressive Thought Refinement (PTR), a framework that enables LLMs to progressively refine their responses. PTR operates in two phases: (1) Thought data construction stage: We propose a weak and strong model collaborative selection strategy to build a high-quality progressive refinement dataset to ensure logical consistency from thought to answers, and the answers are gradually refined in each round. (2) Thought-Mask Fine-Tuning Phase: We design a training structure to mask the "thought" and adjust loss weights to encourage LLMs to refine prior thought, teaching them to implicitly understand "how to improve" rather than "what is correct." Experimental results show that PTR significantly enhances LLM performance across ten diverse tasks (avg. from 49.6% to 53.5%) without task-specific fine-tuning. Notably, in more open-ended tasks, LLMs also demonstrate substantial improvements in the quality of responses beyond mere accuracy, suggesting that PTR truly teaches LLMs to self-improve over time. Our work is now open-source. https://github.com/cydu24/Progressive-Thought-Refinement
Chengyu Du, Jinyi Han, Yizhou Ying, Aili Chen, Qianyu He, Haokun Zhao, Haoran Guo, Sirui Xia, Jiaqing Liang, Zulong Chen, Liangyue Li, Yanghua Xiao
ICLR5
2025 Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles
abstract
Large Language Models (LLMs), such as OpenAI’s o1 and DeepSeek’s R1, excel at advanced reasoning tasks like math and coding via Reinforcement Learning with Verifiable Rewards (RLVR), but still struggle with puzzles solvable by humans without domain knowledge. We introduce ENIGMATA, the first comprehensive suite tailored for improving LLMs with puzzle reasoning skills. It includes 36 tasks across 7 categories, each with: 1) a generator that produces unlimited examples with controllable difficulty, and 2) a rule-based verifier for automatic evaluation. This generator-verifier design supports scalable, multi-task RL training, fine-grained analysis, and seamless RLVR integration. We further propose ENIGMATA-Eval, a rigorous benchmark, and develop optimized multi-task RLVR strategies. Our trained model, Qwen2.5-32B-ENIGMATA, consistently surpasses o3-mini-high and o1 on the puzzle reasoning benchmarks like ENIGMATA-Eval, ARC-AGI (32.8%), and ARC-AGI 2 (0.6%). It also generalizes well to out-of-domain puzzle benchmarks and mathematical reasoning, with little multi-tasking trade-off. When trained on larger models like Seed1.5-Thinking (20B activated parameters and 200B total parameters), puzzle data from ENIGMATA further boosts SoTA performance on advanced math and STEM reasoning tasks such as AIME (2024-2025), BeyondAIME and GPQA (Diamond), showing nice generalization benefits of ENIGMATA. This work offers a unified, controllable framework for advancing logical reasoning in LLMs. Project page: https://seed-enigmata.github.io.
Jiangjie Chen, Qianyu He, Aili Chen, Zhicheng Cai, Weinan Dai, Hongli Yu, Jiaze Chen, Qiying Yu, Hao Zhou 0012, Mingxuan Wang
NeurIPS2
2025 KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
abstract
Recent advancements in large language models (LLMs) underscore the need for more comprehensive evaluation methods to accurately assess their reasoning capabilities. Existing benchmarks are often domain-specific and thus cannot fully capture an LLM’s general reasoning potential. To address this limitation, we introduce the **Knowledge Orthogonal Reasoning Gymnasium (KORGym)**, a dynamic evaluation platform inspired by KOR-Bench and Gymnasium. KORGym offers over fifty games in either textual or visual formats and supports interactive, multi-turn assessments with reinforcement learning scenarios. Using KORGym, we conduct extensive experiments on 19 LLMs and 8 VLMs, revealing consistent reasoning patterns within model families and demonstrating the superior performance of closed-source models. Further analysis examines the effects of modality, reasoning strategies, reinforcement learning techniques, and response length on model performance. We expect KORGym to become a valuable resource for advancing LLM reasoning research and developing evaluation methodologies suited to complex, interactive environments.
Jiajun Shi, Jian Yang 0037, Xingyuan Bu, Jiangjie Chen, Junting Zhou, Kaijing Ma, Zhoufutu Wen, Bingli Wang, Yancheng He, Hualei Zhu, Wei Zhang 0021, Ruibin Yuan, Yunli Wang, Siyuan Fang, Qianyu He, Robert Tang, Yingshui Tan, Wangchunshu Zhou, Zhaoxiang Zhang 0001, Zhoujun Li 0001, Wenhao Huang 0001, Ge Zhang 0009
NeurIPS22
2024 Xiezhi: An Ever-Updating Benchmark for Holistic Domain Knowledge Evaluation
abstract
New Natural Langauge Process~(NLP) benchmarks are urgently needed to align with the rapid development of large language models (LLMs). We present Xiezhi, the most comprehensive evaluation suite designed to assess holistic domain knowledge.Xiezhi comprises multiple-choice questions across 516 diverse disciplines ranging from 13 different subjects with 249,587 questions and accompanied by Xiezhi-Specialty with 14,041 questions and Xiezhi-Interdiscipline with 10,746 questions. We conduct evaluation of the 47 cutting-edge LLMs on Xiezhi. Results indicate that LLMs exceed average performance of humans in science, engineering, agronomy, medicine, and art, but fall short in economics, jurisprudence, pedagogy, literature, history, and management. All the evaluation code and data are open sourced in https://github.com/MikeGu721/XiezhiBenchmark
Zhouhong Gu, Xiaoxuan Zhu, Haoning Ye, Jianchen Wang, Sihang Jiang 0001, Zhuozhi Xiong, Weijie Wu, Qianyu He, Rui Xu 0026, Shusen Wang, Weiguo Zheng, Hongwei Feng, Yanghua Xiao
AAAI11
2024 Small Language Model Can Self-Correct
abstract
Generative Language Models (LMs) such as ChatGPT have exhibited remarkable performance across various downstream tasks. Nevertheless, one of their most prominent drawbacks is generating inaccurate or false information with a confident tone. Previous studies have devised sophisticated pipelines and prompts to induce large LMs to exhibit the capability for self-correction. However, large LMs are explicitly prompted to verify and modify their answers separately rather than completing all steps spontaneously like humans. Moreover, these complex prompts are extremely challenging for small LMs to follow. In this paper, we introduce the Intrinsic Self-Correction (ISC) in generative language models, aiming to correct the initial output of LMs in a self-triggered manner, even for those small LMs with 6 billion parameters. Specifically, we devise a pipeline for constructing self-correction data and propose Partial Answer Masking (PAM), aiming to endow the model with the capability for intrinsic self-correction through fine-tuning. We conduct experiments using LMs with parameters sizes ranging from 6 billion to 13 billion in two tasks, including commonsense reasoning and factual knowledge reasoning. Our experiments demonstrate that the outputs generated using ISC outperform those generated without self-correction. We believe that the output quality of even small LMs can be further improved by empowering them with the ability to intrinsic self-correct.
Haixia Han, Jiaqing Liang, Jie Shi 0010, Qianyu He, Yanghua Xiao
AAAI4
2024 Can Large Language Models Understand Real-World Complex Instructions?
abstract
Large language models (LLMs) can understand human instructions, showing their potential for pragmatic applications beyond traditional NLP tasks. However, they still struggle with complex instructions, which can be either complex task descriptions that require multiple tasks and constraints, or complex input that contains long context, noise, heterogeneous information and multi-turn format. Due to these features, LLMs often ignore semantic constraints from task descriptions, generate incorrect formats, violate length or sample count constraints, and be unfaithful to the input text. Existing benchmarks are insufficient to assess LLMs’ ability to understand complex instructions, as they are close-ended and simple. To bridge this gap, we propose CELLO, a benchmark for evaluating LLMs' ability to follow complex instructions systematically. We design eight features for complex instructions and construct a comprehensive evaluation dataset from real-world scenarios. We also establish four criteria and develop corresponding metrics, as current ones are inadequate, biased or too strict and coarse-grained. We compare the performance of representative Chinese-oriented and English-oriented models in following complex instructions through extensive experiments. Resources of CELLO are publicly available at https://github.com/Abbey4799/CELLO.
Qianyu He, Jie Zeng 0003, Lina Chen, Qianxi He, Xunzhe Zhou, Jiaqing Liang, Yanghua Xiao
AAAI1
2024 Enhancing Quantitative Reasoning Skills of Large Language Models through Dimension Perception
abstract
Quantities are distinct and critical components of texts that characterize the magnitude properties of entities, providing a precise perspective for the understanding of natural language, especially for reasoning tasks. In recent years, there has been a flurry of research on reasoning tasks based on large language models (LLMs), most of which solely focus on numerical values, neglecting the dimensional concept of quantities with units despite its importance. We argue that the concept of dimension is essential for precisely understanding quantities and of great significance for LLMs to perform quantitative reasoning. However, the lack of dimension knowledge and quantity-related benchmarks has resulted in low performance of LLMs. Hence, we present a framework to enhance the quantitative reasoning ability of language models based on dimension perception. We first construct a dimensional unit knowledge base (DimUnitKB) to address the knowledge gap in this area. We propose a benchmark DimEval consisting of seven tasks of three categories to probe and enhance the dimension perception skills of LLMs. To evaluate the effectiveness of our methods, we propose a quantitative reasoning task and conduct experiments. The experimental results show that our dimension perception method dramatically improves accuracy (43.55%→50.67%) on quantitative reasoning tasks compared to GPT-4.
Yuncheng Huang, Qianyu He, Jiaqing Liang, Sihang Jiang 0001, Yanghua Xiao, Yunwen Chen
ICDE2
2023 MAPS-KB: A Million-Scale Probabilistic Simile Knowledge Base
abstract
The ability to understand and generate similes is an imperative step to realize human-level AI. However, there is still a considerable gap between machine intelligence and human cognition in similes, since deep models based on statistical distribution tend to favour high-frequency similes. Hence, a large-scale symbolic knowledge base of similes is required, as it contributes to the modeling of diverse yet unpopular similes while facilitating additional evaluation and reasoning. To bridge the gap, we propose a novel framework for large-scale simile knowledge base construction, as well as two probabilistic metrics which enable an improved understanding of simile phenomena in natural language. Overall, we construct MAPS-KB, a million-scale probabilistic simile knowledge base, covering 4.3 million triplets over 0.4 million terms from 70 GB corpora. We conduct sufficient experiments to justify the effectiveness and necessity of the methods of our framework. We also apply MAPS-KB on three downstream tasks to achieve state-of-the-art performance, further demonstrating the value of MAPS-KB. Resources of MAPS-KB are publicly available at https://github.com/Abbey4799/MAPS-KB.
Qianyu He, Xintao Wang 0001, Jiaqing Liang, Yanghua Xiao
AAAI1
2023 HAUSER: Towards Holistic and Automatic Evaluation of Simile Generation
abstract
Similes play an imperative role in creative writing such as story and dialogue generation.Proper evaluation metrics are like a beacon guiding the research of simile generation (SG).However, it remains under-explored as to what criteria should be considered, how to quantify each criterion into metrics, and whether the metrics are effective for comprehensive, efficient, and reliable SG evaluation.To address the issues, we establish HAUSER, a holistic and automatic evaluation system for the SG task, which consists of five criteria from three perspectives and automatic metrics for each criterion.Through extensive experiments, we verify that our metrics are significantly more correlated with human ratings from each perspective compared with prior automatic metrics.Resources of HAUSER are publicly available at https://github.com/Abbey4799/HAUSER.
Qianyu He, Yikai Zhang 0004, Jiaqing Liang, Yuncheng Huang, Yanghua Xiao, Yunwen Chen
ACL (1)1
2022 Can Pre-trained Language Models Interpret Similes as Smart as Human?
abstract
Simile interpretation is a crucial task in natural language processing.Nowadays, pre-trained language models (PLMs) have achieved stateof-the-art performance on many tasks.However, it remains under-explored whether PLMs can interpret similes or not.In this paper, we investigate the ability of PLMs in simile interpretation by designing a novel task named Simile Property Probing, i.e., to let the PLMs infer the shared properties of similes.We construct our simile property probing datasets from both general textual corpora and humandesigned questions, containing 1,633 examples covering seven main categories.Our empirical study based on the constructed datasets shows that PLMs can infer similes' shared properties while still underperforming humans.To bridge the gap with human performance, we additionally design a knowledge-enhanced training objective by incorporating the simile knowledge into PLMs via knowledge embedding methods.Our method results in a gain of 8.58% in the probing task and 1.37% in the downstream task of sentiment classification.The datasets and code are publicly available at https://github.com/Abbey4799/PLMs- Interpret-Simile.
Qianyu He, Sijie Cheng, Zhixu Li, Rui Xie 0005, Yanghua Xiao
ACL (1)1
2022 A Context-Enhanced Generate-then-Evaluate Framework for Chinese Abbreviation Prediction
abstract
As a popular form of lexicalization, abbreviation is widely used in both oral and written language and plays an important role in various Natural Language Processing applications. However, current approaches cannot ensure that the predicted abbreviation preserves the meaning of its full form and maintains fluency. In this paper, we introduce a fresh perspective to evaluate the quality of abbreviations within their textual contexts with pre-trained language model. To this end, we propose a novel two-stage generate-then-evaluate framework enhanced by context, which consists of a generation model to generate multiple candidate abbreviations and an evaluation model to evaluate their quality within their contexts. Experimental results show that our framework consistently outperforms all the existing approaches, achieving 53.2% [email protected] performance with a 5.6 points improvement compared to its previous best result. Our code and data are publicly available at https://github.com/HavenTong/CEGE.
Hanwen Tong, Chenhao Xie 0002, Jiaqing Liang, Qianyu He, Zhiang Yue, Yanghua Xiao
CIKM4
2022 Language Models as Knowledge Embeddings
abstract
Knowledge embeddings (KE) represent a knowledge graph (KG) by embedding entities and relations into continuous vector spaces. Existing methods are mainly structure-based or description-based. Structure-based methods learn representations that preserve the inherent structure of KGs. They cannot well represent abundant long-tail entities in real-world KGs with limited structural information. Description-based methods leverage textual information and language models. Prior approaches in this direction barely outperform structure-based ones, and suffer from problems like expensive negative sampling and restrictive description demand. In this paper, we propose LMKE, which adopts Language Models to derive Knowledge Embeddings, aiming at both enriching representations of long-tail entities and solving problems of prior description-based methods. We formulate description-based KE learning with a contrastive learning framework to improve efficiency in training and evaluation. Experimental results show that LMKE achieves state-of-the-art performance on KE benchmarks of link prediction and triple classification, especially for long-tail entities.
Xintao Wang 0001, Qianyu He, Jiaqing Liang, Yanghua Xiao
IJCAI2