Yichun Yin

dblp:180/5934 · DBLP profile ↗
← Back
22ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0001-5567-8861ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Discover and Prove: An Open-source Agentic Framework for Hard Mode Automated Theorem Proving in Lean 4
abstract
Chengwu Liu, Yichun Yin, Ye Yuan, Jiaxuan Xie, Botao Li, Siqi Li, Jianhao Shen, Yan Xu, Lifeng Shang, Ming Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Chengwu Liu 0001, Yichun Yin, Ye Yuan 0016, Jiaxuan Xie, Botao Li, Jianhao Shen, Lifeng Shang, Ming Zhang 0004
ACL (1)2
2026 MATCH: Modulating Attention via In-Context Retrieval for Long-Context Transformers
abstract
Linrui Ma, Chun Hei Lo, Xinyu Wang, Peng Lu, Xihao Yuan, Hanting Chen, Kai Han, Xinghao Chen, Chengjun Zhan, Hanlin xu, Yichun Yin, Lifeng Shang, Feng Wen, Boxing Chen, Yufei Cui. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Linrui Ma, Chun Hei Lo, Xinyu Wang 0061, Peng Lu 0006, Xihao Yuan, Hanting Chen, Kai Han 0002, Xinghao Chen 0001, Chengjun Zhan, Hanlin Xu, Yichun Yin, Lifeng Shang, Boxing Chen, Yufei Cui
ACL (1)11
2025 Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification
abstract
Chain-of-Thought (CoT) prompting has become the de facto method to elicit reasoning capabilities from large language models (LLMs). However, to mitigate hallucinations in CoT that are notoriously difficult to detect, current methods such as process reward models (PRMs) or self-consistency operate as opaque boxes and do not provide checkable evidence for their judgments, possibly limiting their effectiveness. To address this issue, we draw inspiration from the idea that “the gold standard for supporting a mathematical claim is to provide a proof”. We propose a retrospective, step-aware formal verification framework Safe. Rather than assigning arbitrary scores, we strive to articulate mathematical claims in formal mathematical language Lean 4 at each reasoning step and provide formal proofs to identify hallucinations. We evaluate our framework Safe across multiple language models and various mathematical datasets, demonstrating a significant performance improvement while offering interpretable and verifiable evidence. We also propose FormalStep as a benchmark for step correctness theorem proving with 30,809 formal statements. To the best of our knowledge, our work represents the first endeavor to utilize formal mathematical language Lean 4 for verifying content generated by LLMs, aligning with the reason why formal mathematical languages were created in the first place: to provide a robust foundation for hallucination-prone human-written proofs.
Chengwu Liu 0001, Ye Yuan 0016, Yichun Yin, Zaoyu Chen, Yasheng Wang, Lifeng Shang, Qun Liu 0001, Ming Zhang 0004
ACL (1)3
2025 RidgeLoRA: Matrix Ridge Enhanced Low-Rank Adaptation of Large Language Models
abstract
As one of the state-of-the-art parameter-efficient fine-tuning~(PEFT) methods, Low-Rank Adaptation (LoRA) enables model optimization with reduced computational cost through trainable low-rank matrix. However, the low-rank nature makes it prone to produce a decrease in the representation ability, leading to suboptimal performance. In order to break this limitation, we propose RidgeLoRA, a lightweight architecture like LoRA that incorporates novel architecture and matrix ridge enhanced full-rank approximation, to match the performance of full-rank training, while eliminating the need for high memory and a large number of parameters to restore the rank of matrices. We provide a rigorous mathematical derivation to prove that RidgeLoRA has a better upper bound on the representations than vanilla LoRA. Furthermore, extensive experiments across multiple domains demonstrate that RidgeLoRA achieves better performance than other LoRA variants, and can even match or surpass full-rank training.
Junda Zhu 0003, Jun Ai, Yichun Yin, Yasheng Wang, Lifeng Shang, Qun Liu 0001
NeurIPS4
2024 Preparing Lessons for Progressive Training on Language Models
abstract
The rapid progress of Transformers in artificial intelligence has come at the cost of increased resource consumption and greenhouse gas emissions due to growing model sizes. Prior work suggests using pretrained small models to improve training efficiency, but this approach may not be suitable for new model structures. On the other hand, training from scratch can be slow, and progressively stacking layers often fails to achieve significant acceleration. To address these challenges, we propose a novel method called Apollo, which prepares lessons for expanding operations by learning high-layer functionality during training of low layers. Our approach involves low-value-prioritized sampling (LVPS) to train different depths and weight sharing to facilitate efficient expansion. We also introduce an interpolation method for stable model depth extension. Experiments demonstrate that Apollo achieves state-of-the-art acceleration ratios, even rivaling methods using pretrained models, making it a universal and efficient solution for training deep models while reducing time, financial, and environmental costs.
Yu Pan 0005, Ye Yuan 0016, Yichun Yin, Jiaxin Shi, Zenglin Xu, Ming Zhang 0004, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001
AAAI3
2024 DQ-LoRe: Dual Queries with Low Rank Approximation Re-ranking for In-Context Learning
abstract
Recent advances in natural language processing, primarily propelled by Large Language Models (LLMs), have showcased their remarkable capabilities grounded in in-context learning. A promising avenue for guiding LLMs in intricate reasoning tasks involves the utilization of intermediate reasoning steps within the Chain-of-Thought (CoT) paradigm. Nevertheless, the central challenge lies in the effective selection of exemplars for facilitating in-context learning. In this study, we introduce a framework that leverages Dual Queries and Low-rank approximation Re-ranking (DQ-LoRe) to automatically select exemplars for in-context learning. Dual Queries first query LLM to obtain LLM-generated knowledge such as CoT, then query the retriever to obtain the final exemplars via both question and the knowledge. Moreover, for the second query, LoRe employs dimensionality reduction techniques to refine exemplar selection, ensuring close alignment with the input question's knowledge. Through extensive experiments, we demonstrate that DQ-LoRe significantly outperforms prior state-of-the-art methods in the automatic selection of exemplars for GPT-4, enhancing performance from 92.5\% to 94.2\%. Our comprehensive analysis further reveals that DQ-LoRe consistently outperforms retrieval-based approaches in terms of both performance and adaptability, especially in scenarios characterized by distribution shifts. DQ-LoRe pushes the boundaries of in-context learning and opens up new avenues for addressing complex reasoning challenges.
Chuanyang Zheng, Zhijiang Guo, Yichun Yin, Enze Xie, Qingxing Cao, Xiongwei Han, Jing Tang 0004, Xiaodan Liang
ICLR5
2023 One Cannot Stand for Everyone! Leveraging Multiple User Simulators to train Task-oriented Dialogue Systems
abstract
Yajiao Liu, Xin Jiang, Yichun Yin, Yasheng Wang, Fei Mi, Qun Liu, Xiang Wan, Benyou Wang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yajiao Liu, Xin Jiang 0002, Yichun Yin, Yasheng Wang, Fei Mi, Qun Liu 0001, Benyou Wang
ACL (1)3
2023 DT-Solver: Automated Theorem Proving with Dynamic-Tree Sampling Guided by Proof-level Value Function
abstract
Haiming Wang, Ye Yuan, Zhengying Liu, Jianhao Shen, Yichun Yin, Jing Xiong, Enze Xie, Han Shi, Yujun Li, Lin Li, Jian Yin, Zhenguo Li, Xiaodan Liang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Zhengying Liu, Jianhao Shen, Yichun Yin, Enze Xie, Jian Yin 0001, Zhenguo Li, Xiaodan Liang
ACL (1)5
2023 TRIGO: Benchmarking Formal Mathematical Proof Reduction for Generative Language Models
abstract
Jing Xiong, Jianhao Shen, Ye Yuan, Haiming Wang, Yichun Yin, Zhengying Liu, Lin Li, Zhijiang Guo, Qingxing Cao, Yinya Huang, Chuanyang Zheng, Xiaodan Liang, Ming Zhang, Qun Liu. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Jianhao Shen, Ye Yuan 0016, Yichun Yin, Zhengying Liu, Zhijiang Guo, Qingxing Cao, Yinya Huang, Chuanyang Zheng, Xiaodan Liang, Ming Zhang 0004, Qun Liu 0001
EMNLP5
2023 Reusing Pretrained Models by Multi-linear Operators for Efficient Training
abstract
Training large models from scratch usually costs a substantial amount of resources. Towards this problem, recent studies such as bert2BERT and LiGO have reused small pretrained models to initialize a large model (termed the ``target model''), leading to a considerable acceleration in training. Despite the successes of these previous studies, they grew pretrained models by mapping partial weights only, ignoring potential correlations across the entire model. As we show in this paper, there are inter- and intra-interactions among the weights of both the pretrained and the target models. As a result, the partial mapping may not capture the complete information and lead to inadequate growth. In this paper, we propose a method that linearly correlates each weight of the target model to all the weights of the pretrained model to further enhance acceleration ability. We utilize multi-linear operators to reduce computational and spacial complexity, enabling acceptable resource requirements. Experiments demonstrate that our method can save 76\% computational costs on DeiT-base transferred from DeiT-small, which outperforms bert2BERT by +12\% and LiGO by +21\%, respectively.
Yu Pan 0005, Ye Yuan 0016, Yichun Yin, Zenglin Xu, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001
NeurIPS3
2022 bert2BERT: Towards Reusable Pretrained Language Models
abstract
Cheng Chen, Yichun Yin, Lifeng Shang, Xin Jiang, Yujia Qin, Fengyu Wang, Zhi Wang, Xiao Chen, Zhiyuan Liu, Qun Liu. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Yichun Yin, Lifeng Shang, Xin Jiang 0002, Yujia Qin, Zhi Wang 0001, Xiao Chen 0012, Zhiyuan Liu 0001, Qun Liu 0001
ACL (1)2
2022 G-MAP: General Memory-Augmented Pre-trained Language Model for Domain Tasks
abstract
Recently, domain-specific PLMs have been proposed to boost the task performance of specific domains (e.g., biomedical and computer science) by continuing to pre-train general PLMs with domain-specific corpora.However, this Domain-Adaptive Pre-Training (DAPT; Gururangan et al. ( 2020)) tends to forget the previous general knowledge acquired by general PLMs, which leads to a catastrophic forgetting phenomenon and sub-optimal performance.To alleviate this problem, we propose a new framework of General Memory-Augmented Pre-trained Language Model (G-MAP), which augments the domain-specific PLM by a memory representation built from the frozen general PLM without losing any general knowledge.Specifically, we propose a new memory-augmented layer, and based on it, different augmented strategies are explored to build the memory representation and then adaptively fuse it into the domain-specific PLM.We demonstrate the effectiveness of G-MAP on various domains (biomedical and computer science publications, news, and reviews) and different kinds (text classification, QA, NER) of tasks, and the extensive results show that the proposed G-MAP 1 can achieve SOTA results on all tasks.
Zhongwei Wan, Yichun Yin, Wei Zhang 0196, Jiaxin Shi, Lifeng Shang, Guangyong Chen, Xin Jiang 0002, Qun Liu 0001
EMNLP2
2021 AutoTinyBERT: Automatic Hyper-parameter Optimization for Efficient Pre-trained Language Models
abstract
Yichun Yin, Cheng Chen, Lifeng Shang, Xin Jiang, Xiao Chen, Qun Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yichun Yin, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001
ACL/IJCNLP (1)1
2021 Extract then Distill: Efficient and Effective Task-Agnostic BERT Distillation
Yichun Yin, Lifeng Shang, Zhi Wang 0001, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001
ICANN (3)2
2021 Improving task-agnostic BERT distillation with layer mapping search
Xiaoqi Jiao, Huating Chang, Yichun Yin, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Linlin Li 0001, Fang Wang 0001, Qun Liu 0001
Neurocomputing3
2020 Dialog State Tracking with Reinforced Data Augmentation
abstract
Neural dialog state trackers are generally limited due to the lack of quantity and diversity of annotated training data. In this paper, we address this difficulty by proposing a reinforcement learning (RL) based framework for data augmentation that can generate high-quality data to improve the neural state tracker. Specifically, we introduce a novel contextual bandit generator to learn fine-grained augmentation policies that can generate new effective instances by choosing suitable replacements for specific context. Moreover, by alternately learning between the generator and the state tracker, we can keep refining the generative policies to generate more high-quality training data for neural state tracker. Experimental results on the WoZ and MultiWoZ (restaurant) datasets demonstrate that the proposed framework significantly improves the performance over the state-of-the-art models, especially with limited training data.
Yichun Yin, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001
AAAI1
2020 PoD: Positional Dependency-Based Word Embedding for Aspect Term Extraction
abstract
Dependency context-based word embedding jointly learns the representations of word and dependency context, and has been proved effective in aspect term extraction.In this paper, we design the positional dependency-based word embedding (POD) which considers both dependency context and positional context for aspect term extraction.Specifically, the positional context is modeled via relative position encoding.Besides, we enhance the dependency context by integrating more lexical information (e.g., POS tags) along dependency paths.Experiments on SemEval 2014/2015/2016 datasets show that our approach outperforms other embedding methods in aspect term extraction.
Yichun Yin, Chenguang Wang 0001, Ming Zhang 0004
COLING1
2020 TernaryBERT: Distillation-aware Ultra-low Bit BERT
abstract
Transformer-based pre-training models like BERT have achieved remarkable performance in many natural language processing tasks.However, these models are both computation and memory expensive, hindering their deployment to resource-constrained devices.In this work, we propose TernaryBERT, which ternarizes the weights in a fine-tuned BERT model.Specifically, we use both approximation-based and loss-aware ternarization methods and empirically investigate the ternarization granularity of different parts of BERT.Moreover, to reduce the accuracy degradation caused by the lower capacity of low bits, we leverage the knowledge distillation technique (Jiao et al., 2019) in the training process.Experiments on the GLUE benchmark and SQuAD show that our proposed TernaryBERT outperforms the other BERT quantization methods, and even achieves comparable performance as the fullprecision model while being 14.9x smaller.
Wei Zhang 0196, Lu Hou 0002, Yichun Yin, Lifeng Shang, Xiao Chen 0012, Xin Jiang 0002, Qun Liu 0001
EMNLP (1)3
2020 The Solution of Huawei Cloud & Noah's Ark Lab to the NLPCC-2020 Challenge: Light Pre-Training Chinese Language Model for NLP Task
Yichun Yin, Qun Liu 0001
NLPCC (2)4
2017 Document-Level Multi-Aspect Sentiment Classification as Machine Comprehension
abstract
Document-level multi-aspect sentiment classification is an important task for customer relation management.In this paper, we model the task as a machine comprehension problem where pseudo questionanswer pairs are constructed by a small number of aspect-related keywords and aspect ratings.A hierarchical iterative attention model is introduced to build aspectspecific representations by frequent and repeated interactions between documents and aspect questions.We adopt a hierarchical architecture to represent both word level and sentence level information, and use the attention operations for aspect questions and documents alternatively with the multiple hop mechanism.Experimental results on the TripAdvisor and BeerAdvocate datasets show that our model outperforms classical baselines.
Yichun Yin, Yangqiu Song, Ming Zhang 0004
EMNLP1
2017 Socialized Word Embeddings
abstract
Word embeddings have attracted a lot of attention. On social media, each user’s language use can be significantly affected by the user’s friends. In this paper, we propose a socialized word embedding algorithm which can consider both user’s personal characteristics of language use and the user’s social relationship on social media. To incorporate personal characteristics, we propose to use a user vector to represent each user. Then for each user, the word embeddings are trained based on each user’s corpus by combining the global word vectors and local user vector. To incorporate social relationship, we add a regularization term to impose similarity between two friends. In this way, we can train the global word vectors and user vectors jointly. To demonstrate the effectiveness, we used the latest large-scale Yelp data to train our vectors, and designed several experiments to show how user vectors affect the results.
Ziqian Zeng, Yichun Yin, Yangqiu Song, Ming Zhang 0004
IJCAI2
2016 Unsupervised Word and Dependency Path Embeddings for Aspect Term Extraction
Yichun Yin, Furu Wei, Li Dong 0004, Kaimeng Xu, Ming Zhang 0004, Ming Zhou 0001
IJCAI1