Jiaxi Yang 0004

dblp:293/9901-4 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0002-7710-1489ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 From Completion to Editing: Unlocking Context-Aware Code Infilling via Search-and-Replace Instruction Tuning
abstract
Jiajun Zhang, Zeyu Cui, Jiaxi Yang, Lei Zhang, Yuheng Jing, Zeyao Ma, Tianyi Bai, Zilei Wang, Qiang Liu, Liang Wang, Binyuan Hui, Junyang Lin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiajun Zhang 0012, Zeyu Cui, Jiaxi Yang 0004, Lei Zhang 0201, Yuheng Jing, Zeyao Ma, Tianyi Bai, Zilei Wang, Qiang Liu 0006, Liang Wang 0001, Binyuan Hui, Junyang Lin
ACL (1)3
2025 Hierarchical Context Pruning: Optimizing Real-World Code Completion with Repository-Level Pretrained Code LLMs
abstract
Some of the latest released Code Large Language Models (Code LLMs) have been trained on repository-level code data, enabling them to perceive repository structures and utilize cross-file code information. This capability allows us to directly concatenate the content of repository code files in prompts to achieve repository-level code completion. However, in real development scenarios, directly concatenating all code repository files in a prompt can easily exceed the context window of Code LLMs, leading to a significant decline in completion performance. Additionally, overly long prompts can increase completion latency, negatively impacting the user experience. In this study, we conducted extensive experiments, including completion error analysis, topology dependency analysis, and cross-file content analysis, to investigate the factors affecting repository-level code completion. Based on the conclusions drawn from these preliminary experiments, we proposed a strategy called **Hierarchical Context Pruning (HCP)** to construct high-quality completion prompts. We applied the **HCP** to six Code LLMs and evaluated them on the CrossCodeEval dataset. The experimental results showed that, compared to previous methods, the prompts constructed using our **HCP** strategy achieved higher completion accuracy on five out of six Code LLMs. Additionally, the **HCP** managed to keep the prompt length around 8k tokens (whereas the full repository code is approximately 50k tokens), significantly improving completion throughput. Our code and data will be publicly available.
Lei Zhang 0201, Yunshui Li, Jiaming Li 0004, Xiaobo Xia, Jiaxi Yang 0004, Run Luo, Minzheng Wang 0001, Longze Chen, Junhao Liu 0001, Qiang Qu 0001, Min Yang 0007
AAAI5
2025 START: Self-taught Reasoner with Tools
abstract
Chengpeng Li, Mingfeng Xue, Zhenru Zhang, Jiaxi Yang, Beichen Zhang, Bowen Yu, Binyuan Hui, Junyang Lin, Xiang Wang, Dayiheng Liu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Chengpeng Li 0001, Mingfeng Xue, Zhenru Zhang, Jiaxi Yang 0004, Bowen Yu 0002, Binyuan Hui, Junyang Lin, Xiang Wang 0010, Dayiheng Liu
EMNLP4
2025 CodeArena: Evaluating and Aligning CodeLLMs on Human Preference
abstract
Jian Yang, Jiaxi Yang, Wei Zhang, Jin Ke, Yibo Miao, Lei Zhang, Liqun Yang, Zeyu Cui, Yichang Zhang, Zhoujun Li, Binyuan Hui, Junyang Lin. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Jian Yang 0003, Jiaxi Yang 0004, Wei Zhang 0021, Yibo Miao, Lei Zhang 0201, Liqun Yang, Zeyu Cui, Yichang Zhang, Zhoujun Li 0001, Binyuan Hui, Junyang Lin
EMNLP2
2025 Synthesizing Software Engineering Data in a Test-Driven Manner
abstract
We introduce **SWE-Flow**, a novel data synthesis framework grounded in Test-Driven Development (TDD). Unlike existing software engineering data that rely on human-submitted issues, **SWE-Flow** automatically infers incremental development steps directly from unit tests, which inherently encapsulate high-level requirements. The core of **SWE-Flow** is the construction of a Runtime Dependency Graph (RDG), which precisely captures function interactions, enabling the generation of a structured, step-by-step *development schedule*. At each step, **SWE-Flow** produces a partial codebase, the corresponding unit tests, and the necessary code modifications, resulting in fully verifiable TDD tasks. With this approach, we generated 16,061 training instances and 2,020 test instances from real-world GitHub projects, creating the **SWE-Flow-Eval** benchmark. Our experiments show that fine-tuning open model on this dataset significantly improves performance in TDD-based coding. To facilitate further research, we release all code, datasets, models, and Docker images at [Github](https://github.com/Hambaobao/SWE-Flow).
Lei Zhang 0201, Jiaxi Yang 0004, Min Yang 0007, Jian Yang 0003, Mouxiang Chen, Jiajun Zhang 0012, Zeyu Cui, Binyuan Hui, Junyang Lin
ICML2
2025 Parallel Scaling Law for Language Models
abstract
It is commonly believed that scaling language models should commit a significant space or time cost, by increasing the parameters (parameter scaling) or output tokens (inference-time scaling). We introduce another and more inference-efficient scaling paradigm: increasing the model's parallel computation during both training and inference time. We apply $P$ diverse and learnable transformations to the input, execute forward passes of the model in parallel, and dynamically aggregate the $P$ outputs. This method, namely parallel scaling (ParScale), scales parallel computation by reusing existing parameters and can be applied to any model structure, optimization procedure, data, or task. We theoretically propose a new scaling law and validate it through large-scale pre-training, which shows that a model with $P$ parallel streams is similar to scaling the parameters by $\mathcal O(\log P)$ while showing superior inference efficiency. For example, ParScale can use up to 22$\times$ less memory increase and 6$\times$ less latency increase compared to parameter scaling that achieves the same performance improvement. It can also recycle an off-the-shelf pre-trained model into a parallelly scaled one by post-training on a small amount of tokens, further reducing the training budget. The new scaling law we discovered potentially facilitates the deployment of more powerful models in low-resource scenarios, and provides an alternative perspective for the role of computation in machine learning. Our code and 67 trained model checkpoints are publicly available at https://github.com/QwenLM/ParScale and https://huggingface.co/ParScale.
Mouxiang Chen, Binyuan Hui, Zeyu Cui, Jiaxi Yang 0004, Dayiheng Liu, Jianling Sun, Junyang Lin, Zhongxin Liu 0002
NeurIPS4
2024 One-Shot Learning as Instruction Data Prospector for Large Language Models
abstract
Yunshui Li, Binyuan Hui, Xiaobo Xia, Jiaxi Yang, Min Yang, Lei Zhang, Shuzheng Si, Ling-Hao Chen, Junhao Liu, Tongliang Liu, Fei Huang, Yongbin Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yunshui Li, Binyuan Hui, Xiaobo Xia, Jiaxi Yang 0004, Min Yang 0007, Lei Zhang 0201, Shuzheng Si, Junhao Liu 0001, Tongliang Liu, Fei Huang 0002, Yongbin Li 0001
ACL (1)4
2024 Iterative Forward Tuning Boosts In-Context Learning in Language Models
abstract
Jiaxi Yang, Binyuan Hui, Min Yang, Bailin Wang, Bowen Li, Binhua Li, Fei Huang, Yongbin Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Jiaxi Yang 0004, Binyuan Hui, Min Yang 0007, Bailin Wang, Bowen Li 0002, Binhua Li, Fei Huang 0002, Yongbin Li 0001
ACL (1)1
2024 Synthesizing Text-to-SQL Data from Weak and Strong LLMs
abstract
The capability gap between open-source and closed-source large language models (LLMs) remains challenging in text-to-SQL tasks.In this paper, we introduce a synthetic data approach that amalgamates strong data generated by larger, more potent models (strong models) with weak data produced by smaller, less wellaligned models (weak models).Our approach contributes to the improvement of domain generalization in text-to-SQL models and investigates the potential of weak data supervision through preference learning.Moreover, we utilize the synthetic data approach for instruction tuning on open-source LLMs, yielding SENSE, a specialized text-to-SQL model.The effectiveness of SENSE is substantiated by achieving state-of-the-art results on the SPIDER and BIRD benchmarks, thereby mitigating the performance disparity between open-source models and the methods derived from closed-source models.
Jiaxi Yang 0004, Binyuan Hui, Min Yang 0007, Jian Yang 0003, Junyang Lin, Chang Zhou 0005
ACL (1)1
2024 Marathon: A Race Through the Realm of Long Context with Large Language Models
abstract
Lei Zhang, Yunshui Li, Ziqiang Liu, Jiaxi Yang, Junhao Liu, Longze Chen, Run Luo, Min Yang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Lei Zhang 0201, Yunshui Li, Jiaxi Yang 0004, Junhao Liu 0001, Longze Chen, Run Luo, Min Yang 0007
ACL (1)4
2023 Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs
abstract
Text-to-SQL parsing, which aims at converting natural language instructions into executable SQLs, has gained increasing attention in recent years. In particular, GPT-4 and Claude-2 have shown impressive results in this task. However, most of the prevalent benchmarks, i.e., Spider, and WikiSQL, focus on database schema with few rows of database contents leaving the gap between academic study and real-world applications. To mitigate this gap, we present BIRD, a BIg benchmark for laRge-scale Database grounded in text-to-SQL tasks, containing 12,751 pairs of text-to-SQL data and 95 databases with a total size of 33.4 GB, spanning 37 professional domains. Our emphasis on database values highlights the new challenges of dirty database contents, external knowledge between NL questions and database contents, and SQL efficiency, particularly in the context of massive databases. To solve these problems, text-to-SQL models must feature database value comprehension in addition to semantic parsing. The experimental results demonstrate the significance of database values in generating accurate text-to-SQLs for big databases. Furthermore, even the most popular and effective text-to-SQL models, i.e. GPT-4, only achieve 54.89% in execution accuracy, which is still far from the human result of 92.96%, proving that challenges still stand. We also provide an efficiency analysis to offer insights into generating text-to-efficient-SQLs that are beneficial to industries. We believe that BIRD will contribute to advancing real-world applications of text-to-SQL research.The leaderboard and source code are available: https://bird-bench.github.io/.
Jinyang Li 0003, Binyuan Hui, Ge Qu, Jiaxi Yang 0004, Binhua Li, Bowen Li 0002, Bailin Wang, Bowen Qin, Ruiying Geng, Nan Huo, Xuanhe Zhou, Chenhao Ma 0001, Guoliang Li 0001, Kevin Chen-Chuan Chang, Fei Huang 0002, Reynold Cheng, Yongbin Li 0001
NeurIPS4
2023 User-Specific Adaptive Fine-Tuning for Cross-Domain Recommendations
abstract
Making accurate recommendations for cold-start users has been a longstanding and critical challenge for recommender systems (RS). Cross-domain recommendations (CDR) offer a solution to tackle such a cold-start problem when there is no sufficient data for the users who have rarely used the system. An effective approach in CDR is to leverage the knowledge (e.g., user representations) learned from a related but different domain and transfer it to the target domain. Fine-tuning works as an effective transfer learning technique for this objective, which adapts the parameters of a pre-trained model from the source domain to the target domain. However, current methods are mainly based on the global fine-tuning strategy: the decision of which layers of the pre-trained model to freeze or fine-tune is taken for all users in the target domain. In this paper, we argue that users in RS are personalized and should have their own fine-tuning policies for better preference transfer learning. As such, we propose a novel User-specific Adaptive Fine-tuning method (UAF), selecting which layers of the pre-trained network to fine-tune, on a per-user basis. Specifically, we devise a policy network with three alternative strategies to automatically decide which layers to be fine-tuned and which layers to have their parameters frozen for each user. Extensive experiments show that the proposed UAF exhibits significantly better and more robust performance for user cold-start recommendation.
Lei Chen 0072, Fajie Yuan, Jiaxi Yang 0004, Xiangnan He 0001, Chengming Li 0004, Min Yang 0007
IEEE Trans. Knowl. Data Eng.3
2021 A User-Adaptive Layer Selection Framework for Very Deep Sequential Recommender Models
abstract
Sequential recommender systems (SRS) have become a research hotspot in recent studies. Because of the requirement in capturing user's dynamic interests, sequential neural network based recommender models often need to be stacked with more hidden layers (e.g., up to 100 layers) compared with standard collaborative filtering methods. However, the high network latency has become the main obstacle when deploying very deep recommender models into a production environment. In this paper, we argue that the typical prediction framework that treats all users equally during the inference phase is inefficient in running time, as well as sub-optimal in accuracy. To resolve such an issue, we present SkipRec, an adaptive inference framework by learning to skip inactive hidden layers on a per-user basis. Specifically, we devise a policy network to automatically determine which layers should be retained and which layers are allowed to be skipped, so as to achieve user-specific decisions. To derive the optimal skipping policy, we propose using gumbel softmax and reinforcement learning to solve the non-differentiable problem during backpropagation. We perform extensive experiments on three real-world recommendation datasets, and demonstrate that SkipRec attains comparable or better accuracy with much less inference time.
Lei Chen 0072, Fajie Yuan, Jiaxi Yang 0004, Xiang Ao 0001, Chengming Li 0004, Min Yang 0007
AAAI3