Xiaofei Ma 0001

dblp:120/1069-1 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
12since 2021 · last 2025
0009-0000-3163-0310ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 11 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 LibEvolutionEval: A Benchmark and Study for Version-Specific Code Generation
abstract
Sachit Kuhar, Wasi Uddin Ahmad, Zijian Wang, Nihal Jain, Haifeng Qian, Baishakhi Ray, Murali Krishna Ramanathan, Xiaofei Ma, Anoop Deoras. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Sachit Kuhar, Wasi Uddin Ahmad, Zijian Wang 0002, Nihal Jain, Haifeng Qian, Baishakhi Ray, Murali Krishna Ramanathan, Xiaofei Ma 0001, Anoop Deoras
NAACL (Long Papers)8
2025 Approximately Aligned Decoding
abstract
It is common to reject undesired outputs of Large Language Models (LLMs); however, current methods to do so require an excessive amount of computation to re-sample after a rejection, or distort the distribution of outputs by constraining the output to highly improbable tokens. We present a method, Approximately Aligned Decoding (AprAD), to balance the distortion of the output distribution with computational efficiency, inspired by algorithms from the speculative decoding literature. AprAD allows for the generation of long sequences of text with difficult-to-satisfy constraints, while amplifying low probability outputs much less compared to existing methods. We show through a series of experiments that the task-specific performance of AprAD is comparable to methods that do not distort the output distribution, while being much more computationally efficient.
Daniel Melcer, Sujan K. Gonugondla, Pramuditha Perera, Haifeng Qian, Wen-Hao Chiang, Nihal Jain, Pranav Garg 0001, Xiaofei Ma 0001, Anoop Deoras
NeurIPS9
2025 UTFix: Change Aware Unit Test Repairing using LLM
abstract
Software updates, including bug repair and feature additions, are frequent in modern applications but they often leave test suites outdated, resulting in undetected bugs and increased chances of system failures. A recent study by Meta revealed that 14%-22% of software failures stem from outdated tests that fail to reflect changes in the codebase. This highlights the need to keep tests in sync with code changes to ensure software reliability. In this paper, we present UTFix , a novel approach for repairing unit tests when their corresponding focal methods undergo changes. UTFix addresses two critical issues: assertion failure and reduced code coverage caused by changes in the focal method. Our approach leverages language models to repair unit tests by providing contextual information such as static code slices, dynamic code slices, and failure messages. We evaluate UTFix on our generated synthetic benchmark (Syn-Bench), and real-world benchmark. In our experiment, UTFix successfully repaired 89.2% of assertion failures and achieved 100% code coverage for 96 tests out of 369 unit tests. On the real-world benchmarks, UTFix repaired 60% of assertion failures while achieving 100% code coverage for 19 out of 30 unit tests. To the best of our knowledge, this is the first comprehensive study focused on unit test in evolving Python projects. Our contributions include the development of UTFix , the creation of Syn-Bench and real-world benchmarks, and the demonstration of the effectiveness of LLM-based methods in addressing unit test failures due to software evolution.
Shanto Rahman, Sachit Kuhar, Berk Çirisci, Pranav Garg 0001, Shiqi Wang 0002, Xiaofei Ma 0001, Anoop Deoras, Baishakhi Ray
Proc. ACM Program. Lang.6
2024 Lightweight reranking for language model generations
abstract
Large Language Models (LLMs) can exhibit considerable variation in the quality of their sampled outputs.Reranking and selecting the best generation from the sampled set is a popular way of obtaining strong gains in generation quality.In this paper, we present a novel approach for reranking LLM generations.Unlike other techniques that might involve additional inferences or training a specialized reranker, our approach relies on easy to compute pairwise statistics between the generations that have minimal compute overhead.We show that our approach can be formalized as an extension of self-consistency and analyze its performance in that framework, theoretically as well as via simulations.We show strong improvements for selecting the best k generations for code generation tasks as well as robust improvements for the best generation for the tasks of autoformalization, summarization, and translation.While our approach only assumes black-box access to LLMs, we show that additional access to token probabilities can improve performance even further.
Siddhartha Jain 0001, Xiaofei Ma 0001, Anoop Deoras, Bing Xiang
ACL (1)2
2024 Code Representation Learning at Scale
abstract
Recent studies have shown that code language model at scale demonstrate significant performance gains on downstream tasks, i.e., code generation. However, most of the existing works on code representation learning train models at a hundred million parameter scale using very limited pretraining corpora. In this work, we fuel code representation learning with a vast amount of code data via a two-stage pretraining scheme. We first train the encoders via a mix that leverages both randomness in masking language modeling and implicit structure and semantic aspects of programming language. We then enhance the representations via contrastive learning with hard negative and hard positive constructed in an unsupervised manner. We establish an off-the-shelf encoder model that persistently outperforms the existing models on a wide variety of downstream tasks by large margins. To comprehend the factors contributing to successful code representation learning, we conduct detailed ablations and share our findings on (i) a customized and effective token-level denoising scheme for source code; (ii) the importance of hard negatives and hard positives; (iii) how the proposed bimodal contrastive learning boost the cross-lingual semantic search performance; and (iv) how the pretraining schemes decide the downstream task performance scales with the model size.
Dejiao Zhang, Wasi Uddin Ahmad, Hantian Ding, Ramesh Nallapati, Dan Roth 0001, Xiaofei Ma 0001, Bing Xiang
ICLR7
2024 Repoformer: Selective Retrieval for Repository-Level Code Completion
abstract
Recent advances in retrieval-augmented generation (RAG) have initiated a new era in repository-level code completion. However, the invariable use of retrieval in existing methods exposes issues in both efficiency and robustness, with a large proportion of the retrieved contexts proving unhelpful or harmful to code language models (code LMs). In this paper, we propose a selective RAG framework to avoid retrieval when unnecessary. To power this framework, we design a self-supervised learning approach to enable a code LM to accurately self-evaluate whether retrieval can improve its output quality and robustly leverage the potentially noisy retrieved contexts. Using this LM as both the selective RAG policy and the generation model, our framework achieves state-of-the-art repository-level code completion performance on diverse benchmarks including RepoEval, CrossCodeEval, and CrossCodeLongEval, a new long-form code completion benchmark. Meanwhile, our analyses show that selectively retrieving brings as much as 70% inference speedup in the online serving setting without harming the performance. We further demonstrate that our framework is able to accommodate different generation models, retrievers, and programming languages. These advancements position our framework as an important step towards more accurate and efficient repository-level code completion.
Di Wu 0054, Wasi Uddin Ahmad, Dejiao Zhang, Murali Krishna Ramanathan, Xiaofei Ma 0001
ICML5
2024 LeDex: Training LLMs to Better Self-Debug and Explain Code
abstract
In the domain of code generation, self-debugging is crucial. It allows LLMs to refine their generated code based on execution feedback. This is particularly important because generating correct solutions in one attempt proves challenging for complex tasks. Prior works on self-debugging mostly focus on prompting methods by providing LLMs with few-shot examples, which work poorly on small open-sourced LLMs. In this work, we propose LeDex, a training framework that significantly improves the self-debugging capability of LLMs. Intuitively, we observe that a chain of explanations on the wrong code followed by code refinement helps LLMs better analyze the wrong code and do refinement. We thus propose an automated pipeline to collect a high-quality dataset for code explanation and refinement by generating a number of explanations and refinement trajectories from the LLM itself or a larger teacher model and filtering via execution verification. We perform supervised fine-tuning (SFT) and further reinforcement learning (RL) on both success and failure trajectories with a novel reward design considering code explanation and refinement quality. SFT improves the pass@1 by up to 15.92\% and pass@10 by 9.30\% over four benchmarks. RL training brings additional up to 3.54\% improvement on pass@1 and 2.55\% improvement on pass@10. The trained LLMs show iterative refinement ability and can keep refining code continuously. Lastly, our human evaluation shows that the LLMs trained with our framework generate more useful code explanations and help developers better understand bugs in source code.
Nan Jiang 0012, Xiaopeng Li 0002, Shiqi Wang 0002, Qiang Zhou 0009, Soneya Binta Hossain, Baishakhi Ray, Xiaofei Ma 0001, Anoop Deoras
NeurIPS8
2023 Multitask Pretraining with Structured Knowledge for Text-to-SQL Generation
abstract
Robert Giaquinto, Dejiao Zhang, Benjamin Kleiner, Yang Li, Ming Tan, Parminder Bhatia, Ramesh Nallapati, Xiaofei Ma. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Robert Giaquinto, Dejiao Zhang, Benjamin Kleiner, Parminder Bhatia, Ramesh Nallapati, Xiaofei Ma 0001
ACL (1)8
2023 ContraCLM: Contrastive Learning For Causal Language Model
abstract
Nihal Jain, Dejiao Zhang, Wasi Uddin Ahmad, Zijian Wang, Feng Nan, Xiaopeng Li, Ming Tan, Ramesh Nallapati, Baishakhi Ray, Parminder Bhatia, Xiaofei Ma, Bing Xiang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Nihal Jain, Dejiao Zhang, Wasi Uddin Ahmad, Zijian Wang 0002, Feng Nan, Xiaopeng Li 0002, Ramesh Nallapati, Baishakhi Ray, Parminder Bhatia, Xiaofei Ma 0001, Bing Xiang
ACL (1)11
2023 Efficient Shapley Values Estimation by Amortization for Text Classification
abstract
Despite the popularity of Shapley Values in explaining neural text classification models, computing them is prohibitive for large pretrained models due to a large number of model evaluations.In practice, Shapley Values are often estimated with a small number of stochastic model evaluations.However, we show that the estimated Shapley Values are sensitive to random seed choices -the top-ranked features often have little overlap across different seeds, especially on examples with longer input texts.This can only be mitigated by aggregating thousands of model evaluations, which on the other hand, induces substantial computational overheads.To mitigate the trade-off between stability and efficiency, we develop an amortized model that directly predicts each input feature's Shapley Value without additional model evaluations.It is trained on a set of examples whose Shapley Values are estimated from a large number of model evaluations to ensure stability.Experimental results on two text classification datasets demonstrate that our amortized model estimates Shapley Values accurately with up to 60 times speedup compared to traditional methods.Furthermore, the estimated values are stable as the inference is deterministic.We release our code at https://github.com/yangalan123/ Amortized-Interpretability.
Chenghao Yang 0001, Fan Yin, He He 0001, Kai-Wei Chang 0001, Xiaofei Ma 0001, Bing Xiang
ACL (1)5
2023 STREET: A Multi-Task Structured Reasoning and Explanation Benchmark
Danilo Neves Ribeiro, Shen Wang 0005, Xiaofei Ma 0001, Henghui Zhu, Deguang Kong, Juliette Burger, Anjelica Ramos, Zhiheng Huang, William Yang Wang, George Karypis, Bing Xiang, Dan Roth 0001
ICLR3
2022 Learning Dialogue Representations from Consecutive Utterances
abstract
Zhihan Zhou, Dejiao Zhang, Wei Xiao, Nicholas Dingwall, Xiaofei Ma, Andrew Arnold, Bing Xiang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Zhihan Zhou 0001, Dejiao Zhang, Wei Xiao 0001, Nicholas Dingwall, Xiaofei Ma 0001, Andrew O. Arnold, Bing Xiang
NAACL-HLT5
2020 Beyond [CLS] through Ranking by Generation
abstract
Generative models for Information Retrieval, where ranking of documents is viewed as the task of generating a query from a document's language model, were very successful in various IR tasks in the past.However, with the advent of modern deep neural networks, attention has shifted to discriminative ranking functions that model the semantic similarity of documents and queries instead.Recently, deep generative models such as GPT2 and BART have been shown to be excellent text generators, but their effectiveness as rankers have not been demonstrated yet.In this work, we revisit the generative framework for information retrieval and show that our generative approaches are as effective as state-of-the-art semantic similarity-based discriminative models for the answer selection task.Additionally, we demonstrate the effectiveness of unlikelihood losses for IR.
Cícero Nogueira dos Santos, Xiaofei Ma 0001, Ramesh Nallapati, Zhiheng Huang, Bing Xiang
EMNLP (1)2
2019 Multi-passage BERT: A Globally Normalized BERT Model for Open-domain Question Answering
abstract
Zhiguo Wang, Patrick Ng, Xiaofei Ma, Ramesh Nallapati, Bing Xiang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Zhiguo Wang 0006, Patrick Ng, Xiaofei Ma 0001, Ramesh Nallapati, Bing Xiang
EMNLP/IJCNLP (1)3