VLDB 2026 Research / reviewers in the wild / expert
Zexiong Ma
dblp:359/6950
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0000-2846-7399ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Progressively Mitigating API Hallucination in LLM-Generated Code via Knowledge Graph Reasoning
Zexiong Ma, Yanzhen Zou, Lihan Yang |
SANER | 2 |
| 2025 | SoRFT: Issue Resolving with Subtask-oriented Reinforced Fine-TuningabstractMainstream issue-resolving frameworks predominantly rely on commercial models, leading to high costs and privacy concerns.Existing training approaches for issue resolving struggle with poor generalization and fail to fully leverage open-source development resources.We propose Subtask-oriented Reinforced Fine-Tuning (SoRFT), a novel training approach to enhance the issue resolving capability of LLMs.We decomposes issue resolving into structured subtasks: file localization, function localization, line localization, and code edit generation.SoRFT consists of two training stages:(1) rejection-sampled supervised fine-tuning, Chain of Thought (CoT) data is filtered using ground-truth before fine-tuning the LLM, and (2) rule-based reinforcement learning, which leverages PPO with ground-truth based rewards.We evaluate the SoRFT-trained model on SWE-Bench Verified and SWE-Bench Lite, achieving state-of-the-art (SOTA) performance among open-source models (e.g., resolve 21.4% issues on SWE-Bench Verified with SoRFT-Qwen-7B).The experimental results demonstrate that SoRFT significantly enhances issue-resolving performance, improves model generalization, and provides a cost-efficient alternative to commercial models. Zexiong Ma, Chao Peng 0002, Xiangxin Meng, Yanzhen Zou |
ACL (1) | 1 |
| 2024 | Compositional API Recommendation for Library-Oriented Code GenerationabstractLarge language models (LLMs) have achieved exceptional performance in code generation. However, the performance remains unsatisfactory in generating library-oriented code, especially for the libraries not present in the training data of LLMs. Previous work utilizes API recommendation technology to help LLMs use libraries: it retrieves APIs related to the user requirements, then leverages them as context to prompt LLMs. However, developmental requirements can be coarse-grained, requiring a combination of multiple fine-grained APIs. This granularity inconsistency makes API recommendation a challenging task. Zexiong Ma, Shengnan An, Zeqi Lin |
ICPC | 1 |
| 2024 | Make Your LLM Fully Utilize the ContextabstractWhile many contemporary large language models (LLMs) can process lengthy input, they still struggle to fully utilize information within the long context, known as the *lost-in-the-middle* challenge.
We hypothesize that it stems from insufficient explicit supervision during the long-context training, which fails to emphasize that any position in a long context can hold crucial information.
Based on this intuition, our study presents **information-intensive (IN2) training**, a purely data-driven solution to overcome lost-in-the-middle.
Specifically, IN2 training leverages a synthesized long-context question-answer dataset, where the answer requires (1) **fine-grained information awareness** on a short segment (~128 tokens) within a synthesized long context (4K-32K tokens), and (2) the **integration and reasoning** of information from two or more short segments.
Through applying this information-intensive training on Mistral-7B, we present **FILM-7B** (FIll-in-the-Middle).
To thoroughly assess the ability of FILM-7B for utilizing long contexts, we design three probing tasks that encompass various context styles (document, code, and structured-data context) and information retrieval patterns (forward, backward, and bi-directional retrieval).
The probing results demonstrate that FILM-7B can robustly retrieve information from different positions in its 32K context window.
Beyond these probing tasks, FILM-7B significantly improves the performance on real-world long-context tasks (e.g., 23.5->26.9 F1 score on NarrativeQA), while maintaining a comparable performance on short-context tasks (e.g., 59.3->59.2 accuracy on MMLU). Shengnan An, Zexiong Ma, Zeqi Lin, Nanning Zheng 0001, Jian-Guang Lou, Weizhu Chen |
NeurIPS | 2 |