Yuheng Tang

dblp:287/8692 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
2 papers
Debugging and program repair · 94% Program synthesis and code generation · 6%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Debugging and program repair
automated program repair
1.722025
Co-PatcheR: Collaborative Software Patching with Component-specific Small Reasoning Models · NeurIPS 2025
PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification · ICML 2025
Debugging and program repair › automated program repair
patch validation
1.722025
Co-PatcheR: Collaborative Software Patching with Component-specific Small Reasoning Models · NeurIPS 2025
PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification · ICML 2025
Program synthesis and code generation
code generation with language models
0.312025
Co-PatcheR: Collaborative Software Patching with Component-specific Small Reasoning Models · NeurIPS 2025
Debugging and program repair
fault localization
0.312025
PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification · ICML 2025

Methods — techniques the papers use, named apart from their topics

small reasoning models · 0.9rule-based planning workflow · 0.9majority vote · 0.9large language model · 0.9critique-based generation · 0.9
YearPublicationVenuePosition
2025 PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification
abstract
Recent research builds various patching agents that combine large language models (LLMs) with non-ML tools and achieve promising results on the state-of-the-art (SOTA) software patching benchmark, SWE-bench. Based on how to determine the patching workflows, existing patching agents can be categorized as agent-based planning methods, which rely on LLMs for planning, and rule-based planning methods, which follow a pre-defined workflow. At a high level, agent-based planning methods achieve high patching performance but with a high cost and limited stability. Rule-based planning methods, on the other hand, are more stable and efficient but have key workflow limitations that compromise their patching performance. In this paper, we propose PatchPilot, an agentic patcher that strikes a balance between patching efficacy, stability, and cost-efficiency. PatchPilot proposes a novel rule-based planning workflow with five components: reproduction, localization, generation, validation, and refinement (where refinement is unique to PatchPilot). We introduce novel and customized designs to each component to optimize their effectiveness and efficiency. Through extensive experiments on the SWE-bench benchmarks, PatchPilot shows a superior performance than existing open-source methods while maintaining low cost (less than 1\$ per instance) and ensuring higher stability. We also conduct a detailed ablation study to validate the key designs in each component. Our code is available at https://github.com/ucsb-mlsec/PatchPilot.
Hongwei Li 0025, Yuheng Tang, Shiqi Wang 0032, Wenbo Guo 0002
ICML2
2025 SECODEPLT: A Unified Benchmark for Evaluating the Security Risks and Capabilities of Code GenAI
abstract
Existing benchmarks for evaluating the security risks and capabilities (e.g., vulnerability detection) of code-generating large language models (LLMs) face several key limitations:(1) limited coverage of risk and capabilities;(2) reliance on static evaluation metrics such as LLM judgments or rule-based detection, which lack the precision of dynamic analysis; and(3) a trade-off between data quality and benchmark scale.To address these challenges, we introduce a general and scalable benchmark construction framework that begins with manually validated, high-quality seed examples and expands them via targeted mutations.Each mutated sample retains the seed’s security semantics while providing diverse, unseen instances. The resulting benchmark bundles every artifact required for dynamic evaluation, including prompts, vulnerable and patched code, test cases, and ground-truth proofs of concept, enabling rigorous measurement of insecure coding, vulnerability detection, and patch generation. Applying this framework to Python, C/C++, and Java, we build SECODEPLT, a dataset of more than 5.9k samples spanning 44 CWE-based risk categories and three security capabilities. Compared with state-of-the-art benchmarks, SECODEPLT offers broader coverage, higher data fidelity, and substantially greater scale. We use SECODEPLT to evaluate leading code-generation LLMs and agents, revealing their strengths and weaknesses in both generating secure code and identifying or fixing vulnerabilities.We provide our code in \url{https://github.com/ucsb-mlsec/SeCodePLT}, data in \url{https://huggingface.co/datasets/UCSB-SURFI/SeCodePLT}
Yuzhou Nie, Zhun Wang, Yu Yang 0007, Ruizhe Jiang, Yuheng Tang, Xander Davies, Yarin Gal, Bo Li 0026, Wenbo Guo 0002, Dawn Song
NeurIPS5
2025 Co-PatcheR: Collaborative Software Patching with Component-specific Small Reasoning Models
abstract
Motivated by the success of general‑purpose large language models (LLMs) in software patching, recent works started to train specialized patching models. Most works trained one model to handle the end‑to‑end patching pipeline (including issue localization, patch generation, and patch validation). However, it is hard for a small model to handle all tasks, as different sub-tasks have different workflows and require different expertise. As such, by using a 70 billion model, SOTA methods can only reach up to 41% resolved rate on SWE-bench-Verified. Motivated by the collaborative nature, we propose Co-PatcheR, the first collaborative patching system with small and specialized reasoning models for individual components. Our key technique novelties are the specific task designs and training recipes. First, we train a model for localization and patch generation. Our localization pinpoints the suspicious lines through a two-step procedure, and our generation combines patch generation and critique. We then propose a hybrid patch validation that includes two models for crafting issue-reproducing test cases with and without assertions and judging patch correctness, followed by a majority vote-based patch selection. Through extensive evaluation, we show that Co-PatcheR achieves 46% resolved rate on SWE-bench-Verified with only 3 x 14B models. This makes Co-PatcheR the best patcher with specialized models, requiring the least training resources and the smallest models. We conduct a comprehensive ablation study to validate our recipes, as well as our choice of training data number, model size, and testing-phase scaling strategy.
Yuheng Tang, Hongwei Li 0025, Kaijie Zhu, Michael Yang, Yangruibo Ding, Wenbo Guo 0002
NeurIPS1