Wei Cheng 0010

dblp:89/2506-10 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0001-6128-7293ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Bootstrapping Code Translation with Weighted Multilanguage Exploration
abstract
Code translation across multiple programming languages is essential yet challenging due to two vital obstacles: scarcity of parallel data paired with executable test oracles, and optimization imbalance when handling diverse language pairs.We propose BootTrans, a bootstrapping method that resolves both obstacles.Its key idea is to leverage the functional invariance and cross-lingual portability of test suites, adapting abundant pivot-language unit tests to serve as universal verification oracles for multilingual reinforcement learning (RL) training.Our method introduces a dual-pool architecture with seed and exploration pools to progressively expand training data via executionguided experience collection.Furthermore, we design a language-aware weighting mechanism that dynamically prioritizes harder translation directions based on relative performance across sibling languages, mitigating optimization imbalance.Extensive experiments on the HumanEval-X and TransCoder-Test benchmarks demonstrate substantial improvements over baseline LLMs across all translation directions, with ablation studies validating the effectiveness of both bootstrapping and weighting components.
Yuhan Wu 0006, Huan Zhang 0019, Wei Cheng 0010, Jingyue Yang, Wei Hu 0007
ACL (1)3
2026 Self-Improving Code Generation via Semantic Entropy and Behavioral Consensus
abstract
Improving the code generation capabilities of large language models (LLMs) typically relies on supervised fine-tuning or preference optimization, both of which require costly external resources such as powerful teacher models or reliable test units. However, in real-world scenarios, it is much harder to obtain reference solutions and test oracles than problem descriptions and test inputs. In this paper, we tackle a challenging yet realistic question: Can a code language model improve itself without access to a superior teacher and a test oracle? To answer this, we propose ConSelf, a self-improving approach built upon two key ideas. First, we introduce code semantic entropy, a novel metric that measures problem-level uncertainty by assessing the functional diversity of program behaviors, enabling a curriculum construction with the most learnable problems. Second, we present consensus-driven direct preference optimization (Con-DPO), a preference-based fine-tuning method that weights each preference pair by its behavioral consensus, thereby mitigating the impact of noisy self-generated supervision. Experiments on various benchmarks and backbone LLMs demonstrate that ConSelf significantly outperforms baselines, validating the effectiveness of semantic entropy-based curriculum construction and consensus-driven optimization in improving code generation without external supervision.
Huan Zhang 0019, Wei Cheng 0010, Wei Hu 0007
ICPC2
2024 Dataflow-Guided Retrieval Augmentation for Repository-Level Code Completion
abstract
Recent years have witnessed the deployment of code language models (LMs) in various code intelligence tasks such as code completion.Yet, it is challenging for pre-trained LMs to generate correct completions in private repositories.Previous studies retrieve cross-file context based on import relations or text similarity, which is insufficiently relevant to completion targets.In this paper, we propose a dataflow-guided retrieval augmentation approach, called DRACO, for repository-level code completion.DRACO parses a private repository into code entities and establishes their relations through an extended dataflow analysis, forming a repo-specific context graph.Whenever triggering code completion, DRACO precisely retrieves relevant background knowledge from the repo-specific context graph and generates well-formed prompts to query code LMs.Furthermore, we construct a large Python dataset, ReccEval, with more diverse completion targets.Our experiments demonstrate the superior accuracy and applicable efficiency of DRACO, improving code exact match by 3.43% and identifier F1-score by 3.27% on average compared to the state-ofthe-art approach.
Wei Cheng 0010, Yuhan Wu 0006, Wei Hu 0007
ACL (1)1
2024 A Pair Programming Framework for Code Generation via Multi-Plan Exploration and Feedback-Driven Refinement
abstract
Large language models (LLMs) have achieved impressive performance on code generation. Although prior studies enhanced LLMs with prompting techniques and code refinement, they still struggle with complex programming problems due to rigid solution plans. In this paper, we draw on pair programming practices to propose PairCoder, a novel LLM-based framework for code generation. PairCoder incorporates two collaborative LLM agents, namely a Navigator agent for high-level planning and a Driver agent for specific implementation. The Navigator is responsible for proposing promising solution plans, selecting the current optimal plan, and directing the next iteration round based on execution feedback. The Driver follows the guidance of Navigator to undertake initial code generation, code testing, and refinement. This interleaved and iterative workflow involves multi-plan exploration and feedback-based refinement, which mimics the collaboration of pair programmers. We evaluate PairCoder with both open-source and closed-source LLMs on various code generation benchmarks. Extensive experimental results demonstrate the superior accuracy of PairCoder, achieving relative pass@1 improvements of 12.00%--162.43% compared to prompting LLMs directly.
Huan Zhang 0019, Wei Cheng 0010, Yuhan Wu 0006, Wei Hu 0007
ASE2
2024 Revisiting Knowledge-Based Inference of Python Runtime Environments: A Realistic and Adaptive Approach
abstract
The reuse and integration of existing code is a common practice for efficient software development. Constantly updated Python interpreters and third-party packages introduce many challenges to Python runtime environment inference. Existing knowledge-based approaches have achieved good performance but still suffer from several limitations in the real world, especially from incomplete domain knowledge. In this paper, we propose ReadPyE, a realistic and adaptive approach to Python runtime environment inference. To leverage the rich code information, we present an automated approach to the construction and maintenance of our designed Python ecosystem knowledge graph (KG). Moreover, we are the first to handle real-world challenges such as complex dependency specifications and incomplete domain knowledge. Specifically, we define a naming similarity measure to match candidate packages for unknown modules and set priorities for multiple candidate packages. ReadPyE solves the optimization problems of candidate package selection and generates compatible runtime environments step by step based on the current Python environment. The inferred environments are iteratively validated and adjusted by matched exception templates in the validation logs. The evaluation results on three real-world datasets show the superior effectiveness and good efficiency of our ReadPyE compared to the existing knowledge-based approaches. ReadPyE solves the environment-related exceptions for 79.75% single-file code snippets, 93% Python projects, and 63.34% program pairs for code integration. We believe ReadPyE can help programmers reduce the time spent on inferring Python runtime environments and facilitate automated software configuration management.
Wei Cheng 0010, Wei Hu 0007, Xiaoxing Ma
IEEE Trans. Software Eng.1
2022 Conflict-aware Inference of Python Compatible Runtime Environments with Domain Knowledge Graph
abstract
Code sharing and reuse is a widespread use practice in software engineering. Although a vast amount of open-source Python code is accessible on many online platforms, programmers often find it difficult to restore a successful runtime environment. Previous studies validated automatic inference of Python dependencies using pre-built knowledge bases. However, these studies do not cover sufficient knowledge to accurately match the Python code and also ignore the potential conflicts between their inferred dependencies, thus resulting in a low success rate of inference. In this paper, we propose PyCRE, a new approach to automatically inferring Python compatible runtime environments with domain knowledge graph (KG). Specifically, we design a domain-specific ontology for Python third-party packages and construct KGs for over 10,000 popular packages in Python 2 and Python 3. PyCRE discovers candidate libraries by measuring the matching degree between the known libraries and the third-party resources used in target code. For the NP-complete problem of dependency solving, we propose a heuristic graph traversal algorithm to efficiently guarantee the compatibility between packages. PyCRE achieves superior performance on a real-world dataset and efficiently resolves nearly half more import errors than previous methods.
Wei Cheng 0010, Xiangrong Zhu 0001, Wei Hu 0007
ICSE1