Haoxuan Chen

dblp:212/7201 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Learning generalizable visual representations with causal diffusion model for controllable editing
Shanshan Huang 0004, Lei Wang 0197, Haoxuan Chen, Yuxuan Liang 0002, Li Liu 0001
Pattern Recognit.3
2025 How Do Developers Use ChatGPT for Software Test Generation? Usage Patterns, Satisfaction, and Code Adoption Analysis
abstract
Large Language Models (LLMs) such as ChatGPT are increasingly used by developers to generate test code, yet little is known about their real-world usage patterns, developer satisfaction, and the extent to which generated code is integrated into software projects. This paper presents an empirical study of 407 user-ChatGPT conversations related to test generation from GitHub, including 152 that contain both test and corresponding source code. We systematically categorize the primary use cases of ChatGPT in test generation, employ sentiment analysis to assess developer satisfaction across different testing scenarios, and analyze the adoption between ChatGPT-generated test code and existing project code using PMD’s Copy-Paste Detector (CPD) tool. Our analysis identifies six ChatGPT usage scenarios in test generation: test case creation, framework demonstration, test enhancement, debugging, CI/CD integration, and mocking solution. Among these, test case creation and framework demonstration emerge as the most frequent. Sentiment analysis reveals that, except for framework demonstration, all other scenarios exhibit a higher proportion of negative feedback than positive. Moreover, adoption analysis indicates that only 25.7% of conversations show evidence of code reuse, suggesting developers are generally cautious when integrating LLM-generated test code. Based on these findings, we propose several practical implications for enhancing the effectiveness of LLM-assisted test generation.
Wanwei Zhan, Junjie Dai, Haoxuan Chen
APSEC5
2025 How Discrete and Continuous Diffusion Meet: Comprehensive Analysis of Discrete Diffusion Models via a Stochastic Integral Framework
abstract
Discrete diffusion models have gained increasing attention for their ability to model complex distributions with tractable sampling and inference. However, the error analysis for discrete diffusion models remains less well-understood. In this work, we propose a comprehensive framework for the error analysis of discrete diffusion models based on Lévy-type stochastic integrals. By generalizing the Poisson random measure to that with a time-independent and state-dependent intensity, we rigorously establish a stochastic integral formulation of discrete diffusion models and provide the corresponding change of measure theorems that are intriguingly analogous to Itô integrals and Girsanov's theorem for their continuous counterparts. Our framework unifies and strengthens the current theoretical results on discrete diffusion models and obtains the first error bound for the $\tau$-leaping scheme in KL divergence. With error sources clearly identified, our analysis gives new insight into the mathematical properties of discrete diffusion models and offers guidance for the design of efficient and accurate algorithms for real-world discrete diffusion model applications.
Yinuo Ren, Haoxuan Chen, Grant M. Rotskoff, Lexing Ying
ICLR2
2025 Collaborative Orchestration with Probabilistic Routing for Dynamic Service Mesh in Clouds
Haonan Ding, Haoxuan Chen, Jianwen He, Menglan Hu, Chao Cai 0001, Kai Peng 0001
INFOCOM3
2025 RealisticCodeBench: Towards More Realistic Evaluation of Large Language Models for Code Generation
abstract
Evaluating the code generation capabilities of Large Language Models (LLMs) remains an open question. Recently, more advanced benchmarks—such as CoderEval, EvoCodeBench, and ClassEval—have been introduced to evaluate LLMs on practical coding tasks from GitHub repositories, such as non-standalone function generation and class-level code generation. However, even the most sophisticated LLMs struggle with these complex tasks; for instance, GPT-4 achieves only a 37.0% pass@1 on ClassEval. Prior studies show that developers often discard LLM-generated code or abandon code generation models when outputs are incorrect or require extensive debugging, which leads them to rely on LLMs primarily for code generation tasks that high-performing models can reliably handle.In response to this gap, we introduce RealisticCodeBench, a benchmark specifically designed to reflect the types of problems developers commonly tackle with LLMs. By mining GitHub repositories for code samples tagged as generated by ChatGPT or Copilot, we collect real-world coding tasks that capture typical LLM usage scenarios. We modify these tasks, generate reference solutions and test cases, and adapt the problems into multiple programming languages. This effort results in RealisticCodeBench, comprising a total of 376 programming problems translated across multiple languages: 361 in Python, 346 in JavaScript, 343 in TypeScript, 307 in Java, and 323 in C++, each with corresponding reference solutions and test cases. We evaluate 12 general-purpose and code-specific LLMs on RealisticCodeBench. Our findings reveal that GPT-4.1 achieves the highest average pass@1 score across languages, closely followed by DeepSeek-V3-671B, suggesting that DeepSeek-V3-671B provides a viable open-source alternative to GPT-4.1 for large companies with sufficient GPU resources and privacy concerns. CodeGeeX4-9B, a cost-effective model, emerges as a suitable substitute for GPT-4o-mini for individual developers and smaller organizations with similar privacy considerations. Additionally, LLM performance discrepancies between HumanEval and RealisticCodeBench suggest that some LLMs are either overly specialized for HumanEval-style problems or insufficiently optimized for real-world coding challenges. Finally, we analyze failed cases, summarize common LLM limitations, and provide implications for researchers and practitioners.
Xiao Yu 0008, Haoxuan Chen, Lei Liu 0062, Xing Hu 0008, Jacky W. Keung, Xin Xia 0001
ASE2
2025 Fast Solvers for Discrete Diffusion Models: Theory and Applications of High-Order Algorithms
abstract
Discrete diffusion models have emerged as a powerful generative modeling framework for discrete data with successful applications spanning from text generation to image synthesis. However, their deployment faces challenges due to the high dimensionality of the state space, necessitating the development of efficient inference algorithms. Current inference approaches mainly fall into two categories: exact simulation and approximate methods such as $\tau$-leaping. While exact methods suffer from unpredictable inference time and redundant function evaluations, $\tau$-leaping is limited by its first-order accuracy. In this work, we advance the latter category by tailoring the first extension of high-order numerical inference schemes to discrete diffusion models, enabling larger step sizes while reducing error. We rigorously analyze the proposed schemes and establish the second-order accuracy of the $\theta$-Trapezoidal method in KL divergence. Empirical evaluations on GSM8K-level math-reasoning, GPT-2-level text, and ImageNet-level image generation tasks demonstrate that our method achieves superior sample quality compared to existing approaches under equivalent computational constraints, with consistent performance gains across models ranging from 200M to 8B. Our code is available at https://github.com/yuchen-zhu-zyc/DiscreteFastSolver
Yinuo Ren, Haoxuan Chen, Grant M. Rotskoff, Molei Tao, Lexing Ying
NeurIPS2
2025 The impact of feature selection and feature reduction techniques for code smell detection: A comprehensive empirical study
Zexian Zhang, Shuang Yin, Haoxuan Chen
Autom. Softw. Eng.6
2025 Large-Scale Service Mesh Orchestration With Probabilistic Routing in Cloud Data Centers
abstract
Service mesh architectures are emerging as a promising microservice paradigm for developing online cloud applications. However, in large-scale microservice scenarios, frequent service communications, intricate call dependencies, and stringent latency requirements bring great pressure to efficient service mesh orchestration. In this case, the problems of service deployment and request routing based on service mesh architectures are tightly-coupled and interdependent, and cannot be effectively optimized individually, enlarging the difficulty for collaborative orchestration. When microservice multiplexing, parallel dependencies, and multi-instance modeling are considered, the difficulty is further aggravated. Nonetheless, most existing work failed to propose appropriate models and methods for the above challenges. Therefore, this article studies the large-scale service mesh orchestration with probabilistic routing and constrained bandwidths for parallel call graphs. We leverage the open Jackson queuing network theory to capture crucial microservices and analyze request processing, queuing, and communication latency for massive user requests in a fine-grained way. Then, this article proposes an efficient three-stage heuristic, which achieves elegant multi-instance consolidation and probabilistic multi-queue routing to reduce response latency and cost. We also provide the algorithm complexity and mathematical analysis of the performance. Finally, extensive trace-driven experiments are performed to validate the superiority of our proposed algorithm over other baselines.
Kai Peng 0001, Haonan Ding, Haoxuan Chen, Liangyuan Wang, Chao Cai 0001, Menglan Hu
IEEE Trans. Serv. Comput.4
2024 On the Relative Value of Feature Selection Techniques for Code Smell Detection
abstract
Machine/deep learning-based code smell detection aims to develop a classification model based on code smell features to predict the presence of code smell in new code instances. To ensure accurate detection, it is crucial to eliminate irrelevant or redundant features that may negatively impact performance. Previous studies have produced inconsistent findings about the impact of feature selection techniques for code smell detection, possibly because they examined only a limited number of different techniques. To address this gap, our study aims to provide a comprehensive analysis of feature selection techniques in code smell detection. We investigate 34 feature selection techniques with 7 classification models to build the code smell detection models on 6 code smell datasets. To assess these effects, we use 3 evaluation metrics, i.e., Precision, Recall, and F-measure, and compare the performance differences using the Scott-Knott effect size difference test and the McNemar's test. The results show that (1) Not all feature selection techniques significantly improve detection performance. The techniques with better performance are chi-square, probabilistic significance, information gain, and symmetrical uncertainty. (2) In general, probabilistic significance should be used as the “generic” feature selection technique because detection models using probabilistic significance can identify more of the same smelly instances compared to models using other methods. (3) The high-frequency features selected by the four highest-performing techniques, which are important for identifying the corresponding code smells, are different for each dataset.
Zexian Zhang, Shuang Yin, Haoxuan Chen
APSEC5
2024 Accelerating Diffusion Models with Parallel Sampling: Inference at Sub-Linear Time Complexity
abstract
Diffusion models have become a leading method for generative modeling of both image and scientific data. As these models are costly to train and \emph{evaluate}, reducing the inference cost for diffusion models remains a major goal. Inspired by the recent empirical success in accelerating diffusion models via the parallel sampling technique~\cite{shih2024parallel}, we propose to divide the sampling process into $\mathcal{O}(1)$ blocks with parallelizable Picard iterations within each block. Rigorous theoretical analysis reveals that our algorithm achieves $\widetilde{\mathcal{O}}(\mathrm{poly} \log d)$ overall time complexity, marking \emph{the first implementation with provable sub-linear complexity w.r.t. the data dimension $d$}. Our analysis is based on a generalized version of Girsanov's theorem and is compatible with both the SDE and probability flow ODE implementations. Our results shed light on the potential of fast and efficient sampling of high-dimensional data on fast-evolving modern large-memory GPU clusters.
Haoxuan Chen, Yinuo Ren, Lexing Ying, Grant M. Rotskoff
NeurIPS1
2023 When can Regression-Adjusted Control Variate Help? Rare Events, Sobolev Embedding and Minimax Optimality
abstract
This paper studies the use of a machine learning-based estimator as a control variate for mitigating the variance of Monte Carlo sampling. Specifically, we seek to uncover the key factors that influence the efficiency of control variates in reducing variance. We examine a prototype estimation problem that involves simulating the moments of a Sobolev function based on observations obtained from (random) quadrature nodes. Firstly, we establish an information-theoretic lower bound for the problem. We then study a specific quadrature rule that employs a nonparametric regression-adjusted control variate to reduce the variance of the Monte Carlo simulation. We demonstrate that this kind of quadrature rule can improve the Monte Carlo rate and achieve the minimax optimal rate under a sufficient smoothness assumption. Due to the Sobolev Embedding Theorem, the sufficient smoothness assumption eliminates the existence of rare and extreme events. Finally, we show that, in the presence of rare and extreme events, a truncated version of the Monte Carlo algorithm can achieve the minimax optimal rate while the control variate cannot improve the convergence rate.
Jose H. Blanchet, Haoxuan Chen, Yiping Lu 0001, Lexing Ying
NeurIPS2
2022 Machine Learning For Elliptic PDEs: Fast Rate Generalization Bound, Neural Scaling Law and Minimax Optimality
Yiping Lu 0001, Haoxuan Chen, Jianfeng Lu 0001, Lexing Ying, Jose H. Blanchet
ICLR2
2022 Preference-oriented partitioning for multiprocessor real-time systems
Qin Xia, Songming Yan, Haoxuan Chen, Dakai Zhu 0001, Hakan Aydin
J. Syst. Archit.3
2019 Modulation and Demodulation of Digital Frequency Shift Keying System Based on Spin Torque Nano Oscillator with Voltage Controlled Magnetic Anisotropy Effect
abstract
In this work, a spin torque nano oscillator (STNO) device whose frequency can be tuned by Voltage Controlled Magnetic Anisotropy effect (VCMA) is proposed. The requirement of magnetic bias field in previous STNO devices is eliminated by the introduction of VCMA effect. Based on VCMA-STNO, a novel architecture is proposed which can compose of a modulation/demodulation digital frequency shift keying (DFSK) communication system. The proposed architecture utilizes VCMA-STNO as core devices and is much simpler comparing with its CMOS counterpart. The proposed VCMA-STNO modulation/demodulation architecture will help to design next generation spintronics DFSK communication system.
Lang Zeng, Zuodong Zhang, Haoxuan Chen, Tianqi Gao, Deming Zhang, Mingzhi Long, Youguang Zhang, Weisheng Zhao 0001
ISCAS3