Runzhe Wang

dblp:153/0092 · DBLP profile ↗
← Back
23ranked-venue papers
3as first author
20since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Software engineering, systems software and programming languages · 8 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2027 FMMCRec: Exploring where you'll go, across spatio-temporal and frequency domains
Sibo Wen, Nan Wang 0024, Runzhe Wang, Yingli Zhong
Expert Syst. Appl.4
2026 Multi-domain Denoising for Attribute-Aware Sequential Recommendation
Pinchao Zhou, Nan Wang 0024, Yingli Zhong, Runzhe Wang
DASFAA (5)4
2026 An interpretable knowledge recommendation method for civil dispute mediation
Shibo Cui, Runzhe Wang, Yongping Yu
Data Knowl. Eng.4
2026 FPS: Frequency-aware polynomial spectral reconstruction for dual-domain learning in long-term time series forecasting
Sibo Wen, Nan Wang 0024, Runzhe Wang, Yingli Zhong
Inf. Sci.4
2025 BAG-RAG: Bidirectional Retrieval-Augmented Generation Based on Multi-Layer Semantic Graphs for Budget Auditing QA
Runzhe Wang, Guilin Qi, Xiaolong Ye, Yongrui Chen 0002, Xinbang Dai, Shenwen Zhong
DASFAA (6)2
2025 Towards Understanding Text Hallucination of Diffusion Models via Local Generation Bias
abstract
Score-based diffusion models have achieved incredible performance in generating realistic images, audio, and video data. While these models produce high-quality samples with impressive details, they often introduce unrealistic artifacts, such as distorted fingers or hallucinated texts with no meaning. This paper focuses on textual hallucinations, where diffusion models correctly generate individual symbols but assemble them in a nonsensical manner. Through experimental probing, we consistently observe that such phenomenon is attributed it to the network's local generation bias. Denoising networks tend to produce outputs that rely heavily on highly correlated local regions, particularly when different dimensions of the data distribution are nearly pairwise independent. This behavior leads to a generation process that decomposes the global distribution into separate, independent distributions for each symbol, ultimately failing to capture the global structure, including underlying grammar. Intriguingly, this bias persists across various denoising network architectures including MLP and transformers which have the structure to model global dependency. These findings also provide insights into understanding other types of hallucinations, extending beyond text, as a result of implicit biases in the denoising models. Additionally, we theoretically analyze the training dynamics for a specific case involving a two-layer MLP learning parity points on a hypercube, offering an explanation of its underlying mechanism.
Rui Lu 0001, Runzhe Wang, Kaifeng Lyu, Xitai Jiang, Gao Huang 0001, Mengdi Wang 0001
ICLR2
2025 MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
abstract
Large language models have demonstrated impressive performance on challenging mathematical reasoning tasks, which has triggered the discussion of whether the performance is achieved by true reasoning capability or memorization. To investigate this question, prior work has constructed mathematical benchmarks when questions undergo simple perturbations – modifications that still preserve the underlying reasoning patterns of the solutions. However, no work has explored hard perturbations, which fundamentally change the nature of the problem so that the original solution steps do not apply. To bridge the gap, we construct MATH-P-Simple and MATH-P-Hard via simple perturbation and hard perturbation, respectively. Each consists of 279 perturbed math problems derived from level-5 (hardest) problems in the MATH dataset (Hendrycks et al., 2021). We observe significant performance drops on MATH-P-Hard across various models, including o1-mini (-16.49%) and gemini-2.0-flash-thinking (-12.9%). We also raise concerns about a novel form of memorization where models blindly apply learned problem-solving skills without assessing their applicability to modified contexts. This issue is amplified when using original problems for in-context learning. We call for research efforts to address this challenge, which is critical for developing more robust and reliable reasoning models. The project is available at https://math-perturb.github.io/.
Kaixuan Huang, Jiacheng Guo, Jiawei Ge 0003, Tianle Cai, Hui Yuan 0002, Runzhe Wang, Ming Yin 0003, Shange Tang, Yangsibo Huang, Chi Jin 0001, Chiyuan Zhang, Mengdi Wang 0001
ICML10
2025 RPG: Linux Kernel Fuzzing Guided by Distribution-Specific Runtime Parameter Interfaces
abstract
The Linux distribution kernel differs significantly from the mainline kernel, incorporating additional features and vendor-specific extensions. Among these additions, many runtime parameter interfaces are unique to distribution kernels, which expands the attack surface and increases the risk of potential vulnerabilities. Fuzzing has been used to assess Linux distributions, but existing tools cannot systematically test these distribution-specific interfaces due to two main challenges: (1) generating test cases for these runtime parameter interfaces, and (2) concentrating test resources on the distribution-specific interface code. To address these challenges, we propose RPG, a distribution-specific runtime parameter-guided kernel fuzzer. RPG operates in three phases: First, RPG extracts distribution-specific runtime parameter interfaces. Then, RPG uses LLM and tuning software databases to model each parameter range to generate meaningful interface test cases. Third, RPG utilizes the distribution kernel’s function control flow graph to guide the fuzzer to generate generic test cases that are more closely related to the distribution-specific interface code. We evaluated RPG on four Linux distribution kernels: Ubuntu 22.04, Fedora 42, OpenAnolis 8.8, and OpenAnolis 23.1. RPG detected 22 previously unknown bugs (13 distribution-specific), of which 15 were confirmed and 10 fixed by kernel maintainers. RPG also achieved 20.4% and 21.2% higher branch coverage than Syzkaller and Healer, respectively.
Yuheng Shen, Guoyu Yin, Runzhe Wang, Tao Ma 0006, Xiaohai Shi, Heyuan Shi
ASE5
2025 LLM-assisted Industrial-Scale Differential Testing of Package Incompatibilities in Linux Distributions
abstract
An open source Linux distribution often undergoes version upgrades and migrations, which is prone to incompatibility issues especially when it comes to large-scale software changes. Although differential testing has been widely used in software testing, it is still challenging to apply it for detecting such incompatibilities in the context of industrial settings. In this paper, we report our experience in leveraging LLMs to address the challenges faced by the Linux distribution community. Specifically, we develop an LLM-based differential testing method called Versify to assist maintainers of Linux distributions in locating incompatibilities during version upgrades and migrations. Its trial operation period within the Linux distribution community shows that it uncovered 8,489 instances of differing behavior, of which 644 were prioritized for attention by developers. After deduplication and filtering, 39 unique compatibility reports were identified. Feedback from Linux distributions developers indicates that our reports have provided valuable recommendations for package selection in future OS releases.
Chijin Zhou, Runzhe Wang, Weibo Zhang, Yuheng Shen, Xiaohai Shi, Tao Ma 0006, Zhe Wang 0015, Heyuan Shi
ASE3
2025 DragonRadar: Fuzzing Linux Kernel Deployed in Cloud-Native Environment
abstract
Kata Containers is a secure container runtime with lightweight virtual machines and a customized Linux kernel optimized for cloud-native workloads, which is important for cloud-native systems. Fuzzing is a widely-used technique for detecting kernel vulnerability. However, current kernel fuzzers can't be simply applied to kernels in cloud-native environments because of the discrepancies between test and actual deployment scenarios. This paper introduces DragonRadar, a kernel fuzzing tool adapted for Kata Containers, which aligns the testing environment with cloud deployment realities. We extend to support kernel fuzzing in cloud-native environments by integrating Syzkaller's capability with a lightweight virtual machine manager called Dragonball. The evaluation shows that DragonRadar effectively identifies 25 kernel vulnerabilities in the mainline Linux kernel used in the Kata Containers environment, while maintaining code coverage similar to vanilla Syzkaller. DragonRadar is available at https://github.com/TOBESTONG//DragonRadar.
Heyuan Shi, Weibo Zhang, Runzhe Wang, Xiaohai Shi, Guoyu Yin, Jianzhong Liu, Yuheng Shen
SANER3
2024 Attributed Triple Extraction by Combination Under Contrastive Learning
Runzhe Wang, Guilin Qi, Yongrui Chen 0002, Songlin Zhai, Rihui Jin, Nijun Li, Qianren Wang
DASFAA (7)1
2024 The Marginal Value of Momentum for Small Learning Rate SGD
abstract
Momentum is known to accelerate the convergence of gradient descent in strongly convex settings without stochastic gradient noise. In stochastic optimization, such as training neural networks, folklore suggests that momentum may help deep learning optimization by reducing the variance of the stochastic gradient update, but previous theoretical analyses do not find momentum to offer any provable acceleration. Theoretical results in this paper clarify the role of momentum in stochastic settings where the learning rate is small and gradient noise is the dominant source of instability, suggesting that SGD with and without momentum behave similarly in the short and long time horizons. Experiments show that momentum indeed has limited benefits for both optimization and generalization in practical training regimes where the optimal learning rate is not very large, including small- to medium-batch training from scratch on ImageNet and fine-tuning language models on downstream tasks.
Runzhe Wang, Sadhika Malladi, Tianhao Wang 0017, Kaifeng Lyu, Zhiyuan Li 0005
ICLR1
2024 Industry Practice of Directed Kernel Fuzzing for Open-source Linux Distribution
abstract
Directed grey-box fuzzing is a widely used automatic testing technique that has helped developers test specific code space in the target program. Although many directed fuzzers are designed to test the Linux kernel, challenges still remain due to the complexity of industrial requirements and deployment environments. In this paper, we collaborate with developers from Alibaba and the OpenAnolis community to conduct an industry practice of directed kernel fuzzing for open-source Linux distribution. We highlight typical challenges in deploying directed kernel fuzzing, including target-related kernel configuration options being disabled, unrelated initial seeds limiting fuzzing startup performance, no support for kernel feature interface fuzzing, independent fuzzer execution limiting fuzzing effectiveness, much manual work to triage and analyze crashes, and hard to integrate into the existing fuzzing framework. We provide solutions to these challenges, which allowed us to discover 11 previously unknown kernel bugs related to cloud-native features, io_uring, and other components in the OpenAnolis Linux distribution.
Heyuan Shi, Runzhe Wang, Weibo Zhang, Yuheng Shen, Xiaohai Shi, Yu Jiang 0001
ASE3
2024 PatchBert: Continuous Stable Patch Identification for Linux Kernel via Pre-trained Model Fine-tuning
abstract
Stable patch identification is crucial in merging patches into stable versions, which helps ensure the stability of the Linux kernel. Although many tools have been proposed to mitigate the manual effort of stable patch identification, challenges still arise because they neglect continuous stable patch tracking and advanced Natural Language Processing (NLP) pre-training techniques. In this paper, in collaboration with developers from the openAnolis Linux operating system distribution community, we present a stable patch identification model called PatchBERT. It utilizes BERT and CodeBERT to capture the semantic patch representation from the commit message and code changes in a patch. We then perform patch classification and output the probability that the patch should be merged into the stable versions. We perform experiments on the dataset used by the previous methods. The experimental results show the superior performance of PatchBERT over state-of-the-art baselines. Additionally, it is common practice to train the model using the latest Linux patches and implement it in a real-world industrial setting. In this exercise, we randomly select 10,000 patches for identification, accurately identifying 8,617 patches and incorrectly identifying 1,383 patches. This practical outcome further confirms the effectiveness and utility of PatchBERT in real-world scenarios.
Heyuan Shi, Runzhe Wang, Yuheng Shen, Yuao Chen, Xiaohai Shi, Yu Jiang 0001
SANER4
2023 KeenTune: Automated Tuning Tool for Cloud Application Performance Testing and Optimization
abstract
The performance testing and optimization of cloud applications is challenging, because manual tuning of cloud computing stacks is tedious and automated tuning tools are rare used for cloud services. To address this issue, we introduce KeenTune, an automated tuning tool designed to optimize application performance and facilitate performance testing. KeenTune is a lightweight and flexible tool that can be deployed with to-be-tuned applications with negligible impact on their performance. Specifically, KeenTune uses a surrogate model that can be implemented with machine learning models to filter out less relevant parameters for efficient tuning. Our empirical evaluation shows that KeenTune significantly enhances the throughput performance of Nginx web servers, resulting in performance improvements of up to 90.43% and 117.23% in certain cases. This study highlights the benefits of using KeenTune for achieving efficient and effective performance testing of cloud applications. The video and source code for KeenTune are provided as supplementary materials.
Qinglong Wang 0003, Runzhe Wang, Xiaohai Shi, Zheng Liu 0022, Tao Ma 0006, Houbing Song, Heyuan Shi
ISSTA2
2023 Brief Industry Paper: Directed Kernel Fuzz Testing on Real-time Linux
abstract
Rt-Linux contains critical modifications that are much less tested than the vanilla kernel, thus placing many systems at risk. In this paper, we present DRLF, a directed fuzzer targeted towards fuzzing any code area in Rt- Linux, thus allowing for more efficient tests on Rt-Linux's unique code sections. DRLF performs directed fuzzing through a kernel-level weighted callgraph construction technique, and prioritizing input sequences that exhibit less distance to the target code. Evaluations show that DRLF delivers better cover speed while achieving a 24.70% coverage increase for the targeting code areas. DRLF also found 11 previously unknown bugs within Rt-Linux, and has been integrated into Alibaba's CI/CD pipeline.
Yuheng Shen, Jianzhong Liu, Yiru Xu, Runzhe Wang, Heyuan Shi, Yu Jiang 0001
RTSS6
2022 Industry practice of configuration auto-tuning for cloud applications and services
abstract
Auto-tuning attracts increasing attention in industry practice to optimize the performance of a system with many configurable parameters. It is particularly useful for cloud applications and services since they have complex system hierarchies and intricate knob correlations. However, existing tools and algorithms rarely consider practical problems such as workload pressure control, the support for distributed deployment, and expensive time costs, etc., which are utterly important for enterprise cloud applications and services. In this work, we significantly extend an open source tuning tool – KeenTune to optimize several typical enterprise cloud applications and services. Our practice is in collaboration with enterprise users and tuning tool developers to address the aforementioned problems. Specifically, we highlight five key challenges from our experiences and provide a set of solutions accordingly. Through applying the improved tuning tool to different application scenarios, we achieve 2%-14% improvements for the performance of MySQL, OceanBase, nginx, ingress-nginx, and 5%-70% improvements for the performance of ACK cloud container service.
Runzhe Wang, Qinglong Wang 0003, Heyuan Shi, Yuheng Shen, Zheng Liu 0022, Xiaohai Shi, Yu Jiang 0001
ESEC/SIGSOFT FSE1
2021 Going Beyond Linear RL: Sample Efficient Neural Function Approximation
abstract
Deep Reinforcement Learning (RL) powered by neural net approximation of the Q function has had enormous empirical success. While the theory of RL has traditionally focused on linear function approximation (or eluder dimension) approaches, little is known about nonlinear RL with neural net approximations of the Q functions. This is the focus of this work, where we study function approximation with two-layer neural networks (considering both ReLU and polynomial activation functions). Our first result is a computationally and statistically efficient algorithm in the generative model setting under completeness for two-layer neural networks. Our second result considers this setting but under only realizability of the neural net function class. Here, assuming deterministic dynamics, the sample complexity scales linearly in the algebraic dimension. In all cases, our results significantly improve upon what can be attained with linear (or eluder dimension) methods.
Baihe Huang, Kaixuan Huang, Sham M. Kakade, Jason D. Lee, Runzhe Wang, Jiaqi Yang 0001
NeurIPS6
2021 Optimal Gradient-based Algorithms for Non-concave Bandit Optimization
abstract
Bandit problems with linear or concave reward have been extensively studied, but relatively few works have studied bandits with non-concave reward. This work considers a large family of bandit problems where the unknown underlying reward function is non-concave, including the low-rank generalized linear bandit problems and two-layer neural network with polynomial activation bandit problem.For the low-rank generalized linear bandit problem, we provide a minimax-optimal algorithm in the dimension, refuting both conjectures in \cite{lu2021low,jun2019bilinear}. Our algorithms are based on a unified zeroth-order optimization paradigm that applies in great generality and attains optimal rates in several structured polynomial settings (in the dimension). We further demonstrate the applicability of our algorithms in RL in the generative model setting, resulting in improved sample complexity over prior approaches.Finally, we show that the standard optimistic algorithms (e.g., UCB) are sub-optimal by dimension factors. In the neural net setting (with polynomial activation functions) with noiseless reward, we provide a bandit algorithm with sample complexity equal to the intrinsic algebraic dimension. Again, we show that optimistic approaches have worse sample complexity, polynomial in the extrinsic dimension (which could be exponentially worse in the polynomial degree).
Baihe Huang, Kaixuan Huang, Sham M. Kakade, Jason D. Lee, Runzhe Wang, Jiaqi Yang 0001
NeurIPS6
2021 Gradient Descent on Two-layer Nets: Margin Maximization and Simplicity Bias
abstract
The generalization mystery of overparametrized deep nets has motivated efforts to understand how gradient descent (GD) converges to low-loss solutions that generalize well. Real-life neural networks are initialized from small random values and trained with cross-entropy loss for classification (unlike the "lazy" or "NTK" regime of training where analysis was more successful), and a recent sequence of results (Lyu and Li, 2020; Chizat and Bach, 2020; Ji and Telgarsky, 2020) provide theoretical evidence that GD may converge to the "max-margin" solution with zero loss, which presumably generalizes well. However, the global optimality of margin is proved only in some settings where neural nets are infinitely or exponentially wide. The current paper is able to establish this global optimality for two-layer Leaky ReLU nets trained with gradient flow on linearly separable and symmetric data, regardless of the width. The analysis also gives some theoretical justification for recent empirical findings (Kalimeris et al., 2019) on the so-called simplicity bias of GD towards linear or other "simple" classes of solutions, especially early in training. On the pessimistic side, the paper suggests that such results are fragile. A simple data manipulation can make gradient flow converge to a linear classifier with suboptimal margin.
Kaifeng Lyu, Zhiyuan Li 0005, Runzhe Wang, Sanjeev Arora
NeurIPS3
2019 Industry practice of coverage-guided enterprise Linux kernel fuzzing
abstract
Coverage-guided kernel fuzzing is a widely-used technique that has helped kernel developers and testers discover numerous vulnerabilities. However, due to the high complexity of application and hardware environment, there is little study on deploying fuzzing to the enterprise-level Linux kernel. In this paper, collaborating with the enterprise developers, we present the industry practice to deploy kernel fuzzing on four different enterprise Linux distributions that are responsible for internal business and external services of the company. We have addressed the following outstanding challenges when deploying a popular kernel fuzzer, syzkaller, to these enterprise Linux distributions: coverage support absence, kernel configuration inconsistency, bugs in shallow paths, and continuous fuzzing complexity. This leads to a vulnerability detection of 41 reproducible bugs which are previous unknown in these enterprise Linux kernel and 6 bugs with CVE IDs in U.S. National Vulnerability Database, including flaws that cause general protection fault, deadlock, and use-after-free.
Heyuan Shi, Runzhe Wang, Xiaohai Shi, Xun Jiao 0002, Houbing Song, Yu Jiang 0001, Jia-Guang Sun 0001
ESEC/SIGSOFT FSE2
2019 Vulnerable Code Clone Detection for Operating System Through Correlation-Induced Learning
abstract
Vulnerable code clones in the operating system (OS) threaten the safety of smart industrial environment, and most vulnerable OS code clone detection approaches neglect correlations between functions that limits the detection effectiveness. In this article, we propose a two-phase framework to find vulnerable OS code clones by learning on correlations between functions. On the training phase, functions as the training set are extracted from the latest code repository and function features are derived by their AST structure. Then, external and internal correlations are explored by graph modeling of functions. Finally, the graph convolutional network for code clone detection (GCN-CC) is trained using function features and correlations. On the detection phase, functions in the to-be-detected OS code repository are extracted and the vulnerable OS code clones are detected by the trained GCN-CC. We conduct experiments on five real OS code repositories, and experimental results show that our framework outperforms the state-of-the-art approaches.
Heyuan Shi, Runzhe Wang, Yu Jiang 0001, Jian Dong 0001, Jia-Guang Sun 0001
IEEE Trans. Ind. Informatics2
2015 FastID: An undeceived router for real-time identification of WiFi terminals
abstract
In recent past, the rapid developing of mobile internet inspires the widespread use of WiFi (IEEE 802.11) technology. In WiFi, the access control of a terminal to the router remains a significant challenge because the PIN (password) and MAC address are easy to guess and forge. In this paper, we present FastID - a practical system that identifies WiFi terminals in real-time by fingerprinting their clocks. Previous approaches of clock fingerprinting require tens of minutes or even hours for clock data collection, and thus cannot be applied into real-time WiFi terminal identification. Even worse, unstable wireless communications and unknown status of terminals' OSes may further degrade the accuracy of fingerprint computation. In comparison, FastID performs fast clock fingerprinting based on the timestamps carried by terminals' ICMP packets. Moreover, FastID employs simple but efficient techniques to remove outliers of collected clock data and differentiate terminals based on the similarity of their distributions, making it suitable for fast terminals identification. FastID is implemented on an off-the-shelf commercial WiFi router and extensively evaluated based on 10 commodity WiFi terminals. Experimental results show that FastID is able to identify terminals with high accuracy and low cost within several seconds.
Li Lu 0001, Runzhe Wang, Wubin Mao, Hongzi Zhu
Networking2