Guoheng Sun

dblp:353/7384 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
10since 2021 · last 2026
0009-0004-4346-8516ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Efficient and distributed learning · 56% Trustworthy machine learning · 23% Language models and text generation · 14%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Electronic design automation · 94% Performance modeling and evaluation · 6%
Network and information security
1 paper
Security and privacy of machine learning · 67% Authentication and access control · 33%
Databases, data mining, and information retrieval
1 paper
Data integration and cleaning · 50% Machine learning and data management · 50%
Computer networks
1 paper
Edge and fog computing · 100%

Topics — the 26 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
1.622025
Router-Tuning: A Simple and Effective Approach for Dynamic Depth · EMNLP 2025
FLoRA: Federated Fine-Tuning Large Language Models with Heterogeneous Low-Rank Adaptations · NeurIPS 2024
Security and privacy of machine learning
adversarial example
1.012026
Enhancing the Security of Large Character Set CAPTCHAs Using Transferable Adversarial Examples · IEEE Trans. Dependable Secur. Comput. 2026
Authentication and access control › human interactive proofs
CAPTCHA
1.012026
Enhancing the Security of Large Character Set CAPTCHAs Using Transferable Adversarial Examples · IEEE Trans. Dependable Secur. Comput. 2026
Security and privacy of machine learning › adversarial attack
transferable adversarial attack
1.012026
Enhancing the Security of Large Character Set CAPTCHAs Using Transferable Adversarial Examples · IEEE Trans. Dependable Secur. Comput. 2026
Machine learning › Trustworthy machine learning › fairness
causal fairness
0.912025
Towards counterfactual fairness through auxiliary variables · ICLR 2025
Machine learning › Trustworthy machine learning › fairness › causal fairness
counterfactual fairness
0.912025
Towards counterfactual fairness through auxiliary variables · ICLR 2025
Machine learning › Efficient and distributed learning › adaptive computation
dynamic layer skipping
0.912025
Router-Tuning: A Simple and Effective Approach for Dynamic Depth · EMNLP 2025
Machine learning › Trustworthy machine learning
fairness
0.912025
Towards counterfactual fairness through auxiliary variables · ICLR 2025
Machine learning › Deep learning architectures and training
mixture of experts
0.912025
Router-Tuning: A Simple and Effective Approach for Dynamic Depth · EMNLP 2025
Electronic design automation › hardware verification and test › formal verification
equivalence checking
0.912025
SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic Reasoning · NeurIPS 2025
Electronic design automation
hardware verification and test
0.912025
SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic Reasoning · NeurIPS 2025
Electronic design automation
logic synthesis
0.912025
SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic Reasoning · NeurIPS 2025
Electronic design automation › logic synthesis › digital system synthesis
RTL optimization
0.912025
SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic Reasoning · NeurIPS 2025
Natural language and speech › Language models and text generation
dataset refinement
0.812024
SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning · NeurIPS 2024
Machine learning › Efficient and distributed learning › federated learning
federated fine-tuning
0.812024
FLoRA: Federated Fine-Tuning Large Language Models with Heterogeneous Low-Rank Adaptations · NeurIPS 2024
Machine learning › Efficient and distributed learning
federated learning
0.812024
FLoRA: Federated Fine-Tuning Large Language Models with Heterogeneous Low-Rank Adaptations · NeurIPS 2024
Natural language and speech › Language models and text generation
instruction tuning
0.812024
SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning · NeurIPS 2024
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
low-rank adaptation
0.812024
FLoRA: Federated Fine-Tuning Large Language Models with Heterogeneous Low-Rank Adaptations · NeurIPS 2024
Machine learning › Efficient and distributed learning › large-scale learning
model scaling
0.812024
Model-GLUE: Democratized LLM Scaling for A Large Model Zoo in the Wild · NeurIPS 2024
Data integration and cleaning
data curation
0.812024
SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning · NeurIPS 2024
Machine learning and data management
data selection
0.812024
SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning · NeurIPS 2024
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.312025
EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices · MobiSys 2025
Electronic design automation › logic synthesis
finite state machine optimization
0.312025
SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic Reasoning · NeurIPS 2025
Performance modeling and evaluation › state space exploration
state aggregation
0.312025
SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic Reasoning · NeurIPS 2025
Machine learning › Efficient and distributed learning
data-efficient learning
0.212024
SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning · NeurIPS 2024
Machine learning › Efficient and distributed learning
model reuse
0.212024
Model-GLUE: Democratized LLM Scaling for A Large Model Zoo in the Wild · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

batch inference · 1.7adapter caching · 1.7LoRA · 1.6dataset refinement · 1.5ensemble generation · 1.0adversarial perturbation · 1.0symbolic reasoning · 0.9router fine-tuning · 0.9retrieval-augmented generation · 0.9moe layer skipping · 0.9lora · 0.9large language model · 0.9formal equivalence checking · 0.9exogenous variable · 0.9causal reasoning · 0.9auxiliary variables · 0.9attention layer skipping · 0.9abstract syntax tree · 0.9
YearPublicationVenuePosition
2026 VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
abstract
Automating Register Transfer Level (RTL) code generation with Large Language Models (LLMs) can reduce manual hardware design effort. However, current LLM-based approaches face four challenges: limited availability of high-quality training data, weak alignment between natural language specifications and generated code, lack of built-in verification mechanisms, and difficulty in adapting general-purpose models to RTL-specific constraints. Inspired by DeepSeek-R1, which combines reinforcement learning with reasoning capabilities, we introduce VeriReason, a framework that integrates supervised fine-tuning with Group Relative Policy Optimization (GRPO) for RTL code generation. Using high-quality training examples, a feedback-driven reward model, testbench evaluation, and structural heuristics, VeriReason improves specification-code alignment, reduces hallucinations, and strengthens reasoning traces and first-attempt functional correctness. To our knowledge, VeriReason is the first system that successfully integrates explicit reasoning capabilities with reinforcement learning for Verilog generation. On VerilogEval-Machine, VeriReason reaches 83.1% pass@5, while consistently outperforming comparable-sized open-source models. Our approach demonstrates up to a 2.8 × increase in first-attempt functional correctness compared to baseline methods.
Guoheng Sun, Wanghao Ye, Gang Qu 0001, Ang Li 0005
ACM Great Lakes Symposium on VLSI2
2026 Enhancing the Security of Large Character Set CAPTCHAs Using Transferable Adversarial Examples
abstract
The large character set CAPTCHA is an important extension of the traditional text-based CAPTCHA with larger alphabet languages to defend against automated attack programs. However, the state-of-the-art deep learning attacks have cracked such CAPTCHA. Existing defenses against such threats increase the complexity of CAPTCHA, thus decreasing usability. We propose ACG (Adversarial Large Character Set CAPTCHA Generation), a framework with two modules: aFine-grained Generation Module, combining three novel strategies to prevent attackers from recognizing characters, and anEnsemble Generation Moduleto generate global perturbations in CAPTCHAs. It not only strengthens defense against recognition attacks but also improves robustness against diverse detection architectures through adversarial perturbations. Additionally, we develop a toolkit, Adv-Eval, consisting of CAPTCHA datasets from 10 of the most popular Chinese CAPTCHA schemes and benchmarking various attacks. We conduct extensive experiments using Adv-Eval to demonstrate ACG's efficacy, especially manifesting a significant decrease in the average success rate of diverse attacks from 51.52% to 2.56%. To the best of our knowledge, ACG is the first framework to defend large character set CAPTCHAs against detection attacks using transferable adversarial examples.
Guoheng Sun, Yucheng Fu, Juntian Huang, Ruimei Zhang, Haizhou Wang 0001
IEEE Trans. Dependable Secur. Comput.1
2025 Router-Tuning: A Simple and Effective Approach for Dynamic Depth
abstract
The Mixture of Depths (MoD) was introduced to improve computational efficiency by dynamically skipping less important layers, reducing redundant computation while maintaining model capacity.Despite its promise, existing MoD approaches remain under-explored and face two main challenges: (1) high training costs due to the need to train the entire model along with the routers that determine which layers to skip, and (2) performance degradation when important layers are bypassed.In response to the first issue, we propose Router-Tuning, which fine-tunes only the routers on a small dataset, drastically reducing the computational overhead associated with full model training.For the second challenge, we investigate Router-Tuning across different architectures and granularities, demonstrating its effectiveness on Attention layers and MoE layers.This method preserves the model's performance while significantly enhancing computational and memory efficiency.Extensive experiments demonstrate that our approach delivers competitive results while dramatically improving the computation efficiency, e.g., 21% speedup and only a 0.2% performance drop.
Shwai He, Tao Ge 0001, Guoheng Sun, Bowei Tian, Xiaoyang Wang 0001, Dong Yu 0001
EMNLP3
2025 Towards counterfactual fairness through auxiliary variables
abstract
The challenge of balancing fairness and predictive accuracy in machine learning models, especially when sensitive attributes such as race, gender, or age are considered, has motivated substantial research in recent years. Counterfactual fairness ensures that predictions remain consistent across counterfactual variations of sensitive attributes, which is a crucial concept in addressing societal biases. However, existing counterfactual fairness approaches usually overlook intrinsic information about sensitive features, limiting their ability to achieve fairness while simultaneously maintaining performance. To tackle this challenge, we introduce EXOgenous Causal reasoning (EXOC), a novel causal reasoning framework motivated by exogenous variables. It leverages auxiliary variables to uncover intrinsic properties that give rise to sensitive attributes. Our framework explicitly defines an auxiliary node and a control node that contribute to counterfactual fairness and control the information flow within the model. Our evaluation, conducted on synthetic and real-world datasets, validates EXOC's superiority, showing that it outperforms state-of-the-art approaches in achieving counterfactual fairness without sacrificing accuracy. Our code is available at https://github.com/CASE-Lab-UMD/counterfactual_fairness_2025.
Bowei Tian, Shwai He, Wanghao Ye, Guoheng Sun, Yucong Dai, Yongkai Wu, Ang Li 0005
ICLR5
2025 EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
abstract
Large Language Models (LLMs) have gained significant attention due to their versatility across a wide array of applications. Fine-tuning LLMs with parameter-efficient adapters, such as Low-Rank Adaptation (LoRA), enables these models to efficiently adapt to downstream tasks without extensive retraining. Deploying fine-tuned LLMs on multi-tenant edge devices offers substantial benefits, such as reduced latency, enhanced privacy, and personalized responses. However, serving LLMs efficiently on resource-constrained edge devices presents critical challenges, including the complexity of adapter selection for different tasks, memory overhead from frequent adapter swapping. Moreover, given the multiple requests in the multi-tenant settings, processing requests sequentially will result in underutilization of computational resources and significant latency. This paper introduces EdgeLoRA, an efficient system for serving LLMs on edge devices in multi-tenant environments. EdgeLoRA incorporates three key innovations: (1) an adaptive adapter selection mechanism to streamline the adapter configuration process; (2) heterogeneous memory management, leveraging intelligent adapter caching and pooling to mitigate memory operation overhead; and (3) batch LoRA inference, which enables efficient batch processing to significantly reduce computational latency. Comprehensive evaluations using the Llama3.1-8B model demonstrates that EdgeLoRA significantly outperforms the status quo (i.e., llama.cpp) in terms of both latency and throughput. The results demonstrates EdgeLoRA could achieve up to 4× boost in throughput with less energy consumption. Even more impressively, it manages to serve several orders of magnitude more adapters simultaneously without sacrificing inference performance. These results highlight EdgeLoRA's potential to transform edge deployment of LLMs in multi-tenant scenarios, offering a scalable and efficient solution for resource-constrained environments.
Zheyu Shen, Yexiao He, Guoheng Sun, Wanghao Ye, Ang Li 0005
MobiSys5
2025 SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic Reasoning
abstract
Optimizing Register Transfer Level (RTL) code is crucial for improving the efficiency and performance of digital circuits in the early stages of synthesis. Manual rewriting, guided by synthesis feedback, can yield high-quality results but is time-consuming and error-prone. Most existing compiler-based approaches have difficulty handling complex design constraints. Large Language Model (LLM)-based methods have emerged as a promising alternative to address these challenges. However, LLM-based approaches often face difficulties in ensuring alignment between the generated code and the provided prompts. This paper introduces SymRTLO, a neuron-symbolic framework that integrates LLMs with symbolic reasoning for the efficient and effective optimization of RTL code. Our method incorporates a retrieval-augmented system of optimization rules and Abstract Syntax Tree (AST)-based templates, enabling LLM-based rewriting that maintains syntactic correctness while minimizing undesired circuit behaviors. A symbolic module is proposed for analyzing and optimizing finite state machine (FSM) logic, allowing fine-grained state merging and partial specification handling beyond the scope of pattern-based compilers. Furthermore, a fast verification pipeline, combining formal equivalence checks with test-driven validation, further reduces the complexity of verification. Experiments on the RTL-Rewriter benchmark with Synopsys Design Compiler and Yosys show that SymRTLO improves power, performance, and area (PPA) by up to 43.9%, 62.5%, and 51.1%, respectively, compared to the state-of-the-art methods. We will release the code as open source upon the paper's acceptance.
Wanghao Ye, Ping Guo 0007, Yexiao He, Bowei Tian, Shwai He, Guoheng Sun, Zheyu Shen, Ankur Srivastava 0001, Qingfu Zhang 0001, Gang Qu 0001, Ang Li 0005
NeurIPS8
2024 SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning
abstract
The pre-trained Large Language Models (LLMs) can be adapted for many downstream tasks and tailored to align with human preferences through fine-tuning. Recent studies have discovered that LLMs can achieve desirable performance with only a small amount of high-quality data, suggesting that a large portion of the data in these extensive datasets is redundant or even harmful. Identifying high-quality data from vast datasets to curate small yet effective datasets has emerged as a critical challenge. In this paper, we introduce SHED, an automated dataset refinement framework based on Shapley value for instruction fine-tuning. SHED eliminates the need for human intervention or the use of commercial LLMs. Moreover, the datasets curated through SHED exhibit transferability, indicating they can be reused across different LLMs with consistently high performance. We conduct extensive experiments to evaluate the datasets curated by SHED. The results demonstrate SHED's superiority over state-of-the-art methods across various tasks and LLMs; notably, datasets comprising only 10% of the original data selected by SHED achieve performance comparable to or surpassing that of the full datasets.
Yexiao He, Zheyu Shen, Guoheng Sun, Yucong Dai, Yongkai Wu, Hongyi Wang 0001, Ang Li 0005
NeurIPS4
2024 FLoRA: Federated Fine-Tuning Large Language Models with Heterogeneous Low-Rank Adaptations
abstract
The rapid development of Large Language Models (LLMs) has been pivotal in advancing AI, with pre-trained LLMs being adaptable to diverse downstream tasks through fine-tuning. Federated learning (FL) further enhances fine-tuning in a privacy-aware manner by utilizing clients' local data through in-situ computation, eliminating the need for data movement. However, fine-tuning LLMs, given their massive scale of parameters, poses challenges for clients with constrained and heterogeneous resources in FL. Previous methods employed low-rank adaptation (LoRA) for efficient federated fine-tuning but utilized traditional FL aggregation strategies on LoRA adapters. This approach led to mathematically inaccurate aggregation noise, reducing fine-tuning effectiveness and failing to address heterogeneous LoRAs. In this work, we first highlight the mathematical incorrectness of LoRA aggregation in existing federated fine-tuning methods. We introduce a new approach called FLoRA that enables federated fine-tuning on heterogeneous LoRA adapters across clients through a novel stacking-based aggregation method. Our approach is noise-free and seamlessly supports heterogeneous LoRAs. Extensive experiments demonstrate FLoRA's superior performance in both homogeneous and heterogeneous settings, surpassing state-of-the-art methods. We envision this work as a milestone for efficient, privacy-preserving, and accurate federated fine-tuning of LLMs.
Zheyu Shen, Yexiao He, Guoheng Sun, Hongyi Wang 0001, Lingjuan Lyu, Ang Li 0005
NeurIPS4
2024 Model-GLUE: Democratized LLM Scaling for A Large Model Zoo in the Wild
Guoheng Sun, Ruisi Cai, Pingzhi Li, Peihao Wang, Bowen Tan, Yexiao He, Beidi Chen, Binhang Yuan, Hongyi Wang 0001, Ang Li 0005, Zhangyang Wang, Tianlong Chen 0001
NeurIPS2
2023 Fighting Attacks on Large Character Set CAPTCHAs Using Transferable Adversarial Examples
abstract
Over a long period, large character set CAPTCHAs are widely used to defend against automated attack programs on the Internet. However, with the development of deep learning techniques, some attacks for large character set CAPTCHAs have been proposed, proving that they are no longer secure. To defend against black-box attacks on these CAPTCHAs, we propose a novel defense method based on transferable adversarial example techniques. On the one hand, we defend against character recognition attacks by adding adversarial perturbations to the characters of CAPTCHAs combining three strategies: Gradient-based Attacks, Input Transformations and Attention Mechanism. On the other hand, we defend against character detection attacks by leveraging an ensemble method to generate adversarial perturbations on the background of CAPTCHAs. To the best of our knowledge, this is the first study to improve the security of large character set CAPTCHAs against black-box attacks based on transferable adversarial example techniques. Using the eight most popular Chinese CAPTCHA schemes as examples, we conduct comprehensive experiments. Results show that our method improves the security of large character set CAPTCHAs by making the average success rate of black-box attacks significantly drop from 53.33% to 3.49%. Overall, our method can be helpful to the design of more secure large character set CAPTCHAs.
Yucheng Fu, Guoheng Sun, Juntian Huang, Haizhou Wang 0001
IJCNN2