Yexiao He

dblp:309/9333 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0002-4675-7733ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Computer networks · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Efficient and distributed learning · 75% Language models and text generation · 25%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Electronic design automation · 94% Performance modeling and evaluation · 6%
Databases, data mining, and information retrieval
1 paper
Data integration and cleaning · 50% Machine learning and data management · 50%
Computer networks
1 paper
Edge and fog computing · 100%

Topics — the 18 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation › hardware verification and test › formal verification
equivalence checking
0.912025
SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic Reasoning · NeurIPS 2025
Electronic design automation
hardware verification and test
0.912025
SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic Reasoning · NeurIPS 2025
Electronic design automation
logic synthesis
0.912025
SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic Reasoning · NeurIPS 2025
Electronic design automation › logic synthesis › digital system synthesis
RTL optimization
0.912025
SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic Reasoning · NeurIPS 2025
Natural language and speech › Language models and text generation
dataset refinement
0.812024
SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning · NeurIPS 2024
Machine learning › Efficient and distributed learning › federated learning
federated fine-tuning
0.812024
FLoRA: Federated Fine-Tuning Large Language Models with Heterogeneous Low-Rank Adaptations · NeurIPS 2024
Machine learning › Efficient and distributed learning
federated learning
0.812024
FLoRA: Federated Fine-Tuning Large Language Models with Heterogeneous Low-Rank Adaptations · NeurIPS 2024
Natural language and speech › Language models and text generation
instruction tuning
0.812024
SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning · NeurIPS 2024
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
low-rank adaptation
0.812024
FLoRA: Federated Fine-Tuning Large Language Models with Heterogeneous Low-Rank Adaptations · NeurIPS 2024
Machine learning › Efficient and distributed learning
model compression
0.812024
FLoRA: Federated Fine-Tuning Large Language Models with Heterogeneous Low-Rank Adaptations · NeurIPS 2024
Machine learning › Efficient and distributed learning › large-scale learning
model scaling
0.812024
Model-GLUE: Democratized LLM Scaling for A Large Model Zoo in the Wild · NeurIPS 2024
Data integration and cleaning
data curation
0.812024
SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning · NeurIPS 2024
Machine learning and data management
data selection
0.812024
SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning · NeurIPS 2024
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.312025
EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices · MobiSys 2025
Electronic design automation › logic synthesis
finite state machine optimization
0.312025
SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic Reasoning · NeurIPS 2025
Performance modeling and evaluation › state space exploration
state aggregation
0.312025
SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic Reasoning · NeurIPS 2025
Machine learning › Efficient and distributed learning
data-efficient learning
0.212024
SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning · NeurIPS 2024
Machine learning › Efficient and distributed learning
model reuse
0.212024
Model-GLUE: Democratized LLM Scaling for A Large Model Zoo in the Wild · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

batch inference · 1.7adapter caching · 1.7LoRA · 1.6shapley value · 1.5dataset refinement · 1.5symbolic reasoning · 0.9retrieval-augmented generation · 0.9lora · 0.9large language model · 0.9formal equivalence checking · 0.9abstract syntax tree · 0.9stacking-based aggregation · 0.8heterogeneous low-rank adaptations · 0.8
YearPublicationVenuePosition
2025 EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
abstract
Large Language Models (LLMs) have gained significant attention due to their versatility across a wide array of applications. Fine-tuning LLMs with parameter-efficient adapters, such as Low-Rank Adaptation (LoRA), enables these models to efficiently adapt to downstream tasks without extensive retraining. Deploying fine-tuned LLMs on multi-tenant edge devices offers substantial benefits, such as reduced latency, enhanced privacy, and personalized responses. However, serving LLMs efficiently on resource-constrained edge devices presents critical challenges, including the complexity of adapter selection for different tasks, memory overhead from frequent adapter swapping. Moreover, given the multiple requests in the multi-tenant settings, processing requests sequentially will result in underutilization of computational resources and significant latency. This paper introduces EdgeLoRA, an efficient system for serving LLMs on edge devices in multi-tenant environments. EdgeLoRA incorporates three key innovations: (1) an adaptive adapter selection mechanism to streamline the adapter configuration process; (2) heterogeneous memory management, leveraging intelligent adapter caching and pooling to mitigate memory operation overhead; and (3) batch LoRA inference, which enables efficient batch processing to significantly reduce computational latency. Comprehensive evaluations using the Llama3.1-8B model demonstrates that EdgeLoRA significantly outperforms the status quo (i.e., llama.cpp) in terms of both latency and throughput. The results demonstrates EdgeLoRA could achieve up to 4× boost in throughput with less energy consumption. Even more impressively, it manages to serve several orders of magnitude more adapters simultaneously without sacrificing inference performance. These results highlight EdgeLoRA's potential to transform edge deployment of LLMs in multi-tenant scenarios, offering a scalable and efficient solution for resource-constrained environments.
Zheyu Shen, Yexiao He, Guoheng Sun, Wanghao Ye, Ang Li 0005
MobiSys2
2025 SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic Reasoning
abstract
Optimizing Register Transfer Level (RTL) code is crucial for improving the efficiency and performance of digital circuits in the early stages of synthesis. Manual rewriting, guided by synthesis feedback, can yield high-quality results but is time-consuming and error-prone. Most existing compiler-based approaches have difficulty handling complex design constraints. Large Language Model (LLM)-based methods have emerged as a promising alternative to address these challenges. However, LLM-based approaches often face difficulties in ensuring alignment between the generated code and the provided prompts. This paper introduces SymRTLO, a neuron-symbolic framework that integrates LLMs with symbolic reasoning for the efficient and effective optimization of RTL code. Our method incorporates a retrieval-augmented system of optimization rules and Abstract Syntax Tree (AST)-based templates, enabling LLM-based rewriting that maintains syntactic correctness while minimizing undesired circuit behaviors. A symbolic module is proposed for analyzing and optimizing finite state machine (FSM) logic, allowing fine-grained state merging and partial specification handling beyond the scope of pattern-based compilers. Furthermore, a fast verification pipeline, combining formal equivalence checks with test-driven validation, further reduces the complexity of verification. Experiments on the RTL-Rewriter benchmark with Synopsys Design Compiler and Yosys show that SymRTLO improves power, performance, and area (PPA) by up to 43.9%, 62.5%, and 51.1%, respectively, compared to the state-of-the-art methods. We will release the code as open source upon the paper's acceptance.
Wanghao Ye, Ping Guo 0007, Yexiao He, Bowei Tian, Shwai He, Guoheng Sun, Zheyu Shen, Ankur Srivastava 0001, Qingfu Zhang 0001, Gang Qu 0001, Ang Li 0005
NeurIPS4
2024 SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning
abstract
The pre-trained Large Language Models (LLMs) can be adapted for many downstream tasks and tailored to align with human preferences through fine-tuning. Recent studies have discovered that LLMs can achieve desirable performance with only a small amount of high-quality data, suggesting that a large portion of the data in these extensive datasets is redundant or even harmful. Identifying high-quality data from vast datasets to curate small yet effective datasets has emerged as a critical challenge. In this paper, we introduce SHED, an automated dataset refinement framework based on Shapley value for instruction fine-tuning. SHED eliminates the need for human intervention or the use of commercial LLMs. Moreover, the datasets curated through SHED exhibit transferability, indicating they can be reused across different LLMs with consistently high performance. We conduct extensive experiments to evaluate the datasets curated by SHED. The results demonstrate SHED's superiority over state-of-the-art methods across various tasks and LLMs; notably, datasets comprising only 10% of the original data selected by SHED achieve performance comparable to or surpassing that of the full datasets.
Yexiao He, Zheyu Shen, Guoheng Sun, Yucong Dai, Yongkai Wu, Hongyi Wang 0001, Ang Li 0005
NeurIPS1
2024 FLoRA: Federated Fine-Tuning Large Language Models with Heterogeneous Low-Rank Adaptations
abstract
The rapid development of Large Language Models (LLMs) has been pivotal in advancing AI, with pre-trained LLMs being adaptable to diverse downstream tasks through fine-tuning. Federated learning (FL) further enhances fine-tuning in a privacy-aware manner by utilizing clients' local data through in-situ computation, eliminating the need for data movement. However, fine-tuning LLMs, given their massive scale of parameters, poses challenges for clients with constrained and heterogeneous resources in FL. Previous methods employed low-rank adaptation (LoRA) for efficient federated fine-tuning but utilized traditional FL aggregation strategies on LoRA adapters. This approach led to mathematically inaccurate aggregation noise, reducing fine-tuning effectiveness and failing to address heterogeneous LoRAs. In this work, we first highlight the mathematical incorrectness of LoRA aggregation in existing federated fine-tuning methods. We introduce a new approach called FLoRA that enables federated fine-tuning on heterogeneous LoRA adapters across clients through a novel stacking-based aggregation method. Our approach is noise-free and seamlessly supports heterogeneous LoRAs. Extensive experiments demonstrate FLoRA's superior performance in both homogeneous and heterogeneous settings, surpassing state-of-the-art methods. We envision this work as a milestone for efficient, privacy-preserving, and accurate federated fine-tuning of LLMs.
Zheyu Shen, Yexiao He, Guoheng Sun, Hongyi Wang 0001, Lingjuan Lyu, Ang Li 0005
NeurIPS3
2024 Model-GLUE: Democratized LLM Scaling for A Large Model Zoo in the Wild
Guoheng Sun, Ruisi Cai, Pingzhi Li, Peihao Wang, Bowen Tan, Yexiao He, Beidi Chen, Binhang Yuan, Hongyi Wang 0001, Ang Li 0005, Zhangyang Wang, Tianlong Chen 0001
NeurIPS8
2024 Dynamic relay node selection and routing for cloud-native Software Defined WANs
Chenyu Fan, Yangming Zhao, Yexiao He
Comput. Networks4
2022 Joint optimization of Service Chain Graph Design and Mapping in NFV-enabled networks
Yexiao He, Zixiang Xia, Keshav Sood, Shui Yu 0001
Comput. Networks1