Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Siyuan Chen 0007

dblp:84/5999-7 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2025
0009-0003-8454-0804ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Hardware accelerators and domain-specific architectures · 55% GPUs and heterogeneous computing · 12% Performance modeling and evaluation · 9%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%
Artificial intelligence
3 papers
Efficient and distributed learning · 84% Deep learning architectures and training · 16%

Topics — the 12 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › dataflow optimization
operator fusion
1.322023
TileFlow: A Framework for Modeling Fusion Dataflow via Tree-based Analysis · MICRO 2023
Chimera: An Analytical Optimizing Framework for Effective Compute-intensive Operators Fusion · HPCA 2023
Machine learning › Efficient and distributed learning
model compression
0.912025
Practical Offloading for Fine-Tuning LLM on Commodity GPU via Learned Sparse Projectors · AAAI 2025
Compilers and program optimization
machine learning compiler
0.712023
Chimera: An Analytical Optimizing Framework for Effective Compute-intensive Operators Fusion · HPCA 2023
Compilers and program optimization › deep learning compiler
operator fusion
0.712023
Chimera: An Analytical Optimizing Framework for Effective Compute-intensive Operators Fusion · HPCA 2023
Cloud and datacenter computing › datacenter workloads
batch processing
0.712023
ED-Batch: Efficient Automatic Batching of Dynamic Neural Networks via Learned Finite State Machines · ICML 2023
Electronic design automation › system-level design
dataflow design
0.712023
TileFlow: A Framework for Modeling Fusion Dataflow via Tree-based Analysis · MICRO 2023
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN accelerator
0.712023
TileFlow: A Framework for Modeling Fusion Dataflow via Tree-based Analysis · MICRO 2023
Hardware accelerators and domain-specific architectures › neural network mapping
DNN mapping
0.712023
Memory and Computation Coordinated Mapping of DNNs onto Complex Heterogeneous SoC · DAC 2023
Hardware accelerators and domain-specific architectures › accelerator architecture
heterogeneous accelerator
0.712023
Memory and Computation Coordinated Mapping of DNNs onto Complex Heterogeneous SoC · DAC 2023
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.712023
Chimera: An Analytical Optimizing Framework for Effective Compute-intensive Operators Fusion · HPCA 2023
Machine learning › Efficient and distributed learning
dynamic neural network
0.212023
ED-Batch: Efficient Automatic Batching of Dynamic Neural Networks via Learned Finite State Machines · ICML 2023
Parallel and multicore computing › locality optimization
data locality optimization
0.212023
Chimera: An Analytical Optimizing Framework for Effective Compute-intensive Operators Fusion · HPCA 2023

Methods — techniques the papers use, named apart from their topics

learned sparse compressors · 1.7layer-wise communication scheduling · 1.7reinforcement learning · 1.3finite state machine · 1.3dynamic programming · 1.3analytical modeling · 1.3PQ tree · 1.3tree-based analysis · 0.7simulated annealing · 0.7greedy algorithm · 0.7genetic algorithm · 0.7dataflow grouping · 0.7
YearPublicationVenuePosition
2025 Practical Offloading for Fine-Tuning LLM on Commodity GPU via Learned Sparse Projectors
abstract
Fine-tuning large language models (LLMs) requires significant memory, often exceeding the capacity of a single GPU. A common solution to this memory challenge is offloading compute and data from the GPU to the CPU. However, this approach is hampered by the limited bandwidth of commodity hardware, which constrains communication between the CPU and GPU, and by slower matrix multiplications on the CPU. In this paper, we present an offloading framework, LSP-Offload, that enables near-native speed LLM fine-tuning on commodity hardware through learned sparse projectors. Our data-driven approach involves learning efficient sparse compressors that minimize communication with minimal precision loss. Additionally, we introduce a novel layer-wise communication schedule to maximize parallelism between communication and computation. As a result, our framework can fine-tune a 1.3 billion parameter model on a 4GB laptop GPU and a 6.7 billion parameter model on an NVIDIA RTX 4090 GPU with 24GB memory. Compared to state-of-the-art offloading frameworks, our approach reduces end-to-end fine-tuning time by 33.1%-62.5% when converging to the same accuracy.
Siyuan Chen 0007, Zhuofeng Wang, Zelong Guan, Phillip B. Gibbons
AAAI1
2023 Memory and Computation Coordinated Mapping of DNNs onto Complex Heterogeneous SoC
abstract
The DNN models are now pervasively used for various applications. Meanwhile, the computing hardware has shifted towards heterogeneous system composed of various accelerators. The intertwined complexity of DNN models and hardware makes it challenging for mapping DNN models. Existing mapping frameworks suffer from inefficiencies due to under utilization of computation and bandwidth in heterogeneous SoC. In this paper, we propose COMB, a mapping framework that coordinates the memory and computation and data transfer overhead of heterogeneous accelerators to achieve latency improvement and energy efficiency with two optimizations: dataflow grouping and accelerator mapping. Dataflow grouping maps multiple independent DNN layers to the same accelerator at the same time to spatially share the hardware resources; accelerator mapping finds the optimized placement of the layer groups to accelerators to reduce data transfer overhead. These two optimizations provide a huge design space for heterogeneous DNN mapping. To explore the space efficiently, we present a hybrid scheduling algorithm by combining greedy algorithm and genetic algorithm. In evaluation, COMB achieves 1.28× and 1.37× speedup for latency compared to MAGMA and H2H; COMB also reduces 22.7% and 29.2% energy consumption compared to MAGMA and H2H.
Size Zheng 0001, Siyuan Chen 0007, Yun Liang 0001
DAC2
2023 Chimera: An Analytical Optimizing Framework for Effective Compute-intensive Operators Fusion
abstract
Machine learning models with various tensor operators are becoming ubiquitous in recent years. There are two types of operators in machine learning: compute-intensive operators (e.g., GEMM and convolution) and memory-intensive operators (e.g., ReLU and softmax). In emerging machine learning models, compute-intensive operators are usually organized in a chain structure. With the continual specialization of hardware, the gap between computing performance and memory bandwidth has become more prominent. Consequently, the implementations of many compute-intensive operator chains are bounded by memory bandwidth, and generating fused kernels to improve locality for these compute-intensive operators becomes necessary. But in existing machine learning compilers, there lack both precise analysis and efficient optimization for compute-intensive operator chains on different accelerators. As a result, they usually produce sub-optimal performance for these operator chains.In this paper, we propose Chimera, an optimizing framework that can efficiently improve the locality of compute-intensive operator chains on different hardware accelerators. In Chimera, each compute-intensive operator is composed of a series of computation blocks. To generate efficient fused kernels for the operator chains, optimizations for both inter-block and intra-block are required. For inter-block optimization, Chimera decides the optimized block execution order by minimizing the data movement volume among blocks using an analytical model. For intra-block optimization, Chimera uses unified replaceable micro kernels to apply hardware-specific optimizations for different accelerators. Finally, Chimera generates fused kernels for compute-intensive operator chains. Evaluation of batch GEMM chains and convolution chains on CPU, GPU, and NPU shows that Chimera achieves up to 2.87×, 2.29×, and 2.39× speedups to hand-tuned libraries. Compared to state-of-the-art compilers, the speedups are up to 2.29×, 1.64×, and 1.14× for CPU, GPU, and NPU.
Size Zheng 0001, Siyuan Chen 0007, Peidi Song, Renze Chen, Shengen Yan, Dahua Lin, Jingwen Leng, Yun Liang 0001
HPCA2
2023 ED-Batch: Efficient Automatic Batching of Dynamic Neural Networks via Learned Finite State Machines
abstract
Batching has a fundamental influence on the efficiency of deep neural network (DNN) execution. However, for dynamic DNNs, efficient batching is particularly challenging as the dataflow graph varies per input instance. As a result, state-of-the-art frameworks use heuristics that result in suboptimal batching decisions. Further, batching puts strict restrictions on memory adjacency and can lead to high data movement costs. In this paper, we provide an approach for batching dynamic DNNs based on finite state machines, which enables the automatic discovery of batching policies specialized for each DNN via reinforcement learning. Moreover, we find that memory planning that is aware of the batching policy can save significant data movement overheads, which is automated by a PQ tree-based algorithm we introduce. Experimental results show that our framework speeds up state-of-the-art frameworks by on average 1.15x, 1.39x, and 2.45x for chain-based, tree-based, and lattice-based DNNs across CPU and GPU. The framework is open-sourced at https://github.com/gulang2019/ED-Batch.git.
Siyuan Chen 0007, Pratik Fegade, Tianqi Chen 0001, Phillip B. Gibbons, Todd C. Mowry
ICML1
2023 TileFlow: A Framework for Modeling Fusion Dataflow via Tree-based Analysis
abstract
With the increasing size of DNN models and the growing discrepancy between compute performance and memory bandwidth, fusing multiple layers together to reduce off-chip memory access has become a popular approach in dataflow design. However, designing such dataflows requires flexible and accurate performance models to facilitate evaluation, architecture analysis, and design space exploration. Unfortunately, current state-of-the-art performance models are limited to the dataflows of single operator acceleration, making them inapplicable to operator fusion dataflows.
Size Zheng 0001, Siyuan Chen 0007, Liancheng Jia, Guangyu Sun 0003, Runsheng Wang, Yun Liang 0001
MICRO2