EDBT 2026 Demo / reviewers in the wild / expert
Huiying Lan
dblp:191/7785
· DBLP profile ↗
14ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0003-3120-5773ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Data-Driven Dynamic Execution Orchestration ArchitectureabstractDomain-specific accelerators deliver exceptional performance on their target workloads through fabrication-time orchestrated datapaths. However, such specialized architectures often exhibit performance fragility when exposed to new kernels or irregular input patterns. In contrast, programmable architectures like FPGAs, CGRAs, and GPUs rely on compile-time orchestration to support a broader range of applications; but they are typically less efficient under irregular or sparse data. Pushing the boundaries of programmable architectures requires designs that can achieve efficiency and high-performance on par with specialized accelerators while retaining the agility of general-purpose architectures. Zhenyu Bai, Pranav Dangi, Rohan Juneja, Zhaoying Li 0004, Zhanglu Yan, Huiying Lan, Tulika Mitra |
ASPLOS (1) | 6 |
| 2026 | HighP: In-Memory Acceleration of SpGEMM With High Bank-Level ParallelismabstractGeneralized sparse matrix-matrix multiplication (SpGEMM) is a critical computational primitive that is highly memory-bound due to its inherent irregular data-dependent access pattern. Near-bank processing-in-memory (PIM) is a promising technique to overcome the memory bottleneck of SpGEMM by performing computations near the bank where the data is stored. However, earlier PIM studies fail to fully utilize the high memory bandwidth when performing SpGEMM due to low bank-level parallelism. As a result, 80% memory bandwidth is wasted as observed in our in-depth experimental analysis.Our key insight in this paper is that non-conflicting matrix columns in SpGEMM, where each row of these columns has no more than one non-zero element, can be processed simultaneously in different banks. We hence propose HighP, a near-bank PIM accelerator for SpGEMM with high bank-level parallelism. We first propose a set-based search mechanism, which finds non-conflicting columns through set operations automatically. We then develop a DIMM-based PIM architecture with detailed hardware and workflow designs for SpGEMM. Set operation logic and unified scratchpad memory management are designed to perform set operations with high computational parallelism and to enhance data reuse, respectively. HighP provides up to 17.88× performance improvement compared to the state-of-th-eart SpGEMM accelerator and achieves up to 8.19× performance improvement over the state-of-the-art PIM solution. Dan Chen 0006, Huize Li, Huiying Lan, Zhaoying Li 0004, Pengcheng Yao, Tulika Mitra |
IEEE Trans. Computers | 3 |
| 2025 | Efficient Partitioning Deep Learning Models for Medical Image Analysis on Iot DevicesabstractDeep learning models, due to their strong capabilities in learning from data and representing features, have been widely deployed, particularly in the medical field. However, given the limitations of medical devices and deployment scenarios, more and more researchers are focusing on how to better deploy deep learning models on resource-constrained edge devices to enable real-time medical data analysis. Nevertheless, such resourceconstrained IoT devices present significant challenges, making it difficult to achieve both high accuracy and low inference latency. To address this problem, we propose a novel framework, RC-DLM, designed to split and execute complex deep learning models on IoT devices. Specifically, we partition a deep learning model into multiple sub-models according to the computational capacity of each device, with each sub-model responsible for handling a subset of classes. To further reduce computation overhead and inference latency, we integrate a class-wise pruning method to shrink the size of each sub-model. Through large-scale experiments conducted on four popular datasets with three model architectures, we demonstrate that our approach significantly reduces inference latency and model size by up to 5.72 times and 57.5 times, respectively. We further deploy our method on real-world edge devices and compare it with state-of-theart approaches, evaluating the three most important aspects: accuracy, inference time, and model size. The comprehensive experimental results of our RC-DLM confirm the effectiveness of our proposed method. Xiang Liu 0017, Mengyao Zheng, Junyong Cao, Dehui Wei, Kang Lai, Huiying Lan, Yijun Song, Xia Li 0005 |
BIBM | 6 |
| 2025 | High Performance Computing Framework for Secure Variable Selection on Genome-Wide Association Studies with Adaptive Vertical Federated LearningabstractVariable selection for genome-wide association studies (GWAS) has long been a central focus in academic research. However, with the advent of the big data era and the rapid growth of biomedical and healthcare data, scientists are increasingly challenged to extract meaningful information from massive datasets. Worse still, such data are often distributed across multiple parties, making collaborative analysis necessary while also requiring strong privacy preservation. To date, there is still no effective framework that can support high-dimensional data analysis, ensure data privacy across collaborators, and simultaneously capture the relatedness between explanatory and response variables. To address these challenges, we introduce the first high-performance computing framework for variable selection in GWAS with Vertical Federated Learning, termed VS-VFL. Our approach leverages Vertical Federated Learning to enable seamless multi-party collaboration while maintaining data privacy and security. Furthermore, we integrate a wide range of state-of-the-art methods, allowing collaborators to apply their preferred techniques, and we explicitly account for the noni.i.d. nature of biomedical data when analyzing the relatedness between explanatory and response variables. In addition, we employ novel optimization strategies and adaptive algorithms to efficiently handle high-dimensional data with sparse features. This framework empowers researchers to conduct comprehensive analyses and perform accurate linkage mapping of gene associations. Our framework is implemented in Python, can be easily deployed on any platform, and is designed to make advanced GWAS analysis accessible to a broader community of researchers. Mengyao Zheng, Xiang Liu 0017, Junyong Cao, Huiying Lan, Liangxi Liu, Xia Li 0005 |
BIBM | 5 |
| 2025 | Efficient Partitioning Vision Transformer on Edge Devices for Distributed InferenceabstractDeep learning models are increasingly utilized on resource-constrained edge devices for real-time data analytics. Recently, Vision Transformer and their variants have shown exceptional performance in various computer vision tasks. However, their substantial computational requirements and low inference latency create significant challenges for deploying such models on resource-constrained edge devices. To address this issue, we propose a novel framework, ED-ViT, which is designed to efficiently split and execute complex Vision Transformers across multiple edge devices. Our approach involves partitioning Vision Transformer models into several sub-models, while each dedicated to handling a specific subset of data classes. To further reduce computational overhead and inference latency, we introduce a class-wise pruning technique that decreases the size of each sub-model. Through extensive experiments conducted on five datasets using three model architectures and actual implementation on edge devices, we demonstrate that our method significantly cuts down inference latency on edge devices and achieves a reduction in model size by up to 28.9 times and 34.1 times, respectively, while maintaining test accuracy comparable to the original Vision Transformer. Additionally, we compare ED-ViT with two state-of-the-art methods that deploy CNN and SNN models on edge devices, evaluating metrics such as accuracy, inference time, and overall model size. Our comprehensive evaluation underscores the effectiveness of the proposed ED-ViT framework. Xiang Liu 0017, Yijun Song, Xia Li 0005, Huiying Lan, Linshan Jiang, Jialin Li 0001 |
ICDCS | 5 |
| 2025 | Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCsabstractAs edge-based deep learning applications become more complex, optimizing performance on heterogeneous System-on-Chips (SoCs) presents unique challenges. Traditional pipelining techniques distributing the computation across different on-chip processing units, while effective for throughput, do not address the latency demands posed by modern neural networks with complex interdependencies and extensive operator parallelism. There is a potential in leveraging operator parallelism to enable concurrent execution across multiple processing units, thereby reducing inference latency. However, prioritizing pipelining or parallel execution often necessitates a compromise, where optimizing one performance metric adversely impacts the other. This paper introduces Para-Pipe, a hierarchical mapping framework that integrates intra-and inter-stage operator parallelism within a pipelined architecture. Para-Pipe navigates the trade-off between throughput and latency by selectively fine-tuning parallelism levels within and across pipeline stages. This strategy can significantly reduce inter-processor communication overhead, significantly improving energy efficiency. Our evaluation demonstrates that Para-Pipe generates multiple Pareto-optimal configurations, achieving a balance between throughput and latency on an Amlogic SoC equipped with ARM big.LITTLE CPUs and GPU, as well as the Black Sesame Technology SoC featuring a deep learning accelerator and two DSPs. More importantly, throughput-optimized configurations under Para-Pipe on Amlogic SoC show an average energy efficiency improvement of 11.0% over purely pipelined strategies and 23.3% relative to non-pipelined parallel execution. Yujie Zhang 0007, Huiying Lan, Ehsan Aghapour, Peng Zan, Weidong Shao, Anuj Pathania, Tulika Mitra |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | FusionFrame: A Fusion Dataflow Scheduling Framework for DNN Accelerators via Analytical Modeling
Liutao Zheng, Huiying Lan, Xiang Liu 0017, Linshan Jiang, Xuehai Zhou |
ICA3PP (6) | 2 |
| 2019 | Addressing Sparsity in Deep Neural NetworksabstractNeural networks (NNs) have been demonstrated to be useful in a broad range of applications, such as image recognition, automatic translation, and advertisement recommendation. State-of-the-art NNs are known to be both computationally and memory intensive, due to the ever-increasing deep structure, i.e., multiple layers with massive neurons and connections (i.e., synapses). Sparse NNs have emerged as an effective solution to reduce the amount of computation and memory required. Though existing NN accelerators are able to efficiently process dense and regular networks, they cannot benefit from the reduction of synaptic weights. In this paper, we propose a novel accelerator, Cambricon-X, to exploit the sparsity and irregularity of NN models for increased efficiency. The proposed accelerator features a processing element (PE)-based architecture consisting of multiple PEs. An indexing module efficiently selects and transfers needed neurons to connected PEs with reduced bandwidth requirement, while each PE stores irregular and compressed synapses for local computation in an asynchronous fashion. With 16 PEs, our accelerator is able to achieve at most 544 GOP/s in a small form factor (6.38 mm2and 954 mW at 65 nm). Experimental results over a number of representative sparse networks show that our accelerator achieves, on average, $7.23\times$ speedup and $6.43\times$ energy saving against the state-of-the-art NN accelerator. We further investigate possibilities of leveraging activation sparsity and multi-issue controller, which improve the efficiency of Cambricon-X. To ease the burden of programmers, we also propose a high efficient library-based programming environment for our accelerator. Xuda Zhou, Zidong Du, Shijin Zhang, Lei Zhang 0008, Huiying Lan, Shaoli Liu, Ling Li 0001, Qi Guo 0001, Tianshi Chen 0002, Yunji Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2018 | DLIR: An Intermediate Representation for Deep Learning Processors
Huiying Lan, Zidong Du |
NPC | 1 |
| 2018 | Leveraging Subgraph Extraction for Performance Portable Programming Frameworks on DL Accelerators
Huiying Lan, Tian Zhi |
NPC | 2 |
| 2018 | BenchIP: Benchmarking Intelligence Processors
Jinhua Tao, Zidong Du, Qi Guo 0001, Huiying Lan, Lei Zhang 0008, Shengyuan Zhou, Lingjie Xu, Shan Tang, Allen Rush, Willian Chen, Shaoli Liu, Yunji Chen, Tianshi Chen 0002 |
J. Comput. Sci. Technol. | 4 |
| 2018 | An Instruction Set Architecture for Machine LearningabstractMachine Learning (ML) are a family of models for learning from the data to improve performance on a certain task. ML techniques, especially recent renewed neural networks (deep neural networks), have proven to be efficient for a broad range of applications. ML techniques are conventionally executed on general-purpose processors (such as CPU and GPGPU), which usually are not energy efficient, since they invest excessive hardware resources to flexibly support various workloads. Consequently, application-specific hardware accelerators have been proposed recently to improve energy efficiency. However, such accelerators were designed for a small set of ML techniques sharing similar computational patterns, and they adopt complex and informative instructions (control signals) directly corresponding to high-level functional blocks of an ML technique (such as layers in neural networks) or even an ML as a whole. Although straightforward and easy to implement for a limited set of similar ML techniques, the lack of agility in the instruction set prevents such accelerator designs from supporting a variety of different ML techniques with sufficient flexibility and efficiency. In this article, we first propose a novel domain-specific Instruction Set Architecture (ISA) for NN accelerators, called Cambricon, which is a load-store architecture that integrates scalar, vector, matrix, logical, data transfer, and control instructions, based on a comprehensive analysis of existing NN techniques. We then extend the application scope of Cambricon from NN to ML techniques. We also propose an assembly language, an assembler, and runtime to support programming with Cambricon, especially targeting large-scale ML problems. Our evaluation over a total of 16 representative yet distinct ML techniques have demonstrated that Cambricon exhibits strong descriptive capacity over a broad range of ML techniques and provides higher code density than general-purpose ISAs such as x86, MIPS, and GPGPU. Compared to the latest state-of-the-art NN accelerator design DaDianNao [7] (which can only accommodate three types of NN techniques), our Cambricon-based accelerator prototype implemented in TSMC 65nm technology incurs only negligible latency/power/area overheads, with a versatile coverage of 10 different NN benchmarks and 7 other ML benchmarks. Compared to the recent prevalent ML accelerator PuDianNao, our Cambricon-based accelerator is able to support all the ML techniques as well as the 10 NNs but with only approximate 5.1% performance loss. Yunji Chen, Huiying Lan, Zidong Du, Shaoli Liu, Jinhua Tao, Qi Guo 0001, Ling Li 0001, Yuan Xie 0001, Tianshi Chen 0002 |
ACM Trans. Comput. Syst. | 2 |
| 2017 | DLPlib: A Library for Deep Learning Processor
Huiying Lan, Linyang Wu, Jinhua Tao, Xunyu Chen, Bingrui Wang, Yu-Qing Wang, Qi Guo 0001, Yunji Chen |
J. Comput. Sci. Technol. | 1 |
| 2016 | Cambricon-X: An accelerator for sparse neural networksabstractNeural networks (NNs) have been demonstrated to be useful in a broad range of applications such as image recognition, automatic translation and advertisement recommendation. State-of-the-art NNs are known to be both computationally and memory intensive, due to the ever-increasing deep structure, i.e., multiple layers with massive neurons and connections (i.e., synapses). Sparse neural networks have emerged as an effective solution to reduce the amount of computation and memory required. Though existing NN accelerators are able to efficiently process dense and regular networks, they cannot benefit from the reduction of synaptic weights. In this paper, we propose a novel accelerator, Cambricon-X, to exploit the sparsity and irregularity of NN models for increased efficiency. The proposed accelerator features a PE-based architecture consisting of multiple Processing Elements (PE). An Indexing Module (IM) efficiently selects and transfers needed neurons to connected PEs with reduced bandwidth requirement, while each PE stores irregular and compressed synapses for local computation in an asynchronous fashion. With 16 PEs, our accelerator is able to achieve at most 544 GOP/s in a small form factor (6.38 mm2and 954 mW at 65 nm). Experimental results over a number of representative sparse networks show that our accelerator achieves, on average, 7.23x speedup and 6.43x energy saving against the state-of-the-art NN accelerator. Shijin Zhang, Zidong Du, Lei Zhang 0008, Huiying Lan, Shaoli Liu, Ling Li 0001, Qi Guo 0001, Tianshi Chen 0002, Yunji Chen |
MICRO | 4 |