EDBT 2026 Demo / reviewers in the wild / expert
Tianhao Cai
dblp:340/3915
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2025
0009-0008-9601-113XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Hardware accelerators and domain-specific architectures · 55% Memory systems · 24% Processor architecture and microarchitecture · 11% | |
| Artificial intelligence
2 papers |
Efficient and distributed learning · 54% 3D vision · 46% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
cache management |
0.9 | 1 | 2025 | CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs · DAC 2025 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
neural processing unit |
0.9 | 1 | 2025 | CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs · DAC 2025 |
Memory systems › cache management › cache partitioning
shared cache partitioning |
0.9 | 1 | 2025 | CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs · DAC 2025 |
Processor architecture and microarchitecture › computer arithmetic
bit-serial architecture |
0.8 | 1 | 2024 | BitNN: A Bit-Serial Accelerator for K-Nearest Neighbor Search in Point Clouds · ISCA 2024 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN accelerator |
0.8 | 1 | 2024 | ReDas: A Lightweight Architecture for Supporting Fine-Grained Reshaping and Multiple Dataflows on Systolic Array · IEEE Trans. Computers 2024 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
k-nearest neighbor search accelerator |
0.8 | 1 | 2024 | BitNN: A Bit-Serial Accelerator for K-Nearest Neighbor Search in Point Clouds · ISCA 2024 |
Hardware accelerators and domain-specific architectures › vision accelerator
point cloud accelerator |
0.8 | 1 | 2024 | BitNN: A Bit-Serial Accelerator for K-Nearest Neighbor Search in Point Clouds · ISCA 2024 |
Reconfigurable computing and FPGAs › reconfigurable computing
reconfigurable dataflow |
0.8 | 1 | 2024 | ReDas: A Lightweight Architecture for Supporting Fine-Grained Reshaping and Multiple Dataflows on Systolic Array · IEEE Trans. Computers 2024 |
Hardware accelerators and domain-specific architectures
systolic array |
0.8 | 1 | 2024 | ReDas: A Lightweight Architecture for Supporting Fine-Grained Reshaping and Multiple Dataflows on Systolic Array · IEEE Trans. Computers 2024 |
Computer vision › 3D vision
point cloud processing |
0.2 | 1 | 2024 | BitNN: A Bit-Serial Accelerator for K-Nearest Neighbor Search in Point Clouds · ISCA 2024 |
Methods — techniques the papers use, named apart from their topics
dynamic cache allocation · 1.7cache-aware mapping · 1.7point-wise data layout · 1.5early termination · 1.5dimension-wise encoding · 1.5bit-serial computation · 1.5reconfigurable data paths · 0.8multi-mode data buffers · 0.8mapper · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUsabstractWith the rapid development of DNN applications, multi-tenant execution, where multiple DNNs are co-located on a single SoC, is becoming a prevailing trend. Although many methods are proposed in prior works to improve multi-tenant performance, the impact of shared cache is not well studied. This paper proposes CaMDN, an architecture-scheduling co-design to enhance cache efficiency for multi-tenant DNNs on integrated NPUs. Specifically, a lightweight architecture is proposed to support model-exclusive, NPU-controlled regions inside shared cache to eliminate unexpected cache contention. Moreover, a cache scheduling method is proposed to improve shared cache utilization. In particular, it includes a cache-aware mapping method for adaptability to the varying available cache capacity and a dynamic allocation algorithm to adjust the usage among co-located DNNs at runtime. Compared to prior works, CaMDN reduces the memory access by 33.4% on average and achieves a model speedup of up to 2.56 × (1.88 × on average). Tianhao Cai, Liang Wang 0020, Limin Xiao 0001, Xiaojian Liao |
DAC | 1 |
| 2024 | BitNN: A Bit-Serial Accelerator for K-Nearest Neighbor Search in Point CloudsabstractPoint cloud-based machine perception applications have achieved great success in various scenarios. In this work, we focus on point cloud k-Nearest Neighbor (kNN) search, an important kernel for point clouds. Existing kNN acceleration techniques have overlooked the operation-level optimization in the Euclidean distance computation operations, which suffer from low efficiency due to a number of unnecessary computations and various data precision requirements.We reconsider point cloud kNN search from a new bitserial computation perspective and propose BitNN, a bit-serial architecture for point cloud kNN search. BitNN supports adaptive precision processing and unnecessary computing reduction, significantly improving the performance and power efficiency of kNN search. To achieve that, we first propose a bit-serial computation method for kNN search, which derives a recursive expression to compute the Euclidean distance bit by bit. Then, the dimension-wise point cloud encoding method and point-wise data layout method are proposed to enable adaptive precision processing based on bit-serial computation. Furthermore, we present an early termination mechanism for bit-serial kNN search. By estimating the lower bound of distance based on a few bits, a number of unnecessary computations can be reduced. Finally, we design an efficient bit-serial accelerator for kNN search. The accelerator exploits the massive parallelism to improve computing efficiency.We evaluate BitNN with several widely used point cloud datasets. BitNN achieves up to $6.6 \times$ speedup and $3.6 \times$ power efficiency compared to a comparable sized architecture. Moreover, BitNN can be easily integrated into existing bit-parallel kNN accelerators. We enhance the state-of-the-art kNN accelerator, ParallelNN, with bit-serial computation techniques, achieving up to $4.4 \times$ speedup and $2.9 \times$ power efficiency Liang Wang 0020, Limin Xiao 0001, Hao Zhang 0205, Tianhao Cai, Xiangrong Xu 0002 |
ISCA | 5 |
| 2024 | ReDas: A Lightweight Architecture for Supporting Fine-Grained Reshaping and Multiple Dataflows on Systolic ArrayabstractThe systolic accelerator is one of the premier architectural choices for DNN acceleration. However, the conventional systolic architecture suffers from low PE utilization due to the mismatch between the fixed array and diverse DNN workloads. Recent studies have proposed flexible systolic array architectures to adapt to DNN models. However, these designs support only coarse-grained reshaping or significantly increase hardware overhead. In this study, we propose ReDas, a flexible and lightweight systolic array that supports dynamic fine-grained reshaping and multiple dataflows. First, ReDas integrates lightweight and reconfigurable roundabout data paths, which achieve fine-grained reshaping using only short connections between adjacent PEs. Second, we redesign the PE microarchitecture and integrate a set of multi-mode data buffers around the array. The PE structure enables additional data bypassing and flexible data switching. Simultaneously, the multi-mode buffers facilitate fine-grained reallocation of on-chip memory resources, adapting to various dataflow requirements. ReDas can dynamically reconfigure to up to 129 different logical shapes and 3 dataflows for a 128 × 128 array. Finally, we propose an efficient mapper to generate appropriate configurations for each layer of DNN workloads. Compared to the conventional systolic array, ReDas can achieve about 4.6× speedup and 8.3× energy-delay product (EDP) reduction. Liang Wang 0020, Limin Xiao 0001, Tianhao Cai, Xiangrong Xu 0002 |
IEEE Trans. Computers | 4 |