Tianhao Cai

dblp:340/3915 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
0009-0008-9601-113XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Hardware accelerators and domain-specific architectures · 55% Memory systems · 24% Processor architecture and microarchitecture · 11%
Artificial intelligence
2 papers
Efficient and distributed learning · 54% 3D vision · 46%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache management
0.912025
CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs · DAC 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
neural processing unit
0.912025
CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs · DAC 2025
Memory systems › cache management › cache partitioning
shared cache partitioning
0.912025
CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs · DAC 2025
Processor architecture and microarchitecture › computer arithmetic
bit-serial architecture
0.812024
BitNN: A Bit-Serial Accelerator for K-Nearest Neighbor Search in Point Clouds · ISCA 2024
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN accelerator
0.812024
ReDas: A Lightweight Architecture for Supporting Fine-Grained Reshaping and Multiple Dataflows on Systolic Array · IEEE Trans. Computers 2024
Hardware accelerators and domain-specific architectures › machine learning accelerator
k-nearest neighbor search accelerator
0.812024
BitNN: A Bit-Serial Accelerator for K-Nearest Neighbor Search in Point Clouds · ISCA 2024
Hardware accelerators and domain-specific architectures › vision accelerator
point cloud accelerator
0.812024
BitNN: A Bit-Serial Accelerator for K-Nearest Neighbor Search in Point Clouds · ISCA 2024
Reconfigurable computing and FPGAs › reconfigurable computing
reconfigurable dataflow
0.812024
ReDas: A Lightweight Architecture for Supporting Fine-Grained Reshaping and Multiple Dataflows on Systolic Array · IEEE Trans. Computers 2024
Hardware accelerators and domain-specific architectures
systolic array
0.812024
ReDas: A Lightweight Architecture for Supporting Fine-Grained Reshaping and Multiple Dataflows on Systolic Array · IEEE Trans. Computers 2024
Computer vision › 3D vision
point cloud processing
0.212024
BitNN: A Bit-Serial Accelerator for K-Nearest Neighbor Search in Point Clouds · ISCA 2024

Methods — techniques the papers use, named apart from their topics

dynamic cache allocation · 1.7cache-aware mapping · 1.7point-wise data layout · 1.5early termination · 1.5dimension-wise encoding · 1.5bit-serial computation · 1.5reconfigurable data paths · 0.8multi-mode data buffers · 0.8mapper · 0.8
YearPublicationVenuePosition
2025 CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs
abstract
With the rapid development of DNN applications, multi-tenant execution, where multiple DNNs are co-located on a single SoC, is becoming a prevailing trend. Although many methods are proposed in prior works to improve multi-tenant performance, the impact of shared cache is not well studied. This paper proposes CaMDN, an architecture-scheduling co-design to enhance cache efficiency for multi-tenant DNNs on integrated NPUs. Specifically, a lightweight architecture is proposed to support model-exclusive, NPU-controlled regions inside shared cache to eliminate unexpected cache contention. Moreover, a cache scheduling method is proposed to improve shared cache utilization. In particular, it includes a cache-aware mapping method for adaptability to the varying available cache capacity and a dynamic allocation algorithm to adjust the usage among co-located DNNs at runtime. Compared to prior works, CaMDN reduces the memory access by 33.4% on average and achieves a model speedup of up to 2.56 × (1.88 × on average).
Tianhao Cai, Liang Wang 0020, Limin Xiao 0001, Xiaojian Liao
DAC1
2024 BitNN: A Bit-Serial Accelerator for K-Nearest Neighbor Search in Point Clouds
abstract
Point cloud-based machine perception applications have achieved great success in various scenarios. In this work, we focus on point cloud k-Nearest Neighbor (kNN) search, an important kernel for point clouds. Existing kNN acceleration techniques have overlooked the operation-level optimization in the Euclidean distance computation operations, which suffer from low efficiency due to a number of unnecessary computations and various data precision requirements.We reconsider point cloud kNN search from a new bitserial computation perspective and propose BitNN, a bit-serial architecture for point cloud kNN search. BitNN supports adaptive precision processing and unnecessary computing reduction, significantly improving the performance and power efficiency of kNN search. To achieve that, we first propose a bit-serial computation method for kNN search, which derives a recursive expression to compute the Euclidean distance bit by bit. Then, the dimension-wise point cloud encoding method and point-wise data layout method are proposed to enable adaptive precision processing based on bit-serial computation. Furthermore, we present an early termination mechanism for bit-serial kNN search. By estimating the lower bound of distance based on a few bits, a number of unnecessary computations can be reduced. Finally, we design an efficient bit-serial accelerator for kNN search. The accelerator exploits the massive parallelism to improve computing efficiency.We evaluate BitNN with several widely used point cloud datasets. BitNN achieves up to $6.6 \times$ speedup and $3.6 \times$ power efficiency compared to a comparable sized architecture. Moreover, BitNN can be easily integrated into existing bit-parallel kNN accelerators. We enhance the state-of-the-art kNN accelerator, ParallelNN, with bit-serial computation techniques, achieving up to $4.4 \times$ speedup and $2.9 \times$ power efficiency
Liang Wang 0020, Limin Xiao 0001, Hao Zhang 0205, Tianhao Cai, Xiangrong Xu 0002
ISCA5
2024 ReDas: A Lightweight Architecture for Supporting Fine-Grained Reshaping and Multiple Dataflows on Systolic Array
abstract
The systolic accelerator is one of the premier architectural choices for DNN acceleration. However, the conventional systolic architecture suffers from low PE utilization due to the mismatch between the fixed array and diverse DNN workloads. Recent studies have proposed flexible systolic array architectures to adapt to DNN models. However, these designs support only coarse-grained reshaping or significantly increase hardware overhead. In this study, we propose ReDas, a flexible and lightweight systolic array that supports dynamic fine-grained reshaping and multiple dataflows. First, ReDas integrates lightweight and reconfigurable roundabout data paths, which achieve fine-grained reshaping using only short connections between adjacent PEs. Second, we redesign the PE microarchitecture and integrate a set of multi-mode data buffers around the array. The PE structure enables additional data bypassing and flexible data switching. Simultaneously, the multi-mode buffers facilitate fine-grained reallocation of on-chip memory resources, adapting to various dataflow requirements. ReDas can dynamically reconfigure to up to 129 different logical shapes and 3 dataflows for a 128 × 128 array. Finally, we propose an efficient mapper to generate appropriate configurations for each layer of DNN workloads. Compared to the conventional systolic array, ReDas can achieve about 4.6× speedup and 8.3× energy-delay product (EDP) reduction.
Liang Wang 0020, Limin Xiao 0001, Tianhao Cai, Xiangrong Xu 0002
IEEE Trans. Computers4