Aojie Jiang

dblp:141/9979 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
2since 2021 · last 2024
0009-0006-1271-4475ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 50% Hardware reliability and fault tolerance · 25% Hardware accelerators and domain-specific architectures · 25%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
0.812024
GroupQ: Group-Wise Quantization With Multi-Objective Optimization for CNN Accelerators · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024
Machine learning › Efficient and distributed learning › model compression
quantization
0.812024
GroupQ: Group-Wise Quantization With Multi-Objective Optimization for CNN Accelerators · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024
Compilers and program optimization
compiler infrastructure
0.812024
A Compilation Framework for SRAM Computing-in-Memory Systems With Optimized Weight Mapping and Error Correction · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024
Hardware accelerators and domain-specific architectures › machine learning accelerator
CNN accelerator
0.812024
GroupQ: Group-Wise Quantization With Multi-Objective Optimization for CNN Accelerators · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024
Memory systems › processing-in-memory
computing-in-memory
0.812024
A Compilation Framework for SRAM Computing-in-Memory Systems With Optimized Weight Mapping and Error Correction · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024
Hardware reliability and fault tolerance
error correction
0.812024
A Compilation Framework for SRAM Computing-in-Memory Systems With Optimized Weight Mapping and Error Correction · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024
Memory systems › processing-in-memory › computing-in-memory
in-SRAM computing
0.812024
A Compilation Framework for SRAM Computing-in-Memory Systems With Optimized Weight Mapping and Error Correction · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024

Methods — techniques the papers use, named apart from their topics

weight mapping · 1.5multi-objective optimization · 1.5lookup table · 1.5error correction · 1.5clustering · 1.5calibration · 1.5
YearPublicationVenuePosition
2024 A Compilation Framework for SRAM Computing-in-Memory Systems With Optimized Weight Mapping and Error Correction
abstract
Deploying convolution-based algorithms into SRAM computing-in-memory (CIM) systems faces various challenges, such as operator incompatibility and intrinsic non-ideal error. This paper proposes a compilation framework to address this issue. Efficient weight mapping strategies are introduced to improve the utilization of SRAM-CIM macro. The intrinsic non-ideal errors of SRAM-CIM macro are also taken into consideration, and two efficient error correction schemes are proposed, which include calibration of computation voltage linear error (CCVLE) and the mitigation of analog-to-digital quantization error (MAQE). In addition, bit-width flexibility and signed-unsigned reconfigurability are also supported to facilitate the deployment of various convolution-based algorithms. ResNet18, finite impulse response (FIR) filtering, and Gaussian image filtering are deployed into a multi-macro SRAM-CIM system. These algorithms serve as deployment representatives of convolutional neural network (CNN), digital signal processing (DSP), and digital image processing (DIP), respectively. The results show that the introduced weight mapping strategies improve the macro utilization by 63.29% and 21.10% for two types of frequently used convolution layers compared to the traditional strategy. Moreover, the proposed error correction schemes achieve similar algorithm accuracy to the floating-point results, and the deployment result of ResNet18 achieves 66.3%~70.1% top-1 classification accuracy evaluated on the ImageNet dataset with different throughput tradeoffs.
Yichuan Bai, Yaqing Li, Heng Zhang 0024, Aojie Jiang, Yuan Du
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2024 GroupQ: Group-Wise Quantization With Multi-Objective Optimization for CNN Accelerators
abstract
Mixed-precision Neural Networks achieve high energy efficiency and throughput for hardware deployment. The most common mixed-precision methods adopt layer-wise granularity. However, the layer-wise method does not quantize the model to its limit because the optimal bit precision that preserves accuracy for different kernels can be different. To address this issue, this paper presents GroupQ, a group-wise quantization method with multi-objective optimization for CNN accelerators. Group-wise divides the convolutional kernels in a layer into several groups by clustering, and each group shares the same bit precision. The multi-objective optimization algorithm is used to optimize the quantization policy automatically based on the selected quantization objectives, such as model accuracy, model size, or computation cost. The experiments show that GroupQ significantly outperforms the existing layer-wise retraining-free methods, even better than some training-based methods. Specifically, GroupQ achieves a 0.49% higher accuracy with up to 28.1% smaller Bit Operations (BOPs) on ResNet-18 compared to HAWQ-V3 and can quantize MobileNetV2 to 1.65MB model size with 71.75% top-1 accuracy. This paper shows that GroupQ is friendly for hardware deployment by a lookup table (LUT)-based mixed-precision processing element (LMPE) proposed for CNN accelerators. LMPE provides power reduction of up to 3.6%, up to 3.9% lower area, compared to conventional implementation.
Aojie Jiang, Yuan Du
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2013 Rate-efficient error robustness for IDR frames through edge-based redundancy maps
abstract
This paper proposes an efficient way of using edge based redundant information for improving the error resilience of IDR frames. The proposed method generates spatial error concealment mode and edge direction data at the encoder which are sent to the decoder to guide the concealment process. Results show that the method offers very good performance with only a very small increase in complexity at the decoder and the introduction of negligible overhead to the coded stream.
Aojie Jiang, Dimitris Agrafiotis
ICIP1