Jun Wang 0175

dblp:125/8189-175 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
0009-0003-1907-2979ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 65% Processor architecture and microarchitecture · 35%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
SIMD
0.812024
SOPHGO BM1684X: A Commercial High Performance Terminal AI Processor with Large Model Support · MICRO 2024
Hardware accelerators and domain-specific architectures › tensor accelerator
tensor processing unit
0.812024
SOPHGO BM1684X: A Commercial High Performance Terminal AI Processor with Large Model Support · MICRO 2024
Processor architecture and microarchitecture
instruction set architecture
0.312025
OASIS: A Commercial High Performance Terminal AI Processor Supporting RISC-V Tensor Extension Instructions · MICRO 2025
Processor architecture and microarchitecture › instruction set architecture
RISC-V
0.312025
OASIS: A Commercial High Performance Terminal AI Processor Supporting RISC-V Tensor Extension Instructions · MICRO 2025

Methods — techniques the papers use, named apart from their topics

tensor extension instructions · 0.9crossbar · 0.8TPU-MLIR toolchain · 0.8SIMD · 0.8
YearPublicationVenuePosition
2025 OASIS: A Commercial High Performance Terminal AI Processor Supporting RISC-V Tensor Extension Instructions
Peng Gao 0016, Yang Liu 0038, Haonan Sun 0006, Jun Wang 0175, Zonghui Hong, Jiali Qu
MICRO5
2024 SOPHGO BM1684X: A Commercial High Performance Terminal AI Processor with Large Model Support
abstract
This paper presents BM1684X, a cutting-edge AI processor from SOPHGO designed to meet the demanding requirements of broad AI applications. Firstly, we employ SIMD architecture with very large data width to design our TPU to reduce the area ratio of the instruction unit and greatly improves the computing power density. Secondly, the customization of special acceleration instructions within the EU enables the dynamic pipeline execution, leading to a reduction in the total number of instructions and execution time. This customization enhances the performance of TPU in processing RQ and DQ operations, crucial for AI computations. Thirdly, the CUBE array within the TPU implements the multiplication and addition operations of 64 pairs of INT8 operands in the channel dimension of the feature map. By utilizing an addition tree instead of a conventional adder, the implementation significantly reduces both area and power consumption, optimizing the efficiency of TPU. Additionally, the BM1684X processor incorporates a 64-input, 64-output, 8-bit crossbar within the lane, facilitating high-performance data gathering. This crossbar design enhances data gathering capabilities, enabling efficient data processing and manipulation within the TPU architecture. Furthermore, BM1684X offers three distinct memory access modes, showing the processor's versatility in addressing a wide range of AI processing needs and optimizing DRAM utilization for various tasks and workloads. Finally, we design a TPU-MLIR toolchain, highlighting its rich features such as unified processing of multiple frameworks, hierarchical design of model abstractions, correctness guarantees, and traceability of each transformation step. BM1684X excels in providing high-performance computing for a variety of AI models including large models, demonstrating its capabilities through comprehensive evaluations with industry-leading peers.
Peng Gao 0016, Yang Liu 0038, Jun Wang 0175, Wanlin Cai, Guangchong Shen, Zonghui Hong, Jiali Qu
MICRO3