Yifan Hao 0001

dblp:118/3316-1 · DBLP profile ↗
← Back
35ranked-venue papers
1as first author
35since 2021 · last 2026
0009-0002-9823-2573ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 22 · 1 first-author · 22 since 2021Artificial intelligence and machine learning · 9 · 9 since 2021Software engineering, systems software and programming languages · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 COMET: An FP32 Matrix Multiplication Accelerator by Extending INT8-Based Arrays
Ruiyang Xia, Yifan Hao 0001, Yongwei Zhao 0001, Zidong Du, Xing Hu 0001, Qi Guo 0001
APPT2
2026 Hardwired-Neuron Language Processing Units as General-Purpose Cognitive Substrates
abstract
The rapid advancement of Large Language Models (LLMs) has established language as a core general-purpose cognitive substrate, driving the demand for specialized Language Processing Units (LPUs) tailored for LLM inference. To overcome the growing energy consumption of LLM inference systems, this paper proposes a Hardwired-Neurons Language Processing Unit (HNLPU), which physically hardwires LLM weight parameters into the computational fabric, achieving several orders of magnitude computational efficiency improvement by extreme specialization. However, a significant challenge still lies in the scale of modern LLMs. A straightforward hardwiring of GPT-OSS-120B would require fabricating photomask sets valued at over 6 billion dollars, rendering this straightforward solution economically impractical.
Yang Liu 0466, Yongwei Zhao 0001, Yifan Hao 0001, Zifu Zheng, Weihao Kong, Zhangmai Li, Dongchen Jiang, Ruiyang Xia, Zhihong Ma, Zisheng Liu, Zhaoyong Wan, Yunqi Lu, Hongrui Guo, Zhe Wang 0017, Tianrui Ma, Mo Zou, Rui Zhang 0040, Ling Li 0001, Xing Hu 0001, Zidong Du, Zhiwei Xu 0002, Qi Guo 0001, Tianshi Chen 0002, Yunji Chen
ASPLOS (2)4
2026 Cambricon-CIM: Enabling Energy-Efficient and Error-Resilient Analog CIM Acceleration via Reformation of Coding Bases
abstract
Recently, multi-bit slicing has emerged as a promising technique to improve the energy efficiency of charge-domain Compute-In-Memory (CIM) accelerators by reducing the number of Analog-to-Digital (A/D) conversions. However, multi-bit slicing requires shift-and-add operations to reconstruct outputs, which exponentially amplify errors and cause significant accuracy degradation. Existing works mainly rely on hardware-aware retraining or noise-suppression techniques, incurring considerable design or power overhead. Thus, multi-bit CIM designs often face the dilemma of trading off energy efficiency for error resilience. In this paper, we propose Cambricon-CIM, a charge-domain multi-bit CIM accelerator that achieves both high energy efficiency and strong error resilience, without requiring retraining. The core insight is that the error amplification is proportional to digit weights; and by redefining these digit weights with smaller non-binary coding bases, it is possible to reduce the total error amplification. Leveraging this principle, CambriconCIM dynamically selects the minimal coding bases for every analog dot-product. With novel circuit and architectural support, Cambricon-CIM enables fast, low-overhead reconfiguration of coding bases at runtime. Experimental results show that Cambricon-CIM achieves 2.27× energy efficiency and 3.06× performance over RAELLA, a state-of-the-art error-resilient multi-bit slicing CIM architecture.
Hongrui Guo, Tianrui Ma, Zidong Du, Mo Zou, Yifan Hao 0001, Yongwei Zhao 0001, Rui Zhang 0040, Wei Li 0008, Xing Hu 0001, Zhiwei Xu 0002, Qi Guo 0001, Tianshi Chen 0002
HPCA5
2026 Sycophancy in vision-language models: A systematic analysis and an inference-time mitigation framework
Yunpu Zhao, Rui Zhang 0040, Junbin Xiao, Changxin Ke, Ruibo Hou, Yifan Hao 0001, Ling Li 0001
Neurocomputing6
2026 Cambricon-QM: A Hybrid Architecture for Microscaling Format Training
Yongwei Zhao 0001, Chang Liu 0021, Zidong Du, Xing Hu 0001, Yimin Zhuang, Yifan Hao 0001, Xinkai Song, Wei Li 0008, Xishan Zhang, Ling Li 0001, Zhiwei Xu 0002, Tianshi Chen 0002, Qi Guo 0001
IEEE Trans. Computers8
2026 AGON: Automated Design Framework for Customizing Processors From ISA Documents
abstract
Customized processors are essential for domain-specific applications such as the Internet of Things (IoT) and multi-media embedded systems, yet their design often requires extensive expert intervention. Traditional approaches, including hardware design using encapsulated abstractions (e.g., Chisel) and high-level synthesis (HLS) from languages like C or SystemC, reduce some manual efforts but remain either costly or suboptimal. Recent explorations into leveraging Large Language Models (LLMs) to generate RTL from natural language specifications have shown promise, but these methods still struggle with generating complex and high-performance processors mainly due to the complicated low-level details in the RTL code. In this work, we introduce AGON, a novel framework designed to facilitate the development of customized processor RTL from instruction set architecture (ISA) documents using LLMs. The framework comprises two layers: a functional description layer and a hardware implementation layer. At the functional layer, AGON employs a nano-operator (nOP)-based Intermediate Representation (IR) that abstracts basic instruction operations, thereby reducing the semantic gap between natural language and RTL code. This abstraction significantly shortens the descriptive code required for LLM generation, improving the generation accuracy in single-pass. At the hardware layer, AGON offers three abstraction levels (i.e. instruction, ISA, and processor) along with rule-based primitives to systematically lower the nOP-based IR into a fully optimized processor implementation. This decoupled design not only ensures correctness-by-construction but also enables automated, PPA-aware performance optimization. We evaluate AGON by designing high-performance out-of-order processors that correctly execute practical programs. Experimental results demonstrate that processors generated with AGON achieve an average speedup of 4.51× on specific tasks compared to expert-designed general-purpose CPUs while requiring minimal design effort.
Chongxiao Li, Pengwei Jin, Tianyun Ma, Husheng Han, Shuyao Cheng, Yifan Hao 0001, Yongwei Zhao 0001, Guanglin Xu, Zidong Du, Rui Zhang 0040, Xiaqing Li, Yuanbo Wen 0001, Xing Hu 0001, Qi Guo 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2026 DASA: Distribution-Aware Sparse Attention for Accelerating Diffusion Transformer
abstract
Diffusion Transformers (DiTs) have demonstrated remarkable success in text-to-video generation. However, the self-attention mechanism in DiTs imposes significant computational and memory burdens, particularly when handling long patch sequences like high-resolution or long-time videos. While sparse attention shows promise in reducing self-attention costs, existing approaches struggle to deliver performance gains due to the unique challenges in DiTs,i.e., varied sparse patterns across layers and timesteps, and the cumulative nature of inference errors over timesteps. In this paper, we propose DASA, an algorithm-hardware co-design that effectively addresses these challenges of attention sparsification in DiTs. Specifically, leveraging the insight that the generation quality is primarily influenced by overall distribution drift rather than changes in specific values, we introduce a novel Distribution-Aware Filtering (DAF) mechanism for sparsification. To further accelerate the process, we design a specialized Filtering Unit that enables fast candidate selection based on the proposed DAF mechanism. Experimental results show that DASA achieves 2.52× speed up compared to A100 GPU, and up to 1.22× speedup over state-of-the-art accelerators for self-attention computation.
Tianyun Ma, Jiaming Guo, Xinkai Song, Husheng Han, Pengwei Jin, Xiangtao Guan, Yifan Hao 0001, Yuanbo Wen 0001, Shuyao Cheng, Zidong Du, Rui Zhang 0040, Xing Hu 0001, Qi Guo 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.9
2025 Mosaic: Exploiting Instruction-Level Parallelism on Deep Learning Accelerators with iTex Tessellation
abstract
Deep learning has achieved great success in numerous application areas at the cost of high computational complexity. To meet the ever-increasing computational demand, commodity hardware platforms (e.g., CPUs and GPUs) offer abundant computing resources including scalar, vector, and tensor units for deep learning that could execute in parallel. However, existing top-down tiling-based deep learning compilers often generate a homogeneous mapping from the given tensor computation task to hardware arithmetic instructions, failing to utilize different computing units simultaneously to achieve higher performance.
Jianxing Xu, Yuanbo Wen 0001, Ruibai Xu, Tingfeng Ruan, Jun Bi, Rui Zhang 0040, Xinkai Song, Yifan Hao 0001, Xing Hu 0001, Zidong Du, Chongqing Zhao, Jiang Jie, Qi Guo 0001
ASPLOS (2)10
2025 SEEN-DA: SEmantic ENtropy guided Domain-aware Attention for Domain Adaptive Object Detection
abstract
Domain adaptive object detection (DAOD) aims to generalize detectors trained on an annotated source domain to an unlabelled target domain. Traditional works focus on aligning visual features between domains to extract domain-invariant knowledge, and recent VLM-based DAOD methods leverage semantic information provided by the textual encoder to supplement domain-specific features for each domain. However, they overlook the role of semantic information in guiding the learning of visual features that are beneficial for adaptation. To solve the problem, we propose semantic entropy to quantify the semantic information contained in visual features, and design SEmantic ENtropy guided Domain-aware Attention (SEEN-DA) to adaptively refine visual features with the semantic information of two domains. Semantic entropy reflects the importance of features based on semantic information, which can serve as attention to select discriminative visual features and suppress semantically irrelevant redundant information. Guided by semantic entropy, we introduce domain-aware attention modules into the visual encoder in SEEN-DA. It utilizes an inter-domain attention branch to extract domain-invariant features and eliminate redundant information, and an intra-domain attention branch to supplement the domain-specific semantic information discriminative on each domain. Comprehensive experiments validate the effectiveness of SEEN-DA, demonstrating significant improvements in cross-domain object detection performance.
Haochen Li 0002, Rui Zhang 0040, Hantao Yao, Xin Zhang 0062, Yifan Hao 0001, Xinkai Song, Shaohui Peng, Yongwei Zhao 0001, Ling Li 0001
CVPR5
2025 Automated Superscalar Processor Design by Learning Data Dependencies
abstract
Automated processor design, which can significantly reduce human efforts and accelerate design cycles, has received considerable attention. While recent advancements have automatically designed single-cycle processors that execute one instruction per cycle, their performance cannot compete with modern superscalar processors that execute multiple instructions per cycle. Previous methods fail on superscalar processor design because they cannot address inter-instruction data dependencies, leading to inefficient sequential instruction execution. This paper proposes a novel approach to automatically designing superscalar processors using a hardware-friendly model called the Stateful Binary Speculation Diagram (State-BSD). We observe that processor parallelism can be enhanced through on-the-fly inter-instruction dependent data predictors, reusing the processor's internal states to learn the data dependency. To meet the challenge of both hardware-resource limitation and design functional correctness, State-BSD consists of two components: 1) a lightweight state-selector trained by simulated annealing method to detect the most reusable processor states and store them in a small buffer; and 2) a highly precise state-speculator trained by BSD expansion method to predict the inter-instruction dependent data using the selected states. It is the first work to achieve the automated superscalar processor design, i.e. QiMeng-CPU-v2, which improves the performance by about 380x than the state-of-the-art automated design and is comparable to human-designed superscalar processors such as ARM Cortex A53.
Shuyao Cheng, Rui Zhang 0040, Wenkai He, Pengwei Jin, Chongxiao Li, Zidong Du, Xing Hu 0001, Yifan Hao 0001, Guanglin Xu, Yuanbo Wen 0001, Ling Li 0001, Qi Guo 0001, Yunji Chen
IJCAI8
2025 Cambricon-SR: An Accelerator for Neural Scene Representation with Sparse Encoding Table
abstract
Neural Scene Representation (NSR) is a promising technique for representing real scenes.By learning from dozens of 2D photos captured from different viewpoints, NSR computes the 3D representation of real scenes.However, the performance of NSR processing running on GPU is insufficient for applications.Cambricon-R achieves high performance of more than 60 scenes per second, but at the cost of modeling quality.
Tianbo Liu 0006, Xinkai Song, Zhifei Yue, Xing Hu 0001, Zhuoran Song, Yuanbo Wen 0001, Yifan Hao 0001, Wei Li 0008, Zidong Du, Rui Zhang 0040, Jiaming Guo, Shaohui Peng, Guangzhong Sun, Qi Guo 0001, Tianshi Chen 0002
ISCA8
2025 QiMeng-Xpiler: Transcompiling Tensor Programs for Deep Learning Systems with a Neural-Symbolic Approach
Shouyang Dong, Jun Bi, Jiaming Guo, Jianxing Xu, Ruibai Xu, Xinkai Song, Yifan Hao 0001, Ling Li 0001, Xuehai Zhou, Tianshi Chen 0002, Qi Guo 0001, Yunji Chen
OSDI8
2025 SaaP: Rearchitect SoC-as-a-Processor to Orchestrate Hardware Heterogeneity
abstract
Due to the end of Moore’s Law and Dennard Scaling, Domain-Specific Accelerators (DSAs) have come to a Cambrian explosion. Especially when advancing into the intelligent era, more and more DSAs are integrated into System-on-Chips (SoCs) as intellectual property (IP) blocks to provide high performance and efficiency. Currently, IPs usually expose IP-dependent hardware interfaces, requiring SoCs to manage them as isolated devices with software running on the host CPU. However, such software-managed heterogeneity in CPU-centric SoCs leads to low IP utilization. This inefficiency arises from the dependence on software optimization, coupled with the control and data exchange overheads. To improve IP utilization of heterogeneous SoCs, in this article, we rearchitect the SoC as a processor (i.e., SaaP) to orchestrate hardware heterogeneity. SaaP features an orchestration pipeline where DSAs are integrated as execution units and managed directly by the hardware pipeline to conceal the hardware heterogeneity from software. Moreover, SaaP redesigns the register file and data paths to implement an IP-level data-forwarding mechanism, avoiding the costly control and data exchange in the CPU-centric execution model. Block data dependence among different DSAs is carefully resolved to exploit mixed-level parallelism and inter-IP data exchange. SaaP abstracts tasks as mixed-scale instructions, where each instruction can be mapped to different IPs. Experimental results show that compared against Xavier on six fully software-optimized benchmarks from different domains, SaaP-rearchitected Xavier achieves a$2.08{\times }$speedup, with an 8.21% area reduction and only 2.98% increase in power consumption.
Pengwei Jin, Zhe Fan, Yongwei Zhao 0001, Zidong Du, Hongrui Guo, Ziyuan Nan, Yifan Hao 0001, Chongxiao Li, Tianyun Ma, Xiaqing Li, Wei Li 0008, Xing Hu 0001, Qi Guo 0001, Zhiwei Xu 0002, Tianshi Chen 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2025 VariPar: Variation-Aware Workload Partitioning in Chiplet-Based DNN Accelerators
abstract
Chiplet-based DNN accelerators have been extensively explored to save design and manufacturing costs. Previous works regard all chiplets as identical and employ uniform workload partitioning strategies. These workload partitioning strategies overlook various real-world factors that contribute to remarkable performance variations among chiplets, including manufacturing process variation, thermal condition, physical placement, and power supply condition. When considering these performance variations, a variation-aware workload partitioning can achieve superior performance. This paper introduces VariPar, a systematic framework to employ variation-aware partitioning strategy in chiplet-based DNN accelerators. VariPar models performance variations for each chiplet and partition workloads accordingly. VariPar includes a simulator with multi-factor variation modeling and a heuristic search engine to generate near-optimal partitioning within a reasonable time. Experiment results show that VariPar achieves 1.45× performance and 1.82× energy efficiency improvement on average when compared to uniform partitioning strategy.
Yongwei Zhao 0001, Mo Zou, Yang Liu 0466, Yifan Hao 0001, Xiaqing Li, Rui Zhang 0040, Yuanbo Wen 0001, Xing Hu 0001, Zidong Du, Qi Guo 0001, Tianshi Chen 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2024 Emergent Communication for Numerical Concepts Generalization
abstract
Research on emergent communication has recently gained significant traction as a promising avenue for the linguistic community to unravel human language's origins and explore artificial intelligence's generalization capabilities. Current research has predominantly concentrated on recognizing qualitative patterns of object attributes(e.g., shape and color) and paid little attention to the quantitative relationship among object quantities which is known as the part of numerical concepts. The ability to generalize numerical concepts, i.e., counting and calculations with unseen quantities, is essential, as it mirrors humans' foundational abstract reasoning abilities. In this work, we introduce the NumGame, leveraging the referential game framework, forcing agents to communicate and generalize the numerical concepts effectively. Inspired by the human learning process of numbers, we present a two-stage training approach that sequentially fosters a rudimentary numerical sense followed by the ability of arithmetic calculation, ultimately aiding agents in generating semantically stable and unambiguous language for numerical concepts. The experimental results indicate the impressive generalization capabilities to unseen quantities and regularity of the language emergence from communication.
Enshuai Zhou, Yifan Hao 0001, Rui Zhang 0040, Zidong Du, Xishan Zhang, Xinkai Song, Chao Wang 0003, Xuehai Zhou, Jiaming Guo, Qi Yi, Shaohui Peng, Ruizhi Chen, Qi Guo 0001, Yunji Chen
AAAI2
2024 TensorTEE: Unifying Heterogeneous TEE Granularity for Efficient Secure Collaborative Tensor Computing
abstract
Heterogeneous collaborative computing with NPU and CPU has received widespread attention due to its substantial performance benefits. To ensure data confidentiality and integrity during computing, Trusted Execution Environments (TEE) is considered a promising solution because of its comparatively lower overhead. However, existing heterogeneous TEE designs are inefficient for collaborative computing due to fine and different memory granularities between CPU and NPU. 1) The cacheline granularity of CPU TEE intensifies memory pressure due to its extra memory access, and 2) the cacheline granularity MAC of NPU escalates the pressure on the limited memory storage. 3) Data transfer across heterogeneous enclaves relies on the transit of non-secure regions, resulting in cumbersome re-encryption and scheduling.
Husheng Han, Xinyao Zheng, Yuanbo Wen 0001, Yifan Hao 0001, Erhu Feng, Ling Liang 0003, Jianan Mu, Xiaqing Li, Tianyun Ma, Pengwei Jin, Xinkai Song, Zidong Du, Qi Guo 0001, Xing Hu 0001
ASPLOS (4)4
2024 Automated CPU Design by Learning from Input-Output Examples
Shuyao Cheng, Pengwei Jin, Qi Guo 0001, Zidong Du, Rui Zhang 0040, Xing Hu 0001, Yongwei Zhao 0001, Yifan Hao 0001, Xiangtao Guan, Husheng Han, Zhengyue Zhao, Xishan Zhang, Yuejie Chu, Weilong Mao, Tianshi Chen 0002, Yunji Chen
IJCAI8
2024 Cambricon-D: Full-Network Differential Acceleration for Diffusion Models
abstract
Diffusion models have made significant progress in current image generation tasks, thus becoming a prominent area of research. Diffusion models necessitate repetitive iterations on minimally altered input data across timesteps, each timestep requiring the recalculation of the entire model, resulting in a remarkable computational redundancy and substantial hardware expenditures.Performing differential computing on input data seems to be a feasible approach for addressing such computational redundancy and improving hardware efficacy. However, non-linear operations (particularly activation functions) necessitate the merging of deltas (i.e., differential values) with raw inputs repeatedly to ensure computational correctness, leading to significant memory access for loading raw inputs, which fragmentedly blocks the forwarding of deltas throughout the network and undermines performance.To solve this problem, we propose Cambricon-D, a fullnetwork differential computing architecture with concise memory access. While maintaining the computational efficiency brought by differential computing, Cambricon-D employs a sign-mask dataflow, which requires only the loading of 1-bit signs (instead of large bitwidth raw inputs), thereby facilitating the seamless forwarding of deltas and effectively mitigating memory access overheads. Experimental results show that, compared to Diffy, Cambricon-D’s dataflow reduces 66% ~ 82% off-chip memory access. In total, Cambricon-D achieves 1.46× ~ 2.38× speedup over A100 on various diffusion models with different resolutions.
Weihao Kong, Yifan Hao 0001, Qi Guo 0001, Yongwei Zhao 0001, Xinkai Song, Xiaqing Li, Mo Zou, Zidong Du, Rui Zhang 0040, Chang Liu 0021, Yuanbo Wen 0001, Pengwei Jin, Xing Hu 0001, Wei Li 0008, Zhiwei Xu 0002, Tianshi Chen 0002
ISCA2
2024 Cambricon-C: Efficient 4-Bit Matrix Unit via Primitivization
abstract
Deep learning trends to use low precision numeral formats to cope with the ever-growing model sizes. For example, the large language model LLaMA2 has been widely deployed in 4-bit precision. With larger models and fewer unique values caused by low precision, an increasing proportion of arithmetic in matrix multiplication is repeating. Although discussed in prior works, such value redundancy has not been fully exploited, and the cost to leverage the value redundancy often offsets any advantages. In this paper, we propose to primitivize the matrix multiplication, that is decomposing it down to the 1-ary successor function (a.k.a. counting) to merge repeating arithmetic. We revisited various techniques to propose Cambricon-C SA, a 4-bit primitive matrix multiplication unit that doubles the energy efficiency over conventional systolic arrays. Experimental results show that Cambricon-C SA can achieve$\mathbf{1}.\mathbf{95}\times$energy efficiency improvement compared with MAC-based systolic array.
Yongwei Zhao 0001, Yifan Hao 0001, Yuanbo Wen 0001, Yuntao Dai, Xiaqing Li, Yang Liu 0466, Rui Zhang 0040, Mo Zou, Xinkai Song, Xing Hu 0001, Zidong Du, Huaping Chen 0001, Qi Guo 0001, Tianshi Chen 0002
MICRO3
2024 Cambricon-M: A Fibonacci-Coded Charge-Domain SRAM-Based CIM Accelerator for DNN Inference
abstract
Charge-domain SRAM-based Computing-in-memory (CIM) proves to be a promising method for DNN inference, and benefits from avoiding data movement between computing units and memory. However, the high resolution Analog-to-Digital Converters (ADCs) dominates the energy consumption (up to 64%), limiting the energy efficiency of SRAM-CIM architectures. The main reason is the wide range of input analog values, requiring high resolution ADCs to convert the high precision averaged analog voltages into high bitwidth digital data. In this paper, to reduce the ADC overhead, we propose Cambricon-M, a novel Fibonacci-coded SRAM-based charge-domain CIM accelerator for DNN inference. Cambricon-M features the Fibonacci coding, which guarantees low density of ‘1’ in operands (i.e., the adjacent two bits of each ‘1’ are both ‘0’), narrowing the output voltage range and enabling low resolution ADCs. Further, Cambricon-M exploits the high bit-level sparsity to address the extra energy and area overhead caused by the larger bitwidth in Fibonacci coding. Specifically, Cambricon-M proposes zero-skipping methods to reduce ineffectual input/output, and the bit-slice based compression method to reduce memory capacity/bandwidth pressure. Experimental results show that Cambricon-M reduces ADC energy by 68.7%, and improves the energy efficiency 3.48× and 1.62× compared to TPUv4 and an ISAAC-based charge-domain SRAM-CIM accelerator.
Hongrui Guo, Mo Zou, Yifan Hao 0001, Zidong Du, Erxiang Ren, Yang Liu 0466, Yongwei Zhao 0001, Tianrui Ma, Rui Zhang 0040, Xing Hu 0001, Fei Qiao, Zhiwei Xu 0002, Qi Guo 0001, Tianshi Chen 0002
MICRO3
2024 SRender: Boosting Neural Radiance Field Efficiency via Sensitivity-Aware Dynamic Precision Rendering
abstract
Neural Radiance Field (NeRF) holds immense promise for generating photo-realistic images and videos. How-ever, the computation and memory demands significantly impede its applicability in real-time tasks such as virtual reality and augmented reality. While accelerators dedicatedly designed for NeRF have emerged, they may not be the optimal solution as they only consider the optimization of full-precision NeRF models, leading to unsatisfactory memory and computation saving. This paper proposes Sensitivity-Aware Dynamic Precision Rendering, abbreviated as SRender, an algorithm-hardware co-design framework that borrows the idea of Adaptive Rendering commonly used in conventional graphics rendering to accelerate the rendering process of NeRF exploiting the fine-grained per-ray and per-point levels sensitivity. The fundamental idea behind SRender is that rays and points with high sensitivity require high-precision data types, while the insensitive rays and points do not, allowing them to be sacrificed for greater acceleration. Accordingly, the algorithm part of SRender dynamically configures the precision of the hash table and the multilayer perceptron (MLP), and the hardware part features a specialized accelerator that exploits such mixed-precision memory access and computation patterns. Specifically, we predict the sensitive rays and points based on low-precision hash encoding and feature computation. Then, we trigger a high-precision rendering again on the sensitive ones to guarantee the rendering quality. The hardware design of SRender includes a sensitivity prediction engine, coarse-then-fine-grained encoding engines, and a recovery engine, which translates the computational saving in the algorithm to real speedup. Using the coarse-then-fine-grained encoding engine, we maximize the data reuse on-chip, thereby reducing the stress of off-chip memory bandwidth. Experiments show that the edge-level SRender can achieve 3.94x and 34.06x speedup over state-of-the-art NeRF accelerator NeuRex and Jetson AGX Xavier GPU with negligible rendering quality loss.
Zhuoran Song, Houshu He, Fangxin Liu, Yifan Hao 0001, Xinkai Song, Li Jiang 0002, Xiaoyao Liang
MICRO4
2024 Cambricon-LLM: A Chiplet-Based Hybrid Architecture for On-Device Inference of 70B LLM
abstract
Deploying advanced large language models on edge devices, such as smartphones and robotics, is a growing trend that enhances user data privacy and network connectivity resilience while preserving intelligent capabilities. However, such a task exhibits single-batch computing with incredibly low arithmetic intensity, which poses the significant challenges of huge memory footprint and bandwidth demands on limited edge resources. To address these issues, we introduce Cambricon-LLM, a chiplet-based hybrid architecture with NPU and a dedicated NAND flash chip to enable efficient on-device inference of 70B LLMs. Such a hybrid architecture utilizes both the high computing capability of NPU and the data capacity of the NAND flash chip, with the proposed hardware-tiling strategy that minimizes the data movement overhead between NPU and NAND flash chip. Specifically, the NAND flash chip, enhanced by our innovative in-flash computing and on-die ECC techniques, excels at performing precise lightweight on-die processing. Simultaneously, the NPU collaborates with the flash chip for matrix operations and handles special function computations beyond the flash's on-die processing capabilities. Overall, Cambricon-LLM enables the on-device inference of 70B LLMs at a speed of 3.44 token/s, and 7B LLMs at a speed of 36.34 token/s, which is over 22× to 45× faster than existing flash-offloading technologies, showing the potentiality of deploying powerful LLMs in edge devices.
Zhongkai Yu, Shengwen Liang, Tianyun Ma, Yunke Cai, Ziyuan Nan, Xinkai Song, Yifan Hao 0001, Jie Zhang 0048, Tian Zhi, Yongwei Zhao 0001, Zidong Du, Xing Hu 0001, Qi Guo 0001, Tianshi Chen 0002
MICRO8
2024 DA-Ada: Learning Domain-Aware Adapter for Domain Adaptive Object Detection
abstract
Domain adaptive object detection (DAOD) aims to generalize detectors trained on an annotated source domain to an unlabelled target domain. As the visual-language models (VLMs) can provide essential general knowledge on unseen images, freezing the visual encoder and inserting a domain-agnostic adapter can learn domain-invariant knowledge for DAOD. However, the domain-agnostic adapter is inevitably biased to the source domain. It discards some beneficial knowledge discriminative on the unlabelled domain, \ie domain-specific knowledge of the target domain. To solve the issue, we propose a novel Domain-Aware Adapter (DA-Ada) tailored for the DAOD task. The key point is exploiting domain-specific knowledge between the essential general knowledge and domain-invariant knowledge. DA-Ada consists of the Domain-Invariant Adapter (DIA) for learning domain-invariant knowledge and the Domain-Specific Adapter (DSA) for injecting the domain-specific knowledge from the information discarded by the visual encoder. Comprehensive experiments over multiple DAOD tasks show that DA-Ada can efficiently infer a domain-aware visual encoder for boosting domain adaptive object detection. Our code is available at https://github.com/Therock90421/DA-Ada.
Haochen Li 0002, Rui Zhang 0040, Hantao Yao, Xin Zhang 0062, Yifan Hao 0001, Xinkai Song, Xiaqing Li, Yongwei Zhao 0001, Yunji Chen, Ling Li 0001
NeurIPS5
2024 Environmental Condition Aware Super-Resolution Acceleration Framework in Server-Client Hierarchies
abstract
In the current landscape, high-resolution (HR) videos have gained immense popularity, promising an elevated viewing experience. Recent research has demonstrated that the video super-resolution (SR) algorithm, empowered by deep neural networks (DNNs), can substantially enhance the quality of HR videos by processing low-resolution (LR) frames. However, the existing DNN models demand significant computational resources, posing challenges for the deployment of SR algorithms on client devices. While numerous accelerators have proposed solutions, their primary focus remains on client-side optimization. In contrast, our research recognizes that the HR video is originally stored in the cloud server and presents an untapped opportunity for achieving both high accuracy and performance improvements. Building on this insight, this article introduces an end-to-end video CODEC-assisted super-resolution (E 2 SR+) algorithm, which tightly integrates the cloud server with the client device to deliver a seamless and real-time video viewing experience. We propose the motion vector search algorithm executed in a cloud server, which can search the motion vectors and residuals for a part of the HR video frames and then pack them as add-ons. We also design an auto-encoder algorithm to down-sample the residuals to save the bitstream cost while guaranteeing the quality of the residuals. Lastly, we propose a reconstruction algorithm performed in the client to quickly reconstruct the corresponding HR frames using the add-ons to skip part of the DNN computations. To implement the E 2 SR+ algorithm, we design corresponding E 2 SR+ architecture in the client, which achieves significant speedup with minimal hardware overhead. Given that the environmental condition varies in the server–client hierarchies, we believe that simply applying E 2 SR+ to all frames is irrational. Accordingly, we offer an environmental condition–aware system to chase the best performance while adapting to the diverse environment. In the system, we design a linear programming (LP) model to simulate the environment and allocate frames to three existing mechanisms. Our experimental results demonstrate that the E 2 SR+ algorithm enhances the peak signal-to-noise ratio by 1.2, 2.5, and 2.3 compared with the state-of-the-art (SOTA) methods EDVR, BasicVSR, and BasicVSR++, respectively. In terms of performance, the E 2 SR+ architecture offers significant improvements over existing SOTA methods. For instance, while BasicVSR++ requires 98 ms on an NVIDIA V100 graphics processing unit (GPU) to generate a 1,280 × 720 HR frame, the E 2 SR+ architecture reduces the execution time to just 39 ms, highlighting the efficiency and effectiveness of our proposed method. Overall, the E 2 SR+ architecture respectively achieves 1.4×, 2.2×, 4.6×, and 442.0× performance improvement compared with ADAS, ISRAcc, the NVIDIA V100 GPU, and a central processing unit. Lastly, the proposed system showcases its superiority and surpasses all the existing mechanisms in terms of execution time when varying environmental conditions.
Zhuoran Song, Zhongkai Yu, Xinkai Song, Yifan Hao 0001, Li Jiang 0002, Naifeng Jing, Xiaoyao Liang
ACM Trans. Archit. Code Optim.4
2024 Tyche: An Efficient and General Prefetcher for Indirect Memory Accesses
abstract
Indirect memory accesses (IMAs, i.e., A [ f ( B [ i ])]) are typical memory access patterns in applications such as graph analysis, machine learning, and database. IMAs are composed of producer-consumer pairs, where the consumers’ memory addresses are derived from the producers’ memory data. Due to the built-in value-dependent feature, IMAs exhibit poor locality, making prefetching ineffective. Hindered by the challenges of recording the potentially complex graphs of instruction dependencies among IMA producers and consumers, current state-of-the-art hardware prefetchers either (a) exhibit inadequate IMA identification abilities or (b) rely on the run-ahead mechanism to prefetch IMAs intermittently and insufficiently. To solve this problem, we propose Tyche, 1 an efficient and general hardware prefetcher to enhance IMA performance. Tyche adopts a bilateral propagation mechanism to precisely excavate the instruction dependencies in simple chains with moderate length (rather than complex graphs). Based on the exact instruction dependencies, Tyche can accurately identify various IMA patterns, including nonlinear ones, and generate accurate prefetching requests continuously. Evaluated on broad benchmarks, Tyche achieves an average performance speedup of 16.2% over the state-of-the-art spatial prefetcher Berti. More importantly, Tyche outperforms the state-of-the-art IMA prefetchers IMP, Gretch, and Vector Runahead, by 15.9%, 12.8%, and 10.7%, respectively, with a lower storage overhead of only 0.57 KB.
Chenji Han, Xinyu Li 0010, Junliang Wu, Yifan Hao 0001, Zidong Du, Qi Guo 0001, Fuxin Zhang
ACM Trans. Archit. Code Optim.7
2024 Real-Time Robust Video Object Detection System Against Physical-World Adversarial Attacks
abstract
DNN-based video object detection (VOD) powers autonomous driving and video surveillance industries with rising importance and promising opportunities. However, adversarial patch attack yields huge concern in live vision tasks because of its practicality, feasibility, and powerful attack effectiveness. This work proposes Themis, a software/hardware system to defend against adversarial patches for real-time robust VOD. We observe that adversarial patches exhibit extremely localized superficial feature importance in a small region with nonrobust predictions, and thus propose the adversarial region detection algorithm for adversarial effect elimination. Themis also proposes a systematic design to efficiently support the algorithm by eliminating redundant computations and memory traffics. Experimental results show that the proposed methodology can effectively recover the system from the adversarial attack with negligible hardware overhead.
Husheng Han, Xing Hu 0001, Yifan Hao 0001, Kaidi Xu, Pucheng Dang, Ying Wang 0001, Yongwei Zhao 0001, Zidong Du, Qi Guo 0001, Yanzhi Wang 0001, Xishan Zhang, Tianshi Chen 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 REACT: Remainder Adaptive Compensation for Domain Adaptive Object Detection
abstract
Domain adaptive object detection (DAOD) aims to infer a robust detector on the target domain with the labelled source datasets. Recent studies utilize a feature extractor shared on the source and target domains to capture the domain-invariant features and the task-relevant information with both feature-alignment constraint and source annotations. However, the feature extractor shared across domains discards partial task-relevant information of the target domain due to the domain gap and lack of target annotations, leading to compromised discrimination capabilities within target domain. To this end, we propose a novel REmainder Adaptive CompensaTion network (REACT) to adaptively compensate the extracted features with the remainder features for generating task-relevant features. The key insight is that the remainder features contain the discarded task-relevant information, so they can be adapted to compensate for the inadequate target features. Especially, REACT introduces an additional remainder branch to regain the remainder features, and then adaptively utilizes them to compensate for the discarded task-relevant information, improving discrimination on the target domain. Extensive experiments over multiple cross-domain adaptation tasks with three baselines demonstrate that our approach gains significant improvements and achieves superior performance compared with highly-optimized state-of-the-art methods.
Haochen Li 0002, Rui Zhang 0040, Hantao Yao, Xin Zhang 0062, Yifan Hao 0001, Xinkai Song, Ling Li 0001
IEEE Trans. Image Process.5
2024 FastTuning: Enabling Fast and Efficient Hyper-Parameter Tuning With Partitioning and Parallelism of Search Space
abstract
Hyper-parameter tuning (HPT) for deep learning (DL) models is prohibitively expensive. Sequential model-based optimization (SMBO) emerges as the state-of-the-art (SOTA) approach to automatically optimize HPT performance due to its heuristic advantages. Unfortunately, focusing on algorithm optimization rather than a large-scale parallel HPT system, existing SMBO-based approaches still cannot effectively remove their strong sequential nature, posing two performance problems: (1)extremely low tuning speedand (2)sub-optimal model quality. In this paper, we propose FastTuning, a fast, scalable, and generic system aiming at parallelly accelerating SMBO-based HPT for large DL/ML models. The key is to partition the highly complex search space into multiple smaller sub-spaces, each of which is assigned to and optimized by a different tuning worker in parallel. However, determining the right level of resource allocation to strike a balance between quality and cost remains a challenge. To address this, we further propose NIMBLE, a dynamic scheduling strategy that is specially designed for FastTuning, including (1) Dynamic Elimination Algorithm, (2) Sub-space Re-division, and (3) Posterior Information Sharing. Finally, we incorporate 6 SOTAs (i.e., 3 tuning algorithms and 3 parallel tuning tools) into FastTuning. Experimental results, on ResNet18, VGG19, ResNet50, and ResNet152, show that FastTuning can consistently offer much faster tuning speed (up to$80\times$) with better accuracy (up to 4.7% improvement), thereby enabling the application of automatic HPT to real-life DL models.
Xiaqing Li, Qi Guo 0001, Guangyan Zhang, Siwei Ye, Guanhua He, Yiheng Yao, Rui Zhang 0040, Yifan Hao 0001, Zidong Du
IEEE Trans. Parallel Distributed Syst.8
2023 BALTO: fast tensor program optimization with diversity-based active learning
Jun Bi, Xiaqing Li, Qi Guo 0001, Rui Zhang 0040, Yuanbo Wen 0001, Xing Hu 0001, Zidong Du, Xinkai Song, Yifan Hao 0001, Yunji Chen
ICLR9
2023 Cambricon-U: A Systolic Random Increment Memory Architecture for Unary Computing
abstract
Unary computing, whose arithmetics require only one logic gate, has enabled efficient DNN processing, especially on strictly power-constrained devices. However, unary computing still confronts the power efficiency bottleneck for buffering unary bitstreams. The buffering of unary bitstreams requires accumulating bits into large bitwidth binary numbers. The large bitwidth binary number needs to activate all bits per cycle in case of carry propagation. As a result, the accumulation process accounts for 32%-70% of the power budget.
Hongrui Guo, Yongwei Zhao 0001, Zhangmai Li, Yifan Hao 0001, Chang Liu 0021, Xinkai Song, Xiaqing Li, Zidong Du, Rui Zhang 0040, Qi Guo 0001, Tianshi Chen 0002, Zhiwei Xu 0002
MICRO4
2023 Emergent Communication for Rules Reasoning
abstract
Research on emergent communication between deep-learning-based agents has received extensive attention due to its inspiration for linguistics and artificial intelligence. However, previous attempts have hovered around emerging communication under perception-oriented environmental settings, that forces agents to describe low-level perceptual features intra image or symbol contexts. In this work, inspired by the classic human reasoning test (namely Raven's Progressive Matrix), we propose the Reasoning Game, a cognition-oriented environment that encourages agents to reason and communicate high-level rules, rather than perceived low-level contexts. Moreover, we propose 1) an unbiased dataset (namely rule-RAVEN) as a benchmark to avoid overfitting, 2) and a two-stage curriculum agent training method as a baseline for more stable convergence in the Reasoning Game, where contexts and semantics are bilaterally drifting. Experimental results show that, in the Reasoning Game, a semantically stable and compositional language emerges to solve reasoning problems. The emerged language helps agents apply the extracted rules to the generalization of unseen context attributes, and to the transfer between different context attributes or even tasks.
Yifan Hao 0001, Rui Zhang 0040, Enshuai Zhou, Zidong Du, Xishan Zhang, Xinkai Song, Yuanbo Wen 0001, Yongwei Zhao 0001, Xuehai Zhou, Jiaming Guo, Qi Yi, Shaohui Peng, Ruizhi Chen, Qi Guo 0001, Yunji Chen
NeurIPS2
2023 Learning Domain-Aware Detection Head with Prompt Tuning
abstract
Domain adaptive object detection (DAOD) aims to generalize detectors trained on an annotated source domain to an unlabelled target domain. However, existing methods focus on reducing the domain bias of the detection backbone by inferring a discriminative visual encoder, while ignoring the domain bias in the detection head. Inspired by the high generalization of vision-language models (VLMs), applying a VLM as the robust detection backbone following a domain-aware detection head is a reasonable way to learn the discriminative detector for each domain, rather than reducing the domain bias in traditional methods. To achieve the above issue, we thus propose a novel DAOD framework named Domain-Aware detection head with Prompt tuning (DA-Pro), which applies the learnable domain-adaptive prompt to generate the dynamic detection head for each domain. Formally, the domain-adaptive prompt consists of the domain-invariant tokens, domain-specific tokens, and the domain-related textual description along with the class label. Furthermore, two constraints between the source and target domains are applied to ensure that the domain-adaptive prompt can capture the domains-shared and domain-specific knowledge. A prompt ensemble strategy is also proposed to reduce the effect of prompt disturbance. Comprehensive experiments over multiple cross-domain adaptation tasks demonstrate that using the domain-adaptive prompt can produce an effectively domain-related detection head for boosting domain-adaptive object detection. Our code is available at https://github.com/Therock90421/DA-Pro.
Haochen Li 0002, Rui Zhang 0040, Hantao Yao, Xinkai Song, Yifan Hao 0001, Yongwei Zhao 0001, Ling Li 0001, Yunji Chen
NeurIPS5
2023 Chip design with machine learning: a survey from algorithm perspective
Wenkai He, Xiaqing Li, Xinkai Song, Yifan Hao 0001, Rui Zhang 0040, Zidong Du, Yunji Chen
Sci. China Inf. Sci.4
2023 Hardware Acceleration for SLAM in Mobile Systems
Zhe Fan, Yifan Hao 0001, Tian Zhi, Qi Guo 0001, Zidong Du
J. Comput. Sci. Technol.2
2022 Cambricon-P: A Bitflow Architecture for Arbitrary Precision Computing
abstract
Arbitrary precision computing (APC), where the digits vary from tens to millions of bits, is fundamental for scientific applications, such as mathematics, physics, chemistry, and biology. APC on existing platforms (e.g., CPUs and GPUs) is achieved by decomposing the original data into small pieces to accommodate to the low-bitwidth (e.g., 32-/64-bit) functional units. However, such fine-grained decomposition inevitably introduces large amounts of intermediates, bringing in intensive on-chip data traffic and long, complex dependency chains, so that causing low hardware utilization.To address this issue, we propose Cambricon-P, a bitflow architecture supporting monolithic large and flexible bitwidth operations for efficient APC processing, which avoids generating large amounts of intermediates from decomposition. Cambricon- P features a tightly-integrated computational architecture for processing different bitflows in parallel, where full bit-serial data paths are deployed. The bit-serial scheme still needs to eliminate the dependency chain of APC for exploiting parallelism within one monolithic large-bitwidth operation. For this purpose, Cambricon-P adopts a carry parallel computing mechanism, which enables recursively transforming the multiplication into smaller inner-products that can be performed in parallel between bit-indexed IPUs (Inner-Product Units). Furthermore, to improve the computing efficiency of APC, Cambricon- P employs a bit-indexed inner-product processing scheme, namely BIPS, to eliminate intra-IPU bit-level redundancy. Compared to Intel Xeon 6134 CPU, Cambricon-P achieves 100.98$\times$ performance on monolithic long multiplication, and 23.41$\times$/30.16$\times$ speedup and energy benefit over four real-world APC applications on average. Compared to NVidia V100 GPU, Cambricon-P also delivers the same throughput, as well as 430$\times$/60.5$\times$ lesser area and power, respectively, on batch-processing multiplications.
Yifan Hao 0001, Yongwei Zhao 0001, Chenxiao Liu, Zidong Du, Shuyao Cheng, Xiaqing Li, Xing Hu 0001, Qi Guo 0001, Zhiwei Xu 0002, Tianshi Chen 0002
MICRO1