Zichen Fan

dblp:206/2833 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CEN-RTDETR: A Co-Enhancement-Based Real-Time Single-Domain Generalized Object Detection for Road Scenes
abstract
ABSTRACT In urban road scenes, the cross‐domain data distribution differences caused by light and weather changes make the generalization performance of single‐domain trained object detectors in unknown weather scenarios significantly degraded (e.g., daytime‐sunny trained models in dusk‐rainy scenarios reduce the detection accuracy by more than 46% on average); moreover, faster R‐CNN suffers from insufficient generalization capability and inefficient real‐time inference due to architectural constraints. To address the above challenges, this paper proposes the CEN‐RTDETR method to improve the single‐domain generalization capability through a collaborative enhancement strategy: CP‐Mix color channel permutation dynamically simulates the RGB color channel permutation to simulate the multi‐sky color bias phenomenon, and enhances the color robustness of the input data; NP feature normalization perturbation applies a random feature perturbation to the channel statistics of the shallow feature map to optimize the extraction of texture, color and other basic style features; CORAL Loss minimizes the difference in feature distribution between the source domain and the virtual target domain through second‐order statistical matching. The experimental results show that CEN‐RTDETR achieves significant performance improvement on the cross‐weather scenario dataset DWD: The mean average precision ([email protected]) across different weather scenarios increases from 39.66% to 42.72% (+3.06%), especially for dusk‐rainy and night‐rainy scenarios, where [email protected] rises from 32.9% and 18.6% to 39.0% (+6.1%) and 24.1% (+5.5%) in extreme weather. The method in this paper effectively solves the cross‐domain generalization and efficiency problems in single‐domain generalized object detection, which provides new technical possibilities for real‐time detection in complex urban road scenes.
Huantong Geng, Long Fang, Yingrui Wang, Zichen Fan
IET Image Process.5
2025 SQ-DM: Accelerating Diffusion Models with Aggressive Quantization and Temporal Sparsity
abstract
Diffusion models have gained significant popularity in image generation tasks. However, generating high-quality content remains notably slow because it requires running model inference over many time steps. To accelerate these models, we propose to aggressively quantize both weights and activations, while simultaneously promoting significant activation sparsity. We further observe that the stated sparsity pattern varies among different channels and evolves across time steps. To support this quantization and sparsity scheme, we present a novel diffusion model accelerator featuring a heterogeneous mixed-precision dense-sparse architecture, channel-last address mapping, and a time-step-aware sparsity detector for efficient handling of the sparsity pattern. Our 4-bit quantization technique demonstrates superior generation quality compared to existing $\mathbf{4}$-bit methods. Our custom accelerator achieves $6.91 \times$ speed-up and 51.5% energy reduction compared to traditional dense accelerators.
Zichen Fan, Steve Dai, Rangharajan Venkatesan, Dennis Sylvester, Brucek Khailany
DAC1
2024 ConSmax: Hardware-Friendly Alternative Softmax with Learnable Parameters
abstract
The self-attention mechanism distinguishes transformer-based large language models (LLMs) apart from convolutional and recurrent neural networks. Despite the performance improvement, achieving real-time LLM inference on silicon remains challenging due to the extensive use of Softmax in self-attention. In addition to the non-linearity, the low arithmetic intensity significantly limits processing parallelism, especially when working with longer contexts. To address this challenge, we propose Constant Softmax (ConSmax), a software-hardware co-design that serves as an efficient alternative to Softmax. ConSmax utilizes differentiable normalization parameters to eliminate the need for maximum searching and denominator summation in Softmax. This approach enables extensive parallelization while still executing the essential functions of Softmax. Moreover, a scalable ConSmax hardware design with a bitwidth-split look-up table (LUT) can achieve lossless non-linear operations and support mixed-precision computing. Experimental results show that ConSmax achieves a minuscule power consumption of 0.2mW and an area of 0.0008mm2 at 1250MHz working frequency in 16nm FinFET technology. For open-source contribution, we further implement our design with the OpenROAD toolchain under SkyWater's 130nm CMOS technology. The corresponding power is 2.69mW and the area is 0.007mm2. ConSmax achieves 3.35× power savings and 2.75× area savings in 16nm technology, and 3.15× power savings and 4.14× area savings with the open-source EDA toolchain. In the meantime, it also maintains comparable accuracy on the GPT-2 model and the WikiText103 dataset. The project is available at https://github.com/ReaLLMASIC/ConSmax.
Shiwei Liu 0002, Guanchen Tao, Yifei Zou, Derek Chow, Zichen Fan, Kauna Lei, Bangfei Pan, Dennis Sylvester, Gregory Kielian, Mehdi Saligane
ICCAD5
2023 SONA: An Accelerator for Transform-Domain Neural Networks with Sparse-Orthogonal Weights
abstract
Recent advances in model pruning have enabled sparsity-aware deep neural network accelerators that improve the energy-efficiency and performance of inference tasks. We introduce SONA, a novel transform-domain neural network accelerator in which convolution operations are replaced by element-wise multiplications with sparse-orthogonal weights. SONA employs an output stationary dataflow coupled with an energy-efficient memory organization to reduce the overhead of sparse-orthogonal transform-domain kernels that are concurrently processed without any conflicts. Weights in SONA are non-uniformly quantized with bit-sparse canonical-signed-digit representations to reduce multiplications to simple additions. Moreover, for sparse fully-connected layers (FCLs), SONA introduces column-based-block structured pruning, which is integrated into the same architecture that maintains full multiply-and-accumulate (MAC) array utilization. Compared to prior dense and sparse neural networks accelerators, SONA can reduce inference energy by$5.1\times$and$2.4 \times$and increase performance by$5.2\times$and$2.1\times$, respectively, for convolution layers. For sparse FCLs, SONA can reduce inference energy by$2.4\times$and increase performance by$2\times$compared to prior work.
Pierre Abillama, Zichen Fan, Yu Chen 0070, Hyochan An, Qirui Zhang 0001, Seungkyu Choi, David T. Blaauw, Dennis Sylvester, Hun-Seok Kim
ASAP2
2023 Efficient Computation Sharing for Multi-Task Visual Scene Understanding
abstract
Solving multiple visual tasks using individual models can be resource-intensive, while multi-task learning can conserve resources by sharing knowledge across different tasks. Despite the benefits of multi-task learning, such techniques can struggle with balancing the loss for each task, leading to potential performance degradation. We present a novel computation- and parameter-sharing framework that balances efficiency and accuracy to perform multiple visual tasks utilizing individually-trained single-task transformers. Our method is motivated by transfer learning schemes to reduce computational and parameter storage costs while maintaining the desired performance. Our approach involves splitting the tasks into a base task and the other sub-tasks, and sharing a significant portion of activations and parameters/weights between the base and sub-tasks to decrease inter-task redundancies and enhance knowledge sharing. The evaluation conducted on NYUD-v2 and PASCAL-context datasets shows that our method is superior to the state-of-the-art transformer-based multi-task learning techniques with higher accuracy and reduced computational resources. Moreover, our method is extended to video stream inputs, further reducing computational costs by efficiently sharing information across the temporal domain as well as the task domain. Our codes are available at https://github.com/sarashoouri/EfficientMTL.
Sara Shoouri, Mingyu Yang 0002, Zichen Fan, Hun-Seok Kim
ICCV3
2023 TaskFusion: An Efficient Transfer Learning Architecture with Dual Delta Sparsity for Multi-Task Natural Language Processing
abstract
The combination of pre-trained models and task-specific fine-tuning schemes, such as BERT, has achieved great success in various natural language processing (NLP) tasks. However, the large memory and computation costs of such models make it challenging to deploy them in edge devices. Moreover, in real-world applications like chatbots, multiple NLP tasks need to be processed together to achieve higher response credibility. Running multiple NLP tasks with specialized models for each task increases the latency and memory cost latency linearly with the number of tasks. Though there have been recent works on parameter-shared tuning that aim to reduce the total parameter size by partially sharing weights among multiple tasks, computation remains intensive and redundant despite different tasks using the same input. In this work, we identify that a significant portion of activations and weights can be reused among different tasks, to reduce cost and latency for efficient multi-task NLP. Specifically, we propose TaskFusion, an efficient transfer learning software-hardware co-design that exploits delta sparsity in both weights and activations to boost data sharing among tasks. For training, TaskFusion uses ℓ1 regularization on delta activation to learn inter-task data redundancies. A novel hardware-aware sub-task inference algorithm is proposed to exploit the dual delta sparsity. We then designed a dedicated heterogeneous architecture to accelerate multi-task inference with an optimized scheduling to increase hardware utilization and reduce off-chip memory access. Extensive experiments demonstrate that TaskFusion can reduce the number of floating point operations (FLOPs) by over 73% in multi-task NLP with negligible accuracy loss, while adding a new task at the cost of only < 2% parameter size increase. With the proposed architecture and optimized scheduling, Task-Fusion can achieve 1.48--2.43× performance and 1.62--3.77× energy efficiency than those using state-of-the-art single-task accelerators for multi-task NLP applications.
Zichen Fan, Qirui Zhang 0001, Pierre Abillama, Sara Shoouri, Changwoo Lee 0001, David T. Blaauw, Hun-Seok Kim, Dennis Sylvester
ISCA1
2023 Global Localization of Energy-Constrained Miniature RF Emitters using Low Earth Orbit Satellites
abstract
Daily tracking of small objects or animals anywhere on earth for long time-periods is a long sought-after goal. Recently, the emergence of low earth orbit (LEO) satellites offers a unique pathway to achieve this goal. However, to date, LEO trackers have not achieved cm-size. While the integrated chip can be readily scaled to sub-cm size, the size of trackers remains limited by their battery and antenna size. To address these two fundamental size limiting factors, this paper presents a LEO satellite localization system that is specifically optimized to reduce antenna size and transmit power, thereby reducing battery size. To reduce power, a new cooperative waveform is designed which enhances the localization accuracy, combined with an increased packet length to enable low transmit power while maintaining packet energy. However, this long packet length introduces a intra-packet Doppler shift which we address by proposing a localization algorithm that includes a Doppler shift correction. The final result is a 50 kHz periodic BPSK signal with 23 dBm equivalent isotropic radiation power (EIRP), and 120 ms packet length (> 10 × longer than conventional), at a 60 s interval. The proposed solution enables 7 months operation on a 2.5 × 1.2 cm LiPo battery within a North American search area. To address the antenna size, the optimal transmit frequency was studied and a 1 cm loop antenna with 65% radiation efficiency was designed with internal matching to 50 Ohm. Using the proposed techniques, three satellite flyover experiments were performed to confirm the accuracy of the proposed tracking system and localization algorithms using a USRP-X310, a custom 1 cm-size antenna, and a commercial satellite cluster. The measured average localization error is 320 - 840 m depending on satellite trajectories, demonstrating an improved accuracy in real life measurements compared to prior art with experimental result while simultaneously achieving 15 -- 26 dB lower transmit power and > 3 × lower packet energy.
Demba Komma, Jaechan Lim, Zichen Fan, Chien-Wei Tseng, Hun-Seok Kim, David T. Blaauw
SenSys6
2022 Fast Trajectory Generation and Asteroid Sequence Selection in Multispacecraft for Multiasteroid Exploration
abstract
As an increasing number of asteroids are being discovered, detecting them using limited propulsion resources and time has become an urgent problem in the aerospace field. However, there is no universal fast asteroid sequence selection method that finds the trajectories for multiple low-thrust spacecraft for detecting a large number of asteroids. Furthermore, the calculation efficiency of the traditional trajectory optimization method is low, and it requires a large number of iterations. Therefore, this study combines Monte Carlo tree search (MCTS) with spacecraft trajectory optimization. A fast MCTS pruning algorithm is proposed, which can quickly complete asteroid sequence selection and trajectory generation for multispacecraft exploration of multiple asteroids. By combining the Bezier shape-based (SB) method and MCTS, this study realizes the fast search of the exploration sequence and the efficient optimization of the continuous transfer trajectories. In the simulation example, compared with the traversal algorithm, the MCTS pruning algorithm obtained the global optimal detection sequence of the search tree in a very short time. Under the same conditions, the Bezier SB method obtained the transfer trajectory with a better performance index faster than the finite Fourier series SB method. Performances of the proposed method are illustrated through a complex asteroid multiflyby mission design.
Naiming Qi, Zichen Fan, Mingying Huo, Desong Du, Ce Zhao
IEEE Trans. Cybern.2
2020 RED: A ReRAM-Based Efficient Accelerator for Deconvolutional Computation
abstract
Deconvolution is a key component in contemporary neural networks, especially, generative adversarial networks (GANs) and fully convolutional networks (FCNs). Due to extra operations of deconvolution compared to convolution, considerable degradation of performance, as well as energy efficiency is incurred when implementing deconvolution on the existing resistive random access memory (ReRAM)-based processing-in-memory (PIM) accelerators. In this article, we propose an ReRAM-based accelerator design, RED, for providing high-performance and low-energy deconvolution. We analyze the deconvolution execution on the existing ReRAM-based PIMs and utilize its interior computation pattern for design optimization. RED includes two major contributions: 1) pixel-wise mapping scheme and 2) zero-skipping data flow. Pixel-wise mapping scheme removes the zero insertion and performs convolutions over several ReRAM arrays and thus enables parallel computations with nonzero inputs. Zero-skipping data flow, assisted with customized input buffers design, enhances the computation parallelism and input data reuse. In evaluation, we compare RED against the existing ReRAM-based PIMs and CMOS-based counterpart with a variety of GAN and FCN models, each of which contains multiple deconvolution layers. The experimental results show that RED achieves a$4.0\times $–$56.16\times $speedup and a$1.05\times $–$18.17\times $energy efficiency improvement over previous related accelerator designs.
Ziru Li, Bing Li 0017, Zichen Fan, Hai Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 ASP-SIFT: Using Analog Signal Processing Architecture to Accelerate Keypoint Detection of SIFT Algorithm
abstract
The scale-invariant feature transform (SIFT) algorithm is still one of the most reliable image feature extraction methods. Despite its excellent robustness on various image transformations, SIFT's intensive computational burden has been severely preventing it from being used in real-time and energy-efficient embedded machine vision systems. To reduce processing time and energy cost while executing SIFT, an analog signal processing architecture, analog signal processing (ASP)SIFT, is proposed in this article. In ASP-SIFT, the Gaussian pyramid construction, difference-of-Gaussian (DoG) pyramid construction and keypoint locating, which are the primary steps of the keypoint detection part of the SIFT algorithm, are done directly with analog circuit networks. Thus, by completing keypoint detection in the analog domain, the total processing time is approximately equal to the settling time of the circuit network. Besides, by adopting a current-mode circuit network operating in the subthreshold region, the power dissipation would be very low. Simulation results show that the total processing speed for a typical video graphics array (VGA)-format (640 × 480) image is up to 2.3 kframes per second, which is at least 3.26× faster than the state-of-the-art digital hardware accelerators, while the system power is 94.5 mW and the energy consumption is only 40 μJ per frame.
Zichen Fan, Zheyu Liu, Zheng Qu 0002, Fei Qiao, Qi Wei 0001, Shuzheng Xu, Huazhong Yang
IEEE Trans. Very Large Scale Integr. Syst.1
2019 RED: A ReRAM-based Deconvolution Accelerator
abstract
Deconvolution has been widespread in neural networks. For example, it is essential for performing unsupervised learning in generative adversarial networks or constructing fully convolutional networks for semantic segmentation. Resistive RAM (ReRAM)-based processing-in-memory architecture has been widely explored in accelerating convolutional computation and demonstrates good performance. Performing deconvolution on existing ReRAM-based accelerator designs, however, suffers from long latency and high energy consumption because deconvolutional computation includes not only convolution but also extra add-on operations. To realize the more efficient execution for deconvolution, we analyze its computation requirement and propose a ReRAM-based accelerator design, namely, RED. More specific, RED integrates two orthogonal methods, the pixel-wise mapping scheme for reducing redundancy caused by zero-inserting operations and the zero-skipping data flow for increasing the computation parallelism and therefore improving performance. Experimental evaluations show that compared to the state-of- the-art ReRAM-based accelerator, RED can speed up operation 3.69~31.15× and reduce 8%~88.36% energy consumption.
Zichen Fan, Ziru Li, Bing Li 0017, Yiran Chen 0001, Hai Li 0001
DATE1