EDBT 2026 Demo / reviewers in the wild / expert
Yuan Du
dblp:26/8831
· DBLP profile ↗
60ranked-venue papers
3as first author
46since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 39 · 1 first-author · 31 since 2021Artificial intelligence and machine learning · 10 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Computer networks · 7 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Decomposing the Neurons: Activation Sparsity via Mixture of Experts for Continual Test Time AdaptationabstractContinual Test-Time Adaptation (CTTA), which aims to adapt the pre-trained model to ever-evolving target domains, emerges as an important task for vision models. As current vision models appear to be heavily biased towards texture, continuously adapting the model from one domain distribution to another can result in serious catastrophic forgetting. Drawing inspiration from the the encoding characteristics of neuron activation in neural networks, we propose the Mixture-of-Activation-Sparsity-Experts (MoASE) for the CTTA task. Given the distinct reaction of neurons with low and high activation to domain-specific and agnostic features, MoASE decomposes the neural activation into high-activation and low-activation components in each expert with a Spatial Differentiable Dropout (SDD). Based on the decomposition, we devise a Domain-Aware Router (DAR) that utilizes domain information to adaptively weight experts that process the post-SDD sparse activations, and the Activation Sparsity Gate (ASG) that adaptively assigns feature selection thresholds of the SDD for different experts for more precise feature decomposition. Finally, we introduce a Homeostatic-Proximal (HP) loss to maintain update consistency between the teacher and student experts to prevent error accumulation. Extensive experiments substantiate that MoASE achieves state-of-the-art performance in both classification and segmentation tasks. Rongyu Zhang, Aosong Cheng, Yulin Luo, Gaole Dai, Huanrui Yang, Jiaming Liu 0003, Ran Xu 0013, Dan Wang 0002, Yuan Du |
AAAI | 10 |
| 2026 | MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot ManipulationabstractVision-Language-Action (VLA) models enable robotic systems to perform embodied tasks but face deployment challenges due to the high computational demands of the dense Large Language Models (LLMs), with existing early-exit-based sparsification methods often overlooking the critical semantic role of final layers in downstream tasks. Aligning with the recent breakthrough of the Shallow Brain Hypothesis (SBH) in neuroscience and the mixture of experts in model sparsification, we conceptualize each LLM layer as an expert and propose a Mixture-of-LayEr Vision Language Action model (MoLe-VLA or simply MoLe) architecture for dynamic LLM layer activation. Specifically, we introduce a Spatial-Temporal Aware Router (STAR) for MoLe to selectively activate only parts of the layers based on the robot’s current state, mimicking the brain's distinct signal pathways specialized for cognition and causal reasoning. Additionally, to compensate for the cognition ability of LLM lost during the layer-skipping, we devise a Cognitive self-Knowledge Distillation (CogKD) to enhance the understanding of task demands and generate task-relevant action sequences by leveraging cognition features. Extensive experiments in RLBench simulations and real-world environments demonstrate the superiority of MoLe-VLA in both efficiency and performance, improving the mean success rate by 9.7% across ten simulation tasks while accelerating inference by 36.8% over OpenVLA. Rongyu Zhang, Menghang Dong, Yuan Zhang 0020, Liang Heng, Xiaowei Chi, Gaole Dai, Dan Wang 0002, Yuan Du, Shanghang Zhang |
AAAI | 9 |
| 2026 | A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP SystemsabstractMixture-of-Experts (MoE) models facilitate edge deployment by decoupling model capacity from active computation, yet their large memory footprint drives the need for GPU systems with near-data processing (NDP) capabilities that offload experts to dedicated processing units. However, deploying MoE models on such edge-based GPU-NDP systems faces three critical challenges: 1) severe load imbalance across NDP units due to non-uniform expert selection and expert parallelism, 2) insufficient GPU utilization during expert computation within NDP units, and 3) extensive data pre-profiling necessitated by unpredictable expert activation patterns for pre-fetching. To address these challenges, this paper proposes an efficient inference framework featuring three key optimizations. First, the underexplored tensor parallelism in MoE inference is exploited to partition and compute large expert parameters across multiple NDP units simultaneously towards edge low-batch scenarios. Second, a load-balancing-aware scheduling algorithm distributes expert computations across NDP units and GPU to maximize resource utilization. Third, a dataset-free pre-fetching strategy proactively loads frequently accessed experts to minimize activation delays. Experimental results show that our framework enables GPU-NDP systems to achieve 2.41× on average and up to 2.56× speedup in end-to-end latency compared to state-of-the-art approaches, significantly enhancing MoE inference efficiency in resource-constrained environments. Chao Fang 0005, Yichuan Bai, Yuan Du |
DATE | 7 |
| 2026 | UniMAC: Multi-Precision MAC with a Unified Computing Architecture for Efficient Large Language Model Inference
Xiao Cong, Xingyuan Hu, Xulong Zhang 0007, Yuan Du |
ISCAS | 7 |
| 2026 | A 30Gb/s/pin PAM3 Single-Ended Transceiver with Crosstalk-Cancelling Coding for Memory Interfaces
Yiheng Hui, Yuan Du |
ISCAS | 6 |
| 2026 | A Compilation Framework for Domain Specific Multi-Chiplet Systems with Double-Layer Genetic Algorithms
Yichuan Bai, Yuan Du |
ISCAS | 4 |
| 2026 | GIL-DDI: multi-view graph invariant learning for unknown drug-drug interaction prediction
Yuanxian Li, Yuan Du, Zhenli He, Xin Jin 0005, Cheng Xie 0001 |
Knowl. Inf. Syst. | 2 |
| 2026 | An integrated framework for enhancing small AprilTag pose accuracy under long-distance conditions
Hezhi Zhang, Yuan Du, Qin Zhang 0010, Guanwen Huang |
Pattern Recognit. | 2 |
| 2026 | A High-Speed FPGA Implementation for IVF-PQ Index ConstructionabstractThe Inverted File with Product Quantization (IVF-PQ) is a widely used method for Approximate Nearest Neighbor Search (ANNS), playing a critical role in AI-driven applications such as search engines, recommendation systems, and advertising platforms. With the advent of Large Language Models (LLMs), the demand for efficient and real-time index construction has significantly increased, especially for edge-side personal applications. In this paper, we propose a scalable and high-speed FPGA implementation of IVF-PQ index construction, significantly reducing indexing latency and making it feasible for edge scenarios. First, we optimize the original index construction algorithm by introducing batch-mode centroid updates and replacing floating-point division with hardware-efficient operations, while maintaining competitive recall performance (with less than 5% degradation and up to 12.5% improvement compared to the original algorithm). Next, based on the modified algorithm, we design a flexible and scalable hardware architecture that supports two distance metrics (L2 and Inner Product), six PQ configurations, and input data with up to 1024 dimensions, all without necessitating hardware recompilation. Our implementation maximizes computational efficiency through finely tuned parallelism and dataflow, ensuring full pipeline utilization. Finally, we implement our design in Verilog and evaluate it on the Xilinx XCU280-FSVH2892-2L-E FPGA platform. Experimental results show that our accelerator achieves up to$30\times $speedup over a high-end server CPU (Intel Xeon Gold 6248R), reducing the indexing time from hours to minutes. Yifeng Song, Yuan Du, Zhongfeng Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2026 | MG-PM: A Mixed-Granularity Polynomial Multiplier for High-Efficient Computation Across Variable DegreesabstractPolynomial multiplication across variable degrees has been a dominant bottleneck in cryptographic applications, including post-quantum cryptography (PQC) to fully homomorphic encryption (FHE). Yet, none of the existing approaches can achieve both high performance and scalability across variable degrees. Number theoretic transform (NTT) excels at large degrees but suffers from performance degradation when the polynomial is decomposed into small degrees, while the iterative convolution algorithm (ICA) scales poorly to large degrees. Reconciling two disparate approaches has been hindered by a mathematical barrier: NTT decomposition produces heterogeneous polynomial rings that defy uniform ICA processing. This work overcomes this barrier by establishing a new algorithm, based on isomorphic mapping, to transform heterogeneous rings into a fixed-size cyclic ring. Leveraging this algorithm, we propose a scalable architecture called MG-PM, a mixed-granularity polynomial multiplier that, for the first time, enables efficient acceleration across variable degrees. Implemented on a Xilinx VCU118 FPGA, MG-PM supports degrees from 256 to 65 536, achieving up to$\mathbf {9.1\times }$latency reduction and$\mathbf {2.2\times }$area reduction under different configurations, as well as 92.3% improvement in area-time product (ATP) compared with other state-of-the-art designs, establishing MG-PM as a highly efficient and scalable solution for accelerating polynomial multiplication across variable degrees. Yiping Shi, Liulu He, Yuan Du |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | PAT: Pruning-Aware Tuning for Large Language ModelsabstractLarge language models (LLMs) excel in language tasks, especially with supervised fine-tuning after pre-training. However, their substantial memory and computational requirements hinder practical applications. Structural pruning, which reduces less significant weight dimensions, is one solution. Yet, traditional post-hoc pruning often leads to significant performance loss, with limited recovery from further fine-tuning due to reduced capacity. Since the model fine-tuning refines the general and chaotic knowledge in pre-trained models, we aim to incorporate structural pruning with the fine-tuning, and propose the Pruning-Aware Tuning (PAT) paradigm to eliminate model redundancy while preserving the model performance to the maximum extend. Specifically, we insert the innovative Hybrid Sparsification Modules (HSMs) between the Attention and FFN components to accordingly sparsify the upstream and downstream linear modules. The HSM comprises a lightweight operator and a globally shared trainable mask. The lightweight operator maintains a training overhead comparable to that of LoRA, while the trainable mask unifies the channels to be sparsified, ensuring structural pruning. Additionally, we propose the Identity Loss which decouples the transformation and scaling properties of the HSMs to enhance training robustness. Extensive experiments demonstrate that PAT excels in both performance and efficiency. For example, our Llama2-7b model with a 25% pruning ratio achieves 1.33x speedup while outperforming the LoRA-finetuned model by up to 1.26% in accuracy with a similar training cost. Yijiang Liu, Huanrui Yang, Youxin Chen, Rongyu Zhang, Yuan Du |
AAAI | 6 |
| 2025 | DiffCkt: A Diffusion Model-Based Hybrid Neural Network Framework for Automatic Transistor-Level Generation of Analog CircuitsabstractAnalog circuit design consists of the pre-layout and layout phases. Among them, the pre-layout phase directly decides the final circuit performance, but heavily depends on experienced engineers to do manual design according to specific application scenarios. To overcome these challenges and automate the analog circuit pre-layout design phase, we introduce DiffCkt: a diffusion model-based hybrid neural network framework for the automatic transistor-level generation of analog circuits, which can directly generate corresponding circuit structures and device parameters tailored to specific performance requirements. To more accurately quantify the efficiency of circuits generated by DiffCkt, we introduce the Circuit Generation Efficiency Index (CGEI), which is determined by both the figure of merit (FOM) of a single generated circuit and the time consumed. Compared with relative research, DiffCkt has improved CGEI by a factor of 2.21 ~ 8365× , reaching a state-of-the-art (SOTA) level. In conclusion, this work shows that the diffusion model has the remarkable ability to learn and generate analog circuit structures and device parameters, providing a revolutionary method for automating the pre-layout design of analog circuits. The circuit dataset is now available at https://github.com/CjLiu-NJU/DiffCkt. Yabing Feng, Yuan Du |
ICCAD | 6 |
| 2025 | FBQuant: FeedBack Quantization for Large Language ModelsabstractDeploying Large Language Models (LLMs) on edge devices is increasingly important, as it eliminates reliance on network connections, reduces expensive API calls, and enhances user privacy. However, on-device deployment is challenging due to the limited computational resources of edge devices. In particular, the key bottleneck stems from memory bandwidth constraints related to weight loading. Weight-only quantization effectively reduces memory access, yet often induces significant accuracy degradation. Recent efforts to incorporate sub-branches have shown promise for mitigating quantization errors, but these methods either lack robust optimization strategies or rely on suboptimal objectives. To address these gaps, we propose FeedBack Quantization (FBQuant), a novel approach inspired by negative feedback mechanisms in automatic control. FBQuant inherently ensures that the reconstructed weights remain bounded by the quantization process, thereby reducing the risk of overfitting. To further offset the additional latency introduced by sub-branches, we develop an efficient CUDA kernel that decreases 60% of extra inference time. Comprehensive experiments demonstrate the efficiency and effectiveness of FBQuant across various LLMs. Notably, for 3-bit Llama2-7B, FBQuant improves zero-shot accuracy by 1.2%. Yijiang Liu, Hengyu Fang, Liulu He, Rongyu Zhang, Yichuan Bai, Yuan Du |
IJCAI | 6 |
| 2025 | A 65nm 100MS/s-400MS/s Multi-Column SAR/SS ADC with Reconfigurable 2-8 Bit Resolution for Optical and In-memory ComputingabstractComputing-in-Memory (CIM) architecture effectively reduces power consumption caused by data movement between analog computing elements and memory. Optoelectronic computing offers a promising alternative to traditional analog computing by bypassing the energy-bandwidth tradeoff and reducing latency. In analog computing architectures, analog-to-digital converters (ADCs) are essential for quantizing column-parallel computing results. However, ADCs with fixed resolution can result in unnecessary power consumption during low-resolution requirements. In this paper, we propose a hybrid ADC with a configurable resolution from 2 to 8-bit. The ADC integrates successive approximation register (SAR) and single-slope (SS) circuits to minimize the area and power consumption. The proposed ADC architecture was fabricated in 65nm CMOS process with a 1 V power supply. The prototype occupies an area of 177μm × 11μm and consumes 238μW in the 8-bit conversion mode. Measured results show that the ADC achieves an effective number of bits (ENOB) of 7.35 and a Walden figure of merit (FoMw)of 14.6 fJ/conv. Anying Jiang, Wenhe Yin, Yize Wang, Mingqian Yang, Yuan Du |
ISCAS | 8 |
| 2025 | Low-Latency DAE: A Configurable Lightweight Hybrid Data and Address Encryption Engine for IoT Real-Time NVM ProtectionabstractIn Internet of Things (IoT) real-time systems and edge computing applications, memory encryption engines (MEEs) are used for real-time memory encryption to protect program code and sensitive data in nonvolatile memories (NVMs) and mitigate some side-channel attacks. However, for resource-constrained devices, it is a considerable challenge to realize a lightweight solution with low-hardware overhead, high security, and flexible bit-width. This article presents a lightweight full-MEE, called data&address encryption (DAE), which employs hybrid data and address encryption to protect NVMs on-the-fly with low-logic latency, flexible width adaptation, and enhanced security in some aspects. The security analyses show that DAE performs effective mitigation in some side-channel attacks, such as Remanence attack, and provides better security than data-only ciphers in resisting the brute-force attack, etc. Evaluated with TSMC’s 40-nm standard CMOS technology, 128-bit DAE has a lightweight feature of 8.703 KGates, which is only 5.46% of 128-bit advanced encryption standard (AES-128), and performs a$5.8\times $throughput, a$105.9\times $area efficiency, and a$48.1\times $energy efficiency. In the experiments on SoC simulation and field programmable gate array platform with an embedded RISC-V core, the results show that DAE causes little or no loss of system frequency and throughput. Xuewen He, Yuan Du |
IEEE Internet Things J. | 3 |
| 2025 | BE-NPU: A Bandwidth-Efficient Neural Processing Unit With Adaptive Processing Schemes for Reduced Off-Chip Bandwidth DemandabstractExisting neural processing units (NPUs) mainly focus on the optimized multiply-accumulate (MAC) arrays for efficient inference of convolutional neural networks (CNNs). However, off-chip data transmission usually keeps NPUs waiting during CNN inference, causing up to 38.4GB/s off-chip bandwidth (OCB) demand for mobile AI devices. And none of the previous benchmarks quantitatively evaluate the bandwidth efficiency of different NPU architectures. In addition, CNNs exhibit distinct characteristics of off-chip data transmission when applied to different fields, and it has become a challenging task for NPUs to support different CNNs efficiently with reasonable OCB demand. To address the aforementioned issues, this paper proposes the Bandwidth-Peak Performance Ratio for n percentages of ideal frame rate (BPPR-n%) to demonstrate the normalized OCB demand of different NPU architectures. A bandwidth-efficient NPU (BE-NPU) is introduced with adaptive processing schemes to reduce the OCB demand during inference of different CNNs. The adaptive processing schemes include both instruction-level and thread-level schemes. For the instruction-level scheme, decoupled execute/access is introduced into depth-first (DF) and layer-first (LF) schemes to improve the concurrency between NPU calculation (CAL) and direct memory access (DMA) instructions. For the thread-level scheme, DF and LF threads are hybridly processed to further improve overall NPU efficiency. Compared with state-of-the-art works, BE-NPU achieves 48.1%~80.6% reduction of BPPR-80% and 67.0%~95.1% reduction of BPPR-95%. The proposed architecture is synthesized with TSMC 28nm technology node. BE-NPU utilizes 14.3% additional logic gates compared with baseline implementation. Yichuan Bai, Yaqing Li, Yuan Du |
IEEE Trans. Computers | 5 |
| 2025 | ISCAS Guest Editorial Special Issue Based on the 2025 IEEE International Symposium on Circuits and Systems
Jiafeng Xie, Yuan Du, Xinmiao Zhang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2025 | BEVUDA++: Geometric-Aware Unsupervised Domain Adaptation for Multi-View 3D Object DetectionabstractVision-centric Bird’s Eye View (BEV) perception holds considerable promise for autonomous driving. Recent studies have prioritized efficiency or accuracy enhancements, yet the issue of domain shift has been overlooked, leading to substantial performance degradation upon transfer. We identify major domain gaps in real-world cross-domain scenarios and initiate the first effort to address the Domain Adaptation (DA) challenge in multi-view 3D object detection for BEV perception. Given the complexity of BEV perception approaches with their multiple components, domain shift accumulation across multi-geometric spaces (e.g., 2D, 3D Voxel, BEV) poses a significant challenge for BEV domain adaptation. In this paper, we introduce an innovative geometric-aware teacher-student framework, BEVUDA++, to diminish this issue, comprising a Reliable Depth Teacher (RDT) and a Geometric Consistent Student (GCS) model. Specifically, RDT effectively blends target LiDAR with dependable depth predictions to generate depth-aware information based on uncertainty estimation, enhancing the extraction of Voxel and BEV features that are essential for understanding the target domain. To collaboratively reduce the domain shift, GCS maps features from multiple spaces into a unified geometric embedding space, thereby narrowing the gap in data distribution between the two domains. Additionally, we introduce a novel Uncertainty-guided Exponential Moving Average (UEMA) to further reduce error accumulation due to domain shifts informed by previously obtained uncertainty guidance. To demonstrate the superiority of our proposed method, we execute comprehensive experiments in four cross-domain scenarios, securing state-of-the-art performance in BEV 3D object detection tasks, e.g., 12.9% NDS and 9.5% mAP enhancement on Day-Night adaptation. Rongyu Zhang, Jiaming Liu 0003, Xiaoqi Li 0009, Xiaowei Chi, Dan Wang 0002, Yuan Du, Shanghang Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | RepCaM++: Exploring Transparent Visual Prompt With Inference-Time Re-Parameterization for Neural Video DeliveryabstractRecently, content-aware methods have been employed to reduce bandwidth and enhance the quality of Internet video delivery. These methods involve training distinct content-aware super-resolution (SR) models for each video chunk on the server, subsequently streaming the low-resolution (LR) video chunks with the SR models to the client. Prior research has incorporated additional partial parameters to customize the models for individual video chunks. However, this leads to parameter accumulation and can fail to adapt appropriately as video lengths increase, resulting in increased delivery costs and reduced performance. In this paper, we introduce RepCaM++, an innovative framework based on a novel Re- parameterization Content-aware Modulation (RepCaM) module that uniformly modulates video chunks. The RepCaM framework integrates extra parallel-cascade parameters during training to accommodate multiple chunks, subsequently eliminating these additional parameters through re- parameterization during inference. Furthermore, to enhance RepCaM's performance, we propose the Transparent Visual Prompt (TVP), which includes a minimal set of zero-initialized image-level parameters (e.g., less than 0.1%) to capture fine details within video chunks. We conduct extensive experiments on the VSD4K dataset, encompassing six different video scenes, and achieve state-of-the-art results in video restoration quality and delivery bandwidth compression. Rongyu Zhang, Xize Duan, Jiaming Liu 0003, Yuan Du, Dan Wang 0002, Shanghang Zhang, Fangxin Wang 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | A Hybrid-Structured Lossless Compression-Decompression Engine for Intermediate Feature Maps in Vision Neural NetworksabstractWith the continuous evolution of vision neural networks, the off-chip transmission and storage of intermediate feature maps have become a major bottleneck during inference, especially in resource-constrained edge devices. Lossless compression of the intermediate feature maps exhibits the plug-and-play feature without requiring additional evaluation or retraining. However, previous lossless compression hardware engines face challenges in the tradeoff between hardware complexity and compression ratio. To address this issue, this article proposes a hybrid-structured lossless compression and decompression engine of intermediate feature maps in vision neural networks. The proposed work combines multidimensional prediction, run-length encoding (RLE), and extended encoding (EE). In the predictive stage, delta and contextual prediction are employed to enhance sparsity for compression efficiency. In the encoding stage, RLE minimizes redundancy between adjacent bytes, and EE is used to select various compression methods to decrease bit-level redundancy dynamically. The average compression ratio of widely used vision neural networks is 48.42%, which is better than that of Huffman coding. Compared with the state-of-the-art works, it improves the average compression ratio by 20.90%. For hardware implementation, hardware reuse and data reuse are employed to achieve 17.95% and 21.18% reduction of the gate count and the power, respectively. Experimental results show that we achieve a throughput-per-area of 1.10 [(bits/cycle)/K GCs] and a throughput-per-power of 1.19 [(bits/cycle)/mW] in the 28-nm process node. Junyong Hua, Yichuan Bai, Yuan Du |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | A Fast-Convergence Near-Memory-Computing Accelerator for Solving Partial Differential EquationsabstractSolving partial differential equations (PDEs) is omnipresent in scientific research and engineering and requires expensive numerical iteration for memory and computation. The primary concerns for solving PDEs are convergence speed, data movement, and power consumption. This work proposed the first fast-convergence PDE solver with an automatic adjustment multiple-stride iteration method, significantly increasing the PDE convergence speed. A dynamic-precision near-memory-computing architecture with booth encoding is proposed to reduce iterated intermediate data movement. A customized 32T compressor and a 14T full adder are designed to reduce the power and hardware cost of the solver. The processor is fabricated using 65-nm CMOS technology and occupies a 6.25 mm2 die area. It can achieve a convergence speedup by$4\times $compared with the existing work. Chenjia Xie, Xingyuan Hu, Yuan Du |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | Corrections to "An Efficient Two-Stage Pipelined Compute-in-Memory Macro for Accelerating Transformer Feed-Forward Networks"abstractIn the above article [1], the die photograph on the right side of original Fig. 9 was inadvertently mirrored horizontally, as shown in Fig. 1. This occurred during the annotation process, where the image used had already been flipped without our awareness. As a result, the internal layout labeling (e.g., CIMA1, CIMA2, and ADC) appeared in reverse orientation relative to the actual die.Fig. 1.Difference clarification between the original Fig. 9 of our published article and the revised Fig. 9. Fig. 9.Die photograph and measure setup for the proposed chip. Heng Zhang 0024, Wenhe Yin, Sunan He, Yuan Du |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2024 | Efficient Deweahter Mixture-of-Experts with Uncertainty-Aware Feature-Wise Linear ModulationabstractThe Mixture-of-Experts (MoE) approach has demonstrated outstanding scalability in multi-task learning including low-level upstream tasks such as concurrent removal of multiple adverse weather effects. However, the conventional MoE architecture with parallel Feed Forward Network (FFN) experts leads to significant parameter and computational overheads that hinder its efficient deployment. In addition, the naive MoE linear router is suboptimal in assigning task-specific features to multiple experts which limits its further scalability. In this work, we propose an efficient MoE architecture with weight sharing across the experts. Inspired by the idea of linear feature modulation (FM), our architecture implicitly instantiates multiple experts via learnable activation modulations on a single shared expert block. The proposed Feature Modulated Expert (FME) serves as a building block for the novel Mixture-of-Feature-Modulation-Experts (MoFME) architecture, which can scale up the number of experts with low overhead. We further propose an Uncertainty-aware Router (UaR) to assign task-specific features to different FM modules with well-calibrated weights. This enables MoFME to effectively learn diverse expert functions for multiple tasks. The conducted experiments on the multi-deweather task show that our MoFME outperforms the state-of-the-art in the image restoration quality by 0.1-0.2 dB while saving more than 74% of parameters and 20% inference time over the conventional MoE counterpart. Experiments on the downstream segmentation and classification tasks further demonstrate the generalizability of MoFME to real open-world applications. Rongyu Zhang, Yulin Luo, Jiaming Liu 0003, Huanrui Yang, Zhen Dong 0003, Denis A. Gudovskiy, Tomoyuki Okuno, Yohei Nakata, Kurt Keutzer, Yuan Du, Shanghang Zhang |
AAAI | 10 |
| 2024 | SFC: Achieve Accurate Fast Convolution under Low-precision ArithmeticabstractFast convolution algorithms, including Winograd and FFT, can efficiently accelerate convolution operations in deep models. However, these algorithms depend on high-precision arithmetic to maintain inference accuracy, which conflicts with the model quantization. To resolve this conflict and further improve the efficiency of quantized convolution, we proposes SFC, a new algebra transform for fast convolution by extending the Discrete Fourier Transform (DFT) with symbolic computing, in which only additions are required to perform the transformation at specific transform points, avoiding the calculation of irrational number and reducing the requirement for precision. Additionally, we enhance convolution efficiency by introducing correction terms to convert invalid circular convolution outputs of the Fourier method into effective ones. The numerical error analysis is presented for the first time in this type of work and proves that our algorithms can provide a 3.68× multiplication reduction for 3×3 convolution, while the Winograd algorithm only achieves a 2.25× reduction with similarly low numerical errors. Experiments carried out on benchmarks and FPGA show that our new algorithms can further improve the computation efficiency of quantized models while maintaining accuracy, surpassing both the quantization-alone method and existing works on fast convolution quantization. Liulu He, Yuan Du |
ICML | 4 |
| 2024 | Optoelectronic Computing Evaluation and Deployment Platform Based on a 256-MAC Silicon Photonic ChipabstractThe deceleration of Moore's Law has led to increasing difficulties in advancing the computational speed and power efficiency of Complementary-Metal-Oxide-Semiconductor (CMOS) chips. As a solution to this challenge, optical computing emerges as a promising technology, boasting low energy consumption, high processing speed, and extensive bandwidth. Yet, a critical obstacle remains: the absence of a co-simulation platform that incorporates both photonic chips and peripheral electrical circuits. This paper addresses this gap by introducing a hybrid optoelectronic computing evaluation and deployment platform utilizing Simulink tools. Based on the measured data from the silicon optical computing chip, we have deployed an image filtering algorithm and a convolutional neural network onto this platform. The optical computing chip achieves an accuracy of 86.4% on the ImageNet image dataset. Through evaluation, we have identified the most substantial impacts on calculation results. To achieve an image classification accuracy of 80%, the signal-to-noise ratio (SNR) of the low-speed DAC must be a minimum of 52 dB. These findings provide crucial insights into the optimization of optical computing systems. Likai Li, Yichuan Bai, Shengping Liu, Sunan He, Yaqing Li, Yuan Du |
ISCAS | 8 |
| 2024 | VeCAF: Vision-language Collaborative Active Finetuning with Training Objective AwarenessabstractFinetuning a pretrained vision model (PVM) is a common technique for learning downstream vision tasks. The conventional finetuning process with the randomly sampled data points results in diminished training efficiency. To address this drawback, we propose a novel approach, Vision- languag e C ollaborative A ctive F inetuning (VeCAF). VeCAF optimizes a parametric data selection model by incorporating the training objective of the model being tuned. Effectively, this guides the PVM towards the performance goal with improved data and computational efficiency.With the ever-growing feasibility of acquiring labels and natural language annotations of image data through web-scale crawling, we exploit the inherent semantic richness of the text embedding space and utilize text embeddings of image annotations to augment PVM image features for better data selection and finetuning. Furthermore, the flexibility of text-domain augmentation gives VeCAF the unique ability to handle out-of-distribution scenarios without external augmented data. Extensive experiments show the leading performance and high efficiency of VeCAF that is superior to baselines in both in-distribution and out-of-distribution image classification tasks. On ImageNet, VeCAF needs up to 3.3× less training batches to reach the target performance compared to full fine-tuning and achieves an accuracy improvement of 2.8% over active SOTA fine-tuning methods with the same number of batches. Our code is now available at https://github.com/RoyZry98/VeCAF-Pytorch. Rongyu Zhang, Zefan Cai, Huanrui Yang, Denis A. Gudovskiy, Tomoyuki Okuno, Yohei Nakata, Kurt Keutzer, Baobao Chang, Yuan Du, Shanghang Zhang |
ACM Multimedia | 10 |
| 2024 | Low-Latency PAE: Permutation-Based Address Encryption Hardware Engine for IoT Real-Time Memory ProtectionabstractIn Internet of Things (IoT) endpoint devices, some data or address ciphers are used for real-time memory protection to mitigate some side-channel attacks against memories. To better meet the requirements of real-time memory protection, this article proposes a hardware engine of permutation-based address encryption (PAE) to implement memory address encryption with flexible width adaptation, low latency, and low hardware overhead. When evaluated with TSMC’s 40-nm standard CMOS technology, PAE features lightweight characteristics with a gate count of 0.589 KGates, which is only 0.37% of advanced encryption standard (AES) and 33.50% of address cipher Galois field encryption (GF-Enc). The security of PAE in memory protection is quantitatively proven through both logic cryptanalysis and side-channel attacks. The results show that PAE performs effective mitigation in some side-channel attacks and provides better security than other address ciphers in resisting the brute-force attack, chosen-plaintext attack, and the differential attack. A RISC-V system with PAE and AES is deployed on an field-programmable gate array platform to analyze the impact on performance. The evaluation data show that PAE has no impact on system throughput in the case analysis, while AES reduces system throughput by 89.47%. Xuewen He, Yichuan Bai, Zhongfeng Wang 0001, Yuan Du |
IEEE Internet Things J. | 6 |
| 2024 | A Compilation Framework for SRAM Computing-in-Memory Systems With Optimized Weight Mapping and Error CorrectionabstractDeploying convolution-based algorithms into SRAM computing-in-memory (CIM) systems faces various challenges, such as operator incompatibility and intrinsic non-ideal error. This paper proposes a compilation framework to address this issue. Efficient weight mapping strategies are introduced to improve the utilization of SRAM-CIM macro. The intrinsic non-ideal errors of SRAM-CIM macro are also taken into consideration, and two efficient error correction schemes are proposed, which include calibration of computation voltage linear error (CCVLE) and the mitigation of analog-to-digital quantization error (MAQE). In addition, bit-width flexibility and signed-unsigned reconfigurability are also supported to facilitate the deployment of various convolution-based algorithms. ResNet18, finite impulse response (FIR) filtering, and Gaussian image filtering are deployed into a multi-macro SRAM-CIM system. These algorithms serve as deployment representatives of convolutional neural network (CNN), digital signal processing (DSP), and digital image processing (DIP), respectively. The results show that the introduced weight mapping strategies improve the macro utilization by 63.29% and 21.10% for two types of frequently used convolution layers compared to the traditional strategy. Moreover, the proposed error correction schemes achieve similar algorithm accuracy to the floating-point results, and the deployment result of ResNet18 achieves 66.3%~70.1% top-1 classification accuracy evaluated on the ImageNet dataset with different throughput tradeoffs. Yichuan Bai, Yaqing Li, Heng Zhang 0024, Aojie Jiang, Yuan Du |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2024 | GroupQ: Group-Wise Quantization With Multi-Objective Optimization for CNN AcceleratorsabstractMixed-precision Neural Networks achieve high energy efficiency and throughput for hardware deployment. The most common mixed-precision methods adopt layer-wise granularity. However, the layer-wise method does not quantize the model to its limit because the optimal bit precision that preserves accuracy for different kernels can be different. To address this issue, this paper presents GroupQ, a group-wise quantization method with multi-objective optimization for CNN accelerators. Group-wise divides the convolutional kernels in a layer into several groups by clustering, and each group shares the same bit precision. The multi-objective optimization algorithm is used to optimize the quantization policy automatically based on the selected quantization objectives, such as model accuracy, model size, or computation cost. The experiments show that GroupQ significantly outperforms the existing layer-wise retraining-free methods, even better than some training-based methods. Specifically, GroupQ achieves a 0.49% higher accuracy with up to 28.1% smaller Bit Operations (BOPs) on ResNet-18 compared to HAWQ-V3 and can quantize MobileNetV2 to 1.65MB model size with 71.75% top-1 accuracy. This paper shows that GroupQ is friendly for hardware deployment by a lookup table (LUT)-based mixed-precision processing element (LMPE) proposed for CNN accelerators. LMPE provides power reduction of up to 3.6%, up to 3.9% lower area, compared to conventional implementation. Aojie Jiang, Yuan Du |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | An 11T1C Bit-Level-Sparsity-Aware Computing- in-Memory Macro With Adaptive Conversion Time and Computation VoltageabstractA static random-access memory (SRAM)-based computing-in-memory (CiM) is a promising architecture for efficiently performing high-precision integer (INT) multiplication and accumulation (MAC) operations. In this work, we propose a charge-domain bit-level-sparsity-aware analog CiM (ACiM) macro for an area-energy-efficient convolutional neural network (CNN). An 11T1C ACiM bit-cell is proposed to dynamically remove the computation capacitors during the accumulation phase based on the weight value (W) for improving the partial sums and analog computing accuracy margin (ACAM). The computation voltage is dynamically adjusted according to column-wise sparsity by the bit-level-sparsity-aware controller to improve energy efficiency. To digitize the MAC computing results, a 2-8bit column-parallel time-interleaved hybrid analog-to-digital converter (ADC) is designed by sharing the voltage reference generator, which achieves a low unit pitch size. A$256\times 64~11$T1C ACiM macro prototype with hybrid ADCs is implemented using 55nm CMOS process. The silicon measurement results show that the proposed ACiM achieves a throughput of 51.2-153.6 GOPS, core area efficiency reaching 112-336GOPS/mm2, and energy efficiency ranging from 17 to 111 TOPS/W with 8bit weights and 8bit inputs. Yuandong Li, Heng Zhang 0024, Jingjing Lv, Anying Jiang, Yuan Du |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2024 | An Efficient GCN Accelerator Based on Workload Reorganization and Feature ReductionabstractThe irregular adjacency matrix and the mismatched computation patterns of Aggregation and Combination phases make Graph Neural Networks (GNNs) challenging to compute efficiently. This paper proposes a software and hardware co-design system to reduce computational latency and memory access based on workload reorganization and feature reduction. In software, the adjacency matrix is preprocessed, and the workload in both feature and node dimensions is concentrated to optimize memory access and hardware utilization. The interlayer nodes are analyzed using Principal Component Analysis (PCA) to explore the minimum feature vector length based on information redundancy, and a unique weight initialization is utilized for retraining to trim the feature vector to the minimum length. In hardware, an efficient GCN accelerator is designed to fully support the reorganized workload by reconfigurable output node computation. The hardware accelerator is implemented using 28-nm CMOS technology. It achieves 3.3 TOPS peak throughput and 2.6 TOPS/W energy efficiency. Compared with HyGCN, this result shows that the proposed method can improve the overall performance by$5\times $with a negligible accuracy loss of less than 0.5%. Chenjia Xie, Zihan Ning, Liang Chang 0002, Yuan Du |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2024 | An Energy-Efficient Spiking Neural Network Accelerator Based on Spatio-Temporal Redundancy ReductionabstractThe neurons of spiking neural networks (SNNs) carry sparsity to the activation from temporal and spatial. To achieve high energy efficiency, this work proposes layer-wise configurable timesteps (LCTs) and channel-regrouped sparse convolution (CRSC) to explore and exploit the redundancy of temporal and spatial dimensions. In the temporal dimension, principal component analysis (PCA) is utilized to analyze redundant information in each layer, and LCT is adopted to balance and eliminate variable redundancies between layers. In the spatial dimension, channels are regrouped by spike activation frequencies to ensure workload balance, and sparse convolution is implemented to further accelerate SNN’s computation flow. A layer-fuse method that embeds the parameter of batch normalization in leaky integrate-and-fire (BLIF) is also proposed to reduce data movement and processing time. An energy-efficient SNN accelerator integrating the above methods is designed. Compared with the baseline, the hardware equipped with the proposed LCT can achieve$3.22\times $acceleration by redundancy elimination. The CRSC and BLIF can also reduce the hardware computation and memory access by$3.24\times $during inference. Implemented in TSMC 28 nm technology, this accelerator can achieve an energy efficiency of 36.89 TOPS/W at 650 MHz. Chenjia Xie, Yuan Du |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2024 | An Efficient Two-Stage Pipelined Compute-in-Memory Macro for Accelerating Transformer Feed-Forward NetworksabstractTransformer architectures have achieved state-of-the-art performance in various applications. However, deploying transformer models on resource-constrained platforms is still challenging due to its dynamic workloads, intensive computations, and substantial memory access. In this article, we propose a two-stage pipelined compute-in-memory (CIM) macro for effectively deploying and accelerating the feed-forward network (FFN) layers of transformer models. Two independent CIM arrays are designed to execute the two distinct linear projections in FFN layers, which are interconnected by co-designed analog rectified linear unit (ReLU) circuits to realize the nonlinear activation function. The analog multiply-and-add (MAC) results from the first CIM array are streamed directly to the analog ReLU circuits, and subsequently to the next CIM array for performing another linear projection. This architecture eliminates the need for analog-to-digital converters (ADCs) and digital-to-analog converters (DACs) for internal results’ staging, thereby enhancing overall macro efficiency and reducing computing latency. A proof-of-concept macro is fabricated using TSMC 65-nm process and achieves 4.096 TOPS peak throughput, 4.39 TOPS/mm2 area efficiency, and 49.83 TOPS/W energy efficiency. To map transformer models onto the proposed macro, we quantize the FFN layers of BERTMINI model under per-token granularity for activations and per-tensor granularity for weights using quantization-aware training (QAT), which exhibits excellent accuracy across multiple benchmarks. Heng Zhang 0024, Wenhe Yin, Sunan He, Yuan Du |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2023 | QD-BEV : Quantization-aware View-guided Distillation for Multi-view 3D Object DetectionabstractMulti-view 3D detection based on BEV (bird-eye-view) has recently achieved significant improvements. However, the huge memory consumption of state-of-the-art models makes it hard to deploy them on vehicles, and the nontrivial latency will affect the real-time perception of streaming applications. Despite the wide application of quantization to lighten models, we show in our paper that directly applying quantization in BEV tasks will 1) make the training unstable, and 2) lead to intolerable performance degradation. To solve these issues, our method QD-BEV enables a novel view-guided distillation (VGD) objective, which can stabilize the quantization-aware training (QAT) while enhancing the model performance by leveraging both image features and BEV features. Our experiments show that QD-BEV achieves similar or even better accuracy than previous methods with significant efficiency gains. On the nuScenes datasets, the 4-bit weight and 6-bit activation quantized QD-BEV-Tiny model achieves 37.2% NDS with only 15.8 MB model size, outperforming BevFormer-Tiny by 1.8% with an 8× model compression. On the Small and Base variants, QD-BEV models also perform superbly and achieve 47.9% NDS (28.2 MB) and 50.9% NDS (32.9 MB), respectively. Zhen Dong 0003, Huanrui Yang, Ming Lu 0002, Cheng-Ching Tseng, Yuan Du, Kurt Keutzer, Shanghang Zhang |
ICCV | 6 |
| 2023 | Siamese Network Representation for Active LearningabstractActive learning is a crucial part of machine learning aiming to reduce the amount of labeled data by selecting the most informative data to be annotated. Most of the previous proposed active learning methods are based on aleatoric or epistemic uncertainties obtained by learning models while ignoring relationships within the data itself. We propose an efficient similarity-based active learning method using siamese convolutional neural networks. Pairs of image data are sent into the siamese network and similarity between them is computed on their output features. We evaluate our method on image classification, and validate the method on CIFAR10/100 and Caltech101 dataset. Our method outperforms at most 3.17% accuracy than Bayesian-based method and 6.31% than random sample. In addition, we propose a hierarchical clustering method for pool-based sampling strategies, which will boost the representation stage of our method. We also conduct an ablation study to fully explore the efficiency of our method. Yuan Du |
ICIP | 2 |
| 2023 | Characterization of Charge-Trap-Transistor (CTT) Threshold Voltage Degradation and Differential-Pair-Based Memory DesignabstractThis paper characterizes the threshold voltage$\boldsymbol{(V_{th})}$degradation of programmed charge trap transistors (CTTs). Logarithmic mathematical modeling of$\boldsymbol{V_{th}}$degradation versus time is proposed and fits well with the experiment result. The measurement is conducted on CTTs in TSMC 28-nm technology node. A CTT differential-pair-based memory architecture is proposed to cancel$\boldsymbol{V_{th}}$degradation. Charge trapping and de-trapping are utilized to modify CTTs'$\boldsymbol{V_{th}{}^{\prime}\mathrm{s}}$, then analog values are written as time-invariant differential$\boldsymbol{V_{th}}$'s into CTT -based pairs. The memory implementation proposed in this paper uses CMOS-only technologies without adding additional process. The experiment shows that the data retention trend varies less than$\boldsymbol{10\mu\mathrm{V}}$per hour, making it suitable for long-term analog values storage. Yuan Du |
ISCAS | 4 |
| 2023 | An Efficient CNN Inference Accelerator Based on Intra- and Inter-Channel Feature Map CompressionabstractDeep convolutional neural networks (CNNs) generate intensive inter-layer data during inference, which results in substantial on- chip memory size and off-chip bandwidth. To solve the memory constraint, this paper proposes an accelerator adopting a compression technique that can reduce the inter-layer data by removing both intra- and inter-channel redundant information. Principal component analysis (PCA) is utilized in the compression process to concentrate inter-channel information. The spatial differences, truncation, and reconfigurable bit-width coding are implemented inside every feature map to eliminate the intra-channel data redundancy. Moreover, a particular data arrangement is introduced to enhance data continuity to optimize PCA analysis and improve compression performance. A CNN accelerator with the proposed compression technique is designed to support the on- the-fly compression process by pipelining the reconstruction, CNN computation, and compression operation. The prototype accelerator is implemented using 28-nm CMOS technology. It achieves 819.2GOPS peak throughput and 3.75TOPS/W energy efficiency with 218.5mW. Experiments show that the proposed compression technique achieves compression ratios of 21.5%$\sim $43.0% (8-bit mode) and 9.8%$\sim $19.3% (16-bit mode) on state-of-the-art CNNs with a negligible accuracy loss. Chenjia Xie, Yuan Du |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2023 | SSM-CIM: An Efficient CIM Macro Featuring Single-Step Multi-bit MAC Computation for CNN Edge InferenceabstractCompute-in-memory (CIM) is a promising approach to solving the memory-wall problem existing in traditional computing architectures. In this paper, we introduce SSM-CIM, a charge-domain, static random-access memory (SRAM)-based CIM macro designed for area-energy-efficient convolutional neural network (CNN) inference. SSM-CIM utilizes an original sign-magnitude data encoding method for both inputs and weights. By codesigning four adjacent SRAM computing cells and employing a 3-bit digital-to-analog converter (DAC), SSM-CIM performs accurate 4-bit multiply-and-accumulate (MAC) computation in a single step, eliminating the peripheral digital shift-and-add circuits. To digitize the MAC computing results, a dedicated multi-reference assisted SAR ADC is designed by reusing the reference voltages from the DAC, which offers significant power and area savings. In addition, analog computing errors and quantization errors are analyzed to ensure the multi-bit computing accuracy of SSM-CIM. SSM-CIM is implemented and evaluated using 28-nm global foundry process. The post-layout simulation results validate the excellent computing linearity and accuracy of SSM-CIM. Benefitting from the compact layout design and fully parallel computing flow, the$144\times 256$macro achieves a peak throughput of 2.3 TOPS, an area efficiency of 10.2 TOPS/mm2, and an energy efficiency of 205.4 TOPS/W with 4-bit weights and 4-bit inputs. Heng Zhang 0024, Sunan He, Xinjie Guo, Shaodi Wang, Yuan Du |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2023 | QPA: A Quantization-Aware Piecewise Polynomial Approximation Methodology for Hardware-Efficient ImplementationsabstractPiecewise polynomial approximation (PPA) on nonlinear functions plays an important role in high-precision computing. In this article, we proposed QPA, an integration of error-flattened quantization-aware PPA methods, to generate the optimized coefficients for efficient hardware implementations targeting any polynomial order. QPA incorporated four key features to minimize the fitting error and the hardware cost, including using the Remez algorithm to compute the minimax fitting polynomial, combining the fitting and quantization operations to get an error-flattened characteristic, assigning specific coefficient bit width to each multiplier to reduce the hardware cost, and fine-tuning the truncated coefficients to further reduce the fitting error. Experimental results showed that our methods consistently achieved the lowest fitting error compared with the state-of-the-art error-flattened piecewise approximation methods. We synthesized the proposed designs with 28-nm TSMC CMOS technology. The results showed that the proposed designs achieved up to 37.0% area reduction and 50.5% power consumption reduction compared to the state-of-the-art error-flattened piecewise linear (PWL) method, and up to 27.0% area reduction, 21.4% delay reduction, and 20.8% power consumption reduction compared to the state-of-the-art error-flattened piecewise quadratic (PWQ) method. Yuan Du |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2022 | Prototype-Voxel Contrastive Learning for LiDAR Point Cloud Panoptic SegmentationabstractLiDAR point cloud panoptic segmentation, including both semantic and instance segmentation, plays a critical role in meticulous scene understanding for autonomous driving. Existing 3D voxelized approaches either utilize 3D sparse convolution that only focuses on local scene understanding, or add extra and time-consuming PointNet branch to capture global feature structures. To address these limitations, we propose an end-to-end Prototype-Voxel Contrastive Learning (PVCL) framework for learning stable and discriminative semantic representations, which includes voxel-level and prototype-level contrastive learning (CL). The voxel-level CL decreases intra-class distance and increases inter-class distance among sample representations, while the prototype-level CL further reduces the dependence of CL on negative sampling and avoids the influence of outliers from the same class, enabling PVCL to be more effective for outdoor point cloud panoptic segmentation. Extensive experiments are conducted on the public point cloud panoptic segmentation datasets, Semantic-KITTI and nuScenes, where evaluations and ablation studies demonstrate PVCL achieves superior performance compared with the state-of-the-art. Our approach ranks the top on the public leaderboard of Semantic-KITTI at the time of submission, and surpasses the published 2nd rank, EfficientLPS, by 1.7% in PQ. Minzhe Liu, Hengshuang Zhao, Jianing Li 0001, Yuan Du, Kurt Keutzer, Shanghang Zhang |
ICRA | 5 |
| 2022 | A 3-8bit Reconfigurable Hybrid ADC Architecture with Successive-approximation and Single-slope Stages for Computing in MemoryabstractComputing in Memory (CIM) is reported as one of the most promising non-Von-Neumann computing architectures to replace the existing digital AI processor architecture. Compared with digital-based computation, CIM shows advantages in computing density, energy efficiency, and throughput. However, it requires a large analog-to-digital array to quantize the column-parallel analog Multiply-Accumulate (MAC) results with tight area, high speed, and low power requirements. In this paper, we first discuss the design trade-off between different ADC architectures for CIM accelerators. To improve the overall system efficiency, a 3-8bit reconfigurable hybrid ADC architecture with successive-approximation and single-slope stages is proposed, particularly emphasizing the reconfigurability of the conversion speed and bit-resolution for different computation mode. A prototype is designed and simulated in 65-nm CMOS, which occupies an area of l90$\mu$m $\times$ 5$\mu$m and consumes a power of 48$\mu$W at 8-bit conversion mode, achieving 7.87-bit ENOB and 10.2 fJ/conv. Wuyu Fan, Yuandong Li, Likai Li, Yuan Du |
ISCAS | 5 |
| 2022 | Deep Neural Network Interlayer Feature Map Compression Based on Least-Squares FittingabstractDeep convolutional neural networks (CNNs) have brought a significant amount of interlayer data during computation, resulting in a large data-exchange delay and power consumption. This paper proposes a Least-Squares Fitting Compression (LSFC) method to compress the interlayer data to resolve the above problem. In LSFC, the feature maps are firstly divided into block groups; then, two base blocks are selected for each block group. Finally, the LSFC core is applied to get the fitting parameters, and the fitting parameters are selectively stored in the on-chip memory according to the mean-squared error (MSE) results. The proposed compression method is hardware-implemented and integrated into an AI accelerator to support the on-the-fly compression process with a slight hardware overhead and latency. Experiments show that the LSFC can reduce the required on-chip storage space by 21.9% $\sim$ 33.6% during CNN computation without loss of network prediction. Chenjia Xie, Yuan Du, Zhongfeng Wang 0001 |
ISCAS | 6 |
| 2022 | An X-band Phase Detector Based on Quadrature Modulation in 28-nm CMOSabstractA phase detector (PD) based on quadrature modulation is presented, which is used to perform a phase delay measurement due to signal path at X-band. The X-band signal phase difference is converted to the baseband signal difference through complex frequency conversion. In this paper, an improved active balun and a two-stage tunable poly-phase filter (PPF) are used to generate broadband in-phase and quadrature (I/Q) signals. L-C resonance-based double-balanced Gilbert cells are used as mixers for modulation and demodulation. The proposed X-band PD is implemented in 28-nm CMOS technology and occupies 0.33mm2. The simulated maximal phase error is less than 0.75° over 8-12GHz. The PD consumes 3.5 mW with 0.9V power supply, achieving 0°-180° extended linear phase detection range. Chengqiang Zhao, Wuyu Fan, Jingjing Lv, Yuan Du |
ISCAS | 5 |
| 2022 | Memory-Efficient CNN Accelerator Based on Interlayer Feature Map CompressionabstractExisting deep convolutional neural networks (CNNs) generate massive interlayer feature data during network inference. To maintain real-time processing in embedded systems, large on-chip memory is required to buffer the interlayer feature maps. In this paper, we propose an efficient hardware accelerator with an interlayer feature compression technique to significantly reduce the required on-chip memory size and off-chip memory access bandwidth. The accelerator compresses interlayer feature maps through transforming the stored data into frequency domain using hardware-implemented$8\times 8$discrete cosine transform (DCT). The high-frequency components are removed after the DCT through quantization. Sparse matrix compression is utilized to further compress the interlayer feature maps. The on-chip memory allocation scheme is designed to support dynamic configuration of the feature map buffer size and scratch pad size according to different network-layer requirements. The hardware accelerator combines compression, decompression, and CNN acceleration into one computing stream, achieving minimal compressing and processing delay. A prototype accelerator is implemented on an FPGA platform and also synthesized in TSMC 28-nm COMS technology. It achieves 403GOPS peak throughput and$1.4\times \sim 3.3\times $interlayer feature map reduction by adding light hardware area overhead, making it a promising hardware accelerator for intelligent IoT devices. Yuan Du, Huadong Wei, Chenjia Xie, Zhongfeng Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2022 | An Efficient High-Throughput Structured-Light Depth EngineabstractIn this article, an efficient high-throughput depth engine is proposed to generate high-quality 3-D depth maps for speckle-pattern structured-light depth cameras. A dynamic-binarization (DB) method is introduced with a significant reduction of computational complexity in contrast to the sum-of-absolute-distance (SAD) method. The depth map evaluation shows good robustness compared with other window-based correlation methods. Parallel architecture and reuse of intermediate results are employed for efficient hardware implementation. Our design is verified on a field-programmable gate array (FPGA) and implemented in the SMIC 55-nm CMOS technology, achieving a frame rate of 1731.77 fps ($640\times480$) with an area efficiency of 3.75 fps/KGE. The proposed engine shows a$2.71\times $promotion of area efficiency in contrast to the SAD-based implementation. In addition, the subpixel estimation algorithm deployed in postprocessing is optimized for efficient hardware implementation, reducing the gate count by 69.2% without significant performance loss. Yichuan Bai, Mingzhe Jiang, Qingyu Zhu, Yuan Du, Zhongfeng Wang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2021 | A DNN Optimization Framework with Unlabeled Data for Efficient and Accurate Reconfigurable Hardware InferenceabstractOpen-source deep-learning frameworks are prevalent in designing, training, and deploying deep neural networks (DNNs) on general-purpose computing devices, such as CPU, GPU, and DSP. However, for custom-designed reconfigurable hardware accelerators, there is no existing universal framework, capable of optimizing DNN deployment configuration and guiding the hardware design with specific accuracy and efficiency requirements. In the paper, we proposed a cross- platform framework, which can convert deep-learning models from popular open-source frameworks to intermediate representation and optimize weight/activation dynamic ranges and quantization strategy to achieve better efficiency and accuracy based on a baseline reference design of hardware accelerator. With a few unlabeled data, the proposed framework can analyze the statistical inference information, compare different bit-width impacts, and optimize network structure. We further illustrate the detailed experiment results using the framework, showing mAP and top-1 accuracy loss is less than 1.5% and 1.2% with 12-bit and 8-bit activation-constrained quantization schemes respectively for object detection and image classification. Yuan Du, Xingyu Gu, Zhongfeng Wang 0001 |
ISCAS | 3 |
| 2020 | In-Memory Computing: The Next-Generation AI Computing ParadigmabstractTo overcome the memory bottleneck of von-Neuman architecture, various memory-centric computing techniques are emerging to reduce the latency and energy consumption caused by data communication. The great success of artificial intelligence (AI) algorithms, which involve a large number of computations and data movements, has motivated and accelerated the recent researches of in-memory computing (IMC) techniques to significantly reduce or even diminish the accesses of off-chip data, where memory is not only storing data but can also directly output computation results. For example, the multiply-and-accumulate (MAC) operations in deep learning algorithms can be realized by accessing the memory using the input activations. This paper will investigate the recent trends of IMC from techniques (SRAM, flash, RRAM and other types of non-volatile memory) to architecture and to applications, which will serve as a guide to the future advances on computing in-memory (CIM). Yufei Ma 0002, Yuan Du, Jun Lin 0001, Zhongfeng Wang 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2019 | A 7.5-mW 10-Gb/s 16-QAM wireline transceiver with carrier synchronization and threshold calibration for mobile inter-chip communications in 16-nm FinFETabstractA compact energy-efficient 16-QAM wireline transceiver with carrier synchronization and threshold calibration is proposed to leverage high-density fine-pitch interconnects. Utilizing frequency-division multiplexing, the transceiver transfers four-bit data through one RF band to reduce intersymbol interferences. A forwarded clock is also transmitted through the same interconnect with the data simultaneously to enable low-power PVT-insensitive symbol clock recovery. A carrier synchronization algorithm is proposed to overcome nontrivial current and phase mismatches by including DC offset calibration and dedicated I/Q phase adjustments. Along with this carrier synchronization, a threshold calibration process is used for the transceiver to tolerate channel and circuit variations. The transceiver implemented in 16-nm FinFET occupies only 0.006-mm2 and achieves 10 Gb/s with 0.75-pJ/bit efficiency and <2.5-ns latency. Jieqiong Du, Chien-Heng Wong, Yo-Hao Tu, Wei-Han Cho, Yilei Li, Yuan Du, Po-Tsang Huang, Sheau Jiung Lee, Mau-Chung Frank Chang |
NOCS | 6 |
| 2019 | An Analog Neural Network Computing Engine Using CMOS-Compatible Charge-Trap-Transistor (CTT)abstractAn analog neural network computing engine based on CMOS-compatible charge-trap transistor (CTT) is proposed in this paper. CTT devices are used as analog multipliers. Compared to digital multipliers, CTT-based analog multiplier shows significant area and power reduction. The proposed computing engine is composed of a scalable CTT multiplier array and energy efficient analog-digital interfaces. By implementing the sequential analog fabric, the engine's mixed-signal interfaces are simplified and hardware overhead remains constant regardless of the size of the array. A proof-of-concept 784 by 784 CTT computing engine is implemented using TSMC 28-nm CMOS technology and occupies 0.68 mm2. The simulated performance achieves 76.8 TOPS (8-bit) with 500 MHz clock frequency and consumes 14.8 mW. As an example, we utilize this computing engine to address a classic pattern recognition problem-classifying handwritten digits on MNIST database and obtained a performance comparable to state-of-the-art fully connected neural networks using 8-bit fixed-point resolution. Yuan Du, Xuefeng Gu, Jieqiong Du, X. Shawn Wang, Boyu Hu, Mingzhe Jiang, Xiaoliang Chen 0001, Subramanian S. Iyer, Mau-Chung Frank Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2018 | A Single Layer 3-D Touch Sensing System for Mobile Devices ApplicationabstractTouch sensing has been widely implemented as a main methodology to bridge human and machine interactions. The traditional touch sensing range is 2-D and therefore limits the user experience. To overcome these limitations, we propose a novel 3-D contactless touch sensing called Airtouch system, which improves user experience by remotely detecting single/multi-finger position. A single layer touch panel with triangle-shaped electrodes is proposed to achieve multitouch detection capability as well as manufacturing cost reduction. Moreover, an oscillator-based-capacitive touch sensing circuit is implemented as the sensing hardware with the bootstrapping technique to eliminate the interchannel coupling effects. To further improve the system accuracy, a grouping algorithm is proposed to group the useful channels' data and filter out hardware noise impact. Finally, improved algorithms are proposed to eliminate the fringing capacitance effect and achieve accurate finger position estimation. EM simulation proved that the proposed algorithm reduced the maximum systematic error by 11 dB in the horizontal position detection. The proposed system consumes 2.3 mW and is fully compatible with existing mobile device environments. A prototype is built to demonstrate that the system can successfully detect finger movement in a vertical direction up to 6 cm and achieve a horizontal resolution up to 0.6 cm at 1 cm finger-height. As a new interface for human and machine interactions, this system offers great potential in finger movement detection and gesture recognition for small-sized electronics and advanced human interactive games for mobile device. Yan Zhang 0050, Yilei Li, Yuan Du, Yen-Cheng Kuan, Mau-Chung Frank Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2018 | A Novel Fully Synthesizable All-Digital RF Transmitter for IoT ApplicationsabstractIn this paper, a fully synthesizable all-digital transmitter (ADTX) is first proposed. This transmitter (TX) uses Cartesian architecture and supports wide-band quadratic-amplitude modulation with wide carrier frequency range. Furthermore, the design methodology for ADTX and corresponding bandpass filter is discussed. This TX is synthesized with digital register transfer level-graphic database system flow, and can be easily implemented in any standard CMOS technology. An exemplary TX is synthesized by TSMC 28-nm standard cell library with extremely small area (0.0009 mm2) and supports carrier frequency as high as 6 GHz with excellent error vector magnitude (<;-30 dB). To the best of the authors' knowledge, this is the first work on a fully synthesizable design of RF transistors, allowing easy technology migration and portability. Yilei Li, Kirti Dhwaj, Chien-Heng Wong, Yuan Du, Yiwu Tang, Yiyu Shi 0001, Tatsuo Itoh, Mau-Chung Frank Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2017 | Retention-Aware Hybrid Main Memory (RAHMM): Big DRAM and Little SCMabstractHybrid memory comprised of a big SCM and a little DRAM (BSLD) is widely studied to address the growing power consumption challenge of pure DRAM. However, the performance degradation, limited endurance and immature mass production of ultra-high-density SCM are still the painful points of BSLD. Here we propose a Retention-Aware Hybrid Main Memory (RAHMM) architecture with a big DRAM and a little SCM (BDLS) for the first time. DRAM is refreshed at a much longer interval by using SCM to store the small quantity of leaky tail bits in DRAM. A two-step search technology combined with outcome forecasting is put forward to get ultra-fast read access, as well as to diminish the power and performance overheads. A hidden buffer strategy (HBS) is proposed to optimize write performance and endurance hurt. The experimental results show 45 percent reduction of power consumption and 30 percent performance optimization, which are significantly improved compared to that of both serial and parallel BSLD with a counterpart capacity Weiliang Jing, Yinyin Lin, Beomseop Lee, Sangkyu Yoon, Yuan Du, Bomy Chen |
IEEE Trans. Computers | 7 |
| 2017 | An R2R-DAC-Based Architecture for Equalization-Equipped Voltage-Mode PAM-4 Wireline Transmitter DesignabstractThis brief presents a wireline transmitter architecture, enabling multilevel signaling with feedforward equalization (FFE) in voltage-mode. A compact R2R-DAC-based front end is proposed and analyzed in terms of its speed, power consumption, and linearity. A voltage-mode PAM-4 transmitter with 2-tap FFE utilizing the proposed architecture is implemented in the 65-nm CMOS technology. It achieves a data rate of 34 Gb/s and an energy efficiency of 2.7 mW/Gb/s. Boyu Hu, Yuan Du, Rulin Huang, Jeffrey Lee, Young-Kai Chen, Mau-Chung Frank Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | Aspectual Properties of Conversational ActivitiesabstractSegmentation of spoken discourse into distinct conversational activities has been applied to broadcast news, meetings, monologs, and two-party dialogs. This paper considers the aspectual properties of discourse segments, meaning how they transpire in time. Classifiers were con-structed to distinguish between segment boundaries and non-boundaries, where the sizes of utterance spans to represent data instances were varied, and the locations of segment boundaries relative to these in-stances. Classifier performance was better for representations that included the end of one discourse segment combined with the beginning of the next. In addition, classi-fication accuracy was better for segments in which speakers accomplish goals with distinctive start and end points. 1 Rebecca J. Passonneau, Boxuan Guan, Cho Ho Yeung, Yuan Du, Emma Conner |
SIGDIAL Conference | 4 |
| 2013 | ConSub: Incentive-Based Content Subscribing in Selfish Opportunistic Mobile NetworksabstractRecently, content-based publish/subscribe (pub/sub) services have become a significant research field in opportunistic mobile networks (OppNets). Pub/sub is an asynchronous messaging paradigm, in which content transmissions are guided by the interest. Since selfish behavior is common in reality, nodes often behave selfishly with an aim to maximize their own utilities without considering performance of other nodes. Therefore, how to encourage nodes to collect, store and share network content efficiently is one of the key challenges under this paradigm. In this paper, we propose an incentive-based pub/sub scheme, called ConSub, for OppNets. In ConSub, Tit-For-Tat (TFT) mechanism is employed to deal with selfish behavior. ConSub also implements a content exchange protocol between two interacting node, thus encouraging them to play as businessmen and carry contents to satisfy each other's interest. Specifically, the exchange order is determined by the content utility, which is calculated by contact probability and cooperation level between the current node and its neighbors subscribing to the interest. Extensive realistic trace-driven simulation results show that ConSub is superior to existing schemes in terms of delivered packets and transmission hops with reasonable transmission cost. Huan Zhou 0002, Jiming Chen 0001, Jialu Fan, Yuan Du, Sajal K. Das 0001 |
IEEE J. Sel. Areas Commun. | 4 |
| 2013 | Geocommunity-Based Broadcasting for Data Dissemination in Mobile Social NetworksabstractIn this paper, we consider the issue of data broadcasting in mobile social networks (MSNets). The objective is to broadcast data from a superuser to other users in the network. There are two main challenges under this paradigm, namely 1) how to represent and characterize user mobility in realistic MSNets; 2) given the knowledge of regular users' movements, how to design an efficient superuser route to broadcast data actively. We first explore several realistic data sets to reveal both geographic and social regularities of human mobility, and further propose the concepts of geocommunity and geocentrality into MSNet analysis. Then, we employ a semi-Markov process to model user mobility based on the geocommunity structure of the network. Correspondingly, the geocentrality indicating the “dynamic user density” of each geocommunity can be derived from the semi-Markov model. Finally, considering the geocentrality information, we provide different route algorithms to cater to the superuser that wants to either minimize total duration or maximize dissemination ratio. To the best of our knowledge, this work is the first to study data broadcasting in a realistic MSNet setting. Extensive trace-driven simulations show that our approach consistently outperforms other existing superuser route design algorithms in terms of dissemination ratio and energy efficiency. Jialu Fan, Jiming Chen 0001, Yuan Du, Wei Gao 0006, Jie Wu 0001, Youxian Sun |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2011 | Tilt & touch: mobile phone for 3D interactionabstractMobile phones are becoming de facto pervasive devices for people's daily use. This demonstration illustrates a new interaction, Tilt & Touch, to enable a smart phone to be a 3D controller. It exploits capacitive touchscreen and built-in MEMS motion sensors. When people want to navigate in a virtual reality environment on a large display, they can tilt the phone for viewpoint transforming, touch the phone screen for avatar moving, and pinch screen for viewing camera zooming. The virtual objects in the virtual reality environment can be rotated accordingly by tilting the phone. Yuan Du, Haoyi Ren, Gang Pan 0001, Shijian Li |
UbiComp | 1 |
| 2011 | Broadcast yourself: understanding YouTube uploadersabstractYouTube uploaders are the central agents in the YouTube phenomenon. We conduct extensive measurement and analysis of YouTube uploaders. We estimate YouTube scale and examine the uploading behavior of YouTube users. We demonstrate the positive reinforcement between on-line social behavior and uploading behavior. Furthermore, we examine whether YouTube users are truly broadcasting themselves, via characterizing and classifying videos as either user generated or user copied. Yuan Ding 0003, Yuan Du, Yingkai Hu, Zhengye Liu, Luqin Wang, Keith W. Ross, Anindya Ghose |
Internet Measurement Conference | 2 |
| 2011 | Experimental analysis of user mobility pattern in mobile social networksabstractMobility pattern of device users plays a crucial role in a wide range of mobile computing applications, including data forwarding, content sharing, information search and advertising. Hence, it is important to characterize the mobility path information of users, so as to accurately predict user mobility. In this paper, we introduce two typical user mobility patterns: standard Markov and semi-Markov models. Especially, we experimentally explore the correlation of community and geography information in Mobile Social Networks (MSNets), and analyze user sojourn time distribution over communities. Both of theoretical analysis and trace-driven simulation results show that semi-Markov model is more effective in characterizing user mobility pattern and further making more accurate mobility prediction compared with standard Markov model. Yuan Du, Jialu Fan, Jiming Chen 0001 |
WCNC | 1 |
| 2010 | Geography-aware active data dissemination in mobile social networksabstractIn mobile social networks (MSNets), data dissemination is an important topic, which has not been widely investigated yet. Active data dissemination is a networking paradigm where a superuser intentionally facilitates the connectivity in the network. One of the key challenges under this paradigm is how to design the most efficient superuser route to achieve certain properties of end-to-end connectivity. Most existing solutions only focus on the network with stationary users or strongly constrained node mobility, and assume the superuser always moves with a fixed route. In this paper, we propose a flexible approach to design the superuser routes, considering the realistic user movements in MSNets. To the best of our knowledge, this work is the first to study active data dissemination from the social network perspective. We explore the geographic regularity of human mobility in the network, employ a semi-Markov analytical model to describe such mobility pattern, and hence formulate the superuser route design as a combinational optimization problem of Convex Optimization and Traveling Salesman Problem by exploiting social network concepts including communities and centrality. Extensive trace-driven simulations show that our approach consistently outperforms other existing superuser route design algorithms in terms of delivery ratio and energy efficiency. Jialu Fan, Yuan Du, Wei Gao 0006, Jiming Chen 0001, Youxian Sun |
MASS | 2 |