EDBT 2026 Demo / reviewers in the wild / expert
Le Ye
dblp:94/9850
· DBLP profile ↗
45ranked-venue papers
4as first author
32since 2021 · last 2026
0000-0003-0599-7762ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 42 · 3 first-author · 30 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RV-CIM: Energy-Delay Optimized Mapping and Architecture Co-Design for a RISC-V Multi-core SoC with Configurable DCIM Cluster
Ninghui Shang, Ying Liu 0069, Zecheng Zhou, Jiyong Hu, Zhiqiang Guo, Zhiyuan Chen 0009, Guoxiang Li, Yufei Ma 0002, Le Ye |
APPT | 11 |
| 2026 | Quartet: A Digital Compute-in-Memory Versatile AI Accelerator With Heterogeneous Tensor Engines and Off-Chip-Less DataflowabstractAlthough most AI core operations can be formulated as matrix multiplications (MMs), their characteristics are quite different. In some cases, one input may be a constant weight matrix while the other is a dynamic feature matrix, i.e., FWMM, or both inputs may be dynamic features as FFMM. Furthermore, the data involved can also be sparse, leading to variations like SpFWMM and SpFFMM. To address this challenge, this paper investigates a versatile accelerator architecture for AI algorithms based on heterogeneous tensor engines. For MM operators with varying characteristics, this paper proposes four-quadrant heterogeneous tensor engines to handle FWMM, SpFWMM, FFMM, and SpFFMM, respectively. These four tensor engines are comprised of SRAM-based single address digital compute-in-memory (CIM) array, SRAM-based multi-address digital CIM array, systolic array, and multi-SIMD array, respectively. In addition, to improve the AI execution efficiency, this paper proposes a dual-level multi-issue mechanism to achieve inter-operator and inter-block parallelization along with an off-chip-less dataflow enabled by on-chip unified memory pool. Thanks to the integration of aforementioned innovations, this paper develops a versatile AI acceleration chip Quartet, which achieves exceptional energy efficiency. Specifically, for graph convolutional network on PubMed, it demonstrates a$19.56\times $and$3.47\times $improvement in energy efficiency compared to similar works, ReDCIM and TensorCIM, respectively. Yikan Qiu, Guoxiang Li, Meng Wu 0005, Yifan Jia 0009, Le Ye, Yufei Ma 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | GSNorm: An Efficient 3D Gaussian Rendering Accelerator with Splat Normalization and LUT-assist Rasterizationabstract3D Gaussian Splatting recently emerged as the new SOTA approach for many computer graphic tasks. While Gaussian Splatting has demonstrated impressive rendering quality and performance on GPUs, real-time GS rendering on edge devices is still challenging. We identified the unbalanced Rendering pipeline and the uneven Gaussian distribution as the main obstacles to efficient rendering. To address these problems, we present GSNorm, a rendering accelerator with an online quantization preprocessor for per-gaussian coordinate transformation to normalize Gaussian footprints for pixel-wise calculation reduction. A LUT-based quantized rendering design is also presented to break the pipeline data dependency. Furthermore, a depth-guided cluster-sorting unit is incorporated to improve Gaussian sorting efficiency. GSNorm accelerator is implemented and evaluated in TSMC 22 nm technology with several real-world scenes, providing significant rendering efficiency and performance improvements for real-time applications. Peiran Yan, Yiqi Jing, Le Ye |
ASP-DAC | 4 |
| 2025 | Accelerating Diffusion Transformer via Increment-Calibrated Caching with Channel-Aware Singular Value DecompositionabstractDiffusion transformer (DiT) models have achieved remarkable success in image generation, thanks for their exceptional generative capabilities and scalability. Nonetheless, the iterative nature of diffusion models (DMs) results in high computation complexity, posing challenges for deployment. Although existing cache-based acceleration methods try to utilize the inherent temporal similarity to skip redundant computations of DiT, the lack of correction may induce potential quality degradation. In this paper, we propose increment-calibrated caching, a training-free method for DiT acceleration, where the calibration parameters are generated from the pre-trained model itself with low-rank approximation. To deal with the possible correction failure arising from outlier activations, we introduce channel-aware Singular Value Decomposition (SVD), which further strengthens the calibration effect. Experimental results show that our method always achieve better performance than existing naive caching methods with a similar computation resource budget. When compared with 35-step DDIM, our method eliminates more than 45% computation and improves IS by 12 at the cost of less than 0.06 FID increase. Zhiyuan Chen 0009, Yifan Jia 0009, Le Ye, Yufei Ma 0002 |
CVPR | 4 |
| 2025 | EdgeMM: Multi-Core CPU with Heterogeneous AI-Extension and Activation-aware Weight Pruning for Multimodal LLMs at EdgeabstractEmerging multimodal LLMs (MLLMs) exhibit strong cross-modality perception and reasoning capabilities and hold great potential for various applications at edge. However, MLLMs typically consist of a compute-intensive modality encoder and a memory-bound LLM decoder, leading to distinct bottlenecks for hardware designs. In this work, we present a multi-core CPU solution with heterogeneous AI extensions, which are based on either the compute-centric systolic array or memory-centric digital compute-in-memory (CIM) coprocessors. In addition, dynamic activation-aware weight pruning and bandwidth management are developed to enhance bandwidth efficiency and core utilization, improving overall performance. We implemented our solution using commercial 22nm technology. For representative MLLMs, our evaluations show EdgeMM can achieve $2.84 \times$ performance speedup compared to laptop 3060 GPU. Kangbo Bai, Le Ye, Ru Huang 0001 |
DAC | 2 |
| 2025 | 3D-SubG: A 3D Stacked Hybrid Processing Near/In-Memory Accelerator for Subgraph GNNsabstractSubgraph Graph Neural Networks (GNNs) are emerging as a promising approach to enhance GNN expressiveness, but their more complex graph structures with numerous independent and irregular subgraphs pose significant hardware deployment challenges. In this work, we propose 3D-SubG, a 3D stacked hybrid processing-near/in-memory accelerator for subgraph GNNs. With hybrid bonding packaging technology, a logic die is 3D stacked with a DRAM die for highly parallel memory accesses. The logic die employs digital SRAM-based processing-in-memory (PIM) macros to boost computation density and minimize data transfer. We further propose a bit-level non-zero gathering method to exploit graph sparsity for PIM, a workloadbalanced mapping strategy for subgraph allocation onto different logic-to-DRAM blocks, and a distributed global pooling approach to reduce inter-block data movements. Experimental results show that 3D-SubG achieves average improvements of $146.11 \times$ in performance, $934.18 \times$ in area efficiency, and $1171.80 \times$ in energy efficiency compared to RTX 3090Ti. Guoxiang Li, Runnan Xu, Ruohang Xu, Yikan Qiu, Renati Tuerhong, Muhan Zhang, Le Ye, Yufei Ma 0002 |
DAC | 7 |
| 2025 | Local-GS: An Order-Independent Gaussian Splatting Training Accelerator Exploiting Splat Localityabstract3D Gaussian Splatting has emerged as the SOTA approach for 3D representation and view synthesis. While Gaussian Splatting has demonstrated impressive capability and rendering quality on desktop GPUs, achieving on-demand training on resource-constrained edge devices is still challenging. In this work, we identified the training bottleneck from a few perspectives including algorithm splat locality and the limited memory and hardware under-utilization. To address these problems, we present Local-GS, a 3D Gaussian Splatting training accelerator with order-independent rendering to break the depth-wise data dependency between overlapping Gaussians. We further incorporate a parallel pixel intersection test unit to schedule thread workload based on Gaussian splat locality and improve hardware utilization. A set of unified training-rendering cores are designed to achieve efficient splat-level parallel rendering and gradient propagation. Our Local-GS is implemented in 7 nm and is evaluated by several real-world 3D scenes. Compared to edge Jetson NX GPU, Local-GS achieve 26.9-53 $\times$ training speedup and three orders of magnitude efficiency boost. Qinzhe Zhi, Yiqi Jing, Le Ye, Ru Huang 0001 |
DAC | 4 |
| 2025 | 3D-TokSIM: Stacking 3D Memory with Token-Stationary Compute-in-Memory for Speculative LLM InferenceabstractThe LLM decoding process poses a significant challenge for memory bandwidth due to its autoregressive nature. Prior 2D memory solutions fail to overcome this memory bottleneck due to limited memory-to-logic bandwidth. In this work, we propose 3D-TokSIM, a cross-stack solution by stacking 3D memory on logic die with a specially designed token-stationary compute-in-memory (CIM) to efficiently accelerate speculative decoding. Our CIM is developed with novel token-stationary dataflow to reduce data movement on logic die to save power and balance computation and memory access. To further reduce the buffer requirements, we perform architecture exploration and allocate notable CIM resources for achieving higher decoding parallelism. Compared to RTX 3090 GPU, 3D-TokSIM achieves 15.1 $\times$ throughput and $324 \times$ energy efficiency improvements on speculative Llama2-7B decoding. Boya Lv, Meng Wu 0005, Fengyun Yan, Yufei Ma 0002, Ru Huang 0001, Le Ye |
DAC | 9 |
| 2025 | Leveraging Compute-in-Memory for Efficient Generative Model Inference in TPUsabstractWith the rapid advent of generative models, efficiently deploying these models on specialized hardware has become critical. Tensor Processing Units (TPUs) are designed to accelerate AI workloads, but their high power consumption neces-sitates innovations for improving efficiency. Compute-in-memory (CIM) has emerged as a promising paradigm with superior area and energy efficiency. In this work, we present a TPU architecture that integrates digital CIM to replace conventional digital systolic arrays in matrix multiply units (MXUs). We first establish a CIM-based TPU architecture model and simulator to evaluate the benefits of CIM for diverse generative model inference. Building upon the observed design insights, we further explore various CIM-based TPU architectural design choices. Up to 44.2% and 33.8% performance improvement for large language model and diffusion transformer inference, and 27.3 × reduction in MXU energy consumption can be achieved with different design choices, compared to the baseline TPUv4i architecture. Zhantong Zhu, Hongou Li, Wenjie Ren, Meng Wu 0005, Le Ye, Ru Huang 0001 |
DATE | 5 |
| 2025 | H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM InferenceabstractLarge language models (LLMs) have demonstrated remarkable proficiency in a wide range of natural language processing applications. However, the high energy and latency overhead induced by the KV cache limits the edge deployment, especially for long contexts. Emerging hybrid bonding (HB) technology has been proposed as a promising alternative to conventional near-memory processing (NMP) architectures, offering improved bandwidth efficiency and lower power consumption while exhibiting characteristics of distributed memory.In this paper, we propose H2EAL, an HB-based accelerator with sparse attention algorithm-hardware co-design for efficient LLM inference at the edge. At the algorithm level, we propose a hybrid sparse attention scheme with static and dynamic sparsity for different heads to fully leverage the sparsity with high accuracy. At the hardware level, we co-design the hardware to support hybrid sparse attention and propose memory-compute co-placement to address the distributed memory bottleneck. Since different attention heads exhibit different sparse patterns and the attention structure often mismatches the HB architecture, we further develop a load-balancing scheduler with parallel tiled attention to address workload imbalance and optimize the mapping strategy. Extensive experiments demonstrate H2EAL achieves 5.20 ∼ 48.21× speedup and 6.22 ∼ 73.48× energy efficiency improvement over baseline HB implementation, with a negligible average accuracy drop of 0.87% on multiple benchmarks. Zizhuo Fu, Xiaotian Guo, Wenxuan Zeng, Shuzhang Zhong, Runsheng Wang, Le Ye, Meng Li 0004 |
ICCAD | 8 |
| 2025 | CIMTester: An Agile Golden-Result-Free BIST Compiler for Robust Compute-In-MemoryabstractDigital compute-in-memory (DCIM) is playing an increasingly vital role in efficient AI computing due to its significant efficiency advantages. However, the combination of memory and computation logics in DCIM presents challenges for testing, including extra coupling fault between memory and logic, high overhead for golden result generation and indirect fault location. In this paper, we present a golden-result-free BIST structure together with a computation-coupled CIM BIST algorithm to reduce testing overhead and improve test coverage. In addition, we present CIMTester, a DCIM BIST compiler to adapt to swiftly changing test requirements and DCIM macro sizes and numbers in chips. The template-based generator generates BIST RTL based on proposed BIST structure and modifies the templates according to the architecture parameters of test chip. CIMTester’s iterator analyzes diverse sharing strategy of BIST components to satisfy user specifications of area, test time and fault coverage. We implemented and evaluated a series of TSMC 22nm DCIM macros with the generated BIST circuits, which achieves up to 99.48% fault coverage with only less than 2.44% area overhead. Wenjie Ren, Meng Wu 0005, Le Ye |
ICCAD | 6 |
| 2025 | A 22nm All-digital Fully Synthesizable Adaptive Clock Generator for Fast Frequency Adjustment in AI ProcessorabstractThis paper presents an adaptive clock generator designed in all-digital fully-synthesizable manner for fast frequency management in AI processor. To adapt frequency for dynamic AI workloads, a multi-phase clock chain is designed and a phase picker dynamically switches clock phases to stretch or shrink clock frequency. The digital loop control of PLL employs a coarse exponential and fine phase frequency locking to improve locking time. To ensure the stability, a well-organized floorplan is targeted during the EDA backend flow. Our adaptive clock generator is implemented on 22nm with an AI processor design. The evaluation results demonstrate the design achieves fast clock frequency adjustment within a delay of 1 clock cycle, with advanced area, power, and response time metrics. Le Ye |
ISCAS | 2 |
| 2024 | An In-Memory Computing Accelerator with Reconfigurable Dataflow for Multi-Scale Vision Transformer with Hybrid TopologyabstractTransformer models equipped with multi-head attention (MHA) mechanism have demonstrated promise in computer vision (CV) tasks, i.e., vision transformers (ViTs). Nevertheless, the lack of inductive bias in ViTs leads to substantial computational and storage requirements, hindering their deployment on resource-constrained edge devices. To this end, multi-scale hybrid models are proposed to take the advantages of both transformers and convolutional neural networks (CNNs). However, existing domain-specific architectures focus on the optimization of either convolution or MHA at the expense of flexibility. In this work, an in-memory computing (IMC) accelerator is proposed to efficiently accelerate ViTs with hybrid MHA and convolution topology by introducing pipeline reordering. SRAM-based digital IMC macro is utilized to mitigate memory access bottleneck, while avoiding analog non-ideality. The reconfigurable processing engines and interconnections are investigated to enable the adaptable mapping of both convolution and MHA. Under typical workloads, experimental results exhibit that our proposed IMC architecture delivers 2.20× to 2.52× speedup and 40.6% to 74.8% energy reduction compared with the baseline design. Zhiyuan Chen 0009, Yufei Ma 0002, Yifan Jia 0009, Guoxiang Li, Meng Wu 0005, Le Ye, Ru Huang 0001 |
DAC | 8 |
| 2024 | AIG-CIM: A Scalable Chiplet Module with Tri-Gear Heterogeneous Compute-in-Memory for Diffusion AccelerationabstractThe emergence of Diffusion models has gained significant attention in the field of Artificial Intelligence Generated Content. While Diffusion demonstrates impressive image generation capability, it faces hardware deployment challenges due to its unique model architecture and computation requirement. In this paper, we present a hardware accelerator design, i.e. AIG-CIM, which incorporates tri-gear heterogeneous digital compute-in-memory to address the flexible data reuse demands in Diffusion models. Our framework offers a collaborative design methodology for large generative models from the computational circuit-level to the multi-chip-module system-level. We implemented and evaluated the AIG-CIM accelerator using TSMC 22nm technology. For several Diffusion inferences, scalable AIG-CIM chiplets achieve 21.3× latency reduction, up to 231.2× throughput improvement and three orders of magnitude energy efficiency improvement compared to RTX 3090 GPU. Yiqi Jing, Meng Wu 0005, Yufei Ma 0002, Ru Huang 0001, Le Ye |
DAC | 8 |
| 2024 | SPARK: An Efficient Hybrid Acceleration Architecture with Run-Time Sparsity-Aware Scheduling for TinyML LearningabstractCurrently most TinyML devices only focus on inference, as training requires much more hardware resources. In this paper, we introduce SPARK, an efficient hybrid acceleration architecture with run-time sparsity-aware scheduling for TinyML learning. Besides a standalone accelerator, an in-pipeline acceleration unit is integrated within the CPU pipeline to support simultaneous forward and backward propagation. To better utilize sparsity and improve hardware utilization, a sparsity-aware acceleration scheduler is implemented to schedule the workload between two acceleration units. A unified memory system is also constructed to support transposable data fetch, reducing memory access. We implement SPARK using TSMC 22nm technology and evaluate different TinyML tasks. Compared with the baseline accelerator, SPARK achieves 4.1× performance improvement in average with only 2.27% area overhead. SPARK also outperforms off-shelf edge devices in performance by 9.4× with 446.0× higher efficiency. Qinzhe Zhi, Yanchi Dong, Le Ye |
DAC | 4 |
| 2024 | Hierarchical Power Co-Optimization and Management for LLM Chiplet DesignsabstractThe demand for efficient and high-performance hardware for large language models (LLMs) has driven the development of scalable chiplet design, which requires careful power optimization and management. This paper presents a co-optimization and management methodology for hierarchical power delivery of chiplet designs targeting LLM applications. To model LLM workload mapping and power delivery, we first build a scalable chiplet simulator, which demonstrates different power strategies have notable efficiency impact and require careful and thorough optimizations. We further develop a co-optimization framework ScalePoM for chiplet power management. Based on given LLM model and PPA requirements, ScalePoM can automatically explore the chiplet architecture and workload mapping for optimal hierarchical power delivery. Our co-optimization methodology is evaluated through two scaled LLM chiplets with different interconnect topologies, achieving an average of 45% and up to 62% energy saving for large language model inferences with various sparsity levels. Yanchi Dong, Xiaochen Hao, Yun Liang 0001, Ru Huang 0001, Le Ye |
ICCAD | 6 |
| 2024 | Sparsity-Aware In-Memory Neuromorphic Computing Unit With Configurable Topology of Hybrid Spiking and Artificial Neural NetworkabstractSpiking neural networks (SNNs) have shown great potential in achieving high energy efficiency and low power consumption compared to artificial neural networks (ANNs). However, there remains a significant accuracy gap between SNNs and ANNs. To address this issue, we present an in-memory neuromorphic computing (IMNC) chip that supports hybrid spiking/artificial neural networks (S/ANNs) and sparsity-aware data flows. With the IMNC chip, we aim to improve inference accuracy while simultaneously achieving high energy efficiency through optimization at the algorithm, architecture, and circuit levels. First, at the algorithm level, we note that SNNs extract temporal features from input spikes using time-domain convolution operations. Based on this insight, we efficiently utilize leaky integrate (LI) neurons to hybridize SNNs and ANNs, thereby improving accuracy while maintaining highly sparse operations. Second, at the architecture level, we design a sparsity-aware architecture that supports a hybrid S/ANN topology with varying sparsity. Finally, at the circuit level, we propose a ring-based in-memory computing (IMC) macro, whose energy consumption is inversely proportional to the input sparsity, making it ideal for performing energy-efficient multiplication and accumulation (MAC) operations in both SNNs and ANNs. We evaluate the proposed hybrid S/ANNs on various classification tasks and demonstrate their stronger classification and generalization ability compared with pure SNNs. Notably, our IMNC chip, fabricated using 22 nm CMOS technology, achieves impressive measured accuracy rates of over 95% for voice activity detection (VAD) and ECG anomaly detection. Additionally, our IMNC chip demonstrates superior dynamic energy efficiency of 0.43 pJ per synaptic operation, outperforming related works. Ying Liu 0069, Zhiyuan Chen 0009, Zhixuan Wang, Ru Huang 0001, Le Ye, Yufei Ma 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2024 | DCIM-GCN: Digital Computing-in-Memory Accelerator for Graph Convolutional NetworkabstractGraph convolutional network (GCN) has gained great success in a diverse range of intelligent tasks. However, the hardware performance of GCNs is often bounded by random and non-continuous memory accesses due to the sparse graph data, which incur high latency and high power consumption. The emerging computing-in-memory (CIM) architecture significantly reduces the overhead of data movements, which is suitable for memory-intensive GCN acceleration. Existing analog-based CIM solutions require a large amount of analog-to-digital (AD) and digital-to-analog (DA) conversions, which dominate the overall area and power consumption. Furthermore, the analog non-ideality can degrade accuracy and reliability of CIM. To address these challenges, this work proposes a digital CIM accelerator based on SRAM, called DCIM-GCN, to accelerate GCN algorithm. DCIM-GCN introduces innovations on three levels: circuit, architecture, and algorithm. At the circuit level, digital CIM is proposed with SRAM sub-arrays to eliminate the power and area expensive AD/DA converters. Furthermore, we have incorporated the multi-address feature into the digital CIM, thereby leveraging its ability to efficiently process sparse matrix multiplication. At the architecture level, the sparsity-aware computation engine takes advantage of sparsity in GCNs and leverages CIM to minimize memory accesses and data movements. Finally, at the algorithm level, the balance mapping algorithm tackles workload imbalance issues, while the vertex reorder algorithm reduces idle states for aggregation engines, resulting in increased hardware utilization. Our DCIM-GCN achieves 1.89$\times$and 2.42$\times$speedup and 4.58$\times$and 9.46$\times$energy efficiency improvement on average over other CIM-based graph accelerators, e.g., PASGCN and PIM-GCN, respectively. Yufei Ma 0002, Yikan Qiu, Guoxiang Li, Meng Wu 0005, Le Ye, Ru Huang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2023 | RIMAC: An Array-Level ADC/DAC-Free ReRAM-Based in-Memory DNN Processor with Analog Cache and ComputationabstractBy directly computing in analog domain, processing-in-memory (PIM) is emerging as a promising alternative to overcome the memory bottleneck of traditional von-Neuman architecture, especially for deep neural networks (DNNs). However, the data outside PIM macros in most existing PIM accelerators are stored and operated as digital signals that require massive expensive digital-to-analog (D/A) and analog-to-digital (A/D) converters. In this work, an array-level ADC/DAC-free ReRAM-based in-memory DNN processor named RIMAC is proposed, which accelerates various DNNs in pure analog-domain with analog cache and analog computation modules to eliminate the expensive D/A and A/D conversions. Our experiment result shows the peak energy efficiency is improved by about 34.8×, 97.6×, 10.7×, and 14.0× compared to PRIME, ISAAC, Lattice, and 21'DAC for various DNNs on ImageNet, respectively. Meng Wu 0005, Yufei Ma 0002, Le Ye, Ru Huang 0001 |
ASP-DAC | 4 |
| 2023 | A Model-Specific End-to-End Design Methodology for Resource-Constrained TinyML HardwareabstractTiny machine learning (TinyML) becomes appealing as it enables machine learning on resource-constrained devices with ultra low energy and small form factor. In this paper, a model-specific end-to-end design methodology is presented for TinyML hardware design. First, we introduce an end-to-end system evaluation method using Roofline models, which considering both AI and other general-purpose computing to guide the architecture design choices. Second, to improve the efficiency of AI computation, we develop an enhanced design space exploration framework, TinyScale, to enable optimal low-voltage operation for energy-efficient TinyML. Finally, we present a use case driven design selection method to search the optimal hardware design across a set of application use cases. Our model-specific design methodology is evaluated on both TSMC 22nm and 55nm technology for MLPerf Tiny benchmark and a keyword spotting (KWS) SoC design. With the help of our end-to-end design methodology, an optimal TinyML hardware can be automatically explored with significant energy and EDP improvements for a diverse of TinyML use cases. Yanchi Dong, Kaixuan Du, Yiqi Jing, Qijun Wang, Pixian Zhan, Fengyun Yan, Yufei Ma 0002, Yun Liang 0001, Le Ye, Ru Huang 0001 |
DAC | 11 |
| 2023 | ARES: A Mapping Framework of DNNs Towards Diverse PIMs with General AbstractionsabstractNumerous architectures based on processing-in-memory (PIM) have recently emerged, exhibiting diversity in memory types, compute functions, memory mapping constraints, etc. To effectively utilize PIM hardware for deploying deep neural networks (DNNs), programmers face the challenge of mapping computations and data across multiple memory arrays, scheduling computation and data transfers, while adhering to various hardware constraints. Existing mapping approaches, however, are tailored to specific architectures and lack a general formulation for mapping optimization, limiting their applicability and performance. In this paper, we present ARES, a comprehensive mapping framework designed for diverse PIM architectures. The core of the framework is hardware abstractions for PIMs, which is inspired by the fact that DNNs on PIM hardware can be represented by a tensorized compute function and data layout constraints in the memory array. This abstraction forms the basis for constructing a mapping space that encompasses both compute and memory constraints. Through exploration of this mapping space, we derive efficient mapping strategies tailored to different PIM hardware configurations. Experimental evaluation conducted on four distinct hardware architectures demonstrates that compared to state-of-the-art mapping methods, ARES yields up to a 70% speed improvement for single operator mapping and a 50% speedup for overall network mapping. Xiuping Cui, Size Zheng 0001, Le Ye, Yun Liang 0001 |
ICCAD | 4 |
| 2023 | An Information-Aware Adaptive Data Acquisition System using Level-Crossing ADC with Signal-Dependent Full Scale and Adaptive Resolution for IoT ApplicationsabstractThis paper proposes an information-aware (IA) adaptive data acquisition (ADA) system for the Internet of Things (IoT) applications. The system can obtain valid information adaptively thanks to 1) signal-dependent full-scale feature tracks the amplitude-domain activity of the event; 2) level-crossing (LC) ADC with slope detector delivers the time-domain activity; 3) the IA algorithm determines the quantization resolution according to the detected signal activities. The proposed clock-free event-driven ADA system can reject the redundant data, and compress the valid data from the source, thus saving its power and the power of subsequent data-processing systems. The long-term average power consumption of the system is 128 nW, the resolution varies from 3 to 7 bits according to the input signal state. Compared with conventional ADCs, LC-ADC can compress the data by 2.5x [1]. Further, the proposed system has 15x higher compression ratio (CR) than that of LC-ADC. Yiqi Jing, Zhixuan Wang, Linxiao Shen, Yihan Zhang 0002, Jiayoon Ru, Le Ye |
ISCAS | 7 |
| 2023 | DCIM-3DRec: A 3D Reconstruction Accelerator with Digital Computing-in-Memory and Octree-Based SchedulerabstractLearning-based 3D reconstruction has evolved rapidly with promising quality, while it requires high-performance hardware for interactive applications. In this work, a reconstruction accelerator called DCIM-3DRec is presented which leverages digital computing-in-memory (DCIM) design to facilitate learning-based reconstruction deployment on realtime and low-power edge platforms. The DCIM-3DRec is designed with the following features: a reconfigurable DCIM macro array for high data reuse and macro utilization, and an Octree-based subdivision scheduler for efficient management of 3D space prediction. The DCIM-3DRec accelerator is implemented and evaluated in TSMC 55 nm technology, with a DCIM macro efficiency of 19.4 TOPS/W at INT8. Overall, the DCIM-3DRec accelerator achieves 23× performance gain and four orders of magnitude energy efficiency improvement compared to a Nvidia RTX3090 GPU. Yiqi Jing, Meng Wu 0005, Fengyun Yan, Yufei Ma 0002, Le Ye |
ISLPED | 8 |
| 2023 | Research progress on low-power artificial intelligence of things (AIoT) chip design
Le Ye, Zhixuan Wang, Yufei Ma 0002, Linxiao Shen, Yihan Zhang 0002, Meng Wu 0005, Ying Liu 0069, Yiqi Jing, Hao Zhang 0119, Ru Huang 0001 |
Sci. China Inf. Sci. | 1 |
| 2023 | An 82-nW 0.53-pJ/SOP Clock-Free Spiking Neural Network With 40-μs Latency for AIoT Wake-Up Functions Using a Multilevel-Event-Driven Bionic Architecture and Computing-in-Memory TechniqueabstractThis article presents a clock-free spiking neural network (SNN) intelligent inference engine (IIE) for artificial intelligence of things (AIoT) sensor nodes, which often operate in random-sparse-event (RSE) scenarios. The IIE drastically reduces the system’s long-term average (LTA) power consumption, improves energy efficiency, and achieves microsecond level inference latency. Three techniques are proposed: 1) A clock-free SNN architecture without clock tree, frame generator, and arbiter, is driven by the output spikes, which are encoded with level-crossing (LC) sampling method; the circuit activity is completely related to event activity and spike rates, dramatically reducing the overall power consumption and latency. 2) The bioinspired leaky-integrate-fire (LIF) neurons directly extract the time-domain information from asynchronous spikes, reducing the network size and number of operations. 3) The computing-in-memory (CIM) and mixed-signal synapse-neuron circuits are employed to increase the SNN parallelism and avoid weight movements, thus improving the energy efficiency and response speed. The measured LTA power is bounded at 82 nW while the event-driven chip is on call and waiting for events; the energy efficiency is 0.53 pJ per synapse operation (SOP), only 1/3 that of state-of-the-art methods at 4bit weights even with 180 nm technology. We demonstrate electrocardiogram (ECG) recognition as a typical AIoT application, and the power consumption is less than 350 nW. The measured accuracy of abnormal ECG detection is 90.5%. Moreover, the latency is only$40 \mu \text{s}$to realize real-time NN inference. This work provides an effective solution for AIoT nodes that require both ultralow power and fast response. Ying Liu 0069, Yufei Ma 0002, Zhixuan Wang, Linxiao Shen, Jiayoon Ru, Ru Huang 0001, Le Ye |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2022 | DCIM-GCN: Digital Computing-in-Memory to Efficiently Accelerate Graph Convolutional NetworksabstractComputing-in-memory (CIM) is emerging as a promising architecture to accelerate graph convolutional networks (GCNs) normally bounded by redundant and irregular memory transactions. Current analog based CIM requires frequent analog and digital conversions (AD/DA) that dominate the overall area and power consumption. Furthermore, the analog non-ideality degrades the accuracy and reliability of CIM. In this work, an SRAM based digital CIM system is proposed to accelerate memory intensive GCNs, namely DCIM-GCN, which covers innovations from CIM circuit level eliminating costly AD/DA converters to architecture level addressing irregularity and sparsity of graph data. DCIM-GCN achieves 2.07X, 1.76X, and 1.89× speedup and 29.98×, 1.29×, and 3.73× energy efficiency improvement on average over CIM based PIMGCN, TARe, and PIM-GCN, respectively. Yikan Qiu, Yufei Ma 0002, Meng Wu 0005, Le Ye, Ru Huang 0001 |
ICCAD | 5 |
| 2022 | A 32-ppm/°C 0.9-nW/kHz Relaxation Oscillator with Event-Driven Architecture and Charge Reuse TechniqueabstractThis paper presents a dual-phase RC-based relaxation oscillator (RxO) with low temperature coefficient (TC) and high power efficiency achieved simultaneously for energy-constrained Internet-of-Things (IoT) applications with burst-mode requirements. Its circuit-level event-driven architecture reduces the duty cycle of power-hungry blocks, saving power while posing little performance penalty. In addition, the charge reuse technique further reduces the power consumption for the always-on detecting circuit. Implemented in a 0.18-μm CMOS process, the 180-kHz relaxation oscillator exhibits a frequency deviation of ± 0.26% against temperature (-40 to 125 ° C) from Monte-Carlo simulation (N=30), leading to a low temperature coefficient of 32 ppm/° C. The simulated power consumption is 163 nW, resulting in power efficiency of 0.9 nW/kHz. Xinhang Xu, Siyuan Ye, Jihang Gao, Yihan Zhang 0002, Linxiao Shen, Le Ye |
ISCAS | 6 |
| 2022 | Reliability-Improved Read Circuit and Self-Terminating Write Circuit for STT-MRAM in 16 nm FinFETabstractHigh power consumption is usually required in a spin-torque-transfer magnetoresistive random access memory (STT-MRAM) array’s peripheral circuits for reliable operations. In read, power needs to be spent for the low absolute resistance in the magnetic tunnel junctions (MTJ), and a limited high-state-to-low-state resistance ratio calls for high currents for the same detectable readout voltage under accuracy requirements. In write, the random programming time poses challenges for energy efficient write operations within an acceptable write error rate. To address the issues mentioned above, in this work, we propose a reliability-improved read circuit that consumes only 92.09 fJ/bit read energy while ensuring correct readout values under 4.5 sigma resistance variance, and a self-terminating write peripheral circuit achieving an energy reduction of 82.3% at 1 part-per-million write error rate (WER) under 20 ns write period. Chang Xue, Yihan Zhang 0002, Mingwei Zhu, Tianqiao Wu, Meng Wu 0005, Yandong He, Le Ye |
ISCAS | 8 |
| 2021 | SWIFT: Small-World-based Structural Pruning to Accelerate DNN Inference on FPGAabstractState-of-the-art DNN pruning approaches achieved high sparsity. However, these methods usually do not consider the intrinsic graph property of DNNs, leading to an irregular pruned network. Consequently, hardware accelerators cannot directly benefit from such pruning, suffering additional cost on indexing, control and data paths. Inspired by the observation that the brain and real-world networks follow a Small-World model, we propose a graph-based progressive structural pruning technique, SWIFT, that integrates local clusters and global sparsity in DNNs to benefit the dataflow and workload balance of the accelerators. In particular, we propose an output stationary FPGA architecture to accelerate DNN inference and integrate it with the structural sparsity by SWIFT, so that the communication and computation of clustered zero weights are eliminated. In addition, a full mesh data router is designed to adaptively direct inputs into corresponding processing elements (PEs) for different layer configurations and skipping zero operations. The proposed SWIFT is evaluated with multiple DNNs on different datasets. It achieves sparsity ratio up to 76% for CIFAR-10, 83% for CIFAR-100, 76% for the SVHN datasets. Moreover, our proposed SWIFT FPGA accelerator achieves up to 4.4× improvement in throughput for different dense networks with a marginal hardware overhead. Yufei Ma 0002, Yu Cao 0001, Le Ye, Ru Huang 0001 |
FPGA | 4 |
| 2021 | Ultra-Low-Power and Performance-Improved Logic Circuit Using Hybrid TFET-MOSFET Standard Cells Topologies and Optimized Digital Front-End ProcessabstractTunnel FET is recognized as one of the most promising candidates for ultra-low power applications due to its ultra-low off current and CMOS compatibility. However, some characteristics of TFET caused by asymmetric device structure and special conduction mechanism may make conventional topologies of logic circuits no longer applicable. Our previous work has reported that TFET stacking will result in severe current degradation, which makes traditional logic cells not applicable. In this paper, two solutions are proposed: first, from a logic cell perspective, novel hybrid TFET-MOSFET topologies of standard logic cells are proposed, which achieve more than 2 times lower hardware cost and intrinsic delay, hence up to 4 times lower area-power-delay product (APDP) than that of conventional TFET logic circuits. Compared to MOSFET logic circuits, the designs achieve almost 2 orders of magnitude lower power and up to 34 times lower APDP. Second, from a large-scale circuit perspective, an optimized digital front-end (DFE) is proposed. Taking serial peripheral interface (SPI) as an example, SPI circuit using the optimized DFE achieves 46% lower delay and 4 times lower APDP than that of traditional TFET SPI, and 3 orders of magnitude lower static power and APDP than that of MOSFET SPI. Zhixuan Wang, Le Ye, Kaixuan Du, Zhichao Tan, Yangyuan Wang, Ru Huang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | Re-Assessment of Steep-Slope Device Design From a Circuit-Level Perspective Using Novel Evaluation Criteria and Model-Less MethodabstractPower is becoming a major bottleneck in energy constraint applications such as internet-of-things (IoT). Emerging steep-slope devices such as tunnel FETs (TFET) and negative capacitance (NC) FETs are promising candidates for such type of applications. Nevertheless, due to the time-consuming characterization process and inconsistent evaluation criteria, conventional co-design and co-optimization process between novel devices and logic circuits takes too much time and its results rarely meet expectation. As a result, conventional co-design and co-optimization are quite inefficient. In this paper, for the first time, a new criterion is utilized to evaluate novel steep-slope devices for ultra-low power applications. In addition, an efficient evaluation method is proposed, which not only quantitatively guides device design, but also evaluates devices from a circuit perspective without the need for device compact model and circuit simulation. From a device design perspective, optimal design metrics of novel steep slope devices such as average subthreshold slope (SSavg), off current (IOFF), and on current (ION) can be directly figured out with the help of the proposed evaluation criteria and method. From a circuit design perspective, the proposed evaluation criteria and method can be used to determine application scope. Zhixuan Wang, Le Ye, Yangyuan Wang, Ru Huang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | The Challenges and Emerging Technologies for Low-Power Artificial Intelligence IoT SystemsabstractThe Internet of Things (IoT) is an interface with the physical world that usually operates in random-sparse-event (RSE) scenarios. This article discusses main challenges of IoT chips: power consumption, power supply, artificial intelligence (AI), small-signal acquisition, and evaluation criteria. To overcome these challenges, many works recently aimed at IoT system design have emerged. This work reviews the architecture and circuit innovations that have contributed to IoT developments. This paper does not cover security of IoT. Event-driven architectures and nonuniform sampling ADCs significantly reduce the long-term average power. Besides, embedding AI engines in IoT nodes (AIoT) is one critical trend. The computing-in-memory technique improves the energy efficiency of the AI engine. Asynchronous spike neural networks (ASNNs) AI engines show low power potential. In addition to data processing, small-signal acquisition is also critical. The charge-domain analog-front-end (AFE) techniques such as floating inverter-based amplifiers improve energy efficiency. In addition to the above low power and high energy efficiency technologies, energy harvesting can also enhance the lifetime of AIoT devices. This article discusses recent ambient RF and natural energy harvesting approaches and high-efficiency DC-DC with a wide load range. Finally, novel evaluation criteria are introduced to establish benchmark standards for AIoT chips. Le Ye, Zhixuan Wang, Ying Liu 0069, Hao Zhang 0119, Meng Wu 0005, Linxiao Shen, Yihan Zhang 0002, Zhichao Tan, Yangyuan Wang, Ru Huang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2020 | A 1μW-to-158μW Output Power Pseudo Open-Loop Boost DC-DC with 86.7% Peak Efficiency using Frequency-Programmable Oscillator and Hybrid Zero Current DetectionabstractThis paper proposed a pseudo-open boost DC-DC converter whose input voltage ranges from 300mV-to-500mV and output voltage ranges from 1.2V-to-1.8V. The output power ranges from 1μW to 158μW. Three key structures are designed to realize the high energy conversion efficiency when the load reaches ultra-light: the pseudo open-loop structure, the frequency-programmable oscillator, and the hybrid zero current detection (H-ZCD) circuit. The pseudo open-loop structure eliminates the analog comparators or error amplifiers used in converter and the quiescent power loss is much reduced. The frequency-programmable ring oscillator is able to output 9-bit binary frequency ranges from 500Hz-100KHz, so the power consumption of clock generation is saved when load power reaches to ultra-light. The proposed H-ZCD is designed to reduce the inductor power loss with just tiny power consumption. Accordingly, the high efficiency boost DC-DC works in the discontinuous condition mode (DCM) and the constant on time pulse frequency modulation (PFM) is proposed. As a simulation result, the peak efficiency of the converter reaches to 86.7% while the load power is 13.7μW. Enbin Gong, Hao Zhang 0119, Le Ye, Ru Huang 0001 |
ISCAS | 4 |
| 2020 | 2.4-GHz 16-QAM Passive Backscatter Transmitter for Wireless Self-Power Chips in IoTabstractA 2.4-GHz 16-QAM ultra-low-power passive transmitter for wireless self-power chips in IoT is proposed. To expand wireless sensor networks, it achieves energy harvest and low-power wireless communication with 2.4-GHz infrastructure. Besides, the backscatter technique is employed to reduce power consumption and is compatible with multiple quadrature amplitude modulation (M-QAM) to increase data rate. Measured results show that the transmitter just dissipates 1 μW from 1.8-V supply voltage with 6.46% EVM at 4-Mb/s data rate. Moreover, the transmitter has great availability whether at 100 Mb/s data rate or for the input power of a wide dynamic range. The chip is fabricated in the 0.18-μm 1P6M standard CMOS process and occupies a silicon area of 1465 × 915 μm2without pads. Enbin Gong, Hao Zhang 0119, Le Ye, Ru Huang 0001 |
ISCAS | 4 |
| 2019 | Ultra-Low Power Hybrid TFET-MOSFET Topologies for Standard Logic Cells with Improved Comprehensive PerformanceabstractTunnel FET (TFET) is recognized to be one of the most promising candidates for ultra-low power applications due to its ultra-low off current and high compatibility with CMOS process. However, different from the typical features of MOSFET, some electrical characteristics of TFETs caused by asymmetric device structure and special conduction mechanism may make conventional topologies of circuits no longer applicable. In this paper, it is found that the TFETs stacking will result in severe current degradation behavior, which makes traditional topologies of logic gates may be not applicable. To solve this problem, a set of novel hybrid TFET-MOSFET topologies for standard logic cells are proposed. The proposed designs achieve more than 2 times lower hardware cost and intrinsic delay, and realize up to 4 times lower area-power-delay product (APDP) than that of conventional TFET-based logic circuits. Moreover, the proposed topologies can achieve almost 2 orders of magnitude lower power and up to 34 times lower APDP than that of conventional MOSFET-based logic circuits. The proposed standard logic cells show great superiority for power-constraint applications. Zhixuan Wang, Le Ye, Libo Yang, Yangyuan Wang, Ru Huang 0001 |
ISCAS | 4 |
| 2018 | Combinational Access Tunnel FET SRAM for Ultra-Low Power ApplicationsabstractIn this paper, a novel combinational access topology of Tunnel FET (TFET) SRAM is proposed for ultra-Low Power applications. Since forward p-i-n current of TFET could cause serious damage to SRAM circuit performance, the proposed topology can avoid the forward bias applied to the p-i-n junction, thus increasing SRAM cell read and hold static noise margin (SNM) and decreasing its static power consumption dramatically. At 0.6 V supply voltage, the combinational access TFET SRAM topology presents 26% hold SNM larger than traditional TFET SRAM topologies, 8 orders of magnitude lower static power consumption, and 2 order of magnitude lower power delay product, demonstrating its great potential for ultra-low power applications. Libo Yang, Jiadi Zhu, Zhixuan Wang, Zexue Liu, Le Ye, Ru Huang 0001 |
ISCAS | 7 |
| 2017 | Benchmarking TFET from a circuit level perspective: Applications and guidelineabstractLow power applications have led to a boom in researches on new circuits based on steep-slope transistors, of which the objective is to overcome MOSFET's drawback of inevitable increasing leakage power while maintaining acceptable performance in low voltage operation. Among those emerging transistors, Tunnel FET (TFET) becomes a most promising one due to its low off current and compatibility with CMOS process. In order to guide the application and the improvement of TFET, in this paper from a circuit-level perspective, utilizing a newly defined benchmarking method, we figured out the frequency-VDD range in which Si TFET circuits show low power advantage over their MOSFET counterparts based on HSPICE simulations using calibrated compact model. A systematic and quantitative analysis was then conducted to further enlarge the application scope of TFET circuits, with a Figure of Merit (FOM) and a guideline for future TFET proposed. Lingyi Guo, Le Ye, Libo Yang, Zhu Lv, Xia An, Ru Huang 0001 |
ISCAS | 2 |
| 2014 | A UHF RFID reader transmitter with digital CMOS power amplifierabstractThis paper presents a low TX noise transmitter with digital modulated power amplifier (DPA) for UHF RFID reader System-on-Chip. With I/Q digital signals directly modulated on DPA, the transmitter eliminates the pulse shaping filter, up-conversion mixer and power amplifier driver, which are dominant noise contributors in a conventional RFID transmitter. A switch-mode power amplifier cell is adopted in the DPA, which suppresses the amplitude noise in the TX chain. An LO signal buffer strategy is proposed to minimize uncorrelated phase noise which cannot be cancelled by self-cancellation in the RFID receiver. A low noise self-adaptive LDO is designed for lowering TX noise and stable output power. The transmitter achieves a TX noise density lower than -145dBc/Hz and provides a peak output power of 24 dBm with 48% drain efficiency when transmitting CW signal. Through pre-distortion, the transmitter can achieve a spectrum mask with -50 dBc for ACPR1, -62 dBc for ACPR2 and -65dBc for ACPR3. The proposed transmitter is fabricated in a standard 0.18-μm CMOS process with an area of 1000×640μm2. Long Chen 0009, Le Ye, Xing Zhang 0002, Huailin Liao |
ISCAS | 4 |
| 2013 | SAW-less GNSS front-end amplifier with 80.4-dB GSM blocker suppression using CMOS directional coupler notch filterabstractThis paper presents a SAW-less GNSS front-end amplifier with GSM blocker suppression using CMOS directional coupler notch filter. The front-end amplifier is aimed at the GNSS receiver integrated in cellular phones. Based on our proposed CMOS stacked spiral-coupled (SSC) directional coupler working at the frequency of 900MHz as notch filter, the front end amplifier achieves a NF of 1.7dB and a 80.4-dB suppression of the GSM blocker while provides signal gain of 38.6-dB for the GPS L1-band signal. Yongan Zheng, Le Ye, Long Chen 0009, Huailin Liao, Ru Huang 0001 |
ISCAS | 2 |
| 2013 | A 65 mW fully integrated UHF-band CMMB tuner in 65 nm CMOS process
Junhua Liu 0001, Chen Li 0014, Long Chen 0009, Congyin Shi, Xuankai Weng, Yixiao Wang 0001, Yu Liao, Le Ye, Huailin Liao, Ru Huang 0001 |
Sci. China Inf. Sci. | 9 |
| 2012 | A +21.2 dBm out-of-band IIP3 0.2-3GHz RF front-end using impedance translation techniqueabstractThis paper presents a SAW-less 0.2-3GHz front-end with high out-of-band linearity. Modified impedance translation technique based on N-path current driven mixer is used to improve out-of-band IIP3. At the input node, an 8-path passive mixer switched by 8-phase clocks at the frequency of fLO/2 is utilized to achieve impedance match at fLOand filter out-of-band interferes, while contribute negligible noise. Using a 4-path mixer as the LNA load, switched by 4-phase clocks at the frequency of fLO, out-of-band interferes is further attenuated. Implemented in 65nm CMOS process, the proposed front-end aiming at reconfigurable receivers achieves a NF of 3-5dB from 0.2GHz to 3GHz, out-of-band IIP3 of 21.2dBm, and maximum gain of 45dB, respectively. The front-end consumes 13mA current at 0.2GHz and 27.5mA at 3GHz from a 1.2V voltage supply. Long Chen 0009, Chen Li 0014, Le Ye, Huailin Liao, Ru Huang 0001 |
ISCAS | 4 |
| 2012 | Cost-efficient CMOS RF tunable bandpass filter with active inductor-less biquadsabstractThis paper presents a CMOS RF tunable 4th-order active bandpass filter with the proposed inductor-less biquads. The NMOS cross-coupled pair is utilized in the biquad for the positive-feedback to form the complex pole, which enables the filter working at high frequency of 5GHz with low power of only 4.8mW from 1.2V power supply. Due to the inductor-less topology, the proposed filter only occupies 0.011mm2silicon area, which is cost-efficient and suitable for integration on chip. The center frequency can be tuned from 2GHz to 5GHz, and the Q factor is tuned from 2 to 8 to cover different bandwidth from 250MHz to 2.5GHz, which makes it suitable for the multi-band/multi-mode and SDR applications. The filer is demonstrated in a standard 65nm CMOS process. As for the center frequency of 5GHz and Q of 2, the simulated P1dB is -6.7dBm, and the simulated input referred noise (IRN) density is 14.2nV/sqrt(Hz). Yixiao Wang 0001, Le Ye, Huailin Liao, Ru Huang 0001 |
ISCAS | 2 |
| 2012 | Widely reconfigurable 8th-order chebyshev analog baseband IC with proposed push-pull op-amp for Software-Defined Radio in 65nm CMOSabstractThis paper presents an 8th-order chebyshev active-RC analog baseband IC with tunable cut-off frequency from 500K to 16MHz and adjustable gain from 5.5dB to 70dB for a Software-Defined Radio (SDR) receiver. For the analog baseband, a highly power-efficient push-pull op-amp with two differential-to-single output stages is proposed, which is suitable for the advanced deep-submicron CMOS process. It achieves 45dB gain and 850MHz GBW with only 0.8mA current. I/Q analog baseband IC, consisting of filter, PGA/VGA, and DCOC, is integrated for a SDR receiver, which is fabricated in a standard 65nm CMOS technology. It consumes 9.2mA current from 1.2V power supply, achieves 17.44dBm in-band OIP3, 11.43nV/√Hz input-referred noise (IRN) density, and occupies 0.68mm2silicon area. Le Ye, Yixiao Wang 0001, Long Chen 0009, Huailin Liao, Ru Huang 0001 |
ISCAS | 1 |
| 2011 | -99dBc/Hz@10kHz 1MHz-step dual-loop integer-N PLL with anti-mislocking frequency calibration for global navigation satellite system receiverabstractIn this paper, a low in-band phase noise integer-N CMOS frequency synthesizer is proposed for global navigation satellite system (GNSS) receiver. The synthesizer adopts dual-loop architecture, which consists of a double-balanced mixer and two full PLL loops, to reduce the divide ratio so as to lower the in-band phase noise. It achieves 1MHz resolution and -99 dBc/Hz@10kHz with fixed reference clock of 10MHz, which is compatible to commercial atomic frequency sources. Moreover, a novel adaptive frequency calibration policy is implemented to avoid mis-locking at the unwanted mirror frequency. The PLL is fabricated in 0.18-μm CMOS technology, covers most GPS, Galileo and Beidou-II bands and was integrated in a GNSS receiver with 46MHz intermediate frequency (IF). Congyin Shi, Le Ye, Huailin Liao |
ISCAS | 3 |
| 2011 | A 0.47mW 6th-order 20MHz active filter using highly power-efficient OpampabstractThis paper presents an ultra-low power 6th-order 7MHz-to-20MHz tunable active-RC low-pass filter. Due to the proposed highly power-efficient Opamp, the filter only consumes 0.47 mA power from 1.8 V supply voltage, corresponding to 3.86 pW/Hz/pole normalized power. The Opamp utilizes an adaptive-biased pole-cancellation push-pull buffer to greatly reduce the power consumption. An adaptive bias circuit is proposed to cooperate with the Opamp to tolerate the PVT variations. The filter achieves 20.9 dBm in-band IIP3, and 298 μVrms integrated input-referred noise. The chip is fabricated in a standard 0.18 μm CMOS process, and occupies 0.21 mm2silicon area without ESD/pads. Le Ye, Congyin Shi, Huailin Liao, Ru Huang 0001 |
ISCAS | 1 |