EDBT 2026 Demo / reviewers in the wild / expert
Jun Yang 0006
dblp:y/JunYang6
· DBLP profile ↗
61ranked-venue papers
0as first author
38since 2021 · last 2026
0000-0002-8379-0321ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 43 · 27 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 8 since 2021Software engineering, systems software and programming languages · 5 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 since 2021Computer networks · 4 · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FIXME: Towards End-to-End Benchmarking of LLM-Aided Design VerificationabstractDespite the transformative potential of Large Language Models (LLMs) in hardware design, a comprehensive evaluation of their capabilities in design verification remains underexplored. Current efforts predominantly focus on RTL generation and basic debugging, overlooking the critical domain of functional verification, which is the primary bottleneck in modern design methodologies due to the rapid escalation of hardware complexity. We present FIXME, the first end-to-end, multi-model, and open-source evaluation framework for assessing LLM performance in hardware functional verification (FV) to address this crucial gap. FIXME introduces a structured three-level difficulty hierarchy spanning six verification sub-domains and 180 diverse tasks, enabling in-depth analysis across the design lifecycle. Leveraging a collaborative AI-human approach, we construct a high-quality dataset using 100% silicon-proven designs, ensuring comprehensive coverage of real-world challenges. Furthermore, we enhance the functional coverage by 45.57% through expert-guided optimization. By rigorously evaluating state-of-the-art LLMs such as GPT-4, Claude3, and LlaMA3, we identify key areas for improvement and outline promising research directions to unlock the full potential of LLM-driven automation in hardware design verification. The benchmark is available at https://github.com/ChatDesignVerification/FIXME. Gwok-Waa Wan, Sam-Zaak Wong, Shengchu Su, Chenxu Niu 0001, Ning Wang 0071, Xinlai Wan, Qixiang Chen, Mengnv Xing, Jianmin Ye, Rongchang Song, Qiang Xu 0001, Nan Guan, Zhe Jiang 0004, Xi Wang 0009, Yong Chen 0001, Jun Yang 0006 |
AAAI | 19 |
| 2026 | ChipMind: Retrieval-Augmented Reasoning for Long-Context Circuit Design SpecificationsabstractWhile Large Language Models (LLMs) demonstrate immense potential for automating integrated circuit (IC) development, their practical deployment is fundamentally limited by restricted context windows. Existing context-extension methods struggle to achieve effective semantic modeling and thorough multi-hop reasoning over extensive, intricate circuit specifications. To address this, we introduce ChipMind, a novel knowledge graph-augmented reasoning framework specifically designed for lengthy IC specifications. ChipMind first transforms circuit specifications into a domain-specific knowledge graph (ChipKG) through the Circuit Semantic-Aware Knowledge Graph Construction methodology. It then leverages the ChipKG-Augmented Reasoning mechanism, combining information-theoretic adaptive retrieval to dynamically trace logical dependencies with intent-aware semantic filtering to prune irrelevant noise, effectively balancing retrieval completeness and precision. Evaluated on an industrial-scale specification reasoning benchmark, ChipMind significantly outperforms state-of-the-art baselines, achieving an average improvement of 34.59% (up to 72.73%). Our framework bridges a critical gap between academic research and practical industrial deployment of LLM-aided Hardware Design (LAD). Changwen Xing, Sam-Zaak Wong, Xinlai Wan, Mengli Zhang, Zebin Ma, Lei Qi 0001, Zhengxiong Li, Nan Guan, Zhe Jiang 0004, Xi Wang 0009, Jun Yang 0006 |
AAAI | 12 |
| 2026 | CIM-Tuner: Balancing the Compute and Storage Capacity of SRAM-CIM Accelerator via Hardware-mapping Co-explorationabstractAs an emerging type of AI computing accelerator, SRAM Computing-In-Memory (CIM) accelerators feature high energy efficiency and throughput. However, various CIM designs and under-explored mapping strategies impede the full exploration of compute and storage balancing in SRAM-CIM accelerator, potentially leading to significant performance degradation. To address this issue, we propose CIM-Tuner, an automatic tool for hardware balancing and optimal mapping strategy under area constraint via hardware-mapping co-exploration. It ensures universality across various CIM designs through a matrix abstraction of CIM macros and a generalized accelerator template. For efficient mapping with different hardware configurations, it employs fine-grained two-level strategies comprising accelerator-level scheduling and macro-level tiling. Compared to prior CIM mapping, CIM-Tuner’s extended strategy space achieves 1.58× higher energy efficiency and 2.11× higher throughput. Applied to SOTA CIM accelerators with identical area budget, CIM-Tuner also delivers comparable improvements. The simulation accuracy is silicon-verified and CIM-Tuner tool is open-sourced at https://github.com/champloo2878/CIM-Tuner.git. Jinwu Chen, He Wang 0028, Zhe Jiang 0004, Jun Yang 0006, Xin Si, Zhenhua Zhu 0002 |
DATE | 5 |
| 2026 | ChatTest: Coverage-Enhanced Testbench Generation for Agile Hardware Verification with LLMsabstractThe growing complexity of modern hardware designs has rendered traditional functional verification increasingly time-consuming, with verification costs now dominating the design cycle. While large language models (LLMs) show promise in automating testbench generation, existing approaches struggle with real-world scalability, suffering from poor comprehension of long specifications and complex designs. To address these challenges, we propose ChatTest, a novel, end-to-end, multi-agent LLM framework for coverage-aware, agile hardware verification. Our key innovation lies in a function-mapped, divide-and-conquer architecture that integrates a Verification Description Language (VDL)—a structured, LLM-friendly DSL for precise specification encoding—with Constraint-Aware Segmental Adaptation (CASA) to enable coherent processing of long, heterogeneous design documents. By leveraging retrieval-augmented generation and supervised fine-tuning using multi-hierarchical specification-code alignment, ChatTest ensures accurate translation of functional points into targeted test stimuli. Furthermore, we introduce a coverage-driven feedback loop for automated test augmentation. Evaluated on a new benchmark of 20 complex RTL designs (up to 31K tokens of specification and 4K line-of-code), ChatTest achieves 1.46× higher toggle coverage and 2.28× higher line coverage than SOTA, with a 24.23% improvement in functional coverage, demonstrating its effectiveness in accelerating verification convergence. Gwok-Waa Wan, Shengchu Su, Sam-Zaak Wong, Mengnv Xing, Zhe Jiang 0004, Xi Wang 0009, Jun Yang 0006 |
DATE | 9 |
| 2026 | StatCHAR: Statistical Timing Characterization Framework via Heterogeneous Graph Attention Network and Active Learning With Parasitic RC ReductionabstractStatistical timing characterization for standard cell library poses significant challenges to accuracy and runtime cost. Prior analytical and learning-based methods neglect the profound influence induced by the layout-dependent parasitic resistor and capacitor (RC) network in cell netlist as well as the timing correlation between the topological structures of cells and process, voltage, and temperature (PVT) corners for model training, resulting in tremendous simulation effort and poor accuracy. In this work, a Statistical timing Characterization framework via Heterogeneous graph attention network and Active learning with parasitic RC Reduction (StatCHAR) is proposed, where the transistors and parasitic RC in cell are represented as heterogeneous nodes for graph learning and redundant RC nodes are removed to alleviate node imbalance issue and improve accuracy. The significant training data are selected from the full characterization set with active learning strategy to achieve the optimal balance between simulation overhead for the training set and prediction precision for the remaining test set. The proposed framework was validated with typical standard cells under multiple PVT corners with TSMC 22nm process, which achieves an excellent prediction with a relative Root Mean Square Error (rRMSE) of only 2.43% with only 11.9% of total characterization data for training, demonstrating an accuracy improvement of 2.7$\times \sim 12.1\times $for statical timing analysis on benchmark circuits compared to competitive learning based methods and a characterization runtime reduction by 7.5$\times $. Peng Cao 0002, Zeyuan Deng, Yuhan Dong, Yuyang Ye 0001, Jun Yang 0006 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2026 | A Survey of the First TinyML@ICCAD Contest for Ventricular Arrhythmia Detection by Artificial Intelligence on Low-power MicroprocessorabstractArtificial intelligence has achieved remarkable success in various real-world applications. However, the challenge lies in its implementation on hardware platforms with constrained resources and low power while maintaining real-time capabilities. Edge artificial intelligence, in particular, stands as a pivotal field for the practical deployment of AI. The 41st IEEE/ACM International Conference on Computer-Aided Design introduced the inaugural TinyML Design Contest in 2022. The contest entailed a rigorous, multi-month research and development competition, focusing on the creation of real-time detection algorithms for life-threatening ventricular arrhythmia. These algorithms were required to be deployable on the low-power microprocessor NUCLEO-L432KC. Open to multi-person teams worldwide, the contest garnered 150 teams participation teams from 50+ organizations, with 41 teams successfully completing the challenge. Our SEUer team secured the second place. This article provides a detailed exposition of the contest, offering insights into its structure and objectives. Furthermore, it analyzes and discusses the methods developed by some of the entries as well as representative results. Finally, the article concludes with directions for future improvements. Meng Zhang 0010, Tinghuan Chen, Jun Yang 0006 |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2025 | Truly Pre-Routing Timing Prediction via Considering Power Delivery NetworkabstractFast and accurate pre-routing timing prediction is essential in the chip design flow. However, existing machine learning (ML)assisted pre-routing timing methods often overlook the impact of power delivery networks (PDNs), which contribute to IR drop and routing congestion. This limitation can make these methods less practical for realworld circuit design flows. To address this, we propose two specialized encoders-an IR drop-aware encoder and a routing congestion-aware encoder-that effectively capture PDN effects through multimodal fusion of netlist, layout, and PDN data. To mitigate the challenges of imbalanced multimodal fusion, we further develop a Pareto optimization approach to ensure balanced utilization of all modalities, enhancing timing prediction accuracy. Comprehensive experiments on large-scale open-source designs using TSMC’s 16 nm technology node validate the superiority of our model over state-of-the-art pre-routing timing prediction methods. Yuyang Ye 0001, Mingwei He, Lizheng Ren, Jianwang Zhai, Tinghuan Chen, Jun Yang 0006, Longxing Shi |
DAC | 6 |
| 2025 | LAD: Efficient Accelerator for Generative Inference of LLM with Locality Aware DecodingabstractLarge Language Models (LLMs) have emerged as the cornerstone of content generation applications due to their ability to capture relations between newly generated token and the full preceding context. However, this ability stems from the attention mechanism for decoding that retains the entire generation history as key value cache (KV cache). As the generated sequence lengthens, the KV cache expands, causing a substantial memory access bottleneck. In advanced LLM generation systems running on GPUs, the attention mechanism for decoding accounts for more than 50% of the total inference time when the KV cache length reaches 4096. To address this issue, this paper introduces LAD (Locality Aware Decoding), an LLM generation accelerator with algorithm-hardware enhancements that significantly decrease KV cache access, resulting in considerable speedups and energy savings. A key insight underlying LAD is that when the attention score for a specific position remains fixed over the next several decoding steps, it is unnecessary to repeatedly retrieve the associated key and value at each step to reproduce the computation. Our analysis reveals that numerous positions exhibit notable numerical locality in attention scores through multiple decoding steps. Leveraging these insights, we have designed an innovative attention decoding computation method that decreases the frequency of accessing the key and value for positions demonstrating good locality, all while maintaining decoding accuracy. Extensive experiments show that LAD generates sequences with an average ROUGE-1 similarity of 97% compared to those generated by the original model. When the length of KV cache exceeds 2048, the high configuration of LAD accelerator achieves on average (geomean) $10.7 \times$ speedup and $52.4 \times$ energy efficiency for the attention mechanism compared to the A100 GPU. For end-to-end model inference, it also achieves on average $2.3 \times$ speedup and $13.4 \times$ energy efficiency. Haoran Wang 0012, Ying Wang 0001, Liqi Liu, Jun Yang 0006, Yinhe Han 0001 |
HPCA | 6 |
| 2025 | DiSPlace: Diffusion-Sharing-Driven Transistor-Level Placement Beyond Standard-Cell Boundaries for DTCOabstractAs the increasing demands of design technology co-optimization (DTCO) in advanced nodes, the rigid configurations of standard cells impose significant limitations on wirelength and area optimization. A more flexible alternative is to place transistors directly on the design canvas, allowing for precise transistor-level adjustments that reduce wirelength and minimize design area. In this paper, we propose DiSPlace, a novel diffusion-sharing-driven transistor-level placement algorithm beyond standard-cell boundaries to fully leverage DTCO. We first present an in-cell placement based transistor pairing method to pair PMOS and NMOS transistors with the same gate net, followed by incorporating Gaussian perturbations to generate an initial placement. Then, we propose the first diffusion-sharing-driven global placement framework. It begins with the construction of diffusion sharing nets to guide transistor placement, followed by an analytical model for simultaneously optimizing diffusion sharing, wirelength, and density. Besides, a nonlinear optimization with adaptive penalty adjustment is presented to solve the analytical model effectively and efficiently. Finally, we develop a satisfiability modulo theories (SMT)-based detailed placement method to optimize design area and wirelength while ensuring legal placement. A diffusion-sharing-aware partitioning technique is also developed to enhance the scalability and efficiency of the SMT-based method. Compared to a standard-cell-based placer and the state-of-the-art transistor-level placer, our algorithm achieves significant improvements, reducing wirelength by 18% and 11%, and design area by 24% and 4%, respectively. These results highlight the effectiveness of DiSPlace in achieving high-quality placements for transistor-level designs. Keyu Peng, Yinuo Wu, Zhengzhe Zheng, Ziran Zhu, Chao Wang 0068, Jun Yang 0006 |
ICCAD | 7 |
| 2025 | MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMsabstractThe emergence of multimodal large language models (MLLMs) presents promising opportunities for automation and enhancement in Electronic Design Automation (EDA). However, comprehensively evaluating these models in circuit design remains challenging due to the narrow scope of existing benchmarks. To bridge this gap, we introduce MMCircuitEval, the first multimodal benchmark specifically designed to assess MLLM performance comprehensively across diverse EDA tasks. MMCircuitEval comprises 3614 meticulously curated question-answer (QA) pairs spanning digital and analog circuits across critical EDA stages—ranging from general knowledge and specifications to front-end and back-end design. Derived from textbooks, technical question banks, datasheets, and real-world documentation, each QA pair undergoes rigorous expert review for accuracy and relevance. Our benchmark uniquely categorizes questions by design stage, circuit type, tested abilities (knowledge, comprehension, reasoning, computation), and difficulty level, enabling detailed analysis of model capabilities and limitations. Extensive evaluations reveal significant performance gaps among existing LLMs, particularly in back-end design and complex computations, highlighting the critical need for targeted training datasets and modeling approaches. MMCircuitEval provides a foundational resource for advancing MLLMs in EDA, facilitating their integration into real-world circuit design workflows. Our benchmark is available at https://github.com/cure-lab/MMCircuitEval. Chenchen Zhao 0001, Zhengyuan Shi, Xiangyu Wen 0001, Yi Liu 0081, Yunhao Zhou, Hefei Feng, Yinan Zhu, Gwok-Waa Wan, Yongqi Fu, Chujie Chen, Chenhao Xue, Ying Wang 0001, Yibo Lin, Jun Yang 0006, Ning Xu 0009, Xi Wang 0009, Qiang Xu 0001 |
ICCAD | 18 |
| 2025 | A 22-nm 64-kB lightning-like hybrid computing-in-memory macro with a compressed adder tree and analog-storage quantizers for transformer and CNNs
An Guo 0001, Xi Chen 0107, Fangyuan Dong, Jinwu Chen, Zhihang Yuan, Xing Hu 0010, Guangyu Sun 0003, Arindam Basu, Jun Yang 0006, Xin Si |
Sci. China Inf. Sci. | 10 |
| 2025 | Expansion of the memory pyramid in the era of large models: compute-intensive compute-in-memory and memory-intensive compute-in-memory
Zhican Zhang, Yi Yang 0001, Zhaoyang Zhang 0008, Jinwu Chen, Xin Si, Jun Yang 0006 |
Sci. China Inf. Sci. | 10 |
| 2025 | Routability-Driven Macro Placement Engine for Modern FPGAs With Complex Cascade Shape and Region ConstraintsabstractField-programmable gate array (FPGA) macro placement holds a crucial role within the FPGA physical design flow since it substantially influences the subsequent stages of cell placement and routing. With the increasing number of macros and the complex cascade shape and region constraints imposed by modern FPGAs, the routability and macro placement have become much more challenging. In this paper, we propose an effective and efficient routability-driven macro placement algorithm for modern FPGAs with cascade shape and region constraints. To reserve adequate space for cell placement and guarantee routability, we first develop a routability-driven mixed-size analytical global placement that evenly distributes both macros and cells while considering cascade shape and region constraints. Particularly, the proposed global placement engine integrates a well-trained congestion prediction model, targeting benchmarks with high routing congestion to enhance overall routability. Then, we propose an integer linear programming (ILP)-based cascade shape legalization followed by matching-based macro legalization to remove macro overlaps while satisfying the region constraints. Finally, a routability-driven detailed macro placement is proposed to refine the solution. Compared with the winners of the MLCAD 2023 FPGA macro placement contest and state-of-the-art works, experimental results show that our algorithm achieves the best overall score and routability. Keyu Peng, Jianli Chen, Jun Yang 0006, Ziran Zhu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | Dual Multimodal Fusions With Convolution and Transformer Layers for VLSI Congestion PredictionabstractIn very large scale integration (VLSI) circuit physical design, precise congestion prediction during placement is crucial for enhancing routability and accelerating design processes. Existing congestion prediction models often encounter challenges in handling multimodal information and lack effective fusion of placement and netlist features, limiting their prediction accuracy. In this article, we present a novel congestion prediction model that leverages dual multimodal fusions with convolution and transformer layers to effectively capture the multiscale placement information and enhance congestion prediction accuracy. We first adopt convolutional neural networks (CNNs) to extract grid-based placement features and heterogeneous graph convolutional networks (HGCNs) to extract netlist information. To help the model understand the correlation between different modalities, we then propose an early feature fusion (EFF) to integrate netlist knowledge into multiscale placement features at multimodal interaction subspace. Besides, a deep feature fusion (DFF) method is proposed to further fuse multimodal features, which has multiple vision transformer layers based on adaptive attention enhancement technology. These layers include self-attention (SA) to boost intramodal features and cross-attention (CA) to perform cross-modal feature fusion on netlist and grid-based placement features. Finally, the output features of DFF are sent into the cascaded decoder to recover the congestion map by exploiting several upsampling layers and merging with EFF features. Compared with the existing state-of-the-art congestion prediction models, experimental results demonstrate that our model not only outperforms them in prediction accuracy, but also excels in reducing routing congestion when integrated into the placer DREAMPlace. Youwen Wang, Xinglin Zheng, Keyu Peng, Ziran Zhu, Jianli Chen, Jun Yang 0006 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2025 | A Path Statistical Delay Prediction Framework Based on Global Graph Neural NetworkabstractAs transistor feature size decreases, timing analysis under different PVTs has become a critical challenge. Statistical Static Timing Analysis (SSTA) can achieve higher accuracy than STA. However, SSTA needs to make a lot of repetitive calculations, which is time-consuming. Therefore, we propose a post-routing statistical delay prediction framework, Global Graph Neural Network (Global-GNN). In order to overcome the problem of insufficient information in the cell embedding generated by the limitation of receptive field in traditional GNN, which leads to the poor performance of the prediction model. We optimize the updating cell embedding method so that the cell embedding contains not only local information but also global information. The cell embedding and PVT information is used to predict the path mean delay mean and path squared deviation for the current operating condition, and thus to calculate the post-routing path statistical delay. Experiments are done based on the ISCAS’89, OpenCores, and ITC’99 benchmarks of the TSMC 28nm technology. Under different PVTs the value of$R^{2}$ranges from 0.849 to 0.999, and the speed is at least$215 \times $higher than the traditional post-routing path statistical delay calculation method. At the same time, compared with the competing model, the$R^{2}$of Global-GNN can reach 0.999, and the minimum running speed can reach 0.282 s, which is obviously better than the competing model. Xuejie Ning, Chenfei Hua, Jun Yang 0006, Zhikuang Cai |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | A Hybrid Domain and Pipelined Analog Computing Chain for MVM ComputationabstractIn this article, a stream-architecture and pipelined hybrid computing chain is presented to process matrix-vector multiplication (MVM). In each stage of the computing chain, a primary multiply-accumulate (MAC) stage consisting of charge, time, and digital domain processing units makes signed or unsigned$8\times 1\times 8$bit MAC operations and MSB quantization. Based on the stream architecture, the length of the computing chain can be configured to fit different MVM applications. In the charge-domain MAC unit, a double-plate sampling and weighted capacitor array with writing yield and efficiency enhanced 7T bitcell and three-step weighting scheme is implemented. To utilize the speed and resolution advantages of time-domain computing, a high linearity voltage-to-time converter (VTC) followed by a dynamic tristate delay chain is proposed to transfer and store MAC values from the charge domain in the time domain. To realize fast analog readout, a folding type and distributed time-to-digital converter (TDC) is proposed. To fully eliminate the offset and variation in the distributed TDC, a specific residue readout timing and back-end calibration scheme are applied. In the digital domain, a double-input and double-clock dynamic D flip-flop is built to realize partial sum transmission and accumulation in a single cycle with low energy and area consumption. Post-simulation results show that this computing chain can achieve 20.89–40.72-TOPS/W energy efficiency and 4.498-TOPS/mm2 throughput. Tianzhu Xiong, Yuyang Ye 0001, Xin Si, Jun Yang 0006 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2024 | Variational Label-Correlation Enhancement for Congestion PredictionabstractAs the complexity of Integrated Circuits (ICs) rises, accurate routing and congestion prediction, crucial for identifying early design flaws, become essential to expedite circuit design and conserve resources in the lengthy physical design process. Despite the advancements in current congestion prediction methodologies, an essential aspect that has been largely overlooked is the spatial label-correlation between different grids in congestion prediction. The spatial label-correlation is a fundamental characteristic of circuit design, where the congestion status of a grid is not isolated but inherently influenced by the conditions of its neighboring grids. In order to fully exploit the inherent spatial label-correlation between neighboring grids, we propose a novel approach, VALCE, i.e., VAriational Label-Correlation Enhancement for Congestion Prediction, which considers the local label-correlation in the congestion map, associating the estimated congestion value of each grid with a local label-correlation weight influenced by its surrounding grids. VALCE leverages variational inference techniques to estimate this weight, thereby enhancing the regression model’s performance by incorporating spatial dependencies. Experiment results validate the superior effectiveness of VALCE on the public available ISPD2011 and DAC2012 benchmarks using the superblue circuit line. Congyu Qiao, Ning Xu 0009, Xin Geng 0001, Ziran Zhu, Jun Yang 0006 |
ASPDAC | 6 |
| 2024 | Late Breaking Results: Routability-Driven FPGA Macro Placement Considering Complex Cascade Shape and Region ConstraintsabstractField-programmable gate array (FPGA) macro placement holds a crucial role within the FPGA physical design flow since it substantially influences the subsequent stages of cell placement and routing. In this paper, we propose an effective and efficient routability-driven macro placement algorithm for modern FPGAs with cascade shape and region constraints. To reserve adequate space for cell placement and guarantee routability, we first develop a routability-driven mixed-size analytical global placement (GP) that evenly distributes both macros and cells while considering cascade shape and region constraints. Then, we propose an integer linear programming (ILP)-based cascade shape legalization (LG) followed by matching-based macro legalization to remove macro overlaps while satisfying the region constraints. Finally, a routability-driven detailed macro placement is proposed to refine the solution. Compared with the top contestants of the MLCAD 2023 contest, experimental results show that our algorithm achieves the best overall score and routability. Keyu Peng, Jun Yang 0006, Ziran Zhu |
DAC | 4 |
| 2024 | CATCAM: a 28 nm constant-time alteration TCAM enabling less than 50 ns update latency
Chenchen Deng, Tianzhu Xiong, Zhaoshi Li, Jianfeng Zhu 0001, Jun Yang 0006, Shaojun Wei, Leibo Liu |
Sci. China Inf. Sci. | 7 |
| 2024 | An analysis of TinyML@ICCAD for implementing AI on low-power microprocessor
Meng Zhang 0010, Tinghuan Chen, Jun Yang 0006 |
Sci. China Inf. Sci. | 6 |
| 2024 | High-Performance Placement Engine for Modern Large-Scale FPGAs With Heterogeneity and Clock ConstraintsabstractAs field-programmable gate array (FPGA) architectures continue to evolve and become more complex, the heterogeneity and clock constraints imposed by modern FPGAs have posed significant challenges to FPGA placement. This article proposes a high-performance placement engine for modern large-scale FPGAs with heterogeneity and clock constraints. To improve efficiency and scalability, we develop a clustering method considering both internal/external connectivity and the balance of block types to build the hierarchy. In each hierarchy level, we propose a hybrid penalty and augmented Lagrangian method (HPALM) to convert the FPGA global placement with heterogeneity and clock constraints into a series of unconstrained optimization subproblems, then use the Adam method to solve each subproblem. In particular, we prove that the HPALM is globally convergent for global placement. Besides, a matching-based IP block legalization is developed to legalize the DSPs and RAMs, and a multistage packing is presented to cluster LUTs and FFs into HCLBs. Finally, we propose a history-based legalization to legalize CLBs in an FPGA, and a simulated-annealing-based detailed placement is presented to reduce the wirelength while maintaining legality. Compared with the state-of-the-art works, experimental results based on the ISPD 2017 contest benchmarks show that the proposed algorithm can achieve the shortest routed wirelength in a reasonable runtime. Ziran Zhu, Yangjie Mei, Kangkang Deng, Jianli Chen, Jun Yang 0006, Yao-Wen Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2024 | A 0.65 V 4 dB NF 2.4 GHz Sub-Passive RF Down-Converter With Trans-Frequency Current-Reusing Scheme Achieving Low Flicker Noise and High LinearityabstractIn this article, a 2.4 GHz low-power sub-passive radio frequency (RF) down-converter that employs trans-frequency current-reusing technique is proposed. Part of the RF trans-conductance ($\bf{g_m}$) stage is reused as the bias current source of the trans-impedance amplifier (TIA) to realize trans-frequency current reusing, thereby saving 30% power consumption of the sub-passive down-converter. Compared with a conventional passive down-converter where the flicker noise from the TIA’s bias current source is fully fed into the intermediate frequency (IF) output terminal, the flicker noise originating from the TIA’s current source supplied by the RF$\bf{g_m}$stage is up-converted and filtered out. The flicker noise corner frequency is cut down to 20 KHz, making it possible to apply zero-IF receiver for narrow-band communications. The linearity is also improved by sufficiently cancelling third-order transconductance from PMOS and NMOS$\bf{g_m}$transistors across a wide input swing range, which is contributed by their asymmetrical bias current in the low-noise transconductance amplifier (LNTA). Fabricated in TSMC 40 nm RF CMOS process, the prototype occupies a die area of 0.27mm$^2$. Operating at 2.4 GHz, the sub-passive RF down-converter achieves a measured conversion gain (CG) of 56 dB and a noise figure (NF) of 4 dB, while implementing a 23 dBm output third-order intercept point (OIP3) with 0.8 mW power consumption at a supply voltage of 0.65V. Yan Zhao 0039, Chao Chen 0018, Jun Yang 0006 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | An efficient path delay variability model for wide-voltage-range digital circuits
Weiwei Shan, Yuqiang Cui, Wentao Dai, Xinning Liu, Peng Cao 0002, Jun Yang 0006 |
Sci. China Inf. Sci. | 7 |
| 2023 | From macro to microarchitecture: reviews and trends of SRAM-based compute-in-memory circuits
Zhaoyang Zhang 0008, Jinwu Chen, Xi Chen 0107, An Guo 0001, Bo Wang 0023, Tianzhu Xiong, Yuyao Kong, Xingyu Pu, Shengnan He, Xin Si, Jun Yang 0006 |
Sci. China Inf. Sci. | 11 |
| 2023 | Efficient and Accurate ECO Leakage Optimization Framework With GNN and Bidirectional LSTMabstractEngineering change order (ECO) plays an important role in design flow to perform leakage optimization with gate-sizing and$V_{\mathrm{ th}}$assignment approaches. Unfortunately, it is extremely time consuming due to the iterative nature of cell swap and timing check. Many learning-based methods, especially, graph neural networks (GNNs), have been utilized in leakage optimization to predict$V_{\mathrm{ th}}$assignment, but most of them treat the cells and their neighborhood cells uniformly when aggregating cell-level topology information to gather design-level information and discard the path-level information, suffering from accuracy loss, which could be exploited by bidirectional long short-term memory (BiLSTM) network. In this work, a GNN-BiLSTM-based framework is proposed to perform commercial-quality$V_{\mathrm{ th}}$assignment for leakage optimization by learning design-level and path-level information and is validated with the benchmarks from Opencores and IWLS 2005 under TSMC 28 nm technology. The experimental results demonstrate that the proposed framework achieves the most accurate$V_{\mathrm{ th}}$assignment prediction compared with the competitive models with F1-score ranging from 0.954 to 0.975 for seen designs and from 0.945 to 0.965 for unseen designs, respectively. The divergence between the leakage optimization results of this work and the commercial tool is limited to be between 8.5% and 26.1%, which is reduced by at least$2.2\times $compared with prior works. Owing to efficient training convergence and inference speed, our approach achieves significant runtime improvement by up to$10\times $over commercial tool with similar leakage optimization results. Peng Cao 0002, Guoqing He, Zhanhua Zhang, Jun Yang 0006 |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2022 | High-performance placement for large-scale heterogeneous FPGAs with clock constraintsabstractWith the increasing complexity of the field-programmable gate array (FPGA) architecture, heterogeneity and clock constraints have greatly challenged FPGA placement. In this paper, we present a high-performance placement algorithm for large-scale heterogeneous FPGAs with clock constraints. We first propose a connectivity-aware and type-balanced clustering method to construct the hierarchy and improve the scalability. In each hierarchy level, we develop a novel hybrid penalty and augmented Lagrangian method to formulate the heterogeneous and clock-aware placement as a sequence of unconstrained optimization subproblems and adopt the Adam method to solve each unconstrained optimization subproblem. Then, we present a matching-based IP blocks legalization to legalize the RAMs and DSPs, and a multi-stage packing technique is proposed to cluster FFs and LUTs into HCLBs. Finally, history-based legalization is developed to legalize CLBs in an FPGA. Based on the ISPD 2017 clock-aware FPGA placement contest benchmarks, experimental results show that our algorithm achieves the smallest routed wirelength for all the benchmarks among all published works in a reasonable runtime. Ziran Zhu, Yangjie Mei, Zijun Li 0005, Jingwen Lin, Jianli Chen, Jun Yang 0006, Yao-Wen Chang |
DAC | 6 |
| 2022 | Triple-Skipping Near-MRAM Computing Framework for AIoT EraabstractNear memory computing (NMC) paradigm shows great significance in non-von Neumann architecture to reduce data movement. The normally-off and instance-on characteristics of spin-transfer torque magnetic random access memory (STT-MRAM) promise energy-efficient storage in the AIoT era. To avoid unnecessary memory-related processing, we propose a novel write-read-calculation triple-skipping (TS) NMC for multiply-accumulate (MAC) operation with minimally modified peripheral circuits. The proposed TS-NMC is evaluated with a custom micro control unit (MCU) in 28-nm high-K metal gate (HKMG) CMOS process and foundry announced universal two-transistor two-magnetic tunnel junction (2T-2MTJ) MRAM cell. The framework consists of a sparse flag which is defined in extra STT-MRAM columns with only 0.73% area overhead, and a calculation block for NMC logic with 9.9% overhead. The TS-NMC can successfully work at 0.6-V supply voltage under 20MHz. This Near-MRAM framework can offer up to ~9S.6 % energy saving compared to commercial SRAM refer to ultra-low-power benchmark (ULP-Benchmark). Classification task on MNIST takes 13nJ/pattern. The energy access of memory, calculation, and the total can be reduced by$52.49\times, 2.7\times$, and 11.3 × respectively from the TS scheme. Juntong Chen, Hao Cai 0001, Bo Liu 0019, Jun Yang 0006 |
DATE | 4 |
| 2022 | A Target-Separable BWN Inspired Speech Recognition Processor with Low-power Precision-adaptive Approximate ComputingabstractThis paper proposes a speech recognition processor based on a target-separable binarized weight network (BWN), capable of performing both speaker verification (SV) and keyword spotting (KWS). In traditional speech recognition system, the SV based on traditional model and the KWS based on neural networks (NN) model are two independent hardware modules. In this work, both SV and KWS are processed by the proposed BWN with unified training and optimization framework which can be performed for various application scenarios. By the system-architecture co-design, SV and KWS share most of the network parameters, and the classification part is calculated separately according to different targets. An energy-efficient NN accelerator which can be dynamically reconfigured to process different layers of the BWN with splitting calculation of frequency domain convolution is proposed. SV and KWS can be achieved with only one time calculation of each input speech frame, which greatly improves the computing energy efficiency. The computing units of the NN accelerator are optimized using precision-adaptive approximate computing method with Dual-VDD to further reduce the energy cost. Compared to state-of-the-arts, this work can achieve about 4 × reduction in power consumption while maintaining high system adaptability and accuracy. Bo Liu 0019, Hao Cai 0001, Haige Wu, Anfeng Xue, Zhen Wang 0019, Jun Yang 0006 |
DATE | 8 |
| 2022 | Graph Convolutional Network Empowered Indoor Localization Method via Aggregating MIMO CSIabstractWith the explosive growth of advanced wireless technologies and computing device platforms, mobile sensing has gained huge attention. Indoor localization is actually considered as one of most valuable techniques in the field of contactless sensing. In this paper, we propose a novel graph convolutional network (GCN) empowered indoor localization method, which aggregates channel state information (CSI) features extracted from multiple multiple-input multiple-output (MIMO) links. CSI features from multiple antennas are basically converted into graph nodes in order to adopt GCN classification model. At the same time, graph attention mechanism is introduced to study and transfer spatial and frequency of CSI features. Eventually, output of graph is mapped with multiple measurement points through prediction network to provide final estimate position. 5GHz commercial Wi-Fi equipment is respectively utilized for data collection and experimental evaluation in two representative indoor scenarios. Experimental result shows that the proposed method has better performance in robust localization compared to other state-of-the-art deep learning methods. Jun Yang 0006, Zhengran He, Guan Gui 0001, Haris Gacanin |
GLOBECOM | 2 |
| 2022 | ShareFloat CIM: A Compute-In-Memory Architecture with Floating-Point Multiply-and-Accumulate OperationsabstractCompute-in-memory (CIM) has been widely explored to overcome “Von-Neumann bottleneck” for its high throughput and energy efficiency. However, recent compute-in-memory works can only support integer (INT)-type multiply-and-accumulate (MAC) operations. Floating point MACs (FP-MAC) are highly required to achieve both high performance training and high accuracy inference. In this paper, we proposed a ShareFloat CIM architecture which can support FP-MAC operations. Neural networks with ShareFloat MAC can achieve almost the same accuracy as that with FP64 MAC. A 28nm 64Kb ShareFloat CIM macro was further implemented with an energy efficiency of 18.8 TFLOPS/W and 73.11% accuracy when applied to a VGG-16 network with ShareFloat MAC and CIFAR-100 dataset. An Guo 0001, Yongliang Zhou, Bo Wang 0023, Tianzhu Xiong, Xin Si, Jun Yang 0006 |
ISCAS | 8 |
| 2022 | SNNIM: A 10T-SRAM based Spiking-Neural-Network-In-Memory architecture with capacitance computationabstractSpiking-Neural-Networks (SNN) have natural advantages in high-speed signal processing and big data operation. However, due to the complex implementation of synaptic arrays, SNN based accelerators may face low area utilization and high energy consumption. Computing-In-Memory (CIM) shows great potential in performing intensive and high energy efficient computations. In this work, we proposed a JOT-SRAM based Spiking-Neural-Network-In-Memory architecture (SNNIM) with 28nm CMOS technology node. A compact JOT-SRAM bit-cell was developed to realize signed 5bit synapses arrays and configurable bias arrays (SYBIA). The soma array based standard 8T-SRAM (SMTA) stores the soma membrane voltage and the threshold value. A capacitance computation scheme (CCA) between them was proposed to support various SNN operations. The proposed SNNIM achieved energy efficiency of 25.18 TSyOPSI. And the proposed SNNIM achieved 1.79+× better array efficiency compared with previous works. Bo Wang 0023, Xiang Li 0147, Anran Yin, Zhongyuan Feng, Yuyao Kong, Tianzhu Xiong, Haiming Hsu, Yongliang Zhou, An Guo 0001, Jun Yang 0006, Xin Si |
ISCAS | 13 |
| 2022 | Self-compensation tensor multiplication unit for adaptive approximate computing in low-power CNN processing
Bo Liu 0019, Hao Cai 0001, Reyuan Zhang, Zhen Wang 0019, Jun Yang 0006 |
Sci. China Inf. Sci. | 6 |
| 2022 | An Efficient BCNN Deployment Method Using Quality-Aware Approximate ComputingabstractAs the artificial intelligence and Internet of Things (AIoT) develop rapidly, the deployment of artificial neural networks in edge computing is becoming significant with great challenge. The binarized convolutional neural network (BCNN) is one of the most widely adopted light-weight ANNs in AIoT, which can achieve the balance of system accuracy and hardware resource consumption, compared to others. To achieve high power and area efficiency in BCNN deployment, many approximate computing (AxC) techniques are integrated to make full use of the resilience of BCNN. As the research focused on the integration of AxC in circuit design, the design of AxC itself is not fully considered when applied to specific applications or domains. Based on circuit-architecture-system co-design, this article proposes an efficient BCNN deployment method, including a quality-circuit co-design method for approximate adder generation, a quality-aware intercompensation approach for addition tree, and a computing quality involved retraining approach for BCNN deployment. Experimental results show that the proposed quality model can achieve 86.43% in average accuracy while evaluating nine types of typical approximate adders. The proposed method is conducted on the applications of keyword spotting of GSCD, MNIST, and CIFAR-10, and we can further rise the approximation degree by 50%–75%, while reducing the accuracy by less than 1%. Bo Liu 0019, Xuetao Wang, Anfeng Xue, Qiao Shen 0001, Na Xie, Yu Gong 0002, Zhen Wang 0019, Jun Yang 0006, Hao Cai 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 10 |
| 2022 | Proposal of Analog In-Memory Computing With Magnified Tunnel Magnetoresistance Ratio and Universal STT-MRAM CellabstractIn-memory computing (IMC) is an effective solution for energy-efficient artificial intelligence applications. Analog IMC amortizes the power consumption of multiple sensing amplifiers with an analog-to-digital converter (ADC) and simultaneously completes the calculation of multi-line data with a high parallelism degree. Based on a universal one-transistor one-magnetic tunnel junction (MTJ) spin transfer torque magnetic RAM (STT-MRAM) cell, this paper demonstrates a novel tunneling magnetoresistance (TMR) ratio magnifying method to realize analog IMC. Previous concerns including low TMR ratio and analog calculation nonlinearity are addressed using device-circuit interaction. The TMR is magnified$7500\times $using a latch structure in combination with the device. Peripheral circuits are minimally modified to enable in-memory matrix-vector multiplication. A current mirror with a feedback structure is implemented to enhance analog computing linearity and calculation accuracy. The proposed design maximumly supports 1024 2-bit input and 1-bit weight multiply-and-accumulate (MAC) computations simultaneously. The proposal is simulated using the 28-nm CMOS process and MTJ compact model. The integral nonlinearity is reduced by 57.6% compared with the conventional structure. 9.47-25.4 TOPS/W is realized with 2-bit input, 1-bit weight, and 4-bit output convolution neural network (CNN). Hao Cai 0001, Yanan Guo 0004, Bo Liu 0019, Mingyang Zhou 0002, Juntong Chen, Xinning Liu, Jun Yang 0006 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2022 | More is Less: Domain-Specific Speech Recognition Microprocessor Using One-Dimensional Convolutional Recurrent Neural NetworkabstractLow-power keywords recognition has been a focus of acoustic signal processing for several decades. This work investigates the domain-specific speech recognition microprocessor based on optimized one-dimensional convolutional recurrent neural network (1D-CRNN). Compared to previous DNN based frameworks, the proposed 1D-CRNN can process both the feature extraction and keywords classification, and achieve high recognition accuracy with reduced computation operations under wide range background noise SNRs. An energy-efficient 1D-CRNN accelerator is implemented to dynamically reconfigure and process the different layers. This accelerator has the characteristics of “More is Less” in three aspects: 1) the hybrid network with more complex layers is much more compact and requires less computation; 2) although the weight width quantized to 8 bits requires more memory size and multiplication energy cost, the required network neurons can be reduced and hardware utilization can be improved; 3) an energy-aware self-compensation tensor multiplication unit with dual power supply based on approximation design method can be utilized for 1D-CRNN computing. Compared to the state-of-the-art architectures, the novel more-is-less architecture can achieve a much lower power consumption of$1.4~\mu \text{W}\sim 2.1~\mu \text{W}$(over 80% reduced) under an industry 22nm technology, while maintaining higher system adaptability (support SNRs: −5dB~Clean) for 1~5 real-time keywords recognition. Bo Liu 0019, Hao Cai 0001, Xiaoling Ding, Yu Gong 0002, Weiqiang Liu 0001, Jinjiang Yang, Zhen Wang 0019, Jun Yang 0006 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 10 |
| 2021 | A survey of in-spin transfer torque MRAM computing
Hao Cai 0001, Bo Liu 0019, Juntong Chen, Lirida A. B. Naviner, Yongliang Zhou, Zhen Wang 0019, Jun Yang 0006 |
Sci. China Inf. Sci. | 7 |
| 2021 | Semi-Analytical Path Delay Variation Model With Adjacent Gates Decorrelation for Subthreshold CircuitsabstractThe subthreshold circuit is a practical design style for the ultralow-power applications, but its timing estimation is a challenge due to the increasing local variation effects. The delay variation of adjacent gates is not independent because of input slew variation caused by the precedent gate, so their correlation effects are difficult to model and estimate. This article proposes a semi-analytical statistical delay model considering local variation for the subthreshold region, that is, the combination of analytical and simulation-based method. First, it decorrelates the slew influence between adjacent stages by dividing delay and output slew model into fast/slow input cases and dividing delay variation model into process variation and input slew variation. Then, it can be applied into multi-PVT conditions with a one-time SPICE nominal simulation by analyzing the independence of variability and relative variability of step input gate delay variance with output load capacitance and process, voltage, and temperature (PVT). Finally, experiments are carried out for different benchmarks, processes, voltages, and temperatures (BPVTs). The average errors of variance on different BPVTs are 4.8%, 3.1%, 4.0%, and 4.7%. Compared with other analytical works, the accuracies' improvements of three metrics (variance, variability, and max delay) are 8.3×, 9.6×, and 2.7× by the mean error at all test benchmarks. Compared with industrial method LVF, it has a comparable error in max path delay and runtime, and less three orders of magnitudes than LVF in characterization time and stored data (from TB to GB) at all test benchmarks. Peng Cao 0002, Mengxiao Li, Yu Gong 0002, Zhiyuan Liu 0011, Geng Bai, Jun Yang 0006 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2021 | A Design of Timing Speculation SRAM-Based L1 Caches With PVT Autotracking Under Near-Threshold VoltagesabstractTo improve the performance of SRAM in caches under near-threshold voltages, several timing speculation techniques, such as the cross-sensing SRAM (CS-SRAM), are proposed. Meanwhile, for a given process, voltage, and temperature (PVT) condition, CS-SRAM has an optimal bitline discharging time (TBL) to achieve the lowest average access latency. However, existing timing speculation caches do not track the variations of different PVT conditions to adjust the access timing to the optimal TBL point, on which the system possesses the lowest average memory access time. In this article, we propose a design of CS-SRAM-based L1 caches with a PVT autotracking mechanism, namely TS-PULP, which adjusts both the TBL and the frequency of the system clock to the optimal points. To quantify the improvement of our approach, a cycle-accurate RTL model of CS-SRAM and a field-programmable gate array (FPGA) prototype of the proposed L1 caches with the open-source system on chip (SoC) platform PULP have also been implemented. According to the evaluation results from RTL simulations and the FPGA prototype, the proposed caches can achieve a similar performance (over 80%) of the original standard cell memory (SCM)-based PULP design under TSMC 28 nm 0.5 V and 25 °C with only about 30% chip area. In addition, we introduced the figure of merit (FOM) of million instructions per second (MIPS), area, and energy (MAE) to comprehensively evaluate different approaches. The proposed scheme, TS-PULP, achieves the best FOM of MAE among four architectures (PULP with SCM, PULP with SRAM, TS-cache, and TS-PULP) with different cache sizes. Qingde Lin, Ke Tan 0004, Tianxiang Shao, Shan Shen, Jun Yang 0006 |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2020 | Modeling and Designing of a PVT Auto-tracking Timing-speculative SRAMabstractIn the low supply voltage region, the performance of 6T cell SRAM degrades seriously, which takes more time to achieve the sufficient voltage difference on bitlines. Timing- speculative techniques are proposed to boost the SRAM frequency and the throughput with speculatively reading data in an aggressive timing and correcting timing failures in one or more extended cycles. However, the throughput gains of timing- speculative SRAM are affected by the process, voltage and temperature (PVT) variations, which causes the timing design of speculative SRAM to be either too aggressive or too conservative. This paper first proposes a statistical model to abstract the characteristics of speculative SRAM and shows the presence of an optimal sensing time that maximizes the overall throughput. Then, with the guidance of the performance model, a PVT auto-tracking speculative SRAM is designed and fabricated, which can dynamically self-tune the bitline sensing to the optimal time as the working condition changes. According to the measurement results, the maximum throughput gain of the proposed 28nm SRAM is 1.62X compared to the baseline at 0.6V VDD. Shan Shen, Tianxiang Shao, Jun Yang 0006, Longxing Shi |
DATE | 4 |
| 2020 | Power-Efficient Approximate Multiplier Using Adaptive Error CompensationabstractIn this paper, a design framework is proposed for a power-efficient approximate multiplier using adaptive error compensation and the optimal error compensation values are determined by using probability theory, resulting in a minimal mean-squared-error (MSE). To further pursue the viability of adaptive accuracy, adaptive error compensation scheme using different levels of quantization is utilized for the predetermined compensation values. The simulation results show that the approximate multipliers based on the proposed design framework outperform state-of-the-art designs in both accuracy and circuit measurements. Specifically, with a higher accuracy, the proposed designs save up to 36.84% and 21.61% in power consumption compared to state-of-the-art unsigned and signed approximate $16\times16$ multipliers, respectively. In terms of power-delay-product (PDP), the improvements are up to 42.3% and 21.93% for the unsigned and signed multiplier designs, respectively. Finally, the approximate multipliers are further assessed in the implementation of an FIR filter. It shows that the proposed approximate multiplier achieves a similar filtering quality to the accurate design, with more than 50% reduction in power dissipation. Zhixi Yang, Honglan Jiang, Jun Yang 0006 |
ACM Great Lakes Symposium on VLSI | 4 |
| 2020 | Statistical Timing Model for Subthreshold Circuit with Correlated Variation ConsiderationabstractSubthreshold circuit has the significant advantage in low-power applications but suffers from serious variation increase. Traditional EDA tools could not achieve balance between accuracy and simulation effort for statistical timing analysis in subthreshold region. Many researches have been devoted to statistical timing model to reveal the relation between delay variation and process variation with physical insight. However, the statistical correlation of gate delay in circuit path is hard to capture and not considered appropriately in most prior works. In this paper, a statistical timing model for subthreshold circuit is proposed with the consideration of local process variation and correlated variation from adjacent gates, which is established by deriving variance models for gate delay and output waveform analytically for fast and slow input waveform separately, so that the path delay variation can be translated from the accumulation of the correlated gate delays with input slew to a linear combination of independent step input delays. The proposed model was verified under the process of TSMC28nm technology at the subthreshold supply voltage with the estimation error of less than 6% for circuit path delay variation in benchmark ISCAS99 compared with Monto Carlo simulation results, which outperforms prior works with 5~10X accuracy increase and acceptable simulation cost. Peng Cao 0002, Mengxiao Li, Zhiyuan Liu 0011, Jun Yang 0006 |
ISCAS | 5 |
| 2020 | CATCAM: Constant-time Alteration Ternary CAM with Scalable In-Memory ArchitectureabstractTCAM (Ternary Content-Addressable Memory) is the essential component for high-speed packet classification in modern hardware switches. However, due to its relatively slow update process, recent advances in Software-Defined Network (SDN) regard them as the bottleneck to the agile deployment of network services. Rule installation in commodity switches suffers from non-deterministic delays, ranging from a few milliseconds to nearly half a second. The crux of the problem is that TCAM prioritizes rules based on physical addresses. Corresponding entries have to be reallocated according to the priority of an incoming rule, such that the insertion delay grows linearly with the number of existing rules in a TCAM. In this paper, we present Constant-time Alteration Ternary CAM (CATCAM) that can accomplish both lookup queries and update requests for packet classification in a few nanoseconds. The key to fast update is to decouple rule priorities from physical addresses. We propose a matrix-based priority encoding scheme that records the priority relation between rules and can be implemented in 8T SRAM arrays with the emerging Processing In-Memory (PIM) technique. CATCAM also comes with a hierarchical architecture to scale out, its interval-based scheduling scheme guarantees deterministic update performance in all scenarios. CATCAM is developed under full-custom design in the 28 nm process. Evaluation across benchmark workloads shows that CATCAM provides at least three orders of magnitude speedup over state-of-the-art TCAM update algorithms and offers equivalent search capability to conventional TCAM while incurring 0.3% power and 20% area overhead. Dibei Chen, Zhaoshi Li, Tianzhu Xiong, Jun Yang 0006, Shouyi Yin, Shaojun Wei, Leibo Liu |
MICRO | 5 |
| 2020 | Towards an automated design flow for memristor based VLSI circuits
Hao Cai 0001, Chao Wang 0068, Jun Yang 0006 |
Integr. | 4 |
| 2020 | TS Cache: A Fast Cache With Timing-Speculation Mechanism Under Low Supply VoltagesabstractTo mitigate the ever-worsening “power wall” problem, more and more applications need to expand their working voltage to the wide-voltage range including the nearthreshold region. However, the read delay distribution of the static random access memory (SRAM) cells under the nearthreshold voltage shows a more serious long-tail characteristic than that under the nominal voltage due to the process fluctuation. Such degradation of SRAM delay makes the SRAM-based cache a performance bottleneck of systems as well. To avoid unreliable data reading, circuit-level studies use larger/more transistors in a bitcell by sacrificing chip area and the static power of cache arrays. Architectural studies propose the auxiliary error correction or block disabling/remapping methods in fault-tolerant caches, which worsen both the hit latency and energy efficiency due to the complex accessing logic. This article proposes a timing-speculation (TS) cache to boost the cache frequency and improve energy efficiency under low supply voltages. In the TS cache, the voltage differences of bitlines (BLs) are continuously evaluated twice by a sense amplifier (SA), and the access timing error can be detected much earlier than that in prior methods. According to the measurement results from the fabricated chips, the TS L1 cache aggressively increases its frequency to 1.62× and 1.92× compared with the conventional scheme at 0.5- and 0.6-V supply voltages, respectively. Shan Shen, Tianxiang Shao, Xiaojing Shang, Jun Yang 0006, Longxing Shi |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2019 | A Statistical Current and Delay Model Based on Log-Skew-Normal Distribution for Low Voltage RegionabstractThe increasing performance variation and non-Gaussian distribution pose remarkable challenges to timing analysis for circuits operating in low voltage region. Accurate modeling of the statistical characteristics is urgently required with process variation consideration. In this paper, the statistical models for drain current and gate delay in low voltage region are established in analytical form based on the log-skew-normal (LSN) distribution via moment matching technique. Experimental results show that the probability distribution function (PDF) curves obtained from the proposed models for drain current and gate delay are highly fitted with Monte Carlo (MC) simulation results in sub/near-threshold regions. Moreover, owing to the proposed LSN-based statistical model, less than 8% error is introduced in the predicted sensitivity of gate delay and the maximum/minimum delay indicated by ±3σ percentile points can be calculated more precisely than the LN-based method with up to 3× accuracy improvement for low supply voltage. Peng Cao 0002, Jiangping Wu, Zhiyuan Liu 0011, Jun Yang 0006, Longxing Shi |
ACM Great Lakes Symposium on VLSI | 5 |
| 2019 | Voltage-Controlled Magnetoelectric Memory Bit-cell Design With Assisted Body-bias in FD-SOIabstractVoltage-controlled magnetic anisotropy (VCMA)-magnetic tunnel junction (MTJ) is incorporated into FD-SOI CMOS technology. The design space of 1 transistor-1 MTJ (1T-1M) bit-cell is explored through varied VCMA pulse duration/amplitude and scaling down transistor dimensions. The design point with 1.1 V VCMA pulse amplitude, 0.44 ns pulse duration and W/L = 400 nm/30 nm access transistor shows the ultra low write energy in VCMA-MTJ based bit-cell. It achieves a minimum 3.18 fJ/bit switching energy with 28-nm FD-SOI process. Access transistor sizing is studied, while the ultra low power implementation may lead to MTJ switching failure. Voltage assisted techniques for failure mitigation are proposed based on body-bias generator (BBG). The BBG not only provides VCMA pulse signal to control MTJ barrier, but also generates body-bias to boost the transistor performance. In the presence of forward body-bias (FBB) and increased VCMA pulse level, the proposed strategy is effective in switching failure compensation as well as writing delay improvement. Hao Cai 0001, Menglin Han, Weiwei Shan, Jun Yang 0006, You Wang 0002, Wang Kang 0001, Weisheng Zhao 0001 |
ACM Great Lakes Symposium on VLSI | 4 |
| 2019 | In-memory Processing based on Time-domain CircuitabstractDeep Neural Networks (DNN) have emerged as a dominant algorithm for machine learning (ML). High performance and extreme energy efficiency are critical for deployments of DNN, especially in mobile platforms such as autonomous vehicles, cameras, and other devices of internet of things. However, DNNs lead to massive data movement and memory accesses, which prevents it from being integrated into always-on Internet-of-Things (IoT) devices. Recently, computing in-memory (CIM) architectures embeds analog computation circuits in/near the memory arrays. It significantly reduces data movement energy. This paper summarizes the most recent novel methods on the CIM architectures based on the time-domain computation. Compared with voltage-domain and frequency-domain analog computing method, time-domain computation provides more flexibility, higher accuracy and greater scalability for larger neural networks. Thereafter, the first in-memory binary weight network (BWN) processor based on pulse-width modulation in which the feature is stored in memory is also presented. This work significantly reduces memory accesses (4x), and achieves state-of-the-art peak energy efficiency of 119.7TOPS/W. Yuyao Kong, Jun Yang 0006 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2019 | A Statistical Timing Model for Low Voltage Design Considering Process VariationabstractNear-threshold voltage (NTV) design suffers severe challenge due to the dramatic increase in performance uncertainty introduced by process variation. This paper proposes an analytical approach based on Log-Normal (LN) distribution to characterize the statistical delay for NTV design considering the dominant threshold voltage variation from gate-level to circuit-level. At gate-level, the multivariable threshold voltage variation issue is solved by the equivalent threshold voltage method and equivalent drain current method for generic gates with stack topology and parallel topology, respectively. At circuit-level, a statistical timing model is proposed as the linear combination of the independent statistical delays of all gates in the path with step input by considering the varied correlation between adjacent gates. To the best of our knowledge, we firstly propose a statistical timing model analytically for practical circuit path with physical insights of supply voltage, transistor size, and load capacitance. The characterization effort for each path is only one-time SPICE simulation, which is negligible compared with Monte Carlo (MC) simulation in statistical static timing analysis (SSTA) methods. Experimental results under a commercial 28-nm CMOS process show the proposed models have high accuracy at low supply voltage compared with MC simulations, where the modeling errors for the mean and variance of gate delay can be limited within 1.54% and 11.2%, respectively. Moreover, as for the practical paths in ITC'99 benchmark, the maximum modeling errors of mean, variance, minimum delay, and maximum delay is less than 3.04%, 11.40%, 4.50%, and 2.87%, respectively. Peng Cao 0002, Zhiyuan Liu 0011, Jiangping Wu, Jun Yang 0006, Longxing Shi |
ICCAD | 5 |
| 2019 | Low-Power Centimeter-Level Localization for Indoor Mobile Robots Based on Ensemble Kalman Smoother Using Received Signal StrengthabstractHow to provide a low-cost but accurate localization solution for the indoor mobile robots are essential in many Internet of Things applications, such as smart home and asset tracking. To achieve this goal, this paper originally proposes a modified two-filter smoother based on ensemble Kalman filter (KF) (denoted as EnKS) for the localization of indoor mobile robots. The proposed EnKS algorithm consists of both a forward part of an ensemble KF (EnKF) with statistical linear regression and a backward part of a modified information KF with state error vector. The EnKS based on stochastic sampling with ensemble members can achieve better positioning accuracy than other Kalman smoothers. When compared to EnKF, the proposed EnKS combines a backward filter to compensate for the estimation error of EnKF and further improves the accuracy. Furthermore, the implementation of the proposed EnKS is conducted in the real world visible light positioning (VLP) system using pre-existing LED lights for low-cost robot localization. To make a performance comparison, this paper also uses baseline smoothers based on extended KF and central difference KF in the VLP system. Preliminary experimental results imply that the proposed EnKS is able to achieve the best positioning accuracy, as high as 11.18 cm on average, but with a comparable computational complexity, which enables to meet the demands of many robot applications. Yuan Zhuang 0001, Min Shi 0001, Pan Cao, Longning Qi, Jun Yang 0006 |
IEEE Internet Things J. | 6 |
| 2018 | A fast and robust failure analysis of memory circuits using adaptive importance sampling methodabstractPerformance failure has become a growing concern for the robustness and reliability of memory circuits. It is challenging to accurately estimate the extremely small failure probability when failed samples are distributed in multiple disjoint failure regions. In this paper, we develop an adaptive importance sampling (AIS) method. AIS has several iterations of sampling region adjustments, while existing methods pre-decide a static sampling distribution. By iteratively searching for failure regions, AIS may lead to better efficiency and accuracy. This is validated by our experiments. For SRAM cell with single failure region, AIS uses 5-10X fewer samples and reaches better accuracy when compared to several recent methods. For sense amplifier circuit with multiple failure regions, AIS is 4369X faster than MC without compromising accuracy, while other methods fail to cover all failure regions in our experiment. Xiao Shi 0001, Jun Yang 0006, Lei He 0001 |
DAC | 3 |
| 2018 | Enabling Resilient Voltage-Controlled MeRAM Using Write Assist TechniquesabstractReliability concerns arise in nonvolatile magnetoelectric random access memory (MeRAM) due to continuously nanotechnology scaling down and CMOS-magnetic hybrid integration. The primary objective of this work is to investigate failure mitigation in voltage-controlled magnetic anisotropy-magnetic tunnel junction (VCMA-MTJ) based 1T-1MTJ MeRAM bit-cell, by using MTJ compact model and 28nm fully depleted silicon on insulator (FD-SOI) process design-kit. A comprehensive reliability study is performed considering process variation and aging degradations, including hot carrier injection (HCI), bias temperature instability (BTI), soft breakdown (SBD) and radiation effect. Write assist techniques are proposed to ensure failure resilient MeRAM design. Bit line (BL) boost and negative source line (SL) methods show high efficiency in writing latency improvement and failure mitigation. Hao Cai 0001, You Wang 0002, Wang Kang 0001, Lirida A. B. Naviner, Weiwei Shan, Jun Yang 0006, Weisheng Zhao 0001 |
ISCAS | 6 |
| 2018 | Guest Editorial: Special Issue on Toward Positioning, Navigation, and Location-Based Services (PNLBS) for Internet of ThingsabstractIn the past decade, technological advancements have facilitated the manufacturing of compact, inexpensive, and low-power consuming receivers and sensors for smart devices (e.g., GPS, WiFi, MEMS sensors, RFID, UWB, BLE, etc.). This led to the fast development of positioning, navigation, and location-based services (PNLBS), and much broader new applications than just providing a location or navigation. Yuan Zhuang 0001, Yue Cao 0002, Naser El-Sheimy, Jun Yang 0006 |
IEEE Internet Things J. | 4 |
| 2018 | A Pervasive Integration Platform of Low-Cost MEMS Sensors and Wireless Signals for Indoor LocalizationabstractLocation service is fundamental to many Internet of Things applications such as smart home, wearables, smart city, and connected health. With existing infrastructures, wireless positioning is widely used to provide the location service. However, wireless positioning has the limitations such as highly depending on the distribution of access points (APs); providing a low sample-rate and noisy solution; requiring extensive labor costs to build databases; and having unstable RSS values in indoor environments. To reduce these limitations, this paper proposes an innovative integrated platform for indoor localization by integrating low-cost microelectromechanical systems (MEMS) sensors and wireless signals. This proposed platform consists of wireless AP localization engine and sensor fusion engine, which is suitable for both dense and sparse deployments of wireless APs. The proposed platform can automatically generate wireless databases for positioning, and provide a positioning solution even in the area with only one observed wireless AP, where the traditional trilateration method cannot work. This integration platform can integrate different kinds of wireless APs together for indoor localization (e.g., WiFi, Bluetooth low energy, and radio frequency identification). The platform fuses all of these wireless distances with low-cost MEMS sensors to provide a robust localization solution. A multilevel quality control mechanism is utilized to remove noisy RSS measurements from wireless APs and to further improve the localization accuracy. Preliminary experiments show the proposed integration platform can achieve the average accuracy of 3.30 m with the sparse deployment of wireless APs (1 AP per 800 m2). Yuan Zhuang 0001, Jun Yang 0006, Longning Qi, You Li 0001, Yue Cao 0002, Naser El-Sheimy |
IEEE Internet Things J. | 2 |
| 2017 | Context Management Scheme Optimization of Coarse-Grained Reconfigurable Architecture for Multimedia ApplicationsabstractDue to the combination of flexibility and efficiency, coarse-grained reconfigurable architectures (CGRAs) are suitable for the implementation of computing-intensive applications. However, with the growing performance requirements, the scale of CGRA increases exponentially, which leads to configuration performance degradation and configuration power rise. Based on the analysis of configuration context features, we optimize the context management scheme of CGRA from the aspects of context cache structure and replacement strategy. The context cache is structured hierarchically to reduce the memory overhead without configuration performance degradation and a hybrid context replacement algorithm is proposed to further increase the configuration efficiency with a novel context frequency weight factor. Experimental results show that the proposed context management scheme improves the configuration performance of the base CGRA significantly by 13.6%-20.5% for H.264 decoding and 13.6%-20.5% for MPEG2 decoding with only 43% context cache cost. Compared with other works, the proposed context management scheme shows the advantages of 2.3-6× less normalized context cache size and 2.3-2.7× cache efficiency. Peng Cao 0002, Bo Liu 0019, Jinjiang Yang, Jun Yang 0006, Meng Zhang 0010, Longxing Shi |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2015 | An Energy-Efficient Coarse-Grained Reconfigurable Processing Unit for Multiple-Standard Video DecodingabstractA coarse-grained reconfigurable processing unit (RPU) consisting of 16 ×16 multi-functional processing elements (PEs) interconnected by an area-efficient line-switched mesh connect (LSMC) routing is implemented on a 5.4 mm ×3.1 mm die in TSMC 65 nm LP1P8M CMOS technology. A hierarchical configuration context (HCC) organization scheme is proposed to reduce the implementation overhead and the energy dissipation spent on fast reconfiguration. The proposed RPU is integrated into two system-on-a-chips (SoCs), targeting multiple-standard video decoding. The high-performance chip, comprising two RPU processors (named REMUS_HPP), can decode 1920 ×1080 H.264 video streams at 30 frames per second (fps) under 200 MHz. REMUS_HPP achieves a 25% performance gain over the XPP-III reconfigurable processor with only 280 mW power consumption, resulting in a 14.3 × improvement on energy efficiency. The other chip (named REMUS_LPP), targeting low power applications, integrates only one RPU processor. REMUS_LPP can decode 720 ×480 H.264 video streams at 35fps with 24.5 mW under 75 MHz, achieving a 76% reduction in power dissipation and a 3.96 × improvement on energy efficiency compared with the ADRES reconfigurable processor. Leibo Liu, Dong Wang 0040, Min Zhu 0001, Yansheng Wang, Shouyi Yin, Peng Cao 0002, Jun Yang 0006, Shaojun Wei |
IEEE Trans. Multim. | 7 |
| 2015 | Correction to "An Energy-Efficient Coarse-Grained Reconfigurable Processing Unit for Multiple-Standard Video Decoding"
Leibo Liu, Dong Wang 0040, Min Zhu 0001, Yansheng Wang, Shouyi Yin, Peng Cao 0002, Jun Yang 0006, Shaojun Wei |
IEEE Trans. Multim. | 7 |
| 2014 | A Side-channel Analysis Resistant Reconfigurable Cryptographic Coprocessor Supporting Multiple Block Cipher AlgorithmsabstractA side-channel analysis resistant reconfigurable cryptographic coprocessor is designed and fabricated in 0.18μm CMOS with 1.8V supply and 100MHz frequency, supporting multiple block cipher algorithms of AES, DES, RC6 and IDEA. Our countermeasure utilizes idle processing elements existed in reconfigurable array to do dummy operations to hide leakage information. This method has little impact on area and frequency, and it is flexible after silicon. It resists SPA and DPA without distinguishing the encryption region. And by correlation-based electromagnetic analysis, measurement to disclosure of DES enhances 36 times with partial countermeasures and AES discloses no subkey after more than one million electromagnetic traces with full countermeasures. Weiwei Shan, Longxing Shi, Xingyuan Fu, Chaoxuan Tian, Jun Yang 0006, Jie Li 0057 |
DAC | 7 |
| 2014 | On-Chip Memory Hierarchy in One Coarse-Grained Reconfigurable Architecture to Compress Memory Space and to Reduce Reconfiguration Time and Data-Reference TimeabstractThe coarse-grained reconfigurable architecture (CGRA) is proven to be energy efficient in several specific domains. In CGRAs, the on-chip memory hierarchy, which contains the context memory and the data memory organizations, should be well considered to achieve appropriate tradeoffs among three aspects: 1) performance; 2) area; and 3) power. In this paper, two techniques called the hierarchical configuration context (HCC) and the lifetime-based data-memory organization (LDO) focusing on the context memory and the data memory organizations are proposed to compress the on-chip memory space and to reduce the reconfiguration time and the data-reference time. In the HCC, the contexts are constructed in a hierarchical fashion to completely eliminate the repetitive portions of the contexts, not only reducing the overall context storage, but also alleviating the context transportation overhead. A fast context-indexing mechanism in the HCC is proposed to achieve fast reconfiguration, as the hierarchically organized contexts can be located and accessed conveniently. In the LDO, the on-chip data are classified into two types, based on the lifetime of data. The short-lifetime data are stored in the first in first out to increase the reuse ratio of memory space automatically, whereas the long-lifetime data are stored in the radom access memory for several time references. The HCC and the LDO are used in a CGRA core called as reconfigurable processing unit (RPU). Two RPUs are integrated in a reconfigurable computing processor (RCP) called as REconfigurable MUlti-media System, High-Performance Processor (REMUS_HPP). Because of the HCC, compared with a traditional nonhierarchical system, the total context storage required in H.264 decoding is reduced by 77%. Because of the LDO, the normalized on-chip data memory size at same performance level in the REMUS_HPP is only 23.8% and 14.8% of those in XPP-III (a high-performance RCP) and ADRES (a low-power RCP). REMUS_HPP is implemented on a 48.9-mm2silicon with TSMC 65-nm technology, using a 200-MHz working frequency to achieve 1920 × 1088 at 30 fps H.264 high-profile decoding. Compared with XPP-III, the performance of the REMUS_HPP is 1.81× boosted, whereas the energy efficiency is 4.75× higher. Yansheng Wang, Leibo Liu, Shouyi Yin, Min Zhu 0001, Peng Cao 0002, Jun Yang 0006, Shaojun Wei |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2013 | Hierarchical representation of on-chip context to reduce reconfiguration time and implementation area for coarse-grained reconfigurable architecture
Yansheng Wang, Leibo Liu, Shouyi Yin, Min Zhu 0001, Peng Cao 0002, Jun Yang 0006, Shaojun Wei |
Sci. China Inf. Sci. | 6 |
| 2011 | A Fast Locking All-Digital Phase-Locked Loop via Feed-Forward Compensation TechniqueabstractA fast locking all-digital phase-locked loop (ADPLL) via feed-forward compensation technique is proposed in this paper. The implemented ADPLL has two operation modes which are frequency acquisition mode and phase acquisition mode. In frequency acquisition mode, the ADPLL achieves a fast frequency locking via the proposed feed-forward compensation algorithm. In phase acquisition mode, the ADPLL achieves a finer phase locking. To verify the proposed algorithm and architecture, the ADPLL design is implemented by SMIC 0.18-μm 1P6M CMOS technology. The core size of the ADPLL is 582.2 μm * 343 μm. The frequency range of the ADPLL is from 4 to 416 MHz. The measurement results show that the ADPLL can achieve a frequency locking in two reference cycles when locking to 376 MHz. The corresponding power consumption is 11.394 mW. Xin Chen 0039, Jun Yang 0006, Longxing Shi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | Area-efficient line-based two-dimensional discrete wavelet transform architecture without data bufferabstractAn area-efficient architecture for 2D DWT is proposed in this paper based on novel decomposed lifting scheme, where no data buffer is required to preserve and reorder the intermediate data between the row and column processor. Compared with the reported research, the proposed design could benefit from the reduction of internal memory size and the number of multipliers, adders and registers. The design was implemented for 2D 9/7 and 5/3 DWT in SMIC 0.18 mum CMOS logic fabrication with 15 K equivalent 2-input NAND gates under 150 MHz, which can accommodate up to 512times512 image size with 4 K bytes on-chip dual-port RAM. Peng Cao 0002, Chao Wang 0068, Jun Yang 0006, Longxing Shi |
ICME | 3 |