VLDB 2026 Research / reviewers in the wild / expert
Qiang Xu 0001
dblp:43/1230-1
· DBLP profile ↗
229ranked-venue papers
21as first author
74since 2021 · last 2026
0000-0001-6747-126XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 169 · 21 first-author · 26 since 2021Artificial intelligence and machine learning · 44 · 40 since 2021Software engineering, systems software and programming languages · 28 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 21 since 2021Security and privacy · 7 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DynamicRTL: RTL Representation Learning for Dynamic Circuit BehaviorabstractThere is a growing body of work on using Graph Neural Networks (GNNs) to learn representations of circuits, focusing primarily on their static characteristics. However, these models fail to capture circuit runtime behavior, which is crucial for tasks like circuit verification and optimization. To address this limitation, we introduce DR-GNN (DynamicRTL-GNN), a novel approach that learns RTL circuit representations by incorporating both static structures and multi-cycle execution behaviors. DR-GNN leverages an operator-level Control Data Flow Graph (CDFG) to represent Register Transfer Level (RTL) circuits, enabling the model to capture dynamic dependencies and runtime execution. To train and evaluate DR-GNN, we build the first comprehensive dynamic circuit dataset, comprising over 6,300 Verilog designs and 63,000 simulation traces. Our results demonstrate that DR-GNN outperforms existing models in branch hit prediction and toggle rate prediction. Furthermore, its learned representations transfer effectively to related dynamic circuit tasks, achieving strong performance in power estimation and assertion prediction. Yunhao Zhou, Yi Liu 0081, Zhengyuan Shi, Lingwei Yan, Gang Chen 0023, Qiang Xu 0001, Guojie Luo |
AAAI | 11 |
| 2026 | FIXME: Towards End-to-End Benchmarking of LLM-Aided Design VerificationabstractDespite the transformative potential of Large Language Models (LLMs) in hardware design, a comprehensive evaluation of their capabilities in design verification remains underexplored. Current efforts predominantly focus on RTL generation and basic debugging, overlooking the critical domain of functional verification, which is the primary bottleneck in modern design methodologies due to the rapid escalation of hardware complexity. We present FIXME, the first end-to-end, multi-model, and open-source evaluation framework for assessing LLM performance in hardware functional verification (FV) to address this crucial gap. FIXME introduces a structured three-level difficulty hierarchy spanning six verification sub-domains and 180 diverse tasks, enabling in-depth analysis across the design lifecycle. Leveraging a collaborative AI-human approach, we construct a high-quality dataset using 100% silicon-proven designs, ensuring comprehensive coverage of real-world challenges. Furthermore, we enhance the functional coverage by 45.57% through expert-guided optimization. By rigorously evaluating state-of-the-art LLMs such as GPT-4, Claude3, and LlaMA3, we identify key areas for improvement and outline promising research directions to unlock the full potential of LLM-driven automation in hardware design verification. The benchmark is available at https://github.com/ChatDesignVerification/FIXME. Gwok-Waa Wan, Sam-Zaak Wong, Shengchu Su, Chenxu Niu 0001, Ning Wang 0071, Xinlai Wan, Qixiang Chen, Mengnv Xing, Jianmin Ye, Rongchang Song, Qiang Xu 0001, Nan Guan, Zhe Jiang 0004, Xi Wang 0009, Yong Chen 0001, Jun Yang 0006 |
AAAI | 14 |
| 2026 | Multi-Faceted Attack: Exposing Cross-Model Vulnerabilities in Defense-Equipped Vision-Language ModelsabstractThe growing misuse of Vision-Language Models (VLMs) has led providers to deploy multiple safeguards—alignment tuning, system prompt, and content moderation. Yet the real-world robustness of these defenses against adversarial attack remains underexplored. We introduce Multi-Faceted Attack (MFA), a framework that systematically uncovers general safety vulnerabilities in leading defense-equipped VLMs, including GPT-4o, Gemini-Pro, and LLaMA 4, etc. Central to MFA is the Attention-Transfer Attack (ATA), which conceals harmful instructions inside a meta task with competing objectives. We offer a theoretical perspective grounded in reward-hacking to explain why such an attack can succeed. To maximize cross-model transfer, we introduce a lightweight transfer-enhancement algorithm combined with a simple repetition strategy that jointly evades both input- and output-level filters—without any model-specific fine-tuning. We empirically show that adversarial images optimized for one vision encoder transfer broadly to unseen VLMs, indicating that shared visual representations create a cross-model safety vulnerability. Combined, MFA reaches a 58.5% overall attack success rate, consistently outperforming existing methods. Notably, on state-of-the-art commercial models, MFA achieves a 52.8% success rate, outperforming the second-best attack by 34%. These findings challenge the perceived robustness of current defensive mechanisms, systematically expose general safety loopholes within defense-equipped VLMs, and offer a practical probe for diagnosing and evaluating the safety of VLMs. Chi Harold Liu, Lanqing Hong, Qiang Xu 0001 |
AAAI | 6 |
| 2026 | Solve-Detect-Verify: Inference-Time Scaling with Flexible Generative VerifierabstractFigure 1: The Solve-Detect-Verify (SDV) pipeline transforms linguistic signals into efficiency.Left: On AIME 2024, SDV achieves 83.3% accuracy (vs.63.3% for GenPRM) while using 6x fewer verification tokens by pruning redundant reasoning.Right: The pipeline is powered by FlexiVe , a unified verifier.Unlike process-based verifiers that incur per-step overhead, FlexiVe analyzes traces holistically.It employs a "pragmatic" consensus strategy: parallel "Fast Thinking" checks (∼0.1k tokens) provide an initial semantic intuition, escalating to deliberative "Slow Thinking" (∼4k tokens) only when the model exhibits verbalized uncertainty. Jianyuan Zhong, Zeju Li, Xiangyu Wen 0001, Kezhi Li, Qiang Xu 0001 |
ACL (1) | 6 |
| 2026 | AC-Refiner: Efficient Arithmetic Circuit Optimization Using Conditional Diffusion Models
Chenhao Xue, Kezhi Li, Zhengyuan Shi, Chen Zhang 0001, Yibo Lin, Lining Zhang, Qiang Xu 0001, Guangyu Sun 0003 |
ASP-DAC | 9 |
| 2026 | DeepCut: Structure-Aware GNN Framework for Efficient Cut Timing Prediction in Logic Synthesis
Lingfeng Zhou, Yilong Zhou, Zhengyuan Shi, Qiang Xu 0001, Zhufei Chu |
ASP-DAC | 5 |
| 2026 | MagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street ScenesabstractControllable generative models for images and videos have seen significant success, yet 3D scene generation, especially in unbounded scenarios like autonomous driving, remains underdeveloped. Existing methods lack flexible controllability and often rely on dense view data collection in controlled environments, limiting their generalizability across common datasets (e.g., nuScenes). In this paper, we introduce MagicDrive3D, a novel framework for controllable 3D street scene generation that combines video-based view synthesis with 3D representation (3DGS) generation. It supports multi-condition control, including road maps, 3D objects, and text descriptions. Unlike previous approaches that require 3D representation before training, MagicDrive3D first trains a multi-view video generation model to synthesize diverse street views. This method utilizes routinely collected autonomous driving data, reducing data acquisition challenges and enriching 3D scene generation. In the 3DGS generation step, we introduce Fault-Tolerant Gaussian Splatting to address minor errors and use monocular depth for better initialization, alongside appearance modeling to manage exposure discrepancies across viewpoints. Experiments show that MagicDrive3D generates diverse, high-quality 3D driving scenes, supports any-view rendering, and enhances downstream tasks like BEV segmentation, demonstrating its potential for autonomous driving simulation and beyond. Ruiyuan Gao 0001, Kai Chen 0023, Zhihao Li 0002, Lanqing Hong, Zhenguo Li, Qiang Xu 0001 |
WACV | 6 |
| 2026 | TNT: A Large-Scale P2P Botnet Detection Framework via Communication Topology and Network TrafficabstractWith the booming development of embedded systems and mobile networks, the attack surfaces of botnets are broadened and amplified. Especially recent adversaries tend to leverage peer-to-peer (P2P) manner propagation to construct large-scale botnets because P2P-based schemes eliminate single points of failure. Over the past few decades, the research and industry communities have proposed a variety of solutions to detect botnets, which mainly involve communication topology identification and network traffic analysis. Yet, coping with the large-scale P2P botnets, the former suffer topology indistinguishability, and the latter struggles under massive background traffic. In this paper, we present$\textsf {TNT}$, a large-scale P2P botnet detection framework via communication topology and network traffic. As its core,$\textsf {TNT}$is powered by three tightly-coupled components:$\textsf {(i)}$$\textsf {tScouter}$is responsible for profiling the communication topology;$\textsf {(ii)}$$\textsf {tCommander}$plans the strategy for node inspection; and$\textsf {(iii)}$$\textsf {tPatroller}$investigates the traffic of the corresponding node. Taken together,$\textsf {TNT}$advances the trade-off between detection accuracy (enhance topology-based results via traffic analysis) and overhead (only check part of node traffic according to the planning). Based on 42 groups of combinations involving 6 types of botnets and 7 legitimate P2P traffic, we perform extensive evaluation and demonstrate that$\textsf {TNT}$realizes outstanding detection performance,e.g.,after checking ~20K nodes, achieve ~99.9% accuracy for a communication graph (including >140K nodes). In addition, we develop the expansion experiments in terms of heterogeneous nodes and accuracy loss, as well as provide deep insights into interpretability from the aspect of the attribution matrix. Ziming Zhao 0008, Zhaoxuan Li, Tingting Li 0004, Yu Li 0007, Qiang Xu 0001, Fan Zhang 0010 |
IEEE Trans. Netw. | 6 |
| 2025 | MotionCraft: Crafting Whole-Body Motion with Plug-and-Play Multimodal ControlsabstractWhole-body multimodal motion generation, controlled by text, speech, or music, has numerous applications including video generation and character animation. However, employing a unified model to process different condition modalities presents two main challenges: motion distribution drifts across different tasks (e.g., co-speech gestures and text-driven daily actions) and the complex optimization of mixed conditions with varying granularities (e.g., text and audio). In this paper, we propose MotionCraft, a unified diffusion transformer that crafts whole-body motion with plug-and-play multimodal control. Our framework employs a coarse-to-fine training strategy, starting with the text-to-motion semantic pre-training, followed by the multimodal low-level control adaptation. To effectively learn and transfer motion knowledge across different distributions, we design MC-Attn for parallel modeling of static and dynamic human topology graphs. To overcome the motion format inconsistency of existing benchmarks, we introduce MC-Bench, the first available multimodal whole-body motion generation benchmark based on the unified SMPL-X format. Extensive experiments show that MotionCraft achieves state-of-the-art performance on various standard motion generation tasks. Yuxuan Bian, Ailing Zeng, Xuan Ju, Zhaoyang Zhang 0004, Wei Liu 0005, Qiang Xu 0001 |
AAAI | 7 |
| 2025 | DeepSeq2: Enhanced Sequential Circuit Learning with Disentangled RepresentationsabstractCircuit representation learning is increasingly pivotal in Electronic Design Automation (EDA), serving various downstream tasks with enhanced model efficiency and accuracy. One notable work, DeepSeq, has pioneered sequential circuit learning by encoding temporal correlations. However, it suffers from significant limitations including prolonged execution times and architectural inefficiencies. To address these issues, we introduce DeepSeq2, a novel framework that enhances the learning of sequential circuits, by innovatively mapping it into three distinct embedding spaces---structure, function, and sequential behavior---allowing for a more nuanced representation that captures the inherent complexities of circuit dynamics. By employing an efficient Directed Acyclic Graph Neural Network (DAG-GNN) that circumvents the recursive propagation used in DeepSeq, DeepSeq2 significantly reduces execution times and improves model scalability. Moreover, DeepSeq2 incorporates a unique supervision mechanism that captures transitioning behaviors within circuits more effectively. DeepSeq2 sets a new benchmark in sequential circuit representation learning, outperforming prior works in power estimation and reliability analysis. Sadaf Khan, Zhengyuan Shi, Min Li 0019, Qiang Xu 0001 |
ASP-DAC | 5 |
| 2025 | DynamicSAT: Dynamic Configuration Tuning for SAT Solving
Zhengyuan Shi, Xindi Zhang 0001, Yun Liang 0001, Zhufei Chu, Qiang Xu 0001 |
CP | 7 |
| 2025 | Late Breaking Results: Hybrid Logic Optimization with Predictive Self-SupervisionabstractHybrid optimization is an emerging approach in logic synthesis, focusing on applying diverse optimization methods to different parts of a logic circuit. This paper analyzes the relationship between each vertex and its corresponding optimization method. We extract a subgraph centered on each vertex and quantify the logic optimization results of these subgraphs as vertex features. Based on these features, we propose a circuit partitioning method to cluster the logic circuit, enabling the final optimized circuit to be constructed by merging clusters optimized with their respective methods. Additionally, we introduce a self-supervised prediction model to efficiently obtain vertex features. The experimental results targeting LUT mapping demonstrate that our method achieves improvements of $8.48 \%$ in area and 9.81% in delay compared to the state-of-the-art. Rongliang Fu, Zhengyuan Shi, Yuan Pu 0001, Junying Huang, Qiang Xu 0001, Tsung-Yi Ho |
DAC | 7 |
| 2025 | Logic Optimization Meets SAT: A Novel Framework for Circuit-SAT SolvingabstractThe Circuit Satisfiability (CSAT) problem, a variant of the Boolean Satisfiability (SAT) problem, plays a critical role in integrated circuit design and verification. However, existing SAT solvers, optimized for Conjunctive Normal Form (CNF), often struggle with the intrinsic complexity of circuit structures when directly applied to CSAT instances. To address this challenge, we propose a novel preprocessing framework that leverages advanced logic synthesis techniques and a reinforcement learning (RL) agent to optimize CSAT problem instances. The framework introduces a cost-customized Look-Up Table (LUT) mapping strategy that prioritizes solving efficiency, effectively transforming circuits into simplified forms tailored for SAT solvers. Our method achieves significant runtime reductions across diverse industrial-scale CSAT benchmarks, seamlessly integrating with state-of-the-art SAT solvers. Extensive experimental evaluations demonstrate up to $63 \%$ reduction in solving time compared to conventional approaches, highlighting the potential of EDAdriven innovations to advance SAT-solving capabilities. Zhengyuan Shi, Tiebing Tang, Jiaying Zhu, Sadaf Khan, Hui-Ling Zhen, Mingxuan Yuan, Zhufei Chu, Qiang Xu 0001 |
DAC | 8 |
| 2025 | Speculative Decoding for Verilog: Speed and Quality, All in OneabstractThe rapid advancement of large language models (LLMs) has revolutionized code generation tasks across various programming languages. However, the unique characteristics of programming languages, particularly those like Verilog with specific syntax and lower representation in training datasets, pose significant challenges for conventional tokenization and decoding approaches. In this paper, we introduce a novel application of speculative decoding for Verilog code generation, showing that it can improve both inference speed and output quality, effectively achieving speed and quality all in one. Unlike standard LLM tokenization schemes, which often fragment meaningful code structures, our approach aligns decoding stops with syntactically significant tokens, making it easier for models to learn the token distribution. This refinement addresses inherent tokenization issues and enhances the model’s ability to capture Verilog’s logical constructs more effectively. Our experimental results show that our method achieves up to a $5.05 \times$ speedup in Verilog code generation and increases pass@10 functional accuracy on RTLLM by up to $\mathbf{1 7. 1 9 \%}$ compared to conventional training strategies. These findings highlight speculative decoding as a promising approach to bridge the quality gap in code generation for specialized programming languages. Changran Xu, Yi Liu 0081, Yunhao Zhou, Shan Huang 0010, Ningyi Xu, Qiang Xu 0001 |
DAC | 6 |
| 2025 | WideGate: Beyond Directed Acyclic Graph Learning in Subcircuit Boundary PredictionabstractSubcircuit boundary prediction is an important application of machine learning in logical analysis, effectively supporting tasks such as functional verification and logic optimization. Existing methods often convert circuits into and-inverter graphs and then use directed acyclic graph neural networks to perform this task. However, two key characteristics of subcircuit boundary prediction do not align with the fundamental assumptions of directed acyclic graph (DAG) learning, which limits the model's expressiveness and generalization capabilities. To break these assumptions, we propose WideGate, which includes a receptive field generation module that extends beyond the fanin cone and fanout cone, as well as an adaptive aggregation module that focuses on boundaries. Extensive experiments show that WideGate significantly outperforms existing methods in terms of prediction accuracy and training efficiency for sub circuit boundary prediction. The code is available at https://github.com/BUPT-GAMMA/WideGate. Jiawei Liu 0006, Zhiyan Liu, Jianwang Zhai, Zhengyuan Shi, Qiang Xu 0001, Bei Yu 0001, Chuan Shi 0001 |
DATE | 6 |
| 2025 | Dyve: Thinking Fast and Slow for Dynamic Process VerificationabstractLarge Language Models (LLMs) have advanced significantly in complex reasoning, often leveraging external verifiers to improve multi-step process reliability.However, existing process verification methods face critical limitations: discriminative Process Reward Models (PRMs) often provide overly simplistic binary feedback and struggle with incomplete reasoning traces, while sophisticated Generative Reward Models (GenRMs) can be computationally expensive.Furthermore, curating quality supervision data for process verifier is of challenging.Therefore, we present Dyve, a dynamic process verifier that enhances reasoning error detection in LLMs by integrating fast (System 1) and slow (System 2) thinking, inspired by Kahneman's Systems Theory.Dyve adaptively applies immediate token-level confirmation for straightforward steps and comprehensive analysis for complex ones.To address data challenges and enable its adaptive fast and slow thinking, Dyve employs a novel step-wise consensus-filtered supervision strategy.This strategy leverages Monte Carlo estimation, LLM-as-a-Judge, and specialized reasoning models to extract the high-quality training signals from noisy rollouts.Experimental results on ProcessBench and the MATH dataset confirm that Dyve significantly outperforms existing process-based verifiers and boosts performance in Best-of-N settings, while maintaining computational efficiency through strategic resource allocation.Our code, data and model are released at: https://github.com/ staymylove/ Jianyuan Zhong, Zeju Li, Xiangyu Wen 0001, Qiang Xu 0001 |
EMNLP | 5 |
| 2025 | DeepCell: Self-Supervised Multiview Fusion for Circuit Representation LearningabstractWe introduce DeepCell, a novel circuit representation learning framework that effectively integrates multiview information from both And-Inverter Graphs (AIGs) and Post-Mapping (PM) netlists. At its core, DeepCell employs a self-supervised Mask Circuit Modeling (MCM) strategy, inspired by masked language modeling, to fuse complementary circuit representations from different design stages into unified and rich embeddings. To our knowledge, DeepCell is the first framework explicitly designed for PM netlist representation learning, setting new benchmarks in both predictive accuracy and reconstruction quality. We demonstrate the practical efficacy of DeepCell by applying it to critical EDA tasks such as functional Engineering Change Orders (ECO) and technology mapping. Extensive experimental results show that DeepCell significantly surpasses state-of-the-art open-source EDA tools in efficiency and performance. The code is available at https://github.com/cure-lab/DeepCell. Zhengyuan Shi, Chengyu Ma, Lingfeng Zhou, Hongyang Pan, Fan Yang 0001, Zhufei Chu, Qiang Xu 0001 |
ICCAD | 10 |
| 2025 | Revolution or Hype? Seeking the Limits of Large Models in Hardware DesignabstractRecent breakthroughs in Large Language Models (LLMs) and Large Circuit Models (LCMs) have sparked excitement across the electronic design automation (EDA) community, promising a revolution in circuit design and optimization. Yet, this excitement is met with significant skepticism: Are these AI models a genuine revolution in circuit design, or a temporary wave of inflated expectations? This paper serves as a foundational text for the corresponding ICCAD 2025 panel, bringing together perspectives from leading experts in academia and industry. It critically examines the practical capabilities, fundamental limitations, and future prospects of large AI models in hardware design. The paper synthesizes the core arguments surrounding reliability, scalability, and interpretability, framing the debate on whether these models can meaningfully outperform or complement traditional EDA methods. The result is an authoritative overview offering fresh insights into one of today’s most contentious and impactful technology trends. Qiang Xu 0001, Leon Stok, Rolf Drechsler, Xi Wang 0009, Grace Li Zhang, Igor L. Markov |
ICCAD | 1 |
| 2025 | MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMsabstractThe emergence of multimodal large language models (MLLMs) presents promising opportunities for automation and enhancement in Electronic Design Automation (EDA). However, comprehensively evaluating these models in circuit design remains challenging due to the narrow scope of existing benchmarks. To bridge this gap, we introduce MMCircuitEval, the first multimodal benchmark specifically designed to assess MLLM performance comprehensively across diverse EDA tasks. MMCircuitEval comprises 3614 meticulously curated question-answer (QA) pairs spanning digital and analog circuits across critical EDA stages—ranging from general knowledge and specifications to front-end and back-end design. Derived from textbooks, technical question banks, datasheets, and real-world documentation, each QA pair undergoes rigorous expert review for accuracy and relevance. Our benchmark uniquely categorizes questions by design stage, circuit type, tested abilities (knowledge, comprehension, reasoning, computation), and difficulty level, enabling detailed analysis of model capabilities and limitations. Extensive evaluations reveal significant performance gaps among existing LLMs, particularly in back-end design and complex computations, highlighting the critical need for targeted training datasets and modeling approaches. MMCircuitEval provides a foundational resource for advancing MLLMs in EDA, facilitating their integration into real-world circuit design workflows. Our benchmark is available at https://github.com/cure-lab/MMCircuitEval. Chenchen Zhao 0001, Zhengyuan Shi, Xiangyu Wen 0001, Yi Liu 0081, Yunhao Zhou, Hefei Feng, Yinan Zhu, Gwok-Waa Wan, Yongqi Fu, Chujie Chen, Chenhao Xue, Ying Wang 0001, Yibo Lin, Jun Yang 0006, Ning Xu 0009, Xi Wang 0009, Qiang Xu 0001 |
ICCAD | 21 |
| 2025 | MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control
Ruiyuan Gao 0001, Kai Chen 0023, Lanqing Hong, Zhenguo Li, Qiang Xu 0001 |
ICCV | 6 |
| 2025 | FullDiT: Video Generative Foundation Models with Multimodal Control via Full Attention
Xuan Ju, Weicai Ye, Quande Liu, Qiulin Wang, Xintao Wang 0002, Pengfei Wan 0001, Di Zhang 0026, Kun Gai, Qiang Xu 0001 |
ICCV | 9 |
| 2025 | HiBug2: Efficient and Interpretable Error Slice Discovery for Comprehensive Model DebuggingabstractDespite the significant success of deep learning models in computer vision, they often exhibit systematic failures on specific data subsets, known as error slices. Identifying and mitigating these error slices is crucial to enhancing model robustness and reliability in real-world scenarios. In this paper, we introduce HiBug2, an automated framework for error slice discovery and model repair. HiBug2 first generates task-specific visual attributes to highlight instances prone to errors through an interpretable and structured process. It then employs an efficient slice enumeration algorithm to systematically identify error slices, overcoming the combinatorial challenges that arise during slice exploration. Additionally, HiBug2 extends its capabilities by predicting error slices beyond the validation set, addressing a key limitation of prior approaches. Extensive experiments across multiple domains — including image classification, pose estimation, and object detection — show that HiBug2 not only improves the coherence and precision of identified error slices but also significantly enhances the model repair capabilities. Muxi Chen, Chenchen Zhao 0001, Qiang Xu 0001 |
ICLR | 3 |
| 2025 | DeepRTL: Bridging Verilog Understanding and Generation with a Unified Representation ModelabstractRecent advancements in large language models (LLMs) have shown significant potential for automating hardware description language (HDL) code generation from high-level natural language instructions. While fine-tuning has improved LLMs' performance in hardware design tasks, prior efforts have largely focused on Verilog generation, overlooking the equally critical task of Verilog understanding. Furthermore, existing models suffer from weak alignment between natural language descriptions and Verilog code, hindering the generation of high-quality, synthesizable designs. To address these issues, we present DeepRTL, a unified representation model that excels in both Verilog understanding and generation. Based on CodeT5+, DeepRTL is fine-tuned on a comprehensive dataset that aligns Verilog code with rich, multi-level natural language descriptions.
We also introduce the first benchmark for Verilog understanding and take the initiative to apply embedding similarity and GPT Score to evaluate the models' understanding capabilities. These metrics capture semantic similarity more accurately than traditional methods like BLEU and ROUGE, which are limited to surface-level n-gram overlaps. By adapting curriculum learning to train DeepRTL, we enable it to significantly outperform GPT-4 in Verilog understanding tasks, while achieving performance on par with OpenAI's o1-preview model in Verilog generation tasks. Yi Liu 0081, Changran Xu, Yunhao Zhou, Zeju Li, Qiang Xu 0001 |
ICLR | 5 |
| 2025 | DeepGate4: Efficient and Effective Representation Learning for Circuit Design at ScaleabstractCircuit representation learning has become pivotal in electronic design automation, enabling critical tasks such as testability analysis, logic reasoning, power estimation, and SAT solving. However, existing models face significant challenges in scaling to large circuits due to limitations like over-squashing in graph neural networks and the quadratic complexity of transformer-based models. To address these issues, we introduce \textbf{DeepGate4}, a scalable and efficient graph transformer specifically designed for large-scale circuits. DeepGate4 incorporates several key innovations: (1) an update strategy tailored for circuit graphs, which reduce memory complexity to sub-linear and is adaptable to any graph transformer; (2) a GAT-based sparse transformer with global and local structural encodings for AIGs; and (3) an inference acceleration CUDA kernel that fully exploit the unique sparsity patterns of AIGs. Our extensive experiments on the ITC99 and EPFL benchmarks show that DeepGate4 significantly surpasses state-of-the-art methods, achieving 15.5\% and 31.1\% performance improvements over the next-best models. Furthermore, the Fused-DeepGate4 variant reduces runtime by 35.1\% and memory usage by 46.8\%, making it highly efficient for large-scale circuit analysis. These results demonstrate the potential of DeepGate4 to handle complex EDA tasks while offering superior scalability and efficiency. Shan Huang 0010, Jianyuan Zhong, Zhengyuan Shi, Guohao Dai 0001, Ningyi Xu, Qiang Xu 0001 |
ICLR | 7 |
| 2025 | DeepLayout: Learning Neural Representations of Circuit Placement LayoutabstractRecent advancements have integrated various deep-learning methodologies into physical design, aiming for workflows acceleration and surpasses human-devised solutions. However, prior research has primarily concentrated on developing task-specific networks, which necessitate a significant investment of time to construct large, specialized datasets, and the unintended isolation of models across different tasks. In this paper, we introduce DeepLayout, the first general representation learning framework specifically designed for backend circuit design. To address the distinct characteristics of post-placement circuits, including topological connectivity and geometric distribution, we propose a hybrid encoding architecture that integrates GNN with spatial transformers. Additionally, the framework includes a flexible decoder module that accommodates a variety of task types, supporting multiple hierarchical outputs such as nets and layouts. To mitigate the high annotation costs associated with layout data, we introduce a mask-based self-supervised learning approach designed explicitly for layout representation. This strategy involves a carefully devised masking approach tailored to layout features, precise reconstruction guidance, and most critically—two key supervised learning tasks. We conduct extensive experiments on large-scale industrial datasets, demonstrating that DeepLayout surpasses state-of-the-art (SOTA) methods specialized for individual tasks on two crucial layout quality assessment benchmarks. The experiment results underscore the framework’s robust capability to learn the intrinsic properties of circuits. Zhuomin Chai, Xun Jiang 0002, Qiang Xu 0001, Runsheng Wang, Yibo Lin |
ICML | 4 |
| 2025 | Dependency Matters: Enhancing LLM Reasoning with Explicit Knowledge GroundingabstractLarge language models (LLMs) often produce reasoning steps that are superficially coherent yet internally inconsistent, leading to unreliable outputs. Since such failures typically arise from implicit or poorly-grounded knowledge, we introduce \emph{Grounded Reasoning in Dependency (GRiD)}, a novel dependency-aware reasoning framework that explicitly grounds reasoning steps in structured knowledge. GRiD represents reasoning as a graph consisting of interconnected knowledge extraction nodes and reasoning nodes, enforcing logical consistency through explicit dependencies. Each reasoning step is validated via a lightweight, step-wise verifier that ensures logical correctness relative to its premises. Extensive experiments across diverse reasoning benchmarks—including StrategyQA, CommonsenseQA, GPQA, and TruthfulQA—demonstrate that GRiD substantially improves reasoning accuracy, consistency, and faithfulness compared to recent state-of-the-art structured reasoning methods. Notably, GRiD enhances performance even when applied purely as a lightweight verification module at inference time, underscoring its generalizability and practical utility. Code is available at: https://github.com/cure-lab/GRiD. Xiangyu Wen 0001, Min Li 0019, Junhua Huang, Jianyuan Zhong, Zeju Li, Yongxiang Huang, Mingxuan Yuan, Qiang Xu 0001 |
NeurIPS | 9 |
| 2025 | Functional Matching of Logic Subgraphs: Beyond Structural IsomorphismabstractSubgraph matching in logic circuits is foundational for numerous Electronic Design Automation (EDA) applications, including datapath optimization, arithmetic verification, and hardware trojan detection. However, existing techniques rely primarily on structural graph isomorphism and thus fail to identify function-related subgraphs when synthesis transformations substantially alter circuit topology. To overcome this critical limitation, we introduce the concept of functional subgraph matching, a novel approach that identifies whether a given logic function is implicitly present within a larger circuit, irrespective of structural variations induced by synthesis or technology mapping. Specifically, we propose a two-stage multi-modal framework: (1) learning robust functional embeddings across AIG and post-mapping netlists for functional subgraph detection, and (2) identifying fuzzy boundaries using a graph segmentation approach. Evaluations on standard benchmarks (ITC99, OpenABCD, ForgeEDA) demonstrate significant performance improvements over existing structural methods, with average 93.8% accuracy in functional subgraph detection and a dice score of 91.3% in fuzzy boundary identification. Kezhi Li, Zhengyuan Shi, Qiang Xu 0001 |
NeurIPS | 4 |
| 2025 | Non-Cross Diffusion for Semantic ConsistencyabstractIn diffusion models, deviations from a straight generative flow are a common issue, resulting in semantic inconsistencies and suboptimal generations. To address this challenge, we introduce ‘Non-Cross Diffusion’, an innovative approach in generative modeling for learning ordinary differential equation (ODE) models. Our methodology strategically incorporates an ascending dimension of input to effectively connect points sampled from two distributions with uncrossed paths. This design is pivotal in ensuring enhanced semantic consistency throughout the inference process, which is especially critical for applications reliant on consistent generative flows, including various distillation methods and deterministic sampling, which are fundamental in image editing and interpolation tasks. Our empirical results demonstrate the effectiveness of Non-Cross Diffusion, showing a substantial improvements in semantic consistencies at various inference steps and enhancing the overall performance of diffusion models. Ruiyuan Gao 0001, Qiang Xu 0001 |
WACV | 3 |
| 2025 | FGNN2: A Powerful Pretraining Framework for Learning the Logic Functionality of CircuitsabstractLearning feasible representation from raw gate-level circuits is essential for incorporating machine learning techniques in logic synthesis, physical design, or verification. Existing structure-based learning methods tend to concentrate mainly on the graph topology, often neglecting logic functionality. This oversight frequently results in a failure to capture the underlying semantics, thereby limiting their overall applicability. To address the concern, we propose a novel circuit representation learning framework, FGNN2, that utilizes a contrastive scheme to effectively extract generic functionality knowledge. We construct a comprehensive pretraining dataset through a customized circuit augmentation scheme. We have also developed a novel contrastive loss function to capture the relative functional distance between different circuits, and to generate representations that are invariant to the input order. In addition, we employed a customized graph neural network (GNN) architecture to better align with the above framework. Comprehensive experiments on the multiple complex real-world designs demonstrate that our proposed solution significantly outperforms the state-of-the-art circuit representation learning flows. Ziyi Wang 0010, Zhuolun He, Guangliang Zhang, Qiang Xu 0001, Tsung-Yi Ho, Yu Huang 0005, Bei Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | Vision Transformer Acceleration via a Versatile Attention Optimization FrameworkabstractVision Transformers (ViTs) have achieved remarkable success across various tasks. However, their deployment is hindered by challenges, such as high memory requirements, long inference latency, and significant power consumption. To address these challenges, existing works optimize one of the two key stages on ViT’s critical path: linear projection or self-attention. Regrettably, we have noticed that both linear projection and self-attention can potentially become bottlenecks as the input image resolution varies, which makes the existing approaches lack generality. Accordingly, in this article, we propose a versatile attention optimization framework. On the algorithm side, we present a SpQuant algorithm that sparsifies weight matrices offline and input matrices online during linear projection as well as tunes the bit-width of the probabilities matrix according to their importance. On the hardware side, we design SQArch architectures to improve the performance of the SpQuant algorithm. The proposed SQArch architecture offers a low-cost preprocess module that predicts and prunes nonkey elements of the input matrix on the fly. Moreover, we design a compute module that supports sparse-sparse matrix multiplications (SpMSpM) and multiple precision computations on a single systolic array for generality. Furthermore, we can address the underutilization and workload imbalance problems by 1) decoupling the rows in the systolic array for enough flexibility and 2) proposing a workload balance scheme for SpMSpM that allows the array to accept data of similar sparsity, thereby reducing synchronization between computing units. Extensive experiment results demonstrate that SQArch can achieve satisfactory performance speedups and energy saving compared to state-of-the-art designs. Xuhang Wang, Qiyue Huang, Xing Li 0031, Haozhe Jiang, Qiang Xu 0001, Xiaoyao Liang, Zhuoran Song |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2024 | DetDiffusion: Synergizing Generative and Perceptive Models for Enhanced Data Generation and PerceptionabstractCurrent perceptive models heavily depend on resource-intensive datasets, prompting the need for innovative solutions. Leveraging recent advances in diffusion models, synthetic data, by constructing image inputs from various annotations, proves beneficial for downstream tasks. While prior methods have separately addressed generative and perceptive models, DetDiffusion, for the first time, harmonizes both, tackling the challenges in generating effective data for perceptive models. To enhance image generation with perceptive models, we introduce perception-aware loss (P.A. loss) through segmentation, improving both quality and controllability. To boost the performance of specific perceptive models, our method customizes data augmentation by extracting and utilizing perception-aware attribute (P.A. Attr) during generation. Experimental results from the object detection task highlight DetDiffusion's superior performance, establishing a new state-of-the-art in layout-guided generation. Furthermore, image syntheses from DetDiffusion can effectively augment training data, significantly enhancing downstream detection performance. Yibo Wang 0039, Ruiyuan Gao 0001, Kai Chen 0023, Kaiqiang Zhou, Yingjie Cai, Lanqing Hong, Zhenguo Li, Lihui Jiang, Dit-Yan Yeung, Qiang Xu 0001, Kai Zhang 0008 |
CVPR | 10 |
| 2024 | MMA-Diffusion: MultiModal Attack on Diffusion ModelsabstractIn recent years, Text-to-Image (T2I) models have seen remarkable advancements, gaining widespread adoption. However, this progress has inadvertently opened avenues for potential misuse, particularly in generating inappropriate or Not-Safe-For-Work (NSFW) content. Our work introduces MMA-Diffusion, a framework that presents a significant and realistic threat to the security of T2I models by effectively circumventing current defensive measures in both open-source models and commercial online services. Unlike previous approaches, MMA-Diffusion leverages both textual and visual modalities to bypass safeguards like prompt filters and post-hoc safety checkers, thus exposing and highlighting the vulnerabilities in existing defense mechanisms. Our codes are available at https://github.com/cure-lab/MMA-Diffusion. Ruiyuan Gao 0001, Xiaosen Wang, Tsung-Yi Ho, Nan Xu 0004, Qiang Xu 0001 |
CVPR | 6 |
| 2024 | DeepSeq: Deep Sequential Circuit LearningabstractIn this work, we propose DeepSeq, a novel representation learning framework for sequential netlists. It employs a graph neural network (GNN) with customized propagation to capture temporal correlations. To ensure effective learning, we propose a multi-task training objective with two sets of strongly related supervision: logic probability and transition probability at each logic gate. A novel dual attention aggregation mechanism is introduced to facilitate learning both tasks efficiently. Experimental results validate DeepSeq's superiority over other GNN models in sequential circuit learning. It demonstrates accurate reliability and power estimation across diverse circuits and workloads. Sadaf Khan, Zhengyuan Shi, Min Li 0019, Qiang Xu 0001 |
DATE | 4 |
| 2024 | BrushNet: A Plug-and-Play Image Inpainting Model with Decomposed Dual-Branch Diffusion
Xuan Ju, Xintao Wang 0002, Yuxuan Bian, Ying Shan, Qiang Xu 0001 |
ECCV (20) | 6 |
| 2024 | OriGen: Enhancing RTL Code Generation with Code-to-Code Augmentation and Self-ReflectionabstractRecent studies have demonstrated the significant potential of Large Language Models (LLMs) in generating Register Transfer Level (RTL) code, with notable advancements showcased by commercial models such as GPT-4 and Claude3-Opus. However, these proprietary LLMs often raise concerns regarding privacy and security. While open-source LLMs offer solutions to these concerns, they typically underperform commercial models in RTL code generation tasks, primarily due to the scarcity of high-quality open-source RTL datasets. To address this challenge, we introduce OriGen, a fully open-source framework that incorporates self-reflection capabilities and a novel dataset augmentation methodology for generating high-quality, large-scale RTL code. Our approach employs a code-to-code augmentation technique to enhance the quality of open-source RTL code datasets. Furthermore, OriGen can rectify syntactic errors through a self-reflection process that leverages compiler feedback. Fan Cui, Chenyang Yin, Kexing Zhou, Youwei Xiao, Guangyu Sun 0003, Qiang Xu 0001, Qipeng Guo, Yun Liang 0001, Xingcheng Zhang, Demin Song, Dahua Lin |
ICCAD | 6 |
| 2024 | DeepGate3: Towards Scalable Circuit Representation LearningabstractCircuit representation learning has shown promising results in advancing the field of Electronic Design Automation (EDA). Existing models, such as DeepGate Family, primarily utilize Graph Neural Networks (GNNs) to encode circuit netlists into gate-level embeddings. However, the scalability of GNN-based models is fundamentally constrained by architectural limitations, impacting their ability to generalize across diverse and complex circuit designs. To address these challenges, we introduce DeepGate3, an enhanced architecture that integrates Transformer modules following the initial GNN processing. This novel architecture not only retains the robust gate-level representation capabilities of its predecessor, DeepGate2, but also enhances them with the ability to model subcircuits through a novel pooling transformer mechanism. DeepGate3 is further refined with multiple innovative supervision tasks, significantly enhancing its learning process and enabling superior representation of both gate-level and subcircuit structures. Our experiments demonstrate marked improvements in scalability and generalizability over traditional GNN-based approaches, establishing a significant step forward in circuit representation learning technology. Zhengyuan Shi, Sadaf Khan, Jianyuan Zhong, Min Li 0019, Qiang Xu 0001 |
ICCAD | 6 |
| 2024 | MagicDrive: Street View Generation with Diverse 3D Geometry ControlabstractRecent advancements in diffusion models have significantly enhanced the data synthesis with 2D control. Yet, precise 3D control in street view generation, crucial for 3D perception tasks, remains elusive. Specifically, utilizing Bird's-Eye View (BEV) as the primary condition often leads to challenges in geometry control (e.g., height), affecting the representation of object shapes, occlusion patterns, and road surface elevations, all of which are essential to perception data synthesis, especially for 3D object detection tasks. In this paper, we introduce MagicDrive, a novel street view generation framework, offering diverse 3D geometry controls including camera poses, road maps, and 3D bounding boxes, together with textual descriptions, achieved through tailored encoding strategies. Besides, our design incorporates a cross-view attention module, ensuring consistency across multiple camera views. With MagicDrive, we achieve high-fidelity street-view image & video synthesis that captures nuanced 3D geometry and various scene descriptions, enhancing tasks like BEV segmentation and 3D object detection. Project Website: https://flymin.github.io/magicdrive Ruiyuan Gao 0001, Kai Chen 0023, Enze Xie, Lanqing Hong, Zhenguo Li, Dit-Yan Yeung, Qiang Xu 0001 |
ICLR | 7 |
| 2024 | PnP Inversion: Boosting Diffusion-based Editing with 3 Lines of CodeabstractText-guided diffusion models have revolutionized image generation and editing, offering exceptional realism and diversity. Specifically, in the context of diffusion-based editing, where a source image is edited according to a target prompt, the process commences by acquiring a noisy latent vector corresponding to the source image via the diffusion model. This vector is subsequently fed into separate source and target diffusion branches for editing. The accuracy of this inversion process significantly impacts the final editing outcome, influencing both essential content preservation of the source image and edit fidelity according to the target prompt.
Prior inversion techniques aimed at finding a unified solution in both the source and target diffusion branches. However, our theoretical and empirical analyses reveal that disentangling these branches leads to a distinct separation of responsibilities for preserving essential content and ensuring edit fidelity. Building on this insight, we introduce “PnP Inversion,” a novel technique achieving optimal performance of both branches with just three lines of code. To assess image editing performance, we present PIE-Bench, an editing benchmark with 700 images showcasing diverse scenes and editing types, accompanied by versatile annotations and comprehensive evaluation metrics. Compared to state-of-the-art optimization-based inversion techniques, our solution not only yields superior performance across 8 editing methods but also achieves nearly an order of speed-up. Xuan Ju, Ailing Zeng, Yuxuan Bian, Shaoteng Liu, Qiang Xu 0001 |
ICLR | 5 |
| 2024 | FITS: Modeling Time Series with 10k ParametersabstractIn this paper, we introduce FITS, a lightweight yet powerful model for time series analysis. Unlike existing models that directly process raw time-domain data, FITS operates on the principle that time series can be manipulated through interpolation in the complex frequency domain, achieving performance comparable to state-of-the-art models for time series forecasting and anomaly detection tasks. Notably, FITS accomplishes this with a svelte profile of just about $10k$ parameters, making it ideally suited for edge devices and paving the way for a wide range of applications. The code is available for review at: \url{https://anonymous.4open.science/r/FITS}. Ailing Zeng, Qiang Xu 0001 |
ICLR | 3 |
| 2024 | Multi-Patch Prediction: Adapting Language Models for Time Series Representation LearningabstractIn this study, we present $\text{aL\small{LM}4T\small{S}}$, an innovative framework that adapts Large Language Models (LLMs) for time-series representation learning. Central to our approach is that we reconceive time-series forecasting as a self-supervised, multi-patch prediction task, which, compared to traditional mask-and-reconstruction methods, captures temporal dynamics in patch representations more effectively. Our strategy encompasses two-stage training: (i). a causal continual pre-training phase on various time-series datasets, anchored on next patch prediction, effectively syncing LLM capabilities with the intricacies of time-series data; (ii). fine-tuning for multi-patch prediction in the targeted time-series context. A distinctive element of our framework is the patch-wise decoding layer, which departs from previous methods reliant on sequence-level decoding. Such a design directly transposes individual patches into temporal sequences, thereby significantly bolstering the model’s proficiency in mastering temporal patch-based representations. $\text{aL\small{LM}4T\small{S}}$ demonstrates superior performance in several downstream tasks, proving its effectiveness in deriving temporal representations with enhanced transferability and marking a pivotal advancement in the adaptation of LLMs for time-series analysis. Yuxuan Bian, Xuan Ju, Jiangtong Li, Dawei Cheng, Qiang Xu 0001 |
ICML | 6 |
| 2024 | Be Your Own Neighborhood: Detecting Adversarial Examples by the Neighborhood Relations Built on Self-Supervised LearningabstractDeep Neural Networks (DNNs) are vulnerable to Adversarial Examples (AEs), hindering their use in safety-critical systems. In this paper, we present BEYOND, an innovative AE detection framework designed for reliable predictions. BEYOND identifies AEs by distinguishing the AE’s abnormal relation with its augmented versions, i.e. neighbors, from two prospects: representation similarity and label consistency. An off-the-shelf Self-Supervised Learning (SSL) model is used to extract the representation and predict the label for its highly informative representation capacity compared to supervised learning models. We found clean samples maintain a high degree of representation similarity and label consistency relative to their neighbors, in contrast to AEs which exhibit significant discrepancies. We explain this observation and show that leveraging this discrepancy BEYOND can accurately detect AEs. Additionally, we develop a rigorous justification for the effectiveness of BEYOND. Furthermore, as a plug-and-play model, BEYOND can easily cooperate with the Adversarial Trained Classifier (ATC), achieving state-of-the-art (SOTA) robustness accuracy. Experimental results show that BEYOND outperforms baselines by a large margin, especially under adaptive attacks. Empowered by the robust relationship built on SSL, we found that BEYOND outperforms baselines in terms of both detection ability and speed. Project page: https://huggingface.co/spaces/allenhzy/Be-Your-Own-Neighborhood. Qiang Xu 0001, Tsung-Yi Ho |
ICML | 4 |
| 2024 | Vector Quantization Prompting for Continual LearningabstractContinual learning requires to overcome catastrophic forgetting when training a single model on a sequence of tasks. Recent top-performing approaches are prompt-based methods that utilize a set of learnable parameters (i.e., prompts) to encode task knowledge, from which appropriate ones are selected to guide the fixed pre-trained model in generating features tailored to a certain task. However, existing methods rely on predicting prompt identities for prompt selection, where the identity prediction process cannot be optimized with task loss. This limitation leads to sub-optimal prompt selection and inadequate adaptation of pre-trained features for a specific task. Previous efforts have tried to address this by directly generating prompts from input queries instead of selecting from a set of candidates. However, these prompts are continuous, which lack sufficient abstraction for task knowledge representation, making them less effective for continual learning. To address these challenges, we propose VQ-Prompt, a prompt-based continual learning method that incorporates Vector Quantization (VQ) into end-to-end training of a set of discrete prompts. In this way, VQ-Prompt can optimize the prompt selection process with task loss and meanwhile achieve effective abstraction of task knowledge for continual learning. Extensive experiments show that VQ-Prompt outperforms state-of-the-art continual learning methods across a variety of benchmarks under the challenging class-incremental setting. Qiuxia Lai, Yu Li 0007, Qiang Xu 0001 |
NeurIPS | 4 |
| 2024 | MiraData: A Large-Scale Video Dataset with Long Durations and Structured CaptionsabstractSora's high-motion intensity and long consistent videos have significantly impacted the field of video generation, attracting unprecedented attention. However, existing publicly available datasets are inadequate for generating Sora-like videos, as they mainly contain short videos with low motion intensity and brief captions. To address these issues, we propose MiraData, a high-quality video dataset that surpasses previous ones in video duration, caption detail, motion strength, and visual quality. We curate MiraData from diverse, manually selected sources and meticulously process the data to obtain semantically consistent clips. GPT-4V is employed to annotate structured captions, providing detailed descriptions from four different perspectives along with a summarized dense caption. To better assess temporal consistency and motion intensity in video generation, we introduce MiraBench, which enhances existing benchmarks by adding 3D consistency and tracking-based motion strength metrics. MiraBench includes 150 evaluation prompts and 17 metrics covering temporal consistency, motion strength, 3D consistency, visual quality, text-video alignment, and distribution similarity. To demonstrate the utility and effectiveness of MiraData, we conduct experiments using our DiT-based video generation model, MiraDiT. The experimental results on MiraBench demonstrate the superiority of MiraData, especially in motion strength. Xuan Ju, Yiming Gao 0007, Zhaoyang Zhang 0004, Ziyang Yuan, Xintao Wang 0002, Ailing Zeng, Qiang Xu 0001, Ying Shan |
NeurIPS | 8 |
| 2024 | GuardT2I: Defending Text-to-Image Models from Adversarial PromptsabstractRecent advancements in Text-to-Image models have raised significant safety concerns about their potential misuse for generating inappropriate or Not-Safe-For-Work contents, despite existing countermeasures such as Not-Safe-For-Work classifiers or model fine-tuning for inappropriate concept removal. Addressing this challenge, our study unveils GuardT2I a novel moderation framework that adopts a generative approach to enhance Text-to-Image models’ robustness against adversarial prompts. Instead of making a binary classification, GuardT2I utilizes a large language model to conditionally transform text guidance embeddings within the Text-to-Image models into natural language for effective adversarial prompt detection, without compromising the models’ inherent performance. Our extensive experiments reveal that GuardT2I outperforms leading commercial solutions like OpenAI-Moderation and Microsoft Azure Moderator by a significant margin across diverse adversarial scenarios. Our framework is available at https://github.com/cure-lab/GuardT2I. Ruiyuan Gao 0001, Jianyuan Zhong, Qiang Xu 0001 |
NeurIPS | 5 |
| 2024 | Large circuit models: opportunities and challengesabstractAbstract Within the electronic design automation (EDA) domain, artificial intelligence (AI)-driven solutions have emerged as formidable tools, yet they typically augment rather than redefine existing methodologies. These solutions often repurpose deep learning models from other domains, such as vision, text, and graph analytics, applying them to circuit design without tailoring to the unique complexities of electronic circuits. Such an “AI4EDA” approach falls short of achieving a holistic design synthesis and understanding, overlooking the intricate interplay of electrical, logical, and physical facets of circuit data. This study argues for a paradigm shift from AI4EDA towards AI-rooted EDA from the ground up, integrating AI at the core of the design process. Pivotal to this vision is the development of a multimodal circuit representation learning technique, poised to provide a comprehensive understanding by harmonizing and extracting insights from varied data sources, such as functional specifications, register-transfer level (RTL) designs, circuit netlists, and physical layouts. We champion the creation of large circuit models (LCMs) that are inherently multimodal, crafted to decode and express the rich semantics and structures of circuit data, thus fostering more resilient, efficient, and inventive design methodologies. Embracing this AI-rooted philosophy, we foresee a trajectory that transcends the current innovation plateau in EDA, igniting a profound “shift-left” in electronic design methodology. The envisioned advancements herald not just an evolution of existing EDA tools but a revolution, giving rise to novel instruments of design-tools that promise to radically enhance design productivity and inaugurate a new epoch where the optimization of circuit performance, power, and area (PPA) is achieved not incrementally, but through leaps that redefine the benchmarks of electronic systems’ capabilities. Zhufei Chu, Wenji Fang, Tsung-Yi Ho, Ru Huang 0001, Yu Huang 0005, Sadaf Khan, Yun Liang 0001, Yibo Lin, Guojie Luo, Hongyang Pan, Zhengyuan Shi, Guangyu Sun 0003, Dimitrios Tsaras, Runsheng Wang, Ziyi Wang 0010, Xinming Wei, Zhiyao Xie, Qiang Xu 0001, Chenhao Xue, Junchi Yan, Bei Yu 0001, Mingxuan Yuan, Evangeline F. Y. Young, Xuan Zeng 0001, Haoyi Zhang, Zuodong Zhang, Hui-Ling Zhen, Binwu Zhu, Keren Zhu 0001, Sunan Zou |
Sci. China Inf. Sci. | 25 |
| 2024 | Erratum to: Large circuit models: opportunities and challenges
Zhufei Chu, Wenji Fang, Tsung-Yi Ho, Ru Huang 0001, Yu Huang 0005, Sadaf Khan, Yun Liang 0001, Yibo Lin, Guojie Luo, Hongyang Pan, Zhengyuan Shi, Guangyu Sun 0003, Dimitrios Tsaras, Runsheng Wang, Ziyi Wang 0010, Xinming Wei, Zhiyao Xie, Qiang Xu 0001, Chenhao Xue, Junchi Yan, Bei Yu 0001, Mingxuan Yuan, Evangeline F. Y. Young, Xuan Zeng 0001, Haoyi Zhang, Zuodong Zhang, Hui-Ling Zhen, Binwu Zhu, Keren Zhu 0001, Sunan Zou |
Sci. China Inf. Sci. | 25 |
| 2024 | Spatial attention for human-centric visual understanding: An Information Bottleneck method
Qiuxia Lai, Yongwei Nie, Yu Li 0007, Hanqiu Sun, Qiang Xu 0001 |
Comput. Vis. Image Underst. | 5 |
| 2024 | Self-Supervised Video Representation Learning via Capturing Semantic Changes Indicated by SaccadesabstractIn this paper, we propose a self-supervised video representation learning (video SSL) method by taking inspiration from cognitive science and neuroscience on human visual perception. Different from previous methods that focus on the inherent properties of videos, we argue that humans learn to perceive the world through the self-awareness of the semantic changes or consistency in the input stimuli in the absence of labels, accompanied by representation reorganization during the post-learning rest periods. To this end, we first exploit the presence of saccades as an indicator of semantic changes in a contrastive learning framework, mimicking self-awareness in human representation learning. The saccades are generated by alternating the fixations following the predicted scanpath. Second, we model the semantic consistency in eye fixation by minimizing the prediction error between the predicted and the true state of another time point. Finally, we incorporate prototypical contrastive learning to reorganize the learned representations to enhance the associations among perceptually similar ones. Compared to previous video SSL solutions, our method can capture finer-grained semantics from video instances and further associate similar ones together. Experiments show that the proposed bio-inspired video SSL method significantly improves the Top-1 video retrieval accuracy on UCF101 and achieves superior performance on downstream tasks such as action recognition under comparable settings. Qiuxia Lai, Ailing Zeng, Ye Wang 0011, Lihong Cao, Yu Li 0007, Qiang Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Are Transformers Effective for Time Series Forecasting?abstractRecently, there has been a surge of Transformer-based solutions for the long-term time series forecasting (LTSF) task. Despite the growing performance over the past few years, we question the validity of this line of research in this work. Specifically, Transformers is arguably the most successful solution to extract the semantic correlations among the elements in a long sequence. However, in time series modeling, we are to extract the temporal relations in an ordered set of continuous points. While employing positional encoding and using tokens to embed sub-series in Transformers facilitate preserving some ordering information, the nature of the permutation-invariant self-attention mechanism inevitably results in temporal information loss. To validate our claim, we introduce a set of embarrassingly simple one-layer linear models named LTSF-Linear for comparison. Experimental results on nine real-life datasets show that LTSF-Linear surprisingly outperforms existing sophisticated Transformer-based LTSF models in all cases, and often by a large margin. Moreover, we conduct comprehensive empirical studies to explore the impacts of various design elements of LTSF models on their temporal relation extraction capability. We hope this surprising finding opens up new research directions for the LTSF task. We also advocate revisiting the validity of Transformer-based solutions for other time series analysis tasks (e.g., anomaly detection) in the future. Ailing Zeng, Muxi Chen, Lei Zhang 0001, Qiang Xu 0001 |
AAAI | 4 |
| 2023 | Human-Art: A Versatile Human-Centric Dataset Bridging Natural and Artificial ScenesabstractHumans have long been recorded in a variety of forms since antiquity. For example, sculptures and paintings were the primary media for depicting human beings before the invention of cameras. However, most current human-centric computer vision tasks like human pose estimation and human image generation focus exclusively on natural images in the real world. Artificial humans, such as those in sculptures, paintings, and cartoons, are commonly neglected, making existing models fail in these scenarios. As an abstraction of life, art incorporates humans in both natural and artificial scenes. We take advantage of it and introduce the Human-Art dataset to bridge related tasks in natural and artificial scenarios. Specifically, Human-Art contains 50k high-quality images with over 123k person instances from 5 natural and 15 artificial scenarios, which are annotated with bounding boxes, keypoints, self-contact points, and text information for humans represented in both 2D and 3D. It is, therefore, comprehensive and versatile for various downstream tasks. We also provide a rich set of baseline results and detailed analyses for related tasks, including human detection, 2D and 3D human pose estimation, image generation, and motion transfer. As a challenging dataset, we hope Human-Art can provide insights for relevant research and open up new research questions. Xuan Ju, Ailing Zeng, Qiang Xu 0001, Lei Zhang 0001 |
CVPR | 4 |
| 2023 | On EDA-Driven Learning for SAT SolvingabstractWe present DeepSAT, a novel end-to-end learning framework for the Boolean satisfiability (SAT) problem. Unlike existing solutions trained on random SAT instances with relatively weak supervision, we propose applying the knowledge of the well-developed electronic design automation (EDA) field for SAT solving. Specifically, we first resort to logic synthesis algorithms to pre-process SAT instances into optimized and-inverter graphs (AIGs). By doing so, the distribution diversity among various SAT instances can be dramatically reduced, which facilitates improving the generalization capability of the learned model. Next, we regard the distribution of SAT solutions being a product of conditional Bernoulli distributions. Based on this observation, we approximate the SAT solving procedure with a conditional generative model, leveraging a novel directed acyclic graph neural network (DAGNN) with two polarity prototypes for conditional SAT modeling. To effectively train the generative model, with the help of logic simulation tools, we obtain the probabilities of nodes in the AIG being logic ‘1’ as rich supervision. We conduct comprehensive experiments on various SAT problems. Our results show that, DeepSAT achieves significant accuracy improvements over state-of-the-art learning-based SAT solutions, especially when generalized to SAT instances that are relatively large or with diverse distributions. Min Li 0019, Zhengyuan Shi, Qiuxia Lai, Sadaf Khan, Shaowei Cai 0001, Qiang Xu 0001 |
DAC | 6 |
| 2023 | Integrating Exact Simulation into Sweeping for Datapath Combinational Equivalence CheckingabstractIn the application of IC design for microprocessors, there are often demands for optimizing the implementation of datapath circuits, on which various arithmetic operations are performed. Combinational equivalence checking (CEC) plays an essential role in ensuring the correctness of design optimization. The most prevalent CEC algorithms are based on SAT sweeping, which utilizes SAT to prove the equivalence of the internal node pairs in topological order, and the equivalent nodes are merged. Datapath circuits usually contain equivalent pairs for which the transitive fan-in cones are small but have a high XOR chain density, and proving such node pairs is very difficult for SAT solvers. An exact probability-based simulation (EPS) is suitable for verifying such pairs, while this method is not suitable for pairs with many primary inputs due to the memory cost. We first reduce the memory cost of EPS and integrate it to improve the SAT sweeping method. Considering the complementary abilities of SAT and EPS, we design an engine selection heuristic to dynamically choose SAT or EPS in the sweeping process, according to XOR chain density. Our method is further improved by reducing unnecessary engine calls by detecting regularity. Experiments on a benchmark suite from industrial datapath circuits show that our method is much faster than the state-of-the-art CEC tool namely ABC ‘&cec’ on nearly all instances, and is more than 100× faster on 30% of the instances, 1000× faster on 12% of the instances. Zhihan Chen 0001, Xindi Zhang 0001, Yuhang Qian, Qiang Xu 0001, Shaowei Cai 0001 |
ICCAD | 4 |
| 2023 | SATformer: Transformer-Based UNSAT Core LearningabstractThis paper introduces SATformer, a novel Transformer-based approach for the Boolean Satisfiability (SAT) problem. Rather than solving the problem directly, SATformer approaches the problem from the opposite direction by focusing on unsatisfiability. Specifically, it models clause interactions to identify any unsatisfiable sub-problems. Using a graph neural network, we convert clauses into clause embeddings and employ a hierarchical Transformer-based model to understand clause correlation. SATformer is trained through a multi-task learning approach, using the single-bit satisfiability result and the minimal unsatisfiable core (MUC) for UNSAT problems as clause supervision. As an end-to-end learning-based satisfiability classifier, the performance of SATformer surpasses that of NeuroSAT significantly. Furthermore, we integrate the clause predictions made by SATformer into modern heuristic-based SAT solvers and validate our approach with a logic equivalence checking task. Experimental results show that our SATformer can decrease the runtime of existing solvers by an average of 21.33%. Zhengyuan Shi, Min Li 0019, Yi Liu 0081, Sadaf Khan, Junhua Huang, Hui-Ling Zhen, Mingxuan Yuan, Qiang Xu 0001 |
ICCAD | 8 |
| 2023 | DeepGate2: Functionality-Aware Circuit Representation LearningabstractCircuit representation learning aims to obtain neural repre-sentations of circuit elements and has emerged as a promising research direction that can be applied to various EDA and logic reasoning tasks. Existing solutions, such as DeepGate, have the potential to embed both circuit structural information and functional behavior. However, their capabilities are limited due to weak supervision or flawed model design, resulting in unsatisfactory performance in downstream tasks. In this paper, we introduce Deep Gate2, a novel functionality-aware learning framework that significantly improves upon the original DeepGate solution in terms of both learning effectiveness and efficiency. Our approach involves using pairwise truth table differences between sampled logic gates as training supervision, along with a well-designed and scalable loss function that explicitly considers circuit functionality. Additionally, we consider inherent circuit characteristics and design an efficient one-round graph neural network (GNN), resulting in an order of magnitude faster learning speed than the original DeepGate solution. Experimental results demonstrate significant improvements in two practical downstream tasks: logic synthesis and Boolean satisfiability solving. The code is available at https://github.com/cure-lablDeepGate2. Zhengyuan Shi, Hongyang Pan, Sadaf Khan, Min Li 0019, Yi Liu 0081, Junhua Huang, Hui-Ling Zhen, Mingxuan Yuan, Zhufei Chu, Qiang Xu 0001 |
ICCAD | 10 |
| 2023 | DiffGuard: Semantic Mismatch-Guided Out-of-Distribution Detection using Pre-trained Diffusion ModelsabstractGiven a classifier, the inherent property of semantic Out-of-Distribution (OOD) samples is that their contents differ from all legal classes in terms of semantics, namely semantic mismatch. There is a recent work that directly applies it to OOD detection, which employs a conditional Generative Adversarial Network (cGAN) to enlarge semantic mismatch in the image space. While achieving remarkable OOD detection performance on small datasets, it is not applicable to ImageNet-scale datasets due to the difficulty in training cGANs with both input images and labels as conditions.As diffusion models are much easier to train and amenable to various conditions compared to cGANs, in this work, we propose to directly use pre-trained diffusion models for semantic mismatch-guided OOD detection, named DiffGuard. Specifically, given an OOD input image and the predicted label from the classifier, we try to enlarge the semantic difference between the reconstructed OOD image under these conditions and the original input image. We also present several test-time techniques to further strengthen such differences. Experimental results show that DiffGuard is effective on both Cifar-10 and hard cases of the large-scale ImageNet, and it can be easily combined with existing OOD detection techniques to achieve state-of-the-art OOD detection results. Ruiyuan Gao 0001, Chenchen Zhao 0001, Lanqing Hong, Qiang Xu 0001 |
ICCV | 4 |
| 2023 | HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image GenerationabstractControllable human image generation (HIG) has numerous real-life applications. State-of-the-art solutions, such as ControlNet and T2I-Adapter, introduce an additional learnable branch on top of the frozen pre-trained stable diffusion (SD) model, which can enforce various conditions, including skeleton guidance of HIG. While such a plug-and-play approach is appealing, the inevitable and uncertain conflicts between the original images produced from the frozen SD branch and the given condition incur significant challenges for the learnable branch, which essentially conducts image feature editing for condition enforcement.In this work, we propose a native skeleton-guided diffusion model for controllable HIG called HumanSD. Instead of performing image editing with dual-branch diffusion, we fine-tune the original SD model using a novel heatmap-guided denoising loss. This strategy effectively and efficiently strengthens the given skeleton condition during model training while mitigating the catastrophic forgetting effects. HumanSD is fine-tuned on the assembly of three large-scale human-centric datasets with text-image-pose information, two of which are established in this work. Experimental results show that HumanSD outperforms ControlNet in terms of pose control and image quality, particularly when the given skeleton guidance is sophisticated. Code and data are available at: https://idea-research.github.io/HumanSD/. Xuan Ju, Ailing Zeng, Chenchen Zhao 0001, Lei Zhang 0001, Qiang Xu 0001 |
ICCV | 6 |
| 2023 | Towards Robust Deep Neural Networks Against Design-Time and Run-Time FailuresabstractDeep Neural Networks (DNNs) have gained widespread adoption, but they also exhibit post-deployment failures that pose risks to property and life. Consequently, enhancing DNN robustness in safety-critical areas is crucial. This paper improves the robustness of DNN models against failures that may arise during either design time or run time. For design failures, we introduce TestRank, an efficient method for identifying DNN design issues. Constrained by test resources, TestRank selects high-quality test cases leveraging intrinsic and contextual attributes of test samples. HybridRepair then offers effective failure repair by selectively annotating failure regions and utilizing semi-supervised learning techniques. To counteract run-time fault injection attacks, we propose D2NN and DeepDyve, which introduce the dual modular redundancy concept to models for protection at the neuron and system levels, respectively. This delicate redundancy achieves lightweight protection for DNN-based systems. Evaluation performed on various image classification datasets demonstrates the effectiveness of our approaches. Yu Li 0007, Qiang Xu 0001 |
ITC | 2 |
| 2023 | HiBug: On Human-Interpretable Model DebugabstractMachine learning models can frequently produce systematic errors on critical subsets (or slices) of data that share common attributes. Discovering and explaining such model bugs is crucial for reliable model deployment. However, existing bug discovery and interpretation methods usually involve heavy human intervention and annotation, which can be cumbersome and have low bug coverage.
In this paper, we propose HiBug, an automated framework for interpretable model debugging. Our approach utilizes large pre-trained models, such as chatGPT, to suggest human-understandable attributes that are related to the targeted computer vision tasks. By leveraging pre-trained vision-language models, we can efficiently identify common visual attributes of underperforming data slices using human-understandable terms. This enables us to uncover rare cases in the training data, identify spurious correlations in the model, and use the interpretable debug results to select or generate new training data for model improvement. Experimental results demonstrate the efficacy of the HiBug framework. Muxi Chen, Yu Li 0007, Qiang Xu 0001 |
NeurIPS | 3 |
| 2022 | DeepGate: learning neural representations of logic gatesabstractApplying deep learning (DL) techniques in the electronic design automation (EDA) field has become a trending topic. Most solutions apply well-developed DL models to solve specific EDA problems. While demonstrating promising results, they require careful model tuning for every problem. The fundamental question on "How to obtain a general and effective neural representation of circuits?" has not been answered yet. In this work, we take the first step towards solving this problem. We propose DeepGate, a novel representation learning solution that effectively embeds both logic function and structural information of a circuit as vectors on each gate. Specifically, we propose transforming circuits into unified and-inverter graph format for learning and using signal probabilities as the supervision task in DeepGate. We then introduce a novel graph neural network that uses strong inductive biases in practical circuits as learning priors for signal probability prediction. Our experimental results show the efficacy and generalization capability of DeepGate. Min Li 0019, Sadaf Khan, Zhengyuan Shi, Naixing Wang, Huang Yu, Qiang Xu 0001 |
DAC | 6 |
| 2022 | Functionality matters in netlist representation learningabstractLearning feasible representation from raw gate-level netlists is essential for incorporating machine learning techniques in logic synthesis, physical design, or verification. Existing message-passing-based graph learning methodologies focus merely on graph topology while overlooking gate functionality, which often fails to capture underlying semantic, thus limiting their generalizability. To address the concern, we propose a novel netlist representation learning framework that utilizes a contrastive scheme to acquire generic functional knowledge from netlists effectively. We also propose a customized graph neural network (GNN) architecture that learns a set of independent aggregators to better cooperate with the above framework. Comprehensive experiments on multiple complex real-world designs demonstrate that our proposed solution significantly outperforms state-of-the-art netlist feature learning flows. Ziyi Wang 0010, Zhuolun He, Guangliang Zhang, Qiang Xu 0001, Tsung-Yi Ho, Bei Yu 0001, Yu Huang 0005 |
DAC | 5 |
| 2022 | Out-of-Distribution Detection with Semantic Mismatch Under Masking
Ruiyuan Gao 0001, Qiang Xu 0001 |
ECCV (24) | 3 |
| 2022 | DeciWatch: A Simple Baseline for 10˟ Efficient 2D and 3D Pose Estimation
Ailing Zeng, Xuan Ju, Ruiyuan Gao 0001, Xizhou Zhu, Bo Dai 0002, Qiang Xu 0001 |
ECCV (5) | 7 |
| 2022 | SmoothNet: A Plug-and-Play Network for Refining Human Poses in Videos
Ailing Zeng, Xuan Ju, Jianyi Wang, Qiang Xu 0001 |
ECCV (5) | 6 |
| 2022 | T-WaveNet: A Tree-Structured Wavelet Neural Network for Time Series Signal Analysis
Minhao Liu, Ailing Zeng, Qiuxia Lai, Ruiyuan Gao 0001, Min Li 0019, Harry Qin, Qiang Xu 0001 |
ICLR | 7 |
| 2022 | HybridRepair: towards annotation-efficient repair for deep learning modelsabstractA well-trained deep learning (DL) model often cannot achieve expected performance after deployment due to the mismatch between the distributions of the training data and the field data in the operational environment. Therefore, repairing DL models is critical, especially when deployed on increasingly larger tasks with shifted distributions. Yu Li 0007, Muxi Chen, Qiang Xu 0001 |
ISSTA | 3 |
| 2022 | DeepTPI: Test Point Insertion with Deep Reinforcement LearningabstractTest point insertion (TPI) is a widely used technique for testability enhancement, especially for logic built-in self-test (LBIST) due to its relatively low fault coverage. In this paper, we propose a novel TPI approach based on deep reinforcement learning (DRL), named DeepTpi. Unlike previous learning-based solutions that formulate the TPI task as a supervised-learning problem, we train a novel DRL agent, instantiated as the combination of a graph neural network (GNN) and a Deep Q-Learning network (DQN), to maximize the test coverage improvement. Specifically, we model circuits as directed graphs and design a graph-based value network to estimate the action values for inserting different test points. The policy of the DRL agent is defined as selecting the action with the maximum value. Moreover, we apply the general node embeddings from a pretrained model to enhance node features, and propose a dedicated testability-aware attention mechanism for the value network. Experimental results on circuits with various scales show that DeepTPI significantly improves test coverage compared to the commercial DFT tool. The code of this work is available at https://github.com/cure-lab/DeepTPI. Zhengyuan Shi, Min Li 0019, Sadaf Khan, Liuzheng Wang, Naixing Wang, Yu Huang 0005, Qiang Xu 0001 |
ITC | 7 |
| 2022 | What You See is Not What the Network Infers: Detecting Adversarial Examples Based on Semantic Contradiction
Ruiyuan Gao 0001, Yu Li 0007, Qiuxia Lai, Qiang Xu 0001 |
NDSS | 5 |
| 2022 | SCINet: Time Series Modeling and Forecasting with Sample Convolution and InteractionabstractOne unique property of time series is that the temporal relations are largely preserved after downsampling into two sub-sequences. By taking advantage of this property, we propose a novel neural network architecture that conducts sample convolution and interaction for temporal modeling and forecasting, named SCINet. Specifically, SCINet is a recursive downsample-convolve-interact architecture. In each layer, we use multiple convolutional filters to extract distinct yet valuable temporal features from the downsampled sub-sequences or features. By combining these rich features aggregated from multiple resolutions, SCINet effectively models time series with complex temporal dynamics. Experimental results show that SCINet achieves significant forecasting accuracy improvements over both existing convolutional models and Transformer-based solutions across various real-world time series forecasting datasets. Our codes and data are available at https://github.com/cure-lab/SCINet. Minhao Liu, Ailing Zeng, Muxi Chen, Qiuxia Lai, Lingna Ma, Qiang Xu 0001 |
NeurIPS | 7 |
| 2021 | AppealNet: An Efficient and Highly-Accurate Edge/Cloud Collaborative Architecture for DNN InferenceabstractThis paper presents AppealNet, a novel edge/cloud collaborative architecture that runs deep learning (DL) tasks more efficiently than state-of-the-art solutions. For a given input, AppealNet accurately predicts on-the-fly whether it can be successfully processed by the DL model deployed on the resource-constrained edge device, and if not, appeals to the more powerful DL model deployed at the cloud. This is achieved by employing a two-head neural network architecture that explicitly takes inference difficulty into consideration and optimizes the tradeoff between accuracy and computation/communication cost of the edge/cloud collaborative architecture. Experimental results on several image classification datasets show up to more than 40% energy savings compared to existing techniques without sacrificing accuracy. Min Li 0019, Yu Li 0007, Ye Tian 0010, Li Jiang 0002, Qiang Xu 0001 |
DAC | 5 |
| 2021 | Learning Skeletal Graph Neural Networks for Hard 3D Pose EstimationabstractVarious deep learning techniques have been proposed to solve the single-view 2D-to-3D pose estimation problem. While the average prediction accuracy has been improved significantly over the years, the performance on hard poses with depth ambiguity, self-occlusion, and complex or rare poses is still far from satisfactory. In this work, we target these hard poses and present a novel skeletal GNN learning solution. To be specific, we propose a hop-aware hierarchical channel-squeezing fusion layer to effectively extract relevant information from neighboring nodes while suppressing undesired noises in GNN learning. In addition, we propose a temporal-aware dynamic graph construction procedure that is robust and effective for 3D pose estimation. Experimental results on the Human3.6M dataset show that our solution achieves 10.3% average prediction accuracy improvement and greatly improves on hard poses over state-of-the-art techniques. We further apply the proposed technique on the skeleton-based action recognition task and also achieve state-of-the-art performance. Our code is available at https://github.com/ailingzengzzz/Skeletal-GNN. Ailing Zeng, Xiao Sun 0001, Nanxuan Zhao, Minhao Liu, Qiang Xu 0001 |
ICCV | 6 |
| 2021 | Information Bottleneck Approach to Spatial Attention LearningabstractThe selective visual attention mechanism in the human visual system (HVS) restricts the amount of information to reach visual awareness for perceiving natural scenes, allowing near real-time information processing with limited computational capacity. This kind of selectivity acts as an ‘Information Bottleneck (IB)’, which seeks a trade-off between information compression and predictive accuracy. However, such information constraints are rarely explored in the attention mechanism for deep neural networks (DNNs). In this paper, we propose an IB-inspired spatial attention module for DNN structures built for visual recognition. The module takes as input an intermediate representation of the input image, and outputs a variational 2D attention map that minimizes the mutual information (MI) between the attention-modulated representation and the input, while maximizing the MI between the attention-modulated representation and the task label. To further restrict the information bypassed by the attention map, we quantize the continuous attention scores to a set of learnable anchor values during training. Extensive experiments show that the proposed IB-inspired spatial attention mechanism can yield attention maps that neatly highlight the regions of interest while suppressing backgrounds, and bootstrap standard DNN structures for visual recognition tasks (e.g., image classification, fine-grained recognition, cross-domain classification). The attention maps are interpretable for the decision making of the DNNs as verified in the experiments. Our code is available at this https URL. Qiuxia Lai, Yu Li 0007, Ailing Zeng, Minhao Liu, Hanqiu Sun, Qiang Xu 0001 |
IJCAI | 6 |
| 2021 | Testability-Aware Low Power Controller Design with Evolutionary LearningabstractXORNet-based low power controller is a popular technique to reduce circuit transitions in scan-based testing. However, existing solutions construct the XORNet evenly for scan chain control, and it may result in sub-optimal solutions without any design guidance. In this paper, we propose a novel testability-aware low power controller with evolutionary learning. The XORNet generated from the proposed genetic algorithm (GA) enables adaptive control for scan chains according to their usages, thereby significantly improving XORNet encoding capacity, reducing the number of failure cases with ATPG and decreasing test data volume. Experimental results indicate that under the same control bits, our GA-guided XORNet design can improve the fault coverage by up to 2.11%. The proposed GA-guided XORNets also allows reducing the number of control bits, and the total testing time decreases by 20.78% on average and up to 47.09% compared to the existing design without sacrificing test coverage. Min Li 0019, Zhengyuan Shi, Zezhong Wang 0006, Yu Huang 0005, Qiang Xu 0001 |
ITC | 6 |
| 2021 | TestRank: Bringing Order into Unlabeled Test Instances for Deep Learning TasksabstractDeep learning (DL) systems are notoriously difficult to test and debug due to the lack of correctness proof and the huge test input space to cover. Given the ubiquitous unlabeled test data and high labeling cost, in this paper, we propose a novel test prioritization technique, namely TestRank, which aims at revealing more model failures with less labeling effort. TestRank brings order into the unlabeled test data according to their likelihood of being a failure, i.e., their failure-revealing capabilities. Different from existing solutions, TestRank leverages both intrinsic and contextual attributes of the unlabeled test data when prioritizing them. To be specific, we first build a similarity graph on both unlabeled test samples and labeled samples (e.g., training or previously labeled test samples). Then, we conduct graph-based semi-supervised learning to extract contextual features from the correctness of similar labeled samples. For a particular test instance, the contextual features extracted with the graph neural network and the intrinsic features obtained with the DL model itself are combined to predict its failure-revealing capability. Finally, TestRank prioritizes unlabeled test inputs in descending order of the above probability value. We evaluate TestRank on three popular image classification datasets, and results show that TestRank significantly outperforms existing test prioritization techniques. Yu Li 0007, Min Li 0019, Qiuxia Lai, Yannan Liu, Qiang Xu 0001 |
NeurIPS | 5 |
| 2021 | On Workload-Aware DRAM Failure Prediction in Large-Scale Data CentersabstractDRAM failures are one of the major hardware threats to the reliability of large-scale data centers since the uncorrectable errors in DRAMs may cause servers to shut down. Existing works try to solve this problem by predicting DRAM failures in advance with Machine Learning models. In these works, correctable errors (CEs) are generally deemed as the most important feature. The major reason behind CEs' emergence is the accumulated stress caused by intensive workloads. Moreover, defective DRAMs will not manifest themselves as system errors until the defective cells are accessed by some specific workloads. Therefore, the running workloads on a server are also important for DRAM failure prediction. In this paper, we focus on the impact of workloads on DRAM failures. We design the workload features from both macroscopical and microscopical aspects, i.e. node-level performance metrics and cell-level DRAM access pattern, respectively. Furthermore, we propose Hierarchical DRAM Error Code (HiDEC) to represent the DRAM access pattern. We leverage several Decision Tree-based models for DRAM failure prediction to highlight the generality of our designed features. Experiments are carried out based on the dataset collected from a real-world commercial data center. The results show that both macroscopic and microscopic features can bring significant improvements to the prediction performance. Xingyi Wang, Yu Li 0007, Yiquan Chen, Yin Du, Yuzhong Zhang, Pinan Chen, Wenjun Song, Qiang Xu 0001, Li Jiang 0002 |
VTS | 11 |
| 2020 | DeepDyve: Dynamic Verification for Deep Neural NetworksabstractDeep neural networks (DNNs) have become one of the enabling technologies in many safety-critical applications, e.g., autonomous driving and medical image analysis. DNN systems, however, suffer from various kinds of threats, such as adversarial example attacks and fault injection attacks. While there are many defense methods proposed against maliciously crafted inputs, solutions against faults presented in the DNN system itself (e.g., parameters and calculations) are far less explored. In this paper, we develop a novel lightweight fault-tolerant solution for DNN-based systems, namely DeepDyve, which employs pre-trained neural networks that are far simpler and smaller than the original DNN for dynamic verification. The key to enabling such lightweight checking is that the smaller neural network only needs to produce approximate results for the initial task without sacrificing fault coverage much. We develop efficient and effective architecture and task exploration techniques to achieve optimized risk/overhead trade-off in DeepDyve. Experimental results show that DeepDyve can reduce 90% of the risks at around 10% overhead. Yu Li 0007, Min Li 0019, Bo Luo, Ye Tian 0010, Qiang Xu 0001 |
CCS | 5 |
| 2020 | SRNet: Improving Generalization in 3D Human Pose Estimation with a Split-and-Recombine Approach
Ailing Zeng, Xiao Sun 0001, Fuyang Huang, Minhao Liu, Qiang Xu 0001, Stephen Lin 0001 |
ECCV (14) | 5 |
| 2020 | On Configurable Defense against Adversarial Example AttacksabstractMachine learning systems based on deep neural networks (DNNs) have gained mainstream adoption in many applications. Recently, however, DNNs are shown to be vulnerable to adversarial example attacks with slight perturbations on the inputs. Existing defense mechanisms against such attacks try to improve the overall robustness of the system, but they do not differentiate different targeted attacks even though the corresponding impacts may vary significantly. To tackle this problem, we propose a novel configurable defense mechanism in this work, wherein we are able to flexibly tune the robustness of the system against different targeted attacks to satisfy application requirements. This is achieved by refining the DNN loss function with an attack sensitive matrix to represent the impacts of different targeted attacks. Experimental results on CIFAR-10 data set demonstrate the efficacy of the proposed solution. Bo Luo, Min Li 0019, Yu Li 0007, Qiang Xu 0001 |
ACM Great Lakes Symposium on VLSI | 4 |
| 2020 | DeepFuse: An IMU-Aware Network for Real-Time 3D Human Pose Estimation from Multi-View ImageabstractIn this paper, we propose a two-stage fully 3D network, namely DeepFuse, to estimate human pose in 3D space by fusing body-worn Inertial Measurement Unit (IMU) data and multi-view images deeply. The first stage is designed for pure vision estimation. To preserve data primitiveness of multi-view inputs, the vision stage uses multi-channel volume as data representation and 3D soft-argmax as activation layer. The second one is the IMU refinement stage which introduces an IMU-bone layer to fuse the IMU and vision data earlier at data level. without requiring a given skeleton model a priori, we can achieve a mean joint error of 28.9mm on TotalCapture dataset and 13.4mm on Human3.6M dataset under protocol 1, improving the SOTA result by a large margin. Finally, we discuss the effectiveness of a fully 3D network for 3D pose estimation experimentally which may benefit future research. Fuyang Huang, Ailing Zeng, Minhao Liu, Qiuxia Lai, Qiang Xu 0001 |
WACV | 5 |
| 2020 | Video super-resolution via pre-frame constrained and deep-feature enhanced sparse reconstruction
Qiuxia Lai, Yongwei Nie, Hanqiu Sun, Qiang Xu 0001, Zhensong Zhang, Mingyu Xiao 0001 |
Pattern Recognit. | 4 |
| 2020 | ApproxIt: A Quality Management Framework of Approximate Computing for Iterative MethodsabstractApproximate computing, being able to tradeoff computation quality (e.g., accuracy) and computational effort (e.g., energy) for error-tolerant applications such as media processing and the emerging recognition, mining, and synthesis (RMS) applications, has gained significant traction in recent years. With approximate computing, we expect to obtain acceptable results, but how do we make sure the quality of the final results are good enough? This challenging problems remains largely unexplored. As many of the RMS applications employ iterative methods (IMs) for solution-finding, wherein a sequence of improving approximate solutions are generated before reaching the final converged solution, in this paper, we propose ApproxIt, a novel quality management framework of approximate computing dedicated for IMs with quality guarantees. ApproxIt is comprised of two stages: 1) offline stage and 2) online stage. To be specific, at offline stage, we first analyze the manifold of parameter space to identify the given problem as convex case or nonconvex case at the offline stage. And for each case, we propose the corresponding runtime dynamic quality calibration scheme and reconfiguration control policy. Then during runtime, our proposed lightweight quality estimator will evaluate the intermediate quality at specific calibration iteration, which is determined by the novel Markov model-based calibration scheme. If quality violation occurs, the configuration control policy will select the most appropriate approximate computing mode for the following iterations. With the proposed dynamic effort scaling technique, ApproxIt is able to dramatically improve application energy efficiency under quality guarantees, as demonstrated in our experimental results. Qian Zhang 0020, Qiang Xu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | D2NN: a fine-grained dual modular redundancy framework for deep neural networksabstractDeep Neural Networks (DNNs) have attracted mainstream adoption in various application domains. Their reliability and security are therefore serious concerns in those safety-critical applications such as surveillance and medical systems. In this paper, we propose a novel dual modular redundancy framework for DNNs, namely D2NN, which is able to tradeoff the system robustness with overhead in a fine-grained manner. We evaluate D2NN framework with DNN models trained on MNIST and CIFAR10 datasets under fault injection attacks, and experimental results demonstrate the efficacy of our proposed solution. Yu Li 0007, Yannan Liu, Min Li 0019, Ye Tian 0010, Bo Luo, Qiang Xu 0001 |
ACSAC | 6 |
| 2019 | On Functional Test Generation for Deep Neural Network IPsabstractMachine learning systems based on deep neural networks (DNNs) produce state-of-the-art results in many applications. Considering the large amount of training data and know-how required to generate the network, it is more practical to use third-party DNN intellectual property (IP) cores for many designs. No doubt to say, it is essential for DNN IP vendors to provide test cases for functional validation without leaking their parameters to IP users. To satisfy this requirement, we propose to effectively generate test cases that activate parameters as many as possible and propagate their perturbations to outputs. Then the functionality of DNN IPs can be validated by only checking their outputs. However, it is difficult considering large numbers of parameters and highly non-linearity of DNNs. In this paper, we tackle this problem by judiciously selecting samples from the DNN training set and applying a gradient-based method to generate new test cases. Experimental results demonstrate the efficacy of our proposed solution. Bo Luo, Yu Li 0007, Lingxiao Wei, Qiang Xu 0001 |
DATE | 4 |
| 2019 | Energy-Efficient and Quality-Assured Approximate Computing Framework Using a Co-Training MethodabstractApproximate computing is a promising design paradigm that introduces a new dimension—error—into the original design space. By allowing the inexact computation in error-tolerance applications, approximate computing can gain both performance and energy efficiency. A neural network (NN) is a universal approximator in theory and possesses a high level of parallelism. The emerging deep neural network accelerators deployed with NN-based approximator is thereby a promising candidate for approximate computing. Nevertheless, the approximation result must satisfy the users’ requirement, and the approximation result varies across different applications. We normally deploy an NN-based classifier to ensure the approximation quality. Only the inputs predicted to meet the quality requirement can be executed by the approximator. The potential of these two NNs, however, is fully explored; the involving of two NNs in approximate computing imposes critical optimization questions, such as two NNs’ distinct views of the input data space, how to train the two correlated NNs, and what are their topologies. In this article, we propose a novel NN-based approximate computing framework with quality insurance. We advocate a co-training approach that trains the classifier and the approximator alternately to maximize the agreement of the two NNs on the input space. In each iteration, we coordinate the training of the two NNs with a judicious selection of training data. Next, we explore different selection policies and propose to select training data from multiple iterations, which can enhance the invocation of the approximate accelerator. In addition, we optimize the classifier by integrating a dynamic threshold tuning algorithm to improve the invocation of the approximate accelerator further. The increased invocation of accelerator leads to higher energy efficiency under the same quality requirement. We propose two efficient algorithms to explore the smallest topology of the NN-based approximator and the classifier to achieve the quality requirement. The first algorithm straightforward searches the minimum topology using a greedy strategy. However, the first algorithm incurs too much training overhead. To solve this issue, the second one gradually grows the topology of NNs to match the quality requirement by transferring the learned parameters. Experimental results show significant improvement on the quality and the energy efficiency compared to the existing NN-based approximate computing frameworks. Li Jiang 0002, Zhuoran Song, Haiyue Song, Chengwen Xu, Qiang Xu 0001, Naifeng Jing, Weifeng Zhang 0003, Xiaoyao Liang |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2018 | Towards Imperceptible and Robust Adversarial Example Attacks Against Neural NetworksabstractMachine learning systems based on deep neural networks, being able to produce state-of-the-art results on various perception tasks, have gained mainstream adoption in many applications. However, they are shown to be vulnerable to adversarial example attack, which generates malicious output by adding slight perturbations to the input. Previous adversarial example crafting methods, however, use simple metrics to evaluate the distances between the original examples and the adversarial ones, which could be easily detected by human eyes. In addition, these attacks are often not robust due to the inevitable noises and deviation in the physical world. In this work, we present a new adversarial example attack crafting method, which takes the human perceptual system into consideration and maximizes the noise tolerance of the crafted adversarial example. Experimental results demonstrate the efficacy of the proposed technique. Bo Luo, Yannan Liu, Lingxiao Wei, Qiang Xu 0001 |
AAAI | 4 |
| 2018 | I Know What You See: Power Side-Channel Attack on Convolutional Neural Network AcceleratorsabstractDeep learning has become the de-facto computational paradigm for various kinds of perception problems, including many privacy-sensitive applications such as online medical image analysis. No doubt to say, the data privacy of these deep learning systems is a serious concern. Different from previous research focusing on exploiting privacy leakage from deep learning models, in this paper, we present the first attack on the implementation of deep learning models. To be specific, we perform the attack on an FPGA-based convolutional neural network accelerator and we manage to recover the input image from the collected power traces without knowing the detailed parameters in the neural network. For the MNIST dataset, our power side-channel attack is able to achieve up to 89% recognition accuracy. Lingxiao Wei, Bo Luo, Yu Li 0007, Yannan Liu, Qiang Xu 0001 |
ACSAC | 5 |
| 2018 | Structure-Aware 3D Hourglass Network for Hand Pose Estimation from Single Depth Image
Fuyang Huang, Ailing Zeng, Minhao Liu, Harry Qin, Qiang Xu 0001 |
BMVC | 5 |
| 2018 | Lookup table allocation for approximate computing with memory under quality constraintsabstractComputation kernels in emerging recognition, mining, and synthesis (RMS) applications are inherently error-resilient, where approximate computing can be applied to improve their energy efficiency by trading off computational effort and output quality. One promising approximate computing technique is to perform approximate computing with memory, which stores a subset of function responses in a lookup table (LUT), and avoids redundant computation when encountering similar input patterns. Limited by the memory space, most existing solutions simply store values for those frequently-appeared input patterns, without considering output quality and/or intrinsic characteristic of the target kernel. In this paper, we propose a novel LUT allocation technique for approximate computing with memory, which is able to dramatically improve the hit rate of LUT and hence achieves significant energy savings under given quality constraints. We also present how to apply the proposed LUT allocation solution for multiple computation kernels. Experimental results show the efficacy of our proposed methodology. Ye Tian 0010, Qian Zhang 0020, Ting Wang 0008, Qiang Xu 0001 |
DATE | 4 |
| 2018 | Invocation-driven neural approximate computing with a multiclass-classifier and multiple approximatorsabstractNeural approximate computing gains enormous energy-efficiency at the cost of tolerable quality-loss. A neural approximator can map the input data to output while a classifier determines whether the input data are safe to approximate with quality guarantee. However, existing works cannot maximize the invocation of the approximator, resulting in limited speedup and energy saving. By exploring the mapping space of those target functions, in this paper, we observe a nonuniform distribution of the approximation error incurred by the same approximator. We thus propose a novel approximate computing architecture with a Multiclass-Classifier and Multiple Approximators (MCMA). These approximators have identica network topologies, and thus can share the same hardware resource in an neural processing unit(NPU) clip. In the runtime, MCMA can swap in the invoked approximator by merely shipping the synapse weights from the on-chip memory to the buffers near MAC within a cycle. We also propose efficient co-training methods for such MCMA architecture. Experimental results show a more substantial invocation of MCMA as well as the gain of energy-efficiency. Haiyue Song, Chengwen Xu, Qiang Xu 0001, Zhuoran Song, Naifeng Jing, Xiaoyao Liang, Li Jiang 0002 |
ICCAD | 3 |
| 2018 | Shadow Block: Accelerating ORAM Accesses with Data DuplicationabstractOblivious RAM (ORAM) is a cryptographic primitive designed to hide memory access patterns. To achieve this objective, the intended data block is loaded and evicted back together with other data blocks and dummy blocks in each ORAM access. To further protect the timing pattern, extra dummy ORAM accesses are triggered periodically. Such designs lead to huge memory access overheads. Many techniques have been proposed to mitigate this problem by reducing the total number of ORAM accesses and the number of blocks per access. However, the impact of the access order of intended data block in an ORAM access is not addressed yet. In this work, we argue that higher performance can be achieved by advancing the access to the intended data block in ORAM accesses. However, changing the access order of blocks directly compromises the ORAM security. To solve this problem, we propose a duplication method to advance the access to the intended data blocks without compromising the ORAM security. The method leverages dummy blocks to store extra copies of data blocks, to facilitate early access of intended data blocks. These dummy blocks with valid data duplications are called Shadow blocks in this work. We further introduce two data duplication techniques, called RD-Dup and HD-Dup, to reorder the data block access for different purposes. In addition, we propose ORAM space partitioning to make RD-Dup and HD-Dup cooperate with each other efficiently. Compared with state-of-the-art ORAMs, our design can achieve a 32% reduction in system execution time on average, with negligible hardware overheads. Xian Zhang 0001, Guangyu Sun 0003, Peichen Xie, Chao Zhang 0007, Yannan Liu, Lingxiao Wei, Qiang Xu 0001, Chun Jason Xue |
MICRO | 7 |
| 2018 | SERA: statistical error rate analysis for profit-oriented performance binning of resilient circuits
Qiang Xu 0001, Wen-Ben Jone |
Integr. | 2 |
| 2018 | Hardware Trojan Detection in Third-Party Digital Intellectual Property Cores by Multilevel Feature AnalysisabstractIn modern integrated circuit (IC) designs, intellectual property (IP) cores are often outsourced and designed by third-party vendors, resulting in the partial relinquishment of the control over the IC design flow. Thus, reliable verifications are required to mitigate the threat of hardware Trojans (HTs) which may be inserted into IP cores by malicious vendors. Existing trustiness verification methods cannot take the merit of high efficiency and accuracy at the same time. In this paper, we propose a multilevel fast trustiness verification framework based on feature analysis to detect HTs in third-party digital IP cores. The proposed framework combines flip-flop level and combinational logic level feature analysis to achieve both high efficiency and accuracy. Experimental results demonstrate that both explicitly and implicitly triggered HTs can be detected in very short time with a negligible false positive rate. More importantly, our framework has the unique advantage of being scalable to defend against future and stealthier HTs by adding new features into the framework. Xiaoming Chen 0003, Qiaoyi Liu, Jia Wang 0004, Qiang Xu 0001, Yu Wang 0002, Yongpan Liu, Huazhong Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2017 | On efficient message passing in energy harvesting based distributed systemabstractEnergy harvesting systems recycling energy from ambient environment for extended lifetime have been proposed to be the next step in the evolution of the Internet of Things. One of the largest challenges of energy harvesting is the insufficiency and instability of ambient energy sources. Unlike wireless distributed systems considered to date, frequent power failures need to be taken into account in designing energy-efficient transmission strategies. In this paper, we propose a message passing method between energy harvesting nodes, which predicts the probability of success before message passing and adjusts transmission strategy online according to input power trace. Experimental result shows that our method can reduce energy consumption of transmission significantly compared to traditional ones and keeps effective for various kinds of unstable input power. Ye Tian 0010, Qiang Xu 0001, Chun Jason Xue |
ASP-DAC | 2 |
| 2017 | On resilient task allocation and scheduling with uncertain quality checkersabstractMany emerging applications are inherently error-resilient and hence do not require exact computation. Previous work on resilience-aware task allocation and scheduling problem first generates an initial energy-efficient task schedule on voltage-scalable multiprocessor system at design-time, and then conducts voltage adjustment according to runtime quality checking result. While energy efficiency improvements are quite encouraging, the final quality requirement might be violated because quality checkers are usually designed based on partial information and they are not always correct. In this paper, we propose to address the uncertainty issue of quality checkers in resilient task allocation and scheduling. To be specific, given the initial task schedule and quality checkers, we propose (i) a solution that ensures the final output quality with maximized probability, which can be applied to any approximate computing quality management systems; and (ii) a greedy runtime algorithm to achieve optimized energy efficiency gains. Experimental results on various task graphs demonstrate the efficacy of our proposed technique. Qian Zhang 0020, Ting Wang 0008, Qiang Xu 0001 |
ASP-DAC | 3 |
| 2017 | On Quality Trade-off Control for Approximate Computing Using Iterative TrainingabstractQuality control plays a key role in approximate computing to save the energy and guarantee that the quality of the computation outcome satisfies users' requirement. Previous works proposed a hybrid architecture, composed of a classifier for error prediction and an approximate accelerator for approximate computing using well trained neural-networks. Only inputs predicted to meet the quality are executed by the accelerator. However, the design of this hybrid architecture, relying on one-pass training process, has not been fully explored. In this paper, we propose a novel optimization framework. It advocates an iteratively training process to coordinate the training of the classifier and the accelerator with a judicious selection of training data. It integrates a dynamic threshold tuning algorithm to maximize the invocation of the accelerator (i.e., energy-efficiency) under the quality requirement. At last, we propose an efficient algorithm to explore the topologies of the accelerator and the classifier comprehensively. Experimental results shows significant improvement on the quality and the energy-efficiency compared to the conventional one-pass training method. Chengwen Xu, Wenqi Yin, Qiang Xu 0001, Naifeng Jing, Xiaoyao Liang, Li Jiang 0002 |
DAC | 4 |
| 2017 | RetroDMR: Troubleshooting non-deterministic faults with retrospective DMRabstractThe most notorious faults for diagnosis in post-silicon validation are those that manifest themselves in a non-deterministic manner with system-level functional tests, where errors randomly appear from time to time even when applying the same workloads. In this work, we propose a novel diagnostic framework that resorts to dual-modular redundancy (DMR) for troubleshooting non-deterministic faults, namely RetroDMR. To be specific, we log the essential events (e.g., the sequence of thread migration) in the faulty run to record the mapping relationship between threads and their corresponding execution units. Then in the following diagnosis runs, we apply redundant multithreading (RMT) technique to reduce error detection latency, while at the same time we try to follow the thread migration sequence of the original run whenever possible. By doing so, RetroDMR significantly improves the reproduction rate and diagnosis resolution for non-deterministic faults, as demonstrated in our experimental results. Ting Wang 0008, Yannan Liu, Qiang Xu 0001, Zhaobo Zhang, Xinli Gu |
DATE | 3 |
| 2017 | ApproxQA: A unified quality assurance framework for approximate computingabstractApproximate computing, being able to trade off computation quality and computational effort (e.g., energy) by exploiting the inherent error-resilience of emerging applications (e.g., recognition and mining), has garnered significant attention recently. No doubt to say, quality assurance is indispensable for satisfactory user experience with approximate computing, but this issue has remained largely unexplored in the literature. In this work, we propose a novel framework namely ApproxQA to tackle this problem, in which approximation mode tuning and rollback recovery are considered in a unified manner. To be specific, ApproxQA resorts to a two-level controller, in which the high-level approximation controller tunes approximation modes at a coarse-grained scale based on Q-learning while the low-level rollback controller judiciously determines whether to perform rollback recovery at a fine-grained scale based on the target quality requirement. Experimental results on various benchmark applications demonstrate that it significantly outperforms existing solutions in terms of energy efficiency with quality assurance. Ting Wang 0008, Qian Zhang 0020, Qiang Xu 0001 |
DATE | 3 |
| 2017 | Fault injection attack on deep neural networkabstractDeep neural network (DNN), being able to effectively learn from a training set and provide highly accurate classification results, has become the de-facto technique used in many mission-critical systems. The security of DNN itself is therefore of great concern. In this paper, we investigate the impact of fault injection attacks on DNN, wherein attackers try to misclassify a specified input pattern into an adversarial class by modifying the parameters used in DNN via fault injection. We propose two kinds of fault injection attacks to achieve this objective. Without considering stealthiness of the attack, single bias attack (SBA) only requires to modify one parameter in DNN for misclassification, based on the observation that the outputs of DNN may linearly depend on some parameters. Gradient descent attack (GDA) takes stealthiness into consideration. By controlling the amount of modification to DNN parameters, GDA is able to minimize the fault injection impact on input patterns other than the specified one. Experimental results demonstrate the effectiveness and efficiency of the proposed attacks. Yannan Liu, Lingxiao Wei, Bo Luo, Qiang Xu 0001 |
ICCAD | 4 |
| 2017 | ApproxLUT: A novel approximate lookup table-based acceleratorabstractComputing with memory, which stores function responses of some input patterns into lookup tables offline and retrieves their values when encountering similar patterns (instead of performing online calculation), is a promising energy-efficient computing technique. No doubt to say, with a given lookup table size, the efficiency of this technique depends on which function responses are stored and how they are organized. In this paper, we propose a novel adaptive approximate lookup table based accelerator, wherein we store function responses in a hierarchical manner with increasing fine-grained granularity and accuracy. In addition, the proposed accelerator provides lightweight compensation on output results at different precision levels according to input patterns and output quality requirements. Moreover, our accelerator conducts adaptive lookup table search by exploiting input locality. Experimental results on various computation kernels show significant energy savings of the proposed accelerator over prior solutions. Ye Tian 0010, Ting Wang 0008, Qian Zhang 0020, Qiang Xu 0001 |
ICCAD | 4 |
| 2016 | ApproxMap: On task allocation and scheduling for resilient applicationsabstractMany emerging applications are inherently error-resilient and hence do not require exact computation. In this paper, we consider the task allocation and scheduling problem for mapping such applications to voltage-scalable multiprocessor systems. The proposed solution, namely ApproxMap, judiciously determines the mapping and execution sequence of resilient tasks to minimize the energy consumption of the application while meeting their target quality requirements and timing constraints. To be specific, ApproxMap generates energy-efficient yet flexible task schedule at design-time, and conducts lightweight online adjustment according to runtime dynamics for further energy-efficiency improvement. Experimental results on various task graphs demonstrate the efficacy of ApproxMap. Juan Yi, Qian Zhang 0020, Ye Tian 0010, Ting Wang 0008, Weichen Liu 0001, Edwin H.-M. Sha, Qiang Xu 0001 |
ASP-DAC | 7 |
| 2016 | On Code Execution Tracking via Power Side-ChannelabstractWith the proliferation of Internet of Things, there is a growing interest in embedded system attacks, e.g., key extraction attacks and firmware modification attacks. Code execution tracking, as the first step to locate vulnerable instruction pieces for key extraction attacks and to conduct control-flow integrity checking against firmware modification attacks, is therefore of great value. Because embedded systems, especially legacy embedded systems, have limited resources and may not support software or hardware update, it is important to design low-cost code execution tracking methods that require as little system modification as possible. In this work, we propose a non-intrusive code execution tracking solution via power-side channel, wherein we represent the code execution and its power consumption with a revised hidden Markov model and recover the most likely executed instruction sequence with a revised Viterbi algorithm. By observing the power consumption of the microcontroller unit during execution, we are able to recover the program execution flow with a high accuracy and detect abnormal code execution behavior even when only a single instruction is modified. Yannan Liu, Lingxiao Wei, Zhe Zhou 0001, Kehuan Zhang, Wenyuan Xu 0001, Qiang Xu 0001 |
CCS | 6 |
| 2016 | SlideAcross: A Low-Latency Adaptive Router for Chip Multi-processorabstractThe non-uniform distributed traffic of chip multi-processor (CMP) demands an on-chip communication infrastructure which is able to avoid congestion under high traffic conditions while possessing minimal pipeline delay at low load conditions. In this paper, we propose a low-latency adaptive router with a low-complexity single-cycle bypassing mechanism to meet the communication needs of CMPs. At low loads, this router transmits a flit using dimension-ordered routing (DoR) in the bypass datapath. When the output port required intra-dimension bypassing is not available, the packet is routed adaptively to avoid congestion. The router also has a simplified virtual channel allocation (VA) scheme that yields a non-speculative low-latency pipeline. By combining the low-complexity bypassing technique together with adaptive routing, the proposed router architecture can achieve low-latency communication under various traffic loads. Simulation shows that proposed router can reduce applications' execution time by 16.9% in average compared to low-latency router SWIFT. Wen Zong, Liang Wang 0020, Qiang Xu 0001, Michael Opoku Agyeman |
DSD | 3 |
| 2016 | Approximate Frequent Itemset Mining for streaming data on FPGAabstractFrequent Itemset Mining (FIM) is designed to find frequently occurring itemsets among a series of transactions. It is extremely memory and time expensive. Frequent Itemset Mining from a Data Stream (FIM-DS) is even more challenging since storing the infinite data to memory is infeasible. In recent years, researchers have proposed various approximation algorithms for FIM-DS. However, the computation complexity is still high, and these methods are difficult to be accelerated using hardware accelerators. In this paper, we propose a Space-Saving based approximate algorithm for FIM-DS. It avoids exponential candidates generation and comparisons. We realize a hardware accelerator design and implement it on an FPGA platform. Experimental results show that our algorithm in software implementation achieves up to 8.4× speedup for transactions with small item database, and our hardware accelerator achieves up to 50,000× speedup for transactions with small number of items, and 5.3× speedup for transactions with extremely large number of items. Guohao Dai 0001, Qiang Xu 0001, Yu Wang 0002, Huazhong Yang |
FPL | 4 |
| 2016 | DOART: A low-power and low-latency Network-on-ChipabstractPacket-switched Network-on-Chip (NoC) is the shared global communication infrastructure for future large-scale chip multi-processors (CMPs). Recently, Single-cycle Multi-hop Asynchronous Repeated Traversal (SMART) on repeater-inserted wires to reduce packet delay was proposed. But current NoC with SMART support adds complexity to conventional routers and incurs high power consumption. In this paper, we propose a low-power and low-latency NoC design with SMART support, called Dimension Ordered Asynchronous Repeated Traversal (DOART). First we design a low-power interconnect called Single-cycle Intra-dimension Bridge (SIB) with SMART support, and then we propose an efficient construction framework to connect SIBs generating a large-scale low-power and low-latency NoC. In addition, the proposed DOART supports virtual channel and is protocol and routing-level deadlock-free. Experimental results show that DOART can reduce both the application execution time and network power consumption compared with state-of-the-art NoCs with SMART support. Wen Zong, Qiang Xu 0001 |
ICCD | 2 |
| 2016 | On Effective and Efficient Quality Management for Approximate ComputingabstractApproximate computing, where computation quality is traded off for better performance and/or energy savings, has gained significant tractions from both academia and industry. With approximate computing, we expect to obtain acceptable results, but how do we make sure the quality of the final results are acceptable? This challenging problem remains largely unexplored. In this paper, we propose an effective and efficient quality management framework to achieve controlled quality-efficiency tradeoffs. To be specific, at the offline stage, our solution automatically selects an appropriate approximator configuration considering rollback recovery for large occasional errors with minimum cost under the target quality requirement. Then during the online execution, our framework judiciously determines when and how to rollback, which is achieved with cost-effective yet accurate quality predictors that synergistically combine the outputs of several basic light-weight predictors. Experimental results demonstrate that our proposed solution can achieve 11% to 23% energy savings compared to existing solutions under the target quality requirement. Ting Wang 0008, Qian Zhang 0020, Nam Sung Kim, Qiang Xu 0001 |
ISLPED | 4 |
| 2016 | Defect tolerance for CNFET-based SRAMsabstractSRAMs based on carbon nanotube field-effect transistors (CNFETs) offer a promising alternative to conventional SRAMs due to their high energy efficiency and low leakage. However, the imperfect CNT fabrication process introduces high defect rates and a unique defect distribution; these problems may offset the power/performance benefits of CNFET-based SRAMs and lead to yield degradation. We propose a redundancy architecture with asymmetrically partitioned column blocks and the sharing of spares among column blocks. We also present a analytical model to characterize the distribution of faults, which can guide the design exploration of the proposed redundancy architecture. Simulation results highlight the accuracy of the proposed model, as well as the efficiency and effectiveness of the redundancy architecture. Tianjian Li, Li Jiang 0002, Xiaoyao Liang, Qiang Xu 0001, Krishnendu Chakrabarty |
ITC | 4 |
| 2016 | A Novel Test Method for Metallic CNTs in CNFET-Based SRAMsabstractStatic random access memories (SRAMs) built on carbon nanotube field effect transistors (CNFETs) are promising alternatives to conventional CMOS-based SRAMs, due to their advantages in terms of power consumption and noise immunity. However, the nonideal carbon nanotube (CNT) fabrication process generates metallic-CNTs (m-CNTs) along with semiconductor-CNTs, leading to correlated faulty cells along the growth direction of the m-CNTs. In this paper, we propose a novel low-cost test solution to detect such faults. Instead of using conventional March test to test each and every SRAM cell, we selectively test certain SRAM cells and judiciously skip testing other SRAM cells between the selected cells. To ensure high fault coverage, we propose three jump test algorithms for different CNFET-SRAM layouts. Moreover, we model m-CNT-induced SRAM faults and characterize their distribution in the SRAM array. Experimental results show that the proposed solutions are able to achieve high fault coverage with low test cost. Tianjian Li, Xiaoyao Liang, Qiang Xu 0001, Krishnendu Chakrabarty, Naifeng Jing, Li Jiang 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2015 | On test syndrome merging for reasoning-based board-level functional fault diagnosisabstractMachine learning algorithms are advocated for automated diagnosis of board-level functional failures due to the extreme complexity of the problem. Such reasoning-based solutions, however, remain ineffective at the early stage of the product cycle, simply because there are insufficient historical data for training the diagnostic system that has a large number of test syndromes. In this paper, we present a novel test syndrome merging methodology to tackle this problem. That is, by leveraging the domain knowledge of the diagnostic tests and the board structural information, we adaptively reduce the feature size of the diagnostic system by selectively merging test syndromes such that it can effectively utilize the available training cases. Experimental results demonstrate the effectiveness of the proposed solution. Zelong Sun, Li Jiang 0002, Qiang Xu 0001, Zhaobo Zhang, Xinli Gu |
ASP-DAC | 3 |
| 2015 | Vulnerability analysis for crypto devices against probing attackabstractProbing attack is a severe threat for the security of hardware cryptographic modules (HCMs). In this paper, we make the first step to evaluate the vulnerability of HCMs against probing attack, wherein we investigate the probing complexity and the key candidate reduction capability for probing attack on every signal in the circuit. We also present approximate solutions for the calculation of the proposed metrics to reduce computational complexity. Experimental results demonstrate that the proposed evaluation metric is both effective and efficient. Lingxiao Wei, Jie Zhang 0046, Yannan Liu, Junfeng Fan, Qiang Xu 0001 |
ASP-DAC | 6 |
| 2015 | DERA: yet another differential fault attack on cryptographic devices based on error rate analysisabstractFault-injection attack is a serious threat to the security of cryptographic devices, and various differential fault analysis (DFA) techniques have been presented in the literature over the years. These attacks differ in terms of the underlining assumption on the fault models, the key distinguisher and the complexity of the associated analytical algorithm. In this work, we propose a new DFA technique that uses the inherent bias of the error rates among different signals as the foundation of the key distinguisher design, namely differential error rate analysis (DERA). Compared to existing DFA solutions, DERA is a more efficient and effective attack, in terms of both temporal and spatial needs for the attack, as demonstrated with FPGA emulation in our experiments. Yannan Liu, Jie Zhang 0046, Lingxiao Wei, Qiang Xu 0001 |
DAC | 5 |
| 2015 | Jump test for metallic CNTs in CNFET-based SRAMabstractSRAMs built on Carbon Nanotube Field Transistors (CNFET) are promising alternatives to conventional CMOS-based SRAMs, due to their advantages in terms of both power consumption and noise margin. However, non-ideal Carbon Nanotube (CNT) fabrication process generates metallic-CNTs (m-CNTs) along with semiconductor-CNTs (s-CNTs), rendering correlated faulty cells along the growth direction of the m-CNTs. Based on this phenomenon, we propose a novel testing algorithm for detecting m-CNTs, wherein consecutive write and read operations jump over multiple cells rather than marching through each and every cell, thereby significantly reducing the testing cost. The proposed jump test can be invoked before the march test to screen out those CNFET-SRAMs doomed to failure, and this can reduce the subsequent test overhead. Experimental results show that the proposed solution is able to achieve a high fault coverage with much less testing cost. Xiaoyao Liang, Qiang Xu 0001, Krishnendu Chakrabarty, Naifeng Jing, Li Jiang 0002 |
DAC | 3 |
| 2015 | On the premises and prospects of timing speculation
Rong Ye, Jie Zhang 0046, Qiang Xu 0001 |
DATE | 4 |
| 2015 | ApproxANN: an approximate computing framework for artificial neural network
Qian Zhang 0020, Ting Wang 0008, Ye Tian 0010, Qiang Xu 0001 |
DATE | 5 |
| 2015 | ApproxMA: Approximate Memory Access for Dynamic Precision ScalingabstractMotivated by the inherent error-resilience of emerging recognition, mining, and synthesis (RMS) applications, approximate computing techniques such as precision scaling has been advocated for achieving energy-efficiency gains at the cost of small accuracy loss. Most existing solutions, however, focus on the approximation of on-chip computations without considering that of off-chip data accesses, whose energy consumption may contribute to a significant portion of the total energy. In this work, we propose a novel approximate memory access technique for dynamic precision scaling, namely ApproxMA. To be specific, by taking both runtime data precision constraints and error-resilient capabilities of the application into consideration, ApproxMA determines the precision of data accesses and loads scaled data from off-chip memory for computation. Experimental results with mixture model-based clustering algorithms demonstrate the efficacy of the proposed methodology. Ye Tian 0010, Qian Zhang 0020, Ting Wang 0008, Qiang Xu 0001 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2015 | BoardPUF: Physical Unclonable Functions for Printed Circuit Board AuthenticationabstractPhysical Unclonable Functions (PUFs) are cryptographic primitives that can be used to generate volatile secret keys for cryptographic operations and enable low-cost authentication of integrated circuits. Existing PUF designs mainly exploit variation effects on silicon and hence are not readily applicable for the authentication of printed circuit boards (PCBs). To tackle the above problem, in this paper, we propose a novel PUF device that is able to generate unique and stable IDs for individual PCB, namely BoardPUF. To be specific, we embed a number of capacitors in the internal layer of PCBs and utilize their variations for key generation. Then, by integrating a cryptographic primitive (e.g. hash function) into BoardPUF, we can effectively perform PCB authentication in a challenge-response manner. Our experimental results on fabricated boards demonstrate the efficacy of BoardPUF. Lingxiao Wei, Chaosheng Song, Yannan Liu, Jie Zhang 0046, Qiang Xu 0001 |
ICCAD | 6 |
| 2015 | ApproxEigen: An Approximate Computing Technique for Large-Scale Eigen-DecompositionabstractRecognition, Mining, and Synthesis (RMS) applications are expected to make up much of the computing workloads of the future. Many of these applications (e.g., recommender systems and search engine) are formulated as finding eigenvalues/vectors of large-scale matrices. These applications are inherently error-tolerant, and it is often unnecessary, sometimes even impossible, to calculate all the eigenpairs. Motivated by the above, in this work, we propose a novel approximate computing technique for large-scale eigen-decomposition, namely ApproxEigen, wherein we focus on the practically-used Krylov subspace methods to find finite number of eigenpairs. With ApproxEigen, we provide a set of computation kernels with different levels of approximation for data pre-processing and solution finding, and conduct accuracy tuning under given quality constraints. Experimental results demonstrate that ApproxEigen is able to achieve significant energy-efficiency improvement while keeping high accuracy. Qian Zhang 0020, Ye Tian 0010, Ting Wang 0008, Qiang Xu 0001 |
ICCAD | 5 |
| 2015 | A novel TSV probing technique with adhesive test interposerabstractTSVs can be fabricated with pitch of only tens of μm, and smaller. They can be densely distributed as inter-die interconnect in 3D ICs. However, the huge mismatch between the probe technology, such as the pitch of probe head and the capacity of probe card, and the TSV fabrication technology leads to an insufficient probe on TSV tips. In this paper, we present a novel TSV probing technique that can temporally bond pre-bond die to test interposer using anisotropic conductive adhesive material. On the two sides of the test interposer, TSVs and probe heads make contact with microbumps and C4-bumps, respectively. These two types of bumps are connected using redistribution metal layers, passing through the test interposer, which can bridge the gap between feature sizes of TSVs and probe head. This probing technology is also able to increase the test bandwidth by enlarging the test interposer and redistributing test signals between microbumps and C4-bumps. Moreover, the number of probe-card touchdown can be reduced by sharing the test interposer among multiple dies during the wafer-level testing. Simulation results on the corresponding test structures for TSVs open fault and leakage fault show the great test resolution and robustness considering different design choices and variable design parameters among the test structures. Li Jiang 0002, Xiangwei Huang, Hongfeng Xie, Qiang Xu 0001, Chao Li 0009, Xiaoyao Liang, Huiyun Li |
ICCD | 4 |
| 2015 | On Resilient System Performance BinningabstractBy allowing timing errors to occur and recovering them online, resilient systems can be used to eliminate the voltage/frequency guardband to improve energy efficiency/throughput. Due to the nature of fault tolerance computing, resilient systems have a different binning strategy from traditional circuits. In this paper, we study, for the first time, the binning metrics of resilient systems. We propose a solution for resilient system binning based on structural at-speed delay testing. Then an adaptive clock configuration technique is proposed for yield improvement. Experimental results demonstrate the effectiveness of our proposed binning method, and significant yield improvement accomplished by the adaptive clock configuration technique. Jianghao Guo, Qiang Xu 0001, Wen-Ben Jone |
ISPD | 3 |
| 2015 | On diagnosable and tunable 3D clock network design for lifetime reliability enhancementabstractIn three-dimensional (3D) integrated circuits (IC-s), many clock-TSVs are deployed to deliver clock signals to different tiers with minimum skews. However, these clock-TSVs are prone to aging effects, such as thermal-mechanical stress and electromigration, rendering hard-to-predict clock skews at runtime. These skews have a wide range of influence on the flip-flops, and may violate the safety margins of critical paths in the circuit. Besides the circuit aging effect, the clock-TSV induced skews pose another threat to the circuit lifetime reliability. To tackle this problem, we propose to put tunable buffer for each clock-TSV in the clock network, and introduce an efficient algorithm to place aging sensors in the circuit at design stage. Then, at runtime, we conduct online diagnosis and apply effective clock tuning algorithms based on the triggered alarms in the aging sensors. Experimental results on a post-layout 3D circuit show that the proposed solution is able to significantly improve the lifetime reliability of 3D ICs. Li Jiang 0002, Pu Pang, Naifeng Jing, Sung Kyu Lim, Xiaoyao Liang, Qiang Xu 0001 |
ITC | 6 |
| 2015 | Yield and reliability enhancement for 3D ICs: Dissertation summary: IEEE TTTC E.J. McCluskey doctoral thesis award competition finalistabstractIn order to improve the assembly yield in Through-Silicon Via (TSV) fabrication process, we develop a fault model considering TSV coupling effect that has not been carefully investigated before. It leads our attention to a unique phenomenon, i.e., the faulty TSVs can be clustered. Thus, we propose a novel TSV redundancy architecture, composed of a light-weight switch design and two effective and efficient repair algorithms. Experimental results show that they have much higher repair rate than the existing solutions w.t. or w.o clustered TSV faults. We extend this TSV redundancy architecture to address two critical challenges: catastrophic TSV failure with extremely clustered TSV faults; circuit timing error that occurs in the runtime induced by the latent TSV defects. These works provide a systematic way to improve the yield and reliability for 3D ICs, which is of significance impact, as a percentage increase of the yield means millions of dollars. Li Jiang 0002, Qiang Xu 0001 |
ITC | 2 |
| 2015 | FASTrust: Feature analysis for third-party IP trust verificationabstractThird-party intellectual property (3PIP) cores are widely used in integrated circuit designs. It is essential and important to ensure their trustworthiness. Existing hardware trust verification techniques suffer from high computational complexity, low extensibility, and inability to detect implicitly-triggered hardware trojans (HTs). To tackle the above problems, in this paper, we present a novel 3PIP trust verification framework, named FASTrust, which conducts HT feature analysis on the flip-flop level control-data flow graph (CDFG) of the circuit. FASTrust is not only able to identify existing explicitly-triggered and implicitly-triggered HTs appeared in the literature in an efficient and effective manner, but more importantly, it also has the unique advantage of being scalable to defend against future and more stealthy HTs by adding new features to the system. Xiaoming Chen 0003, Jie Zhang 0046, Qiaoyi Liu, Jia Wang 0004, Qiang Xu 0001, Yu Wang 0002, Huazhong Yang |
ITC | 6 |
| 2015 | Fault-Tolerant 3D-NoC Architecture and Design: Recent Advances and ChallengesabstractIn this paper, we survey recent research work in the design of fault-tolerant three-dimensional (3D) network-on-chip (NoC), which has drawn lots of research attention from both academia and industry. To be specific, we discuss the emerging defects introduced in 3D integration, the state-of-the-art fault-tolerant 3D router designs, various fault-tolerant routing algorithms in three-dimension, as well as the architecture and design methodologies to tolerate defective TSVs in 3D-NoC. Finally, we highlight open challenges and future research directions in this domain. Li Jiang 0002, Qiang Xu 0001 |
NOCS | 2 |
| 2015 | VeriTrust: Verification for Hardware TrustabstractToday's integrated circuit designs are vulnerable to a wide range of malicious alterations, namely hardware Trojans (HTs). HTs serve as backdoors to subvert or augment the normal operation of infected devices, which may lead to functionality changes, sensitive information leakages, or denial of service attacks. To tackle such threats, this paper proposes a novel verification technique for hardware trust, namely VeriTrust, which facilitates to detect HTs inserted at design stage. Based on the observation that HTs are usually activated by dedicated trigger inputs that are not sensitized with verification test cases, VeriTrust automatically identifies such potential HT trigger inputs by examining verification corners. The key difference between VeriTrust and existing HT detection techniques based on “unused circuit identification” is that VeriTrust is insensitive to the implementation style of HTs. Experimental results show that VeriTrust is able to detect all HTs evaluated in this paper (constructed based on various HT design methodologies shown in this paper) at the cost of moderate extra verification time. Jie Zhang 0046, Lingxiao Wei, Yannan Liu, Qiang Xu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2015 | A Low-Cost TSV Test and Diagnosis Scheme Based on Binary Search MethodabstractTesting through-silicon-vias (TSVs) is challenging largely due to the dimension gap between the TSVs and the probe needles. This paper proposes a low-cost and efficient test and diagnosis scheme without extra design for test structure or special probe technologies. A test probe head with current technology is in use, contacting multiple TSVs simultaneously. We propose an efficient binary search-based algorithm to guide the probe card movement and diagnose the faulty TSVs. To evaluate the test efficiency, mathematical analyses are performed considering geometric, probabilistic, and electronic issues including process variations and different probe architectures. The analysis demonstrates that the test efficiency can be largely enhanced by choosing appropriate sizes of probe needles, and by adjusting iteration algorithms according to the distribution of faulty TSVs. Huiyun Li, Li Jiang 0002, Qiang Xu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2014 | DeTrust: Defeating Hardware Trust Verification with Stealthy Implicitly-Triggered Hardware TrojansabstractHardware Trojans (HTs) inserted at design time by malicious insiders on the design team or third-party intellectual property (IP) providers pose a serious threat to the security of computing systems. Researchers have proposed several hardware trust verification techniques to mitigate such threats, and some of them are shown to be able to effectively flag all suspicious HTs implemented in the Trust-Hub hardware backdoor benchmark suite. No doubt to say, adversaries would adjust their tactics of attacks accordingly and it is hence essential to examine whether new types of HTs can be designed to defeat these hardware trust verification techniques. In this paper, we present a systematic HT design methodology to achieve the above objective, namely \emph{DeTrust}. Given an HT design, DeTrust keeps its original malicious behavior while making the HT resistant to state-of-the-art hardware trust verification techniques by manipulating its trigger designs. To be specific, DeTrust implements stealthy implicit triggers for HTs by carefully spreading the trigger logic into multiple sequential levels and combinational logic blocks and combining the trigger logic with the normal logic, so that they are not easily differentiable from normal logic. As shown in our experimental results, adversaries can easily employ DeTrust to evade hardware trust verification. Jie Zhang 0046, Qiang Xu 0001 |
CCS | 3 |
| 2014 | On the Simulation of NBTI-Induced Performance Degradation Considering Arbitrary Temperature and Voltage VariationsabstractWith aggressive CMOS technology scaling, Negative Bias Temperature Instability (NBTI) has emerged as one of the major system lifetime reliability threats, which gradually increases PMOS transistor threshold voltage and hence results in increased circuit delay. NBTI-induced performance degradation depends heavily on time-varying parameters such as temperature, duty cycle and supply voltage. Previous analytical models for NBTI effects, however, cannot cover all these parameters, causing overly optimistic or overly pessimistic analysis. In this work, we propose a comprehensive NBTI analytical model that explicitly takes supply voltage, duty cycle and temperature variations into consideration. The accuracy of the proposed model is validated against cycle-accurate simulation for NBTI effects. In addition, based on the proposed model, we present an efficient simulation framework for system lifetime prediction by running representative workloads only. Experimental results demonstrate the efficacy and efficiency of the proposed solution. Ting Wang 0008, Qiang Xu 0001 |
DAC | 2 |
| 2014 | ApproxIt: An Approximate Computing Framework for Iterative MethodsabstractApproximate computing, being able to tradeoff computation quality (e.g., accuracy) and computational effort (e.g., energy) for error-tolerant applications such as media processing and the emerging Recognition, Mining, and Synthesis (RMS) applications, has gained significant traction in recent years. Many of these applications employ iterative methods for solution-finding, wherein a sequence of improving approximate solutions are generated before reaching the final converged solution. In this work, we propose ApproxIt, a novel approximate computing framework for iterative methods with quality guarantees. To be specific, we present a lightweight quality estimator that is able to capture the solution quality of each iteration and use it to guide the selection of approximate computing mode in the next iteration. With the proposed dynamic effort scaling technique, ApproxIt is able to dramatically improve application energy efficiency under quality guarantees, as demonstrated in our experimental results. Qian Zhang 0020, Rong Ye, Qiang Xu 0001 |
DAC | 4 |
| 2014 | On hybrid memory allocation for FPGA behavioral synthesis (abstract only)abstractFPGA behavioral synthesis has gained significant momentum recently with the growing interests in accelerating high-performance computing applications. While the latest generation of high-level synthesis (HLS) tools has made significant progress, they still lack the support for certain high-level language features such as dynamic memory allocation, despite the fact that efficiently utilization of the on-chip memory resources in FPGAs is critical to achieve the performance and power consumption target for many designs. Qian Zhang 0020, Chenfei Ma, Qiang Xu 0001 |
FPGA | 3 |
| 2014 | On trojan side channel design and identificationabstractTrojan side channels (TSCs) are serious threats to the security of cryptographic systems because they facilitate to leak secret keys to attackers via covert side channels that are unknown to designers. To tackle this problem, we present a new hardware Trojan detection technique for TSCs. To be specific, we first investigate general power-based TSC designs and discuss the tradeoff between their hardware cost and the complexity of the key cracking process. Next, we present our TSC identification technique based on the correlation between the key and the covert physical side channels used by attackers. Experimental results demonstrate the effectiveness of the proposed solution. Jie Zhang 0046, Guantong Su, Yannan Liu, Lingxiao Wei, Guoqiang Bai 0001, Qiang Xu 0001 |
ICCAD | 7 |
| 2014 | Learning-Based Power Management for Multicore Processors via Idle Period ManipulationabstractLearning-based dynamic power management (DPM) techniques, being able to adapt to varying system conditions and workloads, have attracted a lot of research attention recently. To the best of our knowledge, however, none of the existing learning-based DPM solutions are dedicated to power reduction in multicore processors, although they can be utilized by treating each processor core as a standalone entity and conducting DPM for them separately. In this paper, by including task allocation into our learning-based DPM framework for multicore processors, we are able to manipulate idle periods on processor cores to achieve a better tradeoff between power consumption and system performance. Experimental results show that the proposed solution significantly outperforms existing DPM techniques. Rong Ye, Qiang Xu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2013 | On effective and efficient in-field TSV repair for stacked 3D ICsabstractThree-dimensional (3D) integration based on through-silicon-vias (TSVs) is rapidly gaining traction for industry adoption. However, manufacturing processes for TSVs have been shown to introduce new failure mechanisms. In particular, thermo-mechanical stress and electromigration introduce reliability threats for TSVs, e.g., voids and interfacial cracks, which can lead to hard-to-predict timing errors on critical paths with TSVs, thereby resulting in accelerated chip failure in the field. Burn-in for screening latent defects during manufacturing is expensive and its effectiveness for new TSV defect types has yet to be thoroughly characterized. We describe a reconfigurable in-field repair solution that is able to effectively tolerate latent TSV defects through the judicious use of spares. The proposed solution includes a reconfigurable repair architecture that enables spare TSV sharing between TSV grids, and the corresponding in-field repair algorithms. The effectiveness and efficiency of our proposed solution is evaluated using 3D benchmark designs. Li Jiang 0002, Fangming Ye, Qiang Xu 0001, Krishnendu Chakrabarty, Bill Eklow |
DAC | 3 |
| 2013 | Post-placement voltage island generation for timing-speculative circuitsabstractRegion-based multi-supply voltage (MSV) design, by which circuits are partitioned into multiple "voltage islands" and each island operates at a supply voltage that meets its own performance requirement, is an effective technique to tradeoff power and performance. Different from conventional voltage island generation techniques that work in a conservative manner to guarantee "always correct" computation, in this work, we investigate the MSV design problem for timing-speculative circuits, which achieves high energy-efficiency by allowing the occurrence of infrequent timing errors and correcting them online. A novel algorithm based on dynamic programming is developed to tackle this problem. Experimental results on various benchmark circuits demonstrate the effectiveness of the proposed methodology. Rong Ye, Zelong Sun, Wen-Ben Jone, Qiang Xu 0001 |
DAC | 5 |
| 2013 | On testing timing-speculative circuitsabstractBy allowing the occurrence of infrequent timing errors and correcting them online, circuit-level timing speculation is one of the most promising variation-tolerant design techniques. How to effectively test timing-speculative circuits, however, has not been addressed in the literature. This is a challenging problem because conventional scan techniques cannot provide sufficient controllability and observability for such circuits. In this paper, we propose novel techniques to achieve high fault coverage for timing-speculative circuits without incurring high design-for-testability cost. Experimental results on various benchmark circuits demonstrate the effectiveness of the proposed solution. Yannan Liu, Wen-Ben Jone, Qiang Xu 0001 |
DAC | 4 |
| 2013 | InTimeFix: a low-cost and scalable technique for in-situ timing error masking in logic circuitsabstractWith technology scaling, integrated circuits (ICs) suffer from increasing process, voltage, and temperature (PVT) variations and adverse aging effects. In most cases, these reliability threats manifest themselves as timing errors on critical speed-paths of the circuit, if a large design guard band is not reserved. This work presents a novel in-situ timing error masking technique, namely InTimeFix, by introducing fine-grained redundant approximation circuit into the design to provide more timing slack for speed-paths. The synthesis of the redundant circuit relies on simple structural analysis of the original circuit, which is easily scalable to large IC designs. Experimental results show that InTimeFix significantly increases circuit timing slack with low area/power cost. Qiang Xu 0001 |
DAC | 2 |
| 2013 | VeriTrust: verification for hardware trustabstractHardware Trojans (HTs) implemented by adversaries serve as backdoors to subvert or augment the normal operation of infected devices, which may lead to functionality changes, sensitive information leakages, or Denial of Service attacks. To tackle such threats, this paper proposes a novel verification technique for hardware trust, namely VeriTrust, which facilitates to detect HTs inserted at design stage. Based on the observation that HTs are usually activated by dedicated trigger inputs that are not sensitized with verification test cases, VeriTrust automatically identifies such potential HT trigger inputs by examining verification corners. The key difference between VeriTrust and existing HT detection techniques is that VeriTrust is insensitive to the implementation style of HTs. Experimental results show that VeriTrust is able to detect all HTs evaluated in this paper (constructed based on various HT design methodologies shown in the literature) at the cost of moderate extra verification time, which is not possible with existing solutions. Jie Zhang 0046, Lingxiao Wei, Zelong Sun, Qiang Xu 0001 |
DAC | 5 |
| 2013 | Optimization for timing-speculated circuits by redundancy addition and removalabstractIntegrated circuits suffer from severe variation effects with technology scaling, making their timing behavior increasingly unpredictable. Timing speculation is a promising technique to tackle this problem with the help of online timing error detection and correction mechanisms. In this paper, we propose to use redundancy addition and removal (RAR) technique to optimize timing-speculated circuits. By intentionally removing wires on those frequently-exercised critical paths and replacing them with wires on less critical ones (if possible), the proposed technique is able to greatly reduce the timing error rate of the circuit and improve its overall throughput, as shown in our experimental results on various benchmark circuits. Rong Ye, Qiang Xu 0001 |
ETS | 4 |
| 2013 | On reconfiguration-oriented approximate adder design and its applicationabstractApproximate circuit designs allow us to tradeoff computation quality (e.g., accuracy) and computational effort (e.g., energy), by exploiting the inherent error-resilience of many applications. As the computation quality requirement of an application generally varies at runtime, it is preferable to be able to reconfigure approximate circuits to satisfy such needs and save unnecessary computational effort. In this paper, we present a reconfiguration-oriented design methodology for approximate circuits, and propose a reconfigurable approximate adder design that degrades computation quality gracefully. The proposed design methodology enables us to achieve better quality-effort tradeoff when compared to existing techniques, as demonstrated in the application of DCT computing. Rong Ye, Ting Wang 0008, Rakesh Kumar 0002, Qiang Xu 0001 |
ICCAD | 5 |
| 2013 | ForTER: a forward error correction scheme for timing error resilienceabstractWith technology scaling, integrated circuits suffer from increasingly severe static and dynamic variations, which often manifest themselves as infrequent timing errors on circuit speed paths, if a large timing guard-band is not reserved. This paper presents a new forward timing error correction scheme, namely ForTER, which predicts whether the occurrence of timing errors would propagate to the next level of sequential elements and corrects them without necessarily borrowing timing slack. The proposed technique can be combined with other timing error resilient circuit design techniques to further improve circuit performance, as demonstrated in our experimental results with various benchmark circuits. Jie Zhang 0046, Rong Ye, Qiang Xu 0001 |
ICCAD | 4 |
| 2013 | AgentDiag: An agent-assisted diagnostic framework for board-level functional failuresabstractDiagnosing functional failures in complicated electronic boards is a challenging task, wherein debug technicians try to identify defective components by analyzing some syndromes obtained from the application of diagnostic tests. The diagnosis effectiveness and efficiency rely heavily on the quality of the in-house developed diagnostic tests and the debug technicians' knowledge and experience, which, however, have no guarantees nowadays. To tackle this problem, we propose a novel agent-assisted diagnostic framework for board-level functional failures, namely AgentDiag, which facilitates to evaluate the quality of the diagnostic tests and bridge the knowledge gap between the diagnostic programmers who write diagnostic tests and the debug technicians who conduct in-field diagnosis with a lightweight model of the boards and tests. Experimental results on a real industrial board and an OpenRISC design demonstrate the effectiveness of the proposed solution. Zelong Sun, Li Jiang 0002, Qiang Xu 0001, Zhaobo Zhang, Xinli Gu |
ITC | 3 |
| 2013 | On Effective Through-Silicon Via Repair for 3-D-Stacked ICsabstract3-D-stacked integrated circuits (ICs) that employ through-silicon vias (TSVs) to connect multiple dies vertically have gained wide-spread interest in the semiconductor industry. In order to be commercially viable, the assembly yield for 3-D-stacked ICs must be as high as possible, requiring TSVs to be reparable. Existing techniques typically assume TSV faults to be uniformly distributed and use neighboring TSVs to repair faulty ones, if any. In practice, however, clustered TSV faults are quite common due to the fact that the TSV bonding quality depends on surface roughness and cleanness of silicon dies, rendering prior TSV redundancy solutions less effective. Furthermore, existing techniques consume a lot of redundant TSVs that are still costly in the current TSV process. This inefficient TSV redundancy can limit the amount of TSVs that is allowed to use and may even become the obstacle to commercial production. To resolve this problem, we present a novel TSV repair framework, including a hardware redundancy architecture that enables faulty TSVs to be repaired by redundant TSVs that are farther apart, the corresponding repair algorithm and the redundancy architecture construction. By doing so, the manufacturing yield for 3-D-stacked ICs can be dramatically improved, as demonstrated in our experimental results. Li Jiang 0002, Qiang Xu 0001, Bill Eklow |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2013 | On Multiplexed Signal Tracing for Post-Silicon ValidationabstractTrace-based debug techniques have widely been utilized in the industry to eliminate design errors escaped from pre-silicon verification. Existing solutions typically trace the same set of signals throughout each debug run, which is not quite effective for catching design errors. In this paper, we propose a multiplexed signal tracing strategy that is able to significantly increase debuggability of the circuit. That is, we divide the tracing procedure in each debug run into a few periods and trace different sets of signals in each period. We present a trace signal grouping algorithm to maximize the probability of catching the propagated evidences from design errors, considering the trace interconnection fabric design constraints. Moreover, we propose a trace signal selection solution to enhance the error detection capability. Experimental results on benchmark circuits demonstrate the effectiveness of the proposed solution. Xiao Liu 0011, Qiang Xu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2012 | Yield enhancement for 3D-stacked ICs: Recent advances and challengesabstractThree-dimensional (3D) integrated circuits (ICs) that stack multiple dies vertically using through-silicon vias (TSVs) have gained wide interests of the semiconductor industry. The shift towards volume production of 3D-stacked ICs, however, requires their manufacturing yield to be commercially viable. Various techniques have been presented in the literature to address this important problem, including pre-bond testing techniques to tackle the “known good die” problem, TSV redundancy designs to provide defect-tolerance, and wafter/die matching solutions to improve the overall stack yield. In this paper, we survey recent advances in this filed and point out challenges to be resolved in the future. Qiang Xu 0001, Li Jiang 0002, Huiyun Li, Bill Eklow |
ASP-DAC | 1 |
| 2012 | Learning-based power management for multi-core processors via idle period manipulationabstractLearning-based dynamic power management (DPM) techniques, being able to adapt to varying system conditions and workloads, have attracted lots of research attention recently. To the best of our knowledge, however, none of the existing learning-based DPM solutions are dedicated to power reduction in multi-core processors, although they can be utilized by treating each processor core as a standalone entity and conducting DPM for them separately. In this work, by including task allocation into our learning-based DPM framework for multi-core processors, we are able to manipulate idle periods on processor cores to achieve a better tradeoff between power consumption and system performance. Experimental results show that the proposed solution significantly outperforms existing DPM techniques. Rong Ye, Qiang Xu 0001 |
ASP-DAC | 2 |
| 2012 | CODA: A concurrent online delay measurement architecture for critical pathsabstractWith technology scaling, integrated circuits behave more unpredictably due to process variation, environmental changes and aging effects. Various variation-aware and adaptive design methodologies have been proposed to tackle this problem. Clearly, more effective solutions can be obtained if we are able to collect real-time information such as the actual propagation delay of critical paths when the circuit is running in normal function mode. Motivated by the above, in this paper, we propose a novel concurrent online delay measurement architecture for critical paths, namely CODA, to facilitate this task. Experimental results demonstrate high accuracy and practicality of the proposed technique. Yubin Zhang, Haile Yu, Qiang Xu 0001 |
ASP-DAC | 3 |
| 2012 | X-tracer: a reconfigurable X-tolerant trace compressor for silicon debugabstractThe effectiveness of at-speed silicon debug is constrained by the limited trace buffer size and/or trace port bandwidth, requiring highly-efficient trace data compression solutions. As it is usually inevitable to have unknown 'X' values during silicon debug, trace compressor should be equipped with X-tolerance feature in order not to significantly degrade error detection capability. To tackle this problem, this paper presents a novel reconfigurable X-tolerant trace compressor, namely X-Tracer, which is able to tolerate as many X-bits as possible in the trace streams while guaranteeing high compression ratio, at the cost of little extra design-for-debug hardware. Experimental results on benchmark circuits demonstrate the effectiveness of the proposed technique. Xiao Liu 0011, Qiang Xu 0001 |
DAC | 3 |
| 2012 | On effective TSV repair for 3D-stacked ICsabstract3D-stacked ICs that employ through-silicon vias (TSVs) to connect multiple dies vertically have gained wide-spread interest in the semiconductor industry. In order to be commercially viable, the assembly yield for 3D-stacked ICs must be as high as possible, requiring TSVs to be reparable. Existing techniques typically assume TSV faults to be uniformly distributed and use neighboring TSVs to repair faulty ones, if any. In practice, however, clustered TSV faults are quite common due to the fact that the TSV bonding quality depends on surface roughness and cleaness of silicon dies, rendering prior TSV redundancy solutions less effective. To resolve this problem, we present a novel TSV repair framework, including a hardware architecture that enables faulty TSVs to be repaired by redundant TSVs that are farther apart, and the corresponding repair algorithm. By doing so, the manufacturing yield for 3D-stacked ICs can be dramatically improved, as demonstrated in our experimental results. Li Jiang 0002, Qiang Xu 0001, Bill Eklow |
DATE | 2 |
| 2012 | Clock skew scheduling for timing speculationabstractBy assigning intentional clock arrival times to the sequential elements in a circuit, clock skew scheduling (CSS) techniques can be utilized to improve IC performance. Existing CSS solutions work in a conservative manner that guarantees “always correct” computation, and hence their effectiveness is greatly challenged by the ever-increasing process variation effects. By allowing infrequent timing errors and recovering from them with minor performance impact, timing speculation techniques such as Razor have gained wide interests from both academia and industry. In this work, we formulate the clock skew scheduling problem for circuits equipped with timing speculation capability and propose a novel CSS algorithm based on gradient-descent method. Experimental results on various benchmark circuits demonstrate the effectiveness of our proposed methodology. Rong Ye, Hai Zhou 0001, Qiang Xu 0001 |
DATE | 4 |
| 2012 | On logic synthesis for timing speculationabstractBy allowing the occurrence of infrequent timing errors and correcting them with rollback mechanisms, the so-called timing speculation (TS) technique can significantly improve circuit energy-efficiency and hence has become one of the most promising solutions to mitigate the ever-increasing variation effects in nanometer technologies. As timing error recovery incurs non-trivial performance/energy overhead, it is important to reshape the delay distribution of critical paths in timing-speculated circuits to minimize their timing error rates. Most existing TS optimization techniques achieve this objective with post-synthesis techniques such as gate sizing or body biasing. In this work, we propose to conduct logic synthesis for timing-speculated circuits from the ground up. Being able to manipulate circuit structures during logic optimization, the proposed solution is able to dramatically reduce circuit timing error rates and hence improve its throughput, as demonstrated with experimental results on various benchmark circuits. Rong Ye, Rakesh Kumar 0002, Qiang Xu 0001 |
ICCAD | 5 |
| 2012 | On efficient silicon debug with flexible trace interconnection fabricabstractTrace-based debug solutions facilitate to eliminate bugs escaped from pre-silicon verification and have gained wide acceptance in the industry. Generally speaking, a number of “key” signals in the circuit are tapped, but not all of them can be observed at the same time due to the limited trace bandwidth. Therefore, a trace interconnection fabric is utilized to output either a subset of signals with multiplexor (MUX) network or compressed signatures with XOR network to the trace memory/port in each debug run. However, both kinds of trace interconnection fabrics have limitations. On one hand, with MUX-based fabric, the visibility of the circuit is limited and it requires many debug runs to locate errors. On the other hand, with XOR-based fabric, typically clean “golden vectors” (i.e, without unknown bits) are required so that signatures are not corrupted. In this paper, we propose a flexible trace interconnection fabric design that is able to overcome the above limitations, at the cost of little extra design-for-debug hardware. Experimental results on benchmark circuits demonstrate the effectiveness of the proposed technique. Xiao Liu 0011, Qiang Xu 0001 |
ITC | 2 |
| 2012 | On modeling faults in FinFET logic circuitsabstractFinFET transistor has much better short-channel characteristics than traditional planar CMOS transistor and will be widely used in next generation technology. Due to its significant structural difference from conventional planar devices, it is essential to revisit whether existing fault models are applicable to detect faults in FinFET logic gates. In this paper, we study some unique defects in FinFET logic circuits and simulate their faulty behavior. Our simulation study shows that most of the defects can be covered with existing fault models, but they vary under different cases and test strategies may need to be augmented to target them. Qiang Xu 0001 |
ITC | 2 |
| 2012 | On Signal Selection for Visibility Enhancement in Trace-Based Post-Silicon ValidationabstractToday's complex integrated circuit designs increasingly rely on post-silicon validation to eliminate bugs that escape from pre-silicon verification. One effective silicon debug technique is to monitor and trace the behaviors of the circuit during its normal operation. However, due to the associated overhead, designers can only afford to trace a small number of signals in the design. Selecting which signals to trace is therefore a crucial issue for the effectiveness of this technique. This paper proposes an automated trace signal selection strategy with a new probability-based evaluation metric, which is able to dramatically enhance the visibility in post-silicon validation. Experimental results on benchmark circuits show that the proposed technique is more effective than existing solutions. Xiao Liu 0011, Qiang Xu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2012 | On X-Variable Filling and Flipping for Capture-Power Reduction in Linear Decompressor-Based Test Compression EnvironmentabstractExcessive test power consumption and growing test data volume are both serious concerns for the semiconductor industry. Various low-power X-filling techniques and test data compression schemes were developed accordingly to address the above problems. These methods, however, often exploit the very same “don't-care” bits in the test cubes to achieve different objectives and hence may contradict each other. In this paper, we propose novel techniques to reduce scan capture power in linear decompressor-based test compression environment, by employing algorithmic solutions to fill and flip X-variables supplied to the linear decompressor. Experimental results on benchmark circuits demonstrate that our proposed techniques significantly outperform existing solutions. Xiao Liu 0011, Qiang Xu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2012 | Integrated Test-Architecture Optimization and Thermal-Aware Test Scheduling for 3-D SoCs Under Pre-Bond Test-Pin-Count ConstraintabstractWe propose a layout-driven test-architecture design and optimization technique for core-based system-on-chips (SoCs) that are fabricated with three-dimensional (3-D) integration technology. In contrast to prior work, we consider the pre-bond test-pin-count constraint during optimization since these pins occupy large silicon area that cannot be used in functional mode. In addition, the proposed test-architecture design takes the SoC layout into consideration and facilitates the sharing of test wires between pre-bond tests and post-bond test, which significantly reduces the routing cost for test-access mechanisms. In addition, a thermal-aware test scheduling algorithm is proposed to eliminate hot spots during manufacturing test. Experimental results for the ITC'02 SoC benchmarks circuits demonstrate the effectiveness of the proposed solution. Li Jiang 0002, Qiang Xu 0001, Krishnendu Chakrabarty |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2012 | An FPGA Chip Identification Generator Using Configurable Ring OscillatorsabstractPhysically unclonable functions (PUF) are commonly used in applications such as hardware security and intellectual property protection. Various PUF implementation techniques have been proposed to translate chip-specific variations into a unique binary string. It is difficult to maintain repeatability of chip ID generation, especially over a wide range of operating conditions. To address this problem, we propose utilizing configurable ring oscillators and an orthogonal re-initialization scheme to improve repeatability. An implementation on a Xilinx Spartan-3e field-programmable gate array was tested on nine different chips. Experimental results show that the bit flip rate is reduced from 1.5% to approximately 0 at a fixed supply voltage and room temperature. Over a 20°C-80°C temperature range and 25% variation in supply voltage, the bit flip rate is reduced from 1.56% to 3.125×10-7. Haile Yu, Philip H. W. Leong, Qiang Xu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | Compression-aware capture power reduction for at-speed testingabstractTest compression has become a de facto technique in VLSI testing. Meanwhile, excessive capture power of at-speed testing has also become a serious concern. Therefore, it is important to co-optimize test power and compression ratio in at-speed testing. In this paper, a novel X-filling framework is proposed to reduce capture power of both LoC and LoS at-speed testing, which is applicable for different test compression schemes. The proposed technology has been validated by the experimental results on larger ITC'99 benchmark circuits. Jia Li 0022, Qiang Xu 0001 |
ASP-DAC | 2 |
| 2011 | Customer-aware task allocation and scheduling for multi-mode MPSoCsabstractToday's multiprocessor system-on-a-chip (MPSoC) products typically have multiple execution modes, and for each mode, all the products utilize the same task allocation and schedule strategy determined at design stage. As these products experience different usages by customers, such unified solution can at best be optimized for a hypothetical common case. It is hence likely that the product is not reliable or energy-efficient from particular customers' point of view. To tackle this problem, we propose a novel customer-aware task allocation and scheduling technique, wherein we generate an initial task schedule for each execution mode at design stage and then perform online adjustment at regular intervals for lifetime reliability improvement and/or energy reduction according to the specific usage strategy of individual products. Experimental results on several hypothetical MPSoCs with various task graphs demonstrate the effectiveness of the proposed personalized solution. Lin Huang 0002, Rong Ye, Qiang Xu 0001 |
DAC | 3 |
| 2011 | Re-synthesis for cost-efficient circuit-level timing speculationabstractAs the transistor feature size is continuously scaled down, integrated circuits are more vulnerable to process, voltage and temperature (PVT) variations, causing infrequent timing errors. Various techniques have been proposed to tackle this problem and circuit-level timing speculation is one of the most promising solutions. However, directly applying such technique can be quite costly in terms of area overhead and energy consumption. In this paper, we propose cost-efficient re-synthesis solutions to tackle this problem. We try to reduce the number of suspicious flip-flops (FFs) that might have timing errors by retiming techniques, which relocate some suspicious FFs without increasing critical path delay. An efficient and effective algorithm is then utilized to pad those short paths linking the remaining suspicious FFs to ensure the functional correctness of timing speculators. Experimental results show that the proposed solution can achieve significant area reduction for timing speculation. Qiang Xu 0001 |
DAC | 3 |
| 2011 | On multiplexed signal tracing for post-silicon debugabstractTrace-based debug solutions facilitate to eliminate design errors escaped from pre-silicon verification and have gained wide acceptance in the industry. Existing techniques typically trace the same set of signals throughout each debug run, which is not quite effective for catching design errors. In this work, we propose a multiplexed signal tracing strategy that is able to significantly increase debuggability of the circuit. That is, we divide the tracing procedure in each debug run into a few periods and trace different sets of signals in each period. A novel trace signal grouping algorithm is presented to maximize the probability of catching the propagated evidences from design errors, considering the trace interconnection fabric design constraints. Experimental results on benchmark circuits demonstrate the effectiveness of proposed solution. Xiao Liu 0011, Qiang Xu 0001 |
DATE | 2 |
| 2011 | On High-Quality Test Pattern Selection and ManipulationabstractWith a given test set, this paper is concerned about selecting a subset of test patterns and manipulating them into high-quality ones that reduce both under-testing and over-testing, which are both serious concerns for the semiconductor industry with technology scaling. Experimental results on IWLS'05 benchmark demonstrate the effectiveness of the proposed solution. Xiao Liu 0011, Qiang Xu 0001 |
ETS | 3 |
| 2011 | On timing yield improvement for FPGA designs using architectural symmetry (abstract only)abstractAs semiconductor manufacturing technology continues towards reduced feature sizes, timing yield will degrade due to increased process variation. Traditional variation aware design (VAD) methodologies address this problem by using chipwise placement and routing optimizations given the variation distribution is obtained. However, it is very time-consuming to do chipwise variation characterization and optimization. Therefore, this work proposes the use of symmetry in FPGA architectures so that a large range of timing-equivalent configurations can be derived from a single initial implementation by configuration rotation and flipping, allowing the application of post-silicon tuning to mitigate the effects of process variation. Additionally, logic element swaps further improve timing performance. An FPGA design methodology is presented which combines configuration-level redundancy and fine-grained design tuning. The proposed methodology does not need variation characterization and customized placement and routing for each individual FPGA. Compared to other variation aware design methods, it is more cost-efficient in terms of run-time, especially for design implementation on a large amount of FPGAs. Twenty MCNC benchmark circuits in different process technologies were used to show that the proposed method is effective in improving yield and timing in the presence of process variation. Haile Yu, Qiang Xu 0001, Philip H. W. Leong |
FPGA | 2 |
| 2011 | On Timing Yield Improvement for FPGA Designs Using Architectural SymmetryabstractAs semiconductor manufacturing technology continues towards reduced feature sizes, timing yield will degrade due to increased process variation. This work proposes the use of architectural symmetry in FPGA so that multiple timing-equivalent configurations can be derived from a single initial implementation, allowing the application of post-silicon tuning to mitigate process variation effects. Experimental results on twenty MCNC benchmark circuits for various process technologies demonstrate timing yield improvement using the proposed method. Haile Yu, Qiang Xu 0001, Philip H. W. Leong |
FPL | 2 |
| 2011 | Online clock skew tuning for timing speculationabstractThe timing performance and yield of integrated circuits can be improved by carefully assigning intentional clock skews to flip-flops. Due to the ever-increasing process, voltage, and temperature variations with technology scaling, however, traditional clock skew optimization solutions that work in a conservative manner to guarantee “always correct” computation cannot perform as well as expected. By allowing infrequent timing errors and recovering from them with minor performance impact, the concept of timing speculation has attracted lots of research attention since it enables “better than worst-case design”. In this work, we propose a novel online clock skew tuning technique for circuits equipped with timing speculation capability. By observing the occurrence of timing errors at runtime and tuning clock skews accordingly, the proposed technique is able to achieve much better timing performance when compared to existing clock skew optimization solutions. Experimental results on various benchmark circuits demonstrate the effectiveness of the proposed methodology. Rong Ye, Qiang Xu 0001 |
ICCAD | 3 |
| 2011 | Pseudo-functional testing for small delay defects considering power supply noise effectsabstractDetecting small delay defects (SDDs) has become increasingly important to address the quality and reliability concerns of integrated circuits. Without considering functional constraints in the circuits under test, however, existing techniques may generate test patterns that are functionally-unreachable. Such SDD patterns may incur excessive (or limited) power supply noise (PSN) on sensitized paths in test mode, thus leading to over-testing or under-testing of the circuits. In this paper, we propose novel pseudo-functional testing techniques to tackle the above problem. Firstly, by taking the circuit layout information into account, functional constraints related to critical paths are extracted. Then, we generate functionally-reachable test cubes for SDD faults in the circuit. Finally, we use ATPG-like algorithm to justify transitions that pose the maximized PSN effects on sensitized critical paths under the consideration of functional constraints. The effectiveness of the proposed methodology is verified with large ISCAS'89 and ILWS'05 benchmark circuits. Xiao Liu 0011, Qiang Xu 0001 |
ICCAD | 3 |
| 2011 | Capture-power-aware test data compression using selective encoding
Jia Li 0022, Xiao Liu 0011, Yubin Zhang, Yu Hu 0001, Xiaowei Li 0001, Qiang Xu 0001 |
Integr. | 6 |
| 2011 | On Task Allocation and Scheduling for Lifetime Extension of Platform-Based MPSoC DesignsabstractWith the relentless scaling of semiconductor technology, the lifetime reliability of today's multiprocessor system-on-a-chip (MPSoC) designs has become one of the major concerns for the industry. Without explicitly taking this issue into consideration during the task allocation and scheduling process, existing works may lead to imbalanced aging rates among processors, thus reducing the system's service life. To tackle this problem, in this paper, we propose an analytical model to estimate the lifetime reliability of multiprocessor platforms when executing periodical tasks, and we present a novel task allocation and scheduling algorithm that is able to take the aging effects of processors into account, based on the simulated annealing technique. In addition, to speed up the annealing process, several techniques are proposed to simplify the design space exploration process with satisfactory solution quality. Experimental results on various hypothetical multiprocessors and task graphs show that significant system lifetime extension can be achieved by using the proposed approach, especially for heterogeneous platforms with large task graphs. Lin Huang 0002, Qiang Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2010 | On signal tracing in post-silicon validationabstractIt is increasingly difficult to guarantee the first silicon success for complex integrated circuit (IC) designs. Post-silicon validation has thus become an essential step in the IC design flow. Tracing internal signals during circuit's normal operation, being able to provide real-time visibility to the circuit under debug (CUD), is one of the most effective silicon debug techniques and has gained wide acceptance in industrial designs. Trace-based debug solution, however, involves non-trivial design for debug overhead. How to conduct signal tracing effectively for bug elimination is therefore a challenging task for IC designers. In this paper, we provide in-depth discussion for trace-based debug strategy and review recent advancements in this important area. Qiang Xu 0001, Xiao Liu 0011 |
ASP-DAC | 1 |
| 2010 | On Signal Tracing for Debugging Speedpath-Related Electrical Errors in Post-Silicon ValidationabstractOne of the most challenging problems in post-silicon validation is to identify those errors that cause prohibitive extra delay on speed paths in the circuit under debug (CUD) and only expose themselves in a certain electrical environment. To address this problem, we propose a trace-based silicon debug solution, which provides real-time visibility to the speed paths in the CUD during normal operation. Since tracing all speed path-related signals can cause prohibited design for debug (DfD) overhead, we present an automated trace signal selection methodology that maximizes error detection probability under a given constraint. In addition, we develop a novel trace qualification technique that reduces the storage requirement in trace buffers. The effectiveness of the proposed methodology is verified with large benchmark circuits. Xiao Liu 0011, Qiang Xu 0001 |
Asian Test Symposium | 2 |
| 2010 | Performance yield-driven task allocation and scheduling for MPSoCs under process variationabstractWith the ever-increasing transistor variability in CMOS technology, it is essential to integrate variation-aware performance analysis into the task allocation and scheduling process to improve its performance yield when building today's multiprocessor system-on-a-chip (MPSoC). Existing solutions assume that the execution times of tasks performed on different processors are statistically independent, which ignores the spatial correlation characteristics for systematic variation. In addition, a unified task schedule is constructed at design stage and applied to all products with various variation effects, which restricts the maximum performance yield that can be achieved for MPSoC products. To tackle the above problems, in this paper, we present a novel quasi-static scheduling algorithm. Based on a more accurate performance yield estimation method, a set of variation-aware schedules is synthesized off-line and, at run time, the scheduler will select the right one based on the actual variation for each chip, such that the timing constraint can be satisfied whenever possible. Experimental results demonstrate the effectiveness. Lin Huang 0002, Qiang Xu 0001 |
DAC | 2 |
| 2010 | AgeSim: A simulation framework for evaluating the lifetime reliability of processor-based SoCsabstractAggressive technology scaling has an ever-increasing adverse impact on the lifetime reliability of microprocessors. This paper proposes a novel simulation framework for evaluating the lifetime reliability of processor-based system-on-a-chips (SoCs), namely AgeSim, which facilitates designers to make design decisions that affect SoCs' mean time to failure. Unlike existing work, AgeSim can simulate failure mechanisms with arbitrary lifetime distributions and do not require to trace the system's reliability-related factors over its entire lifetime, and hence is more efficient and accurate. Two case studies are conducted to show the flexibility and effectiveness of the proposed methodology. Lin Huang 0002, Qiang Xu 0001 |
DATE | 2 |
| 2010 | Energy-efficient task allocation and scheduling for multi-mode MPSoCs under lifetime reliability constraintabstractIn this paper, we consider energy minimization for multiprocessor system-on-a-chip (MPSoC) under lifetime reliability constraint of the system, which has become a serious concern for the industry with technology scaling. As today's complex embedded systems typically have multiple execution modes, we first identify a set of ¿good¿ task allocation and schedules for each execution mode in terms of lifetime reliability and/or energy consumption, and then we introduce novel techniques to obtain an optimal combination of these single-mode solutions, which is able to minimize the energy consumption of the entire multi-mode system while satisfying given lifetime reliability constraint. Experimental results on several hypothetical MPSoC platforms with various task graphs demonstrate the effectiveness of the proposed approach. Lin Huang 0002, Qiang Xu 0001 |
DATE | 2 |
| 2010 | Layout-aware pseudo-functional testing for critical paths considering power supply noise effectsabstractWhen testing delay faults on critical paths, conventional structural test patterns may be applied in functionally-unreachable states, leading to over-testing or under-testing of the circuits. In this paper, we propose novel layout-aware pseudo-functional testing techniques to tackle the above problem. Firstly, by taking the circuit layout information into account, functional constraints related to delay faults on critical paths are extracted. Then, we generate functionally-reachable test cubes for every true critical path in the circuit. Finally, we fill the don't-care bits in the test cubes to maximize power supply noises on critical paths under the consideration of functional constraints. The effectiveness of the proposed methodology is verified with large ISCAS'89 benchmark circuits. Xiao Liu 0011, Yubin Zhang, Qiang Xu 0001 |
DATE | 4 |
| 2010 | An FPGA chip identification generator using configurable ring oscillatorabstractAn improved chip identification (ID) generator, otherwise known as a physically unclonable function (PUF) is described. Similar to previous designs, a cell, i, is used to obtain a measure of the difference in period of four ring oscillators and obtain the residue Ri, a random variable. Experiments show it is normally distributed with a mean of 0. A binary output value of 0 or 1 assigned depending on the sign of Ri. When |E(Ri)| is large, this scheme consistently gives the same output. Unfortunately, when it is small, the repeatability is compromised, particularly when variations in operating conditions such as supply voltage and temperature are also taken into account, which is a common problem for all previous works. To address this problem, we propose a cell with configurable ring oscillators together with an orthogonal re-initialisation scheme. Together, these two techniques maximise repeatability by causing the distribution of the mean of different Ri's to change from normal to bimodal. We implement this design in the Xilinx Spartan-3e FPGA. Nine FPGA chips are tested, and experimental results show that the new method significantly enhances reliability of ID generation and tolerance to environmental changes. Bit flip rate is reduced from 1.5% to approximately 0 at a fixed supply voltage and room temperature. Over the 20 - 80°C temperature range, and a 25% variation in supply voltage, the bit flip rate is reduced from 1.56% to 3.125 × 10-7, which is a 50000x improvement. Haile Yu, Philip H. W. Leong, Qiang Xu 0001 |
FPT | 3 |
| 2010 | Fine-grained characterization of process variation in FPGAsabstractAs semiconductor manufacturing continues towards reduced feature sizes, yield loss due to process variation becomes increasingly important. To address this issue on FPGA platforms, several variation aware design (VAD) methodologies have been proposed. In this work we present a practical method of process variation characterization (PVC) to facilitate VAD using only intrinsic FPGA resources. The scheme is based on measuring the difference between ring oscillator (RO) delay at different locations within a die, and can be used to perform process variation characterization for LE delays and interconnect delays including direct connection, double wire and hex wires. The difference in loop delays can also be estimated from equations using parameters extracted from primitives and compared with direct measurements. On a Xilinx Spartan-3e device, it was found that the error between the estimated and measured values was on average less than 10%. Haile Yu, Qiang Xu 0001, Philip H. W. Leong |
FPT | 2 |
| 2010 | Characterizing the lifetime reliability of manycore processors with core-level redundancyabstractWith aggressive technology scaling, integrated circuits suffer from ever-increasing wearout effects and their lifetime reliability has become a serious concern for the industry. For manycore processors that integrate a large number of processor cores on a single silicon die, introducing core-level redundancy is an effective way to alleviate this problem. There are, however, many strategies to make use of the redundant cores, which have different implications on the aging effects of embedded processors. How to characterize the lifetime reliability of manycore processors with different usages is therefore an important and relevant problem. In this paper, we propose a novel analytical method to tackle the above problem, which captures the impact of workloads and the associated temperature variations. We then use the proposed model to analyze the lifetime reliability for manycore processors with various redundancy configurations. Finally, the effectiveness of the proposed method is demonstrated with extensive experiments. Lin Huang 0002, Qiang Xu 0001 |
ICCAD | 2 |
| 2010 | Yield enhancement for 3D-stacked memory by redundancy sharing across diesabstractThree-dimensional (3D) memory products are emerging to fulfill the ever-increasing demands of storage capacity. In 3D-stacked memory, redundancy sharing between neighboring vertical memory blocks using short through-silicon vias (TSVs) is a promising solution for yield enhancement. Since different memory dies are with distinct fault bitmaps, how to selectively matching them together to maximize the yield for the bonded 3D-stacked memory is an interesting and relevant problem. In this paper, we present novel solutions to tackle the above problem. Experimental results show that the proposed methodology can significantly increase memory yield when compared to the case that we only bond self-reparable dies together. Li Jiang 0002, Rong Ye, Qiang Xu 0001 |
ICCAD | 3 |
| 2010 | On timing-independent false path identificationabstractThis paper is concerned with finding timing-independent false paths that cannot be sensitized under any signal arrival time condition in integrated circuits. Existing techniques regard a path as a true path as long as a vector pair can be found to sensitize it. This is rather pessimistic since such a path might be activated only with illegal states in the circuit and hence it is actually functionally-unsensitizable. In this paper, we develop novel techniques to take the above issue into consideration when identifying false paths, which facilitates us to find much more false paths than conventional techniques. Experimental results on benchmark circuits demonstrate the effectiveness of the proposed methodology. Qiang Xu 0001 |
ICCAD | 2 |
| 2010 | Modeling TSV open defects in 3D-stacked DRAMabstractThree-dimensional (3D) stacking using through silicon vias (TSVs) is a promising solution to provide low-latency and high-bandwidth DRAM access from microprocessors. The large number of TSVs implemented in 3D DRAM circuits, however, are prone to open defects and coupling noises, leading to new test challenges. Through extensive simulation studies, this paper models the faulty behavior of TSV open defects occurred on the wordlines and the bitlines of 3D DRAM circuits, which serves as the first step for efficient and effective test and diagnosis solutions for such defects. Li Jiang 0002, Yuan Xie 0001, Qiang Xu 0001 |
ITC | 5 |
| 2010 | Economic Analysis of Testing Homogeneous Manycore ChipsabstractThe employment of a large number of structurally identical cores on a single silicon die is generally regarded as a promising solution for tera-scale computation, known as manycore chips. To ensure the product quality of such complex integrated circuits before shipping them to final users, extensive manufacturing tests are necessary and the associated test cost can account for a large share of the total production cost. By introducing spare cores on-chip, the burn-in test time can be shortened and the defect coverage requirements for core tests can be also relaxed, without sacrificing quality of the shipped products. If the above test cost reduction exceeds the manufacturing cost of the extra cores, the total production cost of manycore chips can be reduced. In this paper, we develop novel analytical models to study the above tradeoff and we verify the effectiveness of the proposed test economics model for hypothetical manycore chips with various configurations. Lin Huang 0002, Qiang Xu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2010 | Lifetime Reliability for Load-Sharing Redundant Systems With Arbitrary Failure DistributionsabstractIn this work, a general closed-form expression is presented for the lifetime reliability of load-sharingk-out-of-n:G hybrid redundant systems. In such systems,mcomponents are initially configured as active units. Depending on whether it is performing tasks, an active component can be in either a processing, or a wait state. Each state corresponds to an arbitrary failure distribution. The remaining (n-m) are spares to provide fault tolerance. Each time an active component fails, a spare one converts into active mode, until there are no more spares in the system. Then, the system works in a gracefully degrading manner such that less thanmcomponents share the workload, until the number of good components is less thank. The task allocation, and service are modeled as queueing systems, wherein the utilization ratio essentially affects the aging effect of components. We integrate the various failure distributions for components in different operational states into an analytical model according to the statistical properties of the task allocation mechanisms, and the components' processing capacity, and analyse the lifetime reliability of the entire system. Finally, three special cases, and a series of numerical experiments are discussed in detail to show the practical applicability of the proposed approach. Lin Huang 0002, Qiang Xu 0001 |
IEEE Trans. Reliab. | 2 |
| 2010 | X-Filling for Simultaneous Shift- and Capture-Power Reduction in At-Speed Scan-Based TestingabstractPower consumption during at-speed scan-based testing can be significantly higher than that during normal functional mode in both shift and capture phases, which can cause circuits' reliability concerns during manufacturing test. This paper proposes a novel X-filling technique, namely “iFill”, to address the above issue, by analyzing the impact ofX-bitson switching activities of the circuit nodes in the two different phases. In addition, different from prior$X$-filling methods for shift-power reduction that can only reduce shift-in power, our method is able to cut down power consumptions in both shift-in and shift-out processes. Experimental results on benchmark circuits show that the proposed technique can guarantee the power safety in both shift and capture phases during at-speed scan-based testing. Jia Li 0022, Qiang Xu 0001, Yu Hu 0001, Xiaowei Li 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | Test Pattern Selection for Potentially Harmful Open Defects in Power Distribution NetworksabstractPower distribution network (PDN) designs for today's high performance integrated circuits (ICs) typically occupy a significant share of metal resources in the circuit, and hence defects may be introduced on PDNs during the manufacturing process. Since we cannot afford to over-design the PDNs to tolerate all possible defects, it is necessary to conduct manufacturing test for them. In this paper, we propose novel methodologies to identify those potentially harmful open defects in PDNs and we show how to select a set of patterns that initially target transition faults to achieve high fault coverage for the PDN defects. Experimental results on benchmark circuits demonstrate the effectiveness of the proposed technique. Yubin Zhang, Lin Huang 0002, Qiang Xu 0001 |
Asian Test Symposium | 4 |
| 2009 | Interconnection fabric design for tracing signals in post-silicon validationabstractPost-silicon validation has become an essential step in the design flow of today’s complex integrated circuits. One effective technique that provides real-time visibility to the circuit under debug (CUD) is to monitor and trace internal signals during its normal operation. Typically, a large number of signals are tapped and a subset of them are selected to be observed in each debug process. These trace signals need to be transferred to on-chip buffers and/or off-chip trace ports for analysis. Existing solutions use multiplexer trees or specific access networks to conduct the above duty, which, however, either provide less visibility to the CUD or result in high hardware cost. In this paper, we propose a novel trace signal interconnection fabric design to tackle the above problem. Experimental results on benchmark circuits show the efficacy of the proposed solution. Xiao Liu 0011, Qiang Xu 0001 |
DAC | 2 |
| 2009 | On systematic illegal state identification for pseudo-functional testingabstractThe discrepancy between integrated circuits' activities in normal functional mode and that in structural test mode has an increasing adverse impact on the effectiveness of manufacturing test. Pseudo-functional testing tries to resolve this problem by identifying illegal states in functional mode and avoiding them during the test pattern generation process. Existing methods, however, can only extract a small set of illegal states in the system due to various limitations. In this paper, we first show that illegal states in the system are mainly caused by multi-fanout nets in the circuit, and we develop efficient and effective heuristics to identify them. Experimental results on benchmark circuits demonstrate the effectiveness of our proposed systematic solution. Qiang Xu 0001 |
DAC | 2 |
| 2009 | Lifetime reliability-aware task allocation and scheduling for MPSoC platformsabstractWith the relentless scaling of semiconductor technology, the lifetime reliability of embedded multiprocessor platforms has become one of the major concerns for the industry. If this is not taken into consideration during the task allocation and scheduling process, some processors might age much faster than the others and become the reliability bottleneck for the system, thus significantly reducing the system's service life. To tackle this problem, in this paper, we propose an analytical model to estimate the lifetime reliability of multiprocessor platforms when executing periodical tasks, and we present a novel lifetime reliability-aware task allocation and scheduling algorithm based on simulated annealing technique. In addition, to speed up the annealing process, several techniques are proposed to simplify the design space exploration process with satisfactory solution quality. Experimental results on various multiprocessor platforms and task graphs demonstrate the efficacy of the proposed approach. Lin Huang 0002, Qiang Xu 0001 |
DATE | 3 |
| 2009 | Test architecture design and optimization for three-dimensional SoCsabstractCore-based system-on-chips (SoCs) fabricated on three-dimensional (3D) technology are emerging for better integration capabilities. Effective test architecture design and optimization techniques are essential to minimize the manufacturing cost for such giga-scale integrated circuits. In this paper, we propose novel test solutions for 3D SoCs manufactured with die-to-wafer and die-to-die bonding techniques. Both testing time and routing cost associated with the test access mechanisms in 3D SoCs are considered in our simulated annealing-based technique. Experimental results on ITC'02 SoC benchmark circuits are compared to those obtained with two baseline solutions, which show the effectiveness of the proposed technique. Li Jiang 0002, Lin Huang 0002, Qiang Xu 0001 |
DATE | 3 |
| 2009 | Trace signal selection for visibility enhancement in post-silicon validationabstractToday's complex integrated circuit designs increasingly rely on post-silicon validation to eliminate bugs that escape from pre-silicon verification. One effective silicon debug technique is to monitor and trace the behaviors of the circuit during its normal operation. However, designers can only afford to trace a small number of signals in the design due to the associated overhead. Selecting which signals to trace is therefore a crucial issue for the effectiveness of this technique. This paper proposes an automated trace signal selection strategy that is able to dramatically enhance the visibility in post-silicon validation. Experimental results on benchmark circuits show that the proposed technique is more effective than existing solutions. Xiao Liu 0011, Qiang Xu 0001 |
DATE | 2 |
| 2009 | A generic framework for scan capture power reduction in fixed-length symbol-based test compression environmentabstractGrowing test data volume and overtesting caused by excessive scan capture power are two of the major concerns for the industry when testing large integrated circuits. Various test data compression (TDC) schemes and low-power X-filling techniques were proposed to address the above problems. These methods, however, exploit the very same ldquodon't-carerdquo bits in the test cubes to achieve different objectives and hence may contradict to each other. In this work, we propose a generic framework for reducing scan capture power in test compression environment. Using the entropy of the test set to measure the impact of capture power-aware X-filling on the potential test compression ratio, the proposed holistic solution is able to keep capture power under a safe limit with little compression ratio loss for any fixed-length symbol-based TDC method. Experimental results on benchmark circuits demonstrate the efficacy of the proposed approach. Xiao Liu 0011, Qiang Xu 0001 |
DATE | 2 |
| 2009 | Layout-driven test-architecture design and optimization for 3D SoCs under pre-bond test-pin-count constraintabstractWe propose a layout-driven test-architecture design and optimization technique for core-based system-on-chips (SoCs) that are fabricated using three-dimensional (3D) integration. In contrast to prior work, we consider the pre-bond test-pin-count constraint during optimization since these pins occupy large silicon area that cannot be used in functional mode. In addition, the proposed test-architecture design takes the SoC layout into consideration and facilitates the sharing of test wires between pre-bond tests and post-bond test, which significantly reduces the routing cost for a test-access mechanism in 3D technology. Experimental results for the ITC'02 SoC benchmarks circuits demonstrate the effectiveness of the proposed solution. Li Jiang 0002, Qiang Xu 0001, Krishnendu Chakrabarty |
ICCAD | 2 |
| 2009 | Test economics for homogeneous manycore systemsabstractHomogeneous manycore systems that contain a large number of structurally identical cores are emerging for tera-scale computation. To ensure the required quality and reliability of such complex integrated circuits before supplying them to final users, extensive manufacturing tests need to be conducted and the associated test cost can account for a great share of the total production cost. By introducing spare cores on-chip, the burn-in test time can be shortened and the defect coverage requirements for core tests can be also relaxed, without sacrificing quality of the shipped products. The above test cost reduction is likely to exceed/compensate the manufacturing cost of the extra cores, thus reducing the total production cost of manycore systems. We develop novel analytical models that capture the above tradeoff in this paper and we verify the effectiveness of the proposed test economics model for hypothetical manycore systems with various configurations. Lin Huang 0002, Qiang Xu 0001 |
ITC | 2 |
| 2009 | On simultaneous shift- and capture-power reduction in linear decompressor-based test compression environmentabstractGrowing test data volume and excessive test power consumption in scan-based testing are both serious concerns for the semiconductor industry. Various test data compression (TDC) schemes and low-power X-filling techniques were proposed to address the above problems. These methods, however, exploit the very same ¿don't-care¿ bits in the test cubes to achieve different objectives and hence may contradict to each other. In this work, we propose a generic framework for test power reduction in linear decompressor-based test compression environment, which is able to effectively reduce shift-and capture-power simultaneously. Experimental results on benchmark circuits demonstrate that our proposed techniques significantly outperform existing solutions. Xiao Liu 0011, Qiang Xu 0001 |
ITC | 2 |
| 2009 | Trace signal selection for debugging electrical errors in post-silicon validationabstractDebugging electrical errors is the most challenging problem during the post-silicon validation process. We propose an automated trace signal selection methodology to facilitate this task, in which, by analyzing the layout of the circuit and carefully selecting trace signals, designers are with high probability to identify electrical errors. Xiao Liu 0011, Qiang Xu 0001 |
ITC | 2 |
| 2009 | Compression-aware pseudo-functional testingabstractWith technology scaling, the discrepancy between integrated circuits' activities in normal functional mode and that in structural test mode has an increasing adverse impact on the effectiveness of manufacturing test. By identifying functionally unreachable states in the circuit and avoiding them during the test generation process, pseudo-functional testing is an effective technique to address this problem. Pseudo-functional patterns, however, feature much less don't-care bits when compared to conventional structural patterns, making them less friendly to test compression techniques. In this paper, we propose novel solutions to address the above problem, which facilitate to apply pseudo-functional testing in linear decompressor-based test compression environment. Experimental results on ISCAS'89 benchmark circuits demonstrate the effectiveness of the proposed methodology. Qiang Xu 0001 |
ITC | 2 |
| 2009 | Low-Power Scan Testing for Test Data Compression Using a Routing-Driven Scan ArchitectureabstractA new scan architecture is proposed to reduce peak test power and capture power. Only a subset of scan flip-flops is activated to shift test data or capture test responses in any clock cycle. This can effectively reduce the capture test power and peak test power. Two routing-driven schemes are proposed to reduce the routing overhead. Experimental results show that the proposed scan architecture can effectively reduce peak test power, capture power, test data volume, and test application cost. Dianwei Hu, Qiang Xu 0001, Alex Orailoglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2009 | SOC test-architecture optimization for the testing of embedded cores and signal-integrity faults on core-external interconnectsabstractThe test time for core-external interconnect shorts and opens is typically much less than that for core-internal logic. Therefore, prior work on test-infrastructure design for core-based system-on-a-chip (SOC) has mainly focused on minimizing the test time for core-internal logic. However, as feature sizes shrink for newer process technologies, the test time for signal integrity (SI) faults on interconnects cannot be neglected. The test time for SI faults can be comparable to, or even larger than, the test time for the embedded cores. We investigate the impact of interconnect SI tests on SOC test-architecture design and optimization. A compaction method for SI faults and algorithms for test-architecture optimization are also presented. Experimental results for the ITC'02 benchmarks show that the proposed approach can significantly reduce the overall testing time for core-internal logic and core-external interconnects. Qiang Xu 0001, Yubin Zhang, Krishnendu Chakrabarty |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2009 | On Topology Reconfiguration for Defect-Tolerant NoC-Based Homogeneous Manycore SystemsabstractHomogeneous manycore systems are emerging for tera-scale computation and typically utilize Network-on-Chip (NoC) as the communication scheme between embedded cores. Effective defect tolerance techniques are essential to improve the yield of such complex integrated circuits. We propose to achieve fault tolerance by employing redundancy at the core-level instead of at the microarchitecture level. When faulty cores exist on-chip in this architecture, however, the physical topologies of various manufactured chips can be significantly different. How to reconfigure the system with the most effective NoC topology is a relevant research problem. In this paper, we first show that this problem is an instance of a well known NP-complete problem. We then present novel solutions for the above problem, which not only maximize the performance of the on-chip communication scheme, but also provide a unified topology to Operating System and application software running on the processor. Experimental results show the effectiveness of the proposed techniques. Lei Zhang 0008, Yinhe Han 0001, Qiang Xu 0001, Xiaowei Li 0001, Huawei Li 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2008 | On reducing both shift and capture power for scan-based testingabstractPower consumption in scan-based testing is a major concern nowadays. In this paper, we present a new X-filling technique to reduce both shift power and capture power during scan tests, namely LSC-filling. The basic idea is to use as few as possible X-bits to keep the capture power under the peak power limit of the circuit under test (CUT), while using the remaining X-bits to reduce the shift power to cut down the CUT’s average power consumption during scan tests as much as possible. In addition, by carefully selecting the X-filling order, our X-filling technique is able to achieve lower capture power when compared to existing methods. Experimental results on ISCAS’89 benchmark circuits show the effectiveness of the proposed methodology. Jia Li 0022, Qiang Xu 0001, Yu Hu 0001, Xiaowei Li 0001 |
ASP-DAC | 2 |
| 2008 | A debug probe for concurrently debugging multiple embedded cores and inter-core transactions in NoC-based systemsabstractExisting SoC debug techniques mainly target bus-based systems. They are not readily applicable to the emerging system that use Network-on-Chip (NoC) as on-chip communication scheme. In this paper, we present the detailed design of a novel debug probe (DP) inserted between the core under debug (CUD) and the NoC. With embedded configurable triggers, delay control and timestamping mechanism, the proposed DP is very effective for inter-core transaction analysis as well as controlling embedded cores’ debug processes. Experimental results show the functionalities of the proposed DP and its area overhead. Shan Tang, Qiang Xu 0001 |
ASP-DAC | 2 |
| 2008 | On Reusing Test Access Mechanisms for Debug Data Transfer in SoC Post-Silicon ValidationabstractOne of the main difficulties in post-silicon validation is the limited debug access bandwidth to internal signals. At the same time, SoC devices often contain dedicated bus-based test access mechanisms (TAMs) that are used to transfer test data between external testers and embedded cores. In this paper, we propose to reuse these precious TAM resources for real-time debug data transfer in post-silicon validation. This strategy significantly increases debug bandwidth with negligible routing overhead. To support different TAM architectures and debug scenarios, design for debug (DfD) structures are introduced at both core test wrapper level and system level. Simulation results demonstrate the effectiveness of the proposed approach at low DfD cost. Xiao Liu 0011, Qiang Xu 0001 |
ATS | 2 |
| 2008 | On reliable modular testing with vulnerable test access mechanismsabstractIn modular testing of system-on-a-chip (SoC), test access mechanisms (TAMs) are used to transport test data between the input/output pins of the SoC and the cores under test. Prior work assumes TAMs to be error-free during test data transfer. The validity of this assumption, however, is questionable with the ever-decreasing feature size of today's VLSI technology and the ever-increasing circuit operational frequency. In particular, when functional interconnects such as network-on-chip (NoC) are reused as TAMs, even if they have passed manufacturing test beforehand, failures caused by electrical noise such as crosstalk and transient errors may happen during test data transfer and make good chips appear to be defective, thus leading to undesired test yield loss. To address the above problem, in this paper, we propose novel solutions that are able to achieve reliable modular testing even if test data may sometimes get corrupted during transmission with vulnerable TAMs, by designing a new "jitter-aware" test wrapper and a new "jitter-transparent" ATE interface. Experimental results on an industrial circuit demonstrate the effectiveness of the proposed technique. Lin Huang 0002, Qiang Xu 0001 |
DAC | 3 |
| 2008 | iFill: An Impact-Oriented X-Filling Method for Shift- and Capture-Power Reduction in At-Speed Scan-Based TestingabstractIn scan-based tests, power consumptions in both shift and capture phases may be significantly higher than that in normal mode, which threatens circuits' reliability during manufacturing test. In this paper, by analyzing the impact of X-bits on circuit switching activities, we present an X-filling technique that can decrease both shift- and capture-power to guarantees the reliability of scan tests, called iFill. Moreover, different from prior work on X-filling for shift-power reduction which can only reduce shift-in power, iFill is able to decrease power consumptions during both shift-in and shift-out. Experimental results on ISCAS' 89 benchmark circuits show the effectiveness of the proposed technique. Jia Li 0022, Qiang Xu 0001, Yu Hu 0001, Xiaowei Li 0001 |
DATE | 2 |
| 2008 | In-band Cross-Trigger Event Transmission for Transaction-Based DebugabstractCross-trigger, the mechanism to trigger activities in one debug entity from debug events happened in another debug entity, is a very useful technique for debugging applications involving multiple embedded cores. Existing solutions rely on dedicated interconnects (i.e., different from functional interconnects) to transfer debug events and cannot guarantee the arrival time of the debug events coincides with the arrival time of the data messages between multiple cores. This results in mismatches between the observed system internal operations and the ones that designers expect to watch. To tackle the above problem, in this paper, we propose to package the cross-trigger events and the actual data together into transaction messages and transfer them along the same functional interconnects (namely in- band debug event transmission), with the help of novel design-for- debug circuits. Simulation results on a hypothetical NoC-based systems show the effectiveness of the proposed technique. Shan Tang, Qiang Xu 0001 |
DATE | 2 |
| 2008 | Re-Examining the Use of Network-on-Chip as Test Access MechanismabstractExisting work on testing NoC-based systems advocates to reuse the on-chip network itself as test access mechanism (TAM) to transport test data to/from embedded cores. While this methodology obviously reduces the routing cost when compared to the case that dedicated test buses are introduced as TAMs, it is not clear whether it is beneficial in terms of other important factors that significantly affect test cost, e.g., testing time, test control complexity and test reliability. As a result, in this paper, we re-examine the issue of using NoC as TAM in order to facilitate designers to construct a cost-effective system test architecture based on their requirements. Lin Huang 0002, Qiang Xu 0001 |
DATE | 3 |
| 2008 | Defect Tolerance in Homogeneous Manycore Processors Using Core-Level Redundancy with Unified TopologyabstractHomogeneous manycore processors are emerging for tera-scale computation. Effective defect tolerance techniques are essential to improve the yield of such complex integrated circuits. In this paper, we propose to achieve fault tolerance by employing redundancy at the core-level instead of at the microarchitecture-level. When faulty cores existing on-chip in this architecture, how to reconfigure the processor with the most effective topology is a relevant research problem. We present novel solutions for this problem, which not only maximize the performance of the manycore processor, but also provide a unified topology to operating system and application software running on the processor. Experimental results show the effectiveness of the proposed techniques. Lei Zhang 0008, Yinhe Han 0001, Qiang Xu 0001, Xiaowei Li 0001 |
DATE | 3 |
| 2008 | On capture power-aware test data compression for scan-based testingabstractLarge test data volume and high test power are two of the major concerns for the industry when testing large integrated circuits. With given test cubes in scan-based testing, the ldquodonpsilat-carerdquo bits can be exploited for test data compression and/or test power reduction. Prior work either targets only one of these two issues or considers to reduce test data volume and scan shift power together. In this paper, we propose a novel capture power-aware test compression scheme that is able to keep scan capture power under a safe limit with little loss in test compression ratio. Experimental results on benchmark circuits demonstrate the efficacy of the proposed approach. Jia Li 0022, Xiao Liu 0011, Yubin Zhang, Yu Hu 0001, Xiaowei Li 0001, Qiang Xu 0001 |
ICCAD | 6 |
| 2008 | Is It Cost-Effective to Achieve Very High Fault Coverage for Testing Homogeneous SoCs with Core-Level Redundancy?abstractThis work argues to reduce the production cost of homogeneous SoCs by introducing dedicated test cost-driven redundant cores. By doing so, the fault coverage for each core and hence the SoC test cost can be dramatically reduced, which is able to compensate the manufacturing cost of the extra cores. A case study is presented to demonstrate the effectiveness of the proposed scheme. Lin Huang 0002, Qiang Xu 0001 |
ITC | 2 |
| 2008 | A Generic Framework for Scan Capture Power Reduction in Test Compression EnvironmentabstractThis work proposes a generic framework for reducing scan capture power in test compression environment. Using the entropy of the test set to measure the impact of capture power-aware X-filling on the potential test compression ratio, the proposed method is able to keep capture power under a safe limit with little loss in test compression ratio. Xiao Liu 0011, Qiang Xu 0001 |
ITC | 3 |
| 2008 | SoC Test Architecture Design and Optimization Considering Power Supply Noise EffectsabstractExcessive power supply noise (PSN) during testing can erroneously cause good chips to fail the manufacturing test, thus leading to unnecessary yield loss. While there are some emerging methodologies such as PSN-aware test generation technique to tackle this problem, they are not readily applicable in modular system-on-a-chip (SoC) testing. This is because: (i). embedded core tests are usually prepared by core providers who are not knowledgeable about the SoC power distribution network; (ii) embedded cores are usually tested in parallel to reduce testing time, but the associated inter-core PSN effects are not considered in existing SoC test architecture design and optimization process. In this paper, we present a fast inter-core PSN estimation method and use it to guide the test scheduling process to solve the PSN-induced SoC test yield loss problem. In addition, upon observing that the PSN effects usually manifest themselves only during the capture phase for scan-tested cores, we introduce novel design for test (DfT) structures into SoC test controller to avoid PSN effects with negligible testing time penalty. Experimental results demonstrate the effectiveness of the proposed solution. Qiang Xu 0001 |
ITC | 2 |
| 2008 | On Modeling the Lifetime Reliability of Homogeneous Manycore SystemsabstractAdvancements in technology enable integration of a large number of cores on a single silicon die. At the same time, aggressive technology scaling has an ever-increasing adverse impact on the lifetime reliability of such large integrated circuits. In this work, we model the lifetime reliability of homogeneous manycore systems using a load-sharing nonrepairable k-out-of-n:G system with general failure distributions for embedded cores. In manycore systems, an embedded core can be in operational, cold standby, or warm standby state depending on system redundancy schemes and their workloads. We then use the proposed model to analyze the impact of different redundant schemes and configurations on the lifetime reliability of manycore systems. Lin Huang 0002, Qiang Xu 0001 |
PRDC | 2 |
| 2008 | State-Sensitive X-Filling Scheme for Scan Capture Power ReductionabstractBased on the operation of a state machine, this paper elucidates a comprehensive frame for probability-based primary-input-dominated X-filling methods to minimize the total weighted switching activity (WSA) during the scan capture operation. Experimental results demonstrate that the proposed approach significantly reduces both average and peak WSAs. Jing-Ling Yang, Qiang Xu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2007 | SOC Test Architecture Optimization for Signal Integrity Faults on Core-External InterconnectsabstractThe test time for core-external interconnect shorts/opens is typically much less than that for core-internal logic. Therefore, prior work on test infrastructure design for core-based system-on-a-chip (SOC) has mainly focused on minimizing the test time for core-internal logic. However, as feature sizes shrink for newer process technologies, the test time for interconnect signal integrity (SI) faults cannot be neglected. We investigate the impact of interconnect SI tests on SOC test architecture design and optimization. We present a compaction method for SI faults and algorithms for test architecture optimization. Experimental results for the ITC'02 benchmarks show that the proposed approach can significantly reduce the overall testing time for core-internal logic and core-external interconnects. Qiang Xu 0001, Yubin Zhang, Krishnendu Chakrabarty |
DAC | 1 |
| 2007 | A multi-core debug platform for NoC-based systemsabstractNetwork-on-chip (NoC) is generally regarded as the most promising solution for the future on-chip communication scheme in giga-scale integrated circuits. As traditional debug architecture for bus-based systems is not readily applicable to identify bugs in NoC-based systems, in this paper, we present a novel debug platform that supports concurrent debug access to the cores under debug (CUDs) and the NoC in a unified architecture. By introducing core-level debug probes in between the CUDs and their network interfaces and a system-level debug agent controlled by an off-chip multi-core debug controller, the proposed debug platform provides in-depth analysis features for NoC-based systems, such as NoC transaction analysis, multi-core cross-triggering and global synchronized timestamping. Therefore, the proposed solution is expected to facilitate the designers to identify bugs in NoC-based systems more effectively and efficiently. Experimental results show that the design-for-debug cost for the proposed technique in terms of area and traffic requirements is moderate Shan Tang, Qiang Xu 0001 |
DATE | 2 |
| 2007 | Pattern-directed circuit virtual partitioning for test power reductionabstractFor a large circuit under test (CUT), it is likely that some test patterns result in excessive power dissipations that exceed the CUT’s power rating. Designers may resort to lowpower automatic test pattern generation (ATPG) tools to solve this problem, which, however, usually leads to larger test data volume and requires extra computational effort, even if such tools are available. Another method is to partition the circuit into multiple subcircuits and test them separately. Unfortunately, this usually involves rerunning the time-consuming ATPG for each partitioned subcircuit and solving the problem of how to achieve an acceptable fault coverage for the glue logic between subcircuits. In this paper, we propose a novel low-power virtual test partitioning technique without the above-mentioned shortcomings. The basic idea is to partition the circuit in such way that the faults in the glue logic between subcircuits can be detected by patterns with low power dissipation that are applied at the entire circuit level, while the patterns with high power dissipation can be applied within a partitioned subcircuit without loss of fault coverage. Scan chain routing cost has also been considered during the partitioning process. Experimental results show that the proposed technique is very effective in reducing test power. Qiang Xu 0001, Dianwei Hu |
ITC | 1 |
| 2007 | Test-wrapper designs for the detection of signal-integrity faults on core-external interconnects of SoCsabstractAs feature sizes continue to shrink for newer process technologies, signal integrity (SI) is emerging as a major concern for core-based system-on-a-chip (SoC) integrated circuits. To effectively test SI faults on core-external interconnects, core test wrappers need to be able to generate appropriate transitions at a wrapper output cell (WOC) on the driving side and detect the signal integrity loss at a wrapper input cell on the receiving side. In current wrapper designs, the WOCs for a victim interconnect and its aggressors make transitions at the same time with a common test clock signal in test mode, which is different from the functional mode. This is not adequate for SI test because the time elapsed between the transition of the victim and the transitions of its aggressors significantly affects the behavior of SI-related errors. To address this problem, we propose new IEEE Std. 1500-compliant wrapper designs that are able to apply SI test at functional mode or make transitions with various pre-defined skews between a victim line and its aggressors. We also introduce a novel overshoot detector inside the proposed wrapper. Experimental results show that the proposed wrapper designs are more effective for detecting SI-related errors when compared to existing techniques, with a moderate amount of DFT overhead. Qiang Xu 0001, Yubin Zhang, Krishnendu Chakrabarty |
ITC | 1 |
| 2007 | Test Wrapper Design and Optimization Under Power Constraints for Embedded Cores With Multiple Clock DomainsabstractEven though many embedded cores contain several clock domains, most published methods for wrapper design have been limited to single-frequency cores. Cumbersome and invasive design techniques, such as insertion of test points, are needed to make these methods applicable to current-generation embedded cores. This paper presents a new method for designing test wrappers for embedded cores with multiple clock domains. The proposed 1500-compliant wrapper prevents clock skew and allows scan chains in different clock domains to shift test data at distinct clock frequencies, which enables a better control of power dissipation during test. We present an integer linear programming (ILP) model that can be used to minimize the core testing time under power constraints for small problem instances, and which can be combined with LP-relaxation to obtain lower bounds on the testing time for larger instances. We also present an efficient heuristic method that is applicable to large problem instances, and which yields the same (optimal) testing time as ILP for small problem instances. Qiang Xu 0001, Nicola Nicolici, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2006 | Test/Repair Area Overhead Reduction for Small Embedded SRAMsabstractFor current highly-integrated and memory-dominant system-on-a-chips (SoCs), especially for graphics and networking SoCs, the test/repair area overhead of embedded SRAMs (e-SRAMs) is a big concern. This paper presents various approaches to tackle this problem from a practical point of view. Without sacrificing at-speed testability, diagnosis capability and repairability, the proposed approaches consider partly sharing wrapper for identical memories, sharing memory BIST controllers for e-SRAMs embedded in different functional blocks, test responses compression for wide memories, and various repair strategies for e-SRAMs with different configurations. By combining the above approaches, the test/repair area overhead for e-SRAMs can be significantly reduced. For example, for one benchmark SoC used in our experiments, it can be reduced as much as 10% of the entire memory array Qiang Xu 0001 |
ATS | 2 |
| 2006 | Retention-Aware Test Scheduling for BISTed Embedded SRAMsabstractIn this paper we address the test scheduling problem for Builtin Self-tested (BISTed) embedded SRAMs (e-SRAMs) when Data Retention Faults (DRFs) are considered. The proposed test scheduling algorithm utilizes the "retention-aware"test power model [1] to minimize the total testing time of e- SRAMs while not violating given power constraints. Without losing generality, we consider both cases where the pause time for data retention faults is fixed and cases where it can be varied. Experimental results show that the "retention-aware" test scheduling algorithm can reduce the testing time of e- SRAMs up to more than 98 percent at the computational time within a second. Qiang Xu 0001, Evangeline F. Y. Young |
ETS | 1 |
| 2006 | DFT Infrastructure for Broadside Two-Pattern Test of Core-Based SOCsabstractExisting approaches for modular manufacturing test of core-based system-on-a-chip (SOC) devices do not provide any explicit mechanism for delivering two-pattern tests in the broadside mode, which is necessary to achieve reliable coverage of delay and stuck-open faults. Although wrapper input cells can be enhanced with two memory elements to address this problem, this incur a large test area overhead. This paper proposes a novel architecture for broadside two-pattern test of core-based SOCs without any loss in fault coverage and without increasing the size of the wrapper input cells. The proposed solution combines the dedicated bus-based test access mechanism and functional interconnects for test data transfer in order to provide full controllability of the wrapper input cells in the two consecutive clock cycles required by two-pattern testing. New algorithms for test access mechanism design and test scheduling are proposed and design trade-offs between test area and testing time are discussed using experimental results. Qiang Xu 0001, Nicola Nicolici |
IEEE Trans. Computers | 1 |
| 2006 | Multifrequency TAM design for hierarchical SOCsabstractThe emergence of megacores in hierarchical system-on-a-chip (SOC) presents new challenges to electronic test automation. This paper describes a new framework for designing test access mechanisms (TAMs) for modular testing of hierarchical SOCs. We first explore the concept that TAMs on the same level of design hierarchy employ multiple frequencies for test data transportation. Then we extend this concept to hierarchical SOCs and, by introducing frequency converters at the inputs and outputs of the megacores, the proposed solution not only removes the constraint that the system level TAM width must be wider than the internal TAM width of the megacores, but also facilitates rapid exploration of the tradeoffs between the test application time and the required test area. Experimental results for the ITC'02 SOC Test Benchmarks show that the proposed TAM design algorithms increase the size of the solution space that is explored, which, in turn, will lower the test application time when compared to the existing solutions. Qiang Xu 0001, Nicola Nicolici |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2005 | Register-transfer level functional scan for hierarchical designsabstractThis paper discusses the potential benefits of inserting scan chains (SCs) in hierarchical designs at the register-transfer level (RTL) of design abstraction. Using new algorithms for functional scan chain design, it is shown how tight timing constraints for design-for-test (DFT) planning at RTL can improve the performance of a circuit, when compared to its gate level counterpart, without any loss in testability. Ho Fai Ko, Qiang Xu 0001, Nicola Nicolici |
ASP-DAC | 2 |
| 2005 | Multi-frequency wrapper design and optimization for embedded cores under average power constraintsabstractThis paper presents a new method for designing test wrappers for embedded cores with multiple clock domains. By exploiting the use of multiple shift frequencies, the proposed method improves upon a recent wrapper design method that requires a common shift frequency for the scan elements in the different clock domains. We present an integer linear programming (ILP) model that can be used to minimize the testing time for small problem instances. We also present an efficient heuristic method that is applicable to large problem instances, and which yields the same (optimal) testing time as ILP for small problem instances. Compared to recent work on wrapper design using a single shift frequency, we obtain lower testing times and the reduction in testing time is especially significant under power constraints. Qiang Xu 0001, Nicola Nicolici, Krishnendu Chakrabarty |
DAC | 1 |
| 2005 | On concurrent test of wrapped cores and unwrapped logic blocks in SOCsabstractSystem-on-a-chip designs may contain user defined logic or embedded cores that cannot be wrapped for test purposes due to area constraints or timing violations. This paper discusses how these unwrapped logic blocks can be tested rapidly through the TestRail architecture using only the test control mechanism and the test instructions available through IEEE standard for embedded core test. A new test scheduling algorithm, which facilitates concurrent test of both unwrapped logic blocks and wrapped cores, is proposed and experiments show that it outperforms a previous approach when the available number of tester channels and/or the number of unwrapped logic blocks are small. Qiang Xu 0001, Nicola Nicolici |
ITC | 1 |
| 2005 | Modular SOC testing with reduced wrapper countabstractMotivated by the increasing design for test (DFT) area overhead and potential performance degradation caused by wrapping all the embedded cores for modular system-on-a-chip (SOC) testing, this paper proposes a solution for reducing the number of wrapper boundary register (WBR) cells. By utilizing the functional interconnect topology and the WBRs of the surrounding cores to transfer test stimuli and responses, the WBRs of some cores can be removed without affecting the testability of the SOC. We denote the cores without WBRs as light-wrapped cores and present a new modular SOC test architecture for concurrently testing both the wrapped and the light-wrapped logic cores. Since the WBRs of cores that transfer test stimuli and test responses for light-wrapped cores become shared resources during test, conflicts arise during test scheduling that will negatively impact the test application time. As a consequence, to alleviate this problem, we present a novel test access mechanism (TAM) design algorithm for the proposed SOC test architecture. We conduct experiments on several SOC benchmark circuits and demonstrate that, with an acceptable increase in test application time, the number of WBRs can be significantly decreased. This will ultimately lessen the necessary DFT area for modular SOC testing and reduce the propagation delays between cores. Qiang Xu 0001, Nicola Nicolici |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2005 | Wrapper design for multifrequency IP coresabstractThis paper addresses the testability problems raised by intellectual property cores with multiple clock domains. The proposed solution is based on a novel core wrapper architecture and a new wrapper design algorithm. It is shown how multifrequency at-speed test response capture can be achieved via the design of capture windows without any structural modifications to the logic within the embedded core. The new features in the core wrapper architecture, which introduce limited hardware overhead, can also synchronize the external tester channels with the core's internal scan chains in the shift mode. Thus, the wrapper implementation space can be explored in order to efficiently utilize the available tester bandwidth while meeting the constraints on the maximum internal shift frequency that guarantees low testing time within the given power ratings. Using experimental data, the benefits of the proposed solution are demonstrated by analyzing the tradeoffs between the number of tester channels, testing time, area overhead, and power dissipation. Qiang Xu 0001, Nicola Nicolici |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2005 | Modular and rapid testing of SOCs with unwrapped logic blocksabstractExtensive research has been carried out for test planning of core-based system-on-a-chip devices. Most of the prior work assumes that all of the embedded cores are wrapped for test purpose. However, some designs may contain user-defined logic or cores that cannot be wrapped due to area constraints or timing violations. This paper discusses how these unwrapped logic blocks can be tested rapidly by adapting the TestRail architecture, which uses only the test control mechanism and the test instructions available through the IEEE 1500 standard for embedded core test. A new test scheduling algorithm, which facilitates a concurrent test of both unwrapped logic blocks and IEEE 1500-wrapped cores, is proposed, and experiments show that it outperforms a previous approach when the available number of tester channels and/or the number of unwrapped logic blocks are small. Qiang Xu 0001, Nicola Nicolici |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2004 | Multi-Frequency Test Access Mechanism Design for Modular SOC TestingabstractThis paper investigates the applicability of multi-frequency test access mechanism (TAM) design for reducing the system-on-a-chip (SOC) test application time. Based on the bandwidth matching concept the proposed algorithms explore a larger solution space, which, as shown by experimental data, can lead to improved test application time. Qiang Xu 0001, Nicola Nicolici |
Asian Test Symposium | 1 |
| 2004 | Wrapper Design for Testing IP Cores with Multiple Clock DomainsabstractThis paper addresses the testability problems raised by embedded cores with multiple clock domains. The proposed solution, based on a novel core wrapper architecture, shows how multi-frequency at-speed test response capture can be achieved using low-speed testers synchronized with high-speed on-chip generated clocks. Using experimental data, the trade-offs between the number of tester channels, testing time, area overhead and power dissipation are discussed. Qiang Xu 0001, Nicola Nicolici |
DATE | 1 |
| 2004 | Time/Area Tradeoffs in Testing Hierarchical SOCs With Hard Mega-CoresabstractMotivated by the presence of mega-cores in hierarchical systems-on-a-chip, This work describes a new framework for the design space exploration of multi-level test access mechanisms. Test resources are placed next to the mega-core wrappers, which removes the constraint that upper-level TAM width must be wider than the internal TAM width of the mega-core. The proposed solution can rapidly analyze the tradeoffs between test application time and area overhead and it facilitates test data reuse for hard mega-cores. Qiang Xu 0001, Nicola Nicolici |
ITC | 1 |
| 2003 | Delay Fault Testing of Core-Based Systems-on-a-Chi
Qiang Xu 0001, Nicola Nicolici |
DATE | 1 |
| 2003 | Hardware/Software Co-testing of Embedded Memories in Complex SOCsabstractA novel approach for testing embedded memories in complex systems-on-a-chip (SOCs) is presented. The proposed solution aims to balance the usage of the existing on-chip resources and dedicated design for test (DFT) hardware such that the functional power constraints are not exceeded during test while trading-off the testing time against DFT area and performance overhead. The suitability of software-centric and hardware-centric approaches for embedded memory testing is examined and to combine the advantages of both directions, a new built-in self-test (BIST)-based method, called hardware/software co-testing, is introduced. The proposed solution is programmable, scalable and guarantees low routine overhead. Bai Hong Fang, Qiang Xu 0001, Nicola Nicolici |
ICCAD | 2 |
| 2003 | On Reducing Wrapper Boundary Register Cells in Modular SOC Testing
Qiang Xu 0001, Nicola Nicolici |
ITC | 1 |