VLDB 2026 Research / reviewers in the wild / expert
Kwang-Ting Cheng
dblp:c/KwangTingCheng · also K. T. Tim Cheng, Kwang-Ting (Tim) Cheng, Tim Kwang-Ting Cheng
· DBLP profile ↗
467ranked-venue papers
42as first author
104since 2021 · last 2026
0000-0002-3885-4912ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 338 · 42 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 71 · 30 since 2021Applied, interdisciplinary, general and emerging computing · 55 · 36 since 2021Artificial intelligence and machine learning · 54 · 33 since 2021Software engineering, systems software and programming languages · 37 · 5 since 2021Databases, data management, data science and information retrieval · 4Human-computer interaction and ubiquitous computing · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Relation-R1: Progressively Cognitive Chain-of-Thought Guided Reinforcement Learning for Unified Relation ComprehensionabstractRecent advances in multi-modal large language models (MLLMs) have significantly improved object-level grounding and region captioning. However, they remain limited in visual relation understanding, struggling even with binary relation detection, let alone N-ary relations involving multiple semantic roles. The core reason is the lack of modeling for structural semantic dependencies among multi-entities, leading to over-reliance on language priors (e.g., defaulting to "person drinks a milk" if a person is merely holding it). To this end, we propose Relation-R1, the first unified relation comprehension framework that explicitly integrates cognitive chain-of-thought (CoT)-guided supervised fine-tuning (SFT) and group relative policy optimization (GRPO) within a reinforcement learning (RL) paradigm. Specifically, we first establish foundational reasoning capabilities via SFT, enforcing structured outputs with thinking processes. Then, GRPO is utilized to refine these outputs via multi-rewards optimization, prioritizing visual-semantic grounding over language-induced biases, thereby improving generalization capability. Furthermore, we investigate the impact of various CoT strategies within this framework, demonstrating that a specific-to-general progressive approach in CoT guidance further improves generalization, especially in capturing synonymous N-ary relations. Extensive experiments on widely-used PSG and SWiG datasets demonstrate that Relation-R1 achieves state-of-the-art performance in both binary and N-ary relation understanding. Lin Li 0065, Wei Chen 0070, Jiahui Li 0003, Kwang-Ting Cheng, Long Chen 0016 |
AAAI | 4 |
| 2026 | A Study of Finetuning Video Transformers for Multi-view Geometry TasksabstractThis paper presents an investigation of vision transformer learning for multi-view geometry tasks, such as optical flow estimation, by fine-tuning video foundation models. Unlike previous methods that involve custom architectural designs and task-specific pretraining, our research finds that general-purpose models pretrained on videos can be readily transferred to multi-view problems with minimal adaptation. The core insight is that general-purpose attention between patches learns temporal and spatial information for geometric reasoning. We demonstrate that appending a linear decoder to the Transformer backbone produces satisfactory results, and iterative refinement can further elevate performance to state-of-the-art levels. This conceptually simple approach achieves top cross-dataset generalization results for optical flow estimation with end-point error (EPE) of 0.69, 1.78, and 3.15 on the Sintel clean, Sintel final, and KITTI datasets, respectively. Our method additionally establishes a new record on the online test benchmark with EPE values of 0.79, 1.88, and F1 value of 3.79. Applications to 3D depth estimation and stereo matching also show strong performance, illustrating the versatility of video-pretrained models in addressing geometric vision tasks. Huimin Wu 0001, Kwang-Ting Cheng, Stephen Lin 0001, Zhirong Wu |
AAAI | 2 |
| 2026 | DS-CIM: Digital Stochastic Computing-In-Memory Featuring Accurate OR-Accumulation via Sample Region Remapping for Edge AI ModelsabstractStochastic computing (SC) offers hardware simplicity but suffers from low throughput, while high-throughput Digital Computing-in-Memory (DCIM) is bottlenecked by costly adder logic for matrix-vector multiplication (MVM). To address this trade-off, this paper introduces a digital stochastic CIM (DS-CIM) architecture that achieves both high accuracy and efficiency. We implement signed multiply-accumulation (MAC) in a compact, unsigned OR-based circuit by modifying the data representation. Throughput is enhanced by replicating this low-cost circuit 64 times with only a 1× area increase. Our core strategy, a shared Pseudo Random Number Generator (PRNG) with 2D partitioning, enables single-cycle mutually exclusive activation to eliminate OR-gate collisions. We also resolve the 1s saturation issue via stochastic process analysis and data remapping, significantly improving accuracy and resilience to input sparsity. Our high-accuracy DS-CIM1 variant achieves 94.45% accuracy for INT8 ResNet18 on CIFAR-10 with a root-mean-squared error (RMSE) of just 0.74%. Meanwhile, our high-efficiency DS-CIM2 variant attains an energy efficiency of 3566.1 TOPS/W and an area efficiency of 363.7 TOPS/mm2, while maintaining a low RMSE of 3.81%. The DS-CIM capability with larger models is further demonstrated through experiments with INT8 ResNet50 on ImageNet and the FP8 LLaMA-7B model. Kunming Shao, Jiangnan Yu, Zhipeng Liao, Yi Zou 0001, Kwang-Ting Cheng, Chi-Ying Tsui |
DATE | 7 |
| 2026 | BioSeek: A Design Generation Framework of Biosignal Processors with Large-Language Models for Edge Healthcare ApplicationsabstractDeep neural network (DNN)-based methodologies have shown impressive performance and robustness in the detection of abnormalities and decoding of multi-modal biosignals. While the use of DNNs provides promising classification and decoding capabilities, it also introduces significant design and cost challenges for the implementation of biomedical System on Chips (SoC). To address the increasing demand for advanced and efficient DNN-based healthcare solutions at the edge, we propose BioSeek, an agile design generation framework enhanced by cutting-edge large-language models (LLM). BioSeek offers a comprehensive solution to the design challenges associated with biosignal processors. The effectiveness of BioSeek is evaluated through the design generation of both application-specific and versatile biosignal processors, demonstrating performance that is competitive with existing solutions. Fengshi Tian, Jiakun Zheng, Hui Wu 0010, Zilu Liu, Jinbo Chen 0002, Shiqi Zhao 0001, Jie Yang 0033, Mohamad Sawan, Chi-Ying Tsui, Kwang-Ting Cheng |
ISCAS | 10 |
| 2026 | Prompt-guided foundation model tuning for pathology image classification
Zhengjie Zhu, Kwang-Ting Cheng |
Medical Image Anal. | 3 |
| 2026 | TSAR: A two-stage approach to motion artifact reduction in OCTA images
Benteng Ma, Xiaomeng Li 0001, Dongping Shao, Chubin Ou, Lin An, Kwang-Ting Cheng |
Pattern Recognit. | 9 |
| 2026 | Topology-Preserving retinal vascular segmentation via sparse persistent homology and MoE convolution
Benteng Ma, Xiaomeng Li 0001, Bin Pu, Kwang-Ting Cheng |
Pattern Recognit. | 4 |
| 2026 | Configurable Dataflow and Adaptive Mapping Optimization for Hybrid ReRAM and SRAM Compute-in-Memory AcceleratorabstractHybrid compute-in-memory (CIM) designs have been proposed recently to facilitate the storing of large number of weights of a neural network on-chip. Notably, ReSCIM wang2024res pairs an SRAM cell with a dedicated ReRAM crossbar, allowing ReRAM to serve as the local storage, significantly enhancing the storage capacity of the SRAM-CIM. The SRAM is custom-designed not only to serve as a storage element for CIM but also to function as a sense amplifier to retrieve the data from the ReRAM, which enables super high bandwidth of weight data loading into the CIM engine. However, existing mapping tools for CIM are inadequate for ReSCIM since they do not fully exploit the unique hardware characteristics and advantages of this novel architecture. In this work, we propose an analytical energy and latency model, which incorporates four key factors: hardware, workload, dataflow, and mapping (HWDM), for executing inference of neural network on the ReSCIM accelerator. Specifically, we first characterize the ReSCIM accelerator hardware specifications and the neural network layers. Next, we introduce three dataflows for ReSCIM, leveraging the high weight-loading bandwidth to reduce memory access for various workloads and layer types. Finally, we develop an algorithm to generate optimal mapping and dataflow strategies aimed at minimizing latency or energy consumption. Using our HWDM model, we design a tile-based ReSCIM accelerator and conduct extensive simulations to obtain the cycle-accurate latency and gate-level energy consumption metrics for inference across different neural networks. We conduct design space exploration (DSE) using the HWDM model on a comprehensive set of benchmarks to minimize inference energy or latency. Experimental results show that our optimal ReSCIM accelerator achieves a 44% reduction in EDP reduction compared to the weight-stationary and fixed mapping baseline for SEResNet50. Moreover, our design exhibits 1.74× higher energy efficiency than the state-of-the-art hybrid TL-nvSRAM wang2023tl accelerator on ResNet 18. Jingyu He, Kunming Shao, Kwang-Ting Cheng, Chi-Ying Tsui |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2026 | FedHAC: Towards Robust Federated Multi-Lesion Segmentation With Heterogeneous Annotation CompletenessabstractFederated learning (FL) has emerged as a promising paradigm for collaborative medical image segmentation across institutions while preserving data privacy. Despite great efforts in addressing cross-client annotation heterogeneity FL, the prevalent annotation completeness heterogeneity in clinical practice due to varying diagnostic priorities has been completely overlooked, hindering the deployment of FL. In this paper, we formulate such a challenge and propose FedHAC for incompleteness-robust medical image segmentation. FedHAC consists of three modules, i.e., Global Class Prototype Alignment (GCPA), Annotation Completeness-Aware Aggregation (ACAA), and GMM-driven Progressive Correction (GPC). Specifically, GCPA constructs a noise-resilient warm-up model through proximal-term regularization and prototype alignment. ACAA estimates client-wise annotation completeness and dynamically prioritizes high-quality clients. GPC groups clients into "noisy" and "clean" via GMM for progressive annotation correction to minimize error propagation. Extensive comparison experiments and ablation studies on public datasets demonstrate the superiority of FedHAC over state-of-the-art methods under various levels of annotation incompleteness. Yangyang Xiang, Li Yu 0003, Kwang-Ting Cheng, Zengqiang Yan |
IEEE J. Biomed. Health Informatics | 4 |
| 2026 | Balancing FP8 Computation Accuracy and Efficiency on Digital CIM via Shift-Aware On-the-Fly Aligned-Mantissa Bitwidth PredictionabstractFP8 low-precision formats have gained significant adoption in transformer inference and training. However, existing digital compute-in-memory (DCIM) architectures face challenges in supporting variable FP8 aligned-mantissa bitwidths, as unified alignment strategies and fixed-precision multiply accumulate (MAC) units struggle to handle input data with diverse distributions. This work presents a flexible FP8 DCIM accelerator with three innovations: 1) a dynamic shift-aware bitwidth prediction (DSBP) with on-the-fly input prediction that adaptively adjusts weight (2/4/6/8b) and input ($2\sim 12$b) aligned-mantissa precision; 2) a FIFO-based input alignment unit (FIAU) replacing complex barrel shifters with pointer-based control; and 3) a precision-scalable INT MAC array achieving flexible weight precision with minimal overhead. Implemented in 28-nm CMOS with a$64~\times ~96$CIM array, the design achieves 20.4 TFLOPS/W for fixed E5M7, demonstrating$2.8\times $higher FP8 efficiency than previous work while supporting all FP8 formats. Results on Llama-7b show that the DSBP achieves higher efficiency than fixed bitwidth mode at the same accuracy level on both BoolQ and Winogrande datasets, with configurable parameters enabling flexible accuracy–efficiency tradeoffs. Kunming Shao, Zhipeng Liao, Xijie Huang, Kwang-Ting Cheng, Chi-Ying Tsui, Yi Zou 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | A 10.60 μW 150 GOPS Mixed-Bit-Width Sparse CNN Accelerator for Life-Threatening Ventricular Arrhythmia DetectionabstractThis paper proposes an ultra-low power, mixed-bit-width sparse convolutional neural network (CNN) accelerator to accelerate ventricular arrhythmia (VA) detection. The chip achieves 50% sparsity in a quantized 1D CNN using a sparse processing element (SPE) architecture. Measurement on the prototype chip TSMC 40nm CMOS low-power (LP) process for the VA classification task demonstrates that it consumes 10.60 μW of power while achieving a performance of 150 GOPS and a diagnostic accuracy of 99.95%. The computation power density is only 0.57 μW/mm2, which is 14.23× smaller than state-of-the-art works, making it highly suitable for implantable and wearable medical devices. Zhenge Jia, Zheyu Yan, Jay Mok, Manto Yung, Yu Liu 0007, Wujie Wen, Luhong Liang, Kwang-Ting Cheng, Xiaobo Sharon Hu, Yiyu Shi 0001 |
ASP-DAC | 10 |
| 2025 | SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and TrainingabstractExisting text-to-image (T2I) diffusion models face several limitations, including large model sizes, slow runtime, and low-quality generation on mobile devices. This paper aims to address all of these challenges by developing an extremely small and fast T2I model that generates high-resolution and high-quality images on mobile platforms. We propose several techniques to achieve this goal. First, we systematically examine the design choices of the network architecture to reduce model parameters and latency, while ensuring high-quality generation. Second, to further improve generation quality, we employ cross-architecture knowledge distillation from a much larger model, using a multi-level approach to guide the training of our model from scratch. Third, we enable a few-step generation by integrating adversarial guidance with knowledge distillation. For the first time, our model SnapGen, demonstrates the generation of 10242px images on a mobile device around 1.4 seconds. On ImageNet-1K, our model, with only 372M parameters, achieves an FID of 2.06 for 2562px generation. On T2I benchmarks (i.e., GenEval and DPG-Bench), our model with merely 379M parameters, surpasses large-scale models with billions of parameters at a significantly smaller size (e.g., 7× smaller than SDXL, 14× smaller than IF-XL). Jierun Chen, Dongting Hu, Xijie Huang, Huseyin Coskun, Arpit Sahni, Aarush Gupta, Anujraaj Goyal, Dishani Lahiri, Yerlan Idelbayev, Junli Cao, Yanyu Li, Kwang-Ting Cheng, Shueng-Han Gary Chan, Mingming Gong, Sergey Tulyakov, Anil Kag, Yanwu Xu 0003, Jian Ren 0005 |
CVPR | 13 |
| 2025 | APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-DesignabstractDNN accelerators, significantly advanced by model compression and specialized dataflow techniques, have marked considerable progress. However, the frequent access of highprecision partial sums (PSUMs) leads to excessive memory demands in architectures utilizing input/weight stationary dataflows. Traditional compression strategies have typically overlooked PSUM quantization, which may account for 69% of power consumption. This study introduces a novel Additive Partial Sum Quantization (APSQ) method, seamlessly integrating PSUM accumulation into the quantization framework. A grouping strategy that combines APSQ with PSUM quantization enhanced by a reconfigurable architecture is further proposed. The APSQ performs nearly lossless on NLP and CV tasks across BERT, Segformer, and EfficientViT models while compressing PSUMs to INT8. This leads to a notable reduction in energy costs by $\mathbf{2 8-8 7 \%}$. Extended experiments on LLaMA2-7B demonstrate the potential of APSQ for large language models. Code is available at https://github.com/Yonghao-Tan/APSQ. Yonghao Tan, Pingcheng Dong, Yongkun Wu, Yu Liu 0007, Shih-Yang Liu, Xijie Huang, Luhong Liang, Kwang-Ting Cheng |
DAC | 11 |
| 2025 | SynDCIM: A Performance-Aware Digital Computing-in-Memory Compiler with Multi-Spec-Oriented Subcircuit SynthesisabstractDigital Computing-in-Memory (DCIM) is an innovative technology that integrates multiply-accumulation (MAC) logic directly into memory arrays to enhance the performance of modern AI computing. However, the need for customized memory cells and logic components currently necessitates significant manual effort in DCIM design. Existing tools for facilitating DCIM macro designs struggle to optimize subcircuit synthesis to meet user-defined performance criteria, thereby limiting the potential system-level acceleration that DCIM can offer. To address these challenges and enable the agile design of DCIM macros with optimal architectures, we present SynDCIM - a performance-aware DCIM compiler that employs multi-spec-oriented subcircuit synthesis. SynDCIM features an automated performance-to-layout generation process that aligns with user-defined performance expectations. This is supported by a scalable subcircuit library and a multi-spec-oriented searching algorithm for effective subcircuit synthesis. The effectiveness of SynDCIM is demonstrated through extensive experiments and validated with a test chip fabricated in a 40nm CMOS process. Testing results reveal that designs generated by SynDCIM exhibit competitive performance when compared to state-of-the-art manually designed DCIM macros. Kunming Shao, Fengshi Tian, Jiakun Zheng, Jia Chen 0032, Jingyu He, Hui Wu 0010, Jinbo Chen 0002, Xihao Guan, Fengbin Tu, Jie Yang 0033, Mohamad Sawan, Kwang-Ting Cheng, Chi-Ying Tsui |
DATE | 14 |
| 2025 | CoXplorer: Multi-Staged Co-Exploration Framework for AI Model Compression and Accelerator DesignabstractThe rapid evolution of artificial intelligence (AI) algorithms demands efficient computing chips, positioning algorithm-hardware co-design as a crucial optimization strategy. However, automating the co-design process remains challenging due to the lack of a unified exploration framework for both algorithmic and hardware domains, as existing tools - hardware design space exploration (DSE) and compression neural architecture search (Compression NAS) - operate independently, relying entirely on manual collaboration. This paper presents CoXplorer, a co-exploration framework that connects model-compression optimization space and architecture design space. We make three key contributions: (1) a multi-staged co-design space decomposition method that enables systematic exploration of compression-hardware design choices with reduced complexity, (2) an AC-Copilot toolchain enhanced with multi-grained performance modeling driven by hardware simulation-compilation hierarchical cooperation to fulfill various evaluation requirements of co-exploration, enabling balanced simulation accuracy-efficiency trade-offs, and (3) a co-exploration workflow with hierarchical and bottleneck-guided search for harmonizing optimization objectives of both model and hardware design spaces, resulting in improved search efficiency. We validate the CoXplorer on two edge chips, which achieve 53.7% throughput and 45.8% energy efficiency improvements for the CNN acceleration, and 7.5× speedup with 9.9× energy efficiency boost for the Transformer acceleration. A case study on large language model acceleration shows CoXplorer’s extensibility to emerging workloads, enhancing LLAMA2-7B inference throughput from 6.75 to 25.46 tokens/s via co-optimization with compression and near-memory computing architecture. Songchen Ma, Yonghao Tan, Pingcheng Dong, Di Pang, Yu Liu 0007, Luhong Liang, Kwang-Ting Cheng, Fengbin Tu |
ICCAD | 10 |
| 2025 | Memory Efficient Transformer Adapter for Dense PredictionsabstractWhile current Vision Transformer (ViT) adapter methods have shown promising accuracy, their inference speed is implicitly hindered by inefficient memory access operations, e.g., standard normalization and frequent reshaping. In this work, we propose META, a simple and fast ViT adapter that can improve the model's memory efficiency and decrease memory time consumption by reducing the inefficient memory access operations. Our method features a memory-efficient adapter block that enables the common sharing of layer normalization between the self-attention and feed-forward network layers, thereby reducing the model's reliance on normalization operations. Within the proposed block, the cross-shaped self-attention is employed to reduce the model's frequent reshaping operations. Moreover, we augment the adapter block with a lightweight convolutional branch that can enhance local inductive biases, particularly beneficial for the dense prediction tasks, e.g., object detection, instance segmentation, and semantic segmentation. The adapter block is finally formulated in a cascaded manner to compute diverse head features, thereby enriching the variety of feature representations. Empirically, extensive evaluations on multiple representative datasets validate that META substantially enhances the predicted quality, while achieving a new state-of-the-art accuracy-efficiency trade-off. Theoretically, we demonstrate that META exhibits superior generalization capability and stronger adaptability. Pingcheng Dong, Kwang-Ting Cheng |
ICLR | 4 |
| 2025 | DPE-CIM: Compute-In-Memory Accelerator using Dynamic Posit Encoding and Speculative AlignmentabstractIn this study, we propose two novel approaches to address the memory wall of AI accelerators. First, based on Posit, we introduce a new format called dynamic Posit encoding (DPE), which dynamically extends the dynamic range of its representation at run time with minimal hardware overhead. Using two exponent encoding schemes, DPE accommodates the data distribution with lower quantization error compared to regular Posit. Second, we propose a compute-in-memory (CIM) architecture to implement DPE multiply-and-accumulate (MAC) computation to reduce weight data movement. Traditional CIM proposed for floating-point-alike MAC computation uses a comparator tree (CT) to compute the maximum exponent, enabling the CIM to locus on integer MAC. However, the CT-based design has poor scalability as the number of inputs increases. To address this, we propose a speculative input alignment design that significantly reduces the delay, area, and power consumption for the max exponent computation. We show that DPE outperforms state-of-the-art quantization approaches across various neural network models through software evaluations. Hardware synthesis and simulation results further illustrate that our approach achieves significant energy efficiency and area efficiency improvement compared to the state-of-the-art posit processing element. Jingyu He, Kwang-Ting Cheng, Chi-Ying Tsui |
ISCAS | 2 |
| 2025 | A Flexible Precision Scaling Deep Neural Network Accelerator with Efficient Weight CombinationabstractDeploying mixed-precision neural networks on edge devices is friendly to hardware resources and power consumption. To support fully mixed-precision neural network inference, it is necessary to design flexible hardware accelerators for continuous varying precision operations. However, the previous works have issues on hardware utilization and overhead of reconfigurable logic. In this paper, we propose an efficient accelerator for 2 ∼ 8-bit precision scaling with serial activation input and parallel weight preloaded. First, we set two loading modes for the weight operands and decompose the weight into the corresponding bitwidths, which extends the weight precision support efficiently. Then, to improve hardware utilization of low-precision operations, we design the architecture that performs bit-serial MAC operation with systolic dataflow, and the partial sums are combined spatially. Furthermore, we designed an efficient carry save adder tree supporting both signed and unsigned number summation across rows. The experiment result shows that the proposed accelerator, synthesized with TSMC 28nm CMOS technology, achieves peak throughput of 4.09TOPS and peak energy efficiency of 68.94TOPS/W at 2/2-bit operations. Kunming Shao, Fengshi Tian, Kwang-Ting Cheng, Chi-Ying Tsui, Yi Zou 0001 |
ISCAS | 4 |
| 2025 | NeuroEye: A 54.59mW, 12200FPS Event-Driven Near-Sensor Eye-Tracking Processor with Pipelined Spatial-Temporal Spike-StreamingabstractThis paper presents a design of an eye tracking system based on neuromorphic computing to enhance user interaction in augmented reality (AR) and virtual reality (VR) environments. Traditional methods face challenges of high computational demands and power consumption. To address these issues, we propose a fully-spike eye-tracking system that utilizes dynamic vision sensors (DVS) for asynchronous pixel-level change detection, thereby reducing data redundancy and improving temporal resolution. We proposed a pipelined processor specifically tailored for handling DVS events and Spiking Neural Network (SNN) computations. Our spatial-temporal spike-streaming architecture enables cascaded computation across all layers, achieving high energy efficiency and high frame rate in eye-tracking tasks. Implemented in a 40nm CMOS process, NeuroEye demonstrates up to 12200 frame-per-second (FPS) and 4.47uJ/frame energy efficiency with 54.59mW power consumption in post-layout evaluations. Jiakun Zheng, Fengshi Tian, Jinbo Chen 0002, Chaoming Fang, Jie Yang 0033, Mohamad Sawan, Kwang-Ting Cheng, Chi-Ying Tsui |
ISCAS | 8 |
| 2025 | DIRC-RAG: Accelerating Edge RAG with Robust High-Density and High-Loading-Bandwidth Digital In-ReRAM ComputationabstractRetrieval-Augmented Generation (RAG) enhances large language models (LLMs) by integrating external knowledge retrieval but faces challenges on edge devices due to high storage, energy, and latency demands. Computing-in-Memory (CIM) offers a promising solution by storing document embeddings in CIM macros and enabling in-situ parallel retrievals but is constrained by either low memory density or limited computational accuracy. To address these challenges, we present DIRC-RAG, a novel edge RAG acceleration architecture leveraging Digital In-ReRAM Computation (DIRC). DIRC integrates a high-density multi-level ReRAM subarray with an SRAM cell, utilizing SRAM and differential sensing for robust ReRAM readout and digital multiply-accumulate (MAC) operations. By storing all document embeddings within the CIM macro, DIRC achieves ultra-low-power, single-cycle data loading, substantially reducing both energy consumption and latency compared to off-chip DRAM. A query-stationary (QS) dataflow is supported for RAG tasks, minimizing on-chip data movement and reducing SRAM buffer requirements. We introduce error optimization for the DIRC ReRAM-SRAM cell by extracting the bit-wise spatial error distribution of the ReRAM subarray and applying targeted bit-wise data remapping. An error detection circuit is also implemented to enhance readout resilience against device-and circuit-level variations.Simulation results demonstrate that DIRC-RAG under TSMC 40nm process achieves an on-chip non-volatile memory density of 5.18Mb/mm2and a throughput of 131 TOPS. It delivers a 4MB retrieval latency of 5.6μs/query and an energy consumption of 0.956μJ/query, while maintaining the retrieval precision. Kunming Shao, Zhipeng Liao, Jiangnan Yu, Xijie Huang, Jingyu He, Fengshi Tian, Yi Zou 0001, Kwang-Ting Cheng, Chi-Ying Tsui |
ISLPED | 11 |
| 2025 | SR-SAM: Subspace Regularization for Domain Generalization of Segment Anything Model
Xixi Jiang, Kwang-Ting Cheng, Xin Yang 0008 |
MICCAI (10) | 4 |
| 2025 | MedIAnomaly: A comparative study of anomaly detection in medical images
Yu Cai 0005, Hao Chen 0011, Kwang-Ting Cheng |
Medical Image Anal. | 4 |
| 2025 | Labeled-to-unlabeled distribution alignment for partially-supervised multi-organ medical image segmentation
Xixi Jiang, Kangyi Liu, Kwang-Ting Cheng, Xin Yang 0008 |
Medical Image Anal. | 5 |
| 2025 | Rethinking boundary detection in deep learning-based medical image segmentation
Yi Lin 0009, Kwang-Ting Cheng, Hao Chen 0011 |
Medical Image Anal. | 5 |
| 2025 | Generalized Task-Driven Medical Image Quality Enhancement With Gradient PromotionabstractThanks to the recent achievements in task-driven image quality enhancement (IQE) models like ESTR (Liu et al. 2023), the image enhancement model and the visual recognition model can mutually enhance each other's quantitation while producing high-quality processed images that are perceivable by our human vision systems. However, existing task-driven IQE models tend to overlook an underlying fact-different levels of vision tasks have varying and sometimes conflicting requirements of image features. To address this problem, this paper proposes a generalized gradient promotion (GradProm) training strategy for task-driven IQE of medical images. Specifically, we partition a task-driven IQE system into two sub-models, i.e., a mainstream model for image enhancement and an auxiliary model for visual recognition. During training, GradProm updates only parameters of the image enhancement model using gradients of the visual recognition model and the image enhancement model, but only when gradients of these two sub-models are aligned in the same direction, which is measured by their cosine similarity. In case gradients of these two sub-models are not in the same direction, GradProm only uses the gradient of the image enhancement model to update its parameters. Theoretically, we have proved that the optimization direction of the image enhancement model will not be biased by the auxiliary visual recognition model under the implementation of GradProm. Empirically, extensive experimental results on four public yet challenging medical image datasets demonstrated the superior performance of GradProm over existing state-of-the-art methods. Kwang-Ting Cheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Improving Efficiency in Multi-Modal Autonomous Embedded Systems Through Adaptive GatingabstractThe parallel advancement of AI and IoT technologies has recently boosted the development of multi-modal computing ($M^{2}C$) on pervasive autonomous embedded systems (AES).$M^{2}C$takes advantage of data from different modalities such as images, audio, and text and is able to achieve notable improvements in accuracy. However, achieving these accuracy gains often comes at the cost of increased computational complexity and energy consumption. Furthermore, the presence of numerous advanced sensors in these systems significantly contributes to power consumption, exacerbating the issue of limited power resources. Collectively, these challenges pose difficulties in deploying$M^{2}C$on small embedded devices with scarce energy resources. In this article, we propose anAdaptiveModalityGating technique calledAMGfor in-situ$M^{2}C$applications. The primary objective ofAMGis to conserve energy while preserving the accuracy advantages of$M^{2}C$. To achieve this goal,AMGincorporates two first-of-its-kind designs. Firstly, it introduces a novel semi-gating architecture that enables partial modality sensor power gating. Specifically, we devise the de-centralizedAMG(D-AMG) and centralizedAMG(C-AMG) architecture. The former buffers raw data on sensors while the latter buffers raw data on the computing board, which are suitable for different edge scenarios respectively. Secondly, it facilitates a self-initialization/tuning process on the AES, which is supported by carefully-built analytical model. Extensive evaluations demonstrate the effectiveness ofAMG. It achieves a 1.6x to 3.8x throughput higher than other power management methods and improves the lifespan of AES by 10% to 280% longer within the same energy budget, while satisfying all performance and latency requirements across various scenarios. Xiaofeng Hou, Chao Li 0009, Jiacheng Liu 0001, Xuehan Tang, Kwang-Ting Cheng, Minyi Guo |
IEEE Trans. Computers | 6 |
| 2025 | Exploiting the Memory-Compute-Coupling Feature for CIM Accelerator Design OptimizationabstractSRAM computing-in-memory (CIM) accelerators have evolved as a promising solution to the memory wall problem in neural network (NN) models. By integrating memory and compute resources in each macro, CIM accelerators offer massive in-situ computing parallelism and large memory capacity, enabling spatial mapping with layer fusion and potentially keeping layers stationary in CIM. However, CIM’s memory-compute coupling (MCC) feature poses challenges in designing CIM accelerators. From an architecture aspect, designers must balance CIM’s memory and compute resources by optimizing the macro’s memory-compute ratio (MCR) configuration across diverse scenarios. From a mapping aspect, conventional mappings, which allocate each macro exclusively to each layer, face two major problems: a layer-fusion dilemma (the accelerator suffers from excessive memory access due to layer replications or performance degradation due to load imbalance) and a layer-eviction issue (storing layers stationary in CIM is usually infeasible due to limited CIM capacity). To address these challenges, this paper introduces MCC-DSE, an MCC-aware Design Space Exploration framework for architecture-mapping co-optimization of CIM accelerators. We also propose a three-axis CIM division mapping, which interleaves multiple layers in each macro to concurrently optimize memory access and performance during layer fusion as well as reserves a part of CIM memory in each macro for layer pinning. Compared to baseline architecture and mapping, MCC-DSE shows a 1.4x 8.3x EDP reduction across various workloads and chip areas. Moreover, MCC-DSE provides insights into CIM accelerator optimization, such as selecting optimal MCR and configuring CIM dynamically for different scenarios. Yongkun Wu, Jia Chen 0032, Zhenhua Zhu 0002, Jingyu He, Pingcheng Dong, Yonghao Tan, Xin Zhao 0044, Liang Chang 0002, Yu Wang 0002, Fengbin Tu, Chi-Ying Tsui, Kwang-Ting Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 13 |
| 2025 | Boosting Convolution With Efficient MLP-Permutation for Volumetric Medical Image SegmentationabstractRecently, the advent of Vision Transformer (ViT) has brought substantial advancements in 3D benchmarks, particularly in 3D volumetric medical image segmentation (Vol-MedSeg). Concurrently, multi-layer perceptron (MLP) network has regained popularity among researchers due to their comparable results to ViT, albeit with the exclusion of the resource-intensive self-attention module. In this work, we propose a novel permutable hybrid network for Vol-MedSeg, named PHNet, which capitalizes on the strengths of both convolution neural networks (CNNs) and MLP. PHNet addresses the intrinsic anisotropy problem of 3D volumetric data by employing a combination of 2D and 3D CNNs to extract local features. Besides, we propose an efficient multi-layer permute perceptron (MLPP) module that captures long-range dependence while preserving positional information. This is achieved through an axis decomposition operation that permutes the input tensor along different axes, thereby enabling the separate encoding of the positional information. Furthermore, MLPP tackles the resolution sensitivity issue of MLP in Vol-MedSeg with a token segmentation operation, which divides the feature into smaller tokens and processes them individually. Extensive experimental results validate that PHNet outperformed the state-of-the-art methods with lower computational costs on the widely-used yet challenging COVID-19-20, Synapse, LiTS and MSD BraTS benchmarks. The ablation study also demonstrated the effectiveness of PHNet in harnessing the strengths of both CNNs and MLP. The code is available on Github: https://github.com/xiaofang007/PHNet. Yi Lin 0009, Kwang-Ting Cheng, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | SAMCT: Segment Any CT Allowing Labor-Free Task-Indicator PromptsabstractSegment anything model (SAM), a foundation model with superior versatility and generalization across diverse segmentation tasks, has attracted widespread attention in medical imaging. However, it has been proved that SAM would encounter severe performance degradation due to the lack of medical knowledge in training and local feature encoding. Though several SAM-based models have been proposed for tuning SAM in medical imaging, they still suffer from insufficient feature extraction and highly rely on high-quality prompts. In this paper, we propose a powerful foundation model SAMCT allowing labor-free prompts and train it on a collected large CT dataset consisting of 1.1M CT images and 5M masks from public datasets. Specifically, based on SAM, SAMCT is further equipped with a U-shaped CNN image encoder, a cross-branch interaction module, and a task-indicator prompt encoder. The U-shaped CNN image encoder works in parallel with the ViT image encoder in SAM to supplement local features. Cross-branch interaction enhances the feature expression capability of the CNN image encoder and the ViT image encoder by exchanging global perception and local features from one to the other. The task-indicator prompt encoder is a plug-and-play component to effortlessly encode task-related indicators into prompt embeddings. In this way, SAMCT can work in an automatic manner in addition to the semi-automatic interactive strategy in SAM. Extensive experiments demonstrate the superiority of SAMCT against the state-of-the-art task-specific and SAM-based medical foundation models on various tasks. The code, data, and model checkpoints are available at https://github.com/xianlin7/SAMCT. Xian Lin, Yangyang Xiang, Zhehao Wang, Kwang-Ting Cheng, Zengqiang Yan, Li Yu 0003 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Dynamic Subcluster-Aware Network for Few-Shot Skin Disease ClassificationabstractThis article addresses the problem of few-shot skin disease classification by introducing a novel approach called the subcluster-aware network (SCAN) that enhances accuracy in diagnosing rare skin diseases. The key insight motivating the design of SCAN is the observation that skin disease images within a class often exhibit multiple subclusters, characterized by distinct variations in appearance. To improve the performance of few-shot learning (FSL), we focus on learning a high-quality feature encoder that captures the unique subclustered representations within each disease class, enabling better characterization of feature distributions. Specifically, SCAN follows a dual-branch framework, where the first branch learns classwise features to distinguish different skin diseases, and the second branch aims to learn features, which can effectively partition each class into several groups so as to preserve the subclustered structure within each class. To achieve the objective of the second branch, we present a cluster loss to learn image similarities via unsupervised clustering. To ensure that the samples in each subcluster are from the same class, we further design a purity loss to refine the unsupervised clustering results. We evaluate the proposed approach on two public datasets for few-shot skin disease classification. The experimental results validate that our framework outperforms the state-of-the-art methods by around 2%-5% in terms of sensitivity, specificity, accuracy, and F1-score on the SD-198 and Derm7pt datasets. Shuhan Li, Xiaomeng Li 0001, Xiaowei Xu 0004, Kwang-Ting Cheng |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | DTMFormer: Dynamic Token Merging for Boosting Transformer-Based Medical Image SegmentationabstractDespite the great potential in capturing long-range dependency, one rarely-explored underlying issue of transformer in medical image segmentation is attention collapse, making it often degenerate into a bypass module in CNN-Transformer hybrid architectures. This is due to the high computational complexity of vision transformers requiring extensive training data while well-annotated medical image data is relatively limited, resulting in poor convergence. In this paper, we propose a plug-n-play transformer block with dynamic token merging, named DTMFormer, to avoid building long-range dependency on redundant and duplicated tokens and thus pursue better convergence. Specifically, DTMFormer consists of an attention-guided token merging (ATM) module to adaptively cluster tokens into fewer semantic tokens based on feature and dependency similarity and a light token reconstruction module to fuse ordinary and semantic tokens. In this way, as self-attention in ATM is calculated based on fewer tokens, DTMFormer is of lower complexity and more friendly to converge. Extensive experiments on publicly-available datasets demonstrate the effectiveness of DTMFormer working as a plug-n-play module for simultaneous complexity reduction and performance improvement. We believe it will inspire future work on rethinking transformers in medical image segmentation. Code: https://github.com/iam-nacl/DTMFormer. Zhehao Wang, Xian Lin, Li Yu 0003, Kwang-Ting Cheng, Zengqiang Yan |
AAAI | 5 |
| 2024 | Genetic Quantization-Aware Approximation for Non-Linear Operations in TransformersabstractNon-linear functions are prevalent in Transformers and their lightweight variants, incurring substantial and frequently underestimated hardware costs. Previous state-of-the-art works optimize these operations by piece-wise linear approximation and store the parameters in look-up tables (LUT), but most of them require unfriendly high-precision arithmetics such as FP/INT 32 and lack consideration of integer-only INT quantization. This paper proposed a genetic LUT-Approximation algorithm namely GQA-LUT that can automatically determine the parameters with quantization awareness. The results demonstrate that GQA-LUT achieves negligible degradation on the challenging semantic segmentation task for both vanilla and linear Transformer models. Besides, proposed GQA-LUT enables the employment of INT8-based LUT-Approximation that achieves an area savings of 81.3~81.7% and a power reduction of 79.3~80.2% compared to the high-precision FP/INT 32 alternatives. Code is available at https://github.com/PingchengDong/GQA-LUT. Pingcheng Dong, Yonghao Tan, Tianwei Ni, Yu Liu 0007, Luhong Liang, Shih-Yang Liu, Xijie Huang, Huaiyu Zhu 0004, Fengwei An, Kwang-Ting Cheng |
DAC | 14 |
| 2024 | RWriC: A Dynamic Writing Scheme for Variation Compensation for RRAM-based In-Memory ComputingabstractRRAM-based compute-in-memory (CIM) suffers from programming variation issues, specifically device-to-device variation (DDV) and cycle-to-cycle variation (CCV), which can have a detrimental impact on inference accuracy. To address these variation issues, we propose RWriC, a dynamic Writing scheme for variation Compensation for RRAM-based CIM. RWriC sequentially programs the weights, implemented by multiple RRAM cells, starting from the high significance cell (HSC) and moving towards the low significance cell (LSC). This approach leverages the knowledge of current cumulative errors and the programming targets (PTs) of other RRAM cells to dynamically adjust the PT of the RRAM currently under programming. By shifting the PT of HSC, RWriC enables the LSC to compensate for the programming errors of the HSC. Moreover, when the variation is substantial, RWriC allows the magnitude of LSC to be scaled up, providing an even wider compensation range. Through the combined application of the shifting and scaling techniques, experimental results show that the inference accuracy for ResNet50 on the CIFAR-10 dataset only drops by 0.9% under 18% device variation. In comparison to the conventional writing scheme, our RWriC approach achieves a 5-11x improvement in variation robustness for ResNet50 and Yolov8 across different tasks. Yucong Huang, Jingyu He, Kwang-Ting Cheng, Chi-Ying Tsui, Terry Tao Ye |
DAC | 3 |
| 2024 | AdaP-CIM: Compute-in-Memory Based Neural Network Accelerator Using Adaptive PositabstractThis study proposes two novel approaches to address memory wall issues in AI accelerator designs for large neural networks. The first approach introduces a new format called adaptive Posit (AdaP) with two exponent encoding schemes that dynamically extend the dynamic range of its representation at run time with minimal hardware overhead. The second approach proposes using compute-in-memory (CIM) with speculative input alignment (SAU) to implement the AdaP multiply-and-accumulate (MAC) computation, significantly reducing the delay, area, and power consumption for the max exponent computation. The proposed approaches outperform state-of-the-art quantization methods and achieve significant energy and area efficiency improvements. Jingyu He, Fengbin Tu, Kwang-Ting Cheng, Chi-Ying Tsui |
DATE | 3 |
| 2024 | Fewer is More: Boosting Math Reasoning with Reinforced Context PruningabstractLarge Language Models (LLMs) have shown impressive capabilities, yet they still struggle with math reasoning.In this work, we propose CoT-Influx, a novel approach that pushes the boundary of few-shot Chain-of-Thoughts (CoT) learning to improve LLM mathematical reasoning.Motivated by the observation that adding more concise CoT examples in the prompt can improve LLM reasoning performance, CoT-Influx employs a coarse-to-fine pruner to maximize the input of effective and concise CoT examples.The pruner first selects as many crucial CoT examples as possible and then prunes unimportant tokens to fit the context window.A math reasoning dataset with diverse difficulty levels and reasoning steps is used to train the pruner, along with a math-specialized reinforcement learning approach.As a result, by enabling more CoT examples with double the context window size in tokens, CoT-Influx significantly outperforms various prompting baselines across various LLMs (LLaMA2-7B, 13B, 70B) and 6 math datasets, achieving up to 4.40% absolute improvements.Remarkably, without any fine-tuning, LLaMA2-70B with CoT-Influx surpasses GPT-3.5 and a wide range of larger LLMs (PaLM, Minerva 540B, etc.) on GSM8K.CoT-Influx is a plug-and-play module for LLMs, adaptable in various scenarios.It's compatible with advanced reasoning prompting techniques, such as self-consistency, and supports different long-context LLMs, including Mistral-7B-v0.3-32K and Yi-6B-200K.Codes are available at https://github.com/HuangOwen/CoT-Influx Xijie Huang, Li Lyna Zhang, Kwang-Ting Cheng, Fan Yang 0024, Mao Yang 0004 |
EMNLP | 3 |
| 2024 | ReSCIM: Variation-Resilient High Weight-Loading Bandwidth In-Memory Computation Based on Fine-Grained Hybrid Integration of Multi-Level ReRAM and SRAM CellsabstractSRAM-CIM is a promising approach to implement efficient accelerator architecture as it enables accurate, energy-efficient AI computing, supporting both analog and digital computation. However, it has low area efficiency. On the other hand, Resistive RAM (ReRAM) provides dense on-chip storage, especially with multi-level cells (MLC), but ReRAM-CIM may introduce inaccuracies due to device variation and only supports analog computation. To leverage the strengths of both technologies, a hybrid architecture that combines them at a fine granularity is desirable. Previous hybrid designs incorporate ReRAM resistors into SRAM to improve storage density. However, they face scalability limitations and restricted signal margins for multi-level RRAM readout, leading to degraded computation accuracy. In this work, we propose ReSCIM, a hybrid compute-in-memory (CIM) architecture that seamlessly integrates multi-level ReRAM into SRAM cells at a fine-grained level. By incorporating a compact ReRAM crossbar in each SRAM cell, a dense CIM marco using SRAM-based computation is achieved. We develop an energy-efficient differential sensing scheme that enables parallel weight loading from local ReRAM crossbars to SRAM cells. This scheme allows multi-bit ReRAM data readout using a single SRAM cell and offers resilience to device variations. Furthermore, We designed a ReSCIM accelerator architecture for efficient AI acceleration, fully utilizing the highly scalable storage and exceptional weight-loading bandwidth. We employ a folded weight-mapping approach for MLC ReRAM cells to guarantee accurate classification even under substantial ReRAM device variations. Experimental results show that ReSCIM accelerators based on both analog and digital-based CIM achieve 60% energy savings and 98% latency savings, and 59× higher area efficiency compared to state-of-the-art all-weights-on-chip AI accelerators on AlexNet. Jingyu He, Kunming Shao, Jiakun Zheng, Fengshi Tian, Kwang-Ting Cheng, Chi-Ying Tsui |
ICCAD | 6 |
| 2024 | DoRA: Weight-Decomposed Low-Rank AdaptationabstractAmong the widely used parameter-efficient fine-tuning (PEFT) methods, LoRA and its variants have gained considerable popularity because of avoiding additional inference costs. However, there still often exists an accuracy gap between these methods and full fine-tuning (FT). In this work, we first introduce a novel weight decomposition analysis to investigate the inherent differences between FT and LoRA. Aiming to resemble the learning capacity of FT from the findings, we propose Weight-Decomposed Low-Rank Adaptation (DoRA). DoRA decomposes the pre-trained weight into two components, magnitude and direction, for fine-tuning, specifically employing LoRA for directional updates to efficiently minimize the number of trainable parameters. By employing DoRA, we enhance both the learning capacity and training stability of LoRA while avoiding any additional inference overhead. DoRA consistently outperforms LoRA on fine-tuning LLaMA, LLaVA, and VL-BART on various downstream tasks, such as commonsense reasoning, visual instruction tuning, and image/video-text understanding. The code is available at https://github.com/NVlabs/DoRA. Shih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov 0001, Yu-Chiang Frank Wang, Kwang-Ting Cheng, Min-Hung Chen |
ICML | 6 |
| 2024 | A Tale of Two Domains: Exploring Efficient Architecture Design for Truly Autonomous ThingsabstractAutonomous Things (AuT) refers to a collection of self-sufficient tiny devices capable of performing intelligent computations. Looking ahead, AuT promises to enable ubiquitous deployment of intelligence on many emerging consumer electronics and mission-critical infrastructures. Nevertheless, there is an important research gap to date: architecting efficient AuT systems requires both energy autonomy (EA) and inference autonomy (IA). In other words, practical AuT application scenarios necessitate tailored architectures with significantly expanded inference performance and more efficient use of energy.We present CHRYSALIS, a novel automated EA/IA co-design methodology for autonomous things. It aims to guide the transition from a traditional EA-only and IA-only design approach to a truly AuT-oriented architecture design. To fully understand the interrelationship between the EA domain and the IA domain, CHRYSALIS first introduces an architectural modeling framework encompassing every key AuT module involving energy harvesting, intermittent execution, and accelerator control. Based on the holistic system model, we design an intelligent architecture generation tool that can help find the ideal design for targeted AuT scenarios adhering to different SWaP (Size, Weight and Power) constraints. To validate our work, we use CHRYSALIS for fast construction and exploration of efficient AuT design and pre-RTL design in representative AuT scenarios. Extensive evaluation shows that CHRYSALIS outperforms state-of-the-art designs and our proposed technique shows 56.4% better performance on average. We believe that the methodology and tools developed in this paper will foster the development of more performant and practical architectures in the upcoming AuT era. Xiaofeng Hou, Tongqiao Xu, Chao Li 0009, Jiacheng Liu 0001, Yang Hu 0001, Jieru Zhao, Jingwen Leng, Kwang-Ting Cheng, Minyi Guo |
ISCA | 9 |
| 2024 | BOLS: A Bionic Sensor-direct On-chip Learning System with Direct-Feedback-Through-Time for Personalized Wearable Health MonitoringabstractPrecise bio-signal classification techniques for edge healthcare have been extensively researched, yet the scalability and efficiency of existing studies remain constrained by challenges in sensing, learning, and processing. Additionally, a deficiency in cross-level integration for the development of comprehensive healthcare systems has been observed. To tackle these issues and facilitate ultra-efficient personalized edge healthcare, this paper introduces the pioneering bionic sensor-direct on-chip learning and inference system with direct-feedback-through-time for user-specific cardiac arrhythmia detection, termed BOLS. This innovative system encompasses a compact sensor-direct feature extractor and a pipelined bionic processor, enabling end-to-end on-chip learning and inference. Employing cross-level co-design, our proposed bionic on-chip learning approach attains exceptional classification performance, boasting an accuracy of 98.6%, which ranks among the highest. The entire system has been implemented using 40nm CMOS process and subsequently verified. Remarkably, the proposed BOLS system consumes a mere 1.18mW for inference and 2.57mW for learning, resulting in an impressive power saving of over ×2000 compared to existing commercial training platforms. Fengshi Tian, Jiakun Zheng, Jingyu He, Jinbo Chen 0002, Chaoming Fang, Jie Yang 0033, Mohamad Sawan, Chi-Ying Tsui, Kwang-Ting Cheng |
ISCAS | 10 |
| 2024 | Rethinking Autoencoders for Medical Anomaly Detection from A Theoretical Perspective
Yu Cai 0005, Hao Chen 0011, Kwang-Ting Cheng |
MICCAI (11) | 3 |
| 2024 | Aligning Medical Images with General Knowledge from Large Language Models
Yi Lin 0009, Kwang-Ting Cheng, Hao Chen 0011 |
MICCAI (10) | 4 |
| 2024 | Revisiting Deep Ensemble Uncertainty for Enhanced Medical Anomaly Detection
Yi Lin 0009, Kwang-Ting Cheng, Hao Chen 0011 |
MICCAI (6) | 3 |
| 2024 | Iterative Online Image Synthesis via Diffusion Model for Imbalanced Classification
Shuhan Li, Yi Lin 0009, Hao Chen 0011, Kwang-Ting Cheng |
MICCAI (5) | 4 |
| 2024 | FedMLP: Federated Multi-label Medical Image Classification Under Task Heterogeneity
Zhaobin Sun, Li Yu 0003, Kwang-Ting Cheng, Zengqiang Yan |
MICCAI (10) | 5 |
| 2024 | FedIA: Federated Medical Image Segmentation with Heterogeneous Annotation Completeness
Yangyang Xiang, Li Yu 0003, Xin Yang 0008, Kwang-Ting Cheng, Zengqiang Yan |
MICCAI (10) | 5 |
| 2024 | Multi-Issue Butterfly Architecture for Sparse Convex Quadratic ProgrammingabstractConvex quadratic optimization solvers are extensively utilized in various domains; however, achieving optimal performance in diverse situations remains a significant challenge due to the sparse nature of objective and constraint matrices. General-purpose architectures struggle with hardware utilization when performing critical sparse matrix operations, such as factorization and multiplication. To address this issue, we introduce a pipelined spatial architecture, Multi-Issue Butterfly (MIB), which supports all primitive scalar, vector, and matrix operations required by the Alternating Direction Method of Multipliers (ADMM) based solver algorithm. The proposed architecture features a butterfly computational network with innovative working modes for each node, controlled by runtime instructions. We developed a companion scheduling method for matrix operations based on their sparsity patterns. For factorization, an elimination tree guides the network instructions reordering to avoid data hazards caused by computation dependencies. For matrix-vector multiplication, data prefetching resolves structural hazards caused by read and write conflicts to register files. Instructions without hazards are issued simultaneously to increase pipeline throughput and function unit utilization. We evaluate the proposed architecture using FPGA prototypes, representing the first fully FPGA-based generic QP solver. Our assessment includes extensive performance and efficiency bench-marks across 100 QP problems from five application domains. Compared to the same algorithm variation running on CPU backends, our prototype achieves a geometric mean of$30.5\times$end-to-end speedup,$127.0 \times$greater energy efficiency, and$16.5\times$less runtime jitter. In comparison to GPU backends, the prototype attains a geometric mean of$4.3\times$faster end-to-end speedup,$21.7\times$higher energy efficiency, and$33.4\times$less runtime jitter. Maolin Wang 0002, Ian McInerney, Bartolomeo Stellato, Fengbin Tu, Stephen P. Boyd, Hayden Kwok-Hay So, Kwang-Ting Cheng |
MICRO | 7 |
| 2024 | CAE-GReaT: Convolutional-Auxiliary Efficient Graph Reasoning Transformer for Dense Image Predictions
Yi Lin 0009, Jinhui Tang 0001, Kwang-Ting Cheng |
Int. J. Comput. Vis. | 4 |
| 2024 | Vessel-promoted OCT to OCTA image translation by heuristic contextual constraints
Shuhan Li, Xiaomeng Li 0001, Chubin Ou, Lin An, Yanwu Xu 0001, Weihua Yang, Yanchun Zhang, Kwang-Ting Cheng |
Medical Image Anal. | 9 |
| 2024 | UCTNet: Uncertainty-guided CNN-Transformer hybrid networks for medical image segmentation
Xiayu Guo, Xian Lin, Xin Yang 0008, Li Yu 0003, Kwang-Ting Cheng, Zengqiang Yan |
Pattern Recognit. | 5 |
| 2024 | DyBit: Dynamic Bit-Precision Numbers for Efficient Quantized Neural Network InferenceabstractTo accelerate the inference of deep neural networks (DNNs), quantization with low-bitwidth numbers is actively researched. A prominent challenge is to quantize the DNN models into low-bitwidth numbers without significant accuracy degradation, especially at very low bitwidths (< 8 bits). This work targets an adaptive data representation with variablelength encoding called DyBit. DyBit can dynamically adjust the precision and range of separate bit-fields to be adapted to the DNN weights/activations distribution. We also propose a hardware-aware quantization framework with a mixed-precision accelerator to trade-off the inference accuracy and speedup. Experimental results demonstrate that the ImageNet inference accuracy via DyBit is 1.97% higher than the state-of-the-art at 4-bit quantization, and the proposed framework can achieve up to 8.1× speedup compared with the original ResNet-50 model. Jiajun Zhou 0004, Jiajun Wu 0006, Yizhao Gao 0002, Yuhao Ding, Chaofan Tao, Fengbin Tu, Kwang-Ting Cheng, Hayden Kwok-Hay So, Ngai Wong 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2024 | LENAS: Learning-Based Neural Architecture Search and Ensemble for 3-D Radiotherapy Dose PredictionabstractRadiation therapy treatment planning requires balancing the delivery of the target dose while sparing normal tissues, making it a complex process. To streamline the planning process and enhance its quality, there is a growing demand for knowledge-based planning (KBP). Ensemble learning has shown impressive power in various deep learning tasks, and it has great potential to improve the performance of KBP. However, the effectiveness of ensemble learning heavily depends on the diversity and individual accuracy of the base learners. Moreover, the complexity of model ensembles is a major concern, as it requires maintaining multiple models during inference, leading to increased computational cost and storage overhead. In this study, we propose a novel learning-based ensemble approach named LENAS, which integrates neural architecture search with knowledge distillation for 3-D radiotherapy dose prediction. Our approach starts by exhaustively searching each block from an enormous architecture space to identify multiple architectures that exhibit promising performance and significant diversity. To mitigate the complexity introduced by the model ensemble, we adopt the teacher-student paradigm, leveraging the diverse outputs from multiple learned networks as supervisory signals to guide the training of the student network. Furthermore, to preserve high-level semantic information, we design a hybrid loss to optimize the student network, enabling it to recover the knowledge embedded within the teacher networks. The proposed method has been evaluated on two public datasets: 1) OpenKBP and 2) AIMIS. Extensive experimental results demonstrate the effectiveness of our method and its superior performance to the state-of-the-art methods. Code: github.com/hust-linyi/LENAS. Yi Lin 0009, Hao Chen 0011, Xin Yang 0008, Kai Ma 0002, Yefeng Zheng 0001, Kwang-Ting Cheng |
IEEE Trans. Cybern. | 7 |
| 2024 | MFTrans: Modality-Masked Fusion Transformer for Incomplete Multi-Modality Brain Tumor SegmentationabstractBrain tumor segmentation is a fundamental task and existing approaches usually rely on multi-modality magnetic resonance imaging (MRI) images for accurate segmentation. However, the common problem of missing/incomplete modalities in clinical practice would severely degrade their segmentation performance, and existing fusion strategies for incomplete multi-modality brain tumor segmentation are far from ideal. In this work, we propose a novel framework named M$^{2}$FTrans to explore and fuse cross-modality features through modality-masked fusion transformers under various incomplete multi-modality settings. Considering vanilla self-attention is sensitive to missing tokens/inputs, both learnable fusion tokens and masked self-attention are introduced to stably build long-range dependency across modalities while being more flexible to learn from incomplete modalities. In addition, to avoid being biased toward certain dominant modalities, modality-specific features are further re-weighted through spatial weight attention and channel-wise fusion transformers for feature redundancy reduction and modality re-balancing. In this way, the fusion strategy in M$^{2}$FTrans is more robust to missing modalities. Experimental results on the widely-used BraTS2018, BraTS2020, and BraTS2021 datasets demonstrate the effectiveness of M$^{2}$FTrans, outperforming the state-of-the-art approaches with large margins under various incomplete modalities for brain tumor segmentation. Li Yu 0003, Qimin Cheng, Xin Yang 0008, Kwang-Ting Cheng, Zengqiang Yan |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | BoNuS: Boundary Mining for Nuclei Segmentation With Partial Point LabelsabstractNuclei segmentation is a fundamental prerequisite in the digital pathology workflow. The development of automated methods for nuclei segmentation enables quantitative analysis of the wide existence and large variances in nuclei morphometry in histopathology images. However, manual annotation of tens of thousands of nuclei is tedious and time-consuming, which requires significant amount of human effort and domain-specific expertise. To alleviate this problem, in this paper, we propose a weakly-supervised nuclei segmentation method that only requires partial point labels of nuclei. Specifically, we propose a novel boundary mining framework for nuclei segmentation, named BoNuS, which simultaneously learns nuclei interior and boundary information from the point labels. To achieve this goal, we propose a novel boundary mining loss, which guides the model to learn the boundary information by exploring the pairwise pixel affinity in a multiple-instance learning manner. Then, we consider a more challenging problem, i.e., partial point label, where we propose a nuclei detection module with curriculum learning to detect the missing nuclei with prior morphological knowledge. The proposed method is validated on three public datasets, MoNuSeg, CPM, and CoNIC datasets. Experimental results demonstrate the superior performance of our method to the state-of-the-art weakly-supervised nuclei segmentation methods. Code: https://github.com/hust-linyi/bonus. Yi Lin 0009, Kwang-Ting Cheng, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 4 |
| 2024 | Exploring Feature Representation Learning for Semi-Supervised Medical Image SegmentationabstractThis article presents a simple yet effective two-stage framework for semi-supervised medical image segmentation. Unlike prior state-of-the-art semi-supervised segmentation methods that predominantly rely on pseudo supervision directly on predictions, such as consistency regularization and pseudo labeling, our key insight is to explore the feature representation learning with labeled and unlabeled (i.e., pseudo labeled) images to regularize a more compact and better-separated feature space, which paves the way for low-density decision boundary learning and therefore enhances the segmentation performance. A stage-adaptive contrastive learning method is proposed, containing a boundary-aware contrastive loss that takes advantage of the labeled images in the first stage, as well as a prototype-aware contrastive loss to optimize both labeled and pseudo labeled images in the second stage. To obtain more accurate prototype estimation, which plays a critical role in prototype-aware contrastive learning, we present an aleatoric uncertainty-aware method to generate higher quality pseudo labels. Aleatoric-uncertainty adaptive (AUA) adaptively regularizes prediction consistency by taking advantage of image ambiguity, which, given its significance, is underexplored by existing works. Our method achieves the best results on three public medical image segmentation benchmarks. Huimin Wu 0001, Xiaomeng Li 0001, Kwang-Ting Cheng |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | WASP: Efficient Power Management Enabling Workload-Aware, Self-Powered AIoT DevicesabstractThe wide adoption of edge AI has heightened the demand for various battery-less and maintenance-free smart systems. Nevertheless, emerging Artificial Intelligence of Things (AIoT) are complex workloads showing increased power demand, diversified power usage patterns, and unique sensitivity to power management (PM) approaches. Existing AIoT devices cannot select the most appropriate PM tuning knob, and therefore they often make sub-optimal decisions. In addition, these PM solutions always assume traditional power regulation circuit which incurs non-negligible power loss and control overhead. This can greatly compromise the potential of AIoT efficiency. In this paper, we explore power management optimization for emerging self-powered AIoT devices. We propose WASP, a highly efficient power management scheme for workload-aware, self-powered AIoT devices. The novelty of WASP is two fold. First, it combines offline profiling and light-weight online control to select the most appropriate PM tuning knobs for the given DNN models. Second, it is well tailored to a reconfigurable voltage regulation module that can make the best use of the limited power budget. Our results show that WASP allows AIoT devices to accomplish 65.6% more inference tasks under a stringent power budget without any performance degradation compared with other existing approaches. Xiaofeng Hou, Xuehan Tang, Jiacheng Liu 0001, Chao Li 0009, Luhong Liang, Kwang-Ting Cheng |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2023 | Semi-Supervised Deep Regression with Uncertainty Consistency and Variational Model Ensembling via Bayesian Neural NetworksabstractDeep regression is an important problem with numerous applications. These range from computer vision tasks such as age estimation from photographs, to medical tasks such as ejection fraction estimation from echocardiograms for disease tracking. Semi-supervised approaches for deep regression are notably under-explored compared to classification and segmentation tasks, however. Unlike classification tasks, which rely on thresholding functions for generating class pseudo-labels, regression tasks use real number target predictions directly as pseudo-labels, making them more sensitive to prediction quality. In this work, we propose a novel approach to semi-supervised regression, namely Uncertainty-Consistent Variational Model Ensembling (UCVME), which improves training by generating high-quality pseudo-labels and uncertainty estimates for heteroscedastic regression. Given that aleatoric uncertainty is only dependent on input data by definition and should be equal for the same inputs, we present a novel uncertainty consistency loss for co-trained models. Our consistency loss significantly improves uncertainty estimates and allows higher quality pseudo-labels to be assigned greater importance under heteroscedastic regression. Furthermore, we introduce a novel variational model ensembling approach to reduce prediction noise and generate more robust pseudo-labels. We analytically show our method generates higher quality targets for unlabeled data and further improves training. Experiments show that our method outperforms state-of-the-art alternatives on different tasks and can be competitive with supervised methods that use full labels. Code is available at https://github.com/xmed-lab/UCVME. Weihang Dai, Xiaomeng Li 0001, Kwang-Ting Cheng |
AAAI | 3 |
| 2023 | RVComp: Analog Variation Compensation for RRAM-Based in-Memory ComputingabstractResistive Random Access Memory (RRAM) has shown great potential in accelerating memory-intensive computation in neural network applications. However, RRAM-based computing suffers from significant accuracy degradation due to the inevitable device variations. In this paper, we propose RVComp, a fine-grained analog Compensation approach to mitigate the accuracy loss of in-memory computing incurred by the Variations of the RRAM devices. Specifically, weights in the RRAM crossbar are accompanied by dedicated compensation RRAM cells to offset their programming errors with a scaling factor. A programming target shifting mechanism is further designed with the objectives of reducing the hardware overhead and minimizing the compensation errors under large device variations. Based on these two key concepts, we propose double and dynamic compensation schemes and the corresponding support architecture. Since the RRAM cells only account for a small fraction of the overall area of the computing macro due to the dominance of the peripheral circuitry, the overall area overhead of RVComp is low and manageable. Simulation results show RVComp achieves a negligible 1.80% inference accuracy drop for ResNet18 on the CIFAR-10 dataset under 30% device variation with only 7.12% area and 5.02% power overhead and no extra latency. Jingyu He, Yucong Huang, Miguel Angel Lastras-Montaño, Terry Tao Ye, Chi-Ying Tsui, Kwang-Ting Cheng |
ASP-DAC | 6 |
| 2023 | AutoDCIM: An Automated Digital CIM CompilerabstractDigital Computing-in-Memory (DCIM) is an emerging architecture that integrates digital logic into memory for efficient AI computing. However, current DCIM designs heavily rely on manual efforts. This increases DCIM design time and limits the optimization space, making it challenging to satisfy the user specifications of diverse AI applications. This paper presents AutoDCIM, the first automated DCIM compiler. Au-toDCIM takes the user specifications as inputs and generates a DCIM macro architecture with an optimized layout. AutoDCIM’s template-based generation balances handcrafted cell design and agile macro development. AutoDCIM’s layout exploration loop analyzes diverse DCIM array partitioning schemes to satisfy user specifications. The auto-generated DCIM macros present competitive efficiency results in comparison with state-of-the-art silicon-verified DCIM macros. Jia Chen 0032, Fengbin Tu, Kunming Shao, Fengshi Tian, Xiao Huo, Chi-Ying Tsui, Kwang-Ting Cheng |
DAC | 7 |
| 2023 | PIM-HLS: An Automatic Hardware Generation Tool for Heterogeneous Processing-In-Memory-based Neural Network AcceleratorsabstractProcessing-in-memory (PIM) architectures have shown great abilities for neural network (NN) acceleration on edge devices that demand low latency under severe area constraints. Heterogeneous PIM architectures with different PIM implementation approaches such as RRAM-based PIM and SRAM-based PIM can further improve the performance. However, the automatic generation of heterogeneous PIM architectures faces the following two unresolved problems. First, existing work has not considered the design for heterogeneous PIM-based NN accelerators with multiple memory technologies. Second, for PIM with insufficient memory on edge devices, it is challenging to find the optimal runtime weight scheduling strategy in an O(L!) optimization space for the NN with L layers.In this paper, we propose PIM-HLS, an automatic hardware generation tool for heterogeneous PIM-based NN accelerators. Aiming at the problems above, we first point out that heterogeneous PIM can improve the performance under severe area constraints. Then we optimize the architectures for each NN layer by taking the advantage of different memory technologies. We also define the optimization problem of runtime weight scheduling and mapping for the first time, and propose a dynamic-programming-based weight scheduling algorithm to reduce the optimization space to O(L2). We implement PIM-HLS to automatically generate the hardware code and the instructions. Results show that we achieve an averagely 5.9× speedup with 72.8% less area compared with state-of-the-art PIM designs. Zhenhua Zhu 0002, Guohao Dai 0001, Fengbin Tu, Hanbo Sun, Kwang-Ting Cheng, Huazhong Yang, Yu Wang 0002 |
DAC | 6 |
| 2023 | FoodWise: Food Waste Reduction and Behavior Change on Campus with Data Visualization and GamificationabstractFood waste presents a substantial challenge with significant environmental and economic ramifications, and its severity on campus environments is of particular concern. In response to this, we introduce FoodWise, a dual-component system tailored to inspire and incentivize campus communities to reduce food waste. The system consists of a data storytelling dashboard that graphically displays food waste information from university canteens, coupled with a mobile web application that encourages users to log their food waste reduction actions and rewards active participants for their efforts. Sophia Yi, Leo Yu-Ho Lo, Kento Shigyo, Liwenhan Xie, Jeffry Wicaksana, Kwang-Ting Cheng, Huamin Qu |
COMPASS | 8 |
| 2023 | LLM-FP4: 4-Bit Floating-Point Quantized TransformersabstractWe propose LLM-FP4 for quantizing both weights and activations in large language models (LLMs) down to 4-bit floating-point values, in a post-training manner.Existing posttraining quantization (PTQ) solutions are primarily integer-based and struggle with bit widths below 8 bits.Compared to integer quantization, floating-point (FP) quantization is more flexible and can better handle long-tail or bell-shaped distributions, and it has emerged as a default choice in many hardware platforms.One characteristic of FP quantization is that its performance largely depends on the choice of exponent bits and clipping range.In this regard, we construct a strong FP-PTQ baseline by searching for the optimal quantization parameters.Furthermore, we observe a high interchannel variance and low intra-channel variance pattern in activation distributions, which adds activation quantization difficulty.We recognize this pattern to be consistent across a spectrum of transformer models designed for diverse tasks, such as LLMs, BERT, and Vision Transformer models.To tackle this, we propose per-channel activation quantization and show that these additional scaling factors can be reparameterized as exponential biases of weights, incurring a negligible cost.Our method, for the first time, can quantize both weights and activations in the LLaMA-13B to only 4-bit and achieves an average score of 63.1 on the common sense zero-shot reasoning tasks, which is only 5.8 lower than the full-precision model, significantly outperforming the previous stateof-the-art by 12.7 points.Code is available at: https://github.com/nbasyl/LLM-FP4. Shih-Yang Liu, Zechun Liu, Xijie Huang, Pingcheng Dong, Kwang-Ting Cheng |
EMNLP | 5 |
| 2023 | MMExit: Enabling Fast and Efficient Multi-modal DNN Inference with Adaptive Network Exits
Xiaofeng Hou, Jiacheng Liu 0001, Xuehan Tang, Chao Li 0009, Kwang-Ting Cheng, Li Li 0012, Minyi Guo |
Euro-Par | 5 |
| 2023 | Randomized Quantization: A Generic Augmentation for Data Agnostic Self-supervised LearningabstractSelf-supervised representation learning follows a paradigm of withholding some part of the data and tasking the network to predict it from the remaining part. Among many techniques, data augmentation lies at the core for creating the information gap. Towards this end, masking has emerged as a generic and powerful tool where content is withheld along the sequential dimension, e.g., spatial in images, temporal in audio, and syntactic in language. In this paper, we explore the orthogonal channel dimension for generic data augmentation by exploiting precision redundancy. The data for each channel is quantized through a non-uniform quantizer, with the quantized value sampled randomly within randomly sampled quantization bins. From another perspective, quantization is analogous to channel-wise masking, as it removes the information within each bin, but preserves the information across bins. Our approach significantly surpasses existing generic data augmentation methods, while showing on par performance against modality-specific augmentations. We comprehensively evaluate our approach on vision, audio, 3D point clouds, as well as the DABS benchmark which is comprised of various data modalities. The code is available at https://github.com/microsoft/random_quantize. Huimin Wu 0001, Chenyang Lei, Xiao Sun 0001, Peng-Shuai Wang, Qifeng Chen 0001, Kwang-Ting Cheng, Stephen Lin 0001, Zhirong Wu |
ICCV | 6 |
| 2023 | Oscillation-free Quantization for Low-bit Vision TransformersabstractWeight oscillation is a by-product of quantization-aware training, in which quantized weights frequently jump between two quantized levels, resulting in training instability and a sub-optimal final model. We discover that the learnable scaling factor, a widely-used $\textit{de facto}$ setting in quantization aggravates weight oscillation. In this work, we investigate the connection between learnable scaling factor and quantized weight oscillation using ViT, and we additionally find that the interdependence between quantized weights in $\textit{query}$ and $\textit{key}$ of a self-attention layer also makes ViT vulnerable to oscillation. We propose three techniques correspondingly: statistical weight quantization ($\rm StatsQ$) to improve quantization robustness compared to the prevalent learnable-scale-based method; confidence-guided annealing ($\rm CGA$) that freezes the weights with $\textit{high confidence}$ and calms the oscillating weights; and $\textit{query}$-$\textit{key}$ reparameterization ($\rm QKR$) to resolve the query-key intertwined oscillation and mitigate the resulting gradient misestimation. Extensive experiments demonstrate that our algorithms successfully abate weight oscillation and consistently achieve substantial accuracy improvement on ImageNet. Specifically, our 2-bit DeiT-T/DeiT-S surpass the previous state-of-the-art by 9.8% and 7.7%, respectively. The code is included in the supplementary material and will be released. Shih-Yang Liu, Zechun Liu, Kwang-Ting Cheng |
ICML | 3 |
| 2023 | FedNoRo: Towards Noise-Robust Federated Learning by Addressing Class Imbalance and Label Noise HeterogeneityabstractFederated noisy label learning (FNLL) is emerging as a promising tool for privacy-preserving multi-source decentralized learning. Existing research, relying on the assumption of class-balanced global data, might be incapable to model complicated label noise, especially in medical scenarios. In this paper, we first formulate a new and more realistic federated label noise problem where global data is class-imbalanced and label noise is heterogeneous, and then propose a two-stage framework named FedNoRo for noise-robust federated learning. Specifically, in the first stage of FedNoRo, per-class loss indicators followed by Gaussian Mixture Model are deployed for noisy client identification. In the second stage, knowledge distillation and a distance-aware aggregation function are jointly adopted for noise-robust federated model updating. Experimental results on the widely-used ICH and ISIC2019 datasets demonstrate the superiority of FedNoRo against the state-of-the-art FNLL methods for addressing class imbalance and label noise heterogeneity in real-world FL scenarios. Li Yu 0003, Xuefeng Jiang 0001, Kwang-Ting Cheng, Zengqiang Yan |
IJCAI | 4 |
| 2023 | Architecting Efficient Multi-modal AIoT SystemsabstractMulti-modal computing (M2C) has recently exhibited impressive accuracy improvements in numerous autonomous artificial intelligence of things (AIoT) systems. However, this accuracy gain is often tethered to an incredible increase in energy consumption. Particularly, various highly-developed modality sensors devour most of the energy budget, which would make the deployment of M2C for real-world AIoT applications a difficult challenge. Xiaofeng Hou, Jiacheng Liu 0001, Xuehan Tang, Chao Li 0009, Jia Chen 0032, Luhong Liang, Kwang-Ting Cheng, Minyi Guo |
ISCA | 7 |
| 2023 | Radiomics-Informed Deep Learning for Classification of Atrial Fibrillation Sub-Types from Left-Atrium CT Volumes
Weihang Dai, Xiaomeng Li 0001, Taihui Yu, Jun Shen 0008, Kwang-Ting Cheng |
MICCAI (7) | 6 |
| 2023 | Few Shot Medical Image Segmentation with Cross Attention Transformer
Yi Lin 0009, Kwang-Ting Cheng, Hao Chen 0011 |
MICCAI (2) | 3 |
| 2023 | FedIIC: Towards Robust Federated Learning for Class-Imbalanced Medical Image Classification
Li Yu 0003, Xin Yang 0008, Kwang-Ting Cheng, Zengqiang Yan |
MICCAI (2) | 4 |
| 2023 | Semi-Supervised Contrastive Learning for Deep Regression with Ordinal Rankings from Spectral SeriationabstractContrastive learning methods can be applied to deep regression by enforcing label distance relationships in feature space. However, these methods are limited to labeled data only unlike for classification, where unlabeled data can be used for contrastive pretraining. In this work, we extend contrastive regression methods to allow unlabeled data to be used in a semi-supervised setting, thereby reducing the reliance on manual annotations. We observe that the feature similarity matrix between unlabeled samples still reflect inter-sample relationships, and that an accurate ordinal relationship can be recovered through spectral seriation algorithms if the level of error is within certain bounds. By using the recovered ordinal relationship for contrastive learning on unlabeled samples, we can allow more data to be used for feature representation learning, thereby achieve more robust results. The ordinal rankings can also be used to supervise predictions on unlabeled samples, which can serve as an additional training signal. We provide theoretical guarantees and empirical support through experiments on different datasets, demonstrating that our method can surpass existing state-of-the-art semi-supervised deep regression methods. To the best of our knowledge, this work is the first to explore using unlabeled data to perform contrastive learning for regression. Weihang Dai, Hanru Bai, Kwang-Ting Cheng, Xiaomeng Li 0001 |
NeurIPS | 4 |
| 2023 | SMG: A System-Level Modality Gating Facility for Fast and Energy-Efficient Multimodal ComputingabstractAchieving low-latency and high-efficiency multimodal computing (MMC) is crucial for deploying high-performance autonomous embedded systems (AES) that has limited energy budgets. However, existing methods have mainly focused on optimizing the computing phase and have overlooked the significant energy and latency overhead during the sensing phase. Therefore, we propose SMG, a system-level modality gating facility to optimize this. Our approach introduces a software-defined DSP gating technique that enables MMC tasks to bypass both the sensing and computing phases of unimportant modalities. We also propose a raw data-activated MMC mechanism that comprises a fast modality tester and adaptive modality executor, which adapts to the modality gating architecture and performs energy-efficient MMC. To evaluate SMG, we implement a prototype of SMG by integrating it into existing AES and analyze it with extensive multimodal video recognition workloads. Our experimental results show that SMG outperforms SOTA approaches by adaptively gating some DSP operations, resulting in substantial improvements in both energy consumption and task latency. Xiaofeng Hou, Chao Li 0009, Jiacheng Liu 0001, Kwang-Ting Cheng, Minyi Guo |
RTSS | 6 |
| 2023 | Dual-distribution discrepancy with self-supervised refinement for anomaly detection in medical images
Yu Cai 0005, Hao Chen 0011, Xin Yang 0008, Yu Zhou 0016, Kwang-Ting Cheng |
Medical Image Anal. | 5 |
| 2023 | Nuclei segmentation with point annotations from pathology images via self-supervised learning and co-training
Yi Lin 0009, Zhiyong Qu, Hao Chen 0011, Zhongke Gao, Yuexiang Li, Kai Ma 0002, Yefeng Zheng 0001, Kwang-Ting Cheng |
Medical Image Anal. | 9 |
| 2023 | BATFormer: Towards Boundary-Aware Lightweight Transformer for Efficient Medical Image SegmentationabstractOBJECTIVE: Transformers, born to remedy the inadequate receptive fields of CNNs, have drawn explosive attention recently. However, the daunting computational complexity of global representation learning, together with rigid window partitioning, hinders their deployment in medical image segmentation. This work aims to address the above two issues in transformers for better medical image segmentation. METHODS: We propose a boundary-aware lightweight transformer (BATFormer) that can build cross-scale global interaction with lower computational complexity and generate windows flexibly under the guidance of entropy. Specifically, to fully explore the benefits of transformers in long-range dependency establishment, a cross-scale global transformer (CGT) module is introduced to jointly utilize multiple small-scale feature maps for richer global features with lower computational complexity. Given the importance of shape modeling in medical image segmentation, a boundary-aware local transformer (BLT) module is constructed. Different from rigid window partitioning in vanilla transformers which would produce boundary distortion, BLT adopts an adaptive window partitioning scheme under the guidance of entropy for both computational complexity reduction and shape preservation. RESULTS: BATFormer achieves the best performance in Dice of 92.84 %, 91.97 %, 90.26 %, and 96.30 % for the average, right ventricle, myocardium, and left ventricle respectively on the ACDC dataset and the best performance in Dice, IoU, and ACC of 90.76 %, 84.64 %, and 96.76 % respectively on the ISIC 2018 dataset. More importantly, BATFormer requires the least amount of model parameters and the lowest computational complexity compared to the state-of-the-art approaches. CONCLUSION AND SIGNIFICANCE: Our results demonstrate the necessity of developing customized transformers for efficient and better medical image segmentation. We believe the design of BATFormer is inspiring and extendable to other applications/frameworks. Xian Lin, Li Yu 0003, Kwang-Ting Cheng, Zengqiang Yan |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | Cyclical Self-Supervision for Semi-Supervised Ejection Fraction Prediction From Echocardiogram VideosabstractLeft-ventricular ejection fraction (LVEF) is an important indicator of heart failure. Existing methods for LVEF estimation from video require large amounts of annotated data to achieve high performance, e.g. using 10,030 labeled echocardiogram videos to achieve mean absolute error (MAE) of 4.10. Labeling these videos is time-consuming however and limits potential downstream applications to other heart diseases. This paper presents the first semi-supervised approach for LVEF prediction. Unlike general video prediction tasks, LVEF prediction is specifically related to changes in the left ventricle (LV) in echocardiogram videos. By incorporating knowledge learned from predicting LV segmentations into LVEF regression, we can provide additional context to the model for better predictions. To this end, we propose a novel Cyclical Self-Supervision (CSS) method for learning video-based LV segmentation, which is motivated by the observation that the heartbeat is a cyclical process with temporal repetition. Prediction masks from our segmentation model can then be used as additional input for LVEF regression to provide spatial context for the LV region. We also introduce teacher-student distillation to distill the information from LV segmentation masks into an end-to-end LVEF regression model that only requires video inputs. Results show our method outperforms alternative semi-supervised methods and can achieve MAE of 4.17, which is competitive with state-of-the-art supervised performance, using half the number of labels. Validation on an external dataset also shows improved generalization ability from using our method. Weihang Dai, Xiaomeng Li 0001, Xinpeng Ding, Kwang-Ting Cheng |
IEEE Trans. Medical Imaging | 4 |
| 2023 | The Lighter the Better: Rethinking Transformers in Medical Image Segmentation Through Adaptive PruningabstractVision transformers have recently set off a new wave in the field of medical image analysis due to their remarkable performance on various computer vision tasks. However, recent hybrid-/transformer-based approaches mainly focus on the benefits of transformers in capturing long-range dependency while ignoring the issues of their daunting computational complexity, high training costs, and redundant dependency. In this paper, we propose to employ adaptive pruning to transformers for medical image segmentation and propose a lightweight and effective hybrid network APFormer. To our best knowledge, this is the first work on transformer pruning for medical image analysis tasks. The key features of APFormer are self-regularized self-attention (SSA) to improve the convergence of dependency establishment, Gaussian-prior relative position embedding (GRPE) to foster the learning of position information, and adaptive pruning to eliminate redundant computations and perception information. Specifically, SSA and GRPE consider the well-converged dependency distribution and the Gaussian heatmap distribution separately as the prior knowledge of self-attention and position embedding to ease the training of transformers and lay a solid foundation for the following pruning operation. Then, adaptive transformer pruning, both query-wise and dependency-wise, is performed by adjusting the gate control parameters for both complexity reduction and performance improvement. Extensive experiments on two widely-used datasets demonstrate the prominent segmentation performance of APFormer against the state-of-the-art methods with much fewer parameters and lower GFLOPs. More importantly, we prove, through ablation studies, that adaptive pruning can work as a plug-n-play module for performance improvement on other hybrid-/transformer-based methods. Code is available at https://github.com/xianlin7/APFormer. Xian Lin, Li Yu 0003, Kwang-Ting Cheng, Zengqiang Yan |
IEEE Trans. Medical Imaging | 3 |
| 2023 | FedMix: Mixed Supervised Federated Learning for Medical Image SegmentationabstractThe purpose of federated learning is to enable multiple clients to jointly train a machine learning model without sharing data. However, the existing methods for training an image segmentation model have been based on an unrealistic assumption that the training set for each local client is annotated in a similar fashion and thus follows the same image supervision level. To relax this assumption, in this work, we propose a label-agnostic unified federated learning framework, named FedMix, for medical image segmentation based on mixed image labels. In FedMix, each client updates the federated model by integrating and effectively making use of all available labeled data ranging from strong pixel-level labels, weak bounding box labels, to weakest image-level class labels. Based on these local models, we further propose an adaptive weight assignment procedure across local clients, where each client learns an aggregation weight during the global model update. Compared to the existing methods, FedMix not only breaks through the constraint of a single level of image supervision but also can dynamically adjust the aggregation weight of each local client, achieving rich yet discriminative feature representations. Experimental results on multiple publicly-available datasets validate that the proposed FedMix outperforms the state-of-the-art methods by a large margin. In addition, we demonstrate through experiments that FedMix is extendable to multi-class medical image segmentation and much more feasible in clinical scenarios. The code is available at: https://github.com/Jwicaksana/FedMix. Jeffry Wicaksana, Zengqiang Yan, Xijie Huang, Huimin Wu 0001, Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Medical Imaging | 7 |
| 2023 | Compete to Win: Enhancing Pseudo Labels for Barely-Supervised Medical Image SegmentationabstractThis study investigates barely-supervised medical image segmentation where only few labeled data, i.e., single-digit cases are available. We observe the key limitation of the existing state-of-the-art semi-supervised solution cross pseudo supervision is the unsatisfactory precision of foreground classes, leading to a degenerated result under barely-supervised learning. In this paper, we propose a novel Compete-to-Win method (ComWin) to enhance the pseudo label quality. In contrast to directly using one model’s predictions as pseudo labels, our key idea is that high-quality pseudo labels should be generated by comparing multiple confidence maps produced by different networks to select the most confident one (a compete-to-win strategy). To further refine pseudo labels at near-boundary areas, an enhanced version of ComWin, namely, ComWin$^{+}$, is proposed by integrating a boundary-aware enhancement module. Experiments show that our method can achieve the best performance on three public medical image datasets for cardiac structure segmentation, pancreas segmentation and colon tumor segmentation, respectively. The source code is now available athttps://github.com/Huiimin5/comwin. Huimin Wu 0001, Xiaomeng Li 0001, Yiqun Lin, Kwang-Ting Cheng |
IEEE Trans. Medical Imaging | 4 |
| 2023 | Path-Analysis-Based Reinforcement Learning Algorithm for Imitation FilmingabstractImitation filming has been applied to autonomous filming by mimicking human operators. To imitate the operation of cameramen when filming multiple human actions, existing methods plan the camera motion through time series prediction or train multiple models to handle a particular style in a specific situation. As a result, these methods require various settings to adapt to different scenarios. In this work, we overcome such limitations and propose an end-to-end imitation learning framework for drone cinematography systems. The framework consists of two main components: (1) an efficient motion feature extraction module for generating a compact motion feature space, (2) a path-analysis-based reinforcement learning (PABRL) algorithm for imitating multiple filming styles from demonstrations and incorporating aesthetical features for improved perspective shots. Our PABRL method is based on the actor–critic network, which regards multiple human motion variables, camera translations, and image composition as inputs and then outputs an aesthetical filming strategy related to the subject motion. In addition, we propose an attention mechanism and a long–short-term rewarding function to enhance the motion feature space and the integrity of the generated trajectory, respectively. Extensive experimental results in simulated and real outdoor environments demonstrate that compared with state-of-the-art methods, our method can achieve 69.8% higher performance in terms of trajectory planning accuracy while successfully incorporating aesthetical features into the captured videos. Yuanjie Dang, Chong Huang 0005, Peng Chen 0008, Ronghua Liang, Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Multim. | 6 |
| 2022 | Stereo Neural Vernier CaliperabstractWe propose a new object-centric framework for learning-based stereo 3D object detection. Previous studies build scene-centric representations that do not consider the significant variation among outdoor instances and thus lack the flexibility and functionalities that an instance-level model can offer. We build such an instance-level model by formulating and tackling a local update problem, i.e., how to predict a refined update given an initial 3D cuboid guess. We demonstrate how solving this problem can complement scene-centric approaches in (i) building a coarse-to-fine multi-resolution system, (ii) performing model-agnostic object location refinement, and (iii) conducting stereo 3D tracking-by-detection. Extensive experiments demonstrate the effectiveness of our approach, which achieves state-of-the-art performance on the KITTI benchmark. Code and pre-trained models are available at https://github.com/Nicholasli1995/SNVC. Shichao Li 0002, Zechun Liu, Kwang-Ting Cheng |
AAAI | 4 |
| 2022 | Vision Transformer Slimming: Multi-Dimension Searching in Continuous Optimization SpaceabstractThis paper explores the feasibility of finding an optimal sub-model from a vision transformer and introduces a pure vision transformer slimming (ViT-Slim) framework. It can search a sub-structure from the original model end-to-end across multiple dimensions, including the input tokens, MHSA and MLP modules with state-of-the-art performance. Our method is based on a learnable and unified ℓ1sparsity constraint with pre-defined factors to reflect the global importance in the continuous searching space of different dimensions. The searching process is highly efficient through a single-shot training scheme. For instance, on DeiT-S, ViT-Slim only takes ~43 GPU hours for the searching process, and the searched structure is flexible with diverse dimensionalities in different modules. Then, a budget threshold is employed according to the requirements of accuracy-FLOPs trade-off on running devices, and a retraining process is performed to obtain the final model. The extensive experiments show that our ViT-Slim can compress up to 40% of parameters and 40% FLOPs on various vision transformers while increasing the accuracy by ~0.6% on ImageNet. We also demonstrate the advantage of our searched models on several downstream datasets. Our code is available at https://github.com/Arnav0400/ViT-Slim. Arnav Chavan, Zhuang Liu 0003, Zechun Liu, Kwang-Ting Cheng, Eric P. Xing |
CVPR | 5 |
| 2022 | Nonuniform-to-Uniform Quantization: Towards Accurate Quantization via Generalized Straight-Through EstimationabstractThe nonuniform quantization strategy for compressing neural networks usually achieves better performance than its counterpart, i.e., uniform strategy, due to its superior representational capacity. However, many nonuniform quantization methods overlook the complicated projection process in implementing the nonuniformly quantized weights/activations, which incurs non-negligible time and space overhead in hardware deployment. In this study, we propose Nonuniform-to-Uniform Quantization (N2UQ), a method that can maintain the strong representation ability of nonuniform methods while being hardware-friendly and efficient as the uniform quantization for model inference. We achieve this through learning the flexible inequidistant input thresholds to better fit the underlying distribution while quantizing these real-valued inputs into equidistant output levels. To train the quantized network with learnable input thresholds, we introduce a generalized straight-through estimator (G-STE) for intractable backward derivative calculation w.r.t. threshold parameters. Additionally, we consider entropy preserving regularization to further reduce information loss in weight quantization. Even under this adverse constraint of imposing uniformly quantized weights and activations, our N2UQ outperforms state-of-the-art nonuniform quantization methods by 0.5 ~ 1.7% on ImageNet, demonstrating the contribution of N2UQ design. Code and models are available at: https://github.com/liuzechun/Nonuniform-to-Uniform-Quantization. Zechun Liu, Kwang-Ting Cheng, Dong Huang 0007, Eric P. Xing |
CVPR | 2 |
| 2022 | Data-Free Neural Architecture Search via Recursive Label Calibration
Zechun Liu, Eric P. Xing, Kwang-Ting Cheng, Chas Leichner |
ECCV (24) | 5 |
| 2022 | SDQ: Stochastic Differentiable Quantization with Mixed PrecisionabstractIn order to deploy deep models in a computationally efficient manner, model quantization approaches have been frequently used. In addition, as new hardware that supports various-bit arithmetic operations, recent research on mixed precision quantization (MPQ) begins to fully leverage the capacity of representation by searching various bitwidths for different layers and modules in a network. However, previous studies mainly search the MPQ strategy in a costly scheme using reinforcement learning, neural architecture search, etc., or simply utilize partial prior knowledge for bitwidth distribution, which might be biased and sub-optimal. In this work, we present a novel Stochastic Differentiable Quantization (SDQ) method that can automatically learn the MPQ strategy in a more flexible and globally-optimized space with a smoother gradient approximation. Particularly, Differentiable Bitwidth Parameters (DBPs) are employed as the probability factors in stochastic quantization between adjacent bitwidth. After the optimal MPQ strategy is acquired, we further train our network with the entropy-aware bin regularization and knowledge distillation. We extensively evaluate our method on different networks, hardwares (GPUs and FPGA), and datasets. SDQ outperforms all other state-of-the-art mixed or single precision quantization with less bitwidth, and are even better than the original full-precision counterparts across various ResNet and MobileNet families, demonstrating the effectiveness and superiority of our method. Code will be publicly available. Xijie Huang, Shichao Li 0002, Zechun Liu, Xianghong Hu 0001, Jeffry Wicaksana, Eric P. Xing, Kwang-Ting Cheng |
ICML | 8 |
| 2022 | Dual-Distribution Discrepancy for Anomaly Detection in Chest X-Rays
Yu Cai 0005, Hao Chen 0011, Xin Yang 0008, Yu Zhou 0016, Kwang-Ting Cheng |
MICCAI (3) | 5 |
| 2022 | InsMix: Towards Realistic Generative Data Augmentation for Nuclei Instance Segmentation
Yi Lin 0009, Kwang-Ting Cheng, Hao Chen 0011 |
MICCAI (2) | 3 |
| 2022 | Graph Reasoning Transformer for Image ParsingabstractCapturing the long-range dependencies has empirically proven to be effective on a wide range of computer vision tasks. The progressive advances on this topic have been made through the employment of the transformer framework with the help of the multi-head attention mechanism. However, the attention-based image patch interaction potentially suffers from problems of redundant interactions of intra-class patches and unoriented interactions of inter-class patches. In this paper, we propose a novel Graph Reasoning Transformer (GReaT) for image parsing to enable image patches to interact following a relation reasoning pattern. Specifically, the linearly embedded image patches are first projected into the graph space, where each node represents the implicit visual center for a cluster of image patches and each edge reflects the relation weight between two adjacent nodes. After that, global relation reasoning is performed on this graph accordingly. Finally, all nodes including the relation information are mapped back into the original space for subsequent processes. Compared to the conventional transformer, GReaT has higher interaction efficiency and a more purposeful interaction pattern. Experiments are carried out on the challenging Cityscapes and ADE20K datasets. Results show that GReaT achieves consistent performance gains with slight computational overheads on the state-of-the-art transformer baselines. Jinhui Tang 0001, Kwang-Ting Cheng |
ACM Multimedia | 3 |
| 2022 | One-Shot Imitation Drone Filming of Human Motion VideosabstractImitation learning has recently been applied to mimic the operation of a cameraman in existing autonomous camera systems. To imitate a certain demonstration video, existing methods require users to collect a significant number of training videos with a similar filming style. Because the trained model is style-specific, it is challenging to generalize the model to imitate other videos with a different filming style. To address this problem, we propose a framework that we term "one-shot imitation filming", which can imitate a filming style by "seeing" only a single demonstration video of the target style without style-specific model training. This is achieved by two key enabling techniques: 1) filming style feature extraction, which encodes sequential cinematic characteristics of a variable-length video clip into a fixed-length feature vector; and 2) camera motion prediction, which dynamically plans the camera trajectory to reproduce the filming style of the demo video. We implemented the approach with a deep neural network and deployed it on a 6 degrees of freedom (DOF) drone system by first predicting the future camera motions, and then converting them into the drone's control commands via an odometer. Our experimental results on comprehensive datasets and showcases exhibit that the proposed approach achieves significant improvements over conventional baselines, and our approach can mimic the footage of an unseen style with high fidelity. Chong Huang 0005, Yuanjie Dang, Peng Chen 0008, Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | ReAAP: A Reconfigurable and Algorithm-Oriented Array Processor With Compiler-Architecture Co-DesignabstractParallelism and data reuse are the most critical issues for the design of hardware acceleration in a deep learning processor. Besides, abundant on-chip memories and precise data management are intrinsic design requirements because most of deep learning algorithms are data-driven and memory-bound. In this paper, we propose a compiler-architecture co-design scheme targeting a reconfigurable and algorithm-oriented array processor, named ReAAP. Given specific deep neural networks, the proposed co-design scheme is effective to perform parallelism and data reuse optimization on compute-intensive layers for guiding reconfigurable computing in hardware. Especially, the systemic optimization is performed in our proposed domain-specific compiler to deal with the intrinsic tensions between parallelism and data locality, for the purpose of automatically mapping diverse layer-level workloads onto our proposed reconfigurable array architecture. In this architecture, abundant on-chip memories are software-controlled and its massive data access is precisely handled by compiler-generated instructions. In our experiments, the ReAAP is implemented on an embedded FPGA platform. Experimental results demonstrate that our proposed co-design scheme is effective to integrate software flexibility with hardware parallelism for accelerating diverse deep learning workloads. As a whole system, ReAAP achieves a consistently high utilization of hardware resource for accelerating all the diverse compute-intensive layers in ResNet, MobileNet, and BERT. Jianwei Zheng 0002, Yu Liu 0007, Luhong Liang, Deming Chen, Kwang-Ting Cheng |
IEEE Trans. Computers | 6 |
| 2022 | HyCA: A Hybrid Computing Architecture for Fault-Tolerant Deep LearningabstractHardware faults on the regular 2-D computing array of a typical deep learning accelerator (DLA) can lead to dramatic prediction accuracy loss. Prior redundancy design approaches typically have each homogeneous redundant processing element (PE) to mitigate faulty PEs for a limited region of the 2-D computing array rather than the entire computing array to avoid the excessive hardware overhead. However, they fail to recover the computing array when the number of faulty PEs in any region exceeds the number of redundant PEs in the same region. The mismatch problem deteriorates when the fault injection rate rises and the faults are unevenly distributed. To address the problem, we propose a hybrid computing architecture (HyCA) for fault-tolerant DLAs. It has a set of dot-production processing units (DPPUs) to recompute all the operations that are mapped to the faulty PEs despite the faulty PE locations. According to our experiments, HyCA shows significantly higher reliability, scalability, and performance with less chip area penalty when compared to the conventional redundancy approaches. Moreover, by taking advantage of the flexible recomputing, HyCA can also be utilized to scan the entire 2-D computing array and detect the faulty PEs effectively at runtime. Cheng Liu 0008, Cheng Chu, Dawen Xu 0002, Ying Wang 0001, Huawei Li 0001, Xiaowei Li 0001, Kwang-Ting Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2022 | Customized Federated Learning for Multi-Source Decentralized Medical Image ClassificationabstractThe performance of deep networks for medical image analysis is often constrained by limited medical data, which is privacy-sensitive. Federated learning (FL) alleviates the constraint by allowing different institutions to collaboratively train a federated model without sharing data. However, the federated model is often suboptimal with respect to the characteristics of each client's local data. Instead of training a single global model, we propose Customized FL (CusFL), for which each client iteratively trains a client-specific/private model based on a federated global model aggregated from all private models trained in the immediate previous iteration. Two overarching strategies employed by CusFL lead to its superior performance: 1) the federated model is mainly for feature alignment and thus only consists of feature extraction layers; 2) the federated feature extractor is used to guide the training of each private model. In that way, CusFL allows each client to selectively learn useful knowledge from the federated model to improve its personalized model. We evaluated CusFL on multi-source medical image datasets for the identification of clinically significant prostate cancer and the classification of skin lesions. Jeffry Wicaksana, Zengqiang Yan, Xin Yang 0008, Yang Liu 0165, Lixin Fan, Kwang-Ting Cheng |
IEEE J. Biomed. Health Informatics | 6 |
| 2022 | Adaptive Contrast for Image Regression in Computer-Aided Disease AssessmentabstractImage regression tasks for medical applications, such as bone mineral density (BMD) estimation and left-ventricular ejection fraction (LVEF) prediction, play an important role in computer-aided disease assessment. Most deep regression methods train the neural network with a single regression loss function like MSE or L1 loss. In this paper, we propose the first contrastive learning framework for deep image regression, namely AdaCon, which consists of a feature learning branch via a novel adaptive-margin contrastive loss and a regression prediction branch. Our method incorporates label distance relationships as part of the learned feature representations, which allows for better performance in downstream regression tasks. Moreover, it can be used as a plug-and-play module to improve performance of existing regression methods. We demonstrate the effectiveness of AdaCon on two medical image regression tasks, i.e., bone mineral density estimation from X-ray images and left-ventricular ejection fraction prediction from echocardiogram videos. AdaCon leads to relative improvements of 3.3% and 5.9% in MAE over state-of-the-art BMD estimation and LVEF prediction methods, respectively. Weihang Dai, Xiaomeng Li 0001, Wan Hang Keith Chiu, Michael David Kuo, Kwang-Ting Cheng |
IEEE Trans. Medical Imaging | 5 |
| 2021 | Partial Is Better Than All: Revisiting Fine-tuning Strategy for Few-shot LearningabstractThe goal of few-shot learning is to learn a classifier that can recognize unseen classes from limited support data with labels. A common practice for this task is to train a model on the base set first and then transfer to novel classes through fine-tuning or meta-learning. However, as the base classes have no overlap to the novel set, simply transferring whole knowledge from base data is not an optimal solution since some knowledge in the base model may be biased or even harmful to the novel class. In this paper, we propose to transfer partial knowledge by freezing or fine-tuning particular layer(s) in the base model. Specifically, layers will be imposed different learning rates if they are chosen to be fine-tuned, to control the extent of preserved transferability. To determine which layers to be recast and what values of learning rates for them, we introduce an evolutionary search based method that is efficient to simultaneously locate the target layers and determine their individual learning rates. We conduct extensive experiments on CUB and mini-ImageNet to demonstrate the effectiveness of our proposed method. It achieves the state-of-the-art performance on both meta-learning and non-meta based frameworks. Furthermore, we extend our method to the conventional pre-training + fine-tuning paradigm and obtain consistent improvement. Zechun Liu, Jie Qin 0004, Marios Savvides, Kwang-Ting Cheng |
AAAI | 5 |
| 2021 | Exploring intermediate representation for monocular vehicle pose estimationabstractWe present a new learning-based framework to recover vehicle pose in SO(3) from a single RGB image. In contrast to previous works that map local appearance to observation angles, we explore a progressive approach by extracting meaningful Intermediate Geometrical Representations (IGRs) to estimate egocentric vehicle orientation. This approach features a deep model that transforms perceived intensities to IGRs, which are mapped to a 3D representation encoding object orientation in the camera coordinate system. Core problems are what IGRs to use and how to learn them more effectively. We answer the former question by designing IGRs based on an interpolated cuboid that derives from primitive 3D annotation readily. The latter question motivates us to incorporate geometry knowledge with a new loss function based on a projective invariant. This loss function allows unlabeled data to be used in the training stage to improve representation learning. Without additional labels, our system outperforms previous monocular RGB-based methods for joint vehicle detection and pose estimation on the KITTI benchmark, achieving performance even comparable to stereo methods. Code and pre-trained models are available at this HTTPS URL1. Shichao Li 0002, Zengqiang Yan, Kwang-Ting Cheng |
CVPR | 4 |
| 2021 | S2-BNN: Bridging the Gap Between Self-Supervised Real and 1-Bit Neural Networks via Guided Distribution CalibrationabstractPrevious studies dominantly target at self-supervised learning on real-valued networks and have achieved many promising results. However, on the more challenging binary neural networks (BNNs), this task has not yet been fully explored in the community. In this paper, we focus on this more difficult scenario: learning networks where both weights and activations are binary, meanwhile, without any human annotated labels. We observe that the commonly used contrastive objective is not satisfying on BNNs for competitive accuracy, since the backbone network contains relatively limited capacity and representation ability. Hence instead of directly applying existing self-supervised methods, which cause a severe decline in performance, we present a novel guided learning paradigm from real-valued to distill binary networks on the final prediction distribution, to minimize the loss and obtain desirable accuracy. Our proposed method can boost the simple contrastive learning baseline by an absolute gain of 5.5∼15% on BNNs. We further reveal that it is difficult for BNNs to recover the similar predictive distributions as real-valued models when training without labels. Thus, how to calibrate them is key to address the degradation in performance. Extensive experiments are conducted on the large-scale ImageNet and downstream datasets. Our method achieves substantial improvement over the simple contrastive learning baseline, and is even comparable to many mainstream supervised BNN methods. Code is available at https://github.com/szq0214/S2-BNN. Zechun Liu, Jie Qin 0004, Lei Huang 0015, Kwang-Ting Cheng, Marios Savvides |
CVPR | 5 |
| 2021 | Traffic-Adaptive Power Reconfiguration for Energy-Efficient and Energy-Proportional Optical InterconnectsabstractSilicon microring-based optical interconnects offer great potential for high-bandwidth data communication in future datacenters and high-performance computing systems. However, a lack of effective runtime power management strategies for optical links, especially during idle or low-utilization periods, is devastating to the energy efficiency and the energy proportionality of the network. In this study, we propose Polestar, i.e., POwer LEvel Scaling with Traffic-Adaptive Reconfiguration, for microring-based optical interconnects. Polestar offers a collection of runtime reconfiguration strategies that target the power states of the lasers and the microring tuning circuitry. The reconfiguration mechanism of the power states is traffic-adaptive for exploiting the trade-off between energy saving and application execution time. The evaluation of Polestar with production datacenter traces demonstrates up to 87 % reduction in pJ/b consumption and significant improvements in energy proportionality metrics, notably outperforming existing strategies. Yuyang Wang 0003, Kwang-Ting Cheng |
ICCAD | 2 |
| 2021 | Is Label Smoothing Truly Incompatible with Knowledge Distillation: An Empirical Study
Zechun Liu, Dejia Xu, Zitian Chen, Kwang-Ting Cheng, Marios Savvides |
ICLR | 5 |
| 2021 | How Do Adam and Training Strategies Help BNNs OptimizationabstractThe best performing Binary Neural Networks (BNNs) are usually attained using Adam optimization and its multi-step training variants. However, to the best of our knowledge, few studies explore the fundamental reasons why Adam is superior to other optimizers like SGD for BNN optimization or provide analytical explanations that support specific training strategies. To address this, in this paper we first investigate the trajectories of gradients and weights in BNNs during the training process. We show the regularization effect of second-order momentum in Adam is crucial to revitalize the weights that are dead due to the activation saturation in BNNs. We find that Adam, through its adaptive learning rate strategy, is better equipped to handle the rugged loss surface of BNNs and reaches a better optimum with higher generalization ability. Furthermore, we inspect the intriguing role of the real-valued weights in binary networks, and reveal the effect of weight decay on the stability and sluggishness of BNN optimization. Through extensive experiments and analysis, we derive a simple training scheme, building on existing Adam-based optimization, which achieves 70.5% top-1 accuracy on the ImageNet dataset using the same architecture as the state-of-the-art ReActNet while achieving 1.1% higher accuracy. Code and models are available at https://github.com/liuzechun/AdamBNN. Zechun Liu, Shichao Li 0002, Koen Helwegen, Dong Huang 0007, Kwang-Ting Cheng |
ICML | 6 |
| 2021 | Towards Robust Dual-View Transformation via Densifying Sparse Supervision for Mammography Lesion Matching
Junlin Xian, Zhiwei Wang 0002, Kwang-Ting Cheng, Xin Yang 0008 |
MICCAI (5) | 3 |
| 2021 | Joint Multi-Dimension Pruning via Numerical Gradient UpdateabstractWe present joint multi-dimension pruning (abbreviated as JointPruning), an effective method of pruning a network on three crucial aspects: spatial, depth and channel simultaneously. To tackle these three naturally different dimensions, we proposed a general framework by defining pruning as seeking the best pruning vector (i.e., the numerical value of layer-wise channel number, spatial size, depth) and construct a unique mapping from the pruning vector to the pruned network structures. Then we optimize the pruning vector with gradient update and model joint pruning as a numerical gradient optimization process. To overcome the challenge that there is no explicit function between the loss and the pruning vectors, we proposed self-adapted stochastic gradient estimation to construct a gradient path through network loss to pruning vectors and enable efficient gradient update. We show that the joint strategy discovers a better status than previous studies that focused on individual dimensions solely, as our method is optimized collaboratively across the three dimensions in a single end-to-end training and it is more efficient than the previous exhaustive methods. Extensive experiments on large-scale ImageNet dataset across a variety of network architectures MobileNet V1&V2&V3 and ResNet demonstrate the effectiveness of our proposed method. For instance, we achieve significant margins of 2.5% and 2.6% improvement over the state-of-the-art approach on the already compact MobileNet V1&V2 under an extremely large compression ratio. Zechun Liu, Xiangyu Zhang 0005, Kwang-Ting Cheng, Jian Sun 0001 |
IEEE Trans. Image Process. | 5 |
| 2021 | Variation-Aware Federated Learning With Multi-Source Decentralized Medical Image DataabstractPrivacy concerns make it infeasible to construct a large medical image dataset by fusing small ones from different sources/institutions. Therefore, federated learning (FL) becomes a promising technique to learn from multi-source decentralized data with privacy preservation. However, the cross-client variation problem in medical image data would be the bottleneck in practice. In this paper, we propose a variation-aware federated learning (VAFL) framework, where the variations among clients are minimized by transforming the images of all clients onto a common image space. We first select one client with the lowest data complexity to define the target image space and synthesize a collection of images through a privacy-preserving generative adversarial network, called PPWGAN-GP. Then, a subset of those synthesized images, which effectively capture the characteristics of the raw images and are sufficiently distinct from any raw image, is automatically selected for sharing with other clients. For each client, a modified CycleGAN is applied to translate its raw images to the target image space defined by the shared synthesized images. In this way, the cross-client variation problem is addressed with privacy preservation. We apply the framework for automated classification of clinically significant prostate cancer and evaluate it using multi-source decentralized apparent diffusion coefficient (ADC) image data. Experimental results demonstrate that the proposed VAFL framework stably outperforms the current horizontal FL framework. As VAFL is independent of deep learning architectures for classification, we believe that the proposed framework is widely applicable to other medical image classification tasks. Zengqiang Yan, Jeffry Wicaksana, Zhiwei Wang 0002, Xin Yang 0008, Kwang-Ting Cheng |
IEEE J. Biomed. Health Informatics | 5 |
| 2021 | Fast Depth Prediction and Obstacle Avoidance on a Monocular Drone Using Probabilistic Convolutional Neural NetworkabstractRecent studies employ advanced deep convolutional neural networks (CNNs) for monocular depth perception, which can hardly run efficiently on small drones that rely on low/middle-grade GPU(e.g. TX2 and 1050Ti) for computation. In addition, the methods which can effectively and efficiently produce probabilistic depth prediction with a measure of model confidence have not been well studied. The lack of such a method could yield erroneous, sometimes fatal, decisions in drone applications (e.g. selecting a waypoint in a region with a large depth yet a low estimation confidence). This paper presents a real-time onboard approach for monocular depth prediction and obstacle avoidance with a lightweight probabilistic CNN (pCNN), which will be ideal for use in a lightweight energy-efficient drone. For each video frame, our pCNN can efficiently predict its depth map and the corresponding confidence. The accuracy of our lightweight pCNN is greatly boosted by integrating sparse depth estimation from a visual odometry into the network for guiding dense depth and confidence inference. The estimated depth map is transformed into Ego Dynamic Space (EDS) by embedding both dynamic motion constraints of a drone and the confidence values into the spatial depth map. Traversable waypoints are automatically computed in EDS based on which appropriate control inputs for the drone are produced. Extensive experimental results on public datasets demonstrate that our depth prediction method runs at 12Hz and 45Hz on TX2 and 1050Ti GPU respectively, which is 1.8X~5.6X faster than the state-of-the-art methods and achieves better depth estimation accuracy. We also conducted experiments of obstacle avoidance in both simulated and real environments to demonstrate the superiority of our method to the baseline methods. Xin Yang 0008, Yuanjie Dang, Hongcheng Luo, Yuesheng Tang, Chunyuan Liao, Peng Chen 0008, Kwang-Ting Cheng |
IEEE Trans. Intell. Transp. Syst. | 8 |
| 2021 | R2F: A Remote Retraining Framework for AIoT Processors With Computing ErrorsabstractArtificial Intelligence of Things (AIoT) processors fabricated with newer technology nodes suffer rising soft errors due to the shrinking transistor sizes and lower power supply. Soft errors on the AIoT processors particularly the deep learning accelerators (DLAs) with massive computing may cause substantial computing errors. These computing errors are difficult to be captured by the conventional training on general-purposed processors such as CPUs and GPUs in a server. Applying the offline trained neural network models to the edge accelerators with errors directly may lead to considerable prediction accuracy loss. To address the problem, we propose a remote retraining framework (R2F) for remote AIoT processors with computing errors. It takes the remote AIoT processor with soft errors in the training loop such that the on-site computing errors can be learned with the application data on the server and the retrained models can be resilient to the soft errors. Meanwhile, we propose an optimized partial triple modular redundancy (TMR) strategy to enhance the retraining. According to our experiments, R2F enables elastic design tradeoffs between the model accuracy and the performance penalty. The top-5 model accuracy can be improved by 1.93%–13.73% with 0%–200% performance penalty at high fault error rate. In addition, we notice that the retraining requires massive data transmission and even dominates the training time and propose a sparse increment compression approach for the data transmission optimization, which reduces the retraining time by 38%–88% on average with negligible accuracy loss over straightforward remote retraining. Dawen Xu 0002, Meng He 0012, Cheng Liu 0008, Ying Wang 0001, Long Cheng 0003, Huawei Li 0001, Xiaowei Li 0001, Kwang-Ting Cheng |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2021 | Reliability Evaluation and Analysis of FPGA-Based Neural Network Acceleration SystemabstractPrior works typically conducted the fault analysis of neural network accelerator computing arrays with simulation and focused on the prediction accuracy loss of the neural network models. There is still a lack of systematic fault analysis of the neural network acceleration system that considers both the accuracy degradation and system exceptions, such as system stall and running overtime. To that end, we implemented a representative neural network accelerator and corresponding fault injection modules on a Xilinx ARM-FPGA platform and evaluated the reliability of the system under different fault injection rates when a series of typical neural network models are deployed on the neural network acceleration system. The entire fault injection and reliability evaluation system is open-sourced on GitHub. With comprehensive experiments on the system, we identify the system exceptions based on the various abnormal behaviors of the FPGA-based neural network acceleration system and analyze the underlying reasons. Particularly, we find that the probability of the system exceptions dominates the reliability of the system. The faults also incur accuracy degradation of the neural network models, but the influence depends on the applications of the models and can vary greatly. In addition, we also evaluated the use of conventional triple modular redundancy (TMR) and demonstrated the challenge of TMR with both experiments and analytical models, which may shed light on the reliability design of the FPGA-based neural network acceleration system. Dawen Xu 0002, Ziyang Zhu, Cheng Liu 0008, Ying Wang 0001, Lei Zhang 0008, Huaguo Liang, Huawei Li 0001, Kwang-Ting Cheng |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2020 | Persistent Fault Analysis of Neural Networks on FPGA-based Acceleration SystemabstractThe increasing hardware failures caused by the shrinking semiconductor technologies pose substantial influence on the neural accelerators and improving the resilience of the neural network execution becomes a great design challenge especially to mission-critical applications such as self-driving and medical diagnose. The reliability analysis of the neural network execution is a key step to understand the influence of the hardware failures, and thus is highly demanded. Prior works typically conducted the fault analysis of neural network accelerators with simulation and concentrated on the prediction accuracy loss of the models. There is still a lack of systematic fault analysis of the neural network acceleration system that considers both the accuracy degradation and system exceptions such as system stall and early termination.In this work, we implemented a representative neural network accelerator and fault injection modules on a Xilinx ARM-FPGA platform and conducted fault analysis of the system using four typical neural network models. We had the system open-sourced on github. With comprehensive experiments, we identify the system exceptions based on the various abnormal behaviours of the FPGA-based neural network acceleration system and analyze the underlying reasons. Particularly, we find that the probability of the system exceptions dominates the reliability of the system and they are mainly caused by faults in the DMA, control unit and instruction memory of the accelerators. In addition, faults in these components also incur moderate accuracy degradation of the neural network models other than the system exceptions. Thus, these components are the most fragile part of the accelerators and need to be hardened for reliable neural network execution. Dawen Xu 0002, Ziyang Zhu, Cheng Liu 0008, Ying Wang 0001, Huawei Li 0001, Lei Zhang 0008, Kwang-Ting Cheng |
ASAP | 7 |
| 2020 | Cascaded Deep Monocular 3D Human Pose Estimation With Evolutionary Training DataabstractEnd-to-end deep representation learning has achieved remarkable accuracy for monocular 3D human pose estimation, yet these models may fail for unseen poses with limited and fixed training data. This paper proposes a novel data augmentation method that: (1) is scalable for synthesizing massive amount of training data (over 8 million valid 3D human poses with corresponding 2D projections) for training 2D-to-3D networks, (2) can effectively reduce dataset bias. Our method evolves a limited dataset to synthesize unseen 3D human skeletons based on a hierarchical human representation and heuristics inspired by prior knowledge. Extensive experiments show that our approach not only achieves state-of-the-art accuracy on the largest public benchmark, but also generalizes significantly better to unseen and rare poses. Relevant files and tools are available at the project website. Shichao Li 0002, Lei Ke, Kevin Pratama, Yu-Wing Tai, Chi-Keung Tang, Kwang-Ting Cheng |
CVPR | 6 |
| 2020 | Binarizing MobileNet via Evolution-Based SearchingabstractBinary Neural Networks (BNNs), known to be one among the effectively compact network architectures, have achieved great outcomes in the visual tasks. Designing efficient binary architectures is not trivial due to the binary nature of the network. In this paper, we propose a use of evolutionary search to facilitate the construction and training scheme when binarizing MobileNet, a compact network with separable depth-wise convolution. Being inspired by one-shot architecture search frameworks, we manipulate the idea of group convolution to design efficient 1-Bit Convolutional Neural Networks (CNNs), assuming an approximately optimal trade-off between computational cost and model accuracy. Our objective is to come up with a tiny yet efficient binary neural architecture by exploring the best candidates of the group convolution while optimizing the model performance in terms of complexity and latency. The approach is threefold. First, we modify and train strong baseline binary networks with a wide range of random group combinations at each convolutional layer. This set-up gives the binary neural networks a capability of preserving essential information through layers. Second, to find a good set of hyper-parameters for group convolutions we make use of the evolutionary search which leverages the exploration of efficient 1-bit models. Lastly, these binary models are trained from scratch in a usual manner to achieve the final binary model. Various experiments on ImageNet are conducted to show that following our construction guideline, the final model achieves 60.09% Top-1 accuracy and outperforms the state-of-the-art CI-BCNN with the same computational cost. NhatHai Phan, Zechun Liu, Dang Huynh, Marios Savvides, Kwang-Ting Cheng |
CVPR | 5 |
| 2020 | Robust Design of Large Area Flexible Electronics via Compressed SensingabstractLarge area flexible electronics (FE) is emerging for low-cost, light-weight wearable electronics, artificial skins and IoT nodes, benefiting from its low-cost fabrication and mechanical flexibility. How-ever, the low temperature requirement for fabrication on a flexible substrate and the large-area nature of flexible sensor arrays inevitably result in inadequate device yield, reliability and stability. Therefore, it is essential to develop design methodologies for large area sensing applications which can ensure system robustness with-out relying on highly reliable devices. Based on the observation that most signals sensed by body sensor arrays exhibit sparse statistical characteristics, we propose a system design method which lever-ages the sparse nature via compressed sensing (CS). Specifically, we use flexible circuitry to implement a CS encoder and decode the compressed signal in the silicon side. As a system demonstration, we fabricated the temperature sensor array, shift register and amplifier to illustrate the feasibility of the encoder design using carbon-nanotube-based flexible thin-film transistors. To evaluate the improvement of system robustness achieved by the proposed sensing schema, we conducted two case studies: temperature imaging and tactile-sensor based object recognition. With ~10% sparse errors (due to either device defects or transient errors), we achieved reduction of root-mean-square-error (RMSE) from 0.20 to 0.05 for temperature sensing and boost the classification accuracy from 65% to 84% for tactile-sensing based object recognition. Leilai Shao, Tsung-Ching Huang, Zhenan Bao, Kwang-Ting Cheng |
DAC | 5 |
| 2020 | Characterization and Applications of Spatial Variation Models for Silicon Microring-Based Optical TransceiversabstractPhotonic integrated circuits suffer from large process variations. Effective and accurate characterization of the variation patterns is a critical task for enabling the development of novel techniques to alleviate the variation challenges. In this study, we propose a hierarchical approach that effectively decomposes the spatial variations of silicon microring-based optical transceivers into wafer-level, intra-die, and inter-die components. We then demonstrate that the characterized variation models can be used to generate trustworthy synthetic data for architecture- and system-level solutions for variation alleviation. We further demonstrate the utility of our variation characterization method for accurate yield prediction based on partial measurement data. Yuyang Wang 0003, Jared Hulme, Mudit Jain, M. Ashkan Seyedi, Marco Fiorentino, Raymond G. Beausoleil, Kwang-Ting Cheng |
DAC | 8 |
| 2020 | ReActNet: Towards Precise Binary Neural Network with Generalized Activation Functions
Zechun Liu, Marios Savvides, Kwang-Ting Cheng |
ECCV (14) | 4 |
| 2020 | A Hybrid Computing Architecture for Fault-tolerant Deep Learning AcceleratorsabstractRegular 2D computing array is widely utilized for the processing of the major neural network operations in many deep learning accelerators (DLAs). Hardware failures on the array can lead to considerable computing errors and prediction accuracy loss. Prior works proposed to add homogeneous redundant PEs to each row or column of the regular computing array to mitigate faulty PEs, but they may fail to recover the computing array from faults when the number of faulty PEs in a row or column exceeds the number of redundant PEs in the corresponding row or column. The problem gets worse when the faults are not evenly distributed across the computing array. To address the problem, we propose a hybrid computing architecture (HCA) for fault-tolerant DLAs. Instead of adding homogeneous redundant PEs to the regular computing array of DLAs, it has a dot-production processing unit (DPPU) to recompute the operations that are mapped to the faulty PEs concurrently without performance penalty under moderate fault injection. Even under high fault injection, HCA can be degraded smoothly and remains functional. In addition, DPPU exploits the parallelism within each operation and processes the network operations sequentially, so it can tolerate faulty PEs in arbitrary locations and ensures steady performance under distinct fault distributions. According to our experiments, HCA shows significantly higher reliability and performance under various fault injection with comparable chip area penalty compared to the conventional redundancy approaches. Dawen Xu 0002, Cheng Chu, Cheng Liu 0008, Ying Wang 0001, Lei Zhang 0008, Huaguo Liang, Kwang-Ting Cheng |
ICCD | 8 |
| 2020 | Multi-phase and Multi-level Selective Feature Fusion for Automated Pancreas Segmentation from CT Images
Xixi Jiang, Qingqing Luo, Zhiwei Wang 0002, Xin Li 0001, Kwang-Ting Cheng, Xin Yang 0008 |
MICCAI (4) | 7 |
| 2020 | Bi-Real Net: Binarizing Deep Network Towards Real-Network Performance
Zechun Liu, Wenhan Luo, Baoyuan Wu, Xin Yang 0008, Wei Liu 0005, Kwang-Ting Cheng |
Int. J. Comput. Vis. | 6 |
| 2020 | Semi-supervised mp-MRI data synthesis with StitchLayer and auxiliary distance maximization
Zhiwei Wang 0002, Yi Lin 0009, Kwang-Ting Cheng, Xin Yang 0008 |
Medical Image Anal. | 3 |
| 2020 | Bi-Modality Medical Image Synthesis Using Semi-Supervised Sequential Generative Adversarial NetworksabstractIn this paper, we propose a bi-modality medical image synthesis approach based on sequential generative adversarial network (GAN) and semi-supervised learning. Our approach consists of two generative modules that synthesize images of the two modalities in a sequential order. A method for measuring the synthesis complexity is proposed to automatically determine the synthesis order in our sequential GAN. Images of the modality with a lower complexity are synthesized first, and the counterparts with a higher complexity are generated later. Our sequential GAN is trained end-to-end in a semi-supervised manner. In supervised training, the joint distribution of bi-modality images are learned from real paired images of the two modalities by explicitly minimizing the reconstruction losses between the real and synthetic images. To avoid overfitting limited training images, in unsupervised training, the marginal distribution of each modality is learned based on unpaired images by minimizing the Wasserstein distance between the distributions of real and fake images. We comprehensively evaluate the proposed model using two synthesis tasks based on three types of evaluate metrics and user studies. Visual and quantitative results demonstrate the superiority of our method to the state-of-the-art methods, and reasonable visual quality and clinical significance. Code is made publicly available at https://github.com/hust- linyi/Multimodal-Medical-Image-Synthesis. Xin Yang 0008, Yi Lin 0009, Zhiwei Wang 0002, Xin Li 0001, Kwang-Ting Cheng |
IEEE J. Biomed. Health Informatics | 5 |
| 2020 | Multi-Task Siamese Network for Retinal Artery/Vein Separation via Deep Convolution Along VesselabstractVascular tree disentanglement and vessel type classification are two crucial steps of the graph-based method for retinal artery-vein (A/V) separation. Existing approaches treat them as two independent tasks and mostly rely on ad hoc rules (e.g. change of vessel directions) and hand-crafted features (e.g. color, thickness) to handle them respectively. However, we argue that the two tasks are highly correlated and should be handled jointly since knowing the A/V type can unravel those highly entangled vascular trees, which in turn helps to infer the types of connected vessels that are hard to classify based on only appearance. Therefore, designing features and models isolatedly for the two tasks often leads to a suboptimal solution of A/V separation. In view of this, this paper proposes a multi-task siamese network which aims to learn the two tasks jointly and thus yields more robust deep features for accurate A/V separation. Specifically, we first introduce Convolution Along Vessel (CAV) to extract the visual features by convolving a fundus image along vessel segments, and the geometric features by tracking the directions of blood flow in vessels. The siamese network is then trained to learn multiple tasks: i) classifying A/V types of vessel segments using visual features only, and ii) estimating the similarity of every two connected segments by comparing their visual and geometric features in order to disentangle the vasculature into individual vessel trees. Finally, the results of two tasks mutually correct each other to accomplish final A/V separation. Experimental results demonstrate that our method can achieve accuracy values of 94.7%, 96.9%, and 94.5% on three major databases (DRIVE, INSPIRE, WIDE) respectively, which outperforms recent state-of-the-arts. Zhiwei Wang 0002, Xixi Jiang, Jingen Liu, Kwang-Ting Cheng, Xin Yang 0008 |
IEEE Trans. Medical Imaging | 4 |
| 2020 | Enabling a Single Deep Learning Model for Accurate Gland Instance Segmentation: A Shape-Aware Adversarial Learning FrameworkabstractSegmenting gland instances in histology images is highly challenging as it requires not only detecting glands from a complex background but also separating each individual gland instance with accurate boundary detection. However, due to the boundary uncertainty problem in manual annotations, pixel-to-pixel matching based loss functions are too restrictive for simultaneous gland detection and boundary detection. State-of-the-art approaches adopted multi-model schemes, resulting in unnecessarily high model complexity and difficulties in the training process. In this paper, we propose to use one single deep learning model for accurate gland instance segmentation. To address the boundary uncertainty problem, instead of pixel-to-pixel matching, we propose a segment-level shape similarity measure to calculate the curve similarity between each annotated boundary segment and the corresponding detected boundary segment within a fixed searching range. As the segment-level measure allows location variations within a fixed range for shape similarity calculation, it has better tolerance to boundary uncertainty and is more effective for boundary detection. Furthermore, by adjusting the radius of the searching range, the segment-level shape similarity measure is able to deal with different levels of boundary uncertainty. Therefore, in our framework, images of different scales are down-sampled and integrated to provide both global and local contextual information for training, which is helpful in segmenting gland instances of different sizes. To reduce the variations of multi-scale training images, by referring to adversarial domain adaptation, we propose a pseudo domain adaptation framework for feature alignment. By constructing loss functions based on the segment-level shape similarity measure, combining with the adversarial loss function, the proposed shape-aware adversarial learning framework enables one single deep learning model for gland instance segmentation. Experimental results on the 2015 MICCAI Gland Challenge dataset demonstrate that the proposed framework achieves state-of-the-art performance with one single deep learning model. As the boundary uncertainty problem widely exists in medical image segmentation, it is broadly applicable to other applications. Zengqiang Yan, Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Medical Imaging | 3 |
| 2019 | Bidirectional tuning of microring-based silicon photonic transceivers for optimal energy efficiencyabstractMicroring-based silicon photonic transceivers are promising to resolve the communication bottleneck of future high-performance computing systems. To rectify process variations in microring resonance wavelengths, thermal tuning is usually preferred over electrical tuning due to its preservation of extinction ratios and quality factors. However, the low energy efficiency of resistive thermal tuners results in nontrivial tuning cost and overall energy consumption of the transceiver. In this study, we propose a hybrid tuning strategy which involves both thermal and electrical tuning. Our strategy determines the tuning direction of each resonance wavelength with the goal of optimizing the transceiver energy efficiency without compromising signal integrity. Formulated as an integer programming problem and solved by a genetic algorithm, our tuning strategy yields 32%~53% savings of overall energy per bit for measured data of 5-channel transceivers at 5~10 Gb/s per channel, and up to 24% saving for synthetic data of 30-channel transceivers, generated based on the process variation models built upon measured data. We further investigated a polynomial-time approximation method which achieves over 100x speedup in tuning scheme computation, while still maintaining considerable energy-per-bit savings. Yuyang Wang 0003, M. Ashkan Seyedi, Jared Hulme, Marco Fiorentino, Raymond G. Beausoleil, Kwang-Ting Cheng |
ASP-DAC | 6 |
| 2019 | Learning to Film From Professional Human Motion VideosabstractWe investigate the problem of 6 degrees of freedom (DOF) camera planning for filming professional human motion videos using a camera drone. Existing methods either plan motions for only a pan-tilt-zoom (PTZ) camera, or adopt ad-hoc solutions without carefully considering the impact of video contents and previous camera motions on the future camera motions. As a result, they can hardly achieve satisfactory results in our drone cinematography task. In this study, we propose a learning-based framework which incorporates the video contents and previous camera motions to predict the future camera motions that enable the capture of professional videos. Specifically, the inputs of our framework are video contents which are represented using subject-related feature based on 2D skeleton and scene-related features extracted from background RGB images, and camera motions which are represented using optical flows. The correlation between the inputs and output future camera motions are learned via a sequence-to-sequence convolutional long short-term memory (Seq2Seq ConvLSTM) network from a large set of video clips. We deploy our approach to a real drone cinematography system by first predicting the future camera motions, and then converting them to the drone's control commands via an odometer. Our experimental results on extensive datasets and showcases exhibit significant improvements in our approach over conventional baselines and our approach can successfully mimic the footage of a professional cameraman. Chong Huang 0005, David Chuan-En Lin, Yan Kong, Peng Chen 0008, Xin Yang 0008, Kwang-Ting Cheng |
CVPR | 7 |
| 2019 | Ultra-thin Skin Electronics for High Quality and Continuous Skin-Sensor-Silicon InterfacingabstractSkin-inspired electronics emerges as a new paradigm due to the increasing demands for conformable and high-quality skin-sensor-silicon (SSS) interfacing in wearable, electronic skin and health monitoring applications. Advances in ultra-thin, flexible, stretchable and conformable materials have made skin electronics feasible. In this paper, we prototyped an active electrode (with a thickness ≤ 2 um), which integrates the electrode with a thin-film transistor (TFT) based amplifier, to effectively suppress motion artifacts. The fabricated ultra-thin amplifier can achieve a gain of 32 dB at 20 kHz, demonstrating the feasibility of the proposed active electrode. Using atrial fibrillation (AF) detection for electrocardiogram (ECG) as an application driver, we further develop a simulation framework taking into account all elements including the skin, the sensor, the amplifier and the silicon chip. Systematic and quantitative simulation results indicate that the proposed active electrode can effectively improve the signal quality under motion noises (achieving ≥30 dB improvement in signal-to-noise ratio (SNR)), which boosts classification accuracy by more than 19% for AF detection. Leilai Shao, Sicheng Li 0001, Tsung-Ching Huang, Raymond G. Beausoleil, Zhenan Bao, Kwang-Ting Cheng |
DAC | 7 |
| 2019 | Evaluating Assertion Set Completeness to Expose Hardware Trojans and Verification BlindspotsabstractAssertion-based verification has been adopted by industry as an efficient specification mechanism. Handwritten assertions encode design intent in a parsable format and have been traditionally used to verify an implementation conforms to the properties outlined by the assertions. Our work makes the observation that design behavior not covered by the assertion set is equally revealing and can be leveraged to identify malicious behavior (hardware Trojans) as well as verification blindspots. The difficulty in examining this unspecified and unverified behavior is differentiating between benign functionality that is truly don't care and that which leaks information or violates design intent. Prior work exploring assertion set completeness suffers from this inability to distinguish benign unspecified functionality from actual verification holes, while existing Trojan detection techniques can differentiate these categories, but require unspecified functionality already be characterized. Our technique uses the assertion set and simulation trace data available in most industry design flows to characterize unspecified functionality then separates Trojans and verification blindspots from benign behavior using existing Trojan detection methods. Using our technique, we uncover missing functionality in a first-in first-out (FIFO) queue implementation and demonstrate detection of information leakage Trojans. We also illustrate Trojan detection for a system containing several components connected by an AXI4-Lite bus by analyzing the completeness of the AXI4-Lite assertion set provided by ARM. Nicole Fern, Kwang-Ting Cheng |
DATE | 2 |
| 2019 | Process Design Kit and Design Automation for Flexible Hybrid ElectronicsabstractHigh-performance low-cost flexible hybrid electronics (FHE) are desirable for internet of things (IoT). Carbon-nanotube (CNT) thin-film transistor (TFT) is a promising candidate for high-performance FHE because of its high carrier mobility (25cm2/V.s), superior mechanical flexibility/stretchability, and material compatibility with low-cost printing and solution processes. Flexible sensors and peripheral CNT-TFT circuits, such as decoders, drivers and sense amplifiers, can be printed and integrated with thinned (<;50μm) silicon chips on soft, thin, and flexible substrates for appealing product designs and form factors. Here we report: 1) process design kit (PDK) to enable FHE design automation, from device modeling to physical verification, and 2) open-source and solution-process proven intellectual property (IP) blocks, including Pseudo-CMOS [1] digital logic and analog amplifiers on flexible substrates, as shown in Figure 1. The proposed FHE-PDK and circuit design IP are fully compatible with silicon design EDA tools, and can be readily used for co-design with both CNT-TFT circuits and silicon chips. Tsung-Ching Huang, Leilai Shao, Sridhar Sivapurapu, Madhavan Swaminathan, Sicheng Li 0001, Zhenan Bao, Kwang-Ting Cheng, Raymond G. Beausoleil |
DATE | 8 |
| 2019 | Task Mapping-Assisted Laser Power Scaling for Optical Network-on-ChipsabstractEnergy efficiency of an optical network-on-chip (ONoC) largely relies on an effective laser power management strategy. Addressing the limitations of existing techniques, we propose a Task Mapping-Assisted Laser Power Scaling (TMALPS) framework to optimize the energy consumption and the application execution time of an ONoC. Through the combination of task mapping exploration and runtime laser power reconfiguration applied to a wide range of application benchmarks, our TMALPS framework achieves an average of 66% saving of the energy-delay product, compared to a baseline scenario where the optimization techniques are not applied. Significant improvement over existing techniques was also observed. The hardware overhead required to support our TMALPS framework is minimal with intelligent reuse of existing on-chip hardware resource. Yuyang Wang 0003, Kwang-Ting Cheng |
ICCAD | 2 |
| 2019 | MetaPruning: Meta Learning for Automatic Neural Network Channel PruningabstractIn this paper, we propose a novel meta learning approach for automatic channel pruning of very deep neural networks. We first train a PruningNet, a kind of meta network, which is able to generate weight parameters for any pruned structure given the target network. We use a simple stochastic structure sampling method for training the PruningNet. Then, we apply an evolutionary procedure to search for good-performing pruned networks. The search is highly efficient because the weights are directly generated by the trained PruningNet and we do not need any finetuning at search time. With a single PruningNet trained for the target network, we can search for various Pruned Networks under different constraints with little human participation. Compared to the state-of-the-art pruning methods, we have demonstrated superior performances on MobileNet V1/V2 and ResNet. Codes are available on https://github.com/liuzechun/MetaPruning. Zechun Liu, Haoyuan Mu, Xiangyu Zhang 0005, Zichao Guo, Xin Yang 0008, Kwang-Ting Cheng, Jian Sun 0001 |
ICCV | 6 |
| 2019 | Learning to Capture a Film-Look Video with a Camera DroneabstractThe development of intelligent drones has simplified aerial filming and provided smarter assistant tools for users to capture a film-look footage. Existing methods of autonomous aerial filming either specify predefined camera movements for a drone to capture a footage, or employ heuristic approaches for camera motion planning. However, both predefined movements and heuristically planned motions are hardly able to provide cinematic footages for various dynamic scenarios. In this paper, we propose a data-driven learning-based approach, which can imitate a professional cameraman's intention for capturing a film-look aerial footage of a single subject in real-time. We model the decision-making process of the cameraman with two steps: 1) we train a network to predict the future image composition and camera position, and 2) our system then generates control commands to achieve the desired shot framing. At the system level, we deploy our algorithm on the limited resources of a drone and demonstrate the feasibility of running automatic filming onboard in real-time. Our experiments show how our data-driven planning approach achieves film-look footages and successfully mimics the work of a professional cameraman. Chong Huang 0005, Yan Kong, Peng Chen 0008, Xin Yang 0008, Kwang-Ting Cheng |
ICRA | 6 |
| 2019 | Automated Pulmonary Embolism Detection from CTPA Images Using an End-to-End Convolutional Neural Network
Yi Lin 0009, Jianchao Su, Jingen Liu, Kwang-Ting Cheng, Xin Yang 0008 |
MICCAI (4) | 6 |
| 2019 | Latent Weights Do Not Exist: Rethinking Binarized Neural Network OptimizationabstractOptimization of Binarized Neural Networks (BNNs) currently relies on real-valued latent weights to accumulate small update steps. In this paper, we argue that these latent weights cannot be treated analogously to weights in real-valued networks. Instead their main role is to provide inertia during training. We interpret current methods in terms of inertia and provide novel insights into the optimization of BNNs. We subsequently introduce the first optimizer specifically designed for BNNs, Binary Optimizer (Bop), and demonstrate its performance on CIFAR-10 and ImageNet. Together, the redefinition of latent weights as inertia and the introduction of Bop enable a better understanding of BNN optimization and open up the way for further improvements in training methodologies for BNNs. Koen Helwegen, James Widdicombe, Lukas Geiger, Zechun Liu, Kwang-Ting Cheng, Roeland Nusselder |
NeurIPS | 5 |
| 2019 | Reactive obstacle avoidance of monocular quadrotors with online adapted depth prediction network
Xin Yang 0008, Hongcheng Luo, Yuhao Wu 0010, Chunyuan Liao, Kwang-Ting Cheng |
Neurocomputing | 6 |
| 2019 | A Three-Stage Deep Learning Model for Accurate Retinal Vessel SegmentationabstractAutomatic retinal vessel segmentation is a fundamental step in the diagnosis of eye-related diseases, in which both thick vessels and thin vessels are important features for symptom detection. All existing deep learning models attempt to segment both types of vessels simultaneously by using a unified pixel-wise loss that treats all vessel pixels with equal importance. Due to the highly imbalanced ratio between thick vessels and thin vessels (namely the majority of vessel pixels belong to thick vessels), the pixel-wise loss would be dominantly guided by thick vessels and relatively little influence comes from thin vessels, often leading to low segmentation accuracy for thin vessels. To address the imbalance problem, in this paper, we explore to segment thick vessels and thin vessels separately by proposing a three-stage deep learning model. The vessel segmentation task is divided into three stages, namely thick vessel segmentation, thin vessel segmentation, and vessel fusion. As better discriminative features could be learned for separate segmentation of thick vessels and thin vessels, this process minimizes the negative influence caused by their highly imbalanced ratio. The final vessel fusion stage refines the results by further identifying nonvessel pixels and improving the overall vessel thickness consistency. The experiments on public datasets DRIVE, STARE, and CHASE_DB1 clearly demonstrate that the proposed three-stage deep learning model outperforms the current state-of-the-art vessel segmentation methods. Zengqiang Yan, Xin Yang 0008, Kwang-Ting Cheng |
IEEE J. Biomed. Health Informatics | 3 |
| 2019 | Real-Time Dense Monocular SLAM With Online Adapted Depth Prediction NetworkabstractConsiderable advances have been achieved in estimating the depth map from a single image via convolutional neural networks (CNNs) during the past few years. Combining depth prediction from CNNs with conventional monocular simultaneous localization and mapping (SLAM) is promising for accurate and dense monocular reconstruction, in particular addressing the two long-standing challenges in conventional monocular SLAM: low map completeness and scale ambiguity. However, depth estimated by pretrained CNNs usually fails to achieve sufficient accuracy for environments of different types from the training data, which are common for certain applications such as obstacle avoidance of drones in unknown scenes. Additionally, inaccurate depth prediction of CNN could yield large tracking errors in monocular SLAM. In this paper, we present a real-time dense monocular SLAM system, which effectively fuses direct monocular SLAM with an online-adapted depth prediction network for achieving accurate depth prediction of scenes of different types from the training data and providing absolute scale information for tracking and mapping. Specifically, on one hand, tracking pose (i.e., translation and rotation) from direct SLAM is used for selecting a small set of highly effective and reliable training images, which acts as ground truth for tuning the depth prediction network on-the-fly toward better generalization ability for scenes of different types. A stage-wise Stochastic Gradient Descent algorithm with a selective update strategy is introduced for efficient convergence of the tuning process. On the other hand, the dense map produced by the adapted network is applied to address scale ambiguity of direct monocular SLAM which in turn improves the accuracy of both tracking and overall reconstruction. The system with assistance of both CPUs and GPUs, can achieve real-time performance with progressively improved reconstruction accuracy. Experimental results on public datasets and live application to obstacle avoidance of drones demonstrate that our method outperforms the state-of-the-art methods with greater map completeness and accuracy, and a smaller tracking error. Hongcheng Luo, Yuhao Wu 0010, Chunyuan Liao, Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Multim. | 6 |
| 2019 | Bayesian DeNet: Monocular Depth Prediction and Frame-Wise Fusion With Synchronized UncertaintyabstractUsing deep convolutional neural networks (CNN) to predict the depth from a single image has received considerable attention in recent years due to its impressive performance. However, existing methods process each single image independently without leveraging the multiview information of video sequences in practical scenarios. Properly taking into account multiview information in video sequences beyond individual frames could offer considerable benefits in terms of depth prediction accuracy and robustness. In addition, a meaningful measure of prediction uncertainty is essential for decision making, which is not provided in existing methods. This paper presents a novel video-based depth prediction system based on a monocular camera, named Bayesian DeNet. Specifically, Bayesian DeNet consists of a 59-layer CNN that can concurrently output a depth map and an uncertainty map for each video frame. Each pixel in an uncertainty map indicates the error variance of the corresponding depth estimate. Depth estimates and uncertainties of previous frames are propagated to the current frame based on the tracked camera pose, yielding multiple depth/uncertainty hypotheses for the current frame which are then fused in a Bayesian inference framework for greater accuracy and robustness. Extensive exper-iments on three public datasets demonstrate that our Bayesian DeNet outperforms the state-of-the-art methods for monocular depth prediction. A demo video and code are publicly available.1 Xin Yang 0008, Hongcheng Luo, Chunyuan Liao, Kwang-Ting Cheng |
IEEE Trans. Multim. | 5 |
| 2018 | Process design kit for flexible hybrid electronicsabstractFlexible Electronics (FE) is emerging for wearables and low-cost internet of things (IoT) nodes benefiting from its low-cost fabrication and mechanical flexibility. Combining FE with thinned silicon chips, known as flexible hybrid electronics (FHE), can take advantages of both low-cost printed electronics and high performance silicon chips. To design a FHE system, the process design kit (PDK) offering the capabilities for circuit design, simulation and verification for both FE and silicon chips is needed. The key elements of FHE-PDK include technology files for design rule checking (DRC), layout versus schematic (LVS) and layout parasitics extraction (LPE), as well as SPICE-compatible models for flexible thin-film transistors (TFTs) and passive elements. Wafer scale measurements are used to validate our SPICE models and design rules are derived accordingly to assure a satisfactory yield. With FHE-PDK, circuit and system designers can therefore focus on design innovations and can rely on design tools to produce manufacturable designs. Leilai Shao, Tsung-Ching Huang, Zhenan Bao, Raymond G. Beausoleil, Kwang-Ting Cheng |
ASP-DAC | 6 |
| 2018 | Pairing of microring-based silicon photonic transceivers for tuning power optimizationabstractNanophotonic interconnects have started replacing traditional electrical interconnects in data centers for rack-level communications and demonstrated great potential at board and chip levels. However, microring-based silicon photonic transceivers, an important element for nanophotonic interconnects, are very sensitive to fabrication process variations, and require power hungry wavelength tuning. In this paper, we apply efficient optimization algorithms to mix-and-match a pool of fabricated transceiver devices with the objective of minimizing the overall tuning power. This optimal pairing technique, applied during the production stage, reduce power consumption for wavelength tuning. For two sets of fabricated devices, the pairs of transceivers assigned by the optimal pairing technique reduce the tuning power by 6% to 60%. We further evaluate the method on synthetic data sets that are generated from a well-established process variation model. Our experimental results show that even greater power saving can be achieved when more fabricated devices are available for pairing and the runtime of the optimization algorithm is quite scalable. Rui Wu 0008, M. Ashkan Seyedi, Yuyang Wang 0003, Jared Hulme, Marco Fiorentino, Raymond G. Beausoleil, Kwang-Ting Cheng |
ASP-DAC | 7 |
| 2018 | StitchAD-GAN for Synthesizing Apparent Diffusion Coefficient Images of Clinically Significant Prostate Cancer
Zhiwei Wang 0002, Yi Lin 0009, Chunyuan Liao, Kwang-Ting Cheng, Xin Yang 0008 |
BMVC | 4 |
| 2018 | Compact modeling of carbon nanotube thin film transistors for flexible circuit designabstractCarbon nanotube thin film transistor (CNT-TFT) is a promising candidate for flexible electronics, because of its high carrier mobility and great mechanical flexibility. An accurate and trustworthy device model for CNT-TFTs, however, is still missing. In this paper, we present a SPICE-compatible compact model for CNT-TFT circuit simulation and validate the proposed model based on fabricated CNT-TFTs and Pseudo-CMOS circuits [1][2]. The proposed CNT-TFT model enables circuit designers to explore design space by adjusting device parameters, supply voltages and transistor sizes to optimize the noise margin (NM) and power-delay product (PDP), which are the key merits for larger scale CNT-TFT circuits. We further propose a design framework to effectively optimize the NM and PDP to facilitate greater automation of flexible circuit design based on CNT-TFTs. Leilai Shao, Tsung-Ching Huang, Zhenan Bao, Raymond G. Beausoleil, Kwang-Ting Cheng |
DATE | 6 |
| 2018 | Energy-efficient channel alignment of DWDM silicon photonic transceiversabstractThe comb laser-driven microring-based dense wavelength division multiplexing silicon photonics is a promising candidate for next-generation optical interconnects. However, existing solutions for exploring the power-performance trade-off of such systems have been restricted to a limited design space, resulting from the unnecessary constraints of using an identical spacing for laser comb lines and microring channels, and of utilizing consecutive laser comb lines for data transmission. We propose an energy-efficient channel alignment scheme that aligns the microring channels to a subset of laser comb lines that are non-uniformly distributed in the free spectrum range of the microrings. Based on a well-established process variation model, our simulations show that the proposed scheme significantly reduces the microring tuning power in the presence of denser comb lines. The power saved from microring tuning can improve the overall system energy efficiency despite some power wasted in unused laser comb lines. We further conducted a case study for design space exploration using the proposed channel alignment scheme, seeking the most energy-efficient configuration in order to achieve a target aggregated data rate. Yuyang Wang 0003, M. Ashkan Seyedi, Rui Wu 0008, Jared Hulme, Marco Fiorentino, Raymond G. Beausoleil, Kwang-Ting Cheng |
DATE | 7 |
| 2018 | Bi-Real Net: Enhancing the Performance of 1-Bit CNNs with Improved Representational Capability and Advanced Training Algorithm
Zechun Liu, Baoyuan Wu, Wenhan Luo, Xin Yang 0008, Wei Liu 0005, Kwang-Ting Cheng |
ECCV (15) | 6 |
| 2018 | ACT: An Autonomous Drone Cinematography System for Action ScenesabstractDrones are enabling new forms of cinematography. Aerial filming via drones in action scenes is difficult because it requires users to understand the dynamic scenarios and operate the drone and camera simultaneously. Existing systems allow the user to manually specify the shots and guide the drone to capture footage, while none of them employ aesthetic objectives to automate aerial filming in action scenes. Meanwhile, these drone cinematography systems depend on the external motion capture systems to perceive the human action, which is limited to the indoor environment. In this paper, we propose an Autonomous CinemaTography system “ACT” on the drone platform to address the above the challenges. To our knowledge, this is the first drone camera system which can autonomously capture cinematic shots of action scenes based on limb movements in both indoor and outdoor environments. Our system includes the following novelties. First, we propose an efficient method to extract 3D skeleton points via a stereo camera. Second, we design a real-time dynamical camera planning strategy that fulfills the aesthetic objectives for filming and respects the physical limits of a drone. At the system level, we integrate cameras and GPUs into the limited space of a drone and demonstrate the feasibility of running the entire cinematography system onboard in real-time. Experimental results in both simulation and real-world scenarios demonstrate that our cinematography system “ACT” can capture more expressive video footage of human action than that of a state-of-the-art drone camera system. Chong Huang 0005, Fei Gao 0011, Jie Pan 0004, Weihao Qiu, Peng Chen 0008, Xin Yang 0008, Shaojie Shen, Kwang-Ting Cheng |
ICRA | 9 |
| 2018 | Through-the-Lens Drone FilmingabstractAerial filming in action scenes using a drone is difficult for inexperienced flyers because manipulating a remote controller and meeting the desired image composition are two independent, while concurrent, tasks. Existing systems attempt to utilize wearable GPS-based or infrared-based sensors to track the human movement and to assist in capturing footage. However, these sensors work only in either indoor (infrared-based) or outdoor environments (GPS-based), but not both. In this paper, we introduce a novel drone filming system which integrates monocular 3D human pose estimation and localization into a drone platform to remove the constraints imposed by wearable-sensor-based solutions. Meanwhile, given the estimated position, we propose a novel drone control system, called “through-the-lens drone filming”, to allow a cameraman to conveniently control the drone by manipulating a 3D model in the preview, which closes the gap between the flight control and the viewpoint design. Our system includes two key enabling techniques: 1) subject localization based on visual-inertial fusion, and 2) through-the-lens camera planning. This is the first drone camera system which allows users to capture human actions by manipulating the camera in a virtual environment. From the drone hardware, we integrate a gimbal camera and two GPUs into the limited space of a drone and demonstrate the feasibility of running the entire system onboard with insignificant delays, which are sufficient for filming in our real-time application. Experimental results, in both simulation and real-world scenarios, demonstrate that our techniques can greatly ease camera control and capture better videos. Chong Huang 0005, Yan Kong, Peng Chen 0008, Xin Yang 0008, Kwang-Ting Cheng |
IROS | 6 |
| 2018 | Pre-silicon Formal Verification of JTAG Instruction Opcodes for SecurityabstractWidely implemented standards such as IEEE 1149.1 (JTAG) and 1687 (iJTAG) are essential in providing improved chip and board testability, but it has been demonstrated that undocumented or poorly obfuscated scan and debug instructions can be exploited by hackers to undermine system security. Prior work proposes adding authentication or encryption to JTAG to improve security, but these methods can only protect functionality known to the design and test team. Out-of-spec JTAG functionality can be inserted accidentally or with malicious intent (e.g., hardware Trojans). Our proposed technique can detect anomalous JTAG instructions not present in the specification using commercial formal equivalence checking tools. We demonstrate the effectiveness of our technique by characterizing the entire JTAG instruction set space for the OpenSPARC T2 benchmark in a completely automated manner. In the original design our technique formally proves all undefined opcodes map to the benign bypass instruction and provides the size and location within the design hierarchy of all data registers. In a modified version of the design our technique correctly detects several undefined opcodes that are used to access the L2 cache, as well as extra out-of-spec elements in a data register selected by an existing instruction. Nicole Fern, Kwang-Ting Cheng |
ITC | 2 |
| 2018 | A Deep Model with Shape-Preserving Loss for Gland Instance Segmentation
Zengqiang Yan, Xin Yang 0008, Kwang-Ting Cheng |
MICCAI (2) | 3 |
| 2018 | Monocular Camera Based Real-Time Dense Mapping Using Generative Adversarial NetworkabstractMonocular simultaneous localization and mapping (SLAM) is a key enabling technique for many computer vision and robotics applications. However, existing methods either can obtain only sparse or semi-dense maps in highly-textured image areas or fail to achieve a satisfactory reconstruction accuracy. In this paper, we present a new method based on a generative adversarial network,named DM-GAN, for real-time dense mapping based on a monocular camera. Specifcally, our depth generator network takes a semidense map obtained from motion stereo matching as a guidance to supervise dense depth prediction of a single RGB image. The depth generator is trained based on a combination of two loss functions, i.e. an adversarial loss for enforcing the generated depth maps to reside on the manifold of the true depth maps and a pixel-wise mean square error (MSE) for ensuring the correct absolute depth values. Extensive experiments on three public datasets demonstrate that our DM-GAN signifcantly outperforms the state-of-the-art methods in terms of greater reconstruction accuracy and higher depth completeness. Xin Yang 0008, Zhiwei Wang 0002, Qiaozhe Zhang, Wenyu Liu 0001, Chunyuan Liao, Kwang-Ting Cheng |
ACM Multimedia | 7 |
| 2018 | Robust and real-time pose tracking for augmented reality on mobile devices
Xin Yang 0008, Jiabin Guo, Tangli Xue, Kwang-Ting Cheng |
Multim. Tools Appl. | 4 |
| 2018 | Automated Detection of Clinically Significant Prostate Cancer in mp-MRI Images Based on an End-to-End Deep Neural NetworkabstractAutomated methods for detecting clinically significant (CS) prostate cancer (PCa) in multi-parameter magnetic resonance images (mp-MRI) are of high demand. Existing methods typically employ several separate steps, each of which is optimized individually without considering the error tolerance of other steps. As a result, they could either involve unnecessary computational cost or suffer from errors accumulated over steps. In this paper, we present an automated CS PCa detection system, where all steps are optimized jointly in an end-to-end trainable deep neural network. The proposed neural network consists of concatenated subnets: 1) a novel tissue deformation network (TDN) for automated prostate detection and multimodal registration and 2) a dual-path convolutional neural network (CNN) for CS PCa detection. Three types of loss functions, i.e., classification loss, inconsistency loss, and overlap loss, are employed for optimizing all parameters of the proposed TDN and CNN. In the training phase, the two nets mutually affect each other and effectively guide registration and extraction of representative CS PCa-relevant features to achieve results with sufficient accuracy. The entire network is trained in a weakly supervised manner by providing only image-level annotations (i.e., presence/absence of PCa) without exact priors of lesions' locations. Compared with most existing systems which require supervised labels, e.g., manual delineation of PCa lesions, it is much more convenient for clinical usage. Comprehensive evaluation based on fivefold cross validation using 360 patient data demonstrates that our system achieves a high accuracy for CS PCa detection, i.e., a sensitivity of 0.6374 and 0.8978 at 0.1 and 1 false positives per normal/benign patient. Zhiwei Wang 0002, Chaoyue Liu 0002, Danpeng Cheng, Liang Wang 0052, Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Medical Imaging | 6 |
| 2018 | A Skeletal Similarity Metric for Quality Evaluation of Retinal Vessel SegmentationabstractThe most commonly used evaluation metrics for quality assessment of retinal vessel segmentation are sensitivity, specificity, and accuracy, which are based on pixel-to-pixel matching. However, due to the inter-observer problem that vessels annotated by different observers vary in both thickness and location, pixel-to-pixel matching is too restrictive to fairly evaluate the results of vessel segmentation. In this paper, the proposed skeletal similarity metric is constructed by comparing the skeleton maps generated from the reference and the source vessel segmentation maps. To address the inter-observer problem, instead of using a pixel-to-pixel matching strategy, each skeleton segment in the reference skeleton map is adaptively assigned with a searching range whose radius is determined based on its vessel thickness. Pixels in the source skeleton map located within the searching range are then selected for similarity calculation. The skeletal similarity consists of a curve similarity, which measures the structural similarity between the reference and the source skeleton maps and a thickness similarity, which measures the thickness consistency between the reference and the source vessel segmentation maps. In contrast to other metrics that provide a global score for the overall performance, we modify the definitions of true positive, false negative, true negative, and false positive based on the skeletal similarity, based on which sensitivity, specificity, accuracy, and other objective measurements can be constructed. More importantly, the skeletal similarity metric has better potential to be used as a pixelwise loss function for training deep learning models for retinal vessel segmentation. Through comparison of a set of examples, we demonstrate that the redefined metrics based on the skeletal similarity are more effective for quality evaluation, especially with greater tolerance to the inter-observer problem. Zengqiang Yan, Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Medical Imaging | 3 |
| 2017 | ASP-DAC 2017 keynote speech I-1: Heterogeneous integration of X-tronics: Design automation and educationabstractSummary form only given, as follows. The complete presentation was not made available for publication as part of the conference proceedings. Advances in photonics, flexible electronics, emerging memories, etc. and Si electronics' integration with these devices have enabled new classes of integrated circuits and systems with enhanced functionality, higher performance, or lower power consumption. Driving greater integration of such heterogeneous X-tronics can facilitate the continued proliferation of low-cost micro-/nano-systems for a wide range of applications. However, achieving their large-scale integration will require design ecosystem and design automation tools/methodologies much like those that enabled electronic integration in previous decades. In this talk, I will briefly introduce two recent Manufacturing Innovation Institutes, on Integrated Photonics and on Flexible Hybrid Electronics respectively, and a research center on developing 3D Hybrid CMOS-memristor circuits, which bring together academia, industry, and federal partners to increase U.S. manufacturing competitiveness in these areas. I will then focus on their design automation efforts and highlight the needs, challenges and opportunities of developing a robust design ecosystem for X-tronics integration. I will also share the educational challenges of talent development for X-tronics design automation. Kwang-Ting Cheng |
ASP-DAC | 1 |
| 2017 | Detecting hardware Trojans in unspecified functionality through solving satisfiability problemsabstractFor modern complex designs it is impossible to fully specify design behavior, and only feasible to verify functionally meaningful scenarios. Hardware Trojans modifying only unspecified functionality are not possible to detect using existing verification methodologies and Trojan detection strategies. We propose a detection methodology for these Trojans by 1) precisely defining “suspicious” unspecified functionality in terms of information leakage, and 2) formulating detection as a satisfiability problem that can take advantage of the recent advances in both boolean and satisfiability modulo theory (SMT) solvers. The formulated detection procedure can be applied to a gate-level design using commercial equivalence checking tools, or directly to the Verilog/VHDL code by reasoning about the satisfiability of SMT expressions built from traversing the data-flow graph. We demonstrate the effectiveness of our approach on an adder coprocessor and a UART communication controller infected with Trojans which process information leaked from the on-chip bus during idle cycles using signals with only partially specified behavior. Nicole Fern, Ismail San, Kwang-Ting Cheng |
ASP-DAC | 3 |
| 2017 | DLPS: Dynamic laser power scaling for optical Network-on-ChipabstractOptical Network-on-Chip (NoC), offering the advantages of low energy consumption, high bandwidth, and low latency, is a promising solution for on-chip communications of multi-core systems. However, on-chip lasers, a key element in optical NoCs, are a dominant source of power consumption. In this paper, we propose dynamic laser power scaling (DLPS), a fine-grained control strategy for minimizing laser power consumption while meeting the communication bandwidth required for the application. The proposed DLPS strategy intelligently switches among multiple operation modes based on the communication traffic pattern. Our experiments show that by introducing two new modes (standby, and intermediate data rate), DLPS can further reduce the communication energy for communication-intensive applications, compared to a simple on-off control strategy that dynamically turns lasers either completely on or off. Fan Lan, Rui Wu 0008, Kwang-Ting Cheng |
ASP-DAC | 5 |
| 2017 | An artificial neural network approach for screening test escapesabstractIn this paper we investigate the application of an artificial neural network (ANN) for screening test escapes. Specifically, we propose to train an autoencoder, an ANN, in an unsupervised way to fit the good chip population, i.e. using good chips only as the training set. The autoencoder is designed with both its input and output layers representing a set of features that characterize the test data of the chips under test, where we use the Euclidean distance between the values in the input and output layers as the cost function for training. Based on the trained autoencoder, if the test measurement of a query chip has an abnormally large value for the cost function, the chip is likely to be a test escape because it does not fit the characteristics of the good chip population captured by the model. We demonstrate that an autoencoder-based classification could achieve a higher detection rate for test escapes and a significant reduction in runtime and memory usage, compared with an SVM applied on the same features and some additional proximity features generated from multiple nonlinear transformations. Fan Lin, Kwang-Ting Cheng |
ASP-DAC | 2 |
| 2017 | 3D-DPE: A 3D high-bandwidth dot-product engine for high-performance neuromorphic computingabstractWe present and experimentally validate 3D-DPE, a general-purpose dot-product engine, which is ideal for accelerating artificial neural networks (ANNs). 3D-DPE is based on a monolithically integrated 3D CMOS-memristor hybrid circuit and performs a high-dimensional dot-product operation (a recurrent and computationally expensive operation in ANNs) within a single step, using analog current-based computing. 3D-DPE is made up of two subsystems, namely a CMOS subsystem serving as the memory controller and an analog memory subsystem consisting of multiple layers of high-density memristive crossbar arrays fabricated on top of the CMOS subsystem. Their integration is based on a high-density area-distributed interface, resulting in much higher connectivity between the two subsystems, compared to the traditional interface of a 2D system or a 3D system integrated using through silicon vias. As a result, 3D-DPE's single-step dot-product operation is not limited by the memory bandwidth, and the input dimension of the operations scales well with the capacity of the 3D memristive arrays. To demonstrate the feasibility of 3D-DPE, we designed and fabricated a CMOS memory controller and monolitically integrated 2 layers of titanium-oxide memristive crossbars. Then we performed the analog dot-product operation under different input conditions in two scenarios: (1) with devices within the same crossbar layer and (2) with devices from different layers. In both cases, the devices exhibited low voltage operation and analog switching behavior with high tuning accuracy. Miguel Angel Lastras-Montaño, Bhaswar Chakrabarti, Dmitri B. Strukov, Kwang-Ting Cheng |
DATE | 4 |
| 2017 | Compact modeling and circuit-level simulation of silicon nanophotonic interconnectsabstractNanophotonic interconnects have been playing an increasingly important role in the datacom regime. Greater integration of silicon photonics demands modeling and simulation support for design validation, optimization and design space exploration. In this work, we develop compact models for a number of key photonic devices, which are extensively validated by the measurement data of a fabricated optical network-on-chip (ONoC). Implemented in SPICE-compatible Verilog-A, the models are used in circuit-level simulations of full optical links. The simulation results match well with the measurement data. Our model library and simulation approach enable the electro-optical (EO) co-simulation, allowing designers to include photonic devices in the whole system design space, and to co-optimize the transmitter, interconnect, and receiver jointly. Rui Wu 0008, Yuyang Wang 0003, Clint Schow, John E. Bowers 0001, Kwang-Ting Cheng |
DATE | 7 |
| 2017 | Mining mutation testing simulation traces for security and testbench debuggingabstractUnspecified design functionality can be modified by Hardware Trojans to leak information. Existing methods capable of detecting these Trojans require that unspecified functionality already be characterized, and suggest a manual ad-hoc process to enumerate “don't care” conditions potentially containing security vulnerabilities. Prior work has shown the potential of mutation testing to uncover testbench holes and highlight unspecified functionality, but requires tedious manual analysis of undetected faults to gain useful insight. This work provides the missing link required to fully automate characterization of unspecified functionality and can formally prove the absence of Trojans. Our approach is to mine simulation traces generated during mutation testing to produce assertions characterizing verification holes or unspecified functionality. These assertions can be fed directly to Trojan detection methods making securing unspecified functionality a completely automated process. Our trace mining technique is able to identify unspecified Wishbone bus functionality in a Trojan-free UART core and verify the functionality is benign, while flagging the same functionality in a Trojan-infected version of the design. Nicole Fern, Kwang-Ting Cheng |
ICCAD | 2 |
| 2017 | REDBEE: A visual-inertial drone system for real-time moving object detectionabstractAerial surveillance and monitoring demand both real-time and robust motion detection from a moving camera. Most existing techniques for drones involve sending a video data streams back to a ground station with a high-end desktop computer or server. These methods share one major drawback: data transmission is subjected to considerable delay and possible corruption. Onboard computation can not only overcome the data corruption problem but also increase the range of motion. Unfortunately, due to limited weight-bearing capacity, equipping drones with computing hardware of high processing capability is not feasible. Therefore, developing a motion detection system with real-time performance and high accuracy for drones with limited computing power is highly desirable. In this paper, we propose a visual-inertial drone system for real-time motion detection, namely REDBEE, that helps overcome challenges in shooting scenes with strong parallax and dynamic background. REDBEE, which can run on the state-of-the-art commercial low-power application processor (e.g. Snapdragon Flight board used for our prototype drone), achieves real-time performance with high detection accuracy. The REDBEE system overcomes obstacles in shooting scenes with strong parallax through an inertial-aided dual-plane homography estimation; it solves the issues in shooting scenes with dynamic background by distinguishing the moving targets through a probabilistic model based on spatial, temporal, and entropy consistency. The experiments are presented which demonstrate that our system obtains greater accuracy when detecting moving targets in outdoor environments than the state-of-the-art real-time onboard detection systems. Chong Huang 0005, Peng Chen 0008, Xin Yang 0008, Kwang-Ting Cheng |
IROS | 4 |
| 2017 | Robust design and design automation for flexible hybrid electronicsabstractFlexible electronics is promising for a number of emerging applications such as foldable smartphone, wearables and internet of things (IoT) [1], [2], [3]. However, the key elements of flexible electronics, the thin-film transistors (TFT), often suffer from large process variations and inferior reliability. There is also a lack of trustworthy compact models for these devices. This paper gives an overview of design challenges for the flexible circuits, introduces a robust design style, Pseudo-CMOS, that has been widely used for digital TFT designs, and highlights the development of a flexible hybrid electronics process design kit (FHE-PDK) supporting design automation and verification of flexible hybrid electronics. Tsung-Ching Huang, Leilai Shao, Raymond G. Beausoleil, Zhenan Bao, Kwang-Ting Cheng |
ISCAS | 6 |
| 2017 | Joint Detection and Diagnosis of Prostate Cancer in Multi-parametric MRI Based on Multimodal Convolutional Neural Networks
Xin Yang 0008, Zhiwei Wang 0002, Chaoyue Liu 0002, Hung Le Minh, Kwang-Ting Cheng, Liang Wang 0052 |
MICCAI (3) | 6 |
| 2017 | An Automatic Functional Coverage for Digital Systems Through a Binary Particle Swarm Optimization Algorithm with a Reinitialization Mechanism
Alfonso Martínez-Cruz, Ricardo Barrón, Herón Molina Lozano, Marco A. Ramírez 0001, Luis A. Villa-Vargas, Prometeo Cortés-Antonio, Kwang-Ting Cheng |
J. Electron. Test. | 7 |
| 2017 | Co-trained convolutional neural networks for automated detection of prostate cancer in multi-parametric MRI
Xin Yang 0008, Chaoyue Liu 0002, Zhiwei Wang 0002, Hung Le Minh, Liang Wang 0052, Kwang-Ting Cheng |
Medical Image Anal. | 7 |
| 2017 | Hiding Hardware Trojan Communication Channels in Partially Specified SoC Bus FunctionalityabstractOn-chip bus implementations must be bug-free and secure to provide the functionality and performance required by modern system-on-a-chip (SoC) designs. Regardless of the specific topology and protocol, bus behavior is never fully specified, meaning there exist cycles/conditions where some bus signals are irrelevant, and ignored by the verification effort. We highlight the susceptibility of current bus implementations to Hardware Trojans hiding in this partially specified behavior, and present a model for creating a covert Trojan communication channel between SoC components for any bus topology and protocol. By only altering existing bus signals during the period where their behaviors are unspecified, the Trojan channel is very difficult to detect. We give Trojan channel circuitry specifics for AMBA AXI4 and advanced peripheral bus (APB), then create a simple system comprised of several master and slave units connected by an AXI4-Lite interconnect to quantify the overhead of the Trojan channel and illustrate the ability of our Trojans to evade a suite of protocol compliance checking assertions from ARM. We also create an SoC design running a multiuser Linux OS to demonstrate how a Trojan communication channel can allow an unprivileged user access to root-user data. We then outline several detection strategies for this class of Hardware Trojan. Nicole Fern, Ismail San, Çetin Kaya Koç, Kwang-Ting Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2016 | Hardware Trojans in incompletely specified on-chip bus systems
Nicole Fern, Ismail San, Çetin Kaya Koç, Kwang-Ting Cheng |
DATE | 4 |
| 2016 | Trojans modifying soft-processor instruction sequences embedded in FPGA bitstreamsabstractReconfigurable platforms such as FPGAs and CPLDs are used to implement flexible and lightweight embedded systems often using soft-processors and a fixed instruction sequence stored in block memories. The bitstream format is proprietary for most vendors, however, in this work we demonstrate how to identify and extract block memory contents within the bitstream, allowing an adversary to learn and possibly modify the fixed instruction sequence. Manipulating the instruction sequence by inserting a Trojan in the bitstream as opposed to in the RTL code allows an adversary to bypass many verification steps. Moreover, the proposed Trojans only add extra instructions to the sequence to leak secret information, and do not change the original program behavior, making them virtually impossible to detect using functional tests. We present a case study where a Trojan is injected into a MIPS AES encryption program to leak internal state information by adding extra instructions from the available ones without changing the original program behavior. Ismail San, Nicole Fern, Çetin Kaya Koç, Kwang-Ting Cheng |
FPL | 4 |
| 2016 | A low-power hybrid reconfigurable architecture for resistive random-access memoriesabstractAccess-transistor-free memristive crossbars have shown to be excellent candidates for next generation non-volatile memories. While the elimination of the transistor per memory element enables higher memory densities, it also introduces parasitic currents during the normal operation of the memory that increases both the overall power consumption of the crossbar, and the current requirements of the line drivers. In this work we present a hybrid reconfigurable memory architecture that takes advantage of the fact that a complementary resistive switch (CRS) can behave both as a memristor and as a CRS. By dynamically keeping frequently accessed regions of the memory in the memristive mode and others in the CRS mode, our hybrid memory offer all the benefits that a memristor and a CRS offer individually, without any of their drawbacks. We validate our architecture using the SPEC CPU2006 benchmark and found that our hybrid memory offers average energy savings of 3.6x with respect to a memristive-only memory. In addition, we can offer a memory lifetime that is, on average, 6.4x longer than that of a CRS-only memory. Miguel Angel Lastras-Montaño, Amirali Ghofrani, Kwang-Ting Cheng |
HPCA | 3 |
| 2016 | Process-variation tolerant flexible circuit for wearable electronicsabstractFlexible electronics is a promising technology for wearable applications. The flexible printed thin-film transistor (TFT) circuits, however, often suffer from large process variations and inferior long-term reliability. This paper describes recent progress on robust printed circuits, including a novel design style known as Pseudo-CMOS, which was invented [1] to tackle design challenges for flexible TFT circuits. With post-fabrication tuning capability, Pseudo-CMOS can survive large process variations which are inevitable for the low-cost printing process. Furthermore, degradation of circuit performance, either due to variations or aging, can be recovered by means of an external tuning circuit. The paper also illustrates some design examples of Pseudo-CMOS circuits for applications to energy [2], healthcare [3], biomedical [4], and near-field communication (NFC) tags [5], [6]. Tsung-Ching Huang, Kwang-Ting Cheng, Raymond G. Beausoleil |
ISCAS | 2 |
| 2016 | In-place Repair for Resistive Memories Utilizing Complementary Resistive SwitchesabstractRecent advances in resistive memory technologies have demonstrated their potential to serve as next generation random access memories (RAM) which are fast, low-power, ultra-dense, and nonvolatile. However, owing to their stochastic filamentary nature, several sources of hard errors exist that could affect the lifetime of a resistive RAM (ReRAM). Amirali Ghofrani, Miguel Angel Lastras-Montaño, Yuyang Wang 0003, Kwang-Ting Cheng |
ISLPED | 4 |
| 2016 | Variation and failure characterization through pattern classification of test data from multiple test stagesabstractWe describe a framework for characterizing systematic variations and failures through exploring the hidden patterns of test data from multiple test stages. The framework provides prediction of process variations with a fine resolution based on a limited number of probed process parameters. An unsupervised biclustering technique is then utilized to extract grayscale and binary spatial patterns from process parameters and production test results, respectively, through analyzing both item-to-item and die-to-die correlations in subsets of the test data. A template matching technique exploits these spatial patterns to discover connections between process variations and failures detected by production tests. The proposed framework has been verified by an industrial test dataset of a non-volatile memory product. The discovery of comprehensible correlations between process parameters and some production test items was confirmed by the engineers who have insights to the test dataset. Chun-Kai Hsu, Peter Sarson, Gregor Schatzberger, Friedrich Peter Leisenberger, John M. Carulli Jr., Siddhartha Siddhartha, Kwang-Ting Cheng |
ITC | 7 |
| 2016 | OGB: A Distinctive and Efficient Feature for Mobile Augmented Reality
Xin Yang 0008, Xinggang Wang, Kwang-Ting Cheng |
MMM (1) | 3 |
| 2016 | Printed circuits on flexible substrates: opportunities and challenges (invited paper)abstractPrinted electronics (PE) on flexible substrates is a promising technology for wearables and internet of things (IoT). To implement an integrated system on flexible substrates for applications ranging from medical imaging to disposable thermometers, robust design methodology based on unreliable printed components plays a critical role. This paper reviews robust design of printed circuits based on thin-film transistors (TFT). A novel design style known as Pseudo-CMOS [1], which can tackle several design challenges of TFT circuits, is also introduced. Design examples of Pseudo-CMOS circuits for applications in energy [2], healthcare [3], biomedical [4], and near-field communication (NFC) tags [5], [6] are discussed. Tsung-Ching Huang, Kwang-Ting Cheng, Raymond G. Beausoleil |
NOCS | 2 |
| 2016 | Accurate and efficient pulse measurement from facial videos on smartphonesabstractNon-contact measurement of cardiac pulse signals has attracted high interests due to its convenience and cost effectiveness. However, extracting pulse signals on mobile handheld devices (e.g. smartphones) based on face videos captured by mobile cameras usually suffers from low measurement accuracy due to misalignment errors in face tracking and inevitable illumination changes in a mobile scenario, and low efficiency due to a handheld's limited computing power. We propose two techniques to address these limitations: 1) an accurate and efficient face tracking method based on an Active Shape Model (ASM) and the LDB (Local Difference Binary) feature description; 2) an adaptive temporal filtering method which can detect, and in turn denoise, sharp intensity changes in the source trace. Experimental results demonstrate that the proposed solution can achieve a speedup of 6.2X and is robust to noises in common mobile scenarios. Chong Huang 0005, Xin Yang 0008, Kwang-Ting Cheng |
WACV | 3 |
| 2016 | Renal compartment segmentation in DCE-MRI images
Xin Yang 0008, Hung Le Minh, Kwang-Ting Cheng, Kyung Hyun Sung, Wenyu Liu 0001 |
Medical Image Anal. | 3 |
| 2016 | An Efficient Network-on-Chip Yield Estimation Approach Based on Gibbs SamplingabstractA network-on-chip (NoC), a redundancy-rich and thus relatively robust system-chip, is still vulnerable to defects due to its large-scale integration. Thus, it is desirable to analyze the NoC yield in an early design phase. A Monte Carlo (MC) approach was proposed for the NoC yield analysis at the system level; however, it is inefficient due to the requirement of a large number of simulation runs. In this paper, we propose a Gibbs sampling approach, which can efficiently generate failed NoC instances as simulation samples, for yield estimation. This approach significantly reduces the number of required simulation runs for obtaining an accurate yield estimation. Implementation issues, such as initial sample selection, calculation of conditional distributions, and stop criterion, to customize Gibbs sampling for the NoC yield analysis are discussed. Potential optimization opportunities to further improve Gibbs sampling's efficiency are also explored. Compared to the MC approach, our experimental results show that the proposed approach can reduce the simulation runtime by 5×-100× for a high-yield NoC (a failure rate at 10-2-10-5), while achieving the same level of accuracy for yield estimation. Fan Lan, Kwang-Ting Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2015 | Toward large-scale access-transistor-free memristive crossbarsabstractMemristive crossbars have been shown to be excellent candidates for building an ultra-dense memory system because a per-cell access-transistor may no longer be necessary. However, the elimination of the access-transistor introduces several parasitic effects due to the existence of partially-selected devices during memory accesses, which could limit the scalability of access-transistor-free (ATF) memristive crossbars. In this paper we discuss these challenges in detail and describe some solutions addressing these challenges at multiple levels of design abstraction. Amirali Ghofrani, Miguel Angel Lastras-Montaño, Kwang-Ting Cheng |
ASP-DAC | 3 |
| 2015 | Hardware Trojan detection using exhaustive testing of k-bit subspacesabstractPost-silicon hardware Trojan detection is challenging because the attacker only needs to implement one of many possible design modifications, while the verification effort must guarantee the absence of all imaginable malicious circuitry. Existing test generation strategies for Trojan detection use controllability and observability metrics to limit the modifications targeted. However, for cryptographic hardware, the n plaintext bits are ideal for an attacker to use in Trojan triggering because the size of n prohibits exhaustive testing, and all n bits have identical controllability, making it impossible to bias testing using existing methods. Our detection method addresses this difficult case by observing that an attacker can realistically only afford to use a small subset, k, of all n possible signals for triggering. By aiming to exhaustively cover all possible k subsets of signals, we guarantee detection of Trojans using less than k plaintext bits in the trigger. We provide suggestions on how to determine k, and validate our approach using an AES design. Nicole Lesperance, Shrikant Kulkarni, Kwang-Ting Cheng |
ASP-DAC | 3 |
| 2015 | HReRAM: a hybrid reconfigurable resistive random-access memory
Miguel Angel Lastras-Montaño, Amirali Ghofrani, Kwang-Ting Cheng |
DATE | 3 |
| 2015 | Approximate associative memristive memory for energy-efficient GPUs
Abbas Rahimi, Amirali Ghofrani, Kwang-Ting Cheng, Luca Benini, Rajesh K. Gupta 0001 |
DATE | 3 |
| 2015 | Detecting Hardware Trojans in Unspecified Functionality Using Mutation TestingabstractExisting functional Trojan detection methodologies assume Trojans violate the design specification under carefully crafted rare triggering conditions. We present a new type of Trojan that leaks secret information from the design by only modifying unspecified functionality, meaning the Trojan is no longer restricted to being active only under rare conditions. We provide a method based on mutation testing for detecting this new Trojan type along with mutant ranking heuristics to prioritize analysis of the most dangerous functionality. Applying our method to a UART controller design, we discover unspecified and untested bus functionality with the potential to leak 32 bits of information during hundreds of cycles without being detected! Our method also reveals poorly tested interrupt functionality with information leakage potential. After modifying the specification and test bench to remove the discovered vulnerabilities, we close the verification loop by re-analyzing the design using our methodology and observe the functionality is no longer flagged as dangerous. Nicole Fern, Kwang-Ting Cheng |
ICCAD | 2 |
| 2015 | Pairwise Proximity-Based Features for Test Escape ScreeningabstractTest escapes are chips that pass the chip-level test program but fail system-level test or in the field. It is known that statistical analysis based on chip production test data could identify abnormalities for screening test escapes. It has also been shown that from the chip test data, we can generate revealing features for statistical analysis by comparing the measurement data to different references such as the measurement mean of a wafer, the spatial pattern of a wafer, and the measurements of neighboring chips. Given these existing features as the base features, this paper proposes a new class of transformations which could generate additional informative features based on pairwise proximities between chips on the same wafer. Specifically, we apply multiple distance functions in a feature space composed of the base features and calculate the corresponding pairwise proximities between each pair of chips. Each of the resulting proximities could potentially embed some unique information that reveals the abnormalities of some test escapes. Then we convert the proximities into Euclidean vector spaces using constant shift embedding (CSE), which preserves the cluster structure through the conversion, so that traditional outlier analysis algorithms such as local outlier factor (LOF) can be applied. The LOF value and the first dimension in each embedded space are used as additional features for each sample. These new features, jointly analyzed with the base features, provide more revealing information about test escapes and thus further improve the test escape detection rate in our experiment based on production test data. Fan Lin, Chun-Kai Hsu, Alberto Giovanni Busetto, Kwang-Ting Cheng |
ICCAD | 4 |
| 2015 | Variation-Aware Adaptive Tuning for Nanophotonic InterconnectsabstractShort-reach nanophotonic interconnects are promising to solve the communication bottleneck in data centers and chip-level scenarios. However, the nanophotonic interconnects are sensitive to process and thermal variations, especially for the microring structures, resulting in significant variation of an optical link's bit error rate (BER). In this paper, we propose a power-efficient adaptive tuning approach for nanophotonic interconnects to address the variation issues. During the adaptive tuning process, each nanophotonic interconnect is adaptively allocated just enough power to meet the BER requirement. The proposed adaptive tuning approach could reduce the photonic receiver power by 8% - 34% than the worst-case based fixed design while achieving the same BER. Our evaluation results show that the adaptive tuning approach scales well with the process variation, the thermal variation and the number of communication nodes, and can accommodate different types of NoC architectures and lasers. Rui Wu 0008, Chin-Hui Chen, Tsung-Ching Huang, Fan Lan, John E. Bowers 0001, Raymond G. Beausoleil, Kwang-Ting Cheng |
ICCAD | 10 |
| 2015 | A configurable CMOS memory platform for 3D-integrated memristorsabstractMemristors are emerging as powerful nanoscale devices for diverse applications, such as high-density memories and neuromorphic applications. However, this nascent technology requires considerable advancement before this vision is realized. We present a highly configurable CMOS interface chip which enables the characterization of on-chip memristors, especially for memory applications. The chip was fabricated in On-Semi 3M2P 0.5 μm occupying 2×2 mm2. The chip design allows for post-CMOS fabrication of memristors. The interface between the memristor and the CMOS circuitry was provided via a top metal contact. The chip was designed to support an area-distributed interface decoupling CMOS pitch and memristor pitch, enabling high-density memristor integration. Measurement results on post-CMOS fabricated Ag/SiO2/Pt memristive devices are reported. Though we have shown the results from one memristive material stack, thorough chip characterization demonstrates the versatility of the chip enabling its use with a wide variety of materials stacks. Melika Payvand, Advait Madhavan, Miguel Angel Lastras-Montaño, Amirali Ghofrani, Justin Rofeh, Kwang-Ting Cheng, Dmitri B. Strukov, Luke Theogarajan |
ISCAS | 6 |
| 2015 | Fusion of Vision and Inertial Sensing for Accurate and Efficient Pose Tracking on SmartphonesabstractThis paper aims at accurate and efficient pose tracking of planar targets on modern smartphones. Existing methods, relying on either visual features or motion sensing based on built-in inertial sensors, are either too computationally expensive to achieve realtime performance on a smartphone, or too noisy to achieve sufficient tracking accuracy. In this paper we present a hybrid tracking method which can achieve real-time performance with high accuracy. Based on the same framework of a state-of-the-art visual feature tracking algorithm [5] which ensures accurate and reliable pose tracking, the proposed hybrid method significantly reduces its computational cost with the assistance of a phone's built-in inertial sensors. However, noises in inertial sensors and abrupt errors in feature tracking due to severe motion blurs could result in instability of the hybrid tracking system. To address this problem, we propose to employ an adaptive Kalman filter with abrupt error detection to robustly fuse the inertial and feature tracking results. We evaluated the proposed method on a dataset consisting of 16 video clips with synchronized inertial sensing data. Experimental results demonstrated our method's superior performance and accuracy on smartphones compared to a state-of-the-art vision tracking method [5]. The dataset will be made publicly available with the publication of this paper. Xin Yang 0008, Xun Si, Tangli Xue, Kwang-Ting Cheng |
ISMAR | 4 |
| 2015 | Hardware Trojans hidden in RTL don't cares - Automated insertion and prevention methodologiesabstractDon't cares in RTL code have long plagued chip verification due to hard-to-diagnose “X-bugs” resulting from ambiguous X simulation semantics, yet prevail in modern designs because of enormous opportunities for area/performance/power optimization during synthesis. We analyze don't cares specified at the RTL level from a security perspective and propose a novel class of Hardware Trojans which leak internal circuit node values using only existing design don't cares. Detection of this Trojan class is impossible using either functional simulation/verification or a perfect sequential equivalence checker. We then provide a formal automated X-analysis technique which both prevents the insertion of this new Trojan type and also has the potential to uncover accidental X-bugs as well. We provide several examples, including an Elliptic Curve Processor, illustrating both Trojan insertion and our prevention technique. Nicole Fern, Shrikant Kulkarni, Kwang-Ting Cheng |
ITC | 3 |
| 2015 | AdaTest: An efficient statistical test framework for test escape screeningabstractStatistical analyses based on production test data can help identify test escapes, which are chips that pass the test program but fail later at system-level test or in field. Such analyses do not require extra physical measurements and can be referred to as statistical tests. For designing effective statistical tests, this paper investigates the use of a learning framework based on Adaptive Boosting, which has demonstrated great success in real-time face and object recognition. The framework is composed of a cascade of AdaBoost classifiers, each of which uses a small set of most relevant features that are automatically selected in the training phase, to identify a subset of test escapes. This framework therefore generates only the features that are most relevant for classification and significantly reduces the runtime and memory usage for statistical tests during test application. We also propose a new feature set to characterize the chips under test and demonstrate that including the new feature set as input to the proposed feature selection framework could reveal more test escapes. Fan Lin, Chun-Kai Hsu, Kwang-Ting Cheng |
ITC | 3 |
| 2015 | Automatic Segmentation of Renal Compartments in DCE-MRI Images
Xin Yang 0008, Hung Le Minh, Kwang-Ting Cheng, Kyung Hyun Sung, Wenyu Liu 0001 |
MICCAI (1) | 3 |
| 2015 | Vision-Inertial Hybrid Tracking for Robust and Efficient Augmented Reality on SmartphonesabstractThis paper aims at robust and efficient pose tracking for augmented reality on modern smartphones. Existing methods, relying on either vision analysis or motion sensing, are either too computationally expensive to achieve real-time performance on a smartphone, or too noisy to achieve sufficient robustness. This paper presents a hybrid tracking system which can achieve real-time performance with high robustness. Our system utilizes an efficient featureless method based on pixel-based registration to track the object pose on every frame. The featureless tracking result is revised from time to time by a feature-based method to reduce tracking errors. Both featureless and feature-based tracking results are sensitive to large motion blurs. To improve the robustness, an adaptive Kamlan filter is proposed to fuse the visual tracking results with the inertial tracking results computed form phone's built-in sensors. Our hybrid method is evaluated on a dataset consisting of 16 video clips with synchronized inertial sensing data. Experimental results demonstrated the superior performance of our method to state-of-the-art visual tracking methods [5, 12] on smartphones. The dataset will be made publicly available with the publication of this paper. Xin Yang 0008, Xun Si, Tangli Xue, Liheng Zhang, Kwang-Ting Cheng |
ACM Multimedia | 5 |
| 2015 | A Low-Power Variation-Aware Adaptive Write Scheme for Access-Transistor-Free Memristive MemoryabstractRecent advances in access-transistor-free memristive crossbars have demonstrated the potential of memristor arrays as high-density and ultra-low-power memory. However, with considerable variations in the write-time characteristics of individual memristors, conventional fixed-pulse write schemes cannot guarantee reliable completion of the write operations and waste significant amount of energy. We propose an adaptive write scheme that adaptively adjusts the write pulses to address such variations in memristive arrays, resulting in 7×--11× average energy saving in our case studies. Our scheme embeds an online monitor to detect the completion of a write operation and takes into account the parasitic effect of line-shared devices in access-transistor-free crossbars. This feature also helps shorten the test time of memory march algorithms by eliminating the need of a verifying read right after a write, which is commonly employed in the test sequences of march algorithms. Amirali Ghofrani, Miguel Angel Lastras-Montaño, Siddharth Gaba, Melika Payvand, Wei Lu 0003, Luke Theogarajan, Kwang-Ting Cheng |
ACM J. Emerg. Technol. Comput. Syst. | 7 |
| 2014 | Accurate Vessel Segmentation with Progressive Contrast Enhancement and Canny Refinement
Xin Yang 0008, Kwang-Ting Cheng, Aichi Chien |
ACCV (3) | 2 |
| 2014 | Learning from Production Test Data: Correlation Exploration and Feature EngineeringabstractThe huge amount of test data of a modern chip produced during manufacturing test could be mined for valuable information about the device under test (DUT), far more than the pass/fail information of each test item. Exploring the hidden correlations and patterns in the test data allows better understanding of the DUT and could therefore lead to test cost reduction or test quality improvement. There are several known types of correlations embedded in the test data: spatial correlations, inter-test-item correlations, and temporal correlations, each of which may involve a large number of data dimensions. Deriving and selecting the most relevant features for a specific application is critical for designing an effective and efficient mining solution. This paper provides an overview of recent research efforts on correlation exploration and development of a framework of feature engineering for learning from production test data. Fan Lin, Chun-Kai Hsu, Kwang-Ting Cheng |
ATS | 3 |
| 2014 | Energy-Efficient GPGPU Architectures via Collaborative Compilation and Memristive Memory-Based ComputingabstractThousands of deep and wide pipelines working concurrently make GPGPU high power consuming parts. Energy-efficiency techniques employ voltage overscaling that increases timing sensitivity to variations and hence aggravating the energy use issues. This paper proposes a method to increase spatiotemporal reuse of computational effort by a combination of compilation and micro-architectural design. An associative memristive memory (AMM) module is integrated with the floating point units (FPUs). Together, we enable fine-grained partitioning of values and find high-frequency sets of values for the FPUs by searching the space of possible inputs, with the help of application-specific profile feedback. For every kernel execution, the compiler pre-stores these high-frequent sets of values in AMM modules -- representing partial functionality of the associated FPU-- that are concurrently evaluated over two clock cycles. Our simulation results show high hit rates with 32-entry AMM modules that enable 36% reduction in average energy use by the kernel codes. Compared to voltage overscaling, this technique enhances robustness against timing errors with 39% average energy saving. Abbas Rahimi, Amirali Ghofrani, Miguel Angel Lastras-Montaño, Kwang-Ting Cheng, Luca Benini, Rajesh K. Gupta 0001 |
DAC | 4 |
| 2014 | Joint Virtual Probe: Joint exploration of multiple test items' spatial patterns for efficient silicon characterization and test predictionabstractVirtual Probe (VP), proposed for characterization of spatial variations and for test time reduction, can effectively reconstruct the spatial pattern of a test item for an entire wafer using measurement values from only a small fraction of dies on the wafer. However, VP calculates the spatial signature of each test item separately, one item at a time, resulting in very long runtime for complex chips which often require hundreds, or even thousands, of test items in production. In this paper, we propose a new method, named Joint Virtual Probe (JVP), which can jointly derive spatial patterns of multiple test items. By simultaneously handling a large group of test items, JVP significantly reduces the overall runtime. And the prediction accuracy can also be improved because of JVP's implicit use of inter-test-item correlations in predicting spatial patterns. The experimental results on two industrial products, with 277 and 985 parametric test items in the production test programs respectively, demonstrate that, JVP achieves an average speedup of ~ 170X and ~ 50X over VP in the pre-test analysis and the test application phases respectively, as well as a slightly higher prediction accuracy than VP. Shuangyue Zhang, Fan Lin, Chun-Kai Hsu, Kwang-Ting Cheng |
DATE | 4 |
| 2014 | Geodesic Active Contours with Adaptive Configuration for Cerebral Vessel and Aneurysm SegmentationabstractActive contour is a popular technique for vascular segmentation. However, existing active contour segmentation methods require users to set values for various parameters, which requires insights to the method's mathematical formulation. Manual tuning of these parameters to optimize segmentation results is laborious for clinicians who often lack in-depth knowledge of the segmentation algorithms. Moreover, a global parameter setting applied to all voxels of an input image can hardly achieve optimized results due to vessels' high appearance variability caused by the contrast agent in homogeneity and noises. In this paper, we present a method which adaptively configures parameters for Geodesic Active Contours (GAC). The proposed method leverages shape filtering to produce a parameter image, each voxel of which is used to set parameters of GAC for the corresponding voxel of an input image. An iterative process is further developed to improve the accuracy of the shape-based parameter image. An evaluation study over 8 clinical datasets demonstrates that our method achieves greater segmentation accuracy than two popular active contour methods with manually optimized parameters. Xin Yang 0008, Kwang-Ting Cheng, Aichi Chien |
ICPR | 2 |
| 2014 | Feature engineering with canonical analysis for effective statistical tests screening test escapesabstractIt is known that statistical analysis of test data can help screen potential test escapes without additional physical measurements. Based on analysis of production test data, this paper focuses on feature engineering for statistical tests to screen test escapes. The features are engineered in two aspects: development of effective features and transformation of features into a different space in which the inherent difference between the test escapes and the normal population can be compacted into a small number of features. In feature development, we generate two sets of features to characterize a chip based on the amounts of the chip's test measurements deviated from the measurement means and their amounts deviated from the spatial patterns among dies on the same wafer. In feature transformation, the features are projected into the canonical space, in which the separation between the test escapes and the good chips are encapsulated into the first few dimensions. We show that each set of features reveals a unique set of test escapes, and the transformation of features can result in significant runtime reduction while keeping a comparable differentiating power as that in the original features. Therefore, both sets of features should be utilized and the canonical transformation should be applied when developing statistical tests for test escape reduction. Fan Lin, Chun-Kai Hsu, Kwang-Ting Cheng |
ITC | 3 |
| 2014 | libLDB: a library for extracting ultrafast and distinctive binary feature descriptionabstractThis paper gives an overview of libLDB -- a C++ library for extracting an ultrafast and distinctive binary feature LDB (Local Difference Binary) from an image patch. LDB directly computes a binary string using simple intensity and gradient difference tests on pairwise grid cells within the patch. Relying on integral images, the average intensity and gradients of each grid cell can be obtained by only 4~8 add/subtract operations, yielding an ultrafast runtime. A multiple gridding strategy is applied to capture the distinct patterns of the patch at different spatial granularities, leading to a high distinctiveness of LDB. LDB is very suitable for vision apps which require real-time performance, especially for apps running on mobile handheld devices, such as real-time mobile object recognition and tracking, markerless mobile augmented reality, mobile panorama stitching. This software is available under the GNU General Public License (GPL) v3. Xin Yang 0008, Chong Huang 0005, Kwang-Ting Cheng |
ACM Multimedia | 3 |
| 2014 | Local Difference Binary for Ultrafast and Distinctive Feature DescriptionabstractThe efficiency and quality of a feature descriptor are critical to the user experience of many computer vision applications. However, the existing descriptors are either too computationally expensive to achieve real-time performance, or not sufficiently distinctive to identify correct matches from a large database with various transformations. In this paper, we propose a highly efficient and distinctive binary descriptor, called local difference binary (LDB). LDB directly computes a binary string for an image patch using simple intensity and gradient difference tests on pairwise grid cells within the patch. A multiple-gridding strategy and a salient bit-selection method are applied to capture the distinct patterns of the patch at different spatial granularities. Experimental results demonstrate that compared to the existing state-of-the-art binary descriptors, primarily designed for speed, LDB has similar construction efficiency, while achieving a greater accuracy and faster speed for mobile object recognition and tracking tasks. Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Learning Optimized Local Difference Binaries for Scalable Augmented Reality on Mobile DevicesabstractThe efficiency, robustness and distinctiveness of a feature descriptor are critical to the user experience and scalability of a mobile augmented reality (AR) system. However, existing descriptors are either too computationally expensive to achieve real-time performance on a mobile device such as a smartphone or tablet, or not sufficiently robust and distinctive to identify correct matches from a large database. As a result, current mobile AR systems still only have limited capabilities, which greatly restrict their deployment in practice. In this paper, we propose a highly efficient, robust and distinctive binary descriptor, called Learning-based Local Difference Binary (LLDB). LLDB directly computes a binary string for an image patch using simple intensity and gradient difference tests on pairwise grid cells within the patch. To select an optimized set of grid cell pairs, we densely sample grid cells from an image patch and then leverage a modified AdaBoost algorithm to automatically extract a small set of critical ones with the goal of maximizing the Hamming distance between mismatches while minimizing it between matches. Experimental results demonstrate that LLDB is extremely fast to compute and to match against a large database due to its high robustness and distinctiveness. Compared to the state-of-the-art binary descriptors, primarily designed for speed, LLDB has similar efficiency for descriptor construction, while achieving a greater accuracy and faster matching speed when matching over a large database with 2.3M descriptors on mobile devices. Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2014 | Compact Test Generation With an Influence Input Measure for Launch-On-Capture Transition Fault TestingabstractWe propose a compact test generation method for transition faults based on a conflict avoidance driven scheme. A new measure is proposed to estimate the influence inputs, which is the subset of inputs to be specified, needed for detecting transition faults. The value requirements at the pseudoprimary inputs (PPIs) of the second frame of the automatic test pattern generation circuit model are partitioned into separate subsets. The sequential backtracing scheme backtraces the value requirements on the PPIs of the second frame subset-by-subset. Justification of the necessary value requirements at the PPIs in the second frame is completed using a conflict-driven procedure. With an influence input measure for transition faults under the launch-on-capture scan testing application scheme, a new dynamic test compaction scheme is proposed. The new test compaction scheme tries to compact as many faults as possible into the current test. Experimental results and comparison with existing approaches demonstrate the efficiency and effectiveness of the proposed method. Wenjie Sui, Boxue Yin, Kwang-Ting Cheng |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2014 | Test-Quality Optimization for Variable $n$ -Detections of Transition FaultsabstractAggressive technology scaling in modern chips resulted in complicated faulty timing behaviors, which necessitate undesirable long development cycle and high test volumes to ensure product quality. To reduce the test time, cost-effective and timing-efficient test selection algorithms are used to choose optimal test inputs from a large-volume test set. In this paper, we define an approximate longest sensitized path (ALSP) metric to derive the longest sensitized path for all transition faults (TFs) from the detectability of TFs with very low computational complexity. With the ALSP metric, a general public utilities-based parallel test selection method is proposed to choose a small test set with high delay test quality from the timing-unaware n-detection test set. Our results demonstrate the comparison with a commercial automatic test pattern generation tool and a previous timing-aware test selection method targeting small delay defects, and confirm that our test selection algorithm can achieve better delay test coverage and higher n -detection fault coverage with steeper fault coverage curves of ordered patterns, for the same pattern count. Dawen Xu 0002, Huawei Li 0001, Amirali Ghofrani, Kwang-Ting Cheng, Yinhe Han 0001, Xiaowei Li 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2013 | Mutation analysis with coverage discountingabstractMutation testing is an established technique for evaluating validation thoroughness, but its adoption has been limited by the manual effort required to analyze the results. This paper describes the use of coverage discounting for mutation analysis, where undetected mutants are explained in terms of functional coverpoints, simplifying their analysis and saving effort. Two benchmarks are shown to compare this improved flow against regular mutation analysis. We also propose a confidence metric and simulation ordering algorithm optimized for coverage discounting, potentially reducing overall simulation time. Peter Lisherness, Nicole Lesperance, Kwang-Ting Cheng |
DATE | 3 |
| 2013 | Towards data reliable crossbar-based memristive memoriesabstractA series of breakthroughs in memristive devices have demonstrated the potential of using crossbar-based memristor arrays as ultra-high-density and low-power memory. However, their unique device characteristics could cause data disturbance for both read and write operations resulting in serious data reliability problems. This paper discusses such reliability issues in detail and proposes a comprehensive yet low area-/performance-/energy-overhead solution addressing these problems. The proposed solution applies asymmetric voltages for disturbance confinement, inserts redundancy for disturbance detection, and employs a refreshing mechanism to restore weakened data. The results of a case study show that the average overheads of area, performance and energy consumption for achieving data reliability, over a baseline unreliable memory system, are 3%, 4%, and 19% respectively. Amirali Ghofrani, Miguel Angel Lastras-Montaño, Kwang-Ting Cheng |
ITC | 3 |
| 2013 | Test data analytics - Exploring spatial and test-item correlations in production test dataabstractThe discovery of patterns and correlations hidden in the test data could help reduce test time and cost. In this paper, we propose a methodology and supporting statistical regression tools that can exploit and utilize both spatial and inter-test-item correlations in the test data for test time and cost reduction. We first describe a statistical regression method, called group lasso, which can identify inter-test-item correlations from test data. After learning such correlations, some test items can be identified for removal from the test program without compromising test quality. An extended version of this method, weighted group lasso, allows taking into account the distinct test time/cost of each individual test item in the formulation as a weighted optimization problem. As a result, its solution would favor more costly test items for removal from the test program. We further integrate weighted group lasso with another statistical regression technique, virtual probe, which can learn spatial correlations of test data across a wafer. The integrated method could then utilize both spatial and inter-test-item correlations to maximize the number of test items whose values can be predicted without measurement. Experimental results of a high-volume industrial device show that utilizing both spatial and inter-test-item correlations can help reduce test time by up to 55%. Chun-Kai Hsu, Fan Lin, Kwang-Ting Cheng, Wangyang Zhang, Xin Li 0001, John M. Carulli Jr., Kenneth M. Butler |
ITC | 3 |
| 2013 | Low-Cost Error Tolerance Scheme for 3-D CMOS ImagersabstractThis paper presents an error tolerance scheme for 3-D CMOS imagers that are constructed by stacking a pixel array of imager sensors, an analog-to-digital converter (ADC) array, and an image signal processor (ISP) array using microbumps$(\mu{\rm bumps})$and through silicon vias (TSVs). To deliver high-quality images in the presence of single or multiple$\mu{\rm bump}$, ADC, or TSV failures, we propose to interleave the connections from pixels to ADCs and recover the corrupted data in the ISPs. Key design parameters, such as the interleaving stride and the grouping ratio are determined by analyzing the employed error correction algorithm. Architectural simulation results demonstrate that the error tolerance scheme enhances the effective yield of an exemplar 3-D imager from 44% to 97%. Hsiu-Ming Chang 0001, Jiun-Lang Huang, Ding-Ming Kwai, Kwang-Ting Cheng, Cheng-Wen Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2012 | Platform characterization for Domain-Specific ComputingabstractWe believe that by adapting architectures to fit the requirements of a given application domain, we can significantly improve the efficiency of computation. To validate the idea for our application domain, we evaluate a wide spectrum of commodity computing platforms to quantify the potential benefits of heterogeneity and customization for the domain-specific applications. In particular, we choose medical imaging as the application domain for investigation, and study the application performance and energy efficiency across a diverse set of commodity hardware platforms, such as general-purpose multi-core CPUs, massive parallel many-core GPUs, low-power mobile CPUs and fine-grain customizable FPGAs. This study leads to a number of interesting observations that can be used to guide further development of domain-specific architectures. Alex Bui, Kwang-Ting Cheng, Jason Cong, Luminita A. Vese, Yi-Chu Wang, Yi Zou 0001 |
ASP-DAC | 2 |
| 2012 | On error modeling of electrical bugs for post-silicon timing validationabstractThere is great demand for an accurate and scalable metric to evaluate the functional stimuli, testbench checkers, and DfD (Design-for-Debug) structures used in post-silicon timing validation. In this paper, we show the inadequacy of existing methods (due to either inaccuracy or a lack of scalability) and propose an approach that leverages debug engineers' experience to model timing errors efficiently and with sufficient precision. Experimental results demonstrate that the proposed approach produced an error model six times more accurate than the prior art with a negligible simulation overhead. Peter Lisherness, Kwang-Ting Cheng, Jing-Jia Liou |
ASP-DAC | 3 |
| 2012 | Improving validation coverage metrics to account for limited observabilityabstractIn both pre-silicon and post-silicon validation, the detection of design errors requires both stimulus capable of activating the errors and checkers capable of detecting the behavior as erroneous. Most functional and code coverage metrics evaluate only the activation component of the testbench and ignore propagation and detection. In this paper, we summarize our recent work in developing improved metrics that account for propagation and/or detection of design errors. These works include tools for observability-enhanced code coverage and mutation analysis of high-level designs as well as an analytical method, Coverage Discounting, which adds checker sensitivity to arbitrary functional coverage metrics. Peter Lisherness, Kwang-Ting Cheng |
ASP-DAC | 2 |
| 2012 | Post-fabrication reconfiguration for power-optimized tuning of optically connected multi-core systemsabstractIntegrating optical interconnects into the next-generation multi-/many-core architecture has been considered a viable solution to addressing the limitations in throughput, latency, and power efficiency of electrical interconnects. Optical interconnects also allow the performance growth of inter-core connectivity to keep pace with the growth of the cores' processing ability. However, variations in the fabrication process significantly impair an optical network's communication quality. Existing post-fabrication tuning methods, which are based on adjusting the voltages and temperatures, have very limited tunability and require excessive power to fully compensate for the variation. In this paper, we study the sources and severity of process variation, and propose two methods to enhance the robustness of an on-chip optical network: 1) adding spare modulators and detectors for post-fabrication reconfiguration and low-power tuning, and 2) introducing a combined detector/modulator structure for a more robust network topology. Simulation results show that employing both methods can reduce the tuning power from hundreds of watts to 6W while maintaining a throughput of 99.7%. To maintain a throughput of 50%, the tuning power can be further reduced to only 12mW. Peter Lisherness, Saeed Shamshiri, Amirali Ghofrani, Kwang-Ting Cheng |
ASP-DAC | 6 |
| 2012 | Power-efficient calibration and reconfiguration for on-chip optical communicationabstractOn-chip optical communication infrastructure has been proposed to provide higher bandwidth and lower power consumption for the next-generation high performance multicore systems. Online calibration of these optical components is essential to building a robust optical communication system, which is highly sensitive to process and thermal variation. However, the power consumption of existing tuning methods to properly calibrate the optical devices would be prohibitively high. We propose two calibration and reconfiguration techniques that can significantly reduce the worst case tuning power of a ring-resonator-based optical modulator: 1) a channel re-mapping scheme, with sub-channel redundant resonators, which results in significant reduction in the amount of required tuning, typically within the capability of voltage based tuning, and 2) a dynamic feedback calibration mechanism used to compensate for both process and thermal variations of the resonators. Simulation results demonstrate that these techniques can achieve a 48X reduction in tuning power - less than 10W for a network with 1-million ring resonators. Peter Lisherness, Jock Bovington, Kwang-Ting Cheng |
DATE | 6 |
| 2012 | LDB: An ultra-fast feature for scalable Augmented Reality on mobile devicesabstractThe efficiency, robustness and distinctiveness of a feature descriptor are critical to the user experience and scalability of a mobile Augmented Reality (AR) system. However, existing descriptors are either too compute-expensive to achieve real-time performance on a mobile device such as a smartphone or tablet, or not sufficiently robust and distinctive to identify correct matches from a large database. As a result, current mobile AR systems still only have limited capabilities, which greatly restrict their deployment in practice. In this paper, we propose a highly efficient, robust and distinctive binary descriptor, called Local Difference Binary (LDB). LDB directly computes a binary string for an image patch using simple intensity and gradient difference tests on pairwise grid cells within the patch. A multiple gridding strategy is applied to capture the distinct patterns of the patch at different spatial granularities. Experimental results demonstrate that LDB is extremely fast to compute and to match against a large database due to its high robustness and distinctiveness. Comparing to the state-of-the-art binary descriptor BRIEF, primarily designed for speed, LDB has similar computational efficiency, while achieves a greater accuracy and 5x faster matching speed when matching over a large database with 1.7M+ descriptors. Xin Yang 0008, Kwang-Ting Cheng |
ISMAR | 2 |
| 2012 | 3D CMOS-memristor hybrid circuits: devices, integration, architecture, and applicationsabstractIn this paper, we give an overview of our recent research efforts on monolithic 3D integration of CMOS and memristive nanodevices. These hybrid circuits combine a CMOS subsystem with several layers of nanowire crossbars, consisting of arrays of two-terminal memristors, all connected by an area-distributed interface between the CMOS subsystem and the crossbars. This approach combines the advantages of CMOS technology, including its high flexibility, functionality and yield, with the extremely high density of nanowires, nanodevices and interface vias. As a result, the 3D hybrids can overcome limitations pertinent to other 3D integration techniques (such as through-silicon vias) and enable 3D circuits with unprecedented memory density (up to 1014 bits on a single 1-cm2 chip) and aggregate interlayer communication bandwidth (up to 1018 bits per second per cm2) at manageable power dissipation. Such performance represents a significant step towards addressing the most pressing needs of modern compact electronic systems. Kwang-Ting Cheng, Dmitri B. Strukov |
ISPD | 1 |
| 2012 | Adaptive test selection for post-silicon timing validation: A data mining approachabstractTest failure data produced during post-silicon validation contain accurate design- and process-specific information about the DUD (design-under-debug). Prior research efforts and industry practice focused on feeding this information back to the design flow via bug root-cause analysis. However, the value of this silicon data for helping further improvement of the post-silicon validation process has been largely overlooked. In this paper, we propose an adaptive test selection method to progressively tune the validation plan using knowledge automatically mined from the bug sightings during post-silicon validation. Experimental results demonstrate that the proposed fault-model-free data mining approach can prioritize those tests capable of uncovering more silicon timing errors, resulting in significant reduction of validation time and effort. Peter Lisherness, Kwang-Ting Cheng |
ITC | 3 |
| 2012 | Accelerating SURF detector on mobile devicesabstractRunning a SURF (Speeded Up Robust Features) detector on mobile devices remains too slow to support emerging applications such as mobile augmented reality. Porting it without adapting the algorithm to account for mobile platform limitations could result in significant runtime degradation. In this paper, we identify two mismatches between the SURF algorithm and the mobile hardware that cause substantial slow-down of the point detection process: 1) mismatch between the data access pattern and the small cache size, and 2) mismatch between the huge amount of branches and high pipeline hazard penalty. To address the mismatches, we propose two techniques: tiled SURF and gradient moment based orientation assignment. Tiled SURF improves data locality and greatly reduces memory traffic. A method for determining the optimal tile sizes, named content-aware tiling, is designed to minimize runtime and maximize detection accuracy. To avoid the penalties caused by pipeline hazards, we replace the original orientation operator with branching-free gradient moment computations. The proposed techniques are tested on three mobile platforms. Comparing to the original SURF, the accelerated SURF achieves a 6x~8x speedup without sacrificing recognition accuracy. Meanwhile, it achieves 59%~80% reductions in the runtime ratio of the detector running on mobile platforms compared with on x86-based PCs. Xin Yang 0008, Kwang-Ting Cheng |
ACM Multimedia | 2 |
| 2012 | Comprehensive online defect diagnosis in on-chip networksabstractWe propose a comprehensive yet low-cost solution for online detection and diagnosis of permanent faults in on-chip networks. Using error syndrome collection and packet/flit-counting techniques, high-resolution defect diagnosis is feasible in both datapath and control logic of the on-chip network without injecting any test traffic or incurring significant performance overhead. Amirali Ghofrani, Ritesh Parikh, Saeed Shamshiri, Andrew DeOrio, Kwang-Ting Cheng, Valeria Bertacco |
VTS | 5 |
| 2011 | Post-silicon bug detection for variation induced electrical bugsabstractElectrical bugs, such as those caused by crosstalk or power droop, are a growing concern due to shrinking noise margins and increasing variability. This paper introduces COBE, an electrical bug modeling technique which can be used to evaluate the effectiveness of validation tests and DfD (design-for-debug) structures for detecting these errors in post-silicon validation. COBE first uses gate-level timing details to identify critical flip-flops in which the error effects of electrical bugs are more likely to be captured. Based on RTL simulation traces, the functional tests and corresponding cycles in which these critical flip-flops incur transitions are then recorded as the potential times and locations of bug activation. These selected “bit-flips” are then analyzed through functional simulation to determine if they are propagated to an observation point for detection. Compared to the commonly employed random bit-flip injection technique, COBE provides a significantly more accurate electrical bug model by taking into account the likelihood of bug activation, in terms of both location and time, for bit-flip injection. COBE is experimentally evaluated on an Alpha 21264 processor RTL model. In our simulation-based experiments, the results show that the relative effectiveness of the tests predicted by COBE correlates very well with the tests' electrical bug detection capability, with a correlation factor of 0.921. This method is much more accurate than the random bit-flip injection technique, which has a correlation factor of 0.482. Peter Lisherness, Kwang-Ting Cheng |
ASP-DAC | 3 |
| 2011 | Image quality aware metrics for performance specification of ADC array in 3D CMOS imagersabstractA three-dimensional (3D) CMOS imager constructed from stacking a pixel array of image sensors, an analog-to-digital converter (ADC) array, and an image signal processor (ISP) array is promising for high throughput imaging applications. The design specifications of the ADC array in the imager, which jointly and concurrently converts the pixel data to produce a final image, must consider both intra-ADC linearity and inter-ADC uniformity. In this paper, we investigate the relationship between the image quality and the linearity of individual ADCs as well as the uniformity of neighboring ADCs in the array. With the insights to this relationship, the specification requirements for the ADC array can be derived based on a desired level of image quality. Hsiu-Ming Chang 0001, Kwang-Ting Cheng |
DAC | 2 |
| 2011 | An all-digital built-in self-test technique for transfer function characterization of RF PLLsabstractThis paper presents an all-digital built-in self-test (BIST) technique for characterizing the error transfer function of RF PLLs. This BIST scheme, with on-chip stimulus synthesis and response analysis completely done in the digital domain, achieves high-accuracy characterization and is applicable to a wide range of PLL architectures. For the popular sigma-delta fractional-N RF PLLs, the added circuitry required for this BIST solution is all digital except a bang-bang phase-frequency detector (BB-PFD), which incurs an area of only 0.0001 mm2for our implementation in a 65 nm CMOS technology. The silicon characterization results at 3.6 GHz reported by this BIST solution and by explicit measurement have a root-mean-square difference of 0.375 dB only. Ping-Ying Wang, Hsiu-Ming Chang 0001, Kwang-Ting Cheng |
DATE | 3 |
| 2011 | Test cost reduction through performance prediction using virtual probeabstractThe virtual probe (VP) technique, based on recent breakthroughs in compressed sensing, has demonstrated its ability for accurate prediction of spatial variations from a small set of measurement data. In this paper, we explore its application to cost reduction of production testing. For a number of test items, the measurement data from a small subset of chips can be used to accurately predict the performance of other chips on the same wafer without explicit measurement. Depending on their statistical characteristics, test items can be classified into three categories: highly predictable, predictable, and un-predictable. A case study of an industrial RF radio transceiver with more than 50 production test items shows that a good fraction of these test items (39 out of 51 items) are predictable or highly predictable. In this example, the 3σ error of VP prediction is less than 12% for predictable or highly predictable test items. Applying the VP technique can on average replace 59% of test measurement by prediction and, consequently, reduce the overall test time by 57.6%. Hsiu-Ming Chang 0001, Kwang-Ting Cheng, Wangyang Zhang, Xin Li 0001, Kenneth M. Butler |
ITC | 2 |
| 2011 | End-to-end error correction and online diagnosis for on-chip networksabstractWe propose a comprehensive solution for end-to-end (e2e) error correction and online defect diagnosis for on-chip networks. For e2e error correction, we propose an interleaved error-locality-aware code that efficiently corrects both random and burst errors. We demonstrate that for 64-bit wide network links, interleaving four of the proposed code, 2G4L(26,16), each of which supports 16bit data, can correct as many as two random errors or 16 adjacent errors. In order to maintain the error correction capability of the Error Correcting Code (ECC) for transient and intermittent errors, we further propose an e2e data gathering and online diagnosis approach that locates the defective wires and replaces them with the spare wires embedded in the network. Our analytical and experimental studies show that under heavy noise, high escape rate, uncertainty about routing, and many other harmful effects, the diagnostic data collected by the proposed approach are accurate enough for the purpose of passive diagnosis. Saeed Shamshiri, Amirali Ghofrani, Kwang-Ting Cheng |
ITC | 3 |
| 2011 | Large-scale EMM identification based on geometry-constrained visual word correspondence votingabstractWe present a large-scale Embedded Media Marker (EMM) identification system which allows users to retrieve relevant dynamic media associated with a static paper document via camera-phones. The user supplies a query image by capturing an EMM-signified patch of a paper document through a camera phone. The system recognizes the query and in turn retrieves and plays the corresponding media on the phone. Accurate image matching is crucial for positive user experience in this application. To address the challenges posed by large datasets and variation in camera-phone-captured query images, we introduce a novel image matching scheme based on geometrically consistent correspondences. A hierarchical scheme, combined with two constraining methods, is designed to detect geometric constrained correspondences between images. A spatial neighborhood search approach is further proposed to address challenging cases of query images with a large translational shift. Experimental results on a 200k+ dataset show that our solution achieves high accuracy with low memory and time complexity and outperforms the baseline bag-of-words approach. © 2011 ACM. Xin Yang 0008, Qiong Liu 0003, Chunyuan Liao, Kwang-Ting Cheng, Andreas Girgensohn |
ICMR | 4 |
| 2011 | Time-Multiplexed Online CheckingabstractThere is a growing demand for online hardware checking capability to cope with increasing in-field failures resulting from variability and reliability problems. While many online checking schemes have been proposed, their area overhead remains too high for cost-sensitive applications. In this paper, we introduce a Time-Multiplexed Online Checking (TMOC) scheme using embedded field-programmable blocks for checker implementation, which enables various system parts to be checked dynamically in-field in a time-multiplexed fashion. The test quality analyses using a probabilistic model show that TMOC could maintain high fault coverage that is similar to traditional dedicated checkers. We conducted a case study of an H.264 decoder design that demonstrates our TMOC scheme provides a significant reduction in chip area and power overhead for online checkers at the cost of increased fault detection latency. We have successfully implemented and demonstrated our proposed TMOC scheme using a single Field-Programmable Gate Array (FPGA) chip. Hsiu-Ming Chang 0001, Peter Lisherness, Kwang-Ting Cheng |
IEEE Trans. Computers | 4 |
| 2011 | Modeling Yield, Cost, and Quality of a Spare-Enhanced Multicore ChipabstractIt becomes increasingly difficult to achieve a high manufacturing yield for multicore chips due to larger chip sizes, higher device densities, and greater failure rates. By adding a limited number of spare cores and wires to replace defective cores and wires either before shipment or in the field, the effective yield of the chip and its overall cost can be significantly improved. In this paper, we first model the yield of a multicore chip that incorporates both spare cores and spare wires. Then, we propose a quality metric for an NoC, and model the system yield subject to a given quality constraint. We also model the manufacturing and service costs of a multicore chip and show that a spare scheme can significantly improve the quality, increase the yield, reduce the overall cost, and substitute for the burn-in process. We illustrate that, in a spare-enhance system on a chip with high-quality in-field recovery capability, the reliance on high quality manufacturing testing can be significantly reduced. We also demonstrate that the overall quality of a mesh-based NoC depends more on the reliability of the inner links than the outer links; therefore, nonuniform spare wire distribution is sometimes more effective and cost efficient than a uniform approach. Saeed Shamshiri, Kwang-Ting Cheng |
IEEE Trans. Computers | 2 |
| 2011 | Fast Visual Retrieval Using Accelerated Sequence MatchingabstractWe present an approach to represent, match, and index various types of visual data, with the primary goal of enabling effective and computationally efficient searches. In this approach, an image/video is represented by an ordered list of feature descriptors. Similarities between such representations are then measured by the approximate string matching technique. This approach unifies visual appearance and the ordering information in a holistic manner with joint consideration of visual-order consistency between the query and the reference instances, and can be used for automatically identifying local alignments between two pieces of visual data. This capability is essential for tasks such as video copy detection where only small portions of the query and the reference videos are similar. To deal with large volumes of data, we further show that this approach can be significantly accelerated along with a dedicated indexing structure. Extensive experiments on various visual retrieval and classification tasks demonstrate the superior performance of the proposed techniques compared to existing solutions. Mei-Chen Yeh, Kwang-Ting Cheng |
IEEE Trans. Multim. | 2 |
| 2010 | An error tolerance scheme for 3D CMOS imagersabstractA three-dimensional (3D) CMOS imager constructed by stacking a pixel array of backside illuminated sensors, an analog-to-digital converter (ADC) array, and an image signal processor (ISP) array using micro-bumps (μbumps) and through-silicon vias (TSVs) is promising for high throughput applications. However, due to the direct mapping from pixels to ISPs, the overall yield relies heavily on the correctness of the μbumps, ADCs and TSVs -- a single defect leads to the information loss of a tile of pixels. This paper presents an error tolerance scheme for the 3D CMOS imager that can still deliver high quality images in the presence of μbump, ADC, and/or TSV failures. The error tolerance is achieved by properly interleaving the connections from pixels to ADCs so that the corrupted data, if any, can be recovered in the ISPs. A key design parameter, the interleaving stride, is decided by analyzing the employed error correction algorithm. Architectural simulation results demonstrate that the error tolerance scheme enhances the effective yield of an exemplar 3D imager from 46% to 99%. Hsiu-Ming Chang 0001, Jiun-Lang Huang, Ding-Ming Kwai, Kwang-Ting Cheng, Cheng-Wen Wu |
DAC | 4 |
| 2010 | SCEMIT: a systemc error and mutation injection toolabstractAs high-level models in C and SystemC are increasingly used for verification and even design (through high-level synthesis) of electronic systems, there is a growing need for compatible error injection tools to facilitate further development of coverage metrics and automated diagnosis. This paper introduces SCEMIT, a tool for the automated injection of errors into C/C++/SystemC models. A selection of 'mutation' style errors are supported, and injection is performed though a plugin interface in the GCC compiler, which minimizes the impact of SCEMIT on existing simulation flows. Experimental injected error detection results are presented for the set of OSCI SystemC Example Models as well as the CHStone C High-Level-Synthesis benchmark set. Aside from demonstrating compatibility with these models, the results show the value of high-level error injection as a coverage measure compared to conventional code coverage measures. Peter Lisherness, Kwang-Ting Cheng |
DAC | 2 |
| 2010 | An automatic test generation framework for digitally-assisted adaptive equalizers in high-speed serial linksabstractThis paper presents a new analog ATPG (AATPG) framework that generates near-optimal test stimulus for the digitally-assisted adaptive equalizers in high-speed serial links. Based on the dynamic-signature-based testing scheme developed recently, our AATPG utilizes a Genetic Algorithm (GA) which attempts to maximize the difference between the fault-free and faulty dynamic signatures of the target fault. Our test generation framework takes into account process variations and signal noise in selecting the test stimulus, which minimizes the number of misclassified devices. The experimental results on a 5-tap feed-forward adaptive equalizer demonstrate that the GA-tests generated by our framework can effectively detect faults that are hard to detect by the hand-crafted tests. Mohamed Abbas, Kwang-Ting Cheng, Yasuo Furukawa, Satoshi Komatsu, Kunihiro Asada |
DATE | 2 |
| 2010 | Pseudo-CMOS: A novel design style for flexible electronicsabstractFlexible electronics have attracted much attention since they enable promising applications such as low-cost RFID tags and e-paper. Thin-film transistors (TFTs) are considered as an ideal candidate to implement flexible electronics on low-cost substrates. Most TFT technologies, however, have only mono-type - either n- or p-type - devices and thus modern design technologies for silicon-based electronics cannot be directly applied. In this paper, we propose a novel design style Pseudo-CMOS for flexible electronics that uses only mono-type TFTs while achieving comparable performance with the complementary-type designs. The manufacturing cost and complexity can therefore be significantly reduced while the circuit yield and reliability are also enhanced with the built-in capability of post-fabrication tuning. Some standard cells have been designed and fabricated in p-type organic and n-type InGaZnO (IGZO) TFT technologies which successfully verify the superiority of the proposed Pseudo-CMOS design style. To the best of our knowledge, this is the first design solution that has proven superior performance for both types of TFT technologies. Tsung-Ching Huang, Kenjiro Fukuda, Chun-Ming Lo, Yung-Hui Yeh, Tsuyoshi Sekitani, Takao Someya, Kwang-Ting Cheng |
DATE | 7 |
| 2010 | A portable multi-pitch e-drum based on printed flexible pressure sensorsabstractPressure sensors are ideal candidates for implementing portable digital music instruments. Existing commercial pressure sensors, however, are not optimized to meet both timing and precision requirements for acoustic uses. In this paper, we demonstrate a portable multi-pitch electronic drum (e-drum) system based on large-area (> 15cm in diameter) ring-shaped pressure sensors made with low-cost screen-printing process. This e-drum system, which can accurately generate six different pitches of sounds in the current prototype, has the following key advantages: 1) a light-weight, flexible, bendable, and robust human-instrument interface, 2) real-time sound responses, 3) comparable acoustic sound quality with the conventional drums, and 4) easily expandable to a much larger number of sound-pitches. The digital music synthesis is implemented using a TI-DSP board and can be easily re-configured to realize other percussion instruments such as pianos and xylophones. To the best of our knowledge, this is the first successful demonstration of a portable e-drum based on large-area ring-shaped flexible sensors, whose success could open up many new applications. Chun-Ming Lo, Tsung-Ching Huang, Cheng-Yi Chiang, Johnson Hou, Kwang-Ting Cheng |
DATE | 5 |
| 2010 | Mutation-based diagnostic test generation for hardware design error diagnosisabstractWe propose the use of mutation-based error injection to guide the generation of high-quality diagnostic test patterns. A software-based fault localization technique is employed to derive a ranked candidate list of suspect statements. Experimental results for a set of Verilog designs demonstrate that a finer diagnostic resolution can be achieved by patterns generated by the proposed method. Shujun Deng, Kwang-Ting Cheng, Jinian Bian, Zhiqiu Kong |
ITC | 2 |
| 2010 | nGFSIM : A GPU-based fault simulator for 1-to-n detection and its applicationsabstractWe present nGFSIM, a GPU-based fault simulator for stuck-at faults which can report the fault coverage of one-to n-detection for any specified integer n using only a single run of fault simulation. nGFSIM, which explores the massive parallelism in the GPU architecture and optimizes the memory access and usage, enables accelerated fault simulation without the need of fault dropping. We show that nGFSIM offers a 25X speedup in comparison with a commercial tool and enables new applications in test selection. Huawei Li 0001, Dawen Xu 0002, Yinhe Han 0001, Kwang-Ting Cheng, Xiaowei Li 0001 |
ITC | 4 |
| 2010 | Error-locality-aware linear coding to correct multi-bit upsets in SRAMsabstractHigh-energy cosmic radiation is the major source of soft errors in SRAMs that can cause multi-bit upset around the location of the strike. In this paper, we generalize the coding problem for error detection and correction of both local (burst) and global (random) errors. We suggest using error-locality-aware codes for SRAM memories to correct single-bit or multi-bit upsets as well as physical defects. Solving the coding problem with a SAT-solver, we have found codes to correct double global or multiple (>;=3) local errors for 8, 12, 16, and 24-bit memories. For 16-bit memories, we propose a code that corrects two global or four local errors. With the same cost, our proposed code provides extra reliability than double-error-correcting BCH code. For 12-bit memories, we suggest a code that corrects two global or five local errors and has the same cost as triple-error-correcting Golay code but provides better reliability against multi-bit upsets. For memories of other widths, using syndrome analysis, we demonstrate the possibility of designing codes to correct any arbitrary number of local and global errors. Saeed Shamshiri, Kwang-Ting Cheng |
ITC | 2 |
| 2010 | A GPU-accelerated face annotation system for smartphonesabstractFace annotation makes it easy to share and manage digital photos and videos. While state-of-the-art face recognition algorithms can achieve high accuracy to support automatic face annotation, their implementations on an embedded platform cannot achieve real-time performance due to the demanding computational requirement. However, the availability of an embedded GPU in most smartphones offers the opportunity to use it as an accelerator for the face recognition task. In this demonstration, we show that, with acceleration achieved by the embedded low-power GPU, a real-time face annotation system could be realized on an existing off-the-shelf smartphone. Yi-Chu Wang, Sydney Pang, Kwang-Ting Cheng |
ACM Multimedia | 3 |
| 2010 | Calibration-assisted production testing for digitally-calibrated ADCsabstractThis paper presents a production test strategy for digitally-calibrated analog-to-digital converters (ADCs) that incorporate an equalization-based calibration scheme. By analyzing the data obtained in calibration, devices that fail certain static or dynamic specifications can be identified without any additional testing time beyond calibration. The foundation of this test strategy for the ADCs lies on the strong correlations between calibration and functional testing so that devices which violate specifications can be identified by checking the range of steady-state fluctuation in the calibration data. We further develop calibration stimuli to maximize the failing symptoms for fault detection. Simulation results on a pipelined ADC shows that the proposed strategy can effectively pre-screen a good fraction of defective devices that fail static and dynamic specifications including the gain/offset errors and the effective-number of bits (ENOB). Hsiu-Ming Chang 0001, Kwang-Ting Cheng |
VTS | 3 |
| 2010 | Innovative practices session 2C: Design, fabrication and test of flexible electronicsabstractThe development of inexpensive high-performance electronics requiring low-temperature device processing will enable low-cost, large-area flexible electronics for applications such as large-area displays, sensors, and evolving technologies such as electric paper. The recent developments of thin-film transistor (TFT) backplanes processed on flexible plastic substrates opens the possibility for novel device structures, processes, and applications. Kwang-Ting Cheng |
VTS | 1 |
| 2010 | Design, analysis, and test of low-power and reliable flexible electronicsabstractThis talk discusses some recent progress in robust circuit/system design and test of flexible electronics. We will first give an overview of reliability simulation for predicting TFT degradation under bias-stress. A reliability analysis framework, which has successfully analyzed the reliability of an amorphous-silicon (a-Si) TFT scan driver for TFT-LCD displays, will be discussed [1]. We then discuss solutions that can make TFT circuits operable under a lower supply voltage and can equip them with post-fabrication tunability for reliability and performance enhancement. Specifically, we will present a new design style, named Pseudo-CMOS [2], which has been successfully validated in a p-type organic TFT technology [3] as well as in a n-type InGaZnO (IGZO) TFT technology. Kwang-Ting Cheng, Tsung-Ching Huang |
VTS | 1 |
| 2010 | Modeling yield, cost, and quality of an NoC with uniformly and non-uniformly distributed redundancyabstractIn this paper, we propose a quality metric for an NoC and model the yield and cost of a spare-enhanced multi-core chip subject to a given quality constraint. Our experiments show that the overall quality of a mesh-based NoC depends more on the reliability of the inner links than the outer links; therefore, a non-uniform distribution of spare wires could be more effective and cost efficient than a uniform approach. Saeed Shamshiri, Kwang-Ting Cheng |
VTS | 2 |
| 2010 | Calibration and Test Time Reduction Techniques for Digitally-Calibrated Designs: an ADC Case Study
Hsiu-Ming Chang 0001, Kwang-Ting Cheng |
J. Electron. Test. | 3 |
| 2009 | Calibration as a Functional Test: An ADC Case StudyabstractIn this paper, we analyze the relationship between calibration and linearity testing of a digitally-calibrated pipelined ADC. Simulation results validate that the calibration process, once converged, could have automatically covered the INL testing of the ADC under test. Hsiu-Ming Chang 0001, Kwang-Ting Cheng |
Asian Test Symposium | 3 |
| 2009 | Low Overhead Time-Multiplexed Online Checking: A Case Study of An H.264 DecoderabstractTo cope with increasing in-field failure rates for cost-sensitive electronic products, a low-overhead online checking methodology - Time-Multiplexed Online Checking (TMOC) - was proposed and demonstrated (see IEEE Asian Test Symposium, p.371-6, 2008). In this paper, we study the area overhead required for employing TMOC in an embedded field programmable gate array (eFPGA) core. The overheads caused by the relatively low logic density of eFPGA and the interface routing between a design module and its TMOC checker are examined in detail. In a case study of an H.264 decoder design, TMOC is compared to a dedicated duplication-based online checking scheme, which typically incurs more than 100% area overhead. Experimental results show that TMOC provides significant chip area overhead reduction for online checkers. A reduction of 68% is achieved when one checker is shared by 62 design partitions, for example. TMOC can also help reduce dynamic power overhead of online checking by increasing the number of partitions, at the cost of increased fault detection latency in some partitions. Kwang-Ting Cheng |
Asian Test Symposium | 2 |
| 2009 | Signature-Based Testing for Digitally-Assisted Adaptive Equalizers in High-Speed Serial LinksabstractThis paper presents a cost-effective test methodology for adaptive equalizers which follow the digitally-assisted analog design style. By observing the states in the digital adaptation engine during or after the adaptation process in response to the test stimulus, the health of the adaptive equalizer can be determined. We propose two different types of signatures, namely static and dynamic signatures, based on the states of the digital adaptation engine for fault detection. The static signatures are derived from states after the adaption process converges and the dynamic ones from the state sequences sampled during adaption. Such signatures, combined with a variety of test stimuli, enable the detection of many hard-to-detect faults which cannot be detected by existing approaches. Our experimental results demonstrate the effectiveness and efficiency of the proposed method. Mohamed Abbas, Kwang-Ting Cheng, Yasuo Furukawa, Satoshi Komatsu, Kunihiro Asada |
ETS | 2 |
| 2009 | MyFinder: near-duplicate detection for large image collectionsabstractThe explosive growth of multimedia data poses serious challenges to data storage, management and search. Efficient near-duplicate detection is one of the required technologies for various applications. In this paper, we introduce MyFinder, an image near-duplicate detection system for large image collections. MyFinder consists of three major components: 1) a local-feature-based image representation utilizing the proposed LDP (Local-Difference-Pattern) feature, 2) the Locality-Sensitive-Hashing (LSH) as the core indexing structure to assure the most frequent data access occurred in the main memory, and 3) multi-step verification for queries to best exclude false positives and to increase the precision. Xin Yang 0008, Qiang Zhu 0006, Kwang-Ting Cheng |
ACM Multimedia | 3 |
| 2009 | A compact, effective descriptor for video copy detectionabstractLarge scale video copy detection tasks require a compact and computational-efficient descriptor that is robust to various transformations that are typically applied to generate copies. In this paper, we propose a new frame-level descriptor for such a task. The descriptor encodes the internal structure of a video frame by computing the pair-wise correlations between geometrically pre-indexed blocks. It is conceptually simple, small in size, and fast to compute. Experiments using the MUSCLE VCD benchmark show its superior performance compared to existing approaches. Mei-Chen Yeh, Kwang-Ting Cheng |
ACM Multimedia | 2 |
| 2009 | Calibration and Testing Time Reduction Techniques for a Digitally-Calibrated Pipelined ADCabstractModern mixed-signal/RF circuits with digital calibration capabilities could achieve significant performance improvements once the calibration process is completed; however, the calibration time is often very long - in the order of hundreds of milliseconds or even seconds. As testing such devices would require completion of calibration first, lengthy calibration time would result in unacceptably long testing time. In this paper, we propose design-for-testability modifications and acceleration techniques for adaption algorithms to reduce the calibration time required for testing a digitally-calibrated pipelined ADC. For the pipelined ADC proposed in, simulation results show that the proposed techniques can achieve a 60X reduction in the calibration time. Hsiu-Ming Chang 0001, Chin-Hsuan Chen, Kwang-Ting Cheng |
VTS | 4 |
| 2009 | Yield and Cost Analysis of a Reliable NoCabstractThe yield and cost of a multi-core chip improve significantly through the addition of some spare cores in the system. In this paper, we model the manufacturing and service cost of an NoC with spare wires and routers as well as spare cores. We apply our analysis on an exemplary 9-core processor and on an Intel 80-core processor, and show that a spare scheme can significantly improve the reliability, reduce the cost, and substitute for the burn-in process. Saeed Shamshiri, Kwang-Ting Cheng |
VTS | 2 |
| 2009 | Dynamic Test Compaction for Transition Faults in Broadside Scan Testing Based on an Influence Cone MeasureabstractWe propose a compact test generation method for transition faults, which is driven by a conflict-avoidance scheme employed during test generation. Based on an influence-cone function for transition faults in broadside scan testing, two dynamic test compaction schemes, named selfish test compaction and unselfish test compaction respectively, are proposed. The selfish test compaction tries to compact as many faults as possible into the current test, while the unselfish scheme attempts to compact the tests of the hard-to-compact faults into the current test. Potential conflicts produced by the signal requirements at the pseudo-primary outputs in the first frame are avoided through the use of an input dependency graph. Experimental results and comparison with existing approaches demonstrate the efficiency and effectiveness of the proposed method. Boxue Yin, Kwang-Ting Cheng |
VTS | 3 |
| 2009 | SEChecker: A Sequential Equivalence Checking Framework Based on Kth InvariantsabstractIn recent years, considerable research efforts have been devoted to utilizing circuit structural information to improve the efficiency of Boolean satisfiability (SAT) solving, resulting in several efficient circuit-based SAT solvers. In this paper, we present a sequential equivalence checking framework based on a number of circuit-based SAT solving techniques as well as a novel invariant checker. We first introduce the notion of kth invariants. In contrast to the traditional invariants that hold for all cycles, k th invariants are guaranteed to hold only after the kth cycle from the initial state. We then present a bounded model checker (BMChecker) and an invariant checker (IChecker), both of which are based on circuit SAT techniques. Jointly, BMChecker and IChecker are used to compute the kth invariants, and are further integrated in a sequential circuit SAT solver for checking sequential equivalence. Experimental results demonstrate that the new sequential equivalence checking framework can efficiently verify large industrial designs that cannot be verified by existing solutions. Feng Lu 0002, Kwang-Ting Cheng |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2008 | Digitally-Assisted Analog/RF Testing for Mixed-Signal SoCsabstractWe propose a testing methodology for analog and radio-frequency (RF) circuitry that incorporates digital circuits for performance calibration and adaptation. We explore the reuse of built-in digital calibration circuitry, along with minor digital design-for-testability (DfT) modifications, to test and characterize analog/RF circuit performance. By observing the digital tuning signals captured in the digital calibration circuitry, the analog/RF performance can be closely estimated, thus enabling cost-effective Go/No-Go production testing. In this paper, we illustrate this testing methodology using a case study of a digitally-calibrated Weaver image-reject receiver. Hsiu-Ming Chang 0001, Min-Sheng (Mitchell) Lin, Kwang-Ting Cheng |
ATS | 3 |
| 2008 | Time-Multiplexed Online Checking: A Feasibility StudyabstractThere is growing demand for online hardware checking capability to cope with increasing in-field failures resulting from variability and reliability problems. While many online checking schemes have been proposed, their area overhead remains too high for cost-sensitive applications. We introduce a time-multiplexed online checking scheme using embedded field-programmable blocks for checker implementation, which enables various system parts to be checked dynamically in-field in a time-multiplexed fashion. This incurs less area overhead, and could maintain fault coverage similar to traditional checkers. The test quality is studied using a probabilistic model. The implementation feasibility using a field-programmable gate array (FPGA) is demonstrated. Hsiu-Ming Chang 0001, Peter Lisherness, Kwang-Ting Cheng |
ATS | 4 |
| 2008 | RTL Error Diagnosis Using a Word-Level SAT-SolverabstractWe propose a novel methodology for design error diagnosis in the HDL description using a word-level solver. In this approach, the patterns that result in erroneous responses are first used to limit the number of initial error candidates. The RTL description of the design is then modified by adding one multiplexer to each of the possible error locations. For each of the erroneous pattern, the input pattern and the expected correct response are then imposed as constraints to the modified RTL model, resulting in a formula suitable for word-level satisfiability (SAT) solving. The solutions reported by the word-level SAT solver would indicate the potential error candidates. This constraint solving process iterates for each erroneous pattern, eventually resulting in a small set of error candidates. This method is sufficiently flexible to address both single-error and multiple-error diagnosis. We present experimental results for a set of public RTL benchmark designs to demonstrate the effectiveness of this proposed approach. Saeed Mirzaeian, Feijun (Frank) Zheng, Kwang-Ting Cheng |
ITC | 3 |
| 2008 | A Cost Analysis Framework for Multi-core Systems with SparesabstractIt becomes increasingly difficult to achieve a high manufacturing yield for multi-core chips due to larger chip sizes, higher device densities, and greater failure rates. By adding a limited number of spare cores to replace defective cores either before shipment or in the field, the effective yield of the chip and its overall cost can be significantly improved. In this paper, we propose a yield and cost analysis framework to better understand the dependency of a multi-core chip's cost on key parameters such as the number of cores and spares, core yield, and defect coverage of manufacturing and in-field testing. Our analysis shows that we can eliminate the burn-in process when we have some spare cores for in-field recovery. We demonstrate that a high defect coverage for in-field testing, a necessity for supporting in-field recovery, is essential for overall cost reduction. We also illustrate that, with in-field recovery capability, the reliance on high quality manufacturing testing is significantly reduced. Saeed Shamshiri, Peter Lisherness, Sung-Jui (Song-Ra) Pan, Kwang-Ting Cheng |
ITC | 4 |
| 2008 | A real-time, embedded face-annotation systemabstractFace detection and recognition have numerous multimedia applications of broad interest, one of which is automatic face annotation. There exist many robust algorithms tackling these problems but most of these algorithms are computationally demanding and have only been implemented in PC- or server-based environments. In this demonstration we show a real-time face-annotation system on a commercial PDA development platform. We examine the challenges faced in the design and development of a practical system that can achieve detection and recognition in real-time using limited memory and computational resources which are common constraints for embedded applications. Shih-Wei Chu, Mei-Chen Yeh, Kwang-Ting Cheng |
ACM Multimedia | 3 |
| 2008 | Bit-Error Rate Estimation for Bang-Bang Clock and Data Recovery Circuit in High-Speed Serial LinksabstractClock and data recovery (CDR) circuits incorporating a bang-bang (BB) phase detector have been widely adopted in high-speed serial links due to their advantages in high speed implementations. However, the heavily non-linear nature of the BB phase detector makes the analysis of the CDR loop difficult. In this paper, we propose a new technique for accurate and efficient estimation of the bit-error rate (BER) for BB CDR circuits. The technique estimates the BER based on the spectral information of jitter and the jitter transfer characteristics of the BB CDR circuit. It eliminates the conventional BER measurement process and, thus, substantially accelerates the jitter tolerance test. In addition, this technique offers insights into the behavior of the non-linear CDR loop and the contribution of the jitter to the BER. We present simulation results that demonstrate the potential usefulness of the method. Dongwoo Hong, Kwang-Ting Cheng |
VTS | 2 |
| 2008 | Reliability analysis for flexible electronics: Case study of integrated a-Si: H TFT scan driverabstractFlexible electronics fabricated on thin-film, lightweight, and bendable substrates (e.g., plastic) have great potential for novel applications in consumer electronics such as flexible displays, e-paper, and smart labels; however, the key elements, namely thin-film transistors (TFTs), for implementing flexible circuits often suffer from electrical instability. Therefore, thorough reliability analysis is critical for flexible circuit design to ensure that the circuit will operate reliably throughout its lifetime. In this article we propose a methodology for reliability simulation of hydrogenated amorphous silicon (a-Si:H) TFT circuits. We show that: (1) the threshold voltage ( V TH ) shift of a single TFT can be estimated by analyzing its operating conditions; and (2) the circuit lifetime can be predicted accordingly by using SPICE-like simulators with proper modeling. We also propose an algorithm to reduce the simulation time by orders of magnitude, with good prediction accuracy. To validate our analytical model and simulation methodology, we compare simulation results with the actual circuit measurements of an integrated a-Si:H TFT scan driver fabricated on a glass substrate and we demonstrate very good consistency. Tsung-Ching Huang, Kwang-Ting Cheng, Huai-Yuan Tseng, Chen-Pang Kung |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2007 | An Accurate Jitter Estimation Technique for Efficient High Speed I/O TestingabstractThis paper describes a technique for estimating total jitter that, along with a loopback-based margining test, can be applied to test high speed serial interfaces. We first present the limitations of the existing estimation method, which is based on the dual-Dirac model. The accuracy of the existing method is extremely sensitive to the choice of the fitting region and the ratio of deterministic jitter to random jitter. Then, we propose a high-order polynomial fitting technique and demonstrate its value for a more efficient and accurate total jitter estimation at a very low Bit-Error-Rate level. The estimation accuracy is also analyzed with respect to different numbers of measurement points for fitting. This analysis shows that only a very small number (i.e., 4) of measurement points is needed for achieving accurate estimation. Dongwoo Hong, Kwang-Ting Cheng |
ATS | 2 |
| 2007 | An Efficient Diagnostic Test Pattern Generation Framework Using Boolean SatisfiabilityabstractThis paper presents a diagnostic test pattern generation (DTPG) framework based upon a Boolean Satisfiability engine. We first propose an enhanced miter-based model for distinguishing fault candidates that can achieve greater efficiency as well as can prove a group of undifferentiable faults. The model can also be used to generate diagnostic tests for distinguishing faults of different fault types. Based on this model, we propose a diagnostic pattern compaction strategy. By exploring "don't cares " at the primary inputs, the number of required diagnostic patterns can be reduced. Experimental results show that the proposed method achieves a greater diagnosis resolution when combined with existing approaches. Also, fewer diagnostic test patterns are needed. Feijun Zheng, Kwang-Ting Cheng, Xiaolang Yan, John Moondanos, Ziyad Hanna |
ATS | 2 |
| 2007 | Reliability Analysis for Flexible Electronics: Case Study of Integrated a-Si: H TFT Scan DriverabstractFlexible electronics fabricated with thin-film and bendable substrates (e.g., plastic) have great potential for novel applications in consumer electronics such as flexible displays, e-paper, and smart labels; however, the key elements --- namely thin-film transistors (TFTs) --- often suffer from electrical instability. Therefore, thorough reliability analysis is critical for flexible circuit design to ensure that the circuit would operate reliably throughout its lifetime. In this paper, we propose a methodology for a-Si:H TFT circuits' reliability simulation. We show that: (1) the threshold voltage (VTH) shift of a single TFT can be modeled by analyzing its operating conditions and (2) the circuit lifetime can be predicted accordingly using SPICE. We also propose an algorithm to reduce the simulation time by orders of magnitude with negligible accuracy loss. To validate our analytical model and simulation methodology, we compare the SPICE simulation results with the actual measurements of our integrated a-Si:H TFT scan driver fabricated on the glass substrate and demonstrate very high consistency in the overall results. Tsung-Ching Huang, Huai-Yuan Tseng, Chen-Pang Kung, Kwang-Ting Cheng |
DAC | 4 |
| 2007 | A two-tone test method for continuous-time adaptive equalizers
Dongwoo Hong, Shadi Saberi, Kwang-Ting Cheng, C. Patrick Yue |
DATE | 3 |
| 2007 | Testable design for advanced serial-link transceiversabstractThis paper describes a DfT solution for modern serial-link transceivers. We first summarize the architectures of the crosstalk canceller and the equalizer used in advanced transceivers to which the proposed solution can be applied. The solution addresses the testability and observability issues of the transceiver for both characterization and production testing. Without using sophisticated testing instrument setting, the proposed solution could test the clock and data recovery circuit and characterize the decision-feedback equalizer in the receiver. Our experiments demonstrate that the proposed method has significant higher fault coverage and lower hardware requirement than the conventional approach of probing the eye-opening of the signals inside the transceiver Mitchell Lin, Kwang-Ting Cheng |
DATE | 2 |
| 2007 | A framework for system reliability analysis considering both system error tolerance and component test quality
Sung-Jui (Song-Ra) Pan, Kwang-Ting Cheng |
DATE | 2 |
| 2007 | A hybrid scheme for compacting test responses with unknown valuesabstractThis paper presents a hybrid compaction scheme for test responses containing unknown values, which consists of a space compactor and an unknown-blocking Multiple Input Signature Registers (MISR). The proposed scheme guarantees no coverage loss for the modeled faults. The proposed hybrid scheme can also be tuned to observe any user- specified percentage of responses for controlling the coverage loss for un-modeled faults. The experimental results demonstrate that, in comparison with a space compactor or an unknown-blocking MISR alone, the hybrid compaction scheme achieves a lower coverage loss without demanding more test-data volume. In addition, we propose a quantitative approach to estimate the required percentage of observable responses for the proposed scheme, directly based on a test-quality metric of un-modeled faults. Mango Chia-Tso Chao, Kwang-Ting Cheng, Seongmoon Wang, Srimat T. Chakradhar, Wenlong Wei |
ICCAD | 2 |
| 2007 | Multiple-Fault Diagnosis Based On Adaptive Diagnostic Test Pattern GenerationabstractIn this paper, we propose two fault-diagnosis methods for improving multiple-fault diagnosis resolution. The first method, based on the principle of single-fault activation and single-output observation, employs a new circuit transformation technique in conjunction with the use of a special type of diagnostic test pattern, named single-observation single-location-at-a-time (SO-SLAT) pattern. Given a list of candidate suspects (which could be stuck-at, transition, bridging, or other faults obtained by any existing diagnosis method), we generate a set of SO-SLAT patterns, each of which attempts to activate only one fault in the list and propagate its effects only to a specific observation point. Observing the responses of the circuit under diagnosis to the SO-SLAT patterns helps more precisely determine whether each fault suspect is a true or false candidate. The method can tolerate most of the timing hazards for a more accurate diagnosis of failures caused by timing faults. The second method generates and applies limited-cycle sequential tests, based on a Boolean satisfiability solver, to identify multiple defective signals which can jointly explain the circuit’s faulty behavior. These two methods can be applied independently and/or jointly after any existing state-of-the-art diagnosis process to further improve the diagnosis resolution. The experimental results demonstrate the effectiveness of the proposed methods for diagnosing multiple faults, including timing faults. Yung-Chieh Lin, Feng Lu 0002, Kwang-Ting Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2007 | Silicon Debug for Timing ErrorsabstractDue to various sources of noise and process variations, assuring a circuit to operate correctly at its desired operational frequency has become a major challenge. In this paper, we propose a timing-reasoning-based algorithm and an adaptive test-generation algorithm for diagnosing timing errors in the silicon-debug phase. We first derive three metrics that are strongly correlated to the probability of a candidate's being an actual error source. We analyze the problem of circuit timing uncertainties caused by delay variations and test sampling. Then, we propose a candidate-ranking heuristic, which is robust with respect to such sources of timing uncertainty. Based on the initial ranking result and the timing information, we further propose an adaptive path-selection and test-generation algorithm to generate additional diagnostic patterns for further improvement of the first-hit-rate. The experimental results demonstrate that combining the ranking heuristic and the adaptive test-generation method would result in a very high resolution for timing diagnosis. Kwang-Ting Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2006 | Generation of shorter sequences for high resolution error diagnosis using sequential SATabstractCommonly used pattern sources in simulation-based verification include random, guided random, or design verification patterns. Although these patterns may help bring the design to those hard-to-reach states for activating the errors and for propagating them to observation points, they tend to be very long, which complicates the subsequent diagnosis process. As a key step in reducing the overall diagnosis complexity, we propose a method of generating a shorter error-sequence based on a given long error-sequence. We formulate the problem as a satisfiability problem and employ a SAT solver as the underlying engine for this task. By heuristically selecting an intermediate state S/sub i/ which is reachable by the given long sequence, the task of finding the transfer sequence from the initial state to the target state can be divided into two easier tasks - finding a transfer sequence from the initial state to S/sub i/ and one from S/sub i/ to the target state. Our preliminary experimental results on public benchmark circuits show that the proposed method can achieve significant reduction in the length of the error sequences. Sung-Jui (Song-Ra) Pan, Kwang-Ting Cheng, John Moondanos, Ziyad Hanna |
ASP-DAC | 2 |
| 2006 | Efficient identification of multi-cycle false pathabstractDue to false paths and multicycle paths in a circuit, using only topological delay to determine the clock period could be too conservative. In this paper, we address the timing analysis problem by considering both single-cycle and multicycle operations. We give a precise definition of multicycle false paths and provide the necessary conditions for multicycle sensitizable paths. We then propose an efficient algorithm to identify multicycle false paths. By considering both single-cycle and multicycle false paths, we could derive a shorter clock period than that determined by existing methods. Finally, we propose an algorithm to compute the valid clock period and demonstrate the improvement in clock frequency by taking multicycle false paths into account Kwang-Ting Cheng |
ASP-DAC | 2 |
| 2006 | Fast Human Detection Using a Cascade of Histograms of Oriented GradientsabstractWe integrate the cascade-of-rejectors approach with the Histograms of Oriented Gradients (HoG) features to achieve a fast and accurate human detection system. The features used in our system are HoGs of variable-size blocks that capture salient features of humans automatically. Using AdaBoost for feature selection, we identify the appropriate set of blocks, from a large set of possible blocks. In our system, we use the integral image representation and a rejection cascade which significantly speed up the computation. For a 320 × 280 image, the system can process 5 to 30 frames per second depending on the density in which we scan the image, while maintaining an accuracy level similar to existing methods. Qiang Zhu 0006, Mei-Chen Yeh, Kwang-Ting Cheng, Shai Avidan |
CVPR (2) | 3 |
| 2006 | Unknown-tolerance analysis and test-quality control for test response compaction using space compactorsabstractFor a space compactor, degradation of fault detection capability caused by the masking effects from unknown values is much more serious than that caused by error masking (i.e. aliasing). In this paper, we first propose a mathematical framework to estimate the percentage of observable responses under unknown-induced masking for a space compactor. We further develop a prediction scheme which can correlate the percentage of observable responses with the modeled-fault coverage and with a n-detection metric for a given test set. As a result, the quality of a space compactor can be measured directly based on its test quality, instead of based on indirect metrics such as the number of tolerated unknowns or the aliasing probability. With the prediction scheme above, we propose a construction flow for space compactors to achieve the desired level of test quality while maximizing the compaction ratio. Mango Chia-Tso Chao, Kwang-Ting Cheng, Seongmoon Wang, Srimat T. Chakradhar, Wenlong Wei |
DAC | 2 |
| 2006 | Coverage loss by using space compactors in presence of unknown valuesabstractThe presence of unknown values in simulation is the great est barrier to effective test response compaction. For space compactors, some responses may not be observable due to the masking effect caused by unknown values. This paper reports on experiments conducted to evaluate the impact on the test quality of various percentages of observable responses for both modeled and un-modeledfaults. Mango Chia-Tso Chao, Seongmoon Wang, Srimat T. Chakradhar, Wenlong Wei, Kwang-Ting Cheng |
DATE | 5 |
| 2006 | Multiple-fault diagnosis based on single-fault activation and single-output observationabstractIn this paper, we propose a new circuit transformation technique in conjunction with the use of a special diagnostic test pattern, named SO-SLAT pattern, to achieve higher multiple-fault diagnosis resolutions. For a given list of candidate faults, which could be stuck-at, transition, bridging, or other faults, we generate a set of SO-SLAT patterns, each of which attempts to activate only one fault in the list and propagate its effects to only one observation point. Observing the responses to SO-SLAT patterns helps more precisely identify fault candidates. The method can also tolerate most of the timing hazards for more accurate diagnosis of failures caused by timing faults. The experimen tal results demonstrate the effectiveness of the proposed method for diagnosing multiple faults. Yung-Chieh Lin, Kwang-Ting Cheng |
DATE | 2 |
| 2006 | Timing-reasoning-based delay fault diagnosisabstractIn this paper, we propose a timing-reasoning algorithm to improve the resolution of delay fault diagnosis. In contrast to previous approaches which identify candidates by utilizing only logic conditions, we propose a timing-simulation-based method to perform the candidate reasoning. Based on the circuit timing information, we identify invalid candidates which cannot maintain the consistency of failure behaviors. By eliminating those invalid candidates, the diagnosis resolution can be improved. We then analyze the problem of circuit timing uncertainty caused by the delay variation and the simulation model. We calculate a metric, named invalid-probability, for each candidate. Then we propose a candidate-ranking heuristic which is robust with respect to such sources of timing uncertainty. By ranking the candidates based on their invalid-probability, we can improve the candidate first-hit-rate of the traditional critical path tracing (CPT) technique. To demonstrate the efficiency of the proposed method, we have developed a timing diagnosis framework which can simulate the real diagnosis process to evaluate and compare different algorithms Kwang-Ting Cheng |
DATE | 2 |
| 2006 | Bit Error Rate Estimation for Improving Jitter Testing of High-Speed Serial LinksabstractThis paper describes a bit error rate (BER) estimation technique for high-speed serial links, which utilizes the jitter spectral information extracted from the transmitted data and some key characteristics of the clock and data recovery (CDR) circuit in the receiver. In addition to improving the accuracy of BER prediction, the estimation technique can be used to accelerate the jitter tolerance test by eliminating the conventional BER measurement process. Experimental results comparing the estimated BER and the BERT-measured BER on a 2.5 Gbps commercial CDR circuit demonstrate the high accuracy of the proposed technique Dongwoo Hong, Kwang-Ting Cheng |
ITC | 2 |
| 2006 | A Unified Approach to Test Generation and Test Data Volume ReductionabstractIn this paper, we propose a unified approach to test generation, test stimulus compression, and test response compaction. By integrating virtual models of the decompressor and the compactor with the CUT, a standard ATPG tool can be used to generate compressed tests and their corresponding compacted responses as a unified process. In comparison with the existing solutions which treat them as separate tasks, this unified approach could potentially achieve a higher fault coverage and obtain higher input compression and output compaction ratios. Such a unified approach also offers better flexibility for evaluating various test volume reduction architectures. Our experimental results demonstrate the effectiveness of the proposed method Yung-Chieh Lin, Kwang-Ting Cheng |
ITC | 2 |
| 2006 | Testable Design for Adaptive Linear Equalizer in High-Speed Serial LinksabstractThis paper describes a novel DfT solution to the adaptive linear equalizer in the receiver of a high-speed serial link. We first summarize various equalization architectures, adaptive algorithms, circuit implementations, and their variations to which the proposed solution can be applied. The solution addresses the observability problem of the equalizer and results in a testable design that can be easily characterized and cost-effectively tested in the production line. To validate the proposed method, we conducted experiments for examples with various injected faults, for which the conventional approach of examining the eye opening at the output of the equalizer results in poor fault coverage. Simulation results demonstrate that the proposed method has a significantly higher coverage for those hard-to-detect faults. Mitchell Lin, Kwang-Ting Cheng |
ITC | 2 |
| 2006 | Multimodal fusion using learned text concepts for image categorizationabstractConventional image categorization techniques primarily rely on low-level visual cues. In this paper, we describe a multimodal fusion scheme which improves the image classification accuracy by incorporating the information derived from the embedded texts detected in the image under classification. Specific to each image category, a text concept is first learned from a set of labeled texts in images of the target category using Multiple Instance Learning [1]. For an image under classification which contains multiple detected text lines, we calculate a weighted Euclidian distance between each text line and the learned text concept of the target category. Subsequently, the minimum distance, along with low-level visual cues, are jointly used as the features for SVM-based classification. Experiments on a challenging image database demonstrate that the proposed fusion framework achieves a higher accuracy than the state-of-art methods for image classification. Qiang Zhu 0006, Mei-Chen Yeh, Kwang-Ting Cheng |
ACM Multimedia | 3 |
| 2006 | Guest Editorial
Salvador Mir, Kwang-Ting Cheng, Andrew Richardson 0001 |
J. Electron. Test. | 2 |
| 2006 | Simulation-Based Functional Test Generation for Embedded ProcessorsabstractDeterministic functional test pattern generation has been a long-standing open problem, which is an important problem to be solved for both design verification and manufacturing testing. One key in developing a practical functional test pattern generation approach is to avoid the exponential growth of the test generation complexity in terms of the design size. This work proposes a novel functional test generation approach where simulation results are used to guide the generation of additional tests. Our methodology avoids the complexity growth issue by converting some modules in a design into simpler and more efficient models. Then, these models are used to facilitate the actual test generation process. We develop two sets of techniques to achieve these conversions: Boolean learning for random logic and arithmetic learning for datapath modules. We demonstrate the effectiveness and discuss the. limitations of these techniques through experiments on benchmark circuits. Last, we validate the overall test generation methodology based on the OpenRISC 1200 microprocessor Charles H.-P. Wen, Li-C. Wang, Kwang-Ting Cheng |
IEEE Trans. Computers | 3 |
| 2006 | Pseudofunctional testingabstractRecent research results have shown that the traditional structural testing for delay and signal integrity faults may result in overtesting due to the nontrivial number of such faults that are untestable in the functional mode although testable in the test mode. This paper presents a pseudofunctional-test methodology that attempts to minimize the overtesting problem of the scan-based circuits in automatic test pattern generation (ATPG) and built-in self-test (BIST) test generation approaches. The first pattern of a two-pattern test is still delivered by scan in the test mode but the pattern is generated in such a way that it does not violate the functional constraints extracted from the functional logic. The second pattern is then generated in a functional mode using the functional justification (also called broadside) test application scheme. The authors use a sequential boolean satisfiability solver to extract a set of functional constraints that consists of illegal states and internal signal correlation. The functional constraints are imposed upon an ATPG tool to generate pseudofunctional tests and/or implemented as a monitor in the BIST environment to allow only functional-like patterns generated from the random test pattern generator as tests. The experimental results for delay faults indicate that the percentage of functionally untestable delay faults is nontrivial for many circuits. This finding supports the hypothesis of the overtesting problem in delay testing. In addition, the results indicate the effectiveness of the proposed constraint extraction method and the proposed BIST scheme. Yung-Chieh Lin, Feng Lu 0002, Kwang-Ting Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2005 | Constraint extraction for pseudo-functional scan-based delay testingabstractRecent research results have shown that the traditional structural testing for delay and crosstalk faults may result in over-testing due to the non-trivial number of such faults that are untestable in the functional mode while testable in the test mode. This paper presents a pseudo-functional test methodology that attempts to minimize the over-testing problem of the scan-based circuits for the delay faults. The first pattern of a two-pattern test is still delivered by scan in the test mode but the pattern is generated in such a way that it does not violate the functional constraints extracted from the functional logic. In this paper, we use a SAT solver to extract a set of functional constraints which consists of illegal states and internal signal correlation. Along with the functional justification (also called broad-side) test application scheme, the functional constraints are imposed to a commercial delay-fault ATPG tool to generate pseudo-functional delay tests. The experimental results indicate that the percentage of untestable delay faults is non-trivial for many circuits which support the hypothesis of the over-testing problem in delay testing. The results also indicate the effectiveness of the proposed constraint extraction method. Yung-Chieh Lin, Feng Lu 0002, Kwang-Ting Cheng |
ASP-DAC | 4 |
| 2005 | Structural search for RTL with predicate learningabstractWe present an efficient search strategy for satisfiability checking on circuits represented at the register-transfer-level (RTL). We use the RTL circuit structure by extending concepts from classic automatic test-pattern generation (ATPG) algorithms and interval-arithmetic to guide the search process. We extend the idea of Boolean recursive learning on predicate logic in the RTL using Boolean and interval constraint propagation in the control and data-path of the circuit. This is used as a pre-processing step to derive relations between predicate logic signals that are used to augment the search. We demonstrate experimentally that these methods provide significant improvement over current techniques on sample benchmarks. Ganapathy Parthasarathy, Madhu K. Iyer, Kwang-Ting Cheng, Forrest Brewer |
DAC | 3 |
| 2005 | Efficient Conflict-Based Learning in an RTL Circuit Constraint SolverabstractWe present new techniques for improving search in a hybrid Davis-Putnam-Logemann-Loveland based constraint solver for RTL (register-transfer level) circuits (HDPLL). In earlier work on HDPLL (Parthasarathy, G. et al., 41st DAC, 2004), the authors combined solvers for integer and Boolean domains using finite-domain constraint propagation with heuristic conflict-based learning. We describe a new algorithm that extends the conflict-based unique-implication point learning in Boolean SAT (satisfiability) solvers to hybrid Boolean-integer domains in HDPLL. We describe data-structures for efficient constraint propagation on the hybrid learned relations, similar to two-literal watching in Boolean SAT. We demonstrate that these new techniques provide considerable performance benefits when compared with other combinations of decision theories. Madhu K. Iyer, Ganapathy Parthasarathy, Kwang-Ting Cheng |
DATE | 3 |
| 2005 | An Efficient Sequential SAT Solver With Improved Search StrategiesabstractA sequential SAT solver, Satori, was recently proposed (Iyer, M.K. et al., Proc. IEEE/ACM Int. Conf. on Computer-Aided Design, 2003) as an alternative to combinational SAT in verification applications. This paper describes the design of Seq-SAT, an efficient sequential SAT solver with improved search strategies over Satori. The major improvements include: (1) a new and better heuristic for minimizing the set of assignments to state variables; (2) a new priority-based search strategy and a flexible sequential search framework which integrates different search strategies; (3) a decision variable selection heuristic more suitable for solving the sequential problems. We present experimental results to demonstrate that our sequential SAT solver can achieve orders-of-magnitude speedup over Satori. We plan to release the source code of Seq-SAT. Feng Lu 0002, Madhu K. Iyer, Ganapathy Parthasarathy, Li-C. Wang, Kwang-Ting Cheng, Kuang-Chien Chen |
DATE | 5 |
| 2005 | Response shaper: a novel technique to enhance unknown tolerance for output response compactionabstractThe presence of unknown values in the simulation result is a key barrier to effective output response compaction in practice. This paper proposes a simple circuit module, called a response shaper, to reshape the scan-out responses before feeding them to a space compactor. Along with the proposed reshaping algorithm, response shapers can help the space compactor to reduce the number of undetectable modeled and unmodeled faults in the presence of unknown values. Moreover, the proposed compaction scheme is ATPG-independent and its hardware requirement is pattern-independent. In our experiments, we use a simple XOR compactor as the space compactor to evaluate the effectiveness of the response shaper. The results show that the number of undetectable faults and unobservable scan-out responses can be significantly reduced in comparison with the results of a convolutional compactor. The number of the extra scan-in bits required for the control signals of the response shapers is only a small fraction of the total test data volume. Also, its hardware overhead is acceptable and the runtime of the reshaping algorithm is scalable for large industrial designs. Mango Chia-Tso Chao, Seongmoon Wang, Srimat T. Chakradhar, Kwang-Ting Cheng |
ICCAD | 4 |
| 2005 | RTL SAT simplification by Boolean and interval arithmetic reasoningabstractWe present a method that combines interval-arithmetic (IA) and Boolean reasoning with structural hashing for simplifying SAT problems on circuits expressed at the register-transfer level. We demonstrate that simple transformations based on interval-arithmetic operations can significantly reduce the complexity of the problem. We identify cases where the inherent over-approximations in IA operations can be reduced. We demonstrate that these techniques can significantly reduce RTL-SAT instances in size and runtime. Ganapathy Parthasarathy, Madhu K. Iyer, Kwang-Ting Cheng, Forrest Brewer |
ICCAD | 3 |
| 2005 | ChiYun Compact: A Novel Test Compaction Technique for Responses with Unknown ValuesabstractThis paper proposes a response compactor, named ChiYun compactor, to compact scan-out responses in the presence of unknown values. By adding storage elements into an Xor network, a ChiYun compactor can offer multiple chances for a scan-out response to be observed at ATE channels in one to several scan-shift cycles. We also develop a mathematical analysis to predict the percentage of scan-out responses masked by the unknown values for the ChiYun compactor. With this analysis, we can derive the optimal configuration of a ChiYun compactor for minimizing the masking of scan-out responses. We further propose a selection scheme for the ChiYun compactor to selectively observe partial Xor results for improving the fault coverage. The experimental results demonstrate the effectiveness of the proposed mathematical analysis and the selection scheme. We also demonstrate that the unknown tolerance of a ChiYun compactor is higher than that of a state-of-the-art response compactor proposed in (Wang, 2003). Mango Chia-Tso Chao, Seongmoon Wang, Srimat T. Chakradhar, Kwang-Ting Cheng |
ICCD | 4 |
| 2005 | Accurate Diagnosis of Multiple FaultsabstractIn this paper, we propose a diagnostic test generation method in conjunction with an efficient sequential SAT-based diagnosis procedure to precisely identify multiple defective signals which can jointly explain the circuit's faulty behavior. This method can be applied after any existing state-of-the-art diagnosis process to further improve the diagnosis resolution. The proposed diagnosis method generates limited-cycle sequential tests for SAT-based diagnosis which results in significantly higher diagnosis accuracy for multiple faults. The experimental results demonstrate the effectiveness of the proposed method. Yung-Chieh Lin, Feng Lu 0002, Kwang-Ting Cheng |
ICCD | 3 |
| 2005 | Learning a Sparse, Corner-Based Representation for Time-varying Background ModelingabstractTime-varying phenomenon, such as ripples on water, trees waving in the wind and illumination changes, produces false motions, which significantly compromises the performance of an outdoor-surveillance system. In this paper, we propose a corner-based background model to effectively detect moving-objects in challenging dynamic scenes. Specifically, the method follows a three-step process. First, we detect feature points using a Harris corner detector and represent them as SIFT-like descriptors. Second, we dynamically learn a background model and classify each extracted feature as either a background or a foreground feature. Last, a "Lucas-Kanade" feature tracker is integrated into this framework to differentiate motion-consistent foreground objects from background objects with random or repetitive motion. The key insight of our work is that a collection of SIFT-like features can effectively represent the environment and account for variations caused by natural effects with dynamic movements. Features that do not correspond to the background must therefore correspond to foreground moving objects. Our method is computational efficient and works in real-time. Experiments on challenging video clips demonstrate that the proposed method achieves a higher accuracy in detecting the foreground objects than the existing methods. Qiang Zhu 0006, Shai Avidan, Kwang-Ting Cheng |
ICCV | 3 |
| 2005 | Using visual features for anti-spam filteringabstractUnsolicited commercial email (UCE), also known as spam, has been a major problem on the Internet. In the past, researchers have addressed this problem as a text classification or categorization problem. However, as spammers' techniques continue to evolve and the genre of email content becomes more and more diverse, text-based anti-spam approaches alone are no longer sufficient. In this paper, we propose a novel anti-spam system which utilizes visual clues, in addition to text information in the email body, to determine whether a message is spam. We analyze a large collection of spam emails containing images and identify a number of useful visual features for this application. We then propose using one-class support vector machines (SVM) as the underlying base classifier for anti-spam filtering. The experimental results demonstrate that the proposed system can add significant filtering power to the existing text-based anti-spam filters. Ching-Tung Wu, Kwang-Ting Cheng, Qiang Zhu 0006, Yi-Leh Wu |
ICIP (3) | 2 |
| 2005 | Production-oriented interface testing for PCI-Express by enhanced loop-back techniqueabstractTesting PCI-Express (PCI-E) interface at 2.5Gb/s in production is challenging and very expensive. This paper proposes a low-cost test method which could inject data-dependent and bounded random jitters into conventional external loop-back testing configuration, and also could test the jitter tracking capability of the data recovery circuit. This method has been implemented for production testing for a PCI-E interface device, in which the receiver employs a 3/spl times/-oversampling data recovery circuit. We analyze the measurement data and study their correlation with the simulation model. We also discuss some important issues for achieving high test accuracy and coverage using loop-back testing. Mitchell Lin, Kwang-Ting Cheng, Jimmy Hsu, M. C. Sun, Shelton Lu |
ITC | 2 |
| 2005 | Simulation-based target test generation techniques for improving the robustness of a software-based-self-test methodologyabstractSoftware-based self-test (SBST) was previously proposed as an on-chip functional test methodology. Achieving desired full-chip functional fault coverage has always been a challenge because random test program generation (RTPG) alone may not be sufficient. This work investigates the potential of using target test program generation (TTPG) to supplement the RTPG method. The proposed TTPG method utilizes simulation results to develop learned models for the surrounding modules of the block under test. Then, the learned models replace the surrounding modules around the block in the actual test generation process. Because the learned models are much simpler to handle, this method minimizes the cost of functional TPG. For developing the simulation-based learning scheme, we divide the surrounding modules into two categories: Boolean and arithmetic. We apply different techniques for each category and explain their applicability and limitations. The feasibility and effectiveness of the proposed simulation-based TTPG method in the context of supplementing RTPG for achieving high fault coverage in SBST of a RISC pipelined microprocessor design is demonstrated as well Charles H.-P. Wen, Li-C. Wang, Kwang-Ting Cheng, Wei-Ting Liu, Ji-Jan Chen |
ITC | 3 |
| 2005 | Pseudo-Functional Scan-based BIST for Delay FaultabstractThis paper presents a pseudo-functional BIST scheme that attempts to minimize the over-testing problem of logic BIST for delay and crosstalk-induced failures. The over-testing problem is evident from the non-trivial number of structurally testable while functionally untestable (ST-FU) faults. Such faults can be detected by some scan/BIST patterns but not by any functional pattern. The goal of this BIST scheme is to allow only functional-like patterns generated from the BIST random test pattern generator (RTPG) as tests. This is done by inserting a Monitor at the output of the RTPG, which indicates whether the current pattern violates some pre-extracted functional constraints. In case of violation, the pattern is skipped. In our implementation, a SAT solver is used to analyze and extract a set of functional constraints from the functional logic. These functional constraints are then implemented in hardware as the Monitor. Even though the extracted functional constraints can not be exhausted, the proposed BIST scheme can detect and filter out, in real-time, a substantial subset of the nonfunctional patterns, and thus minimizing the over-testing problem. We present some experimental results to demonstrate the effectiveness of the proposed BIST scheme. Yung-Chieh Lin, Feng Lu 0002, Kwang-Ting Cheng |
VTS | 3 |
| 2005 | On A Software-Based Self-Test Methodology and Its ApplicationabstractSoftware-based self-test (SBST) was originally proposed for cost reduction in SOC test environment. Previous studies have focused on using SBST for screening logic defects. SBST is functional-based and hence, achieving a high full-chip logic defect coverage can be a challenge. This raises the question of SBST's applicability in practice. In this paper, we investigate a particular SBST methodology and study its potential applications. We conclude that the SBST methodology can be very useful for producing speed binning tests. To demonstrate the advantage of using SBST in at-speed functional testing, we develop a SBST framework and apply it to an open source microprocessor core, named OpenRISC 1200. A delay path extraction methodology is proposed in conjunction with the SBST framework. The experimental results demonstrate that our SBST can produce tests for a high percentage of extracted delay paths of which less than half of them would likely be detected through traditional functional test patterns. Moreover, the SBST tests can exercise the functional worst-case delays which could not be reached by even 1M of traditional verification test patterns. The effectiveness of our SBST and its current limitations are explained through these experimental findings. Charles H.-P. Wen, Li-C. Wang, Kwang-Ting Cheng, Wei-Ting Liu, Ji-Jan Chen |
VTS | 3 |
| 2005 | Using 2-domain partitioned OBDD data structure in an enhanced symbolic simulatorabstractIn this article, we propose a symbolic simulation method where Boolean functions can be efficiently manipulated through a 2-domain partitioned OBDD data structure. The functional partition is applied by automatically exploring the key decision points implicitly built inside a circuit. The partition can help to significantly reduce the OBDD sizes, solving problems that could not be solved with monolithic OBDD data structure. We demonstrate the performance of the approach through the symbolic simulation of several benchmark circuits with complex control logics and datapath. The symbolic simulation based on 2-domain partitioned OBDD can be also applied in equivalence checking. It can generate the signature of functions to identify the critical partition points in the optimized gate-level netlist. Tao Feng 0012, Li-C. Wang, Kwang-Ting Cheng, Chih-Chan Lin |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2004 | Improved symbolic simulation by functional-space decomposition
Tao Feng 0012, Li-C. Wang, Kwang-Ting Cheng |
ASP-DAC | 3 |
| 2004 | Jitter spectral extraction for multi-gigahertz signal
Chee-Kian Ong, Dongwoo Hong, Kwang-Ting Cheng, Li-C. Wang |
ASP-DAC | 3 |
| 2004 | Efficient reachability checking using sequential SAT
Ganapathy Parthasarathy, Madhu K. Iyer, Kwang-Ting Cheng, Li-C. Wang |
ASP-DAC | 3 |
| 2004 | TranGen: a SAT-based ATPG for path-oriented transition faults
Kwang-Ting Cheng, Li-C. Wang |
ASP-DAC | 2 |
| 2004 | A Signa-Delta Modulation Based Analog BIST System with a Wide Bandwidth Fifth-Order Analog Response Extractor for Diagnosis PurposeabstractA wide bandwidth /spl Sigma/-/spl Delta/ modulation based analog built-in self-test (BIST) system that can diagnose the prototype is presented. It consists of a low-cost design-for-testability (DJT) switched-capacitor filter as the circuit under test (CUT) and a wide bandwidth analog response extractor (ARE) to digitize the analog responses for final DSP analysis. The first stage of the DJT CUT is reconfigured to accept a repetitive /spl Sigma/-/spl Delta/ modulated bit-steam as its stimulus. This DJT technique reuses every original component and thus provides the advantages of lowering the testing cost, increasing the fault coverage as well as the accuracy, and being able to perform the at-speed tests. The ARE is a cascaded 2-1-1-I fifth-order /spl Sigma/-/spl Delta/ modulator equipped with single-bit quantizers to extend the testing bandwidth while retaining moderate tolerance of circuit imperfections. Our measurement results show that this ARE is able to provide a -95 dB spurious free dynamic range over 1 MHz bandwidth when operates at 30 MHz- A multi-tone test is performed to manifest the wide bandwidth and high accuracy of our BIST system. Based on the BIST results, a method of diagnosing the prototype to speed up the time-to-market is also proposed and demonstrated by our BIST system. Hao-Chiao Hong, Cheng-Wen Wu, Kwang-Ting Cheng |
Asian Test Symposium | 3 |
| 2004 | An efficient finite-domain constraint solver for circuitsabstractWe present a novel hybrid finite-domain constraint solving engine for RTL circuits, that automatically uses data-path abstraction. We describe how DPLL search can be modified by using efficient finite-domain constraint propagation to improve communication between interacting integer and Boolean domains. This enables efficient combination of Boolean SAT and linear integer arithmetic solving techniques. We use conflict-based learning using the variables on the boundary of control and data-path for additional performance benefits. Finally, the hybrid constraint solver is experimentally analyzed using some example circuits. Ganapathy Parthasarathy, Madhu K. Iyer, Kwang-Ting Cheng, Li-C. Wang |
DAC | 3 |
| 2004 | On path-based learning and its applications in delay test and diagnosisabstractThis paper describes the implementation of a novel path-based learning methodology that can be applied for two purposes: (1) In a pre-silicon simulation environment, path-based learning can be used to produce a fast and approximate simulator for statistical timing simulation. (2) In post-silicon phase, path-based learning can be used as a vehicle to derive critical paths based on the pass/fail behavior observed from the test chips. Our path-based learning methodology consists of four major components: a delay test pattern set, a logic simulator, a set of selected paths as the basis for learning, and a machine learner. We explain the key concepts in this methodology and present experimental results to demonstrate its feasibility and applications. Li-C. Wang, Kwang-Ting Cheng, Magdy S. Abadir |
DAC | 3 |
| 2004 | Improved Symoblic Simulation by Dynamic Funtional Space PartitioningabstractIn this paper, we provide a flexible and automatic method to partition the functional space for efficient symbolic simulation. We utilize a 2-tuple list representation as the basis for partitioning the functional space. The partitioning is carried out dynamically during the symbolic simulation based on the sizes of OBDDs. We develop heuristics for choosing the optimal partitioning points. These heuristics intend to balance the tradeoff between the time and space complexity. We demonstrate the effectiveness of our new symbolic simulation approach through experiments based on a floating point adder and a memory management unit. Tao Feng 0012, Li-C. Wang, Kwang-Ting Cheng, Chih-Chan Lin |
DATE | 3 |
| 2004 | Pattern Selection for Testing of Deep Sub-Micron Timing DefectsabstractDue to process variations in deep sub-micron (DSM) technologies, the effects of timing defects are difficult to capture. This paper presents a novel coverage metric for estimating the test quality with respect to timing defects under process variations. Based on the proposed metric and a dynamic timing analyzer, we develop a pattern-selection algorithm for selecting the minimal number of patterns that can achieve the maximal test quality. To shorten the run time in dynamic timing analysis, we propose an algorithm to speed up the Monte-Carlo-based simulation. Our experimental results show that, selecting a small percentage of patterns from a multiple-detection transition fault pattern set is sufficient to maintain the test quality given by the entire pattern set. We present run-time and accuracy comparisons to demonstrate the efficiency and effectiveness of our pattern selection framework. Mango Chia-Tso Chao, Li-C. Wang, Kwang-Ting Cheng |
DATE | 3 |
| 2004 | Random Jitter Extraction Technique in a Multi-Gigahertz SignalabstractIn this paper, we propose a simple technique for estimating the standard deviation of a Gaussian random jitter component in a multi-gigahertz signal. This method may utilize existing on-chip single-shot period measurement techniques to measure the multi-gigahertz signal periods for the estimation. This method does not require an external sampling clock, or any additional measurement beyond existing techniques. Experimental results show that this extraction method can accurately estimate the random jitter variance in a multi-gigahertz signal even with the presence of a few hundred-hertz sinusoidal jitter components. Chee-Kian Ong, Dongwoo Hong, Kwang-Ting Cheng, Li-C. Wang |
DATE | 3 |
| 2004 | A path-based methodology for post-silicon timing validationabstractThis work presents a novel path-based methodology for post-silicon timing validation. In timing validation, the objective is to decide if the timing behavior observed from the silicon is consistent with that predicted by the timing model. At the core of our path-based methodology, we propose a framework to obtain the post-silicon path ranking from observing silicon timing behavior. Then, the consistency is determined by comparing the post-silicon path ranking and the pre-silicon path ranking calculated based on the timing model. Our post-silicon ranking methodology consists of two approaches: ranking optimization and path filtering. We discuss the applications of both approaches and their impacts on the path ranking results. For experiments, we utilize a statistical timing simulator that was developed in the past to derive chip samples and we demonstrate the feasibility of our methodology using benchmark circuits. Leonard Lee, Li-C. Wang, Kwang-Ting Cheng |
ICCAD | 4 |
| 2004 | Static statistical timing analysis for latch-based pipeline designsabstractA latch-based timing analyzer is an essential tool for developing high-speed pipeline designs. As process variations increasingly influence the timing characteristics of DSM designs, a timing analyzer capable of handling process-induced timing variations for latch-based pipeline designs becomes in demand. In this work, we present a static statistical timing analyzer, STAP, for latch-based pipeline designs. Our analyzer propagates statistical worst-case delays as well as critical probabilities across the pipeline stages. We present an efficient method to handle correlations due to re-convergent fanouts. We also demonstrate the impact of not including the analysis of reconvergent fanouts in latch-based pipeline designs. Comparing to a Monte-Carlo based timing analyzer, our experiments show that STAP can accurately evaluate the critical probability that a design violates the timing constraints under a given statistical timing model. The runtime comparison further demonstrates the efficiency of our STAP. Rob A. Rutenbar, Li-C. Wang, Kwang-Ting Cheng, Sandip Kundu |
ICCAD | 3 |
| 2004 | A unified adaptive approach to accurate skin detectionabstractDue to variations of lighting conditions and camera hardware settings and the existence of many ethnic people with a wide range of skin colors, a generic skin model is often inadequate to accurately capture the skin distribution for individual images. In this paper, we propose an adaptive skin detection framework, which allows modeling title skin distribution with significantly higher accuracy and flexibility. First, an adaptive skin model, specific to the image under consideration and refined from the skin-similar space, is derived using a Gaussian mixture model (GMM) and standard expectation maximization (EM) algorithm. Then, we develop a support vector machine (SVM) classifier to identify the skin Gaussian from the trained GMM (with two Gaussian components) by incorporating spatial and shape information of skin pixels. Extensive experimental results performed on large image databases have demonstrated the effectiveness and benefits of the proposed approach. Qiang Zhu 0006, Kwang-Ting Cheng, Ching-Tung Wu |
ICIP | 2 |
| 2004 | SSD tracking using dynamic template and log-polar transformationabstractWith the assumption of small motion displacements in sequences, templates can be matched successfully by SSD (sum-squared-difference) optimization technique. In this paper two crucial modifications are proposed to improve the original SSD tracking. First, templates are dynamically updated for each frame. Thus, instead of designing complicated parametric models to explain various image distortions in tracking sequences, more efficient and compact parametric models, say the pure translation model, can be adopted. Second, rotation and scale can be incorporated into a translation motion model after a log-polar transformation. Extensive experimental results performed on live video sequences show the effectiveness and advantage of our approach. Qiang Zhu 0006, Kwang-Ting Cheng, HongJiang Zhang |
ICME | 2 |
| 2004 | BER Estimation for Serial Links Based on Jitter Spectrum and Clock Recovery CharacteristicsabstractHigh performance serial communication systems often require the bit error rate (BER) to be at the level of 10/sup -12/ or below. The excessive test time for measuring such a low BER is a major hindrance in testing communication systems cost-effectively. We propose a new technique for accurate and efficient estimation of the BER. The proposed technique estimates the BER based on the spectral information of jitter and the characteristics of the clock and data recovery circuit. The method can significantly reduce the production test time for BER testing. Simulation results demonstrate the potential usefulness of the method. Dongwoo Hong, Chee-Kian Ong, Kwang-Ting Cheng |
ITC | 3 |
| 2004 | An adaptive skin model and its application to objectionable image filteringabstractWe propose an adaptive skin-detection method, which allows modelling and detection of the true skin-color pixels with significantly higher accuracy and flexibility than previous methods. In principle, the proposed approach follows a two-step process. For a given image, we first perform a rough skin classification using a generic skin-model which defines the Skin-Similar space. The Skin-Similar space often contains many non-skin pixels due to the inevitable overlap in the color space between skin pixels and some non-skin pixels under the generic skin-model. The objective of the second step is to reduce the false-positive rate by analyzing the image under consideration. Specifically, in the second step, a Gaussian Mixture Model (GMM), specific to the image under consideration and refined from its Skin-Similar space, is derived using the standard Expectation-Maximization (EM) algorithm. We then use a Support Vector Machine (SVM) classifier to identify the skin Gaussian from the trained GMM by incorporating spatial and shape information of the skin pixels. Moreover, we examine how the improvement on skin detection by this adaptive skin-model impacts the detection accuracy in the application of Objectionable Image Filtering. We further propose a two-level classification scheme based on hierarchical bagging to improve the accuracy. Results of extensive experiments on large databases demonstrate the effectiveness and benefits of our adaptive skin-model. Qiang Zhu 0006, Ching-Tung Wu, Kwang-Ting Cheng, Yi-Leh Wu |
ACM Multimedia | 3 |
| 2004 | A Scalable On-Chip Jitter Extraction TechniqueabstractIn this paper, we propose a method for extracting the spectral information of a multi-gigahertz jittery signal. This method utilizes existing on-chip single-shot period measurement techniques to sample and measure the period of multiple cycles of the multi-gigahertz periodic signal for spectral analysis. Since measurements are made on the period of multiple cycles, but not on the period of a single cycle, a lower-speed timing measurement circuitry can be used to measure a higher-speed signal. Therefore, the proposed solution is scalable for even higher-speed signals. This method does not require an external sampling clock, nor any additional measurement beyond existing techniques. Experimental results based on simulation show that this method can accurately estimate the sinusoidal and random jitters of a multi-gigahertz signal. Chee-Kian Ong, Dongwoo Hong, Kwang-Ting Cheng, Li-C. Wang |
VTS | 3 |
| 2004 | Self-referential verification for gate-level implementations of arithmetic circuitsabstractVerification of gate-level implementations of arithmetic circuits is challenging for a number of reasons: the existence of some hard-to-verify arithmetic operators, the use of different operand ordering, the incorporation of merged arithmetic with cross-operator implementations, and the employment of circuit transformations based on arithmetic relations. It is hence a peculiar problem that does not fit well within the existing register-transfer-level-to-gate equivalence-checking methodology. We propose a self-referential functional verification approach which uses the gate-level implementation of the arithmetic circuit under verification to verify itself. The verification task is decomposed into a sequence of equivalence-checking subproblems, each of which compares structurally similar circuit pairs derived from the implementation under verification. These equivalence-checking subproblems represent the functional equations that uniquely define the intended arithmetic function. Based on these self-referential functional equations, a decomposition heuristic using structural information is employed to guide the verification process for better efficiency. Experimental results on a number of implementations of the multipliers, the multiply-add units, and the inner product units with different architectures demonstrate the versatility of this approach. Ying-Tsai Chang, Kwang-Ting Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2004 | Critical path selection for delay fault testing based upon a statistical timing modelabstractCritical path selection is an indispensable step for testing of small-size delay defects. Historically, this step relies on the construction of a set of worst-case paths, where the timing lengths of the paths are calculated based upon discrete-valued timing models. The assumption of discrete-valued timing models may become invalid for modeling delay effects in the deep submicron domain, where the effects of timing defects and process variations are often statistical in nature. This paper studies the problem of critical path selection for testing small-size delay defects, assuming that circuit delays are statistical. We provide theoretical analysis to demonstrate that the new path-selection problem consists of two computationally intractable subproblems. Then, we discuss practical heuristics and their performance with respect to each subproblem. Using a statistical defect injection and timing-simulation framework, we present experimental results to support our theoretical analysis. Li-C. Wang, Jing-Jia Liou, Kwang-Ting Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2003 | Enhanced symbolic simulation for efficient verification of embedded array systemsabstractwas shown to be effective for verifying individual array blocks. However, when applying STE to verify multiple array blocks together as a single system, the run-time OBDD sizes would often blow up. In this paper, we propose using a ”dual-rail ” symbolic simulation scheme to facilitate the application of STE proof methodology for verifying array systems. The proposed scheme implicitly partitions a given design into control domain and datapath domain, and symbolic simulation is carried out on both domains. With this scheme, the run-time OBDD sizes during the symbolic simulation for each domain can be limited. We demonstrate the effectiveness of our approach by verifying the Memory Management Unit (MMU) in Motorola high-performance microprocessors. The verification of MMU as a whole was not possible before because of the OBDD size blow-up problem when an ordinary symbolic simulator was used in the STE proof process. I. Tao Feng 0012, Li-C. Wang, Kwang-Ting Cheng, Magdy S. Abadir |
ASP-DAC | 3 |
| 2003 | Experience in critical path selection for deep sub-micron delay test and timing validationabstractCritical path selection is an indispensable step for AC delay test and timing validation. Traditionally, this step relies on the construction of a set of worse-case paths based upon discrete timing models. However, the assumption of discrete timing models can be invalidated by timing defects and process variation in the deep sub-micron domain, which are often continuous in nature. As a result, critical paths defined in a traditional timing analysis approach may not be truly critical in reality. In this paper, we propose using a statistical delay evaluation framework for estimating the quality of a path set. Based upon the new framework, we demonstrate how the traditional definition of a critical path set may deviate from the true critical path set in the deep sub-micron domain. To remedy the problem, we discuss improvements to the existing path selection strategies by including new objectives. We then compare statistical approaches with traditional approaches based upon experimental analysis of both defect-free and defect-injected cases. Jing-Jia Liou, Li-C. Wang, Angela Krstic, Kwang-Ting Cheng |
ASP-DAC | 4 |
| 2003 | Delta-sigma modulator based mixed-signal BIST architecture for SoCabstractThis paper proposes a mixed-signal Built-In Self-Test (BIST) architecture based on a second-order delta-sigma modulator. This modulator, which incorporates a design-for-testability (DfT) circuitry, is capable of testing/characterizing itself using digital stimulus. This characteristic is attractive for implementing the modulator as an on-chip analog signal analyzer. When applied for mixed-signal BIST, the modulator-based analog signal analyzer is first characterized using digital stimulus. Then the analyzer is utilized to characterize the stimulus generator in the BIST application. Some critical implementation issues of the BIST architecture are also discussed. Chee-Kian Ong, Kwang-Ting Cheng, Li-C. Wang |
ASP-DAC | 2 |
| 2003 | Enhancing diagnosis resolution for delay defects based upon statistical timing and statistical fault modelsabstractIn this paper, we propose a new methodology for diagnosis of delay defects in the deep sub micron domain. The key difference between our diagnosis framework and other traditional diagnosis methods lies in our assumptions of the statistical circuit timing and the statistical delay defect size. Due to the statistical nature of the problem, achieving 100% diagnosis resolution cannot be guaranteed. To enhance diagnosis resolution, we propose a 3-phase diagnosis methodology. In the first phase, our goal is to quickly identify a set of candidate suspect faults that are most likely to cause the failing behavior based on logic constraints. In the second phase, we obtain a much smaller suspect fault set by applying a novel diagnosis algorithm that can effectively utilize the statistical timing information based upon a single defect assumption. In the third phase, our goal is to apply additional fine-tuned patterns to successfully narrow down to more exact suspect defect locations. Using a statistical timing analysis framework, we demonstrate the effectiveness of the proposed methodology for delay defect diagnosis, and discuss experimental results based on benchmark circuits. Angela Krstic, Li-C. Wang, Kwang-Ting Cheng, Jing-Jia Liou |
DAC | 3 |
| 2003 | A signal correlation guided ATPG solver and its applications for solving difficult industrial casesabstractThe developments of efficient SAT solvers have attracted tremendous research interest in recent years. The merits of these solvers are often compared in terms of their performance based upon a wide spread of benchmarks. In this paper, we extend an earlier-proposed solver design concept called (SCGL) Signal Correlation Guided Learning that is ATPG-based into a family of heuristics. Along with this SCGL family of heuristics, we classify benchmark examples according to their performance using the SCGL heuristics. With this study, we identify the class of problems that are uniquely suitable to be solved by using the SCGL approach. In particular, for solving difficult circuit-based problems at INTEL, our SCGL-based ATPG solver is able to achieve at least an order of magnitude speedup over the state-of-the-art SAT solvers. Our conclusion is that SCGL is an unique solver design concept that can complement heuristics proposed by others for solving circuit-oriented difficult problems. Feng Lu 0002, Li-C. Wang, Kwang-Ting Cheng, John Moondanos, Ziyad Hanna |
DAC | 3 |
| 2003 | Delay Defect Diagnosis Based Upon Statistical Timing Models - The First Step
Angela Krstic, Li-C. Wang, Kwang-Ting Cheng, Jing-Jia Liou, Magdy S. Abadir |
DATE | 3 |
| 2003 | A Circuit SAT Solver With Signal Correlation Guided Learning
Feng Lu 0002, Li-C. Wang, Kwang-Ting Cheng, Ric C.-Y. Huang |
DATE | 3 |
| 2003 | SATORI - A Fast Sequential SAT Engine for Circuits
Madhu K. Iyer, Ganapathy Parthasarathy, Kwang-Ting Cheng |
ICCAD | 3 |
| 2003 | The Confluence of Manufacturing Test and Design ValidationabstractThe potential of shared solutions in semiconductor technology is discussed. A study was conducted in the areas of delay testing, timing verification, dynamic verification and online testing. The conclusion was that there is a requirement for design, test and verification techniques that can detect the errors and provide tolerance and reliable computation. Kwang-Ting Cheng |
ITC | 1 |
| 2003 | Diagnosis-Based Post-Silicon Timing Validation Using Statistical Tools and MethodologiesabstractThis paper describes a new post-silicon validation problem for diagnosing systematic timing errors. We illustrate the differences between timing validation and the traditional logic defect diagnosis. The key difference between our validation framework and other traditional diagnosis methods lies in our assumptions of the statistical circuit timing and statistical distribution of the size of timing errors. Different algorithms are proposed and evaluated via statistical timing error injection and simulation. Due to the statistical nature of the problem, 100% diagnosis resolution often cannot be guaranteed. With a statistical timing analysis framework developed in the past, we demonstrate the new concepts in timing validation, and discuss experimental results based upon three types of systematic errors: timing correlation error, crosstalk, and single-site random-size delay perturbation. Angela Krstic, Li-C. Wang, Kwang-Ting Cheng |
ITC | 3 |
| 2003 | Using Logic Models To Predict The Detection Behavior Of Statistical Timing DefectsabstractIn this paper, we study the possibility of using logic defect-level prediction models to predict the detection behavior of statistical timing defects. We compare two known logic models: the Williams-Brown (WB) model and the Mercer-Park-Grimaila-Dworak (MPGD) model. In the WB-model, the defect coverage is replaced by the n-detection transition fault coverage. We first demonstrate that both logic models may fail to predict the detection of statistical timing defects. Then, we propose an improved WB model based upon selection of the hard-to-detect transition faults. We show that, by selecting a proper subset of the hard-to-detect transition faults, the detection behavior of these faults can correlate well to the detection behavior of statistical timing defects. We explain our findings through statistical delay defect injection and simulation, and report results based upon various benchmark circuits. Li-C. Wang, Angela Krstic, Leonard Lee, Kwang-Ting Cheng, M. Ray Mercer, Thomas W. Williams, Magdy S. Abadir |
ITC | 4 |
| 2003 | Diagnosis of Delay Defects Using Statistical Timing ModelsabstractIn this paper, we study the problem of delay defect diagnosis based on statistical timing models. We propose a diagnosis algorithm that can effectively utilize statistical timing information based upon single defect assumption. We evaluate its performance and its applicability to single as well as multiple defect scenarios via statistical defect injection and simulation. With a statistical timing analysis framework developed in the past, we demonstrate the new concept in statistical delay defect diagnosis, and discuss experimental results using benchmark circuits. Angela Krstic, Li-C. Wang, Kwang-Ting Cheng, Jing-Jia Liou |
VTS | 3 |
| 2003 | Embedded Tutorial: Test Consideration for Nanometer Scale CMOS CircuitsabstractThe ITRS (international technology roadmap for semiconductors) predicts aggressive scaling down of device size, transistor threshold voltage and oxide thickness to meet growing demands for performance. Such scaling will result in an exponential increase in leakage current and large variability in threshold voltage both within and across dies. Device counts will increase from about 0.2 B/chip today to approximately 10 B/chip in a decade. This 50/spl times/ increase in device count will increase not only the active power dissipation, but also the standby or the quiescent power. Hence, designers are required to use innovative aggressive power management strategies to meet the power constraints. The exponential increase in leakage, the device parameter variations, and aggressive power management techniques are expected to severely impact the way integrated circuits are tested today. This paper explores test considerations for the scaled CMOS circuits in the nanometer regime. Kaushik Roy 0001, Kwang-Ting Cheng |
VTS | 3 |
| 2003 | Modeling, testing, and analysis for delay defects and noise effects in deep submicron devicesabstractThe performance of deep submicron designs can be affected by various parametric variations, manufacturing defects, noise or modeling errors that are all statistical in nature. In this paper, we propose a methodology to capture the effects of these statistical variations on circuit performance. It incorporates statistical information into timing analysis to compute the performance sensitivity of internal signals subject to a given type of defect, noise or variation sources. Next, we propose a novel path and segment selection methodology for delay testing based on the results of statistical performance sensitivity analysis. The objective of path/segment selection is to identify a small set of paths and segments such that the delay tests for the selected paths/segments guarantee the detection of performance failure. We apply the proposed path selection technique for selection of a set of paths for dynamic timing analysis considering power supply noise effects. Our experimental results demonstrate the difference in estimated circuit performance for the case when power supply noise effects are considered versus when these effects are ignored. Thus, they indicate the need for considering power supply noise effects on delays during path selection and dynamic timing analysis. Jing-Jia Liou, Angela Krstic, Yi-Min Jiang, Kwang-Ting Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2002 | On-chip Analog Response Extraction with 1-Bit ? - ModulatorsabstractBecause of their relative robustness to process variation, /spl Sigma/-/spl Delta/ modulation techniques are particularly suitable for VLSI implementations. In this paper, we propose to employ the 1-bit /spl Sigma/-/spl Delta/ modulation ADC (analog-to-digital converter) as the on-chip analog response extractor for analog/mixed-signal BIST (built-in self-test) applications. To validate the idea, a prototype chip with the proposed BIST circuitry has been designed and fabricated. Performance of the BIST circuitry is validated (up to 87 dB dynamic range), and measurement results of the circuit under test (CUT), a 2nd-order low-pass filter, are presented. Hao-Chiao Hong, Jiun-Lang Huang, Kwang-Ting Cheng, Cheng-Wen Wu |
Asian Test Symposium | 3 |
| 2002 | Self-referential verification of gate-level implementations of arithmetic circuitsabstractVerification of gate-level implementations of arithmetic circuits is challenging due to a number of reasons: the existence of some hard-to-verify arithmetic operators (e.g. multiplication), the use of different operand ordering, the incorporation of merged arithmetic with cross-operator implementations, and the employment of circuit transformations based on arithmetic relations. It is hence a peculiar problem that does not fit quite well into the existing RTL-to-gate equivalence checking methodology. In this paper, we propose a self-referential functional verification approach which uses the gate-level implementation of the arithmetic circuit under verification to verify itself. Specifically, the verification task is decomposed into a sequence of equivalence checking subproblems, each of which compare circuit pairs derived from the implementation under verification based on the proposed self-referential functional equations. A decomposition-based heuristic using structural information is employed to guide the verification process for better efficiency. Experimental results on a number of implementations of the multiply-add units and the inner product units with different architectures demonstrate the versatility of this approach. Ying-Tsai Chang, Kwang-Ting Cheng |
DAC | 2 |
| 2002 | Embedded software-based self-testing for SoC designabstractAt-speed testing of high-speed circuits is becoming increasingly difficult with external testers due to the growing gap between design and tester performance, growing cost of high-performance testers and increasing yield loss caused by inherent tester inaccuracy. Therefore, empowering the chip to test itself seems like a natural solution. Hardware-based self-testing techniques have limitations due to performance and area overhead and problems caused by the application of non-functional patterns.Embedded software-based self-testing has recently become focus of intense research. In this methodology, the programmable cores are used for on-chip test generation, measurement, response analysis and even diagnosis. After the programmable core on a System-on-Chip (SoC) has been self-tested, it can be reused for testing on-chip buses, interfaces and other non-programmable cores. The advantages of this methodology include at-speed testing, low design-for-testability overhead and application of functional patterns in the functional environment. In this paper, we give a survey and outline the roadmap and challenges of this emerging embedded software-based self-testing paradigm. Angela Krstic, Wei-Cheng Lai, Kwang-Ting Cheng, Sujit Dey |
DAC | 3 |
| 2002 | False-path-aware statistical timing analysis and efficient path selection for delay testing and timing validationabstractWe propose a false-path-aware statistical timing analysis framework. In our framework, cell as well as interconnect delays are assumed to be correlated random variables. Our tool can characterize statistical circuit delay distribution for the entire circuit and produce a set of true critical paths. Jing-Jia Liou, Angela Krstic, Li-C. Wang, Kwang-Ting Cheng |
DAC | 4 |
| 2002 | Enhancing test efficiency for delay fault testing using multiple-clocked schemesabstractIn conventional delay testing, the test clock is a single pre-defined parameter that is often set to be the same as the system clock. This paper discusses the potential of enhancing test efficiency by using multiple clock frequencies. The intuition behind our work is that for a given set of AC delay patterns, a carefully-selected, tighter clock would result in higher effectiveness to screen out the potential defective chips. Then, by using a smarter test clock scheme and combining with a second set of AC delay patterns, the overall quality of AC delay test can be enhanced while the cost of including the second pattern set can be minimized. We demonstrate these concepts through analysis and experiments using a statistical timing analysis framework with defect-injected simulation. Jing-Jia Liou, Li-C. Wang, Kwang-Ting Cheng, Jennifer Dworak, M. Ray Mercer, Rohit Kapur, Thomas W. Williams |
DAC | 3 |
| 2002 | On theoretical and practical considerations of path selection for delay fault testingabstractIn current industrial practice, critical path selection is an indispensable step for AC delay test and timing validation. Traditionally, this step relies on the construction of a set of worse-case paths based upon discrete timing models. The assumption of discrete timing models can be invalidated by delay effects in the deep sub-micron domain, where timing defects and process variation are statistical in nature. In this paper, we study the problem of optimizing critical path selection, under both fixed delay and statistical delay assumptions. With a novel problem formulation and new theoretical results, we prove that the problem in both cases are computationally intractable. We then discuss practical heuristics and their theoretical performance bounds, and demonstrate that among all heuristics under consideration, only one is theoretically feasible. Finally, we provide consistent experimental results based upon defect-injected simulation using an efficient statistical timing analysis framework. Jing-Jia Liou, Li-C. Wang, Kwang-Ting Cheng |
ICCAD | 3 |
| 2002 | Analysis of Delay Test Effectiveness with a Multiple-Clock SchemeabstractIn conventional delay testing, two types of tests, transition tests and path delay tests, are often considered. The test clock frequency is usually set to a single pre-determined parameter equal to the system clock. This paper discusses the potential of enhancing test effectiveness by using multiple test sets with multiple clock frequencies. The two intuitions motivating our analysis are 1) multiple test sets can deliver higher test quality than a single test set, and 2) for a given set of AC delay patterns, a carefully-selected, tighter clock would result in higher effectiveness to screen out potentially defective chips. Hence, by using multiple test sets, the overall quality of AC delay test can be enhanced, and by using multiple-clock schemes the cost of adding the additional pattern sets can be minimized. In this paper, we analyze the feasibility of this new delay test methodology with respect to different combinations of pattern sets and to different circuit characteristics. We discuss the pros and cons of multiple-clock schemes through analysis and experiments using a statistical delay evaluation and delay defect-injected framework. Jing-Jia Liou, Li-C. Wang, Kwang-Ting Cheng, Jennifer Dworak, M. Ray Mercer, Rohit Kapur, Thomas W. Williams |
ITC | 3 |
| 2002 | Combining ATPG and Symbolic Simulation for Efficient Validation of Embedded Array SystemsabstractIn the past, symbolic trajectory evaluation (STE) has been shown to be effective for verifying individual array blocks. However, when applying STE to verify multiple array blocks together as a single system, the run-time OBDD (ordered boolean decision diagrams) sizes would often blow up. In this paper, we propose the use of both an ATPG-based justification engine and symbolic simulation to facilitate the application of STE proof methodology for array systems. Our method translates a given verification problem instance into ATPG justification objectives, and partitions a given design into ATPG and symbolic simulation domains. Then, by developing a scheme that enables the ATPG justification engine to work closely with the symbolic simulator, the runtime OBDD sizes during each symbolic simulation run can be limited. We demonstrate the effectiveness of our approach by verifying the memory management units (MMU) in Motorola high-performance microprocessors. The verification of a MMU as a whole was not possible before because of the OBDD size blow-up problem when symbolic simulation is used in the STE proof process. Ganapathy Parthasarathy, Madhu K. Iyer, Tao Feng 0012, Li-C. Wang, Kwang-Ting Cheng, Magdy S. Abadir |
ITC | 5 |
| 2002 | PBIR-MM: multimodal image retrieval and annotationabstractWe demonstrate PBIR-MM, an integrated system that we have built for conducting multimodal image retrieval. The system combines the strengths of content-based soft annotation (CBSA), multimodal relevance feedback through active learning, and perceptual distance formulation and indexing. PBIR-MM supports multimodal query and annotation in any combination of its three basic modes: seed-by-nothing, seed-by-keywords, and seed-by-content. We demonstrate PBIR-MM on a couple of very large image sets provided by image vendors and crawled from the Internet. Wei-Cheng Lai, Chengwei Chang, Edward Y. Chang, Kwang-Ting Cheng, Michael Crandell |
ACM Multimedia | 4 |
| 2002 | Software-Based Weighted Random Testing for IP Cores in Bus-Based Programmable SoCsabstractPresents a software-based weighted random pattern scheme for testing delay faults in IP cores of programmable SoCs. We describe a method for determining static and transition probabilities (profiles) at the inputs of circuits with full-scan using testability metrics based on the targeted fault model, We use a genetic algorithm (GA) based search procedure to determine optimal profiles. We use these optimal profiles to generate a test program that runs on the processor core. This program applies test patterns to the target IP cores in the SoC and analyzes the test responses. This provides the flexibility of applying multiple profiles to the IP core under test to maximize fault coverage. This scheme does not incur the hardware overhead of logic BIST, since the pattern generation and analysis is done by software. We use a probabilistic approach to finding the profiles. We describe our method on transition and path-delay fault models, for both enhanced full-scan and normal full-scan circuits. We present experimental results using the ISCAS 89 benchmarks as IP cores. Madhu K. Iyer, Kwang-Ting Cheng |
VTS | 2 |
| 2002 | Self-Testing Second-Order Delta-Sigma Modulators Using Digital StimulusabstractSingle-bit second-order delta-sigma modulators are commonly used in high-resolution ADCs. Testing this type of modulator requires a high-resolution test stimulus, which is difficult to generate. This paper proposes a novel and robust technique to determine the performance of the modulator, which incorporates simple design-for-testability circuitry. This technique requires only digital stimulus to test the modulator. Hence, it is suitable as an analog signature analyzer used in built-in self-test applications. Simulation results show that this technique is capable of accurately determining the performance of a second-order delta-sigma modulator ADC. Chee-Kian Ong, Kwang-Ting Cheng |
VTS | 2 |
| 2001 | SVM Binary Classifier Ensembles for Image ClassificationabstractWe study how the SVM-based binary classifiers can be effectively combined to tackle the multi-class image classification problem. We study several ensemble schemes, including OPC (one per class), PWC (pairwise coupling), and ECOC (error-correction output coding), that aim to achieve good error correction capability through redundancy. To enhance these ensemble schemes' accuracy, we propose methods that on the one hand boost the margins (i.e., confidence) of the SVM-based binary classifiers, and, on the other hand, remove the noise of irrelevant classifiers from class prediction. From empirical study we show that our margin boosting and noise reduction methods lead to higher classification accuracy than ensemble schemes that are solely designed for maximum error correction capability. Kingshy Goh, Edward Y. Chang, Kwang-Ting Cheng |
CIKM | 3 |
| 2001 | Instruction-Level DFT for Testing Processor and IP Cores in System-on-a-ChipabstractSelf-testing manufacturing defects in a system-on-a-chip (SOC) by running test programs using a programmable core has several potential benefits including, at-speed test-ing, low DfT overhead due to elimination of dedicated test circuitry and better power and thermal management during testing. However, such a self-test strategy might require a lengthy test program and might achieve a high enough fault coverage. We propose a DfT methodlogy to improve the fault coverage and reduce the test program length, by adding test instructions to an on-chip programmable core such as a microprocessor core. This paper discusses a method of identifying effective test instructions which could result in highest benefits with low area/performance over-head. The experimental results show that with the added test instructions, a complete fault coverage for testable path delay faults can be achieved with a greater than 20% reduction in the program size and the program runtime, as compared to the case without instruction-level DfT. Wei-Cheng Lai, Kwang-Ting Cheng |
DAC | 2 |
| 2001 | Fast Statistical Timing Analysis By Probabilistic Event PropagationabstractWe propose a new statistical timing analysis algorithm, which produces arrival-time random variables for all internal signals and primary outputs for cell-based designs with all cell delays modeled as random variables. Our algorithm propagates probabilistic timing events through the circuit and obtains final probabilistic events (distributions) at all nodes. The new algorithm is deterministic and flexible in controlling run time and accuracy. However, the algorithm has exponential time complexity for circuits with reconvergent fanouts. In order to solve this problem, we further propose a fast approximate algorithm. Experiments show that this approximate algorithm speeds up the statistical timing analysis by at least an order of magnitude and produces results with small errors when compared with Monte Carlo methods. Jing-Jia Liou, Kwang-Ting Cheng, Sandip Kundu, Angela Krstic |
DAC | 2 |
| 2001 | Induction-Based Gate-Level Verification of MultipliersabstractWe propose a method based on unrolling the inductive definition of binary number multiplication to verify gate-level implementations of multipliers. The induction steps successively reduce the size of the multiplier under verification. Through induction, the verification of an n-bit multiplier is decomposed into n equivalence checking problems. The resulting equivalence checking problems could be significantly sped up by simple structural analysis. This method could be generalized to the verification of more general arithmetic circuits and the equivalence checking of complex datapaths. Ying-Tsai Chang, Kwang-Ting Cheng |
ICCAD | 2 |
| 2001 | Mining Image Features for Efficient Query ProcessingabstractThe number of features required to depict an image can be very large. Using all features simultaneously to measure image similarity and to learn image query-concepts can suffer from the problem of dimensionality curse, which degrades both search accuracy and search speed. Regarding search accuracy, the presence of irrelevant features with respect to a query can contaminate similarity measurement, and hence decrease both the recall and precision of that query. To remedy this problem, we present a mining method that learns online users' query concepts and identifies important features quickly. Regarding search speed, the presence of a large number of features can slow down query-concept learning and indexing performance. We propose a divide-and-conquer method that divides the concept-learning task into G subtasks to achieve speedup. We notice that a task must be divided carefully, or search accuracy may suffer. We thus propose a genetic-based mining algorithm to discover good feature groupings. Through analysis and mining results, we observe that organizing image features in a multi-resolution manner and minimizing intra-group feature correlation, can speed up query-concept learning substantially while maintaining high search accuracy. Beitao Li, Wei-Cheng Lai, Edward Y. Chang, Kwang-Ting Cheng |
ICDM | 4 |
| 2001 | Delay testing considering crosstalk-induced effectsabstractIncreased noise/interference effects, such as crosstalk, power supply noise, substrate noise and distributed delay variations lead to increased signal integrity problems in deep submicron designs. These problems can cause logic errors and/or performance degradation and must be addressed both in the design for deep submicron and testing for deep submicron phases. Existing delay testing techniques cannot capture the effects of noise on the cell/interconnect delays. In this paper, we address the problem of delay testing considering crosstalk-induced delay effects. We propose solutions for target fault selection and pattern generation. The key elements of our strategy are performance sensitivity analysis with respect to crosstalk noise and a genetic algorithm (GA) based vector generation technique. The role of performance sensitivity analysis is to consider the effects of crosstalk noise during the target fault selection process. Next, for each selected fault consisting of a path and a set of crosstalk noise sources interacting with the path, we apply our iterative GA-based pattern generation process. Our goal is to derive a test that produces a large crosstalk-induced delay effect on the given path. Our technique allows consideration of any number of coupling sources along the target path. Due to its flexibility, efficiency and scalability, the technique can be applied to large circuits. Angela Krstic, Jing-Jia Liou, Yi-Min Jiang, Kwang-Ting Cheng |
ITC | 4 |
| 2001 | PBIR: perception-based image retrieval-a system that can quickly capture subjective image query conceptsabstractWe describe the Perception-Based Image Retrieval (PBIR) system that we have built on our recently developed query-concept learning algorithms, MEGA and SVMActive. We show that MEGA and SVMActive can learn a complex image-query concept in a small number of user iterations (usually three to four) on a large, multi-category, high-dimensional image database. Edward Y. Chang, Kwang-Ting Cheng, Wei-Cheng Lai, Ching-Tung Wu, Chengwei Chang, Yi-Leh Wu |
ACM Multimedia | 2 |
| 2001 | PBIR - Perception-Based Image RetrievalabstractWe demonstrate a system that we have built on our proposed perception-based image retrieval (PBIR) paradigm. This PBIR system achieves accurate similarity measurements by rooting image characterization in human perception and by learning user's query concept through an intelligent sampling process. We show that our system can usually grasp a user's query concept with a small number of labeled instances. Edward Y. Chang, Kwang-Ting Cheng, Lihyuarn L. Chang |
SIGMOD Conference | 2 |
| 2001 | An On-Chip Short-Time Interval Measurement Technique for Testing High-Speed Communication LinksabstractIn this paper, we present a BIST scheme for on-chip short-time interval measurement intended for characterizing the time-domain specifications, e.g., the rise/fall time of modern high-speed communication transceivers. To reduce hardware overhead, the proposed BIST technique uses the coherent under-sampling principle, and measures implicitly the time interval in a two-pass manner. Simulation results are shown to validate the proposed technique. Jiun-Lang Huang, Kwang-Ting Cheng |
VTS | 2 |
| 2001 | A Self-Test Methodology for IP Cores in Bus-Based Programmable SoCsabstractWe present a novel test methodology for testing IP cores in SoCs with embedded processor cores. A test program is run on the processor core that generates and delivers test patterns to the target IP cores in the SoC and analyzes the test responses. This provides tremendous flexibility in the type of patterns that can be applied to the IP cores without incurring significant hardware overhead. We use a bus based SoC simulation model to validate our test methodology. The test methodology involves addition of a test wrapper that can be configured for specific test needs. The methodology supports at-speed testing for delay faults and stuck-at testing of IP cores implementing full-scan. Jing-Reng Huang, Madhu K. Iyer, Kwang-Ting Cheng |
VTS | 3 |
| 2001 | Embedded-Software-Based Approach to Testing Crosstalk-Induced Faults at On-Chip BusesabstractCrosstalk effects on long interconnects are becoming significant for high-speed circuits. This paper addresses the problem of testing crosstalk-induced faults at on-chip buses in system-on-a-chip (SOC) designs. We propose a method to self-test on-chip buses at-speed, by executing an automatically synthesized program using on-chip processor cores. The test program, executed at system operational speed, can activate and capture the worst-case crosstalk effects on buses and achieve a complete coverage of crosstalk-induced logical and delay faults. This paper discusses the method and the framework for synthesizing such a test program. Based on the bus protocol, the instruction set architecture of an on-chip processor core, and the system specification, the method generates deterministic tests in the form of instruction sequences. The synthesized test program is highly modularized and compact. The experimental results show that, for testing interconnects between a processor core and any other on-chip core, a 3 K-byte program is sufficient to achieve the complete coverage for crosstalk-induced logical and delay faults. Wei-Cheng Lai, Jing-Reng Huang, Kwang-Ting Cheng |
VTS | 3 |
| 2001 | Limitations and challenges of computer-aided design technology for CMOS VLSIabstractAs manufacturing technology moves toward fundamental limits of silicon CMOS processing, the ability to reap the full potential of available transistors and interconnect is increasingly important. Design technology (DT) is concerned with the automated or semi-automated conception, synthesis, verification, and eventual testing of microelectronic systems. While manufacturing technology faces fundamental limits inherent in physical laws or material properties, design technology faces fundamental limitations inherent in the computational intractability of design optimizations and in the broad and unknown range of potential applications within various design processes. In this paper, we explore limitations to how design technology can enable the implementation of single-chip microelectronic systems that take full advantage of manufacturing technology with respect to such criteria as layout density performance, and power dissipation. Randal E. Bryant, Kwang-Ting Cheng, Andrew B. Kahng, Kurt Keutzer, Wojciech Maly, A. Richard Newton, Lawrence T. Pileggi, Jan M. Rabaey, Alberto L. Sangiovanni-Vincentelli |
Proc. IEEE | 2 |
| 2001 | Using word-level ATPG and modular arithmetic constraint-solvingtechniques for assertion property checkingabstractWe present a new approach to checking assertion properties for register-transfer level (RTL) design verification. Our approach combines structural word-level automatic test pattern generation (ATPG) and modular arithmetic constraint-solving techniques to solve the constraints imposed by the target assertion property. Our word-level ATPG and implication technique not only solves the constraints on the control logic, but also propagates the logic implications to the datapath. A novel arithmetic constraint solver based on modular number system is then employed to solve the remaining constraints in datapath. The advantages of the new method are threefold. First, the decision-making process of the word-lever ATPG is confined to the selected control signals only. Therefore, the enumeration of enormous number of choices at the datapath signals is completely avoided. Second, our new implication translation techniques allow word-level logic implication being performed across the boundary of datapath and control logic and, therefore, efficiently cut down the ATPG search space. Third, our arithmetic constraint solver is based on modular instead of integral number systems. It can thus avoid the false-negative effect resulting from the bit-vector value modulation. A prototype system has been built that consists of an industrial front-end hardware description language (HDL) parser, a property-to-constraint converter, and the ATPG/arithmetic constraint-solving engine. The experimental results on some public benchmark and industrial circuits demonstrate the efficiency of our approach and its applicability to large industrial designs. Chung-Yang Huang, Kwang-Ting Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2001 | Pattern generation for delay testing and dynamic timing analysisconsidering power-supply noise effectsabstractNoise effects such as power supply and crosstalk noise can significantly impact the performance of deep submicrometer designs. Existing delay testing and timing analysis techniques cannot capture the effects of noise on the signal/cell delays. Therefore, these techniques cannot capture the worst case timing scenarios and the predicted circuit performance might not reflect the worst case circuit delay. More accurate and efficient timing analysis and delay testing strategies need to be developed to predict and guarantee the performance of deep submicrometer designs. In this paper, we propose a new pattern generation technique for delay testing and dynamic timing analysis that can take into account the impact of the power supply noise on the signal propagation delays. In addition to sensitizing the selected paths, the new patterns also cause high power supply noise on the nodes in these paths. Thus, they also cause longer propagation delays for the nodes along the paths. Our experimental results on benchmark circuits show that the new patterns produce significantly longer delays on the selected paths compared to the patterns derived using existing pattern generation methods. Angela Krstic, Yi-Min Jiang, Kwang-Ting Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2001 | Verifying sequential equivalence using ATPG techniquesabstractIn this paper we address the problem of verifying the equivalence of two sequential circuits. State-of-the-art sequential optimization techniques such as retiming and sequential redundancy removal can handle designs with up to hundreds or even thousands of flip-flops. However, the BDD-based approaches for verifying sequential equivalence can easily run into memory explosion for such designs. In an attempt to handle larger circuits, we modify test pattern-generation techniques for verification. The suggested approach utilizes the popular efficient backward-justification technique used in most sequential ATPG programs. We present several techniques to enhance the efficiency of this approach by (1) identifying equivalent flip-flop pairs using an induction-based algorithm, and (2) generalizing the idea of exploring the structural similarity between circuits to perform verification in stages. This ATPG-based framework is suitable for verifying circuits either with or without a reset state. In order to extend this approach to verify retimed circuits, we introduce a delay-compensation-based algorithm for preprocessing the circuits. The experimental results of verifying the correctness of circuits after sequential redundancy removal and retiming with up to several hundred flip-flops are presented. Shi-Yu Huang, Kwang-Ting Cheng, Kuang-Chien Chen |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2001 | Vector generation for power supply noise estimation and verification of deep submicron designsabstractThis paper presents new techniques for generating a small set of patterns for power network simulation to estimate the maximum power supply noise of the chip, as well as to identify cells/blocks for which the power supply noise at their V/sub dd/ ports exceeds a specified threshold. We first present an efficient, cell-level simulator for estimating power supply noise of any given vectors. Based on this simulator, we then apply the genetic algorithm (GA) to derive a small set of patterns producing high power supply noise. To identify critical nodes with power supply noise exceeding a threshold, the multiobjective GA is adapted for pattern generation. To achieve high coverage of such critical nodes, we model the search criteria as the maximum weighted matching of a bipartite graph, and guide the search direction according to the matching results. The derived patterns will be simulated on a power network simulator to obtain a lower bound of the maximum power supply noise and to identify the critical nodes. Experimental results on public benchmark circuits, as well as some industrial designs, are presented to demonstrate the efficiency and effectiveness of the proposed approaches. Yi-Min Jiang, Kwang-Ting Cheng |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2000 | A sigma-delta modulation based BIST scheme for mixed-signal circuitsabstractIn this work, we present the analysis of a built-in self-test (BIST) scheme for mixed-signal circuits that is intended to provide on-chip stimulus generation and response analysis.Based on the sigma-delta modulation principle, the proposed scheme can produce high-quality stimuli and obtain accurate measurements without the need of precise analog circuitry.Numerical simulations are conducted to validate our idea and the results show that the scheme is a promising BIST approach for mixed-signal circuits. Jiun-Lang Huang, Kwang-Ting Cheng |
ASP-DAC | 2 |
| 2000 | Performance sensitivity analysis using statistical method and its applications to delayabstractThe performance of deep submicron designs can be affected by various parametric variations, manufacturing defects, noise or modeling errors that are all statistical in nature. We propose a statistical framework for analyzing the performance sensitivity of designs to various timing related defects/noise/variations. The core engine of our approach is a highly efficient statistical timing analysis tool. We describe the application of our framework for delay fault modeling and analysis of resistive opens and shorts and as well as interconnect crosstalk. We present experimental results demonstrating the accuracy of our statistical framework as compared to SPICE (for a given set of input patterns) and nominal worst-case analysis. Experimental results for analysis of resistive opens and shorts are also included. Jing-Jia Liou, Angela Krstic, Kwang-Ting Cheng, Deb Aditya Mukherjee, Sandip Kundu |
ASP-DAC | 3 |
| 2000 | A testability metric for path delay faults and its applicationabstractIn this paper, we propose a new testability metric for path delay faults.The metric is computed efficiently using a non-enumerative algorithm.It has been validated through extensive experiments and the results indicate a strong correlation between the proposed metric and the path delay fault testability of the circuit.We further apply this metric to derive a path delay fault test application scheme for scan-based BIST.The selection of the test scheme is guided by the proposed metric.The experimental results illustrate that the derived test application scheme can achieve a higher path delay fault coverage in scan-based BIST.Because of the effectiveness and efficient computation of this metric, it can be used to derive other design-for-testability techniques for path delay faults. Huan-Chih Tsai, Kwang-Ting Cheng, Vishwani D. Agrawal |
ASP-DAC | 2 |
| 2000 | Testing in the Fourth Dimension
Vishwani D. Agrawal, Kwang-Ting Cheng |
Asian Test Symposium | 2 |
| 2000 | Challenges for the Academic Test Community
Melvin A. Breuer, Kwang-Ting Cheng |
Asian Test Symposium | 2 |
| 2000 | Collaboration between Industry and Academia in Test Research
Kwang-Ting Cheng, Vishwani D. Agrawal, Jing-Yang Jou, Li-C. Wang, Chi-Feng Wu, Shianling Wu |
Asian Test Symposium | 1 |
| 2000 | An FPGA-based re-configurable functional tester for memory chipsabstractThe paper presents a prototype re-configurable tester for memory chips. The new tester consists of a memory test-circuitry compiler, a synthesis/mapping CAD tool, and an FPGA-based re-configurable hardware platform. The compiler makes user-specified parameters of memory under test (such as the address and data bus widths, the march test and the background data) as input and generates the test circuitry required to functionally, test the target memory chips. This framework not only enables the automatic synthesis/mapping of the test circuitry into the re-configurable hardware platform, it also guarantees that the hardware platform can correctly operate at the desired clock rate for the user specified parameters. The proposed solution can reduce the memory tester cost by providing hardware re-configurability to support a wide range of memory chips. We demonstrate that the prototype tester can be automatically configured to test SDRAM chips above 100 MHz. Jing-Reng Huang, Chee-Kian Ong, Kwang-Ting Cheng, Cheng-Wen Wu |
Asian Test Symposium | 3 |
| 2000 | Test challenges for deep sub-micron technologiesabstractThe use of deep submicron process technologies presents several new challenges in the area of manufacturing test. While a significant body of work has been devoted to identifying and investigating design challenges in nanometer technologies, the impact on test strategies and methodologies is still not well understood. This paper highlights the challenges to current test methodologies arising from technology driven trends, and will present an overview of emerging techniques that address deep submicron test challenges. Kwang-Ting Cheng, Sujit Dey, Mike Rodgers, Kaushik Roy 0001 |
DAC | 1 |
| 2000 | Assertion checking by combined word-level ATPG and modular arithmetic constraint-solving techniquesabstractWe present a new approach to checking assertion properties for RTI, design verification. Our approach combines structural, word-level automatic test pattern generation (ATPG) and modular arithmetic constraint-solving techniques to solve the constraints imposed by the target assertion property. Our word-level ATPG and implication technique not only solves the constraints on the control logic, but also propagates the logic implications to the datapath. A novel arithmetic constraint solver based on modular number system is then employed to solve the remaining constraints in datapath. The advantages of the new method are threefold. First, the decision-making process of the word-level ATPG is confined to the selected control signals only. Therefore, the enumeration of enormous number of choices at the datapath signals is completely avoided. Second, our new implication translation techniques allow word-level logic implication being performed across the boundary of datapath and control logic, and therefore, efficiently cut down the ATPG search space. Third, our arithmetic constraint solver is based on modular instead of integral number system. It can thus avoid the false negative effect resulting from the bit-vector value modulation. A prototype system has been built which consists of an industrial front-end HDL parser, a property-to-constraint converter and the ATPG/arithmetic constraint-solving engine. The experimental results on some public benchmark and industrial circuits demonstrate the efficiency of our approach and its applicability to large industrial designs. Chung-Yang Huang, Kwang-Ting Cheng |
DAC | 2 |
| 2000 | A BIST Scheme for On-Chip ADC and DAC TestingabstractIn this paper we present a BIST scheme for testing on-chip A/D and D/A converters. We discuss on-chip generation of linear ramps as test stimuli, and propose techniques for measuring the DNL and INL of the converters. We validate the scheme with software simulation-5% LSB (least significant bit) test accuracy can be achieved in the presence of reasonable analog imperfection. Jiun-Lang Huang, Chee-Kian Ong, Kwang-Ting Cheng |
DATE | 3 |
| 2000 | Path Selection and Pattern Generation for Dynamic Timing Analysis Considering Power Supply Noise EffectsabstractNoise effects such as power supply and crosstalk can significantly affect the performance of deep submicron designs. These delay effects are highly input pattern dependent. Existing path selection and timing analysis techniques cannot capture the effects of noise on cell/interconnect delays. Therefore, the selected critical paths may not be the longest paths and predicted circuit performance might not reflect the worst-case circuit delay. In this paper, we propose a path selection technique that can consider power supply noise effects on the propagation delays. Next, for the selected critical paths, we propose a pattern generation technique for dynamic timing analysis such that the patterns produce the worst-case power supply noise effects on the delays of these paths. Our experimental results demonstrate the difference in estimated circuit performance for the case when power supply noise effects are considered vs. when these effects are ignored. Thus, they validate the need for considering power supply noise effects on delays during path selection and dynamic timing analysis. Jing-Jia Liou, Angela Krstic, Yi-Min Jiang, Kwang-Ting Cheng |
ICCAD | 4 |
| 2000 | Testing and characterization of the one-bit first-order delta-sigma modulator for on-chip analog signal analysisabstractDelta-sigma modulation has become popular in modern analog-to-digital modulator design due to its relatively high immunity from process variations. We propose efficient characterization techniques to obtain the key performance parameters of the 1-bit first-order delta-sigma modulator which is intended to be used as an on-chip analog signal digitizer for BIST applications. Numerical simulations have been performed to validate the techniques and the results indicate that accurate estimation of the parameters can be obtained at the presence of noise. Jiun-Lang Huang, Kwang-Ting Cheng |
ITC | 2 |
| 2000 | Static property checking using ATPG vs. BDD techniquesabstractStatic property checking verifies pre-defined functional design rules such as "bus contention", "racing condition"; and "don't-care case". A static property checker typically uses formal verification techniques to prove the property under verification. If the property is proven false, a counter-example is generated for debugging the design. Among the different static property checking approaches, ATPG-based and BDD-based are the most powerful and successful ones. We implement both approaches with several optimization techniques on the same framework to compare their performance. The experimental results on industrial designs show that these two approaches have different strength and weakness in proving the static properties. Furthermore, the results indicate that they often complement each other and therefore a hybrid approach may result in better performance. We propose a static property checker based on combined ATPG and BDD techniques. The experimental results show that this combined approach can prove all the static properties in the test cases while still maintaining comparable performance. Chung-Yang Huang, Bwolen Yang, Huan-Chih Tsai, Kwang-Ting Cheng |
ITC | 4 |
| 2000 | Test program synthesis for path delay faults in microprocessor coresabstractThis paper addresses the problem of testing path delay faults in a microprocessor core using its instruction set. We propose to self-test a processor core by running an automatically synthesized test program which can achieve a high path delay fault coverage. This paper discusses the method and the prototype software framework for synthesizing such a test program. Based on the processor's instruction set architecture, micro-architecture, RTL netlist as well as gate-level netlist on which the path delay faults are modeled, the method generates deterministic tests (in the form of instruction sequences) by cleverly combining structural and instruction-level test generation techniques. The experimental results for two microprocessors indicate that the test instruction sequences can be successfully generated for a high percentage of testable path delay faults. Wei-Cheng Lai, Angela Krstic, Kwang-Ting Cheng |
ITC | 3 |
| 2000 | Efficient test mode selection and insertion for RTL-BISTabstractInserting test logic at the Register Transfer Level (RTL), instead of at the gate-level, offers many advantages. It allows the synthesis process to consider both functional and test logic together for optimization for meeting the timing/area/power goals; thus, this avoids an expensive cycle of re-optimization. It is also a necessary step for supporting RTL signoff. In this paper, we present a Built-in Self-Test (BIST) framework that allows efficient selection and insertion of test points at the RT level for achieving high fault coverage. We discuss the need for a new type of test points, called operator test points, in order to achieve high fault coverage for RTL BIST. The traditional node test points and the operator test points are jointly called test modes in this paper. We present a test-mode selection algorithm at the RT level, which uses a hybrid cost function derived from controllability, observability (C/O) and testability gradients of signals. Experimental results on some industrial designs indicate that high fault coverage can be achieved for various implementations of an RTL design by selecting and inserting the test modes at the RT-level. Subrata Roy, Gokhan Guner, Kwang-Ting Cheng |
ITC | 3 |
| 2000 | On Testing the Path Delay Faults of a Microprocessor Using its Instruction SetabstractThis paper addresses the problem of testing path delay faults in a microprocessor using instructions. It is observed that a structurally testable path (i.e., a path testable through at-speed scan) in a microprocessor might not be testable by its instructions simply because no instruction sequence can produce the desired test sequence which can sensitize the paths and capture the fault effect into the destination output/flip-flop at-speed. These paths are called functionally untestable paths. We discuss the impact of delay defects on the functionally untestable paths on the overall circuit performance and illustrate that they do not need to be tested if the delay defect does not cause the path delay to exceed twice the clock period. Identification of such paths helps determine the achievable path delay fault coverage and reduce the subsequent test generation effort. The experimental results for two microprocessors (Parwan and DLX) indicate that a significant percentage of structurally testable paths are functionally untestable and thus need not be tested. Wei-Cheng Lai, Angela Krstic, Kwang-Ting Cheng |
VTS | 3 |
| 2000 | Path Selection for Delay Testing of Deep Sub-Micron Devices Using Statistical Performance Sensitivity AnalysisabstractThe performance of deep sub-micron designs can be affected by various parametric variations, manufacturing defects, noise or even modeling errors that are all statistical in nature. In order to capture the effects of these statistical variations on circuit performance, we incorporate statistical information in timing analysis to compute the performance sensitivity of internal signals subject to a given type of defect, noise or variation sources. We further propose a novel path and segment selection methodology for delay testing based on the results of statistical performance sensitivity analysis. The objective of path/segment selection is to identify a small set of paths and segments such that the delay tests for the selected paths/segments guarantee the detection of performance failure caused by the target type of defect, noise or variation source. This new path selection methodology defines a new path/segment searching paradigm for detecting delay faults in deep sub-micron devices. Jing-Jia Liou, Kwang-Ting Cheng, Deb Aditya Mukherjee |
VTS | 2 |
| 2000 | Characterization of a Pseudo-Random Testing Technique for Analog and Mixed-Signal Built-in-Self-TestabstractIn this paper, we characterize and evaluate the effectiveness of a pseudo-random-based implicit functional testing technique for analog and mixed-signal circuits. The analog test problem is transformed into the digital domain by embedding the device-under-test (DUT) between a digital-to-analog-converter and an analog-to-digital converter. The pseudo-random testing technique uses band-limited digital white noise (pseudo-random-patterns) as input stimulus. The signature is constructed by computing the cross-correlation between the digitized output response and the pseudo-random input sequence. We have implemented a DSP-based hardware testbed to evaluate the effectiveness of the pseudo-random testing technique. Our results show that we can achieve close to 100% yield and fault coverages by carefully selecting only two cross-correlation samples. Noise level and total harmonic distortion below 0.1% and 0.5%, respectively, do not affect the classification accuracy. Jan Arild Tofte, Chee-Kian Ong, Jiun-Lang Huang, Kwang-Ting Cheng |
VTS | 4 |
| 2000 | AQUILA: An Equivalence Checking System for Large Sequential DesignsabstractIn this paper, we present a practical method for verifying the functional equivalence of two synchronous sequential designs. This tool is based on our earlier framework that uses Automatic Test Pattern Generation (ATPG) techniques for verification. By exploring the structural similarity between the two designs under verification, the complexity can be reduced substantially. We enhance our framework by three innovative features. First, we develop a local BDD-based technique which constructs Binary Decision Diagram (BDD) in terms of some internal signals, for identifying equivalent signal pairs. Second, we incorporate a technique called partial justification to explore not only combinational similarity, but also sequential similarity. This is particularly important when the two designs have a different number of flip-flops. Third, we extend our gate-to-gate equivalence checker for RTL-to-gate verification. Two major issues are considered in this extension: (1) how to model and utilize the external don't care information for verification; and (2) how to extract a subset of unreachable states to speed up the verification process. Compared with existing approaches based on symbolic Finite State Machine (FSM) traversal techniques, our approach is less vulnerable to the memory explosion problem and, therefore, is more suitable for a lot of real-life designs. Experimental results of verifying designs with hundreds of flip-flops will be presented to demonstrate the effectiveness of this approach. Shi-Yu Huang, Kwang-Ting Cheng, Kuang-Chien Chen, Chung-Yang Huang, Forrest Brewer |
IEEE Trans. Computers | 2 |
| 2000 | On improving test quality of scan-based BISTabstractIn this paper, we explore two techniques, under the existing scan-based built-in self-test (BIST) architectures, for improving the test quality with practically no additional overhead. The proposed techniques are an almost-full-scan BIST strategy and a general scan-based BIST test application scheme. We first demonstrate that under the scan-based BIST architecture, full scan may not result in the highest fault coverage (FC) and unscanning a small number of scan flip-flops may increase the BIST FC. We then present an algorithm for identifying those not-to-be-scanned flip-flops. We further show that the proposed general scan-based BIST test application scheme could also result in higher BIST FC and only requires a minor modification to the BIST controller. Experiments have been conducted using an industrial tool, psb2, on benchmark circuits to illustrate the effectiveness of the proposed techniques and algorithms. The results have demonstrated that both techniques are able to maximize the FC and reduce the test application time without additional test hardware compared to the conventional scan-based BIST architectures. Huan-Chih Tsai, Kwang-Ting Cheng, Sudipta Bhawmik |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2000 | Estimation for maximum instantaneous current through supply lines for CMOS circuitsabstractWe present new techniques for estimating the maximum instantaneous current through the power supply lines for CMOS circuits. We investigate four different approaches: (1) timed-ATPG-based approach; (2) probability-based approach; (3) genetic algorithm-based approach; and (4) integer linear programming (ILP) approach. The first three approaches produce a tight lower bound on the maximum current. The ILP-based approach produces the exact solutions for small circuits, and tight upper bounds of the solutions for large circuits. Our experimental results show that the upper bounds produced by the ILP approach combined with the lower bounds produced by the other three approaches confine the exact solution for the maximum instantaneous current to a small range. Yi-Min Jiang, Angela Krstic, Kwang-Ting Cheng |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 1999 | Analysis of Performance Impact Caused by Power Supply Noise in Deep Submicron DevicesabstractThe paper addresses the problem of analyzing the performance degradation caused by noise in power supply lines for deep submicron CMOS devices.We first propose a statistical modeling technique for the power supply noise including inductive ∆I noise and power net IR voltage drop.The model is then integrated with a statistical timing analysis framework to estimate the performance degradation caused by the power supply noise.Experimental results of our analysis framework, validated by HSPICE, for benchmark circuits implemented on both 0.25 µ, 2.5 V and 0.55 µ, 3.3 V technologies are presented and discussed.The results show that on average, with the consideration of this noise effect, the circuit critical path delays increase by 33% and 18%, respectively for circuits implemented on these two technologies. Yi-Min Jiang, Kwang-Ting Cheng |
DAC | 2 |
| 1999 | Improving the Test Quality for Scan-Based BIST Using a General Test Application SchemeabstractIn this paper, we propose a general test application scheme for existing scan-based BIST architectures.The objective is to further improve the test quality without inserting additional logic to the Circuit Under Test (CUT).The proposed test scheme divides the entire test process into multiple test sessions.A different number of capture cycles is applied after scanning in a test pattern in each test session to maximize the fault detection for a distinct subset of faults.We present a procedure to find the optimal number of capture cycles following each scan sequence for every fault.Based on this information, the number of test sessions and the number of capture cycles after each scan sequence are determined to maximize the random testability of the CUT.We conduct experiments on ISCAS89 benchmark circuits to demonstrate the effectiveness of our approach. Huan-Chih Tsai, Kwang-Ting Cheng, Sudipta Bhawmik |
DAC | 2 |
| 1999 | VIP - an input pattern generator for indentifying critical voltage drop for deep sub-micron designsabstractWe present a novel input pattern generator for dynamic power network simulation.The obtained patterns successfully identia critical voltage drop areas for a set of industrial designs, which are dificult to be found using functional vectors.The search engine of the pattern generator for worst-case IR voltage drop is based on the multiobjective genetic algorithm.To achieve high coverage for critical voltage drop cells, we propose to model the search criteria into the maximum weighted matching of a bipartite graph, and guide the search direction according to the matching results.Experimental results show that, compared with the other approaches, our patterns give a higher coverage of critical voltage drop cells. Yi-Min Jiang, Tak K. Young, Kwang-Ting Cheng |
ISLPED | 3 |
| 1999 | Delay testing considering power supply noise effectsabstractWe propose a new delay test generation technique that can take into account the impact of the power supply noise on the signal propagation delays. This is different from existing delay fault models and test generation techniques that ignore the dependence of path delays on the applied test patterns and cannot capture the worst-case timing scenarios in deep submicron designs. In addition to sensitizing the fault and propagating the fault effects to the primary outputs, our new tests also produce the worst-case power supply noise on the nodes in the target path. Thus, the tests also cause the worst-case propagation delay for the nodes along the target path. Our experimental results on benchmark circuits show that the new delay tests produce significantly longer delays on the tested paths compared to the tests derived using existing delay testing methods. Yi-Min Jiang, Angela Krstic, Kwang-Ting Cheng |
ITC | 3 |
| 1999 | Specification Back-Propagation and Its Application to DC Fault Simulation for Analog/Mixed-Signal CircuitsabstractIn this paper we present the specification backpropagation technique which enables one to derive the constraint of an internal functional block with respect to a given DC specification for an analog/mixed-signal system. Based on this technique, we implement an efficient fault simulator which reduces the required efforts by (1) removing undetectable faults from the fault list, and (2) performing fault simulation only locally for the fault block. Simulation results on an industrial design show a speedup factor of 7.2 with 98% correct classification of detected and undetected faults as compared with full-chip DC fault simulation. Jiun-Lang Huang, Chen-Yang Pan, Kwang-Ting Cheng |
VTS | 3 |
| 1999 | Testing High Speed VLSI Devices Using Slower TestersabstractThe speed of new VLSI designs is rapidly increasing. Assuring the performance of the circuit requires that the circuit be tested at its intended operating speed. The high cost of high speed testers makes it impossible for the testers to follow the designs in terms of speed increase. This gap between the speed of the new circuits and the speed of the testers is not likely to disappear. In this paper, we focus on at-speed strategies for testing high speed designs on slower testers. Conventional at-speed testing strategies assume that the primary inputs/outputs can be applied/observed at the circuit rated speed. This requires a high speed tester. Our assumption is that a fast clock matching the speed of the designs is available. We describe two classes of at-speed strategies that can be used on a low speed tester. The first class consists of testing schemes for which the test generation procedure is independent of the speed of the tester. These methods apply multiple input patterns in one tester cycle and the test application time for them can be long. The strategies in the second class of at-speed testing schemes integrate the tester's speed limitations with the test generation process. Due to constraints placed at the test generation process, these schemes might result in a reduced fault coverage. To increase the fault coverage and reduce the test application time, the slow-fast-slow and at-speed strategies can be combined for testing high speed designs on slower testers. We present preliminary experimental results for at-speed schemes for slow testers for transition faults. Angela Krstic, Kwang-Ting Cheng, Srimat T. Chakradhar |
VTS | 2 |
| 1999 | A New Bare Die Test MethodologyabstractWhile multichip module technology has been developed for high performance IC applications, the technology is not widely adopted due to economical reasons. One of the reasons that makes the technology economically unattractive is the problems and the high cost associated with testing and diagnosing each individual un-packaged IC in the system and the MCM module itself. The low MCM system yield prevents the technology from being used other than in high cost and high performance applications. In this paper, we propose a new methodology using ideas of tester-on-a-chip and a pressure contact technology to test bare dies. This methodology can reduce the IC testing cost and overall cost of the MCM module. It can also be considered as an alternative to high speed wafer probe. We designed an experiment for SRAM dies to examine the feasibility of this new method. Zao Yang, Kwang-Ting Cheng, King L. Tai |
VTS | 2 |
| 1999 | Fault emulation: A new methodology for fault gradingabstractIn this paper, we introduce a method that uses the field programmable gate array (FPGA)-based emulation system for fault grading. The real-time simulation capability of a hardware emulator could significantly improve the performance of fault grading, which is one of the most time consuming tasks in the circuit design and test process. We employ a serial fault emulation algorithm enhanced by two speed-up techniques. First, a set of independent faults can be injected and emulated at the same time. Second, multiple dependent faults can be simultaneously injected within a single FPGA-configuration by adding extra circuitry. Because the reconfiguration time of mapping the numerous faulty circuits into the FPGA's is pure overhead and could be the bottleneck of the entire process, using extra circuitry for injecting a large number of faults can reduce the number of FPGA-reconfigurations and, thus, improving the performance significantly. In addition, we address the issue of handling potentially detected faults in this hardware emulation environment by using the dual-railed logic. The performance estimation shows that this approach could be several orders of magnitude faster than the existing software approaches for large sequential designs. Kwang-Ting Cheng, Shi-Yu Huang, Wei-Jin Dai |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1999 | ErrorTracer: design error diagnosis based on fault simulation techniquesabstractThis paper addresses the problem of locating error sources in an erroneous combinational or sequential circuit. We use a fault simulation-based technique to approximate each internal signal's correcting power. The correcting power of a particular signal is measured in terms of the signal's correctable set, namely, the maximum set of erroneous input vectors or sequences that can be corrected by resynthesizing the signal. Only the signals that can correct every given erroneous input vector or sequence are considered as a potential error source. Our algorithm offers three major advantages over existing methods. First, unlike symbolic approaches, it is applicable for large circuits. Second, it delivers more accurate results than other simulation-based approaches because it is based on a more stringent condition for identifying potential error sources. Third, it can be generalized to identify multiple errors theoretically. Experimental results on diagnosing combinational and sequential circuits with one and two random errors are presented to show the effectiveness and efficiency of this new approach. Shi-Yu Huang, Kwang-Ting Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1999 | AutoFix: a hybrid tool for automatic logic rectificationabstractWe address the problem of rectifying an erroneous combinational circuit. Based on the symbolic binary decision diagram techniques, we consider the rectification process as a sequence of partial corrections. Each partial correction reduces the size of the input vector set that produces error responses. Compared with the existing approaches, this approach is more general, and thus, suitable for circuits with multiple errors and for the engineering change problem. Also, we derive the necessary and sufficient condition of general single-gate correction to improve the quality of rectification. To handle larger circuits, we develop a hybrid approach that makes use of the information of structural correspondence between specification and implementation. Experiments are performed on a suite of industrial examples as well as the entire set of ISCAS'85 benchmark circuits to demonstrate its effectiveness. Shi-Yu Huang, Kuang-Chien Chen, Kwang-Ting Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 1999 | Primitive delay faults: identification, testing, and design for testabilityabstractWe investigate two strategies to guarantee temporal correctness of a combinational circuit. We first propose a new technique to identify and test primitive faults. A primitive fault is a path delay fault that has to be tested to guarantee the performance of the circuit. Primitive faults can consist of single- (SPDF's) or multiple path delay faults (MPDF's). Testing strategies for single primitive faults exist. In this paper, we focus on identifying and testing multiple primitive faults. Identification and testing of these faults is important for at least two reasons: (1) a large percentage of paths in production circuits remain untestable under the SPDF model, and (2) distributed manufacturing defects usually adversely affect more than one path and these defects can be detected only by analyzing multiple affected paths. The SPDF's contained in a multiple primitive fault have to merge at some gate(s). Our methodology can quickly (1) rule out a large number of gates as possible merging gates for primitive faults, and (2) prune the combinations of paths that can never belong to any primitive fault. Our identification procedure also finds a test for the fault. We present a complete algorithm for identifying and testing double path delay faults, Identifying and testing all primitive faults is impractical for large designs. This is because no efficient methods are known for testing primitive faults that include a large number of paths. However, to guarantee that the performance of a digital circuit is not affected by timing defects, it is necessary to test all primitive faults. Our second contribution is a new design for testability method. Our method guarantees that only primitive faults with at most two paths can exist in the circuit in the test mode. The main idea is to efficiently identify a small set of signals for inserting test points to eliminate primitive faults with more than two paths. Our test points only provide controllability. Addition of a single test point can lower the cardinality of several primitive faults. Our approach efficiently re-evaluates primitive delay fault testability of the circuit after insertion of a test point. After a few iterations only primitive faults with at most two paths can exist in the circuit in the test mode. Experimental results on several multilevel combinational benchmark circuits are included to demonstrate the usefulness of our techniques. Angela Krstic, Kwang-Ting Cheng, Srimat T. Chakradhar |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1998 | A Hybrid Power Model for RTL Power EstimationabstractWe propose a hybrid power model for estimating the power dissipation of a design at the RT-level. This new model combines the advantages of both RT-level and gate-level approaches. We investigate the relationship between steady-state transition power and overall power dissipation. We observe that, statistically, two input sequences causing similar amount of steady-state transitions will exhibit similar overall power dissipation for an RTL module. Based on this observation, we propose a method to construct a hybrid power model for RTL modules. We further propose a hierarchical power estimation method for estimating the power dissipation of data-path consisting of RTL modules. Experimental results show that, for full-chip power estimation, the estimation time of the technique based on our power models is on average 275 times faster than directly running a commercial transistor-level power simulator, and the errors are less than 6% as compared to the transistor-level power simulation results. Yi-Min Jiang, Shi-Yu Huang, Kwang-Ting Cheng, Deborah C. Wang, ChingYen Ho |
ASP-DAC | 3 |
| 1998 | Fault-Simulation Based Design Error Diagnosis for Sequential CircuitsabstractThis paper addresses the problem of locating design errors in a sequential circuit. For single-error circuits, we consider a signal ƒ as a potential error source only if the circuit can be completely rectified by re-synthesizing ƒ (i.e., changing the function of signal ƒ). In order to handle larger circuits, we do not rely on Binary Decision Diagram. Instead, we search for potential error sources by a modified sequential fault simulation process. The main contributions of this paper are two-fold: (1) we derive the necessary and sufficient condition of whether an erroneous input sequence (i.e., an input sequence producing erroneous responses) can be corrected by changing the function of a particular internal signal; and (2) we propose a modified fault simulation procedure to check this condition. Our approach does not rely on any error model, and thus, is suitable for general types of errors. Furthermore, it can be easily extended to identify multiple errors. Experimental results on ISCAS89 benchmark circuits are presented to demonstrate its capability. Shi-Yu Huang, Kwang-Ting Cheng, Kuang-Chien Chen, Juin-Yeu Joseph Lu |
DAC | 2 |
| 1998 | Functional Scan Chain TestingabstractFunctional scan chains are scan chains that have scan paths through a circuit's functional logic and flip-flops. Establishing functional scan paths by test point insertion (TPI) has been shown to be an effective technique to reduce the scan overhead. However, once the scan chain is allowed to go through functional logic, the traditional alternating test sequence is no longer enough to ensure the correctness of the scan chain. We identify the faults that affect the functional scan chain, and show a methodology to find tests for these faults. Our results have the number of undetected faults at only 0.006% of the total number of faults, or 0.022% of the faults affecting the scan chain. Douglas Chang, Kwang-Ting Cheng, Malgorzata Marek-Sadowska, Mike Tien-Chien Lee |
DATE | 2 |
| 1998 | Exact and Approximate Estimation for Maximum Instantaneous Current of CMOS CircuitsabstractWe present an integer-linear-programming-based approach for estimating the maximum instantaneous current through the power supply lines for CMOS circuits. It produces the exact solutions for the maximum instantaneous current for small circuits, and tight upper bounds for large circuits. We formulate the maximum instantaneous current estimation problem as an integer linear programming (ILP) problem, and solve the corresponding ILP formulae to obtain the exact solution. For large circuits we propose to partition the circuits, and apply our ILP-based approach for each sub-circuit. The sum of the exact solutions of all sub-circuits provides an upper bound of the exact solution for the entire circuit. Our experimental results show that the upper bounds produced by our approach combined with the lower bounds produced by a genetic-algorithm-based approach confine the exact solution to a small range. Yi-Min Jiang, Kwang-Ting Cheng |
DATE | 2 |
| 1998 | Estimation of maximum power supply noise for deep sub-micron designsabstractWe propose a new technique for generating a small set of patterns to estimate the maximum power supply noise of deep sub-micron designs. We first build the charge/discharge current and output voltage waveform libraries for each cell, taking power and ground pin characteristics, the power net RC and other input characteristics as parameters. Based on the cells' current and voltage libraries, the power supply noise of a 2-vector sequence can be estimated efficiently by a cell-level waveform simulator. We then apply the Genetic Algorithm based on the efficient waveform simulator to generate a small set of patterns producing high power supply noise. Finally, the results are validated by simulating the obtained patterns using a transistor level simulator. Our experimental results show that the patterns generated by our approach produce a tight lower bound on the maximum power supply noise. Yi-Min Jiang, Kwang-Ting Cheng, An-Chang Deng |
ISLPED | 2 |
| 1998 | LIBRA - a library-independent framework for post-layout performance optimizationabstractIn this paper we present a post-layout timing optimization framework which (1) is library-independent such that it can take the logic-optimized Verilog file as its input netlist, (2) provides a prototype interface which can communicate with any vendor's physical design tools to obtain the accurate timing, topological and physical information, and perform ECO placement and routing, and (3) has fast and powerful rewiring routines that offer an extra solution space beyond the existing physical-level optimization methodologies. We conduct the post-layout performance optimization experiments on some benchmark circuits which are originally optimized by Synopsys's Design Compiler, (with high timing effort), followed by Avant!'s timing-driven place-and-route tool, Apollo. The optimization strategies we used include rewiring, buffer insertion, and cell sizing. To study the trade-offs between these transformations and the benefits of mixing them together, they are applied both separately and closely integrated by some heuristic cost functions. The result shows that by using all these strategies, post-layout timing optimization can further achieve up to 23.9% of improvement after global routing. We also discuss the pros and cons for our proposed procedures applied after global routing versus after detail routing. Some factors that can affect the quality of rewiring such as level of recursive learning and type of rewiring will also be addressed. Chung-Yang Huang, Kwang-Ting Cheng |
ISPD | 3 |
| 1998 | National Science Foundation Workshop on Future Research Directions in Testing of Electronic Circuits and Systems: executive summary of workshop reportabstractA two-day meeting, sponsored by National Science Foundation, was held in Santa Barbara, California on May 12 and 13, 1998 to discuss the academic research topics and education issues in testing of electronic circuits and systems. The goals of the meeting were (1) to identify emerging and mature research areas within the VLSI testing field, in order to help focus the field on research necessary to develop algorithms and techniques for testing VLSI circuits and systems designed using future technologies, (2) to address issues related to increasing the impact of the VLSI test field on education in electrical and computer engineering, and (3) to identify models and mechanisms for enhancing the interaction, collaboration and data-sharing between industry and academia. Four working groups were formed in the meeting: two groups focusing on the identification of emerging and mature research topics, one on industry and university interaction/collaboration, and one on test education. A brief summary of a report containing findings and recommendations from these four working groups is presented. Kwang-Ting Cheng |
ITC | 1 |
| 1998 | An almost full-scan BIST solution-higher fault coverage and shorter test application timeabstractThis paper illustrates that for existing scan-based Built-In Self-Test (BIST) architectures under the pseudo-random testing scheme, scanning all flip-flops may not be the best strategy for achieving high fault coverage with a practical limit on test length. In general, for scan-based BIST, not scanning flip-flops with relatively high pseudo-random observabilities through the primary outputs may indeed improve the fault coverage as well as the test application time. We illustrate the issues and present a flip-flop selection strategy for scan-based BIST to maximize the fault coverage and reduce the test application time. Experiments have been conducted based on an industrial tool psb2 for several benchmark circuits. The results show that the almost-full-scan circuits based on our flip-flop selection strategy can achieve higher fault coverages and significantly shorter test application time as compared with the full-scan circuits. Huan-Chih Tsai, Sudipta Bhawmik, Kwang-Ting Cheng |
ITC | 3 |
| 1998 | A hybrid methodology for switching activities estimationabstractIn this paper, we propose a hybrid approach for estimating the switching activities of the internal nodes in logic circuits. The new approach combines the advantages of the simulation-based techniques and the probability-based techniques. We use the user-specified control sequence for simulation, and treat the weakly correlated data inputs using the probabilistic model. The new approach, on one hand, is more accurate than the probabilistic approaches because the strong temporal and spatial correlations among control inputs are well taken into consideration. On the other hand, the new approach is much more efficient than the simulation-based approaches because the weakly correlated data inputs are not explicitly simulated. We also discuss the situation where BDD's are built in terms of internal nodes so that large circuits can he handled. Extensive experimental results are presented to show the effectiveness and efficiency of our algorithms. David Ihsin Cheng, Kwang-Ting Cheng, Deborah C. Wang, Malgorzata Marek-Sadowska |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1998 | Test-point insertion: scan paths through functional logicabstractConventional scan design imposes considerable area and delay overheads. To establish a scan chain in the test mode, multiplexers at the inputs of flip-flops and scan wires are added to the actual design. We propose a low-overhead scan design methodology that employs a new test-point insertion technique. Unlike the conventional test-point insertion, where test points are used directly to increase the controllability and observability of the selected signals, the test points are used here to establish scan paths through the functional logic. The proposed technique reuses the functional logic for scan operations; as a result, the design-for-testability overhead on area or timing can be minimized. We show an algorithm that uses the new test-point insertion technique to reduce the area overhead for the full-scan design. We also discuss its application to the timing-driven partial-scan design. Chih-Chang Lin, Malgorzata Marek-Sadowska, Kwang-Ting Cheng, Mike Tien-Chien Lee |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 1998 | Efficient test-point selection for scan-based BISTabstractWe propose a test point selection algorithm for scan-based built-in self-test (BIST). Under a pseudorandom BIST scheme, the objectives are (1) achieving a high random pattern fault coverage, (2) reducing the computational complexity, and (3) minimizing the performance as well as the area overheads due to the insertion of test points. The proposed algorithm uses a hybrid approach to accurately estimate the profit of the global random testability of a test point candidate. The timing information is fully integrated into the algorithm to access the performance impact of a test point. In addition, a symbolic procedure is proposed to compute testability measures more efficiently for circuits with feedbacks so that the test point selection algorithm can be applied to partial-scan circuits. The experimental results show the proposed algorithm achieves higher fault coverages than previous approaches,with a significant reduction of computational complexity. By taking timing information into consideration, the performance degradation can he minimized with possibly more test points. Huan-Chih Tsai, Kwang-Ting Cheng, Chih-Jen Lin, Sudipta Bhawmik |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1997 | AQUILA: An equivalence verifier for large sequential circuitsabstractIn this paper, we address the problem of verifying the equivalence of two sequential circuits. A hybrid approach that combines the advantages of BDD-based and ATPG-based approaches is introduced. Furthermore, we incorporate a technique called partial justification to explore the sequential similarity between the two circuits under verification to speed up the verification process. Compared with existing approaches, our method is much less vulnerable to the memory explosion problem, and therefore can handle larger designs. The experimental results show that in a few minutes of CPU time, our tool can verify the sequential equivalence of an intensively optimized benchmark circuit with hundreds of flip-flops against its original version. Shi-Yu Huang, Kwang-Ting Cheng, Kuang-Chien Chen |
ASP-DAC | 2 |
| 1997 | A Test Synthesis Approach to Reducing BALLAST DFT OverheadabstractIn this paper, we present a test synthesis approach which integratesBALLAST (BALAnced structure Scan Test) withan enhanced test point insertion (TPI) algorithm to functionallyscan the flip-flops chosen by BALLAST.BALLASTis an attractive partial scan technique in that it offers combinationalATPG efficiency while promising to reduce full scanoverhead.However, the practical problem with BALLASTis it typically requires more scan flip-flops than other partialscan techniques.The TPI enhancements enable TPI toaim at the reduction of BALLAST overhead.The enhancementsinclude a more flexible test point insertion heuristic,a modified gain function which enables TPI to target a selectedset of flip-flops, and a more efficient procedure toremove redundant test points.The experimental results onnine benchmark circuits show the proposed test synthesisapproach can achieve on average 38% area saving comparedto full scan, while BALLAST alone achieves 17%. Douglas Chang, Mike Tien-Chien Lee, Malgorzata Marek-Sadowska, Takashi Aikyo, Kwang-Ting Cheng |
DAC | 5 |
| 1997 | Post-Layout Logic Restructuring for Performance OptimizationabstractWe propose a new methodology based on incremental logic restructuring for post-layout performance improvement. The new post-layout logic restructuring technique allows to use accurate interconnection delays for performance optimization, while the incremental nature of the technique guarantees convergence between logic synthesis and layout. The technique can be further integrated with other post-layout optimization techniques such as gate sizing and buffer insertion. Experimental results show that this technique combined with post-layout buffer insertion can achieve an additional 15% improvement in performance compared to designs produced by timing-driven logic optimization followed by pre-layout buffer insertion followed by timing-driven physical design. 1. Introduction Performance-driven logic synthesis followed by performance -driven layout [1] has become a necessity for designing high performance circuits. However, this loosely coupled two-phase timing optimization methodology has... Yi-Min Jiang, Angela Krstic, Kwang-Ting Cheng, Malgorzata Marek-Sadowska |
DAC | 3 |
| 1997 | Vector Generation for Maximum Instantaneous Current Through Supply Lines for CMOS CircuitsabstractWe present two new algorithms for generating a smallset of patterns for estimating the maximum instantaneouscurrent through the power supply lines for CMOScircuits.The first algorithm is based on timed ATPG,while the second is a probability-based approach.Bothalgorithms can handle circuits with arbitrary but knowndelays and they produce a set of 2-vector tests.Experimentalresults demonstrating that the outcome of applyingour algorithms is a small set of patterns producinga current that is a tight lower bound on the maximuminstantaneous current are included. Angela Krstic, Kwang-Ting Cheng |
DAC | 2 |
| 1997 | A Hybrid Algorithm for Test Point Selection for Scan-Based BISTabstractWe propose a new algorithm for test point selection for scan-based BIST. The new algorithm combines the advantages of both explicit-testability-calculation and gradient techniques. The test point selection is guided bya cost function which is partially based on explicit testability recalculation and partially on gradients. With an event-driven mechanism, it can quickly identify a set of nodes whose testability need to be recalculated due to a test point, and then use gradients to estimate the impact of the rest of the circuit. In addition, by incorporating timing information into the cost function, timing penalty caused by test points can be easily avoided. We present the results to illustrate that high fault coverages for both area- and timing-driven test point insertions can be obtained with a small number of test points. The results also indicate a signi#cant reduction of computational complexity while the qualities are similar to the explicitly-testability-calculation method. 1 I... Huan-Chih Tsai, Kwang-Ting Cheng, Chih-Jen Lin, Sudipta Bhawmik |
DAC | 2 |
| 1997 | Analog Fault Diagnosis for Unpowered Circuit BoardsabstractWe present key portions of a method for automatic analog fault diagnosis for unpowered circuit boards. Our work consists of two major parts: (1) test point selection and (2) stimuli selection for diagnostic test generation. For test point selection, we propose an efficient graph-based algorithm achieving a desired level of diagnosibility. The stimuli selection algorithm uses a cost function derived from the sensitivity matrix to select test stimuli and thus avoids expensive circuit simulation. Experimental results of several industrial circuits show that our method is time efficient and promising in selecting high-quality stimuli. Jiun-Lang Huang, Kwang-Ting Cheng |
ITC | 2 |
| 1997 | Error Tracer: A Fault-Simualtion-Based Approach to Design Error DiagnosisabstractThis paper addresses the problem of locating error sources in an erroneous combinational circuit. We use a fault simulation-based technique to approximate each signal's correcting power. The correcting power of a particular signal is measured in terms of the signal's correctable set, namely, the maximum set of erroneous input vectors that can be corrected by re-synthesizing the signal. Only the signals that can correct every erroneous input vector are considered as a potential error source. Our algorithm offers three major advantages over existing methods. First, unlike symbolic approaches, it is applicable for large circuits. Secondly, it delivers more accurate results than other simulation-based approaches because it is based on a more stringent condition for identifying potential error sources. Thirdly, it can be easily generalized to identify multiple errors. Experimental results on diagnosing circuits with one and two random errors are presented to show the effectiveness and efficiency of this new approach. Shi-Yu Huang, Kwang-Ting Cheng, Kuang-Chien Chen, David Ihsin Cheng |
ITC | 2 |
| 1997 | Design for Primitive Delay Fault TestabilityabstractTo guarantee the temporal correctness of a digital circuit a set of multiple path delay faults called primitive faults need to be tested. Primitive faults can contain one or more faulty paths. Existing techniques can identify and test primitive faults containing up to two or three paths. Identifying and testing primitive faults that consist of a larger number of paths is impractical for large designs. We propose a design for testability method that assures the temporal correctness of the circuit without the need to test all primitive faults in the circuit. In the test mode, only primitive faults that contain up to two paths can affect the circuit performance. Our methodology efficiently identifies a small set of potential locations for inserting control points to eliminate primitive faults with more than two paths. Addition of a single control point can lower the cardinality of several primitive faults. Our approach re-evaluates primitive delay fault testability of the circuit after insertion of every control point. After a few iterations only primitive faults with at most two paths can exist in the circuit in the test mode. Experimental results on several circuits are included to demonstrate our method. Angela Krstic, Kwang-Ting Cheng, Srimat T. Chakradhar |
ITC | 2 |
| 1997 | Fault Macromodeling for Analog/Mixed-Signal CircuitsabstractIn this paper we propose an efficient fault macromodeling technique for analog/mixed-signal circuits. We formulate the fault macromodeling problem as a problem of deriving the macro parameter set B based on the performance parameter set P of the transistor-level faulty circuit. The fault macromodel is intended to be used for efficient macro-level fault simulation. In such applications, a common approach to speeding up the macromodeling process is to generate a large number of data pairs (P, B) (the training set) and interpolate an empirical mapping function B=F(P) based on the training set. In our technique, generation of each data pair requires only one run of macro-level simulation, as opposed to multiple runs of macro-level simulation required by iterative fault macromodeling techniques. We also propose a cross-correlation-based technique to select a subset of parameters from the high dimensional parameter set P to speed up function interpolation. We demonstrate the effectiveness and efficiency of our proposed fault macromodeling technique by showing some preliminary, experimental results on an industrial design. Chen-Yang Pan, Kwang-Ting Cheng |
ITC | 2 |
| 1997 | Incremental logic rectificationabstractWe address the problem of rectifying an incorrect combinational circuit against a given specification. Based on the symbolic BDD techniques, we consider the rectification process as a sequence of partial corrections. Each partial correction reduces the size of the input vector set producing error responses. Compared with existing approaches, this approach is more general, and able to handle circuits with multiple errors. We also formulate the necessary and sufficient condition of general single-gate correction to achieve better results for some circuits with a single error. To handle larger circuits, we develop a hybrid approach that makes use of the information of structural correspondence between specification and implementation. Experimental results on industrial examples as well as ISCAS85 benchmark circuits are presented to show the effectiveness of our approach. Shi-Yu Huang, Kuang-Chien Chen, Kwang-Ting Cheng |
VTS | 3 |
| 1997 | Guest Editorial
Kwang-Ting Cheng, Kewal K. Saluja, Hans-Joachim Wunderlich |
J. Electron. Test. | 1 |
| 1997 | Resynthesis of Combinational Circuits for Path Count Reduction and for Path Delay Fault Testability
Angela Krstic, Kwang-Ting Cheng |
J. Electron. Test. | 2 |