Jaeyong Chung

dblp:73/3431 · DBLP profile ↗
← Back
43ranked-venue papers
19as first author
8since 2021 · last 2025
0000-0001-5819-1995ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 40 · 17 first-author · 8 since 2021Software engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-author
YearPublicationVenuePosition
2025 DBC: Drift-aware Binary Code for Drift-tolerant Deep Neural Networks
abstract
Deep neural networks (DNNs) have demonstrated outstanding performance across a wide range of applications. However, their substantial number of weights necessitates scalable and efficient storage solutions. Emerging non-volatile memory technologies, such as phase change memory (PCM) with multi-level cell (MLC) operation, are promising candidates due to their high scalability and non-volatility compared to conventional charge-based storage devices. Despite these advantages, MLC PCM suffers from reliability issues, particularly conductance drift, where the conductance of a PCM cell changes over time. This drift can lead to significant accuracy degradation in DNNs, as their weights are stored in PCM cells. In this paper, we propose Drift-aware Binary Code (DBC), a novel binary code designed to improve the tolerance of DNNs to conductance drift. DBC maps smaller decimal values to less error-prone MLC PCM cell levels and ensures that values shift to smaller magnitudes when conductance drift occurs. This approach helps maintain the accuracy of the DNN over an extended period compared to conventional binary code, as dominant DNN weights are stored at levels less prone to errors and DNNs exhibit better tolerance when weight values decrease rather than increase due to drift. Additionally, DBC requires no additional hardware overhead for auxiliary bits and can be combined with other fault-tolerant approaches, such as error correction code (ECC). Experimental results based on the real PCM device developed by IBM Research demonstrate that DBC improves the drift tolerance of DNNs by up to $55.18 \times$ compared to conventional binary code.
Insu Choi, Jaeyong Chung, Joon-Sung Yang
DAC2
2025 Reducing Errors and Powers in LPDDR for DNN Inference: A Compression and IECC-Based Approach
Jae-Youn Hong, Je-Woo Jang, Sung-Hyuk Cho, Youngbae Kong, Sungkyu Kim, Youngjung Kang, Jaehyung Ko, Jaeyong Chung, Joon-Sung Yang
J. Syst. Archit.8
2025 AGD: Analytic Gradient Descent for Discrete Optimization in EDA and its Use to Gate Sizing
abstract
In electronic design automation (EDA), simulation models are often non-differentiable, and many design choices are discrete. As a result, greedy optimization methods based on numerical gradients are widely used, although they often lead to suboptimal solutions. In contrast, analytical methods may provide better solutions but require significant research effort. Reinforcement learning (RL) has been employed to address this problem; however, RL also suffers from notorious sample inefficiency, which is exaggerated in EDA because data sampling in EDA is very expensive due to slow simulations. This article proposes an alternative to RL for EDA, namely analytic gradient descent (AGD). Our method starts with a differentiable performance model, which can be either a learned surrogate or a static model. It then applies transformations similar to Shannon decomposition for each design variable in the performance model. Finally, one design option for each variable is selected using a one-hot variable, which is trained via a straight-through estimator (STE) through gradient descent. We demonstrate AGD on the well-known gate sizing problem using both a learned surrogate and a static model across 20 industrial benchmark circuits. Our experimental results show that the proposed method can outperform a several-decade-old commercial tool in the gate sizing task for 19 out of the 20 circuits.
Phuoc Pham, Tae-Min Park 0001, Sung-Hyuk Cho, Tayyeb Mahmood, Joon-Sung Yang, Jaeyong Chung
ACM Trans. Design Autom. Electr. Syst.6
2024 FPGA-assisted Design Space Exploration of Parameterized AI Accelerators: A Quickloop Approach
Kashif Inayat, Fahad Bin Muslim, Tayyeb Mahmood, Jaeyong Chung
J. Syst. Archit.4
2024 Factored Systolic Arrays Based on Radix-8 Multiplication for Machine Learning Acceleration
abstract
Systolic arrays (SAs) are re-gaining the attention as the heart to accelerate machine learning workloads. This article shows that a large design space exists at the logic level despite the simple structure of SAs and proposes two novel SAs based on factoring and radix-$8$multipliers: The first factored SA (FSA) extracts out the booth encoding and the hard-multiple generation which is common across all processing elements (PEs), reducing the delay and the area of the whole SA. This factoring is done at the cost of an increased number of registers; however, the reduced pipeline register requirement in radix-$8$offsets this effect. Our second proposed FSA compresses the interconnections further with two steps of hard-multiple addition. In the first part, carries are computed column-wise outside PEs, and in the second part, early-generated carries are used for hard-multiple final addition inside PEs. We called it hard-multiple carry portioned (HCP) FSA (HCP FSA). The first proposed factored$16$-bit multiplier achieves up to$15$%,$13$%, and$23$% better delay, area, and power, respectively, compared with the radix-$4$multipliers even if the register overhead is included. And first proposed FSA architecture improves delay, area, and power up to$11$%,$20$%, and$31$%, respectively, for different bitwidths when compared with the conventional radix-$4$SA. In addition, the second HCP FSA design eliminates the additional registered overhead associated with the first proposed FSA by reducing interconnections and shows further reductions in the area up to 11.7% and power up to 16.7% with little increase in delay for various sizes of SAs.
Kashif Inayat, Inayat Ullah, Jaeyong Chung
IEEE Trans. Very Large Scale Integr. Syst.3
2023 Quickloop: An Efficient, FPGA-Accelerated Exploration of Parameterized DNN Accelerators
abstract
Quickloop is a design-space exploration (DSE) framework of parameterized RTL generators, their software stack, and their simulation on FPGA. FPGAs are recently accelerating RTL simulations due to their rapid turnaround times (TAT), compared to ASIC. However, this TAT is still restrictive in DSE. We adopt a data-driven approach to optimize Quickloop's TAT and leverage this framework to extensively search the design space of an open source DNN accelerator. We show that our approach effectively slashes the TAT by above 30%, compared to conventional toolflow.
Tayyeb Mahmood, Kashif Inayat, Jaeyong Chung
PACT3
2023 AGD: A Learning-based Optimization Framework for EDA and its Application to Gate Sizing
abstract
In electronic design automation (EDA), most simulation models are not differentiable, and many design choices are discrete. As a result, greedy optimization methods based on numerical gradients have been used widely, although it suffers from suboptimal solutions. On the other hand, analytic methods may offer better solutions, but at the cost of enormous research efforts. Reinforcement learning (RL) has been leveraged to tackle this problem owing to its generality; however, RL also suffers from notorious sample inefficiency, which is exaggerated in EDA because data sampling in EDA is very expensive due to slow simulations. This paper proposes an alternative to RL for EDA, namely analytic gradient descent (AGD). Our method calculates analytic gradients of a design objective with respect to continuous and discrete design choices through a neural network learned by a simulation model. Then it performs a gradient descent procedure optimizing the design objective directly. We demonstrate AGD on the well-known gate sizing problem and show that our method can be very close to an industry-leading commercial tool in terms of design quality of result (QoR), while it only takes several person-months in comparison to dedicated efforts of human engineering over decades to develop. In addition, we also show that AGD can generalize to unseen circuits, with less training specific in a small amount of execution time.
Phuoc Pham, Jaeyong Chung
DAC2
2022 Hybrid Accumulator Factored Systolic Array for Machine Learning Acceleration
abstract
Deep learning applications have become ubiquitous in today’s era and it has led to vast development in machine learning (ML) accelerators. Systolic arrays have been a primary part of ML accelerator architecture. To fully leverage the systolic arrays, it is required to explore the computer arithmetic data-path components and their tradeoffs in accelerators. We present a novel factored systolic array (FSA) architecture, in which the carry propagation adder (CPA) and carry-save adder (CSA) perform hybrid accumulation on least significant bit (LSB) bits and most significant bits (MSB) bits, respectively, inside each processing element. In addition, a small CPA to complete accumulation for MSB bits along with rounding logic for each column of the array is placed, which not only reduces the area, delay, and power but also balances the combinational and sequential area tradeoffs. We demonstrate the hybrid accumulator with partial CPA factoring in “Gemmini,” an open-source practical systolic array accelerator and factoring technique does not change the functionality of the base design. We implemented three baselines, original Gemmini and two variants of it, and show that the proposed approach leads to overall significant reduction in area within the range 12.8% – 50.2% and in power within the range 18.6% – 41% with improved or similar delay in comparison to the baselines.
Kashif Inayat, Jaeyong Chung
IEEE Trans. Very Large Scale Integr. Syst.2
2020 Factored Radix-8 Systolic Array for Tensor Processing
abstract
Systolic arrays are re-gaining the attention as the heart to accelerate machine learning workloads. This paper shows that a large design space exists at the logic level despite the simple structure of systolic arrays and proposes a novel systolic array based on factoring and radix-8 multipliers. The factored systolic array (FSA) extracts out the booth encoding and the hard-multiple generation which is common across all processing elements, reducing the delay and the area of the whole systolic array. This factoring is done at the cost of an increased number of registers, however, the reduced pipeline register requirement in radix-8 offsets this effect. The proposed factored 16-bit multiplier achieves up to 15%, 13%, and 23% better delay, area, and power, respectively, compared with the radix-4 multipliers even if the register overhead is included. The proposed FSA architecture improves delay, area, and power up to 11%, 20% and 31%, respectively, for different bitwidths when compared with the conventional radix-4 systolic array.
Inayat Ullah, Kashif Inayat, Joon-Sung Yang, Jaeyong Chung
DAC4
2020 Reliable and Lightweight PUF-based Key Generation using Various Index Voting Architecture
abstract
Physical Unclonable Functions (PUFs) can be utilized for secret key generation in security applications. Since the inherent randomness of PUF can degrade its reliability, most of the existing PUF architectures have designed post-processing logic to enhance the reliability such as an error correction function for guaranteeing reliability. However, the structures incur high cost in terms of implementation area and power consumption. This paper introduces a Various Index Voting Architecture (VIVA) that can enhance the reliability with a low overhead compared to the conventional schemes. The proposed architecture is based on an index-based scheme with simple computation logic units and iterative operations to generate multiple indices for the accuracy of key generation. Our evaluation results show that the proposed architecture reduces the hardware implementation overhead by 2 to more than 5 times, without losing a key generation failure probability compared to conventional approaches.
Jeong-Hyeon Kim, Ho-Jun Jo, Kyung-Kuk Jo, Sung-Hee Cho, Jaeyong Chung, Joon-Sung Yang
DATE5
2020 ER-TCAM: A Soft-Error-Resilient SRAM-Based Ternary Content-Addressable Memory for FPGAs
abstract
Static random access memory (SRAM)-based ternary content-addressable memory (TCAM) on field-programmable gate arrays (FPGAs) is used for packet classification in software-defined networking (SDN) and OpenFlow applications. SRAMs implementing TCAM contents constitute the major part of a TCAM design on FPGAs, which are vulnerable to soft errors. The protection of SRAM-based TCAMs against soft errors is challenging without compromising critical path delay and maintaining a high search performance. This brief presents a lowcost and low-response-time technique for the protection of SRAM-based TCAMs. This technique uses simple, single-bit parity for fault detection which has a minimal critical path overhead. This technique exploits the binary-encoded TCAM table maintained in SRAM-based TCAMs for update purposes to implement a low-response-time error-correction mechanism at low cost. The error-correction process is carried out in the background, allowing lookup operations to be performed simultaneously, thus maintaining a high search performance. The proposed technique provides protection against soft errors with a response time of 293 ns, whereas maintaining a search rate of 222 million searches per second on a 1024 × 40 size TCAM on Artix-7 FPGA.
Inayat Ullah, Joon-Sung Yang, Jaeyong Chung
IEEE Trans. Very Large Scale Integr. Syst.3
2019 DeepRT: predictable deep learning inference for cyber-physical systems
Woochul Kang, Jaeyong Chung
Real Time Syst.2
2019 Simplifying Deep Neural Networks for FPGA-Like Neuromorphic Systems
abstract
Deep learning using deep neural networks is taking machine intelligence to the next level in computer vision, speech recognition, natural language processing, etc. Brain-like hardware platforms for the brain-inspired computational models are being studied, but the maximum size of neural networks they can evaluate is often limited by the number of neurons and synapses equipped with the hardware. This paper presents two techniques, factorization and pruning, that not only compress the models but also maintain the form of the models for the execution on neuromorphic architectures. We also propose a novel method to combine the two techniques. The proposed method shows significant improvements in reducing the number of model parameters over standalone use of each method while maintaining the performance. Our experimental results show that the proposed method can achieve 30× reduction rate within 1% budget of accuracy for the largest layer of AlexNet.
Jaeyong Chung, Taehwan Shin, Joon-Sung Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2019 Weight Partitioning for Dynamic Fixed-Point Neuromorphic Computing Systems
abstract
Neuromorphic computing systems consist of neurons and synapses with limited programmability, and neural networks are modified to be mapped for such a system. In order to map a perceptron with a large number of connections into a hardware neuron with a fixed, small number of synapses, it is decomposed into a tree of perceptrons, which substantially affects the neuron usage and predictive performance. In this paper, we propose two decomposition algorithms that take advantage of plastic connections and the dynamic scaling capability of neurons. One algorithm based on sorting considers the neuron usage first, and the other algorithm based on packing considers the predictive performance first. Our experimental results on two popular deep convolutional neural networks showed that the sorting-based algorithm substantially improved the accuracy at no cost of neurons compared to a previous work, and the packing-based algorithm improved it even further at a small cost of neurons.
Yongshin Kang, Joon-Sung Yang, Jaeyong Chung
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2018 READ: Reliability Enhancement in 3D-Memory Exploiting Asymmetric SER Distribution
abstract
3D-memory is one of promising applications in 3D-IC technology. With a 3D integration technology, the effective density of memories can increase and the interconnect distance from processor to memory can be shortened. Due to its stacked structure, the upper dies behave as shields blocking outer particles from reaching lower dies, and it makes error rate of the top layer largest among all layers. From a heat perspective, the lower dies would suffer from reliability problems since the lower dies are placed on top of logic die. The heat dissipation can more influence lower dies than upper dies. This creates unequal a reliability distribution for each layer in 3D-memories. A novel ECC organization scheme for 3D-memory to secure reliable operations under soft error rate (SER) profiles is introduced in this paper. The proposed scheme does not require additional redundant arrays. Instead, it utilizes unused spare columns of relatively reliable layer memories to store additional check-bits of less reliable layer memories. It forms a heterogeneous ECC organization across different layers which enhances ECC capabilities in less reliable layers. In addition, redundancy sharing scheme for yield enhancement can be implemented with the proposed scheme. Experimental results show that a memory with the proposed method can tolerate more than three times of a bit-error rate compared to the conventional memory.
Hyunseung Han, Jaeyong Chung, Joon-Sung Yang
IEEE Trans. Computers2
2018 Mitigating Observability Loss of Toggle-Based X-Masking via Scan Chain Partitioning
abstract
The Toggle-based X-masking method requires a single toggle at a given cycle, there is a chance that non-Xvalues are also masked. Hence, the non-Xvalue over-masking problem may cause a fault coverage degradation. In this paper, a scan chain partitioning scheme is described to alleviate non-Xbit over-masking problem arising from Toggle-based X-Masking method. The scan chain partitioning method finds a scan chain combination that gives the least toggling conflicts. The experimental results show that the amount of over-masked bits is significantly reduced, and it is further reduced when the proposed method is incorporated with X-canceling method. However, as the number of scan chain partitions increases, the control data for decoder increases. To reduce a control data overhead, this paper exploits a Huffman coding based data compression. Assuming two partitions, the size of control bits is even smaller than the conventional X-toggling method that uses only one decoder. In addition, selection rules of X-bits delivered to X-Canceling MISR are also proposed. With the selection rules, a significant test time increase can be prevented.
Sae-Eun Kim, Jaeyong Chung, Joon-Sung Yang
IEEE Trans. Computers2
2017 Synthesis of activation-parallel convolution structures for neuromorphic architectures
abstract
Convolutional neural networks have demonstrated continued success in various visual recognition challenges. The convolutional layers are implemented in the activation-serial or fully parallel manner on neuromorphic computing systems. This paper presents an unrolling method that generates parallel structures for the convolutional layers depending on a required level of parallel processing. We analyze the resource requirements for the unrolling of the two-dimensional filters, and propose methods to deal with practical considerations such as stride, borders, and alignment. We apply the propose methods to practical convolutional neural networks including AlexNet and the generated structures are mapped onto a recent neuromorphic computing system. This demonstrates that the proposed methods can improve the performance or reduce the power consumption significantly even without area penalty.
Seban Kim, Jaeyong Chung
DATE2
2017 Energy-efficient response time management for embedded databases
Woochul Kang, Jaeyong Chung
Real Time Syst.2
2016 Simplifying deep neural networks for neuromorphic architectures
abstract
Deep learning using deep neural networks is taking machine intelligence to the next level in computer vision, speech recognition, natural language processing, etc. Brain-like hardware platforms for the brain-inspired computational models are being studied, but none of such platforms deals with the huge size of practical deep neural networks. This paper presents two techniques, factorization and pruning, that not only compress the models but also maintain the form of the models for the execution on neuromorphic architectures. We also propose a novel method to combine the two techniques. The proposed method shows significant improvements in reducing the number of model parameters over standalone use of each method while maintaining the performance. Our experimental results show that the proposed method can achieve 31× reduction rate without loss of accuracy for the largest layer of AlexNet.
Jaeyong Chung, Taehwan Shin
DAC1
2016 Live demonstration: Real-time image classification on a neuromorphic computing system with zero off-chip memory access
abstract
This demo shows a neuromorphic computing system called INsight that classifies images fed by an OV7670 camera in real-time into 10 categories such as dogs, cats, trucks, etc. A 6-layer deep convolutional neural network with 0.7M parameters is trained using the training set of CIFAR-10, and it is compressed and converted into a time delay neural network (TDNN) with 18,000 connections. Then, the TDNN is mapped one-to-one into silicon neurons and synapses on INsight implemented with Xilinx Kintex 7 325T FPGA. This system achieves 80.23% accuracy for the test set of CIFAR-10 and can process 4882 images per second consuming 2.14W power.
Taehwan Shin, Yongshin Kang, Seungho Yang, Seban Kim, Jaeyong Chung
ISCAS5
2016 Defect Diagnosis via Segment Delay Learning
abstract
This brief presents a scalable, high-quality delay defect diagnosis method based on segment delay estimation. Several recent studies have assumed that accurate segment delay recovery leads to a better diagnosis method, and they predict segment delays using Gaussian prior distributions on them. We show that the assumption is not necessarily true and propose to rank segments by the probability of the occurrence for the estimated delays. When random localized defects are considered, prior distributions on segment delays can have a long tail, but all the previous studies fail to model this region properly. We propose to modify the standard deviations of the Gaussian priors depending on the defect size. Our experiment shows that one of our methods achieves ~14% first hit rank improvements and 100× speedup on average over a previous method based on linear programming.
Jaeyong Chung, Woochul Kang
IEEE Trans. Very Large Scale Integr. Syst.1
2015 Bit-Width Optimization by Divide-and-Conquer for Fixed-Point Digital Signal Processing Systems
abstract
This paper presents a novel approach to fractional bit-width optimization of fixed-point designs. We first propose a divide-and-conquer algorithm that can assign optimal fractional bit-widths to a special class of designs that does not have reconvergent paths starting from an internal signal. General designs are partitioned into designs of that special class and our algorithm is applied to each design. The algorithm recursively breaks down a given design into sub-designs and finds Pareto optimal solutions to each sub-design. Those solutions are merged to form Pareto optimal solutions to a larger design. In addition, two pruning methods based on area and error, respectively, are proposed, speeding up the algorithm. The optimization process is guided by static maximum absolute error analysis, and functional correctness is guaranteed for all possible input stimuli. Our approach is demonstrated in five case studies including polynomial approximation and RGB-to-YCbCr conversion, for which the divide-and-conquer algorithm produces the optimal solutions.
Jaeyong Chung
IEEE Trans. Computers1
2015 Segment Delay Learning From Quantized Path Delay Measurements
abstract
Our understanding on a silicon chip is limited due to low measurement resolution or model-silicon miscorrelation including variations. This paper shows that chips are better understood by combining noisy measurement results and model information through a mathematical algorithm. Our proposed method learns segment delays in logic circuits from quantized path delay measurements using ridge regression. During the learning process, we take advantage of both nominal segment delays and the delay sensitivity with respect to variations. We also interpret the ridge regression in Bayesian context and in doing so, propose an analytic formula to set the regularization parameter of the ridge regression. For the silicon measurement environments where low measurement resolution is the dominant source of measurement noise, this formula allows us to predict post-silicon results more accurately and speed up the algorithm eliminating inefficient and inaccurate cross-validation. We also demonstrate our method in enhancing the resolution of already measured path delays. We learn segment delays from quantized path delay measurements and predict the path delays prior to the quantization. Our simulation results show that the predicted path delays are much closer to actual values than the measured values and the nominal values.
Jaeyong Chung, Jibum Kim
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2014 3-D Probe: Low-Cost Variation Modeling Using Intertest-Item Correlations
abstract
Process variation models for variation tolerant designs are developed through expensive silicon characterization. This paper presents a low-cost variation characterization method that takes advantage of correlations between test items. The proposed method is based on compressed sensing (CS), a new innovative theory in signal processing and information theory, and we formulate the problem of accounting for the correlations in the form of standard CS problems, allowing us to leverage advances in CS theory. We consider wafer-level measurement results for multiple test items a 3-D signal and propose the sparsifying transform that combines the 2-D discrete cosine transform and the Karhunen-Loéve transform. Our experimental results show that the proposed method reduces the number of samples required for the same accuracy up to 2X compared to virtual probe when two test items are used.
Jaeyong Chung, Joon-Sung Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2013 Concurrent Path Selection Algorithm in Statistical Timing Analysis
abstract
Circuit timing is becoming more and more uncertain under greater process variation as technology scales. Given the fault probability of each timing path and their statistical correlation from a statistical timing framework, the path selection problem for delay faults has a nature similar to the problem of designing a portfolio of stocks or assets or determining the size of bets in gambling to minimize risk. This observation allows us to develop a very different path selection approach from the conventional ones. If selection of k paths is required in a set of paths, we partition the set into two path sets and determine how many paths should be selected in each path set out of the k paths based on the probabilities of each path set containing faulty paths. We recursively continue this process, which results in the paths to be targeted during tests. The partitioning is easily performed because the paths are already grouped into the depth-first search tree based on their suffix or prefix. Experimental results show that the proposed algorithm can effectively use the correlation to generate high-quality path sets. In addition, we study the issues that occur after automatic test pattern generation on the selected paths, and discuss possible solutions to them.
Jaeyong Chung, Jacob A. Abraham
IEEE Trans. Very Large Scale Integr. Syst.1
2013 A Built-In Repair Analyzer With Optimal Repair Rate for Word-Oriented Memories
abstract
This paper presents a built-in self repair analyzer with the optimal repair rate for memory arrays with redundancy. The proposed method requires only a single test, even in the worst case. By performing the must-repair analysis on the fly during the test, it selectively stores fault addresses, and the final analysis to find a solution is performed on the stored fault addresses. To enumerate all possible solutions, existing techniques use depth first search using a stack and a finite-state machine. Instead, we propose a new algorithm and its combinational circuit implementation. Since our formulation for the circuit allows us to use the parallel prefix algorithm, it can be configured in various ways to meet area and test time requirements. The total area of our infrastructure is dominated by the number of content addressable memory entries to store the fault addresses, and it only grows quadratically with respect to the number of repair elements. The infrastructure is also extended to support various types of word-oriented memories.
Jaeyong Chung, Joonsung Park, Jacob A. Abraham
IEEE Trans. Very Large Scale Integr. Syst.1
2012 Temporal Presence Variation in Immersive Computer Games
abstract
Increasingly, the sophistication of modern computer-gaming systems is becoming comparable to that of immersive, virtual reality (VR) environments, and the popular VR research topic of “presence” is now being explored in the context of computer games. The explosion of popularity of networked gameplay and the movement of computing infrastructure on to the Internet and the cloud mean that technical anomalies such as network latency, dropouts, and so on, may increasingly disrupt players' experience of presence. In this article, a study of these “Breaks in Presence” (BIPs), which examines a networked, first-person-shooter game in an immersive virtual environment, is presented. Our study investigates how participants react to BIPs in terms of their impact on participants' levels of presence and in terms of the time needed for participants to recover from BIPs. Four distinct BIP types, which were selected because of their practical significance and because of their relevance to an established model for presence, were tested. The effects of two contrasting game modes (a low-involvement “navigation game” and a high-involvement “combat game”) on the perceptions of these BIPs are analyzed. As part of our experimental procedure a new video-cued-recall slider technique is introduced. Our study shows that participants experience different levels of impact and recovery from BIPs and that the perceptions of impact of BIPs depend on the overall sense of presence as well as being task dependent. The article shows that recovery time seems to be a well-defined concept and that an overall measure of recovery time exhibits believable correlations with overall presence, with the overall number of BIPs (including “spontaneous” BIPs) and with user characteristics. The article also shows that the slider data for recovery time from BIPs provide evidence for a smooth variation of presence during an immersive experience. It is found that perceptions of impact and recovery time behave differently and that, intriguingly, recovery time appears to be much more independent of game mode than impact is. Evidence is found of strong carry-over effects in participants' recollections of the impact of BIPs from one game experience to the next is found. Results from our slider technique appear to show general agreement with results from a postexperiment questionnaire, and our study motivates the usefulness of our slider technique for future experiments.
Jaeyong Chung, Henry J. Gardner
Int. J. Hum. Comput. Interact.1
2012 Refactoring of Timing Graphs and Its Use in Capturing Topological Correlation in SSTA
abstract
Reconvergent paths in circuits have been a nuisance in various computer-aided design (CAD) algorithms, but no elegant solution to deal with them has been found yet. In statistical static timing analysis (SSTA), they cause difficulty in capturing topological correlation. This paper presents a technique that in arbitrary block-based SSTA reduces the error caused by ignoring topological correlation. We interpret a timing graph as an algebraic expression made up of addition and maximum operators. We define the division operation on the expression and propose algorithms that modify factors in the expression without expansion. As a result, the algorithms produce an expression to derive the latest arrival time with better accuracy in SSTA. Existing techniques handling reconvergent fanouts usually use dependency lists, requiring quadratic space complexity. Instead, the proposed technique has linear space complexity by using a new directed acyclic graph search algorithm. Our results show that it outperforms an existing technique in speed and memory usage with comparable accuracy. More important, the proposed technique is not limited to SSTA and is potentially applicable to various issues due to reconvergent paths in timing-related CAD algorithms.
Jaeyong Chung, Jacob A. Abraham
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2012 On Computing Criticality in Refactored Timing Graphs
abstract
The maximum operator in statistical static timing analysis (SSTA) is a decent approximation for timing sign-off, but often causes significant error in SSTA applications. This paper presents a timing criticality computation method based on non-maximum analytic operators in a parameterized SSTA. After an SSTA run, the proposed method computes the criticality for all edges and nodes in a single graph traversal. Although we do not employ the max operator in the computation process, the error in the maximum operator still degrades the accuracy of the computed criticality because the criticality is a joint probability of expressions, including arrival times, which are computed by the maximum operator during SSTA. To address this issue, we employ the refactoring technique, which was recently proposed to reduce common path pessimism in combinational circuits. This paper shows that refactoring is also very useful in reducing the maximum-induced error in arrival times, and how existing graph-based algorithms can be geared toward refactoring. Our experimental results show that the proposed method reduces the error of the criticality significantly compared to the conventional cutset-based method.
Jaeyong Chung, Jacob A. Abraham
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2012 Path Criticality Computation in Parameterized Statistical Timing Analysis Using a Novel Operator
abstract
This paper presents a method to compute criticality probabilities of paths in parameterized statistical static timing analysis. We partition the set of all the paths into several groups and formulate the path criticality into a joint probability of inequalities. Before evaluating the joint probability directly, we simplify the inequalities through algebraic elimination, handling topological correlation. Our proposed method uses conditional probabilities to obtain the joint probability, and statistics of random variables representing process parameters are changed to take into account the conditions. To calculate the conditional statistics of the random variables, we derive analytic formulas by extending Clark's work. This allows us to obtain the conditional probability density function of a path delay, given the path is critical, as well as to compute criticality probabilities of paths. Our experimental results show that the proposed method provides 4.2X better accuracy on average in comparison to the state-of-art method.
Jaeyong Chung, Jinjun Xiong, Vladimir Zolotov, Jacob A. Abraham
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2012 Testability-Driven Statistical Path Selection
abstract
In the face of large-scale process variations, statistical timing methodology has advanced significantly over the last few years, and statistical path selection takes advantage of it in at-speed testing. In deterministic path selection, the separation of path selection and test generation is known to require time consuming iteration between the two processes. This paper shows that in statistical path selection, this is not only the case, but also the quality of results can be severely degraded even after the iteration. To deal with this issue, we consider testability in the first place by integrating a satisfiability (SAT) solver, and this necessitates a new statistical path selection method. We integrate the SAT solver in a novel way that leverages the conflict analysis of modern SAT solvers, which provides more than 4X speedup without special optimizations of the SAT solver for this particular application. Our proposed method is based on a generalized path criticality metric whose properties allow efficient pruning. Our experimental results show that the proposed method achieves 47% better quality of results on average, and up to 361X speedup compared to statistical path selection followed by test generation.
Jaeyong Chung, Jinjun Xiong, Vladimir Zolotov, Jacob A. Abraham
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2011 Path criticality computation in parameterized statistical timing analysis
abstract
This paper presents a method to compute criticality probabilities of paths in parameterized statistical static timing analysis (SSTA). We partition the set of all the paths into several groups and formulate the path criticality into a joint probability of inequalities. Before evaluating the joint probability directly, we simplify the inequalities through algebraic elimination, handling topological correlation. Our proposed method uses conditional probabilities to obtain the joint probability, and statistics of random variables representing process parameters are changed due to given conditions. To calculate the conditional statistics of the random variables, we derive analytic formulas by extending Clark's work. This allows us to obtain the conditional probability density function of a path delay, given the path is critical, as well as to compute criticality probabilities of paths. Our experimental results show that the proposed method provides 4.2X better accuracy on average in comparison to the state-of-art method.
Jaeyong Chung, Jinjun Xiong, Vladimir Zolotov, Jacob A. Abraham
ASP-DAC1
2011 Post-Silicon Timing Validation Method Using Path Delay Measurements
abstract
In the nanometer era, the mismatch between the pre-silicon model and the post-silicon timing behavior is becoming severer. Therefore, it is necessary to validate timing with post-silicon data. We propose a method that estimates all the segment delays in the observed paths of a design from post-silicon path delay measurements. Our method is based on equality-constrained least squares methods, which enable us to find a unique and optimized solution of segment delays from underdetermined systems. Experimental results show that segment delays obtained using our method achieved correlation ranged from 0.848 to 0.992 to the sampled segment delays for different ISCAS-85 benchmark circuits.
Eun Jung Jang, Jaeyong Chung, Anne E. Gattiker, Sani R. Nassif, Jacob A. Abraham
Asian Test Symposium2
2011 Testability driven statistical path selection
abstract
In the face of large-scale process variations, statistical timing methodology has advanced significantly over the last few years, and statistical path selection takes advantage of it in at-speed testing. In deterministic path selection, the separation of path selection and test generation is known to require time consuming iteration between the two processes. This paper shows that in statistical path selection, this is not only the case, but also the quality of results can be severely degraded even after the iteration. To deal with this issue, we consider testability in the first place by integrating a SAT solver, and this necessitates a new statistical path selection method. Our proposed method is based on a generalized path criticality metric which properties allow efficient pruning. Our experimental results show that the proposed method achieves 47% better quality of results on average, and up to 361x speedup compared to statistical path selection followed by test generation.
Jaeyong Chung, Jinjun Xiong, Vladimir Zolotov, Jacob A. Abraham
DAC1
2011 Off-Chip Skew Measurement and Compensation Module (SMCM) Design for Built-Off Test Chip
Kihyuk Han, Joonsung Park, Jae Wook Lee, Jaeyong Chung, Eonjo Byun, Cheol-Jong Woo, Sejang Oh, Jacob A. Abraham
J. Electron. Test.4
2010 At-speed Test of High-Speed DUT Using Built-Off Test Interface
abstract
This paper presents an efficient test framework to extend a use of low-cost ATE (Automatic Test Equipment) to at-speed test of high-speed DUT (Device Under Test). To bridge the speed gap between the ATE and the DUT, an off-chip test interface circuit, called Built-off Test Interface (BOTI), has been developed. Unlike the previous methods which use on-chip or off-chip self-test circuits, in our method, the ATE plays main role in testing high-speed DUTs by actively controlling the BOTI operation, and monitoring the overall test procedure. This makes the presented method flexible to be applied to various test applications without compromising the test coverage. Also, since the BOTI is implemented off-chip, it does not require hardware modifications of the ATE or the DUT except the DUT load board to accommodate the BOTI module. To maintain reliable off-chip signal communication between the BOTI and the DUT, the BOTI measures off-chip channel skew and compensates the measured skew when communicating signals with the DUT. Currently, the BOTI is configured to do the at-speed test of high-speed memory. The measurement results are presented to validate the functionality of the BOTI, and the effectiveness of the presented test framework.
Joonsung Park, Jae Wook Lee, Jaeyong Chung, Kihyuk Han, Jacob A. Abraham, Eonjo Byun, Cheol-Jong Woo, Sejang Oh
Asian Test Symposium3
2010 A Built-In Self-Test scheme for high speed I/O using cycle-by-cycle edge control
abstract
This paper presents a Built-In Self-Test (BIST) circuit for high speed I/O, based on an embedded pattern generator to remove external factors which could affect the I/O parameters. The rising and falling edge positions of the generated patterns can be controlled independently during every cycle. In the basic operation mode, ATE provides the codes for controlling the edge positions, while in extended mode, an embedded counter generates the control codes. The control of both rising and falling edges makes this scheme especially good for systems with Double-Data Rate (DDR) interfaces. Moreover, the cycle-by-cycle control allows us to analyze efficiently the influence of mismatch trees and per-pin skew on I/O performance. The proposed BIST circuit has been simulated using a 0.18-μm process.
Jaeyong Chung, Jacob A. Abraham, Eonjo Byun, Cheol-Jong Woo
ETS2
2010 Reducing test time and area overhead of an embedded memory array built-in repair analyzer with optimal repair rate
abstract
This paper presents a built-in self repair analyzer with the optimal repair rate for embedded memory arrays. The proposed method requires only a single test, even in the worst case. By performing the must-repair analysis on the fly during the test, it selectively stores fault addresses, and the final analysis to find a solution is performed on the stored fault addresses. To enumerate all possible solutions, existing techniques use depth first search using a stack and a FSM. Instead, we propose a new algorithm and its combinational circuit implementation. Since our formulation for the circuit allows us to use the parallel prefix algorithm, it can be configured in various ways to meet area and test time requirements. The total area of our infrastructure is dominated by the number of CAM entries to store the fault addresses, and it only grows quadratically with respect to the number of repair elements.
Jaeyong Chung, Joonsung Park, Jacob A. Abraham, Eonjo Byun, Cheol-Jong Woo
VTS1
2009 LFSR-Based Performance Characterization of Nonlinear Analog and Mixed-Signal Circuits
abstract
This paper presents an efficient pseudorandom (PR) test method to characterize the performance of nonlinear analog and mixed-signal (AMS) circuits including those embedded in SoC devices. Previous applications of the PR test method to BIST have been limited to digital and linear analog circuits. In this paper, we extend the application of PR test to nonlinear AMS circuits. In doing so, we reduce the cost of testing nonlinear circuits, and increase the test coverage of embedded AMS circuits without incurring a large area overhead to accommodate a test stimulus generator. Our method maintains good test accuracy by using a Volterra series model to describe the behavior of the device under test (DUT). A PR sequence generated from a simple LFSR is used to excite the DUTs over a wide range of frequencies and estimate the parameters of the Volterra series, which are then used to predict the performance of DUTs. We present a method to reduce the test time by using a compressed cross-correlation method which reduces the complexity of the presented algorithm. The mathematical background and hardware measurement results are presented to validate our method.
Joonsung Park, Jaeyong Chung, Jacob A. Abraham
Asian Test Symposium2
2009 A hierarchy of subgraphs underlying a timing graph and its use in capturing topological correlation in SSTA
abstract
This paper shows that a timing graph has a hierarchy of specially defined subgraphs, based on which we present a technique that captures topological correlation in arbitrary block-based statistical static timing analysis (SSTA). We interpret a timing graph as an algebraic expression made up of addition and maximum operators. We define the division operation on the expression and propose algorithms that modify factors in the expression without expansion. As a result, they produce an expression to derive the latest arrival time with better accuracy in SSTA. Existing techniques handling reconvergent fanouts usually use dependency lists, requiring quadratic space complexity. Instead, the proposed technique has linear space complexity by using a new directed acyclic graph search algorithm. Our results show that it outperforms an existing technique in speed and memory usage with comparable accuracy.
Jaeyong Chung, Jacob A. Abraham
ICCAD1
2009 Recursive Path Selection for Delay Fault Testing
abstract
This paper presents a new path selection algorithm for delay fault testing in a statistical timing framework. Existing algorithms which consider correlation between paths use an iterative process for each path or defect and require a Monte Carlo simulation for each iteration to calculate the conditional fault probability. The proposed algorithm does not require the iteration process and selects a requested number of paths simultaneously once it performs a statistical timing analysis at the beginning. If selection of k paths is required in a set of paths, it partitions the set into two path sets and determines how many paths should be selected in each path set out of the k paths. It recursively continues this process and ends up with k paths. The partitioning is easily performed during the recursive traversal of a circuit, which produces an imaginary path tree, where paths are already grouped based on their prefix. Experimental results show the proposed algorithm can effectively use structural correlation and spatial correlation to generate high quality path sets.
Jaeyong Chung, Jacob A. Abraham
VTS1
2006 Color Object Tracking System for Interactive Entertainment Applications
abstract
Looking human through a camera is a potentially powerful technique to facilitate human-computer interaction. Many 2D/3D visual human tracking techniques have been introduced in area of computer vision, and applied to smart surveillance, motion analysis, interactive computer graphics and virtual reality applications. For applications emphasizing real time interactivity, the visual tracking techniques need to be fast, robust, and preferably run on an inexpensive hardware system. In this paper, we present a relatively inexpensive (e.g. run on a high-end PC) but reasonably robust real time motion tracking system based on a simple 2D color object segmentation and recognition algorithm. Users grasp color objects for achieving more exact tracking performance on their hands and make predefined motion gestures continuously. Through object segmentation processing based on color information, 2D positions of the objects are computed. And then, a motion gesture corresponding on these 2D position trajectories is found by a simple correlation-based matching algorithm. Also, we demonstrate this system by applying it to a popular interactive entertainment (e.g. TETRIS).
Jaeyong Chung, Kwanghyun Shim
ICASSP (5)1
2006 A Dynamic Load Balancing for Massive Multiplayer Online Game Server
Jungyoul Lim, Jaeyong Chung, Jin Ryong Kim, Kwanghyun Shim
ICEC2