VLDB 2026 Research / reviewers in the wild / expert
Jeong-A Lee
dblp:47/1839
· DBLP profile ↗
43ranked-venue papers
4as first author
19since 2021 · last 2026
0000-0002-5166-0629ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 3 first-author · 10 since 2021Security and privacy · 3 · 1 since 2021Theory of computation · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TRM: An Efficient Hypervisor-Based Framework For Malware Analysis and Memory ReconstructionabstractModern rootkits leverage kernel privileges to hide from analysis tools, while obfuscation techniques render them resistant to static analysis. Reverse engineering malware requires observing memory usage and reconstructing data structures. Existing tools rely on instrumentation or emulation, which introduce high overhead, leave detectable artifacts, and cannot reliably analyze kernel-level malware. We present The Reversing Machine (TRM), a hypervisor-based framework for high-performance introspection of evasive malware. It is the first system to support selective memory tracing and structure reconstruction in the hypervisor. TRM repurposes hardware virtualization to efficiently detect user-kernel mode transitions and obtain memory traces independently of a potentially compromised guest kernel. TRM introduces new insights into leveraging hardware virtualization for runtime memory reconstruction and analysis of data structures while remaining invisible to malware. We demonstrate automatic reconstruction of function signatures and data structures, reduce system call latency overhead from 142% to 57% compared to prior work, and accelerate manual reverse engineering by 43% on average, even for complex kernel objects. TRM shows that hypervisor-level memory tracing makes data structure reconstruction practical in hostile environments, bridging the gap between prior feasibility studies and real-world malware analysis. Mohammad Sina Karvandi, Soroush Meghdadi Zanjani, Sima Arasteh, Saleh Khalaj Monfared, Mohammad K. Fallah, Saeid Gorgin 0001, Jeong-A Lee, Asia Slowinska, Erik van der Kouwe |
AsiaCCS | 7 |
| 2025 | Poster: Integration of Wearable and Affective Computing via Abstraction and Decision Fusion ArchitectureabstractThis paper introduces an efficient emotion detection method to integrate wearable and affective computing paradigms. Our research contributes to advancing emotion detection technologies, offering potential applications in diverse domains such as healthcare, human-computer interaction, and personalized computing experiences. Our approach addresses the increasing need for real-time emotion recognition while minimizing computational demands. By leveraging low-computation techniques, we propose a novel framework that achieves high accuracy in emotion detection. Besides, advanced data abstraction methods are developed to reduce data workload keeping detection performance. Experimental results demonstrate a notable accuracy rate of $89 .77$%, affirming the efficacy of our proposed method. Mohammadreza Najafi, Mohammad K. Fallah, Saeid Gorgin 0001, Ghassem Jaberipur, Jeong-A Lee |
WoWMoM | 5 |
| 2025 | USEFUSE: Uniform stride for enhanced performance in fused layer architecture of deep neural networksabstractConvolutional Neural Networks (CNNs) are crucial in various applications, but deploying them on resource-constrained edge devices poses challenges. This study presents the Sum-of-Products (SOP) units for convolution, which utilize low-latency left-to-right bit-serial arithmetic to minimize response time and enhance overall performance. The study proposes a methodology for fusing multiple convolution layers to reduce off-chip memory communication and increase the overall performance. An effective mechanism detects and skips inefficient convolutions after ReLU layers, minimizing power consumption without compromising accuracy. Additionally, efficient tile movement guarantees uniform access to the fusion pyramid. An analysis demonstrates the uniform stride strategy improves operational intensity. Two designs cater to varied demands: one focuses on minimal response time for mission-critical applications, and another focuses on resource-constrained devices with comparable latency. This approach notably reduced redundant computations, improving the efficiency of CNN deployment on edge devices. Muhammad Sohail Ibrahim, Jeong-A Lee |
J. Syst. Archit. | 3 |
| 2025 | Balanced Modular Addition for the Moduli Set $ \{2^{q},2^{q}\mp 1,2^{2q}+1\}${2q,2q∓1,22q+1} via Moduli-($ 2^{q}\mp \sqrt{-1}$2q∓-1) AddersabstractModuli-set$ \mathbf{\tau}=\{2^{\boldsymbol{q}},2^{\boldsymbol{q}}\pm 1\}$is often the base of choice for realization of digital computations via residue number systems. The optimum arithmetic performance in parallel residue channels, is generally achieved via equal bit-width residues (e.g.,$ \boldsymbol{q}~ \mathbf{i}\mathbf{n}~ \mathbf{\tau}$) that usually leads to equal computation speed within all the residue channels. However, the commonly difficult and costly task of reverse conversion (RC) is often eased in the existence of conjugate moduli. For example,$ 2^{\boldsymbol{q}}\mp 1\in \mathbf{\tau}$, lead to the efficient modulo-($ 2^{2\boldsymbol{q}}-1$) addition, as the bulk of$ \mathbf{\tau}$-RC, via the New-CRT reverse conversion method. Nevertheless, for additional dynamic range,$ \mathbf{\tau}$is augmented with other moduli. In particular,$ \mathbf{\phi}=\mathbf{\tau}\cup \{2^{2\boldsymbol{q}}+1\}$, leads to efficient RC, where the added modulo is conjugate with the product$ 2^{2\boldsymbol{q}}-1$of$ 2^{\boldsymbol{q}}\mp 1\in \mathbf{\tau}$. Therefore, the final step of$ \mathbf{\phi}$-RC would be fast and low cost/power modulo-($ 2^{4\boldsymbol{q}}-1$) addition. However, the$ 2\boldsymbol{q}$-bit channel-width jeopardizes the existing delay-balance in$ \mathbf{\tau}$. As a remedial solution, given that$ 2^{2\boldsymbol{q}}+1=\left(2^{\boldsymbol{q}}-\boldsymbol{j}\right)\left(2^{\boldsymbol{q}}+\boldsymbol{j}\right)$, with$ \boldsymbol{j}=\sqrt{-1}$, we design and implement modulo-($ 2^{2\boldsymbol{q}}+1$) adders via two parallel$ \boldsymbol{q}$-bit moduli-($ 2^{\boldsymbol{q}}\mp \boldsymbol{j}$) adders. The analytical and synthesis based evaluations of the proposed modulo-($ 2^{\boldsymbol{q}}\mp \boldsymbol{j}$) adders show that the delay-balance of$ \mathbf{\tau}$is preserved with no cost overhead vs.$ \mathbf{\phi}$. In particular, the binary-to-complex and complex-to-binary convertors are merely cost-free and immediate. Ghassem Jaberipur, Elham Rahman, Jeong-A Lee |
IEEE Trans. Computers | 3 |
| 2025 | A comprehensive review and classification of micro-scale and macro-scale interconnection network simulators for research and education in network-based computing systems
Atiyeh Gheibi-Fetrat, Fatemeh Serajeh-hassani, Negar Akbarzadeh, Amir Mirzaei, Mahmoud Reza Kheyrati-Fard, Ahmad Javadi Nezhad, Jeong-A Lee, Hamid Sarbazi-Azad |
J. Supercomput. | 7 |
| 2025 | A survey of SSD simulators and emulators
Atiyeh Gheibi-Fetrat, Fatemeh Serajeh-hassani, Masoud Mohammadi-Lak, Amir Mirzaei, Negar Akbarzadeh, Mahmoud Reza Kheyrati-Fard, Mohammad Hosseini 0001, Ahmad Javadi Nezhad, Arash Tavakkol, Jeong-A Lee, Hamid Sarbazi-Azad |
J. Supercomput. | 10 |
| 2025 | Efficient hardware accelerators for k-nearest neighbors classification using most significant digit first arithmetic
Saeid Gorgin 0001, Malik Zohaib Nisar, Jeong-A Lee |
J. Supercomput. | 3 |
| 2025 | MQSimNet: an open-source simulator for next-generation network-based SSDs
Amir Mirzaei, Fatemeh Serajeh-hassani, Atiyeh Gheibi-Fetrat, Mina Zabihi, Sina Ghorbani-Jabbedar, Mahmoud Reza Kheyrati-Fard, Ahmad Javadi Nezhad, Mohammad Hosseini 0001, Negar Akbarzadeh, Jeong-A Lee, Hamid Sarbazi-Azad |
J. Supercomput. | 10 |
| 2024 | Montgomery Modular Multiplication via Single-Base Residue Number Systems
Zabihollah Ahmadpour, Ghassem Jaberipur, Jeong-A Lee |
ARITH | 3 |
| 2024 | Comparative Analysis of Fault-Localization Techniques in AdderabstractAdder is a complex digital circuit because of its interconnectivity, which hinders fault detection and localization. This paper compares two approaches of fault localization in adders based on carry-free addition using signed digits (SD) representation and localized self-checking-full adders for ripple carry adders (RCA). The self-repairing SD adder approach requires computation with standard and shifted inputs toward left and right for fault localization. The resulting complexity caused by the shifting operation is unnecessary for the self-repairing RCA, which simultaneously achieves fault detection and localization. Moreover, the centralized checking mechanism of the SD adder results in system failure if the checker becomes faulty. However, in the case of self-repairing RCA, the failure of individual full adders will not create problems in the self-checking ability of other full adders owing to the distributed fault detection mechanism. It has been observed that the self-checking RCA implemented in FPGA is $\mathbf{6 2 \%}$ more area efficient than the self-checking signed digit adder in terms of LUTs. Muhammad Ali Akbar, Jeong-A Lee, Amine Bermak |
IWCMC | 2 |
| 2024 | An ultra-low-computation model for understanding sign languages
Mohammad K. Fallah, Mohammadreza Najafi, Saeid Gorgin 0001, Jeong-A Lee |
Expert Syst. Appl. | 4 |
| 2023 | Modulo-(2q - 3) Multiplication with Fully Modular Partial Product Generation and ReductionabstractGiven the residue number systems that contain moduli of the form 2q± 1 and 2q± 3, it is desirable to employ delay-balanced adders and multipliers, in order to synchronize the operation of parallel residue channels. The required modulo-(2q± 3) adders, with compatible speed with modulo- (2q± 1) adders, already exist with parallel prefix architectures. However, the previously reported modulo-(2q± 3) multipliers, in one way or another, produce the non-modular products of the residues at the outset and work towards yielding the final modular product. This seems to be the main source of incompatible performance with the existing modulo- (2q± 1) fully modular multipliers. Therefore, as the first endeavor, we were motivated to design and implement efficient modulo-(2q− 3) multipliers with fully modular partial product generation and reduction that are more compatible with their modulo- (2q− 1) counterparts. However, unlike the case of modulo 2q− 1, it turns out that the straightforward modulo-(2q− 3) partial product reduction (e.g., via Wallace-tree reduction with greedy use of full adders and half adders) falls into an infinite loop of reduction stages. Therefore, we undertake a modified reduction algorithm that requires at most two reduction levels more than that of the modulo- (2q− 1) case to converge. To ensure the correct operation of the algorithm and ease the design process, an in-house software program produces the exact composition of reduction cells in each level of partial product reduction. Analytical and synthesis-based evaluations of the proposed design, and the previous ones, exhibit better figures of merit, as regards the delay (≥ 24%), area-delay (≥ 6%) and energy (≥ 10%) measures. Ghassem Jaberipur, Saeid Gorgin 0001, Navid Ahamadian, Jeong-A Lee |
ARITH | 4 |
| 2023 | DSLOT-NN: Digit-Serial Left-to-Right Neural Network AcceleratorabstractWe propose a Digit-Serial Left-tO-righT (DSLOT) arithmetic based processing technique called DSLOT-NN with aim to accelerate inference of the convolution operation in the deep neural networks (DNNs). The proposed work has the ability to assess and terminate the ineffective convolutions which results in massive power and energy savings. The processing engine is comprised of low-latency most-significant-digit-first (MSDF) (also called online) multipliers and adders that processes data from left-to-right, allowing the execution of subsequent operations in digit-pipelined manner. Use of online operators eliminates the need for the development of complex mechanism of identifying the negative activation, as the output with highest weight value is generated first, and the sign of the result can be identified as soon as first nonzero digit is generated. The precision of the online operators can be tuned at runtime, making them extremely useful in situations where accuracy can be compromised for power and energy savings. The proposed design has been implemented on Xilinx Virtex-7 FPGA and is compared with state-of-the-art Stripes on various performance metrics. The results show the proposed design presents power savings, has shorter cycle time, and approximately 50% higher OPS per watt. Muhammad Sohail Ibrahim, Malik Zohaib Nisar, Jeong-A Lee |
DSD | 4 |
| 2023 | A new energy-efficient and temperature-aware routing protocol based on fuzzy logic for multi-WBANs
Danial Javaheri, Pooia Lalbakhsh, Saeid Gorgin 0001, Jeong-A Lee, Mohammad Masdari |
Ad Hoc Networks | 4 |
| 2023 | Fuzzy logic-based DDoS attacks and network traffic anomaly detection methods: Classification, overview, and future perspectives
Danial Javaheri, Saeid Gorgin 0001, Jeong-A Lee, Mohammad Masdari |
Inf. Sci. | 3 |
| 2023 | On the Evolutionary Synthesis of Fault-Resilient Digital CircuitsabstractIn the event of an upset, fault-resilient circuits maintain correct functionality allowing the system to remain fully operational or at least operate with a graceful degradation. Every circuit has a certain level of inherent resilience to faults. Often times, this inherent resilience to faults is insufficient for the given application. This is because conventional synthesis tools generally only focus on optimizing a circuit with respect to area, power, or timing budgets. There is a wide range of applications where faulty circuit behavior can lead to fatal results. Fault injection analyses are reported and show that even a single fault can be critical to the desired circuit operation. To which end, this article presents synthesis of fault-resilient (SYFR) circuits, an evolutionary method for automated synthesis of increased fault-resilience digital circuits suitable for fine-grained use. Test results for synthesis of up to 60 input circuits with SYFR are reported. SYFR can be repeatedly applied to a circuit to obtain various design tradeoffs between fault resilience and implementation costs. SYFR can also be flexibly applied to build circuits, which are selectively fault resilient, i.e., their tolerance to faults is workload aware. In addition, a novel population seeding mechanism to reduce the design space is introduced and experimentally validated. In summary, this article demonstrates that SYFR can be considered a competitive synthesis methodology for constructing fault-resilient circuits. Umar Afzaal, Abdus Sami Hassan, Jeong-A Lee |
IEEE Trans. Evol. Comput. | 4 |
| 2023 | A fast MILP solver for high-level synthesis based on heuristic model reduction and enhanced branch and bound algorithm
Mina Mirhosseini, Mahmood Fazlali, Mohammad K. Fallah, Jeong-A Lee |
J. Supercomput. | 4 |
| 2022 | An Efficient FPGA Implementation of k-Nearest Neighbors via Online Arithmeticabstractk-NN, as one of the well-employed classification algorithms, severely suffers from a computationally intensive nature. This paper exploits the parallelism and digit level pipelining opportunities via FPGA devices and Online arithmetic to offer an efficient k-NN FPGA implementation. All the required operations for computing distances and sorting are applied to serially coming data. Moreover, we dynamically terminate the unnecessary computations once they are detected. To the best of our knowledge, the proposed k-NN implementation is the first one that used FPGA and Online arithmetic effectively. It provides up to 34% speedup compared to the best state-of-the-art design. Saeid Gorgin 0001, MohammadHosein Gholamrezaei, Danial Javaheri, Jeong-A Lee |
FCCM | 4 |
| 2022 | An Energy-Efficient K-means Clustering FPGA Accelerator via Most-Significant Digit First ArithmeticabstractK-means clustering is the most well-known unsupervised learning method that partitions the input dataset into$K$clusters based on the similarity between the data samples. In this paper, to achieve an energy-efficient implementation without sacrificing performance, we take advantage of massive parallelism and digit-level pipelining via FPGA and the most-significant digit first arithmetic. Having the result of the most-significant digits in advance provides the possibility of early termination for unnecessary computations and fetching just the required most-significant part of data points from memory. This early termination technique significantly increases performance and decreases energy consumption. Our experimental results from various datasets and comparisons with the state-of-the-art FPGA accelerators indicate that our proposed design has effectively reduced energy consumption without any performance loss. Saeid Gorgin 0001, MohammadHosein Gholamrezaei, Danial Javaheri, Jeong-A Lee |
FPT | 4 |
| 2020 | Trading the Reliability of Approximate TMR in FPGAs with the Cost of MitigationabstractA number of works have focused on relaxing circuit specification for building partial TMR circuits, but a framework for analysing the effect of circuit degradation on its dependability is missing in the literature. This paper aims to bridge this gap by developing a reliability model for approximate TMR circuits implemented in FPGAs and the parameter definitions for controlling design trade-offs with reliability in the approximation process. The framework is useful in the assignment of the design parameters such that the reliability constraints of the application at hand are satisfied. Reliability curves for different trade-offs show a sharp decline in reliability even at small error thresholds, requiring that for maintaining system operation in the high reliability region, a TMR approximation method must achieve the desired reduction in hardware overheads within tight error constraints. Furthermore, the effect of an unprotected voter on the overall system reliability is also quantified. To which end, it is shown that the simple unprotected voter results in significant degradation on the approximate TMR reliability and therefore using a fault-tolerant voting circuit is essential to a reliable system operation. Umar Afzaal, Jeong-A Lee |
DSD | 2 |
| 2020 | Data Footprint Reduction in DNN Inference by Sensitivity-Controlled Approximations with Online ArithmeticabstractIn deep neural network (DNN) inference, researchers have been trying to reduce the number of computations and connections without performance degradation, departing from a bit-parallel to a bit-serial mode of arithmetic. In this regard, approximations translated as the mixed-precision profile for among-layer-mixed-precision through bit-serial architecture have been adopted in the literature. However, the introduction of within-layer mixed precision through controlled approximations for low-latency DNN architecture is yet to be studied. For DNN inference in this study, we apply an unconventional computation technique of online arithmetic, which serially generates the most significant digits first(MSDF) and then terminates computation according to the required precision. Specifically, Taylor expansion-based sensitivity analysis guides the within-layer-mixed-precision method for the choice of approximation intensity (desired bits) for weights and activations of convolutional layers. In turn, the within-layer-mixed-precision method drives the termination of the convolution operation carried out using an online multiplier. Hence, we aim to reduce the data footprint by early terminations achieved thanks to the insightful nature of within-layer-mixed-precision instead of among-layer-mixedprecision for online convolution. In this manner, convolution operations compute not-more-than-necessary most significant digits to overcome the bottleneck of data footprint for in-demand edge computing devices. Abdus Sami Hassan, Tooba Arifeen, Jeong-A Lee |
DSD | 3 |
| 2019 | AFP-CKSAAP: Prediction of Antifreeze Proteins Using Composition of k-Spaced Amino Acid Pairs with Deep Neural NetworkabstractAntifreeze proteins (AFPs) are the sub-set of ice binding proteins indispensable for the species living in extreme cold weather. These proteins bind to the ice crystals, hindering their growth into large ice lattice that could cause physical damage. There are variety of AFPs found in numerous organisms and due to the heterogeneous sequence characteristics, AFPs are found to demonstrate a high degree of diversity, which makes their prediction a challenging task. Herein, we propose a machine learning framework to deal with this vigorous and diverse prediction problem using the manifolding learning through composition of k-spaced amino acid pairs. We propose to use the deep neural network with skipped connection and ReLU non-linearity to learn the non-linear mapping of protein sequence descriptor and class label. The proposed antifreeze protein prediction method called AFP-CKSAAP has shown to outperform the contemporary methods, achieving excellent prediction scores on standard dataset. The main evaluater for the performance of the proposed method in this study is Youden's index whose high value is dependent on both sensitivity and specificity. In particular, AFP-CKSAAP yields a Youden's index value of 0.82 on the independent dataset, which is better than previous methods. Jeong-A Lee |
BIBE | 2 |
| 2019 | DURE: An Energy- and Resource-Efficient TCAM Architecture for FPGAs With Dynamic UpdatesabstractTernary content-addressable memory (TCAM) designed using static random-access memory (SRAM)-based field-programmable gate arrays (FPGAs) offers a promising lookup performance. However, the update process in a TCAM table poses significant challenges for efficiently employing SRAM-based TCAM. SRAM-based TCAM for FPGAs is designed using block RAM or distributed RAM resources in FPGAs. Such designs suspend search operations during an already high-latency update operation, rendering them infeasible in applications that require high-frequency updates. This paper presents a dynamically updatable energy- and resource-efficient TCAM design (DURE) based on FPGAs. DURE exploits the distributed RAM resources in FPGAs. More specifically, the lookup table RAMs (LUTRAMs) available in SLICEM resources are configured as quad-port RAM, which constitutes the basic memory (BM) block in the implementation of DURE. The contents of the TCAM table are divided into chunks of equal size and mapped onto the LUTRAMs of the proposed BM blocks. DURE implements dynamic updates by reconfiguring the LUTRAMs of only those BM blocks that are associated with the word being updated, thereby allowing search and update operations to be performed simultaneously. This achieves a lookup rate of 335 million lookups per second, with an update rate of 5.15 million updates per second on a 512 × 36 size TCAM on a Virtex-6 FPGA. Compared with the existing SRAM-based TCAMs, DURE has a smaller single-cycle search latency and achieves at least 2.5 times more energy efficiency and a 67% higher performance per area. Inayat Ullah, Zahid Ullah 0001, Umar Afzaal, Jeong-A Lee |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2018 | Effect of FPGA Circuit Implementation on Error Detection Using Logic Implication CheckingabstractAggressive scaling of circuits to achieve smaller feature sizes has led to an increased concern about their reliability as small scale circuits age faster. Thus, an increase in the number of computational errors due to defects is expected in the nanoscale dimensions. Concurrent error detection techniques including logic implication-based checking can detect a partial number of these errors at lower area costs. In this paper, we evaluate the performance of this mode of error detection in implemented circuits, specifically FPGA circuits where it is possible for a single fault to affect multiple logic paths. Fault injection experiments show that the probability of error detection achieved for circuits that are implemented in FPGAs is significantly less than that predicted by fault simulations on their corresponding netlists, almost by half. It is thus shown that the efficiency of implication relationships in detecting errors not only varies from one circuit to another but that it also depends largely on the implementation of the circuit under test as supported through analytic analyses and experimental results. Umar Afzaal, Abdus Sami Hassan, Tooba Arifeen, Jeong-A Lee |
DSD | 4 |
| 2018 | Error Correctable Approximate Multiplier with Area/Power Efficient Design Through Mixed CMOS/PTLabstractThe following topics are dealt with: field programmable gate arrays; embedded systems; learning (artificial intelligence); multiprocessing systems; cryptography; system-on-chip; optimisation; medical signal processing; Internet of Things; logic design. Tooba Arifeen, Abdus Sami Hassan, Jeong-A Lee |
DSD | 3 |
| 2018 | High-Speed Configuration Strategy for Configurable Logic Block-Based TCAM Architecture on FPGAabstractTernary content addressable memory (TCAM)-based high-speed search engines are extensively used to accelerate applications in network routing. Researchers have used field programmable gate arrays (FPGAs) to implement TCAM-based classification engines. The existing RAM-based architectures of TCAM undervalue the massive parallelism capabilities of FPGA which is evident from their relatively high latencies. This paper presents an FPGA-based TCAM design which exploits the rich logic resources available on contemporary FPGAs for an increased throughput parallel implementation compared to existing designs. The encoded bits of the traditional TCAM table are stored in the flip-flop resources, and the comparison operation is performed using the look-up table resources of the slices of configurable logic blocks. The structure and simple mapping methodology of the proposed design enables the implementation of a high-speed update operation in only a single cycle. We implemented our proposed TCAM design on a Virtex-6 XC6VLX760 FPGA device. Our proposed design achieves almost 30 % better throughput compared to existing TCAM designs for FPGAs with an efficient utilization of FPGA resources. Inayat Ullah, Umar Afzaal, Zahid Ullah 0001, Jeong-A Lee |
DSD | 4 |
| 2018 | Generation Methodology for Good-Enough Approximate Modules of ATMR
Abdus Sami Hassan, Tooba Arifeen, Hossein Moradian, Jeong-A Lee |
J. Electron. Test. | 4 |
| 2017 | FPGA-based design of a self-checking TMR voterabstractThe most common error mitigation scheme used for hardening designs against radiation-induced upsets on FPGAs is Triple Modular Redundancy (TMR). In a TMR system, there are three copies of a module and voting circuits that mask errors by voting for the majority. There are several types of voting circuits which can be classified based on their insertion sites in the design, functionality or the type of data structure to be mitigated. These voters are mostly built from Look-up Tables (LUTs) but these voters just like the design that is hardened by applying TMR are also susceptible to radiation-induced effects. In this paper, we present the design of a self-checking LUT based 1-bit voter intended for those sites where a TMR system reduces to a duplex or a simplex system. Voters on these sites make a single point of failure and the proposed design avoids this situation by attaching multiple voter redundancies to the same output. The operation of the proposed voter has been verified through its hardware implementation and timing simulation. Umar Afzaal, Jeong-A Lee |
FPL | 2 |
| 2016 | Probing Approximate TMR in Error Resilient Applications for Better Design TradeoffsabstractApproximate TMR (ATMR) is an approach towards logic masking of soft errors through utilization of approximate circuit modules in order to achieve performance benefits such as area, delay and power consumption. In this work, we propose a novel technique for efficient utilization of acceptable error rate threshold for error resilient applications to perform logic masking. The novelty of the method lies in the systematic methodology for development of ATMR modules with better design tradeoffs. Experimental results and comparative analysis reflect how acceptable error rate threshold is used through a systematic method for ATMR development with minimum overhead. Tooba Arifeen, Abdus Sami Hassan, Hossein Moradian, Jeong-A Lee |
DSD | 4 |
| 2015 | Low-Cost Fault Localization and Error Correction for a Signed Digit Adder Design Utilizing the Self-Dual ConceptabstractThis paper details a low-cost fault localization and error correction technique for binary signed-digit adders that utilizes the self-dual concept. This new design approach will result in a higher reliability, i.e., 100% of the single stuck-at faults can be corrected. Our technique is based on the observation that the proposed method under the existence of any single stuck-at fault, when fed by the complement of its functional input, yields the fault-free complement of the desired output. First, we apply parity-based error detection modules. Upon detection of a fault, this is followed by input inversion, re-computation, and suitable output inversion. By comparing the faulty and fault-free outputs, we can localize the faulty component. The proposed approach has higher reliability and lower complexity compared to previous related works. Hossein Moradian, Jeong-A Lee |
DSD | 2 |
| 2015 | Double phase modular steganography with the help of error images
Masoud Afrakhteh, Inkyu Moon, Jeong-A Lee |
Multim. Tools Appl. | 3 |
| 2015 | Adaptive least significant bit matching revisited with the help of error imagesabstractAbstract State‐of‐the‐art steganographic schemes, such as highly undetectable stego (HUGO) and its extended version, aim at least significant bit (LSB)‐based approaches embedding up to 1 bpp and are more concerned about the undetectability level of the stego image rather than the peak signal‐to‐noise ratio. The complexity of such methods is quite high, too. In this work, a steganographic scheme is proposed in a spatial domain that takes advantage of error images resulting from applying an image quality factor (the same as the ones used in JPEG compression) in order to find the pixels where a slight change could be made. The amount of change is adaptively embedded using LSB matching revisited. We show that our proposed method is less detectable than HUGO and almost as undetectable as the extended HUGO while it has a greater time performance. Copyright © 2014 John Wiley & Sons, Ltd. Masoud Afrakhteh, Jeong-A Lee |
Secur. Commun. Networks | 2 |
| 2015 | Parallel modular steganography using error imagesabstractAbstract The embedding process for most of the state‐of‐the‐art adaptive steganographic methods (not easily detectable) such as highly undetectable steganography (HUGO), edge adaptive image steganography (EA), and Adaptive ±1 Steganography in Extended Noisy Regions, is contingent upon the neighboring pixel values to prevent discovery by steganalysis methods. Hence, the calculation of the possible embedding rate must be suspended until the embedding process is performed pixel by pixel, in sequence. If, however, the cover image is JPEG‐compressed, an error image can be considered as an embedding map implying where and to what extent secret bits can be embedded. In this paper, we present a modular steganography using error images. It is demonstrated that the detectability level of the resulting stego images is lower than that of state‐of‐the‐art schemes and the proposed method executes 2.32 times faster than HUGO. Further, to improve the time performance, a block‐wise parallel algorithm is presented where a single thread is assigned to each block of pixels. Because of the independent behavior of the embedding method and unlike the other methods, it is proved that each block of pixels can be embedded independently. The effectiveness of this parallel approach is analyzed and verified to execute approximately 55 times faster than serial execution. Copyright © 2014 John Wiley & Sons, Ltd. Masoud Afrakhteh, Jeong-A Lee |
Secur. Commun. Networks | 2 |
| 2013 | Self-Checking Carry Select Adder with Fault LocalizationabstractThe common design problem in various approaches for self-checking adders is the fault propagation due to carry. Such a fault can misguide the system to detect the particular faulty module. In this paper, we proposed a self-checking Carry Select Adder (CSA) with fault localization ability. Our scheme can provide minimum area overhead for self-recovery process because instead of replacing the whole system we can now replace the particular faulty modules. The proposed self-checking CSA consumes 12% less area with equal performance as compared to the previously proposed self-checking CSA approach. Muhammad Ali Akbar, Jeong-A Lee |
DSD | 2 |
| 2013 | Area-Time Efficient Self-Checking ALU Based on Scalable Error Detection CodingabstractIn this paper, we propose a self-checking ALU based on Scalable Error Detection Coding (SEDC) algorithm which is capable of detecting 100% unidirectional errors. The SEDC encoded ALU is scalable based on the input data length. The area overhead increases only by a small factor with increase in data length while the computational latency remains constant irrespective of the binary data length. We show that the proposed scheme is faster and easily scalable compared with the previous implementations based on Berger and Bose-Lin codes and also more area efficient than Berger Check Prediction based ALU. Zahid Ali 0001, Park Hui-Jong, Jeong-A Lee |
DSD | 3 |
| 2013 | A novel run-time auto-reconfigurable FPGA architecture for fast fault recovery with backward compatibility (abstract only)abstractA self-repairing fault-tolerant FPGA architecture is developed which is also compatible with existing island-style routing network. Due to this backward compatibility, the proposed architecture can not only be implemented easily in the existing FPGA devices but a new fault-tolerant FPGA device can also be fabricated utilizing the existing island-style routing architecture. A generic fault-tolerant Computation Cell is developed which can be incorporated in existing FPGA CLB (Configurable Logic Block) having 8 LUTs at least. The proposed fault-tolerant FPGA architecture is comprised of Computation Tiles each of which consists of computation cells which are able to heal themselves from transient errors. Computation Tile also contains stem cells which help computation cells to recover from permanent errors all at once. This architecture is centrally controlled by an on-chip fault-tolerant core whose main responsibility is to define the healing priority when an error occurs in more than one of the computation tile at the same time. It also communicates with the external PC software which identifies the faulty tile and reconfigures it through dynamic partial reconfiguration. The robust operation of a proposed architecture is implemented and verified on XILINX Virtex-5 FPGA device. From our proposed fault-tolerant scheme of utilizing the existing routing strategies together with partial reconfiguration of stem cells we achieved a number of benefits, including a fast fault recovery and avoidance of using complicated routing strategies, as compared to recently developed fault-tolerant FPGA architectures. Hasan Baig, Jeong-A Lee |
FPGA | 2 |
| 2012 | An island-style-routing compatible fault-tolerant FPGA architecture with self-repairing capabilitiesabstractIn this work, we have developed a fault-tolerant architecture which is compatible with existing island-style routing network. Due to this compatibility, the proposed architecture can not only be implemented easily in the existing FPGA devices but a new fault-tolerant FPGA device can also be fabricated without refining the existing routing architecture. A generic fault-tolerant Computation Cell is developed which along with its self-checking circuitry also consists of an internal router to route un-faulty function out of the cell. The proposed fault-tolerant FPGA architecture is comprised of Computation Tiles each of which consists of computation cells which are able to heal themselves from transient errors. Computation Tile also contains stem cells which help computation cells to recover from permanent errors all at once. This architecture is centrally controlled by an on-chip fault-tolerant core whose main responsibility is to define the healing priority when an error occurs in more than one of the computation tile at the same time. It also communicates with the external PC software which identifies the faulty tile and reconfigures it through dynamic partial reconfiguration. The robust operation of a proposed architecture is implemented and verified on XILINX Virtex-5 FPGA device. From our proposed fault-tolerant scheme of utilizing the existing routing strategies together with partial reconfiguration of stem cells we achieved a number of benefits, including a fast fault recovery and avoidance of using complicated routing strategies, as compared to recently developed fault-tolerant FPGA architectures. Hasan Baig, Jeong-A Lee |
FPT | 2 |
| 2005 | High performance asynchronous on-chip bus with multiple issue and out-of-order/in-order completionabstractIn this paper, we propose a high performance asynchronous on-chip bus with multiple issue and in-order/out-of-order completion for a Globally Asynchronous Locally Synchronous (GALS) design. The proposed bus implementation can be characterized with distributed and modularized control units based on a layered architecture to support multiple issue and in-order/out-of-order completion. Simulation results reveal that throughputs of asynchronous on-chip buses with multiple issue and in-order/out-of-order completion increases by 31.3% / 34.3%, while power consumption overhead is only 6.76% / 3.98% respectively, compared to a simple asynchronous on-chip bus with only a single issue feature. Eun-Gu Jung, Jeong-Gun Lee, Sanghoon Kwak, Kyoung-Son Jhang, Jeong-A Lee, Dong-Soo Har |
ACM Great Lakes Symposium on VLSI | 5 |
| 1994 | VLSI implementation of CORDIC angle unitsabstractWe design angle units using Lager tools both by a conventional CORDIC algorithm and a fast algorithm called Constant-Factor Redundant CORDIC (CFR-CORDIC) and show that the CFR-CORDIC occupies more than twice the area of a conventional CORDIC but offers good speed-up. We discuss VLSI design issues using Lager such as system partitioning and grouping, floor planning, width and height manipulation to obtain the smallest geometry, and the limitation of the standard cell design approach. In addition, a bit encoding scheme of a signed digit number representation, which simplifies the implementation of the negation is presented.> Jeong-A Lee, Mubashir Ahmad |
Great Lakes Symposium on VLSI | 1 |
| 1992 | Constant-Factor Redundant CORDIC for Angle Calculation and RotationabstractA constant-factor redundant-CORDIC (CFR-CORDIC) scheme, where the scale factor is kept constant while an angle for plane rotations is computed, is developed. The direction of rotation is determined from an estimate of the sign, and convergence is assured by suitably placed correcting iterations. The number of iterations in the CORDIC rotation unit is reduced by about 25% by expressing the direction of the rotation in radix-2 and radix-4, and conversion to conventional representation is done on the fly. The performance of CFR-CORDIC is estimated and compared with that of previously proposed schemes. It is found to provide an execution time similar to that of redundant CORDIC with a variable scaling factor, with a significant saving in area.> Jeong-A Lee, Tomás Lang |
IEEE Trans. Computers | 1 |
| 1991 | SVD by constant-factor-redundant-CORDICabstractA constant-factor-redundant-CORDIC (CFR-CORDIC) scheme is developed where the scale factor is forced to be constant while computing angles for SVD (singular value decomposition). Based on the scheme, a fixed-point implementation of SVD is presented with the following additional features: (1) the final scaling operation is done by shifting; (2) the number of iterations in the CORDIC rotation unit is reduced by about 25% by expressing the direction of the rotation in radix-2 and radix-4; and (3) the conventional number representation of rotated output is obtained on-the-fly, not from a carry-propagate adder. The authors compare this scheme with previously proposed ones and show that it provides an execution time similar to that of redundant CORDIC with variable scaling factor, with significant saving in area.> Jeong-A Lee, Tomás Lang |
IEEE Symposium on Computer Arithmetic | 1 |
| 1991 | Discrete Fourier transform processors using CORDICabstractThe author presents an analysis of the cost-effectiveness of discrete Fourier transform processors, based on CORDIC modules such as the bit-serial, parallel with non-redundant and redundant arithmetic, and pipelined. The performance of each processor is analyzed with respect to the time to process one frequency output and the number of modules required. It is shown that the CORDIC-based DFT processor is a prospective solution in VLSI to be used for a wide range of input bit rate.> Jeong-A Lee, Kiseon Kim |
Great Lakes Symposium on VLSI | 1 |
| 1991 | A Comparison of Redundant CORDIC Rotation EnginesabstractThe CMOS implementation of two high performance rotation processors using redundant CORDIC are reviewed and compared. One of the designs uses a variable scaling factor while the other is with constant scaling. The latter also incorporates some radix-4 CORDIC stages. Characteristics for 1.2 mu m CMOS implementations are given.> John A. Harding, Tomás Lang, Jeong-A Lee |
ICCD | 3 |