VLDB 2026 Research / reviewers in the wild / expert
Jae-Jin Lee
dblp:02/7010
· DBLP profile ↗
28ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0003-3260-1620ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 4 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Computer networks · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FiCABU: A Fisher-Based, Context-Adaptive Machine Unlearning Processor for Edge AIabstractMachine unlearning, driven by privacy regulations and the "right to be forgotten," is increasingly needed at the edge, yet server-centric or retraining-heavy methods are impractical under tight computation and energy budgets. We present FiCABU (Fisher-based Context-Adaptive Balanced Unlearning), a SW–HW co-design that brings unlearning to edge AI processors. FiCABU combines (i) Context-Adaptive Unlearning, which begins edits from back-end layers and halts once the target forgetting is reached, with (ii) Balanced Dampening, which scales dampening strength by depth to preserve retain accuracy. These methods are realized in a full RTL design of a RISC-V edge AI processor that integrates two lightweight IPs for Fisher estimation and dampening into a GEMM-centric streaming pipeline, validated on an FPGA prototype and synthesized in 45 nm for power analysis. Across CIFAR-20 and PinsFaceRecognition with ResNet-18 and ViT, FiCABU achieves random-guess forget accuracy while matching the retraining-free Selective Synaptic Dampening (SSD) baseline on retain accuracy, reducing computation by up to 87.52% (ResNet-18) and 71.03% (ViT). On the INT8 hardware prototype, FiCABU further improves retain preservation and reduces energy to 6.48% (CIFAR-20) and 0.13% (PinsFaceRecognition) of the SSD baseline. In sum, FiCABU demonstrates that back-end–first, depth-aware unlearning can be made both practical and efficient for resource-constrained edge AI devices. Eun-Su Cho, Jeongmin Jin, Jae-Jin Lee |
DATE | 4 |
| 2026 | Dynamic Neural Thresholding on Mixed-Signal Neuromorphic Processors Enabled by Integrated Learning and Hardware DesignabstractSpiking neural networks (SNNs) can improve inference accuracy through joint optimization of synaptic weights and neuronal thresholds. However, mixed-signal neuromorphic processors, which are designed for energy efficiency using analog circuits, face practical limitations. In particular, digital to analog converters (DACs) often lack sufficient resolution to represent the large threshold values required by joint optimization. To address this issue, we propose a mixed-signal neuromorphic processor architecture that shifts threshold control to digital logic. This approach removes the need for high-resolution DACs and allows dynamic threshold adjustment without modifying the analog neural core. We also propose a learning method tailored to this architecture. We evaluate the proposed design on five image classification benchmarks, measuring accuracy, latency, and energy consumption. The results show that our architecture consistently improves accuracy across benchmarks while incurring only minimal latency and energy overhead. This demonstrates that the proven benefits of joint weight and threshold learning can be realized in energy efficient analog hardware. Kyuseung Han, Kwang-Il Oh, Sukho Lee, Hyeonguk Jang, Jae-Jin Lee, Sooyoung Jang |
DATE | 5 |
| 2026 | LoRA-Edge: Tensor-Train-Assisted LoRA for Practical CNN Fine-Tuning on Edge DevicesabstractOn-device fine-tuning of CNNs is essential to with-stand domain shift in edge applications such as Human Activity Recognition (HAR), yet full fine-tuning is infeasible under strict memory, compute, and energy budgets. We present LoRA-Edge, a parameter-efficient fine-tuning (PEFT) method that builds on Low-Rank Adaptation (LoRA) with tensor-train assistance. LoRA-Edge (i) applies Tensor-Train Singular Value Decomposition (TT-SVD) to pre-trained convolutional layers, (ii) selectively updates only the output-side core with zero-initialization to keep the auxiliary path inactive at the start, and (iii) fuses the update back into dense kernels, leaving inference cost unchanged. This design preserves convolutional structure and reduces the number of trainable parameters by up to two orders of magnitude compared to full fine-tuning. Across diverse HAR datasets and CNN backbones, LoRA-Edge achieves accuracy within 4.7% of full fine-tuning while updating at most 1.49% of parameters, consistently outperforming prior parameter-efficient baselines under similar budgets. On a Jetson Orin Nano, TT-SVD initialization and selective-core training yield 1.4–3.8× faster convergence to target F1. LoRA-Edge thus makes structure-aligned, parameter-efficient on-device CNN adaptation practical for edge platforms. Hyunseok Kwak, Kyeongwon Lee, Jae-Jin Lee |
DATE | 3 |
| 2025 | NPX: Automating Neuromorphic Processor Design from Spike-Based Learning to FPGA PrototypingabstractNeuromorphic Processor eXpress (NPX) is a framework designed to facilitate the development of lightweight neuromorphic processors. To evaluate its efficacy, we conducted a case study involving the FPGA-based implementation of a traffic sign recognition system using an NPX-generated neuromorphic processor. The prototype integrates camera input and OLED output, successfully demonstrating full functionality. This case study confirms that NPX substantially streamlines the design and deployment of efficient neuromorphic processors for embedded artificial intelligence applications. Kyuseung Han, Hyeonguk Jang, Sukho Lee, Sung-Eun Kim, Kyudong Hwang, Jae-Jin Lee |
FPL | 6 |
| 2025 | NeuGEMM: A Reordering-Free Unified GEMM-Conv2D Accelerator for Lightweight Neuromorphic ProcessorsabstractNeuromorphic inference applications primarily rely on general matrix-matrix multiplication (GEMM) and twodimensional convolution (Conv2D) operations. When conventional artificial neural network (ANN) acceleration techniques are employed, these computations often necessitate extensive data reordering, which imposes significant overheads, especially in lightweight embedded systems with limited CPU and memory bandwidth. To address this challenge, we propose a unified accelerator architecture executing GEMM and Conv2D operations without data reordering. The accelerator is co-designed with a neuromorphic software framework tailored for lightweight embedded systems. To validate effectiveness, we implement a neuromorphic processor incorporating the proposed accelerator on an FPGA. Evaluation results across four representative neuromorphic applications demonstrate that the proposed design reduces execution time and energy consumption by 69% and 89%, respectively, compared to conventional ANN accelerators. Hyeonguk Jang, Sukho Lee, Jae-Jin Lee, Kyuseung Han |
FPL | 3 |
| 2025 | ASAP-FE: Energy-Efficient Feature Extraction Enabling Multi-Channel Keyword Spotting on Edge ProcessorsabstractMulti-channel keyword spotting (KWS) has become crucial for voice-based applications in edge environments. However, its substantial computational and energy requirements pose significant challenges. We introduce ASAP-FE (Agile Sparsity-Aware Parallelized-Feature Extractor), a hardware-oriented front-end designed to address these challenges. Our framework incorporates three key innovations: (1) Half-overlapped Infinite Impulse Response (IIR) Framing: This reduces redundant data by approximately 25% while maintaining essential phoneme transition cues. (2) Sparsity-aware Data Reduction: We exploit frame-level sparsity to achieve an additional 50% data reduction by combining frame skipping with stride-based filtering. (3) Dynamic Parallel Processing: We introduce a parameterizable filter cluster and a priority-based scheduling algorithm that allows parallel execution of IIR filtering tasks, reducing latency and optimizing energy efficiency. ASAP-FE is implemented with various filter cluster sizes on edge processors, with functionality verified on FPGA prototypes and designs synthesized at 45 nm. Experimental results using TC-ResNet8, DS-CNN, and KWT-1 demonstrate that ASAP-FE reduces the average workload by 62.73% while supporting real-time processing for up to 32 channels. Compared to a conventional fully overlapped baseline, ASAP-FE achieves less than a 1% accuracy drop (e.g., 96.22% vs. 97.13% for DS-CNN), which is well within acceptable limits for edge AI. By adjusting the number of filter modules, our design optimizes the trade-off between performance and energy, with 15 parallel filters providing optimal performance for up to 25 channels. Overall, ASAP-FE offers a practical and efficient solution for multi-channel KWS on energy-constrained edge devices. Jina Park, Jae-Jin Lee, Massoud Pedram |
ISLPED | 3 |
| 2025 | Electromagnetic Interference-Robust Fingerprint Spoof Detection Based on Finger Channel ResponseabstractThis study presents a practically deployable anti-spoofing system for fingerprint biometrics based on electric finger channel response (FCR) to detect falsification attempts using artificial fake fingerprints. Meanwhile, electric signals passed through the human body are particularly vulnerable to distortion from electromagnetic interference (EMI) induced into the body channel by the body antenna effects. The proposed EMI-robust fake fingerprint detection (FFD) employed a neural fully-connected filter trained on the proposed detection features extracted from the amplitude variation and time dispersion characteristics of FCR. A valid FCR dataset for feature analysis was acquired using custom-developed devices in a designated experimental setup involving 20 subjects. Following compatibility verification of the FCR-based FFD with a capacitive-sensing fingerprint scanner in the implemented prototype, the performance evaluation under onsite EMI conditions—primarily modeled as wide-band pulsed-radiated emission and narrow-band sinusoidal wave interferences—showed that the proposed FFD achieved false rejection and acceptance rates of 0.9$\%$and 0.5$\%$, respectively. Kwang-Il Oh, Seong-Eun Kim, Jae-Jin Lee, Sungeun Kim, Kyungjin Byun, Wangrok Oh |
IEEE Trans. Ind. Informatics | 4 |
| 2025 | BCRNet-SNN: Body Channel Response-Aware Spiking Neural Network for User RecognitionabstractSpiking neural networks (SNNs) have been recently highlighted as an attractive approach for implementing artificial intelligence models in resource-constrained edge devices for various industrial applications. These SNNs leverage low-power and wide dynamic range processing through biologically inspired event-driven operations in a massively parallel manner. In this respect, we propose BCRNet-SNN, an SNN model designed to utilize an electric body channel response (BCR) as a biometric feature for user recognition, where each element in the BCR dataset for 15 subjects is a 1-D vector comprising 380 feature points from the measured envelope by applying chirp signals to the body. The network parameters for implementing the posttrained BCRNet-SNN are inherited from the proposed convolutional neural network (CNN) that extensively extracts BCR-based biometric features (BCRNet), followed by applying knowledge distillation-based network lightening process on BCRNet. The performance evaluation results compared to ResNet18 and ResNet6 for the BCR dataset show that BCRNet achieves a greater than 2$\%$and 1.4$\%$improvement, respectively, in the average classification accuracy, while significantly reducing the number of network parameters to less than 1$\%$. The student network distilled from the teacher BCRNet (BCRNet-S) requires only 5.1$\%$network parameters than BCRNet, which can be eligible for implementation in BCRNet-SNN. The proposed structure of BCRNet-SNN to effectively accommodate the CNN-to-SNN-converted parameters from BCRNet-S can achieve up to a maximum accuracy of 98.11$\%$, without observable performance degradation compared to BCRNet. ChanWoo Shin, Jongseok Lee, Jae-Jin Lee, Dong-Gyu Sim, Seong-Eun Kim |
IEEE Trans. Ind. Informatics | 4 |
| 2024 | STARC: Crafting Low-Power Mixed-Signal Neuromorphic Processors by Bridging SNN Frameworks and Analog DesignsabstractDeveloping low-power neuromorphic processors capable of inferring outcomes from SNN Frameworks presents significant challenges, largely due to the gap between frameworks and analog circuit-based SNNs. This paper analyzes the root of this gap as stemming from over/underflow issues and proposes mixed-signal neurons as a solution, further developing a neural core composed of these neurons. In the development of the neural core, we incorporate a design methodology for application-specific neural core optimization. We advance to develop a neural engine as an independent IP, ultimately introducing the snnTorch Architecture (STARC), an integrated mixed-signal neuromorphic processor architecture. The STARC processor, developed as a prototype, demonstrates operational correctness and exceptional low-power performance. Kyuseung Han, Hyunseok Kwak, Kwang-Il Oh, Sukho Lee, Hyeonguk Jang, Jae-Jin Lee |
ISLPED | 6 |
| 2024 | Anti-Spoofing for Fingerprint Recognition Using Electric Body Pulse ResponseabstractThis study presents a highly reliable approach to prevent fingerprint spoofing attacks based on electric body pulse responses (BPRs) in personal Internet of Things (IoT) gadgets. Real fingerprint pulse response (RFPR) and fake fingerprint pulse response (FFPR) data were collected from ten subjects for four weeks. The FFPR was obtained by wearing a fake fingerprint made of artificial substances, such as conductive silicone, over the finger. We analyzed different patterns of FFPR compared to RFPR using an electric circuit model of the proposed fingerprint anti-spoofing system based on BPRs. Simple features comprising ten, five, or three datapoints were selected by the minimum redundancy maximum relevance (MRMR) algorithm and led to reduction in processing complexity. We also validated its robustness to sampling offset errors caused by practical sampling operations in devices based on the evaluation of classification accuracy using machine learning algorithms, such as${k}$-nearest neighbor (KNN) and support vector machine (SVM). Finally, the effectiveness of the selected feature was evaluated using unsupervised anomaly detection algorithms, such as principal component analysis (PCA), one-class SVM (OC-SVM), and variational autoencoder (VAE), in a practical scenario with sampling offset errors in the training and test data. The VAE outperformed PCA and OC-SVM by achieving a detection accuracy of 99.76% using raw data under 100 datapoints and 97.60% with reduced features having only five datapoints, regardless of sampling offset errors. Therefore, the proposed anomaly detection system based on EPRs can provide promising fingerprint spoof detection in IoT devices with limited computing resources. Kwang-Il Oh, Jae-Jin Lee, Sungeun Kim, Wangrok Oh, Seong-Eun Kim |
IEEE Internet Things J. | 3 |
| 2024 | Day-Night architecture: Development of an ultra-low power RISC-V processor for wearable anomaly detectionabstractIn healthcare, anomaly detection has emerged as a central application. This study presents an ultra-low power processor tailored for wearable devices dedicated to anomaly detection. Introducing a unique Day-Night architecture, the processor is bifurcated into two distinct segments: The Day segment and the Night segment, both of which function autonomously. The Day segment, catering to generic wearable applications, is designed to remain largely inactive, awakening only for specific tasks. This approach leads to considerable power savings by incorporating the Main-CPU and system interconnect, both major power consumers. Conversely, the Night segment is dedicated to real-time anomaly detection using sensor data analytics. It comprises a Sub-CPU and a minimal set of IPs, operating continuously but with minimized power consumption. To further enhance this architecture, the paper presents an ultra-lightweight RISC-V core, All-Night core, specialized for anomaly detection applications, replacing the traditional Sub-CPU. To validate the Day-Night architecture, we developed a prototype processor and implemented it on an FPGA board. An anomaly detection application, optimized for this prototype, was also developed to showcase its functional prowess. Finally, when we synthesized the processor prototype using 45 nm process technology, it affirmed our assertion of achieving an energy reduction of up to 57%. Eunjin Choi, Jina Park, Kyeongwon Lee, Jae-Jin Lee, Kyuseung Han |
J. Syst. Archit. | 4 |
| 2024 | Designing Low-Power RISC-V Multicore Processors With a Shared Lightweight Floating Point Unit for IoT EndnodesabstractThe increasing interest in RISC-V from both academia and industry has motivated the development and release of a number of free, open-source cores based on the RISC-V instruction set architecture. Specifically, the use of lightweight RISC-V cores in processors tailored for IoT endnode devices is on the rise. As the range and complexity of these applications grow, there is an increasing demand for multicore processors that can handle floating-point operations. This poses a significant challenge because most lightweight RISC-V cores are integer cores lacking a floating-point unit (FPU). This limitation makes it difficult to design processors optimized for applications that require floating-point operations concurrently with integer operations. While it is inefficient to have a dedicated FPU per core in a multicore processor (because it would give rise to unnecessary power consumption), it is crucial to find a solution that balances performance and energy efficiency. To address this challenge, we propose to utilize an external lightweight FPU that can be added to any RISC-V integer core, along with a low-power multicore architecture that shares the said FPU. We have applied this concept to design a RISC-V processor that integrates these technologies, implemented it on an FPGA device, and completed the fabrication of a System-on-Chip for functional verification. Our experiments, which involved testing various applications on different processor prototypes, demonstrated significant energy savings of up to 79.6% in a quad-core processor prototype, highlighting the potential energy efficiency of our proposed technology. Jina Park, Kyuseung Han, Eunjin Choi, Jae-Jin Lee, Kyeongwon Lee, Massoud Pedram |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2023 | Developing an Ultra-low Power RISC-V Processor for Anomaly DetectionabstractThis paper aims to develop an ultra-low power processor for wearable devices for anomaly detection. To this end, this paper proposes a processor architecture that divides the architecture into a part for general applications running on wearable devices (day part) and a part that performs anomaly detection by analyzing sensor data (night parts), and each part operates completely independently. This day-night architecture allows the day part, which contains the power-hungry main-CPU and system interconnect, to be turned off most of the time except for intermittent work, and the night part, which consists only of the sub-CPU and minimal IPs, can run all the time with low power. By developing a processor based on the proposed processor architecture, the design verification of the proposed technology and the superiority of power saving are demonstrated. Jina Park, Eunjin Choi, Kyungwon Lee, Jae-Jin Lee, Kyuseung Han |
DATE | 4 |
| 2023 | Florian: Developing a Low-Power RISC-V Multicore Processor with a Shared Lightweight FPUabstractAs applications running on lightweight RISC-V processors become increasingly diverse and complex, the need for multicore processors supporting floating-point units (FPUs) is riseing, making processor designs using existing open-source RISC-V cores challenging. With the exception of a very few, most open lightweight RISC-V cores are integer cores without FPUs, which greatly reduces the design exploration space, making it impossible to design a processor optimized for each application. For example, most of these applications mainly perform integer operations, but occasionally perform floating-point operations. For them, a multicore processor with FPU per core is overkill and wastes power, which is a critical problem for processors where low-power design is paramount. To address the problem, we propose an external lightweight FPU that can be attached to any RISC-V integer core and a low-power multicore architecture using the designed FPU. For verification, we designed a RISC-V processor that implements all the proposed technologies, prototyped it on an FPGA device, and finally fabricated it as a System-on-Chip. Through experiments, it was confirmed that the proposed technology can cut energy consumption energy by up to 23%. Jina Park, Kyuseung Han, Eunjin Choi, Sukho Lee, Jae-Jin Lee, Massoud Pedram |
ISLPED | 5 |
| 2021 | Developing TEI-Aware Ultralow-Power SoC Platforms for IoT End NodesabstractRanging from circuit-level characterization to designing a platform architecture, developing a design automation tool, and fabricating a System on Chip (SoC), this article deals with the entire development process for ultralow-power (ULP) SoCs for Internet-of-Things (IoT) end nodes. More precisely, this article first focuses on the unique characteristics of the ULP circuits, the temperature effect inversion (TEI), i.e., the delay of the ULP circuits decreases with increasing temperature. Existing TEI-aware low-power (TEI-LP) techniques have incredible potential to further reduce the power consumption of conventional ULP SoCs, but there is a critical limitation to be widely adopted in real SoCs. To address this limitation and realize the ULP SoCs that can fully benefit from the TEI-LP techniques, this article proposes a new TEI-inspired SoC platform (TIP) architecture. On top of that, taking into account that the highly complex, time consuming, and labor-intensive development process of these ULP SoCs may hinder their widespread use for IoT end nodes, this article presents a new electronic design automation tool to accelerate ULP SoC development, RISC-V express (RVX). Finally, by using the RVX, this article introduces a TIP prototyping chip fabricated in 28-nm FD-SOI technology. This chip demonstrates that power savings of up to 35% can be achieved by lowering the supply voltage from 0.54 to 0.48 V at 25 °C and 0.44 V at 80 °C while continuing to operate at a target 50-MHz clock frequency. Kyuseung Han, Sukho Lee, Kwang-Il Oh, Younghwan Bae, Hyeonguk Jang, Jae-Jin Lee, Massoud Pedram |
IEEE Internet Things J. | 6 |
| 2021 | Energy efficient spiking neural network processing using approximate arithmetic units and variable precision weights
Yi Wang 0064, Hao Zhang 0041, Kwang-Il Oh, Jae-Jin Lee, Seok-Bum Ko |
J. Parallel Distributed Comput. | 4 |
| 2019 | TIP: A Temperature Effect Inversion-Aware Ultra-Low Power System-on-Chip PlatformabstractResearchers have been trying to exploit the temperature effect inversion (TEI) phenomenon to improve energy efficiency of system-on-chip (SoC) designs without sacrificing its performance. However, TEI-aware low power methods have a critical limitation in that they can only be applied to components within the SoC that do not contain long (global) wires. This is because wire delays continue to increase with rising temperatures irrespective of the operating supply voltage level, which tends to cancel out positive effects of the TEI phenomenon in SoCs. To tackle this limitation and thoroughly utilize the TEI-aware methods, this paper presents new TEI-inspired SoC platform (called TIP), which relies on network-on-chip architecture (called μNoC) to realize system interconnects. The μNoC successfully reduces the total number and length of global wires. By fabricating a TIP prototyping chip in Samsung 28nm FD-SOI technology, we verify the effectiveness of TIP. Extensive post-fabrication measurements demonstrate that the chip while continuing to operate at a target 50MHz clock frequency can lower its supply voltage from 0.54V to 0.48V at 25°C and to 0.44V at 80°C, which results in up to 35% power saving. Kyuseung Han, Sukho Lee, Jae-Jin Lee, Massoud Pedram |
ISLPED | 3 |
| 2019 | TEI-ULP: Exploiting Body Biasing to Improve the TEI-Aware Ultralow Power MethodsabstractTemperature effect inversion (TEI) phenomenon in ultralow power (ULP) very large scale integration circuits has been identified as an important effect by both academia and industry. Although a number of ULP methods that attempt to exploit the TEI phenomenon have been proposed, the small size of the design exploration space when applying these methods to ULP circuits hinders them from achieving their full potential. This is mainly due to the limited granularity of the supply voltage level control. Starting with an intuition that the body biasing (BB) technique is a key to overcome this limitation, this paper exploits the BB technique along with the TEI-aware voltage scaling (TEI-VS) method and TEI-aware frequency scaling (TEI-FS) method, so as to substantially increase the design spaces of these methods. Techniques for optimally combining the BB technique with TEI-VS and TEI-FS are introduced. Simulation results with the latest commercial CMOS process technologies for ULP designs demonstrate the effectiveness of the proposed methodology. Jae-Jin Lee, Kyuseung Han, Joongheon Kim, Massoud Pedram |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | TEI-NoC: Optimizing Ultralow Power NoCs Exploiting the Temperature Effect InversionabstractThe era of the Internet of Things (IoT) is upon us. In this era, minimizing power consumption becomes a primary concern of system-on-chip designers. Ultralow power (ULP) very large-scale integration circuits have been receiving considerable interest from both academia and industry as the best-suited techniques for IoT devices, which can take full advantage of power-saving that voltage scaling potentially achieves. Consequently, research on ULP designs has begun to yield tangible outcomes, namely ULP circuits. However, little attention has been paid to ULP network-on-chip (NoC), although the NoC is an essential of the ULP chips, and its power consumption accounts for a significant portion of the total power. This paper focuses on ULP NoCs, and presents a new power management method that exploits delay versus temperature characteristics of ULP circuits. Recent studies on ULP circuits show that delay versus temperature characteristics are fundamentally different from normal circuits, i.e., the delay of the ULP circuits implemented in state-of-the-art bulk CMOS operating at low supply voltages or in FinFET technologies decreases with increasing temperature, a phenomenon known as the temperature effect inversion (TEI). Starting with an intuition that at a certain temperature point, power savings without performance penalty can be achieved by increasing the router frequency to create the opportunity to turn off some routers in ULP NoCs, or by decreasing the NoC supply voltage level, an optimization method is presented to maximize the power savings with minor performance penalty. To validate the proposed method, a concrete ULP NoC simulator, TEI-Noxim, has been developed. Experimental results demonstrate that TEI-aware NoC achieves an average of 36.0% power reduction over 21 applications. Kyuseung Han, Jae-Jin Lee, Jinho Lee 0001, Massoud Pedram |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2015 | Virtual prototype based on Aldebarn CPU coreabstractThis paper proposes a virtual prototype based on the Aldebaran CPU core developed independently by ETRI. The virtual prototype provides instruction and function profiling functionality for software optimization as well as standard integration emulation interface (SystemC, Verilog, Netlist, etc.) compatibility, and architecture performance analysis for efficient adoption into a system-level design environment. Jae-Jin Lee, Kyungjin Byun, Nak-Woong Eum |
VLSI-SoC | 1 |
| 2013 | Skew Compensation Technique for Source-Synchronous Parallel DRAM InterfaceabstractThe interpin skew among the data and the strobe signals of a source-synchronous parallel DRAM interface is compensated by a simple delay-locked loop, which reuses the circuitry of a normal input data path. With the interpin skew compensation, the printed circuit board traces of the data and the strobe signals are allowed to have unequal length. The prototype implemented in a 0.13- μm standard CMOS process shows that the interpin skew is reduced to be less than 26 ps for a 3.2-Gb/s/pin ×8 parallel interface. Jang-Woo Lee, Hong-Jung Kim, Chun-Seok Jeong, Jae-Jin Lee, Changsik Yoo |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2012 | Cost-effective TSV redundancy configuration
Jongpil Jung, Kyungsu Kang, Jae-Jin Lee, Youngjun Yoon, Chong-Min Kyung |
VLSI-SoC | 3 |
| 2010 | New Lookup Tables and Searching Algorithms for Fast H.264/AVC CAVLC DecodingabstractIn this paper, new codeword structures, tables, and searching methods for fast and efficientcoeff_token,total_zeros, andrun_beforedecoding are developed. This new achievement is mainly based on the fact that the context-adaptive variable length coding (CAVLC) decoding can be modeled as a finite state machine. In order to quantitatively evaluate the proposed method in terms of decoding speed and complexity, we define the iteration bound$\left({{1}\over {\mathtilde{\tau}}}\right)$and thecomplexity ratio$(CR)$. Using these gauge variables, we show that the new algorithms reduce${\mathtilde {\tau}}$to about one third andcomplexity ratioto 0.95. This means that the proposed techniques reduce the decoding time to about one third and memory access count by 90% compared to those of the conventional methods without implementation overheads. Multiple-symbol parallel decoding method forrun_beforesyntax element is proposed based on abit-positioningwith the critical path latency of only one multiplexer for the post-combination process. The proposed methods make it possible to implement a fast and efficient CAVLC decoding without losing video quality on any environments. Jae-Jin Lee, Seongmo Park |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Service Impact Analysis Framework Using Service Model for Integrated Service Resource Management of NGN Services
Seung Hee Han, Bom Soo Kim, Chan Kyou Hwang, Jae-Jin Lee |
APNOMS | 4 |
| 2008 | A 100MHz ASIP (application specific instruction processor) for CAVLC of H.264/AVC decoderabstractIn this paper, we implement the configurable processor for CAVLC function module of a H.264/AVC baseline profile decoder as the starting point to implement the H.264/AVC decoder system in a multiprocessor platform. The requirements of the implementations are the low-power processor and speed optimized algorithms tailored to the processor architecture. An arithmetic formula mapping method for fast CAVLC algorithms and a dual-issue VLIW processor architecture with custom instructions are proposed. The experiment results show that the synthesized processor has about 75 K gates and can carry out the decoding of 30-fps CIF (352x288 pixels) images around 120 Mega cycles. Jae-Jin Lee, MooKyoung Jeong, Nak-Woong Eum, Seongmo Park |
ISCAS | 2 |
| 2005 | Design of an application-specific PLD architectureabstractThis paper presents a new application-specific PLD architecture which adopts a bit-level super-systolic array for application-specific arithmetic operation such as MAC. The proposed design offers a significant alternative view on programmable logic device. The bit-level super-systolic array whose cells contain another systolic array is ideal for newly proposed PLD architecture in terms of area efficiency and clock speed as it limits the routing requirement in a PLD to local interconnections between Logic Units and to global interconnections between Logic Modules. The maximum clock cycle is limited only by one AND gate and one full adder. Jae-Jin Lee, Gi-Yong Song |
ASP-DAC | 1 |
| 2004 | Bit-level super-systolic array for FIR filter with a FPGA-based bit-serial semi-systolic multiplierabstractTo achieve higher degree of concurrency in a systolic array, it is desirable to make cells of a systolic array themselves a systolic array as well, leading to a structure called super-systolic array. This paper proposes a bit-level super-systolic FIR filter with a FPGA-based bit-serial semi-systolic multiplier. Compared to the word-level systolic FIR filter and corresponding super-systolic filter, the proposed design is very compact in that it needs only two 1-bit I/O ports in addition to significant improvement on hardware complexity. The input to the implementation proposed in this paper is a sequence of bit-string in contrast to the distributed arithmetic which assumes input of parallel bit-string. Also, the arrangement of the cells of a semi-systolic multiplier is tuned to fit to the structure of the FPGA. (This work was done as a part of Information & Communication fundamental Technology Research Program supported by Ministry of Information & Communication in republic of Korea. Jae-Jin Lee, Gi-Yong Song |
FPGA | 1 |
| 2003 | Implementation of the super-systolic array for convolutionabstractHigh-performance computation on a large array of cells has been an important feature of systolic array. To achieve even higher degree of concurrency, it is desirable to make cells of systolic array themselves systolic array as well. The architecture of systolic array with its cells consisting of another systolic array is to be called super-systolic array.In this paper we propose a scalable super-systolic array architecture which shows high-performance and can be adopted in the VLSI design including regular interconnection and functional primitives that are typical for a systolic architecture. Jae-Jin Lee, Gi-Yong Song |
ASP-DAC | 1 |