Ji-Hoon Kim 0003

dblp:51/4060-3 · also Jihoon Kim 0003 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
8since 2021 · last 2024
0000-0002-9809-1339ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 3 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2024 PA-2SBF: Pattern-Adaptive Two-Stage Bloom Filter for Run-Time Memory Diagnostic Data Compression in Automotive SoCs
abstract
In the realm of safety-critical Automotive System-on-Chip (SoC) design, memory functionality plays a crucial role in determining overall die yield due to its sizable footprint on the chip. Therefore, efficient methodologies for both post-manufacturing offline testing and real-time monitoring are essential to provide timely diagnostic feedback. This paper proses a real-time memory diagnostic data compression technique Pattern-Adaptive Two-Stage Bloom Filter (PA-2SBF) for automotive System-on-Chip (SoC) applications. PA-2SBF is designed to address the challenge of false positives incorporating frequently encountered failure patterns. Bloom filters, a probabilistic data structure renowned for their space-efficiency and quick approximate membership queries, are employed to expedite fault memory diagnosis information lookup and compression. Furthermore, failure patterns are considered to mitigate the false positive rate inherent in Bloom filters. The paper also presents a strategy for leveraging the compressed diagnostic information during run-time. Specifically, it exploits the quick lookup feature of Bloom filter to prohibit CPU access to defective memory regions, enhancing overall system reliability.
Hana Kim, Ji-Hoon Kim 0003
DATE4
2024 Dynamic Resource Management in Reconfigurable SoC for Multi-Tenancy Support
abstract
This study introduces a partially reconfigurable system-on-chip (SoC) platform leveraging dynamic resource management facilitated by a dynamic reconfigurable control processor (DRCP). By addressing the inherent reconfiguration time overheads of reconfigurable SoCs, the study demonstrates performance improvements through a runtime resource management strategy. The introduced management scheme effectively reduces the frequency of reconfigurations, thus lessening the associated overheads and increasing the operational efficiency of the SoC platform, which is designed to support on-chip multi-tenancy. Utilizing DRCP for dedicated resource management, the proposed SoC platform exhibited substantial reductions in reconfiguration times. When the partially reconfigurable SoC platform employed four partial regions (PRs), the reconfiguration counts decreased by 37.6%. Furthermore, upon extending the PRs to eight, there was a notable reduction in reconfiguration counts, achieving a decrease of 47.0%.
Sohyeon Kim, Injun Choi, Minkyu Je, Ji-Hoon Kim 0003
ISCAS4
2024 Algorithm-Hardware Co-Design for Wearable BCIs: An Evolution from Linear Algebra to Transformers
abstract
Recent advancements in brain-computer interface (BCI) technology for steady-state visual evoked potential (SSVEP)-based target identification have shifted from traditional linear algebra (LA) techniques to more sophisticated neural network (NN) approaches, driven by their increased accuracy and consistent performance across different subjects. However, adopting NN-based algorithms has introduced complexities in wearable BCI systems, mainly due to their extensive parameter sets that demand significant memory capacity. Moreover, the computational intensity of these models requires reevaluating hardware architectures. Additionally, the advent of Transformer-based models has further advanced the state of the art, providing even higher accuracy and reduced variability in cross-subject performance, placing greater demands on hardware resources. This paper provides an overview of recent algorithmic progress in SSVEP-based target identification. Also, it proposes considerations for the hardware architecture needed to efficiently support the computation of cutting-edge Transformer-based models in wearable BCIs from the perspective of algorithm-hardware co-design.
Wooseok Byun, Minkyu Je, Ji-Hoon Kim 0003
ISCAS4
2022 SQNR-based Layer-wise Mixed-Precision Schemes with Computational Complexity Consideration
abstract
Recently, AI acceleration is critical for hardware systems from mobile/edge devices to high-performance data servers. For on-device AI, there have been many studies on hardware numerical precision reduction for low hardware complexity considering limited hardware resources of mobile/edge devices [1–2]. However, if taken too far, aggressive quantization with low precision can degrade model accuracy which is the fundamental measure of deep learning quality. With the layer-wise mixed-precision networks where the different precision scheme is applied to each layer, we can reduce off-chip memory bandwidth requirements as well as hardware complexity while retaining model accuracy [3]. However, it takes a long time to determine the optimal precision scheme for each layer due to the repetitive trainings with the different precision configurations. Also, if only the accuracy of the target neural network is considered in the process of the mixed-precision determination, it is difficult to sufficiently reduce the computational complexity, which is the most important in mobile/edge devices [4].
Ha-Na Kim, Hyun Eun, Jung Hwan Choi, Ji-Hoon Kim 0003
ISCAS4
2022 Architectural Supports for Block Ciphers in a RISC CPU Core by Instruction Overloading
abstract
We propose a novel computer architectural concept of instruction overloading to support block ciphers. Instead of adding new instructions, we extend only the execution of some existing instructions. The proposed method allows a central processing unit core to execute different operations for the same instructions, depending on the address of the data, similar to operator overloading in object-oriented languages. We first present an extension for the AES algorithm, then we demonstrate its enhanced applicability with two further extensions supporting multiple block ciphers and hardware masking. The first extension for AES is also applicable to add/AND-rotate-XOR-based block ciphers such as SIMON. The AES and SIMON encryption speed, on this extended core, is at least doubled and is significantly less affected by memory latency. In addition, the AES encryption code requires only 18% of the memory of the previous software implementation. The second extension can further support various block ciphers defined over GF(28), and the SM4 encryption speed is increased by at least 182%. The third extension provides correlation power analysis (CPA) resistance with a 66.6% area overhead but almost no speed overhead, whereas a typical software anti-CPA AES implementation requires at least hundreds of times the execution time.
Piljoo Choi, Won Bae Kong, Ji-Hoon Kim 0003, Mun-Kyu Lee, Dong Kyue Kim
IEEE Trans. Computers3
2022 An Energy-Efficient Domain-Specific Reconfigurable Array Processor With Heterogeneous PEs for Wearable Brain-Computer Interface SoCs
abstract
Recently, there is increasing demand for energy-efficient signal processing in wearable visual-stimuli-based brain-computer interface (V-BCI) devices. For the better accuracy and the reduced latency of the V-BCI system, the target identification (TI) algorithm that analyzes brain signals is being advanced, and the importance of an energy-efficient accelerating chip that processes various linear algebra operations constituting the TI algorithms is growing. In this paper, we propose a domain-specific reconfigurable array processor (RAP) with a dynamically reconfigurable and scalable array including 5-heterogeneous processing elements (PEs) for the energy-efficient acceleration of basic linear algebra subprograms (BLAS) and matrix decompositions. The system-on-chip (SoC), including the proposed RAP, was fabricated in 130-nm CMOS technology with an area of 16.87-mm2 and measured at 1.0 V 90 MHz. The RAP achieved an information transfer rate (ITR) of 139.9-bits/min and a TI accuracy of 95.4% on a fabricated chip through an optimized TI algorithm and scalable array processing. In addition, the RAP has$16.8\times $higher TI energy efficiency than prior work and achieved an energy efficiency of 2144.2-bits/min/mW for information transfer processing rate with the proposed TI algorithm. The RAP supports a greater variety of linear algebra operations and data sizes with hardware reconfiguration than the prior accelerators.
Wooseok Byun, Minkyu Je, Ji-Hoon Kim 0003
IEEE Trans. Circuits Syst. I Regul. Pap.3
2022 A 46-nF/10-MΩ Range 114-aF/0.37-Ω Resolution Parasitic- and Temperature-Insensitive Reconfigurable RC-to-Digital Converter in 0.18-μm CMOS
abstract
This paper presents a 46 nF/10$\text{M}\Omega $-range, digital-intensive, reconfigurable RC-to-digital converter (R2CDC) that can readout multiple C and R sensors in a time-interleaved fashion. Ratio-metric conversion using swing-boosted period-modulation (SB-PM) front-end by the R2CDC results in 114 aFrms/$0.37 \Omega _{\text {rms}}$resolutions and a worst-case temperature-drift of 64.2 ppm/°C over −40 to 125°C. Femto-farad capacitances can be sensed with a relative code-deviation less than 0.16 % even when parasitics vary 30 times the baseline. Implemented in a$0.18 ~\mu \text{m}$standard CMOS process, the R2CDC consumes 140$\mu \text{A}$from a 1 V supply, occupying an active area of$0.175 ~\mu \text{m} ^{\mathrm{ 2}}$.
Arup K. George, Wooyoon Shim, Jaeha Kung 0001, Ji-Hoon Kim 0003, Minkyu Je, Junghyup Lee
IEEE Trans. Circuits Syst. I Regul. Pap.4
2021 ML-Based Humidity and Temperature Calibration System for Heterogeneous MOx Sensor Array in ppm-Level BTEX Monitoring
abstract
Recently, indoor air quality is an important issue for human health and high concentrations of toxic Volatile Organic Compounds (VOCs) gases such as BTEX (Benzene, Toluene, Ethylbenzene, and Xylene) are very harmful to our respiratory system and metabolism. To detect BTEX gases at indoors, Metal Oxide (MOx) sensors are widely used because of their low-cost and high sensitivity. MOx sensors are easily affected by temperature and humidity, hence it is difficult to detect BTEX gases accurately without additional calibration process. In this paper, we present the calibration system for heterogeneous MOx sensor array where machine learning (ML)-based techniques, Linear Regression (LR), Non-Linear Curve Fitting (NLCF), and Artificial Neural Network (ANN), are exploited to reduce the impact of temperature and humidity. For the performance evaluation, we have setup the gas concentration measurement system and recorded the sensor outputs from Temperature-Cycled Operation (TCO) responses of five heterogenous MOx sensors. The proposed calibration system with ANN-based calibration system shows the reduction of gas sensors variation due to temperature and humidity 73% on average, and presents maximum 92% reduction for benzene, 75% for toluene, 83% for ethylbenzene, and 91% for xylene gases, respectively.
Hoyong Sung, Sohyeon Kim, Minkyu Je, Ji-Hoon Kim 0003
ISCAS5
2012 Design of TETRA 2 turbo decoder with minimum memory hardware interleaver
abstract
Terrestrial Trunked Radio (TETRA) is a digital trunked mobile radio standard developed by the European Telecommunications Standards Institute (ETSI). Especially, TETRA 2 supports different channel bandwidths and modulation techniques, and turbo code has been adopted as a channel code with powerful error correcting capability. In this paper, to lower the implementation cost of Takeshita-Costello interleaver adopted in TETRA 2, we present the simple interleaved address generation scheme which requires minimum memory size compared to the conventional memory-based interleaver. Also, overall TETRA 2 turbo decoder architecture including the proposed interleaver is described with its implementation results
Ji-Hoon Kim 0003
ISCAS1
2008 Duo-binary circular turbo decoder based on border metric encoding for WiMAX
abstract
This paper presents a duo-binary circular turbo decoder based on border metric encoding. With the proposed method, the memory size for branch memory is reduced by half and the dummy calculation is removed at the cost of the small-sized memory which holds the encoded border metrics. Based on the proposed SISO decoder and the dedicated hardware interleaver, a duo-binary circular turbo decoder is designed for the WiMAX standard using a 0.13 mum CMOS process, which can support 24.26 Mbps at 200 MHz.
Ji-Hoon Kim 0003, In-Cheol Park
ASP-DAC1
2007 Twiddle factor transformation for pipelined FFT processing
abstract
This paper presents a novel transformation technique that can derive various fast Fourier transform (FFT) in a unified paradigm. The proposed algorithm is to find a common twiddle factor at the input side of a butterfly and migrate it to the output side. Starting from the radix-2 FFT algorithm, the proposed common factor migration technique can generate most of previous FFT algorithms without using mathematical manipulation. In addition, we propose new FFT algorithms derived by applying the proposed twiddle factor moving technique, which reduce the number of twiddle factors significantly compared with the previous algorithms being widely used for pipelined FFT processing.
In-Cheol Park, WonHee Son, Ji-Hoon Kim 0003
ICCD3
2007 Energy-Efficient Double-Binary Tail-Biting Turbo Decoder Based on Border Metric Encoding
abstract
This paper presents an energy-efficient turbo decoder based on border metric encoding, which is especially suitable for non-binary circular turbo codes. In the proposed method, the size of the branch memory is reduced to half and the dummy calculation is removed at the cost of a small-sized memory that holds encoded border metrics. Due to the small size and infrequent access to the border memory, power consumption for soft-input soft-output (SISO) decoding is reduced by 26.0%. Based on the proposed SISO decoder and the dedicated hardware interleaver, a double-binary tail-biting turbo decoder is designed for WiMAX standard using a 0.18 μm CMOS process and it can support 12.14Mbps at operating frequency of 100MHz.
Ji-Hoon Kim 0003, In-Cheol Park
ISCAS1
2006 Low-power hybrid turbo decoding based on reverse calculation
abstract
As turbo decoding is a highly memory-intensive algorithm consuming large power, a major issue to be solved in practical implementation is to reduce power consumption. This paper presents an efficient reverse calculation method to lower the power consumption by reducing the number of memory accesses required in turbo decoding. The reverse calculation method is proposed for the max-log-MAP algorithm, and it is combined with a scaling technique to achieve a new decoding algorithm, called hybrid log-MAP, that results in a similar BER performance to the log-MAP algorithm. For the W-CDMA standard, experimental results show that 80% of memory accesses are reduced through the proposed reverse calculation method. A hybrid log-MAP turbo decoder based on the proposed reverse calculation reduces power consumption and memory size by 34.4% and 39.2%, respectively.
Hye-Mi Choi, Ji-Hoon Kim 0003, In-Cheol Park
ISCAS2