EDBT 2026 Demo / reviewers in the wild / expert
Gain Kim
dblp:158/8085
· DBLP profile ↗
16ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0002-3680-8816ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fast and Accurate SystemVerilog Framework for Mixed-Signal Modeling of PAM-4 Wireline TransceiversabstractThis paper presents a SystemVerilog-based modeling and simulation methodology for a 4-level pulse-amplitude-modulation (PAM-4) transceiver. The framework is developed to address verification challenges in complex analog-to-digital converter (ADC)-based mixed-signal systems, where conventional circuit-level simulations suffer from prohibitively long execution time. Two key modeling techniques to significantly improve simulation efficiency are introduced in this work. For the time-interleaved ADC (TI-ADC), a counter-based equivalent circuit model is employed to capture dominant rank-1 mismatch errors while substantially reducing modeling complexity. In addition, for digital clock and data recovery (CDR), a statistical bit-error-rate (BER) evaluation approach based on XMODEL primitives is adopted as an alternative to conventional time-domain analog-mixed-signal (AMS) simulations. This modeling framework enables accurate verification of critical transceiver performance metrics, such as jitter tolerance (JTOL), within a 30-minute simulation window. Compared to conventional timedomain simulations, which require more than 12 hours to perform, the proposed approach achieves approximately a 24× improvement in simulation speed. Seoyoung Jang, Donggeon Kim, Korkut Kaan Tokgoz, Gain Kim |
ISCAS | 6 |
| 2026 | Fixed-Point Implementation Analysis of MLSE Receiver DSP for High-Speed Wireline Transceivers
Dohyeon Kwon, Donggeon Kim, Yusuf Leblebici, Gain Kim |
ISCAS | 5 |
| 2025 | A Spectral-Efficient Low-Power NRZ/PAM-4 Dual-Mode Wireline Transmitter for Multidrop InterfacesabstractThis paper presents a reconfigurable and energy-efficient digital spectrum shaping signaling (DSSS) for multidrop interfaces, where the output spectrum of the transmitted data is shaped using the 2-times repetitive block transmission to avoid frequency notches in the multidrop channel, thereby achieving a data rate up to 4× the first channel notch frequency. In conventional wireline transceivers (TRX), compensating for frequency notches requires a large number of decision feedback equalizer (DFE) taps at the receiver, resulting in significant area and power overhead. In contrast, the proposed DSSS architecture supports spectrum-efficient reconfigurable dual-mode NRZ/PAM-4, reducing required equalization efforts and improving energy efficiency. The proposed scheme and its transmitter (TX) were first validated through event-driven behavioral simulations using XMODEL and verified with equipment-based measurements. Post-layout simulation results in 28 nm CMOS process demonstrated 4 Gb/s data rate communicating over a channel having its first > 30 dB notch at 1 GHz, with a 235 mV vertical eye opening and a TX energy efficiency of 0.39 pJ/b at 0.8 V supply. Donggeon Kim, Kiarash Gharibdoust, Armin Tajalli, Kyungtae Lee, Gain Kim |
ISLPED | 5 |
| 2025 | An On-Chip Low-Cost Averaging Digital Sampling Scope for 80-GS/s Measurement of Wireline Pulse ResponsesabstractDetermining a channel’s characteristics is a fundamental step for designing a high-speed link system. By identifying the properties of the channel, designers can gain insights into how to transmit a signal with low distortion and optimize a transceiver’s architecture. As the channel’s characteristics can be identified by analyzing its single-bit pulse response (PR), obtaining an accurate PR plot is critical for reliable channel characterization. Therefore, it is preferred to measure the PR in situ to minimize the parasitic effects. In this work, we introduce a novel approach for measuring PR in situ, designed to quickly and accurately generate undistorted plot results. To prove the efficacy of the proposed method, we designed an on-chip sampling scope circuit and fabricated a test chip in 28-nm CMOS technology. While being able to measure a distortion-free PR, the proposed method demonstrates a more than$10^{5}$times faster pulse acquisition rate than prior arts. Won Joon Choi, Myungguk Lee, Junung Choi, Jaeik Cho, Gain Kim, Byungsub Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | A 2-Lane DAC-/ADC-Based 2 × 2 MIMO PAM-4 MMSE-DFE Wireline Transceiver With FEXT Cancellation on RFSoC PlatformabstractThis article presents a 2-lane$2 \times 2$multiple-input, multiple-output (MIMO) 4-level pulse amplitude modulation (PAM-4) minimum mean-squared-error (MMSE)-decision-feedback equalizer (DFE) with far-end crosstalk (FEXT) cancellation for digital-to-analog converter (DAC)-/analog-to-digital converter (ADC)-based high-speed serial links. The receiver (RX) datapath is designed with a 15-tap MIMO feedforward equalizer (FFE) and a one-tap MIMO DFE with the least mean square (LMS), enabling adaptation to channel variation while maintaining the MMSE setting. The RX digital signal processor (DSP) place and route (PnR) in a 28-nm CMOS is estimated to consume 201 mW/lane at a 56-Gb/s/lane data rate while occupying a 0.5-mm2/lane silicon area. We further implement a real-time evaluation platform to verify the functionality of the MIMO PAM-4 MMSE-DFE with rapid bit-error-rate (BER) testing on RFSoC. The measurement result demonstrates that the MIMO MMSE-DFE significantly improves BER performance from 2.75e−3to 1.31e−7compared with equalization without FEXT cancellation when communicating over a channel exhibiting 12.4-dB insertion loss (IL) and 13.2-dB IL-to-crosstalk ratio (ICR) at Nyquist. Seoyoung Jang, Donggeon Kim, Matthias Braendli, Thomas Morf, Marcel A. Kossel, Pier Andrea Francese, Gain Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2025 | A Fast Design Optimization of On-Chip Equalizing Links Using Particle Swarm OptimizationabstractWe propose a fast algorithm to optimize on-chip equalizing link design utilizing a particle swarm optimization (PSO) method. Finding the optimal design parameters of an equalizing link requires too much computation time, because the dependency between design parameters and performances is too complex, while design space is too large. The proposed algorithm greatly reduces the optimization time by utilizing the superior efficiency of PSO in heuristic search. In experiment, on average, the proposed algorithm optimized a link design$168\times $faster than the previous state-of-the-art result, requiring only 1/256 evaluation counts, and reduced computation time from about 2 h to 45 s. Hyoseok Song, Gain Kim, Byungsub Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2024 | DMT 3L4W: A 3-Lane 4-Wire Signaling With Discrete Multitone Modulation for High-Speed Wireline Chip-to-Chip InterconnectsabstractThis work presents a multi-lane transceiver (TRX) architecture with discrete multitone (DMT) modulation for pin-efficient high-bandwidth chip-to-chip wireline communication. The proposed signaling uses 4 wires to transmit 3 lanes of DMT-modulated symbols in parallel including one lane for signal encoding and decoding for the correlated noise cancellation. Compared to the pin-efficiency of 0.5 in differential signaling, the proposed signaling offers a pin-efficiency of 0.75 allowing each TRX lane to operate at a lower speed given a fixed data throughput or to increase the per-pin data rate with the same or even lower per-lane data rate to differential signaling. While each single-ended lane includes random noise, the receiver (RX) can effectively eliminate most of the correlated noise with simple arithmetic in the digital domain. Higher pin efficiency can be achieved with more single-ended data lanes per redundancy lane depending on the link specifications such as raw bit-error-rate (BER), and voltage-domain dynamic range of the transmitter’s driver and the receiver. Simulation results show that 28.5% of throughput gain can be achieved with 3-lane 4-wire configuration given the same channel, analog front-end circuits, and noise condition to multi-lane differential DMT TRXs. Seoyoung Jang, Donggeon Kim, Gain Kim |
ISCAS | 5 |
| 2024 | A 4×4 MIMO Discrete Multitone Wireline Transceiver With Far-End Crosstalk Cancellation For ADC-Based High-Speed Serial LinksabstractThis paper presents an area- and energy-efficient 4-lane far-end crosstalk (FEXT) cancellation wireline transceiver (TRX) with a multiple-input multiple-output (MIMO) discrete multitone (DMT) modulation. The channel estimation (CHEST) is an essential block for DMT TRX to find the MIMO equalizer coefficients at the receiver (RX) side. However, due to the high computational complexity, the matrix inversion in CHEST hinders the generalization to larger MIMO, such as 4×4, considering circuit implementation. In this work, we show that CHEST can be effectively approximated to an element-wise reciprocal instead of an inversion when some properties of the wireline channels are used as constraints. This approximation also simplifies the MIMO equalizer circuit and realizes a decentralized MIMO. Simulation results demonstrated that the FEXT noise from adjacent lanes is sufficiently canceled out even with our approximated CHEST and MIMO equalizer, achieving a symbol error rate (SER) of 2E-4 for communicating over a channel exhibiting insertion loss (IL) of 16 dB and 17 dB of IL-to-crosstalk ratio at Nyquist, while showing SER of 1e-1 when the FEXT is not canceled. Seoyoung Jang, Donggeon Kim, Matthias Braendli, Marcel A. Kossel, Andrea Ruffino, Thomas Morf, Pier Andrea Francese, Gain Kim |
ISCAS | 10 |
| 2024 | A Loop-Break Decision Feedback Equalizer for DAC/ADC-DSP-Based Wireline TransceiversabstractThis paper presents a novel digital decision feedback equalizer (DFE) design that can relax the feedback timing constraints for analog-to-digital converter (ADC)-based high-speed wireline receivers. The proposed technique breaks the loop-unrolled DFE (LU-DFE) chain by computing multiple LU-DFE chains in parallel with all possible seed symbols, and selecting the appropriate output by the post-processing selection logic. The proposed loop-break DFE (LB-DFE) is functionally equivalent to the conventional DFE with any other implementation techniques such as LU-DFE, look-ahead DFE (LA-DFE), or direct DFE. With topographical synthesis in 28nm CMOS process, the proposed LB-DFE achieved up to 54% of DFE area saving as compared to LA-DFE with look-ahead factor (LF) of 16 for 112Gb/s PAM-4 with 875MHz DSP clock speed. The implementation feasibility and functionality are verified using ZCU111 RFSoC platform at 6Gb/s (3GS/s ADC conversion rate) with a channel exhibiting 25dB loss at 1.5GHz, demonstrating the same bit error rate (BER) performance between the LB-DFE and the LA-DFE. Equipment-based measurements using arbitrary waveform generator (AWG) and real-time oscilloscope transmitting/receiving 40GBaud PAM-4 (80Gb/s) to/from the differential cables with software 21-tap feed-forward equalizer (FFE) and LB-DFE on PC was also conducted. Donggeon Kim, Seoyoung Jang, Sungyu Song, Matthias Braendli, Thomas Morf, Marcel A. Kossel, Pier Andrea Francese, Gain Kim |
IEEE Trans. Circuits Syst. I Regul. Pap. | 10 |
| 2024 | Compact Single-Ended Transceivers Demonstrating Flexible Generation of 1/N-Rate Receiver Front-Ends for Short-Reach LinksabstractThis paper presents compact single-ended wireline transceivers with software-generated receiver front-ends. The developed software framework significantly shortens the physical design time of 1/N-rate wireline receiver front-ends. The physical layouts of various receiver front-ends were software-generated in four different CMOS technology nodes (28 nm, 40 nm, 65 nm, and 90 nm) with four different front-end architectures targeting various data rates. In the post-layout simulation, the receiver front-ends generated within a second by the software achieved nearly the same performances as the manually-designed receiver front-ends that require more than about 30 hours of design time. For demonstration, we generated 8 Gb/s full-rate, 10 Gb/s half-rate, 12 Gb/s, and 20 Gb/s quarter-rate receiver front-ends, and fabricated them with a manually-designed feed-forward equalization transmitter in 28 nm CMOS process. The transceivers were measured with the data rate up to 20 Gb/s while consuming 1.39 pJ/b at the channel loss of −9.2 dB. The transceiver with software-generated receiver achieved the highest data rate per area as well as the smallest area among the relevant prior arts while reducing the physical design time of the receiver front-end by more than 140,000 times. Myungguk Lee, Jaeik Cho, Junung Choi, Won Joon Choi, Jiyun Lee, Iksu Jang, Changjae Moon, Gain Kim, Byungsub Kim |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2021 | Link Bit-Error-Rate Requirement Analysis for Deep Neural Network AcceleratorsabstractIn convolutional neural network (CNN) accelerators, the dominant power consumption is caused by the access of external data memory. In addition, power and area occupied by I/O interfaces maintaining low bit-error-rate, e.g., 1e-15, grow as the data rate increases. Considering the inherent error resilience of the inference process in machine learning applications, the requirement of error-free communication in the data-path is controversial. In this paper, a custom CNN accelerator integrating a channel emulator is designed by using an FPGA to analyze the effect of the BER of an I/O transceiver on the image classification accuracy. In order to implement a channel emulator, a digital-domain look-up-table (LUT)-based 12-tap FIR filter is employed to create inter-symbol interference (ISI), and a PRBS31 generator is used as a noise source. The implementation was evaluated by running the ImageNet dataset on the FPGA-based custom accelerator (Virtex Ultrascale+) implementing VGG-16. The results show that the BER up to 1e-4 in the memory access has a negligible impact on the inference accuracy. Gain Kim, Jinho Park 0004, Hyeon-Min Bae |
ISCAS | 2 |
| 2018 | Parallel Implementation Technique of Digital Equalizer for Ultra-High-Speed Wireline ReceiverabstractThis paper presents a parallel implementation technique of digital equalizer for high-speed wireline serial link receiver (RX). In wireline RX, inter-symbol interference (ISI) is mitigated by continuous-time linear equalizer, and the remaining ISI is cancelled out by decision-feedback equalizer (DFE). However, due to the existence of feedback loop in DFE, there is no trivial way to parallelize it, making it difficult to be realized in digital circuits for wireline RX based on analog-to-digital converter (ADC) with ≥ 56 Gb/s data rate. In this work, convolution theorem is applied for achieving parallel digital equalizer implementation. The digital equalizer datapath consists of discrete Fourier transform (DFT) core, inverse-DFT (IDFT) core, complex multipliers between DFT and IDFT cores, and overlap-add circuit. Design considerations for low-area VLSI implementation of such architecture is discussed. Gain Kim, Lukas Kull, Danny Luu, Matthias Braendli, Christian Menolfi, Pier Andrea Francese, Cosimo Aprile, Thomas Morf, Marcel A. Kossel, Alessandro Cevrero, Ilter Özkaya, Thomas Toifl, Yusuf Leblebici |
ISCAS | 1 |
| 2016 | A fully-digital spectrum shaping signaling for serial-data transceiver with crosstalk and ISI reduction property in multi-drop memory interfacesabstractAn efficient signaling scheme for serial-data transceivers (TRXs) has been proposed, which can properly reduce inter-symbol interference (ISI) and crosstalk (Xtalk) in memory interfaces. The proposed architecture relies on fully-digital implementation rather than analog/multi-tone approach, which can offer a very power-efficient and versatile silicon implementation. Moreover, the Xtalk induced noise can be fairly reduced by applying the proposed signaling, and the whole TRX can customize to the communication link trough digital calibration, while the aggregate data rate is kept fixed. Kiarash Gharibdoust, Gain Kim, Armin Tajalli, Yusuf Leblebici |
ISCAS | 2 |
| 2015 | Towards More Efficient Logic Blocks By Exploiting Biconditional Expansion (Abstract Only)abstractNowadays, Field Programmable Gate Arrays (FPGA) exploit Look-Up Tables (LUTs) to generate logic functions. A K-input LUT can implement any Boolean functions with K inputs. Thanks to this flexibility, LUTs remained conceptually unchanged in FPGAs, only the number of inputs increased in time. Unfortunately, the flexibility does not come for free and LUTs have non-negligible costs in both circuit-level performances (large number of memories, area or delay penalties) and logic-level capabilities (limited fan-out). Here, we propose an FPGA fabric based on two novel logic blocks. First, we introduce a new LUT design showing reduced power consumption with no sacrifice in the logic flexibility. Then, we present a block suited to arithmetic functions but preserving enough versatility to implement general logic functions. The two blocks are supported by a recently introduced logic representation called Biconditional Binary Decision Diagrams (BBDDs). Using architectural-level benchmarking, we showed that an FPGA architecture exploiting the novel blocks performs significantly better than current state-of-the-art FPGA architectures at 40nm technological node over a large set of test circuits. While reducing the power consumption of MCNC big20 benchmarks by 29%, the proposed architecture is able to efficiently implement arithmetic circuits as compared to its traditional LUT-based FPGA counterpart. For instance, a 256-bit adder can be realized with a 43% gain in area×delay product. While considering large general and arithmetic logic benchmarks, we observe, on average, 4%, 3% and 10% improvements in area, delay and power respectively. Pierre-Emmanuel Gaillardon, Gain Kim, Xifan Tang, Luca G. Amarù, Giovanni De Micheli |
FPGA | 2 |
| 2015 | Design optimization of polyphase digital down converters for extremely high frequency wireless communicationsabstractIn this paper, an area-optimized polyphase digital down converter (DDC) architecture is introduced, where the mixers can be completely merged into the polyphase decimation filter under certain conditions. We also introduce an interface architecture, called synchronizer, between the back-end of an extremely high-speed time interleaved ADC (TI-ADC) and the front-end of a polyphase DDC. The synchronizer enables safe downsampling for a polyphase DDC, when the TI-ADC's sampling rate is above tens of GS/s. We show that the proposed interface architecture prevents any potential timing constraint violations that might occur in the interface between a TI-ADC and a polyphase DDC for extremely high frequency (EHF) wireless communication applications. Gain Kim, Raffaele Capoccia, Yusuf Leblebici |
VLSI-SoC | 1 |
| 2015 | A Novel FPGA Architecture Based on Ultrafine Grain Reconfigurable Logic CellsabstractIn this paper, we investigate the opportunity brought by controllable-polarity transistors to design efficient reconfigurable circuits. Controllable-polarity transistors are devices whose polarity can be electrostatically programmed to be either n- or p-type. Such devices are used to build ultrafine grain computation cells. These cells are arranged into regular matrices, called MClusters, with a fixed and incomplete interconnection pattern, employed to minimize the reconfigurable interconnection overhead. We subsequently use them into field-programmable gate arrays (FPGAs). To assess this architectural scheme in an efficient and objective manner, we present a complete benchmarking tool flow and focus on the packing algorithm developed to handle the architecture. We finally perform the evaluation with widely used benchmark circuits. Leveraging the ultrafine grain cells compactness from a system-level perspective, we show that FPGAs exploiting MClusters demonstrate average savings of 43% and 23% in area and delay, respectively, as compared with the CMOS lookup table FPGA counterpart at 22-nm technological node. Pierre-Emmanuel Gaillardon, Xifan Tang, Gain Kim, Giovanni De Micheli |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |