VLDB 2026 Research / reviewers in the wild / expert
Wanyeong Jung
dblp:153/9716
· DBLP profile ↗
11ranked-venue papers
0as first author
10since 2021 · last 2025
0000-0002-5671-1341ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TickTockStack: In-Datapath Current Imbalance Elimination Using Clocked Differential Logic in a Voltage Stacked Vector ProcessorabstractIn modern digital chips with high power density, a substantial portion of energy is wasted in the power delivery network (PDN), causing thermal issues and voltage drops, which reduces system reliability. Voltage stacking mitigates this problem by connecting power ports of circuits with similar current profiles in series, reusing their charge, and enabling a reduction in PDN loss. Still, it suffers from voltage ripple and power delivery loss due to the current imbalance of stacked circuits. Even in identical circuits, these imbalances arise from activity variations in CMOS logic. To address this issue, this paper presents TickTockStack, a voltage-stacked vector processor designed through architecture-circuit co-optimization. A vector processor was selected for its synchronized instruction execution across lanes, which minimizes control-driven variation in stacked domains. To mitigate the remaining data-dependent imbalance, the datapath was implemented using clocked dual-rail domino (DRD) logic with constant activity factor and differential structure, the register file was realized as a fully differential latch-based structure with symmetrical multiplexers, and the single-ended control logic uses CMOS gates. The processor was evaluated using standard benchmarks, demonstrating a 6.2× reduction in stack voltage ripple and a 33.5% lower power delivery loss on average compared to a stacked CMOS baseline, enhancing stacking reliability. The area overhead of 18.4% was 4.9x lower than that of a switched-capacitor regulator for equivalent ripple reduction, and unlike subthreshold balancing, the proposed method imposes no limitations on maximum operating frequency. Michal Andrzej Gorywoda, Wanyeong Jung |
ICCAD | 2 |
| 2024 | VVIP: Versatile Vertical Indexing Processor for Edge ComputingabstractThis paper presents a versatile vertical indexing processor (VVIP) based on a single-instruction multiple-data architecture for edge computing. In VVIP, the vertical source and destination indexing instructions are customized for area-efficient computations. The proposed indexing method reorders data within a processing module by using more registers and data-steering logic in the calculations. In particular, VVIP supports multibit-serial multiplication and sparse data operations by leveraging register files as lookup tables or accumulators. The VVIP, verified on a vector processor, has an area overhead of less than 2.8%. It exhibits an average computation rate that is 10.1 times faster than the 1-bit-serial multiplication in linear algebra benchmarks, and 1.2 times average performance improvement in unstructured sparse point-wise convolution tasks when compared to conventional control sequences. Hyungjoon Bae 0001, Da Won Kim, Wanyeong Jung |
DAC | 3 |
| 2024 | SeGen: Automatic Topology Generator for Sequencing ElementsabstractSequencing elements, such as flip-flops (FFs), significantly impact the speed, size, and power consumption of digital integrated circuits. Despite numerous sequencing-element proposals in the past, they have heavily relied on the expertise of designers, and their design method could not search all available FF topologies. This paper introduces a design framework for automatically generating sequencing element designs, named SeGen. Motivated by the fact that the operation of all digital circuits can be represented by a Boolean function, we present SeGen which initially generates all possible Boolean functions for sequencing elements and derives circuit topologies from each of them. By adjusting the functionality qualification step of Boolean functions, SeGen can generate other sequencing elements such as toggle FFs and dual-edge-triggered FFs as well. A total of 40 resulting topologies newly generated by SeGen encompass the entire spectrum of positive-edge-triggered FFs, including master-slave FFs and pulsed latches, and consequentially provide a range of designs suitable for diverse applications. Several generated FFs such as SeGen8 and SeGen28 outperform recent human-crafted FFs and exhibit energy-delay product (EDP) improvement of 203%--257% at 1V supply voltage, compared to conventional transmission-gate FF (TGFF). Kyounghun Kang, Wanyeong Jung |
ICCAD | 2 |
| 2024 | MAC-DO: DRAM-Based Multi-Bit Analog Accelerator Using Output StationaryabstractDRAM-based accelerators have shown their potential in addressing the memory wall challenge of the traditional von Neumann architecture. Such accelerators exploit charge sharing or logic circuits for simple logic operations. As a result, they require many cycles for more complex operations such as a multi-bit multiply-accumulate (MAC) operation, resulting in significant data access and movement and potentially worsening power efficiency.To overcome these limitations, this paper presents MAC-DO, an efficient and low-power DRAM-based accelerator. Compared to previous DRAM-based accelerators, a MAC-DO cell, consisting of two 1T1C DRAM cells, innately supports a multi-bit MAC operation within a single cycle, significantly improving power efficiency while maintaining good linearity and compatibility with existing 1T1C DRAM cell and array structures. This achievement is facilitated by a novel analog computation method utilizing charge steering. As a result, MAC-DO efficiently can accelerate convolutions based on output stationary mapping, supporting the majority of computations performed in deep neural networks.Our evaluation using transistor-level simulation shows that a test MAC-DO array with 16×16 MAC-DO cells achieves 120.96 TOPS/W and 97.07% Top-1 accuracy for MNIST dataset without retraining. Minki Jeong, Wanyeong Jung |
ISCAS | 2 |
| 2024 | A Compact and Low-Power Column Readout Circuit based on Digital Delay ChainabstractThis paper presents a column readout integrated circuit (ROIC) optimized for an analog computing array in terms of size, design simplicity, robustness, and energy efficiency. The digital delay chain with capacitive feedback converts current and charge input to a thermometer code with high accuracy. Adopting dynamic AND gates allows for sequentially changing the negative feedback loop through unit capacitors controlled by a loop-unrolled chain topology. The readout circuit shows superior linearity with essentially no stability problems. Simulated with a 28 nm CMOS technology, the circuit achieves a 5-bit resolution with a DNL of +1.148/-1.147 LSB and an INL of +0.817/-0.677 LSB in the case of current input for 3σ mismatch. The power consumption is 54.8 µW from a 0.9 V supply at the conversion rate of 400 MS/s, and the circuit occupies 29 µm2. Minkyu Yang, Changjoo Park, Wanyeong Jung |
ISCAS | 3 |
| 2023 | 26TSPC: A Low Hold Time, Low Power Flip-Flop With Clock Path OptimizationabstractRecent low-power flip-flops (FFs) synthesize a signal$(\boldsymbol{CKN})$that is activated only when the input data$(\boldsymbol{D})$updates the output state$(\boldsymbol{Q})$. It saves power consumption by avoiding unnecessary transitions in internal nodes. However, the circuit for$\boldsymbol{CKN}$that removes all redundant transitions becomes complex. As a result, its delay worsens timing parameters (setup time, hold time,$\boldsymbol{CK}\_\boldsymbol{Q}$delay) and the robustness of FFs too. By optimizing the delay of the clock insertion paths, the presented flip-flop (26TSPC) maximizes efficiency without speed degradation, while maintaining low hold time. Its design is based on 18TSPC, one of the fast and energy-efficient FFs, and its contention and redundant transition issues are also resolved. Post-layout simulation results based on 65 nm CMOS process show that 26TSPC consumes 83.9%/71.4% less power than that of conventional transmission-gate flop-flop (TGFF) with 10%/20% activity ratio at 1V and 75.4%/65.3% less power than that of TGFF with 10%/20% activity at 0.4V respectively. In addition, the hold time of 26TSPC is negative in all corners, featuring better timing reliability than other low-power FFs. Kyounghun Kang, Wanyeong Jung |
ISCAS | 2 |
| 2023 | An Energy-Efficient Delay Insensitive Asynchronous Interface for Globally Asynchronous Locally Synchronous (GALS) SystemabstractThe process, voltage, and temperature (PVT) variations are significant factors that must be considered when designing chips, especially chips designed for edge devices. Typically, chip designers use an extra timing margin to guarantee the operation of the chip under the worst-case condition. However, this approach can lead to reduced performance and energy efficiency. A globally asynchronous locally synchronous system (GALS) is an alternative approach to ensure chip operation while maintaining performance. This paper presents an energy-efficient asynchronous interface with a delay-insensitive protocol for the GALS system. As the prior delay-insensitive interface consumes significant energy, low-swing signaling by stacking the voltage domain reduces energy consumption at little cost. In addition, adaptive timing control is implemented to support a long wire connection and improve interface robustness. As a result, the energy-efficient asynchronous interface implemented on 65nm technology consumed 1.6x less energy than the prior delay-insensitive asynchronous interface while being able to drive wire up to 1.25mm long. Dalta Imam Maulana, Wanyeong Jung |
ISCAS | 2 |
| 2022 | A Near-Memory Radix Sort Accelerator with Parallel 1-bit SorterabstractSorting is one of the most fundamental operations for many applications. For efficient sorting, data locality can be exploited by processing subdivided data in parallel. This work presents a high-performance and area-efficient near-memory radix sort accelerator where end-to-end sorting is performed locally. With a parallel 1-bit radix sorter, it achieves high throughput by processing multiple keys per cycle. Tested with Xilinx Zynq UltraScale+ ZCU104 FPGA, the experimental result shows up to 10x performance speedup over CPU. It is highly area-efficient and can be integrated into each processing node of a distributed computing system with low area cost. Jihwan Cho, Dalta Imam Maulana, Wanyeong Jung |
FCCM | 3 |
| 2021 | RectBoost: Start-Up Boosting for Rectenna Using an Adaptive Matching NetworkabstractAdvancement in wireless technology has promoted the emergence of the Internet of Things (IoT). Specifically, radio- frequency energy harvesting can accelerate the development of IoT technology by enabling the use of miniaturized battery- less sensor nodes. To ensure efficient energy harvesting, the impedance of the antenna and the rectifier should be matched through an impedance matching network. We present RectBoost, which is a dynamically adaptive matching network, for optimum harvesting performance of rectenna. RectBoost monitors the output voltage of the rectifier and accordingly tunes the impedance matching mode. This technique helps the rectenna overcome early state impedance mismatch by dynamically tuning matching networks. RectBoost provides two main advantages over prior works - output voltage and charging speed boost. Simulation results show that the output voltage is increased by 7 times, and the charging speed is improved by 48 times compared with those of the methods mentioned in prior works. Joonhyuk Cho, Wanyeong Jung |
ISCAS | 2 |
| 2021 | An On-Chip Dual-Output Switched-Capacitor DC- DC Converter with Fine-Grained Output ControlabstractThis paper presents a fully integrated dual-output switched-capacitor DC-DC converter with dynamic four-phase operation. The presented converter is derived from an existing topology known as the algorithmic voltage-feed-in and it maintains the voltage conversion ratio (VCR) reconfiguration capability of the original topology. By using a selective pair of VCRs, the presented converter generates two output voltages with an arbitrary ratio with high efficiency. Additionally, efficiency is maintained for various combinations of load currents via dynamic phase skipping and frequency modulation. A test chip fabricated using a 0.18 pm CMOS process exhibits conversion efficiencies higher than 89% for all VCR combinations with a 0.5 mA load current. Doojin Jang, Unbong Lee, Jeongmyeong Kim, Jangwon Suh, Wanyeong Jung |
ISCAS | 5 |
| 2018 | Edge pursuit comparator with application in a 74.1dB SNDR, 20KS/s 15b SAR ADCabstractThis paper presents a new energy-efficient ring oscillator collapse-based comparator, which is called edge-pursuit comparator (EPC) and demonstrated it in a 15-bit SAR ADC. The comparator automatically adjusts the performance according to its input difference without any control, eliminating unnecessary energy spent on coarse comparisons. The employed SAR ADC supplements a 10-bit differential main CDAC with a 5-bit common-mode CDAC which uses common to differential gain tuning to improves linearity by reducing the effect of switch parasitic capacitance. A test chip fabricated in 40nm CMOS shows 74.12 dB SNDR and 173.4 dB FOMs. The comparator consumes 104 nW with the full ADC consuming 1.17 μW. Minseob Shim, Seokhyeon Jeong, Paul D. Myers, Suyoung Bang, Junhua Shen, Chulwoo Kim, Dennis Sylvester, David T. Blaauw, Wanyeong Jung |
ASP-DAC | 9 |