Hyungjoon Bae 0001

dblp:340/0716-1 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2024
0000-0001-7045-3281ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Integrated circuit design · 50% Processor architecture and microarchitecture · 25% Hardware accelerators and domain-specific architectures · 19%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
edge accelerator
0.812024
VVIP: Versatile Vertical Indexing Processor for Edge Computing · DAC 2024
Processor architecture and microarchitecture
SIMD
0.812024
VVIP: Versatile Vertical Indexing Processor for Edge Computing · DAC 2024
Integrated circuit design › digital circuit design › sequential circuit design
counter design
0.712023
High-Speed Counter With Novel LFSR State Extension · IEEE Trans. Computers 2023
Integrated circuit design
digital circuit design
0.712023
High-Speed Counter With Novel LFSR State Extension · IEEE Trans. Computers 2023
Memory systems
lookup table
0.212024
VVIP: Versatile Vertical Indexing Processor for Edge Computing · DAC 2024
Processor architecture and microarchitecture
register file
0.212024
VVIP: Versatile Vertical Indexing Processor for Edge Computing · DAC 2024

Methods — techniques the papers use, named apart from their topics

multibit-serial multiplication · 0.8
YearPublicationVenuePosition
2024 VVIP: Versatile Vertical Indexing Processor for Edge Computing
abstract
This paper presents a versatile vertical indexing processor (VVIP) based on a single-instruction multiple-data architecture for edge computing. In VVIP, the vertical source and destination indexing instructions are customized for area-efficient computations. The proposed indexing method reorders data within a processing module by using more registers and data-steering logic in the calculations. In particular, VVIP supports multibit-serial multiplication and sparse data operations by leveraging register files as lookup tables or accumulators. The VVIP, verified on a vector processor, has an area overhead of less than 2.8%. It exhibits an average computation rate that is 10.1 times faster than the 1-bit-serial multiplication in linear algebra benchmarks, and 1.2 times average performance improvement in unstructured sparse point-wise convolution tasks when compared to conventional control sequences.
Hyungjoon Bae 0001, Da Won Kim, Wanyeong Jung
DAC1
2023 High-Speed Counter With Novel LFSR State Extension
abstract
This paper presents a high-speed counter architecture associated with novel LFSR state extension. By employing the proposed state extension, an$\emph {m}$-bit LFSR counter with$(\mathrm{2^{\mathit{m}}}-1)$states is modified to cover$\mathrm{2^{\mathit{m}}}$states without degrading the counting rate. Based on the property that only the low-order bits are frequently switched, the proposed counter consists of two sub-counters to achieve a high counting rate and reduce the hardware complexity needed to convert an LFSR state into a binary state. The low-order sub-counter is implemented with the proposed LFSR counter, and the high-order sub-counter is designed by employing the conventional synchronous binary counter. In addition, the implemented counter takes into account the speed degradation caused by the large fan-out of the high-order sub-counter. The proposed counter designed with standard cells operates at 2.08 GHz in a 65 nm CMOS technology, and its counting rate is almost independent of the counter size.
Hyungjoon Bae 0001, Yujin Hyun, Suchang Kim, Sangsoo Park, Jaeyoung Lee 0004, Boseon Jang, Suyoung Choi, In-Cheol Park
IEEE Trans. Computers1
2023 A CNN Inference Accelerator on FPGA With Compression and Layer-Chaining Techniques for Style Transfer Applications
abstract
Recently, convolutional neural networks (CNNs) have actively been applied to computer vision applications such as style transfer that changes the style of a content image into that of a style image. As the style transfer CNNs are based on encoder-decoder network architecture and should deal with high-resolution images that become mainstream these days, the computational complexity and the feature map size are very large, preventing the CNNs from being implemented on an FPGA. This paper proposes a CNN inference accelerator for the style transfer applications, which employs network compression and layer-chaining techniques. The network compression technique is to make a style transfer CNN have low computational complexity and a small amount of parameters, and an efficient data compression method is proposed to reduce the feature map size. In addition, the layer-chaining technique is proposed to reduce the off-chip memory traffic and thus to increase the throughput at the cost of small hardware resources. In the proposed hardware architecture, a neural processing unit is designed by taking into account the proposed data compression and layer-chaining techniques. A prototype accelerator implemented on a FPGA board achieves a throughput comparable to the state-of-the-art accelerators developed for encoder-decoder CNNs.
Suchang Kim, Boseon Jang, Jaeyoung Lee 0004, Hyungjoon Bae 0001, Hyejung Jang, In-Cheol Park
IEEE Trans. Circuits Syst. I Regul. Pap.4