Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Seokjin Jeong

dblp:158/8134 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
0since 2021 · last 2015
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 67% Reconfigurable computing and FPGAs · 33%
Computer graphics and multimedia
1 paper
Image and video coding · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › signal processing accelerator
DCT processor
0.212015
Design of a Loeffler DCT using Xilinx Vivado HLS (Abstract Only) · FPGA 2015
Reconfigurable computing and FPGAs
FPGA-based signal processing
0.212015
Design of a Loeffler DCT using Xilinx Vivado HLS (Abstract Only) · FPGA 2015
Hardware accelerators and domain-specific architectures
signal processing accelerator
0.212015
Design of a Loeffler DCT using Xilinx Vivado HLS (Abstract Only) · FPGA 2015
Image and video coding › transform coding
discrete cosine transform
0.112015
Design of a Loeffler DCT using Xilinx Vivado HLS (Abstract Only) · FPGA 2015

Methods — techniques the papers use, named apart from their topics

shift-and-add · 0.4high-level synthesis · 0.4distributed arithmetic · 0.4
YearPublicationVenuePosition
2015 Design of a Loeffler DCT using Xilinx Vivado HLS (Abstract Only)
abstract
Loeffler discrete cosine transform (DCT) algorithm is recognized as the most efficient one because it requires the theoretically least number of multiplications. However, many applications still encounter difficulty in performing the 11 multiplications required by the algorithm to calculate a 1D eight-point DCT. To avoid expensive multipliers in the hardware, we used two design methods, namely, distributed arithmetic (DA) and shift-and-add (SAA) methods, to design the DCT accelerator. The memory bandwidth is 60 bits: 24 bits for reads of the R(red), G(green), and B(blue) data of a pixel and 36 bits for writes of three corresponding 12-bit DCT coefficients. Thus, the 1D eight-point DCT accelerator for each of R, G, and B can have one 12-bit input port and one 12-bit output port so that it can calculate a 2D DCT by row-column decomposition method. The designs are adjusted to produce the same latency and interval. DA seems promising because Loeffler DCT requires only three small tables with four input bits. However, our experiments using Xilinx Vivado HLS show that the SAA design is better than the DA design for the considered applications. Furthermore, simulation results suggest that the optimal accelerator design can be obtained by adjusting the SAA design to the considered applications. The resultant SAA design requires only 13 adders (per color component) and can calculate one DCT coefficient per clock cycle. The precision of the internal hardware has been adjusted, such that the reconstructed images have PSNR values of at least 39.1 dB for all test images (Lenna, Pepper, House, and Cameraman). If a precision of 13bits is allowed, PSNR becomes at least 44.8 dB. Our presentation describes the architecture and operation of the optimized SAA design.
Seung Yeol Baik, Seokjin Jeong, Hyeong-Cheol Oh
FPGA2