EDBT 2026 Demo / reviewers in the wild / expert
Seongmo Park
dblp:32/5214
· DBLP profile ↗
6ranked-venue papers
2as first author
0since 2021 · last 2018
0000-0001-8656-9094ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Hardware accelerators and domain-specific architectures · 77% Parallel and multicore computing · 23% |
Topics — the 1 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
video coding accelerator |
0.3 | 1 | 2018 | An Efficient Architecture of In-Loop Filters for Multicore Scalable HEVC Hardware Decoders · IEEE Trans. Multim. 2018 |
Methods — techniques the papers use, named apart from their topics
window-based parallel SAO · 0.3skip mode pipelining · 0.3adaptive deblocking filtering order · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | An Efficient Architecture of In-Loop Filters for Multicore Scalable HEVC Hardware DecodersabstractThis paper proposes an efficient architecture of HEVC in-loop filters (ILFs) with the target of providing effective multicore utilization for ultra-high definition video applications. While HEVC allows for a high level of parallelization, the issue of data dependencies at the ILF leads to inefficient parallel processing performance. The novel memory organization and management techniques address the data dependence-related issues between multiple processing units and enable to filter the flexible area on multicore decoder. In addition, we introduce the adaptive deblocking filtering order (ADFO) to minimize the impact of bus congestion when multiple cores interoperate for processing very large data. Furthermore, we design the deblocking filter with skip mode pipelining to achieve the high performance minimizing the increased cost and the power consumption. For SAO, we apply the window-based parallel SAO filtering scheme. The resource sharing is considered throughout the entire architecture. Based on both experimental and analytical results, our proposed design can achieve more than 1.31 Gpixels/s and less than 2.6 Gpixels/s at maximum frequency 660 MHz in single core, and consumes 56.2 Kgates including 10.6 Kgates for memory management architecture, which supports multicore decoder, and about 20.8 mW power on average when synthesizing with the 28 nm CMOS library. Moreover, the skip modes of DF improve both the performance and the power dissipation. The ADFO improves the performance of ~9.17% when decoding 8 K sequence on octacore at 400 MHz frequency. TpG (Throughput per Gate) is the highest among the related works. HyunMi Kim, JeongGil Ko, Seongmo Park |
IEEE Trans. Multim. | 3 |
| 2015 | A hybrid embedded compression codec engine for ultra HD video applicationabstractWe proposed an efficient VLSI hardware architecture of the High Efficiency Video Coding (HEVC) using a hybrid embedded compression algorithm for reducing the frame memory bandwidth. This architecture was designed to reduce the memory bandwidth using an adaptive prediction lossy/lossless algorithm. We saved about 50% of the memory access cycles for the reference data compared to a previous algorithm. The PSNR degradation of 0.12 dB on average was proposed algorithm at the compression ratio of 50%. The architecture was implemented in Verilog HDL and synthesized using a Synopsys Design Compiler with a 65nm cell library; the gate count was about 25,000 gates. Seongmo Park, Kyungjin Byun, Nak-Woong Eum |
VLSI-SoC | 1 |
| 2013 | Fast decision of CU partitioning based on SAO parameter, motion and PU/TU split information for HEVCabstractHigh Efficiency Video Coding (HEVC) has recently been standardized with a significant improvement of coding efficiency compared to its preceding video coding standards. To achieve high coding efficiency improvements, HEVC adopts deeper hierarchical block coding structures of coding unit (CU), prediction unit (PU) and transform unit (TU), to better adapt the complex texture and motion natures of various video sequences. However, this causes a dramatically increased computational and structural complexity of HEVC encoders and decoders. In this paper, we propose a fast decision method of CU partitioning, which significantly reduces total encoding time with negligible RD-performance loss for HEVC. For this, the proposed method effectively utilizes the available side information such as sample adaptive offset (SAO) parameter values, PU sizes, MV sizes and coded block flag (cbf) data, so that the required computational complexity for fast CU split decision is minimized. Our experiment results show that the proposed method reduces the total encoding time to average 43.2% only with average 1.58% bit-rate increase for ten test sequences of HEVC. Sangsoo Ahn, Munchurl Kim, Seongmo Park |
PCS | 3 |
| 2010 | New Lookup Tables and Searching Algorithms for Fast H.264/AVC CAVLC DecodingabstractIn this paper, new codeword structures, tables, and searching methods for fast and efficientcoeff_token,total_zeros, andrun_beforedecoding are developed. This new achievement is mainly based on the fact that the context-adaptive variable length coding (CAVLC) decoding can be modeled as a finite state machine. In order to quantitatively evaluate the proposed method in terms of decoding speed and complexity, we define the iteration bound$\left({{1}\over {\mathtilde{\tau}}}\right)$and thecomplexity ratio$(CR)$. Using these gauge variables, we show that the new algorithms reduce${\mathtilde {\tau}}$to about one third andcomplexity ratioto 0.95. This means that the proposed techniques reduce the decoding time to about one third and memory access count by 90% compared to those of the conventional methods without implementation overheads. Multiple-symbol parallel decoding method forrun_beforesyntax element is proposed based on abit-positioningwith the critical path latency of only one multiplexer for the post-combination process. The proposed methods make it possible to implement a fast and efficient CAVLC decoding without losing video quality on any environments. Jae-Jin Lee, Seongmo Park |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2008 | A 100MHz ASIP (application specific instruction processor) for CAVLC of H.264/AVC decoderabstractIn this paper, we implement the configurable processor for CAVLC function module of a H.264/AVC baseline profile decoder as the starting point to implement the H.264/AVC decoder system in a multiprocessor platform. The requirements of the implementations are the low-power processor and speed optimized algorithms tailored to the processor architecture. An arithmetic formula mapping method for fast CAVLC algorithms and a dual-issue VLIW processor architecture with custom instructions are proposed. The experiment results show that the synthesized processor has about 75 K gates and can carry out the decoding of 30-fps CIF (352x288 pixels) images around 120 Mega cycles. Jae-Jin Lee, MooKyoung Jeong, Nak-Woong Eum, Seongmo Park |
ISCAS | 5 |
| 2000 | An area efficient video/audio codec for portable multimedia applicationabstractIn this paper, we present an area efficient video and audio single chip encoder/decoder for portable multimedia application. The single-chip called as VASP (Video Audio Signal Processor) consists of a video signal processing block and an audio signal processing block. This chip has a mixed hardware/software architecture to combine performance and flexibility. The video signal processing block was designed to implement hardwired solution of pixel input/output, full pixel motion estimation, half pixel motion estimation, discrete cosine transform, quantization, run length coding, host interface, and 16 bit RISC type internal controller. The audio signal processing block is implemented with a software solution using 16 bit fixed point DSP. This chip contains 142,300 gates, 22 kbits FIFO, 107 kbits SRAM, and 556 kbits ROM, and the chip size was 9.02 mm/spl times/9.06 mm which was fabricated using 0.5 micron 3-layers metal CMOS technology. Seongmo Park, Kyeongjin Byeon, Jinjong Cha, Hanjin Cho |
ISCAS | 1 |