Yoshitaka Kidani

dblp:254/8199 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
8since 2021 · last 2026
0009-0002-8855-4106ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 8 since 2021
YearPublicationVenuePosition
2026 High-Efficiency, Low-Complexity Inter-Frame Coding for Video-Based Dynamic Mesh Coding (V-DMC)
abstract
Video-based Dynamic Mesh Coding (V-DMC), being standardized by the MPEG-3DGH group, seeks to establish an efficient standard for compressing dynamic meshes with time-varying vertex positions, connectivity, and attributes. By adopting a subdivision and video-based framework, V-DMC effectively leverages both advanced static mesh and video codecs, achieving state-of-the-art dynamic mesh coding efficiency. However, in V-DMC, inter-frame coding—one of the key coding tools—is applied only to a limited subset of frames and remains computationally expensive. To overcome this limitation, we propose a novel inter-frame coding framework for V-DMC that extends applicability to a much larger portion of frames while achieving both high efficiency and low complexity. Specifically, our method consists of four modules: (a) supervoxel-based shape matching for robust and efficient motion estimation; (b) embedded graph deformation for accurate geometry tracking; (c) early interframe coding mode decision for accelerated rate-distortion optimization; and (d) UV atlas tracking for generating temporally consistent texture images. Experimental results on the MPEG dynamic mesh dataset show that the proposed method achieves BD-rate improvements of −10.1%, −10.1%, −14.7%, −17.6%, and −15.5% for D1, D2, Luma, Cb, and Cr, respectively, along with a 5% decrease in encoding runtime compared to V-DMC reference software.
Xudong Jin, Yoshitaka Kidani, Kei Kawamura
IEEE Trans. Circuits Syst. Video Technol.2
2025 Chained Motion Vector Prediction for Video Coding
abstract
Merge mode has been utilized in advanced video coding standards to facilitate the efficiency of Motion vector (MV) or Block Vector (BV) coding. Merge mode constructs multiple MVs or BVs as the merge candidate list from MV/BV storage and then specifies one within the list by signaling the merge index. Despite various merge candidate derivation methods in prior arts, they do not adequately reach MV/BV, pointing to reference pictures with low quantization noise, leaving room for improved coding performance. This paper proposes a Chained Motion Vector Prediction (CMVP) as a novel merge candidate derivation. The CMVP derives new candidates by accumulating the recursively traced MVs or BVs based on pre-derived merge candidates. Experimental results demonstrate that the proposed method achieves up to 1.01% coding gains with negligible complexity increases on version 12 of Enhanced Compression Model (ECM), a reference software for evaluating promising coding tools beyond Versatile Video Coding (VVC).
Yoshitaka Kidani, Haruhisa Kato, Kei Kawamura
ICASSP1
2025 Block Vector based Intra Prediction Mode Derivation for Beyond VVC
Haruhisa Kato, Yoshitaka Kidani, Kei Kawamura
PCS2
2024 Extended Multiple Cross-Component Linear Models With Adaptive Thresholding and Overlapped Averaging Beyond VVC
abstract
In this paper, we propose an extended multimodel cross-component linear model (MMLM) for video compression beyond the versatile video coding standard. Our proposed method incorporates adaptive thresholding and overlapped averaging to enhance prediction accuracy and reduce discontinuities in multiple linear models. We evaluate our method’s coding gain on various video sequences and demonstrate a notable improvement of up to 1.3% bit-rate savings over the conventional MMLM, validating our method’s efficiency in high-efficiency video compression.
Haruhisa Kato, Yoshitaka Kidani, Kei Kawamura
ICIP2
2024 Bi-Predictive Intra Block Copy for Enhanced Video Coding Beyond VVC
abstract
Intra block copy (IBC), an intra coding tool with a single block vector (BV), has been exploited for significant coding gains of screen content (SC) in advanced video coding standards such as VVC. Several studies have applied IBC to camera-captured content (CC), such as the IBC with fractional-sample-precision BV, which was adopted into the reference software for exploring beyond VVC, i.e., the enhanced compression model (ECM). However, there is room to further achieve the coding gains of IBC because all the conventional methods are uni-predictive IBC with a single BV to generate prediction samples. This paper proposes a bi-predictive IBC using two BVs as a new IBC algorithm for CC and SC, realized by extending the number of BVs in BV storage. In addition, this paper proposes encoder early terminations of applying IBC for CC by comparing coefficients and distortions of the IBC and intra prediction to avoid encoder runtime increases while maintaining coding gains. Experimental results show that the proposed method brings $0.15 \%$ and $0.30 \%$ coding gains for CC and SC over ECM-9 under all-intra configuration, with negligible complexity increases. The proposed method has been adopted into ECM-10.
Yoshitaka Kidani, Haruhisa Kato, Kei Kawamura
ICIP1
2024 Low-complexity learning-based intra prediction with direction-dependent adaptive weights for beyond VVC
abstract
This paper introduces an advanced intra prediction method designed for the Enhanced Compression Model (ECM), which is the reference software for beyond versatile video coding (VVC) standard. It employs a learning-based method to adaptively assign weights for a weighted average across neighboring samples, resulting in more precise prediction samples. The proposed method derives optimized weights for each intra prediction mode, for each block size, and for each sample position. To achieve a reasonable balance between encoding time and prediction accuracy, the conventional intra prediction mode is shared with the proposed method. Experimental evaluations have demonstrated that the proposed method provides bitrate reduction of up to 0.4%.
Haruhisa Kato, Yoshitaka Kidani, Kei Kawamura
VCIP2
2023 Extended Intra Block Copy with Adaptive Filtering and Overlapped Block Averaging
abstract
Next-generation video coding standards are attempting to improve coding performance compared to conventional standards such as VVC by extending technologies such as intra-block copy (IBC). While IBC in VVC has proven effective for screen content, its adaptation to camera-captured content presents challenges regarding sample fluctuations and the continuity of block boundaries. This paper proposes a novel approach to improve IBC performance for camera-captured content by combining adaptive filtering (F-IBC) and overlapped block averaging (OB-IBC). The F-IBC filters IBC prediction samples using filter coefficients derived from adjacent samples to predict sample fluctuations accurately. The OB-IBC is a weighted average of IBC prediction samples of the current block with adjacent samples of the adjacent block’s reference to connect block boundaries smoothly. Following common test conditions in the joint video experts team, experimental results show improved coding performance with a bitrate saving of 0.1 % over the reference software (ECM 7.0) which investigates the enhanced compression beyond VVC capability.
Haruhisa Kato, Yoshitaka Kidani, Kei Kawamura, Sei Naito
VCIP2
2022 Adaptive boundary width of Geometric Partitioning Mode for Beyond Versatile Video Coding
abstract
In order to improve coding efficiency beyond versatile video coding (VVC), we propose an extended geometric partitioning mode (GPM). GPM is a new inter prediction in VVC and is applied to the object boundary between the foreground and background with different motions. Specifically, GPM partitions a rectangular coding block into two regions with 64 predefined types of straight lines, generates inter prediction samples for each partitioned region and then blends them with a fixed boundary width to obtain the final prediction samples. However, the fixed boundary width of GPM is not always optimal for diverse video content. To solve this problem, the proposed method allows GPM to select multiple boundary widths by block-wise signaling. Furthermore, the proposed method also restricts the selectable boundary width according to the short side of the block to reduce the encoding time for selecting the optimal width. Experiment results following common test conditions in JVET showed an improvement in coding efficiency with bitrate savings of 0.11 % and 3.20 % for camera-captured content and for pure screen or video game content, respectively, compared VVC reference software.
Haruhisa Kato, Yoshitaka Kidani, Kei Kawamura, Sei Naito
VCIP2
2020 Block-Size Dependent Overlapped Block Motion Compensation
abstract
Overlapped block motion compensation (OBMC) is one of the inter prediction tools that improves coding performance. OBMC applied to various non-squared blocks has been studied in VVC, which is being standardized by joint video experts team (JVET), to improve coding performance over HEVC. Memory bandwidth, however, is a bottleneck when OBMC is used, and conventional methods have not achieved a good trade-off regarding coding performance and memory bandwidth so far. In this study, interpolation filters and applicable conditions of OBMC depending on block sizes are proposed to achieve the best trade-off. The experimental results show a -0.40% BD-rate gain compared with that of the VVC test model 3 for random access conditions under the common test condition in JVET.
Yoshitaka Kidani, Kei Kawamura, Kyohei Unno, Sei Naito
ICIP1
2019 Blocksize-QP Dependent Intra Interpolation Filters
abstract
Intra interpolation filters for intra angular prediction play an important role in the coding performance. In the intra angular prediction of VVC, which is being standardized by the joint video coding expert team (JVET), block-size based switchable interpolation filters between 4-tap cubic and Gaussian interpolation filters is being studied. Although the two filters have different frequency characteristics, block size-based criteria are insufficient to represent the reference sample characteristics. In this manuscript, switching criteria based on both the block-size and QP value are proposed to improve the coding performance. The experimental results show a -0.45% BD-rate gain compared with that by the VVC test model 2 for all intra conditions under the common test condition (CTC) in JVET.
Yoshitaka Kidani, Kei Kawamura, Kyohei Unno, Sei Naito
ICIP1