Haruhisa Kato

dblp:62/3227 · DBLP profile ↗
← Back
34ranked-venue papers
15as first author
19since 2021 · last 2025
0009-0002-3556-4431ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 31 · 12 first-author · 19 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 first-authorArtificial intelligence and machine learning · 1
YearPublicationVenuePosition
2025 Chained Motion Vector Prediction for Video Coding
abstract
Merge mode has been utilized in advanced video coding standards to facilitate the efficiency of Motion vector (MV) or Block Vector (BV) coding. Merge mode constructs multiple MVs or BVs as the merge candidate list from MV/BV storage and then specifies one within the list by signaling the merge index. Despite various merge candidate derivation methods in prior arts, they do not adequately reach MV/BV, pointing to reference pictures with low quantization noise, leaving room for improved coding performance. This paper proposes a Chained Motion Vector Prediction (CMVP) as a novel merge candidate derivation. The CMVP derives new candidates by accumulating the recursively traced MVs or BVs based on pre-derived merge candidates. Experimental results demonstrate that the proposed method achieves up to 1.01% coding gains with negligible complexity increases on version 12 of Enhanced Compression Model (ECM), a reference software for evaluating promising coding tools beyond Versatile Video Coding (VVC).
Yoshitaka Kidani, Haruhisa Kato, Kei Kawamura
ICASSP2
2025 Multi-Res-3DGS: Multi-Resolution 3d Gaussian Splatting Bound with a Subdivided Mesh Sequence
abstract
Multi-resolution functionality is a powerful tool in video content delivery services, ensuring compatibility across various devices and optimizing bandwidth usage. However, 3D Gaussian Splatting (3DGS), an explicit radiance field for new view synthesis and efficient rendering, currently lacks such functionality. This paper introduces multi-resolution functionality for 3DGS for the first time. Recognizing that the number of 3D Gaussians is a critical factor influencing resource consumption in 3DGS rendering, we propose a method to control this number across various resolutions. We propose a novel pyramidal data structure called Multi-Res-3DGS, which binds 3D Gaussians with a sequence of subdivided meshes. In this framework, the pyramidal meshes guide the multi-resolution through iterative subdivision, while the 3D Gaussians encode details associated with the mesh faces. The Multi-Res-3DGS can be trained using a combination of existing techniques and rendered at the best resolution according to the available resources. Evaluations using the NeRF-Synthetic dataset demonstrate that our approach realizes the multi-resolution functionality with minimal PSNR or LPIPS losses.
Haruhisa Kato, Kei Kawamura
ICIP2
2025 Block Vector based Intra Prediction Mode Derivation for Beyond VVC
Haruhisa Kato, Yoshitaka Kidani, Kei Kawamura
PCS1
2024 Extended Multiple Cross-Component Linear Models With Adaptive Thresholding and Overlapped Averaging Beyond VVC
abstract
In this paper, we propose an extended multimodel cross-component linear model (MMLM) for video compression beyond the versatile video coding standard. Our proposed method incorporates adaptive thresholding and overlapped averaging to enhance prediction accuracy and reduce discontinuities in multiple linear models. We evaluate our method’s coding gain on various video sequences and demonstrate a notable improvement of up to 1.3% bit-rate savings over the conventional MMLM, validating our method’s efficiency in high-efficiency video compression.
Haruhisa Kato, Yoshitaka Kidani, Kei Kawamura
ICIP1
2024 Bi-Predictive Intra Block Copy for Enhanced Video Coding Beyond VVC
abstract
Intra block copy (IBC), an intra coding tool with a single block vector (BV), has been exploited for significant coding gains of screen content (SC) in advanced video coding standards such as VVC. Several studies have applied IBC to camera-captured content (CC), such as the IBC with fractional-sample-precision BV, which was adopted into the reference software for exploring beyond VVC, i.e., the enhanced compression model (ECM). However, there is room to further achieve the coding gains of IBC because all the conventional methods are uni-predictive IBC with a single BV to generate prediction samples. This paper proposes a bi-predictive IBC using two BVs as a new IBC algorithm for CC and SC, realized by extending the number of BVs in BV storage. In addition, this paper proposes encoder early terminations of applying IBC for CC by comparing coefficients and distortions of the IBC and intra prediction to avoid encoder runtime increases while maintaining coding gains. Experimental results show that the proposed method brings $0.15 \%$ and $0.30 \%$ coding gains for CC and SC over ECM-9 under all-intra configuration, with negligible complexity increases. The proposed method has been adopted into ECM-10.
Yoshitaka Kidani, Haruhisa Kato, Kei Kawamura
ICIP2
2024 Minimization of Submesh Boundary Errors In Dynamic Mesh Coding
abstract
The video-based dynamic mesh coding (V-DMC) standard is a cutting-edge technology for the compression of dynamic mesh data. V-DMC enables parallel encoding and partial decoding by introducing submesh frameworks in which dynamic meshes are separated and independently processed. However, V-DMC may raise submesh boundary errors like holes due to misaligning the existence or coordinates of vertices, degrading the objective and subjective qualities of decoded dynamic meshes. To minimize the boundary errors and improve the coding performance, we propose a two-stage boundary error correction method in V-DMC’s preprocessing and encoding/decoding stages. Specifically, the first stage rearranges the preprocessing order to minimize boundary errors, whereas the second stage fills holes based on boundary information. Experimental results show that the proposed method can minimize the boundary errors among the V-DMC decoded meshes, and thus significantly improve the objective and subjective quality compared to the V-DMC reference software.
Koki Kishimoto, Kei Kawamura, Haruhisa Kato
ICIP3
2024 Quantization After Inter Prediction in Displacement Coding of Dynamic Meshes
abstract
Dynamic meshes reasonably represent time-varying 3D objects, but compression is required due to the large amount of data involved. One efficient framework decomposes a dynamic mesh into a base mesh and displacements using decimation and subdivision. The displacements are converted to levels by wavelet transforms and quantization, and they are coded by arithmetic coding. The levels of the current frame are predicted from the reference frame, and only the residuals are coded. However, quantization errors occur two times in the reference frame and the current frame since the coefficients of each frame are quantized before performing inter prediction. In this paper, we propose a method of quantizing the residuals obtained after applying inter prediction in order to reduce the amount of required data. The experimental results show that the proposed method yields improved coding efficiency and that the reconstructed mesh has no quality degradations.
Hitoshi Nishimura, Haruhisa Kato, Kei Kawamura
ICIP2
2024 Temporal Scalable Coding For Dynamic Meshes
abstract
This paper presents the first implementation of temporal scalability in the ongoing standard for Video-based Dynamic Mesh Coding (V-DMC), a crucial enhancement that enables bitstream adaptation to diverse network conditions and device capabilities. While displacement and texture, two of the V-DMC’s sub-bitstreams, already benefit from existing video codec temporal scalability, the non-video basemesh sub-bitstream lacks this feature. To address this gap, we propose an adaptive coding structure designed for the basemesh. Moreover, we propose a novel cost function to adaptively select the frame type between intra-frame and inter-frame in this coding structure. Our experimental results demonstrate significant improvements in coding efficiency compared to the original V-DMC, i.e., the total BD-rates of D1, D2, Luma, Cb, and Cr averaged across all eight test sequences are -15.1%, -15.0%, 0.3%,-9.9%, and -8.4 %, respectively.
Haruhisa Kato, Kei Kawamura
ICIP2
2024 Low-complexity learning-based intra prediction with direction-dependent adaptive weights for beyond VVC
abstract
This paper introduces an advanced intra prediction method designed for the Enhanced Compression Model (ECM), which is the reference software for beyond versatile video coding (VVC) standard. It employs a learning-based method to adaptively assign weights for a weighted average across neighboring samples, resulting in more precise prediction samples. The proposed method derives optimized weights for each intra prediction mode, for each block size, and for each sample position. To achieve a reasonable balance between encoding time and prediction accuracy, the conventional intra prediction mode is shared with the proposed method. Experimental evaluations have demonstrated that the proposed method provides bitrate reduction of up to 0.4%.
Haruhisa Kato, Yoshitaka Kidani, Kei Kawamura
VCIP1
2024 Inter Submesh Border Information Coding with Skip Mode in V-DMC
abstract
Standardization of Dynamic Mesh Coding (V-DMC) has been progressing in MPEG since 2022. The current reference software for V-DMC encodes dynamic meshes by segmenting them into regions (submeshes) and correcting holes occurring at submesh boundaries based on submesh boundary information. However, the encoding performance of submesh boundary information is low in V-DMC because it does not utilize the temporal correlation of submesh boundary information. To address this issue, we propose an inter-coding method for submesh boundary information using reference frame submesh boundary information. Experimental results show that our proposal improves coding performance compared to conventional methods.
Koki Kishimoto, Kei Kawamura, Haruhisa Kato
VCIP3
2024 A High-Efficiency and Low-Complexity SKIP Type for Base Mesh Coding in V-DMC
abstract
Video-based Dynamic Mesh Coding (V-DMC) is an emerging standard for dynamic mesh compression, where the original meshes are decimated into simplified meshes called base meshes. This paper introduces a novel SKIP type for base mesh coding in V-DMC, complementing the existing INTRA and INTER types. When the SKIP type is used in base mesh coding, it directly copies the reconstructed base mesh from the reference frame, eliminating the need for additional data coding. Thus, the reconstructed base mesh in the current frame is identical to that in the reference frame. This significantly reduces the bit rate and decoding time for base meshes. Additionally, this paper employs a Lagrangian cost function using a linear model for bit estimation of INTRA type and L1 norms for distortion approximation of SKIP type to enable the encoder to select the best type for base meshes. Experimental results demonstrate superior BD-rate performance and significantly reduced decoding time for base meshes using the SKIP type, particularly in sequences with minimal object movements.
Haruhisa Kato, Kei Kawamura
VCIP2
2023 Hierarchical Arithmetic Coding of Displacements for Dynamic Mesh Compression
abstract
Dynamic meshes reasonably represent time-varying 3D objects, but compression is required due to the large amount of data. One compression framework decomposes a dynamic mesh into a base mesh and displacements by using decimation and subdivision. The displacements are converted to coefficients by wavelet transforms, quantized, and compressed by video codec, which is well disseminated. However, the abundance of tools in video codec is too complex for uncorrelated displacements. In this paper, we propose hierarchical arithmetic coding, dividing the coefficient levels into blocks and smaller subblocks. When all levels are zero in a block/subblock, a flag is coded instead of the levels. The experimental results show that the coding complexity was significantly reduced while the coding efficiency was maintained.
Hitoshi Nishimura, Haruhisa Kato, Kei Kawamura
ICIP2
2023 Extended Intra Block Copy with Adaptive Filtering and Overlapped Block Averaging
abstract
Next-generation video coding standards are attempting to improve coding performance compared to conventional standards such as VVC by extending technologies such as intra-block copy (IBC). While IBC in VVC has proven effective for screen content, its adaptation to camera-captured content presents challenges regarding sample fluctuations and the continuity of block boundaries. This paper proposes a novel approach to improve IBC performance for camera-captured content by combining adaptive filtering (F-IBC) and overlapped block averaging (OB-IBC). The F-IBC filters IBC prediction samples using filter coefficients derived from adjacent samples to predict sample fluctuations accurately. The OB-IBC is a weighted average of IBC prediction samples of the current block with adjacent samples of the adjacent block’s reference to connect block boundaries smoothly. Following common test conditions in the joint video experts team, experimental results show improved coding performance with a bitrate saving of 0.1 % over the reference software (ECM 7.0) which investigates the enhanced compression beyond VVC capability.
Haruhisa Kato, Yoshitaka Kidani, Kei Kawamura, Sei Naito
VCIP1
2023 1D displacement coding for the displaced subdivision surface
abstract
Compression of a dynamic mesh, which represents the dynamic volumetric data for immersive applications, is an emerging technique. One advanced compression technique is video-based dynamic mesh coding (V-DMC). Therein, an original mesh is decomposed into a decimated base mesh including displacement vectors in the V-DMC framework. Because the mesh represents an object in 3D space, these displacements, which express the detailed information of the mesh, are also represented as 3D vectors over the subdivision surface. However, such representation of a 3D vector is redundant because the displacement direction is almost equal to a normal direction of the subdivision surface. Hence, we present a 1D displacement coding that operates with an existing video codec. This method improves both the encoding/decoding procedure complexity and the coding performance on dynamic mesh coding when compared to competing approaches.
Koki Kishimoto, Kei Kawamura, Haruhisa Kato
VCIP3
2023 Arithmetic Coding of Displacements in Dynamic Meshes with Bypass Mode for Complexity Reduction
abstract
Dynamic meshes reasonably represent time-varying 3D objects, but compression is required due to the large amount of data. One efficient framework decomposes a dynamic mesh into a base mesh and displacements using decimation and subdivision. The displacements are converted to levels by wavelet transforms and quantization, and the levels are coded by block-based hierarchical arithmetic coding. However, the coding complexity is high in the worst case where all coefficients are encoded. In this paper, we propose arithmetic coding of levels with a bypass mode, which has low complexity by skipping context updates. The experimental results show that the coding complexity in the worst case was reduced while coding efficiency was maintained.
Hitoshi Nishimura, Haruhisa Kato, Kei Kawamura
VCIP2
2022 Graph-Based Point Cloud Denoising Using Shape-Aware Consistency For Free-Viewpoint Video
abstract
We propose a novel graph-based denoising method to correct the quantization error (step noise) arising in the process of generating the visual hull, a commonly used technique to synthesize free-viewpoint video. To reduce this step noise effectively, we propose two new notions of consistency, pixel value consistency and normal vector consistency. The resulting denoising method involves a first step of graph construction using the proposed consistency metrics, followed by graph filtering of the 3D point cloud coordinates. Our experiments show that our approach provides visually and quantitatively better performance than state-of-the-art methods.
Keisuke Nonaka, Ryosuke Watanabe, Haruhisa Kato, Tatsuya Kobayashi, Eduardo Pavez, Antonio Ortega
ICASSP3
2022 Point Cloud Denoising Using Normal Vector-Based Graph Wavelet Shrinkage
abstract
Many applications that use point clouds, such as 3D immersive telepresence, suffer from geometric quality degradation. This noise may be caused by measurement errors of the capturing device or by the point cloud estimation method. In this paper, we propose a novel graph-based point cloud denoising approach using the spectral graph wavelet transform (SGWT) and graph wavelet shrinkage. Unlike conventional SGWT-based denoising methods, the proposed wavelet shrinkage thresholds are determined based on the normal vector at each point and are thus based on the local geometric structure of the point cloud. This approach avoids excessive wavelet shrinkage, which can lead to the loss of complex geometric structure. Experimental results show that the proposed method achieves the best accuracy as compared with recent deep-learning-based and graph-based state-of-the-art denoising methods.
Ryosuke Watanabe, Keisuke Nonaka, Haruhisa Kato, Eduardo Pavez, Tatsuya Kobayashi, Antonio Ortega
ICASSP3
2022 Adaptive boundary width of Geometric Partitioning Mode for Beyond Versatile Video Coding
abstract
In order to improve coding efficiency beyond versatile video coding (VVC), we propose an extended geometric partitioning mode (GPM). GPM is a new inter prediction in VVC and is applied to the object boundary between the foreground and background with different motions. Specifically, GPM partitions a rectangular coding block into two regions with 64 predefined types of straight lines, generates inter prediction samples for each partitioned region and then blends them with a fixed boundary width to obtain the final prediction samples. However, the fixed boundary width of GPM is not always optimal for diverse video content. To solve this problem, the proposed method allows GPM to select multiple boundary widths by block-wise signaling. Furthermore, the proposed method also restricts the selectable boundary width according to the short side of the block to reduce the encoding time for selecting the optimal width. Experiment results following common test conditions in JVET showed an improvement in coding efficiency with bitrate savings of 0.11 % and 3.20 % for camera-captured content and for pure screen or video game content, respectively, compared VVC reference software.
Haruhisa Kato, Yoshitaka Kidani, Kei Kawamura, Sei Naito
VCIP1
2021 Split Rendering of the Transparent Channel for Cloud AR
abstract
We are the first to apply split rendering to AR to improve the quality of the transparent channel. The proposed method evaluates a cloud-based AR streaming system that offloads photorealistic rendering to a cloud server and splits the rendering between the cloud server and the mobile device (split rendering). Server-side rendering is capable of rendering photorealistic images, but the quality of the image is degraded by coding when the video is compressed and transmitted. In particular, degradation of the transparent channel significantly reduces the subjective image quality. The proposed method avoids this degradation by rendering the transparent channel at the mobile terminal. In addition, the server improves the coding efficiency by padding the transparent areas. Nonlinear quantization of the difference images contributes to improved subjective image quality. Experimentally, we confirmed that the proposed method has a SSIM gain of 0.022 compared to the conventional method.
Haruhisa Kato, Tatsuya Kobayashi, Masaru Sugano, Sei Naito
MMSP1
2016 Planar Markerless Augmented Reality Using Online Orientation Estimation
Tatsuya Kobayashi, Haruhisa Kato, Masaru Sugano
ACCV (4)2
2014 An adaptive residual decorrelation method for HEVC
abstract
In this paper, we propose an explicit residual decorrelation method to improve the coding performance for 4:4:4 chroma format conforming HEVC framework. The energy from a residual signal is gathered to the primary component by decorrelation of color space. The transform matrix as decorrelation is derived from reference pixel value by using singular value decomposition for each prediction unit. Since the derivation is applied in both encoder and decoder side, the identical matrix is obtained for both sides. Compared to the previous works, the proposed method applies only for the meaningful unit while an enabled flag is explicitly signaled as side information. The proposed method is implemented on HEVC test model. For the RGB/YUV 4:4:4 chroma format sequences, the coding gains in BD-rate are up to 23.3%/4.9%, respectively. Compared to the result by the previous works, average gains are slightly decreased, while each gain of sequence is always better than that by conventional method.
Kei Kawamura, Haruhisa Kato, Sei Naito
ICIP2
2014 3D-Ferns+: Viewpoint-based keypoint classifier for robust 3D object pose detection
abstract
We present a novel pose detection method that can be used in mobile augmented reality (AR) services. Making 3D object pose detection robust against changes in viewpoint is a vitally important but quite difficult task because 3D objects often change their appearance significantly with changes in viewpoint, and the possible range of viewpoints is wide compared with planar targets. 3D-Ferns, which is a keypoint classifier for 3D object pose detection, performs direct 2D-3D matching and handles a wide range of detectable viewpoints, including all rotations. However, many difficult viewpoints still exist for pose detection because of the unevenness of matching performance over all viewpoints. In this paper, we propose a novel class selection strategy that evens out matching performance over all possible viewpoints and improves detection performance from difficult viewpoints by focusing on the per-viewpoint repeatability (PVR) of class 3D points. Experimental results demonstrate the impact of stability of 2D-3D matching on detection performance and the effect of our method, which reduces the detection failures in conventional approaches by over 23% for 3D targets that have various shapes and textures.
Tatsuya Kobayashi, Haruhisa Kato, Hiromasa Yanagihara
ICIP2
2013 Vision-based robust calibration for optical see-through head-mounted displays
abstract
We propose Vision-based Robust Calibration (ViRC) method for OSTHMDs equipped with a camera. In the ViRC method, calibration parameters are decomposed into off-line parameters that remain constant relative to the positional relationship between the camera and the virtual screen, and on-line parameters related to the user's eye. Calculating the off-line parameters beforehand reduces the number of unknown parameters in the on-line phase, giving robust protection against the user's misalignments during calibration. In the off-line phase, the approximate position of the user's eye is calculated using the PnP algorithm. In the online phase, the actual position of the user's eye is estimated from the approximate one by non-linear minimization. In our experiments, we show that the ViRC method can decrease reprojection error by as much as 83% compared with the conventional method based on the DLT algorithm.
Naoya Makibuchi, Haruhisa Kato, Akio Yoneyama
ICIP2
2013 PACMAN UI: vision-based finger detection for positioning and clicking manipulations
abstract
This paper proposes an intuitive input interface that can handle various operations based on finger image recognition. It receives continuous analog input by detecting a knuckle of the user's clenched fist. In contrast to the conventional wireless mouse, whose sensitivity cannot be changed dynamically, the proposed method brings not only stable positioning but also quick clicking with a small finger gesture. In order to evaluate operability, we conducted a user experiment: a time trial for target selection. The subjects completed the task with the proposed controller in 44% less time than with a conventional wireless mouse. We confirmed that the proposed method can reliably follow finger gestures.
Haruhisa Kato, Hiromasa Yanagihara
Mobile HCI1
2013 In-loop colour-space-transform coding based on integered SVD for HEVC range extensions
abstract
Inter colour-component correlation is generally very high in RGB 4:4:4 chroma format. To improve the coding performance of the high efficiency video coding (HEVC) especially for such content, we propose the in-loop colour-space-transform. The colour space is dynamically transformed into un-correlated space by employing singular value decomposition (SVD) for each block at both the encoder and decoder. Signals in transformed colour space are coded with the existing intra / inter coding framework. We utilize the simplified SVD process implemented only by integer operations for the complexity reduction. Compared with HM10.0 as an anchor method, BD-bitrate gain reached 23.8% and 23.4% for the all intra case and the random access case, respectively, while a runtime of the decoder increase 4.8-9.8%.
Kei Kawamura, Haruhisa Kato, Sei Naito
PCS2
2012 A line-based palm-top detector for mobile augmented reality
abstract
We propose a marker-less Augmented Reality (AR) application based on a realtime hand posture estimation technique for smartphones. A conventional marker-less AR system does not have sufficient accuracy and speed in the detection of a mobile device. This paper presents a fast hand posture estimation algorithm based on a combination of feature points and feature lines that consist of the boundary of fingers. The proposed method realized rendering of virtual 3D models on a hand over 12 frames per second (fps) on a smart-phone. Simulation results show that we can archive about 73% complexity reductions and be more accurate than the conventional method.
Haruhisa Kato, Akio Yoneyama
AVI1
2012 Asymmetric partitioning with non-power-of-two transform for intra coding
abstract
HEVC (High Efficiency Video Coding) is an ongoing standardization target as the next generation of video compression technology. HEVC employs a coding tree block, which is a quad-tree structure of a coding unit. It also employs some unit types; coding unit, prediction unit, and transform unit. A coding unit can be divided into smaller units as prediction units. Though an asymmetric unit is used for inter coding, only symmetric units are permitted for intra coding. In this paper, we propose an asymmetric partitioning with a non-power-of-two transform as a prediction and transform unit. While conventional partitioning locates the cross-point of partitioning lines at the center of the coding unit, the proposed method locates the cross-point in places except center. The proposed method reduces 2.0% BD-bitrate compared with HM5.0 under all intra / high efficiency condition. The validity of the proposed method is confirmed by some experimental results.
Kei Kawamura, Haruhisa Kato, Sei Naito
PCS2
2009 A camera-based tangible controller for cellular phones
abstract
This paper proposes a novel easy-to-use camera-based tangible controller for cellular phone applications. It realizes continuous analog input by tracking a marker at the top end of a controller device attached to the embedded camera. In contrast to the conventional keypad which enables limited operability to four discrete directions, the proposed controller brings not only an unconstrained continuous input to arbitrary directions but also continuous input for depth and for a rotation angle. In order to evaluate operability, we conducted a user experiment of time trial for path tracing, and the results showed that the subjects completed the task with the proposed controller in 27.9% less than with the conventional keypad input.
Haruhisa Kato, Tsuneo Kato
Mobile HCI1
2007 Coding Mode Decision for High Quality MPEG-2 to H.264 Transcoding
abstract
This paper proposes a high quality coding mode decision method for MPEG-2 to H.264 transcoding. Here, we evaluate the coding modes of MPEG-2 stream in order to specify whether they should be inherited or changed in H.264 coding since many new coding schemes such as intra prediction have been introduced in H.264. By adaptively determining suitable intra/inter modes according to DCT coefficients, motion vectors (MVs), and neighboring macroblock (MB) modes of input MPEG-2 stream, high quality transcoding has been realized. In the experiment, we observed that the proposed method can improve a peak signal to noise ratio (PSNR) up to 0.29 dB at the cost of 22% more processing time compared with the conventional transcoding method.
Haruhisa Kato, Akio Yoneyama, Yasuhiro Takishima, Yosuke Kaji
ICIP (4)1
2007 A Fast DV to MPEG-4 Transcoder Integrated With Resolution Conversion and Quantization
abstract
We propose a fast transcoder from digital video (DV) to MPEG-4 in the coded domain. Since DV is interlaced sequence whereas MPEG-4 (SIF) is progressive sequence and different discrete cosine transform (DCT) mode $({\hbox {2}}\ast{\hbox {4}}\ast{\hbox {8}}{\hbox {DCT}})$ is used in DV, different compressed domain transcoding method from that of MPEG to MPEG conversion is required. We have exploited matrix conversion reflecting these properties and introduce approximation and integration of resolution conversion and quantization process. Simulation results of DV to MPEG-4 conversion show that the proposed method can achieve very fast conversion while maintaining high transcoding performance when compared with base-band transcoding and siginificant improvement over conventional method is also realized.
Haruhisa Kato, Yasuhiro Takishima, Yasuyuki Nakajima
IEEE Trans. Circuits Syst. Video Technol.1
2006 Fast Intra Mode Decision Method for MPEG to H.264 Transcoding
abstract
This paper proposes a fast MPEG-2 to H.264 intra frame transcoding method. It deploys a novel technology in the refering process of MPEG-2 discrete cosine transform (DCT) coefficients to determine the block size and the prediction mode of intra frame prediction for H.264. The proposed method reduces the coding complexity by approximately 60% while maintaining a peak signal to noise ratio (PSNR) compared with a typical baseband conversion.
Haruhisa Kato, Yasuhiro Takishima, Yosuke Kaji
ICIP1
2005 Resolution conversion integrated with quantization in coded domain for DV to MPEG-4 transcoder
abstract
We propose a new algorithm for the fast conversion of MPEG-4 from DV. DV, which is a video compression coding format, is adopted by commercial and non-commercial digital video cameras and achieves high video quality using a fixed high bitrate of 25 Mbps. However, compared with MPEG-4, which is used extensively in IP or mobile phone applications, DV has a higher storage and transmission cost. This is due to the extremely low coding efficiency of DV required to maintain high quality for frame-accurate editing. In this paper, we analyze the difference between DV and MPEG-4, and propose a new algorithm for the fast conversion of MPEG-4 from DV. Our algorithm accelerates the conversion process in the coded domain using the property of conversion matrices, where resolution conversion, quantization, and inverse quantization are integrated. The experimental results indicate that, while the integration and approximation of a conversion matrix do not affect the video quality, this could greatly reduce the number of operands.
Haruhisa Kato, Akio Yoneyama, Yasuhiro Takishima
ICIP (3)1
2004 Integrated compressed domain resolution conversion with de-interlacing for DV to MPEG-4 transcoding
Haruhisa Kato, Yasuyuki Nakajima, Takashi Sano 0001
ICIP1
2004 Weighting factor determination algorithm for H.264/MPEG-4 AVC weighted prediction
abstract
H.264/MPEG-4 AVC (ISO/IEC 14496-10) adopted the weighted prediction which is particularly useful for coding the fade/dissolve transition scenes. This paper proposes a determination method of weighting factors for the weighted prediction. The simulation results show that over 30% bitrate savings can be achieved when compared with the conventional methods.
Haruhisa Kato, Yasuyuki Nakajima
MMSP1