VLDB 2026 Research / reviewers in the wild / expert
Sei Naito
dblp:12/993
· DBLP profile ↗
41ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0002-6104-4344ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 39 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 2Systems, architecture and hardware · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | No-Reference Textured Mesh Quality Assessment Using Graph-Based Featuresabstract3D mesh models with texture maps play an important role in various 3D applications. Since No-Reference Textured Mesh Quality Assessment (NR-TMQA) methods are essential to identify low-quality mesh models and enhance their quality, this paper presents an accurate NR-TMQA method. Our key contributions are as follows: 1) We propose a novel NR-TMQA scheme that leverages both point cloud and rendered image representations from distorted mesh models. 2) We introduce a novel point cloud representation derived from the faces of a 3D mesh model. It is referred to as a face-based point cloud, which effectively captures face-level features. 3) We present novel graph-based features to estimate mesh quality using the face-based point cloud. Our experimental results demonstrate that Pearson’s linear correlation coefficient and Spearman’s rank-order correlation coefficient in two large-scale TMQA datasets improved by 0.148 (from 0.572 to 0.720) and 0.18 (from 0.525 to 0.705), respectively, compared to the conventional state-of-the-art TMQA method. Ryosuke Watanabe, Hanif Fermanda Putra, Tomoaki Konno, Sei Naito |
ICIP | 4 |
| 2023 | Extended Intra Block Copy with Adaptive Filtering and Overlapped Block AveragingabstractNext-generation video coding standards are attempting to improve coding performance compared to conventional standards such as VVC by extending technologies such as intra-block copy (IBC). While IBC in VVC has proven effective for screen content, its adaptation to camera-captured content presents challenges regarding sample fluctuations and the continuity of block boundaries. This paper proposes a novel approach to improve IBC performance for camera-captured content by combining adaptive filtering (F-IBC) and overlapped block averaging (OB-IBC). The F-IBC filters IBC prediction samples using filter coefficients derived from adjacent samples to predict sample fluctuations accurately. The OB-IBC is a weighted average of IBC prediction samples of the current block with adjacent samples of the adjacent block’s reference to connect block boundaries smoothly. Following common test conditions in the joint video experts team, experimental results show improved coding performance with a bitrate saving of 0.1 % over the reference software (ECM 7.0) which investigates the enhanced compression beyond VVC capability. Haruhisa Kato, Yoshitaka Kidani, Kei Kawamura, Sei Naito |
VCIP | 4 |
| 2022 | Adaptive boundary width of Geometric Partitioning Mode for Beyond Versatile Video CodingabstractIn order to improve coding efficiency beyond versatile video coding (VVC), we propose an extended geometric partitioning mode (GPM). GPM is a new inter prediction in VVC and is applied to the object boundary between the foreground and background with different motions. Specifically, GPM partitions a rectangular coding block into two regions with 64 predefined types of straight lines, generates inter prediction samples for each partitioned region and then blends them with a fixed boundary width to obtain the final prediction samples. However, the fixed boundary width of GPM is not always optimal for diverse video content. To solve this problem, the proposed method allows GPM to select multiple boundary widths by block-wise signaling. Furthermore, the proposed method also restricts the selectable boundary width according to the short side of the block to reduce the encoding time for selecting the optimal width. Experiment results following common test conditions in JVET showed an improvement in coding efficiency with bitrate savings of 0.11 % and 3.20 % for camera-captured content and for pure screen or video game content, respectively, compared VVC reference software. Haruhisa Kato, Yoshitaka Kidani, Kei Kawamura, Sei Naito |
VCIP | 4 |
| 2021 | Facial Action Unit-based Deep Learning Framework for Spotting Macro- and Micro-expressions in Long Video SequencesabstractIn this paper, we utilize facial action units (AUs) detection to construct an end-to-end deep learning framework for the macro- and micro-expressions spotting task in long video sequences. The proposed framework focuses on individual components of facial muscle movement rather than processing the whole image, which eliminates the influence of image change caused by noises, such as body or head movement. Compared with existing models deploying deep learning methods with classical Convolutional Neural Network (CNN) models, the proposed framework utilizes Gated Recurrent Unit (GRU) or Long Short-term Memory (LSTM) or our proposed Concat-CNN models to learn the characteristic correlation between AUs of distinctive frames. The Concat-CNN uses three convolutional kernels with different sizes to observe features of different duration and emphasizes both local and global mutation features by changing dimensionality (max-pooling size) of the output space. Our proposal achieves state-of-the-art performance from the aspect of overall F1-scores: 0.2019 on CAS(ME)2-cropped, 0.2736 on SAMM Long Video, and 0.2118 on CAS(ME)2, which not only outperforms the baseline but is also ranked the 3rd of FME challenge 2021 for combined datasets of CAS(ME)2-cropped and SAMM-LV. Zhiguang Zhou, Megumi Komiya, Koki Kishimoto, Keisuke Nonaka, Toshiharu Horiuchi, Satoshi Komorita, Gen Hattori, Sei Naito, Yasuhiro Takishima |
ACM Multimedia | 11 |
| 2021 | Split Rendering of the Transparent Channel for Cloud ARabstractWe are the first to apply split rendering to AR to improve the quality of the transparent channel. The proposed method evaluates a cloud-based AR streaming system that offloads photorealistic rendering to a cloud server and splits the rendering between the cloud server and the mobile device (split rendering). Server-side rendering is capable of rendering photorealistic images, but the quality of the image is degraded by coding when the video is compressed and transmitted. In particular, degradation of the transparent channel significantly reduces the subjective image quality. The proposed method avoids this degradation by rendering the transparent channel at the mobile terminal. In addition, the server improves the coding efficiency by padding the transparent areas. Nonlinear quantization of the difference images contributes to improved subjective image quality. Experimentally, we confirmed that the proposed method has a SSIM gain of 0.022 compared to the conventional method. Haruhisa Kato, Tatsuya Kobayashi, Masaru Sugano, Sei Naito |
MMSP | 4 |
| 2020 | Block-Size Dependent Overlapped Block Motion CompensationabstractOverlapped block motion compensation (OBMC) is one of the inter prediction tools that improves coding performance. OBMC applied to various non-squared blocks has been studied in VVC, which is being standardized by joint video experts team (JVET), to improve coding performance over HEVC. Memory bandwidth, however, is a bottleneck when OBMC is used, and conventional methods have not achieved a good trade-off regarding coding performance and memory bandwidth so far. In this study, interpolation filters and applicable conditions of OBMC depending on block sizes are proposed to achieve the best trade-off. The experimental results show a -0.40% BD-rate gain compared with that of the VVC test model 3 for random access conditions under the common test condition in JVET. Yoshitaka Kidani, Kei Kawamura, Kyohei Unno, Sei Naito |
ICIP | 4 |
| 2020 | Lossless Video Coding Based On Probability Model Optimization Utilizing Example Search And Adaptive PredictionabstractWe previously proposed a novel lossless coding method that utilizes example search and adaptive prediction within a framework of probability model optimization for still images. Additionally, we also proposed a lossless video coding method where the example search is performed on not only the current but also the previous frames to exploit intra- and inter-frame correlations. In this paper, we integrate these two methods for efficient lossless video coding. Moreover, we extend the adaptive prediction to exploit both spatial and temporal correlations simultaneously. In other words, reference pels used for the prediction are taken from both the current and the motion-compensated previous frames, and their weights, i.e. prediction coefficients, are trained pel-by-pel in a weighted least square manner. The experimental results show that the proposed method achieves better coding performance than the VVC-based lossless video coding scheme. Kyohei Unno, Koji Nemoto, Yusuke Kameda, Ichiro Matsuda, Susumu Itoh, Sei Naito |
ICIP | 6 |
| 2020 | Accurate Background Subtraction Using Dynamic Object Presence Probability in Sports ScenesabstractForeground segmentation technologies play an important role in applications such as free-viewpoint video (FVV) and sports video analysis. In this situation, we propose a new method that achieves accurate foreground silhouette extraction method using dynamic object presence probability (DOPP). Our main contributions are as follows. 1) Object presence probability for each pixel is calculated from the object recognition results based on deep learning. After that, background subtraction is implemented by changing the threshold and the update rate of the background model in response to the object presence probability. Parameter tuning of background subtraction is executed by using the object recognition results to improve the silhouette extraction quality. 2) To calculate more accurate silhouette images, parameters of background subtraction are adjusted by monitoring optical flows between consecutive frames. The object presence probability of the current frame is dynamically updated by using the object presence probability of the previous frame with optical flows. In the experiments, we confirmed that the proposed method achieved more accurate silhouette extraction than conventional methods in three sports sequences. Ryosuke Watanabe, Jun Chen 0017, Tomoaki Konno, Sei Naito |
ICPR | 4 |
| 2019 | Blocksize-QP Dependent Intra Interpolation FiltersabstractIntra interpolation filters for intra angular prediction play an important role in the coding performance. In the intra angular prediction of VVC, which is being standardized by the joint video coding expert team (JVET), block-size based switchable interpolation filters between 4-tap cubic and Gaussian interpolation filters is being studied. Although the two filters have different frequency characteristics, block size-based criteria are insufficient to represent the reference sample characteristics. In this manuscript, switching criteria based on both the block-size and QP value are proposed to improve the coding performance. The experimental results show a -0.45% BD-rate gain compared with that by the VVC test model 2 for all intra conditions under the common test condition (CTC) in JVET. Yoshitaka Kidani, Kei Kawamura, Kyohei Unno, Sei Naito |
ICIP | 4 |
| 2019 | Fast Free-viewpoint Video Synthesis Algorithm for Sports ScenesabstractIn this paper, we report on a parallel free-viewpoint video synthesis algorithm that can efficiently reconstruct a high-quality 3D scene representation of sports scenes. The proposed method focuses on a scene that is captured by multiple synchronized cameras featuring wide-baselines. The following strategies are introduced to accelerate the production of a free-viewpoint video taking the improvement of visual quality into account: (1) a sparse point cloud is reconstructed using a volumetric visual hull approach, and an exact 3D ROI is found for each object using an efficient connected components labeling algorithm. Next, the reconstruction of a dense point cloud is accelerated by implementing visual hull only in the ROIs; (2) an accurate polyhedral surface mesh is built by estimating the exact intersections between grid cells and the visual hull; (3) the appearance of the reconstructed presentation is reproduced in a view-dependent manner that respectively renders the non-occluded and occluded region with the nearest camera and its neighboring cameras. The production for volleyball and judo sequences demonstrates the effectiveness of our method in terms of both execution time and visual quality. Jun Chen 0017, Ryosuke Watanabe, Keisuke Nonaka, Tomoaki Konno, Hiroshi Sankoh, Sei Naito |
IROS | 6 |
| 2018 | Robust Billboard-based, Free-viewpoint Video Synthesis Algorithm to Overcome Occlusions under Challenging Outdoor Sport ScenesabstractThe paper proposes an algorithm to robustly reconstruct an accurate billboard model of an individual object including an occluded one in each camera. Each billboard model is utilized to synthesize high-quality, free-viewpoint video especially for outdoor sport scenes in which roughly calibrated cameras are sparsely placed. The two main contributions of the proposed algorithm are (1) robustness to occlusions caused by overlaps of multiple objects in every camera, that is one of the biggest issues for billboard-based method, and (2) applicability to challenging shooting conditions in which accurate 3D model cannot be reconstructed because of calibration errors, small number of cameras and so on. In order to achieve the contributions above, the algorithm does not try to reproduce an accurate 3D model of each object but utilize a "rough 3D model". The algorithm precisely extracts an individual object region in every camera by reconstructing a "rough 3D model" of each object and back-projecting it to every camera. The 3D coordinate for each billboard to be located is calculated based on the position of a rough 3D model. Experimental results compare the visual quality of free-viewpoint videos synthesized with our proposed method and conventional methods and show the effectiveness of our proposed method in terms of the naturalness of positional relationships and the fineness of the surface textures of all the objects. Hiroshi Sankoh, Sei Naito, Keisuke Nonaka, M. S. Houari Sabirin, Jun Chen 0017 |
ACM Multimedia | 2 |
| 2018 | Fast Plane-Based Free-viewpoint Synthesis for Real-time Live StreamingabstractFree viewpoint technologies that synthesize virtual viewpoint by using multiple actual videos are one of the hottest topics in the video processing field, and would provide immersive experiences for users. To realize this concept, many conventional methods have been proposed. However, these methods require high computational cost to synthesize a virtual viewpoint because they have to calculate huge data to express three-dimensional information, e.g. the shapes of objects, only by using two-dimensional video. This makes them inadequate for end-to-end (from video capture to virtual view rendering) live streaming. To overcome this problem, we propose a simple and fast free-viewpoint synthesis method based on a visual hull, which is a general concept in this field. We calculate the silhouette of an object along planes in virtual space by simple projection from video images to the 3D space, and express the whole shape of the object by integrating the planes. This scheme works very quickly while providing fine quality because it consists of standard functions in general GPU architecture. The experimental results show our method can generate a fine virtual view of an object by using multiple videos in real time. Keisuke Nonaka, Ryosuke Watanabe, Jun Chen 0017, M. S. Houari Sabirin, Sei Naito |
VCIP | 5 |
| 2017 | Fast camera self-calibration for synthesizing Free Viewpoint soccer VideoabstractRecently, non-fixed camera-based free viewpoint sports video synthesis has become very popular. Camera calibration is an indispensable step in free viewpoint video synthesis, and the calibration has to be done frame by frame for a non-fixed camera. Thus, calibration speed is of great significance in real-time application. In this paper, a fast self-calibration method for a non-fixed camera is proposed to estimate the homography matrix between a camera image and a soccer field model. As far as we know, it is the first time to propose constructing feature vectors by analyzing crossing points of field lines in both camera image and field model. Therefore, different from previous methods that evaluate all the possible homography matrices and select the best one, our proposed method only evaluates a small number of homography matrices based on the matching result of the constructed feature vectors. Experimental results show that the proposed method is much faster than other methods with only a slight loss of calibration accuracy that is negligible in final synthesized videos. Akira Kubota, Kaoru Kawakita, Keisuke Nonaka, Hiroshi Sankoh, Sei Naito |
ICASSP | 6 |
| 2016 | Robust moving camera calibration for synthesizing free viewpoint soccer videoabstractIn this paper, a robust moving camera calibration method is proposed in order to synthesize a free viewpoint soccer video with a high degree of accuracy. The main problem in video registration-based moving camera calibration is that the calibration accuracy is very low if the detected feature points are from moving objects. In order to solve this problem, the proposed method tracks the feature points along video frames to construct a trajectory matrix of feature points, and the trajectory matrix is decomposed into a low-rank matrix representing global camera motion and a sparse matrix representing local individual motion. Therefore, according to such decomposition, the individual motions of dynamic feature points that are from moving objects could be suppressed and removed. Experimental results show that the proposed method achieves more accurate calibration result and the visual quality of a synthesized free viewpoint soccer video is also improved by the proposed method. Keisuke Nonaka, Hiroshi Sankoh, Sei Naito |
ICIP | 4 |
| 2016 | Automatic camera self-calibration for immersive navigation of free viewpoint sports videoabstractIn recent years, the demand of immersive experience has triggered a great revolution in the applications and formats of multimedia. Particularly, immersive navigation of free viewpoint sports video has become increasingly popular, and people would like to be able to actively select different viewpoints when watching sports videos to enhance the ultra realistic experience. In the practical realization of immersive navigation of free viewpoint video, the camera calibration is of vital importance. Especially, automatic camera calibration is very significant in real-time implementation and the accuracy of camera parameter directly determines the final experience of free viewpoint navigation. In this paper, we propose an automatic camera self-calibration method based on a field model for free viewpoint navigation in sports events. The proposed method is composed of three parts, namely, extraction of field lines in a camera image, calculation of crossing points, determination of the optimal camera parameter. Experimental results show that the camera parameter can be automatically estimated by the proposed method for a fixed camera, dynamic camera and multi-view cameras with high accuracy. Furthermore, immersive free viewpoint navigation in sports events can also be completely realized based on the camera parameter estimated by the proposed method. Hiroshi Sankoh, Keisuke Nonaka, Sei Naito |
MMSP | 4 |
| 2015 | Accurate silhouette extraction of multiple moving objects for free viewpoint sports video synthesisabstractIn this paper, we propose a new method of automatic silhouette extraction of multiple moving objects with high accuracy for free viewpoint stadium sports video synthesis. The proposed method is basically composed of three parts, including a global extraction based on temporal background subtraction, a classification step based on the constraints of extracted candidates of objects, and local refinement based on the statistical information of the chrominance component of each extracted object. Experimental results show that the proposal outperforms the temporal background subtraction model and Gaussian Mixture Model (GMM) in terms of both objective and subjective evaluations. In addition, the quality of the synthesized free viewpoint sports video is also enhanced by adopting more accurate silhouettes of objects that are extracted by our proposed method. Furthermore, as there is no manual operation in the proposed method, the automatic multiple silhouettes extraction has also been fully realized. Hiroshi Sankoh, M. S. Houari Sabirin, Sei Naito |
MMSP | 4 |
| 2014 | An adaptive residual decorrelation method for HEVCabstractIn this paper, we propose an explicit residual decorrelation method to improve the coding performance for 4:4:4 chroma format conforming HEVC framework. The energy from a residual signal is gathered to the primary component by decorrelation of color space. The transform matrix as decorrelation is derived from reference pixel value by using singular value decomposition for each prediction unit. Since the derivation is applied in both encoder and decoder side, the identical matrix is obtained for both sides. Compared to the previous works, the proposed method applies only for the meaningful unit while an enabled flag is explicitly signaled as side information. The proposed method is implemented on HEVC test model. For the RGB/YUV 4:4:4 chroma format sequences, the coding gains in BD-rate are up to 23.3%/4.9%, respectively. Compared to the result by the previous works, average gains are slightly decreased, while each gain of sequence is always better than that by conventional method. Kei Kawamura, Haruhisa Kato, Sei Naito |
ICIP | 3 |
| 2014 | Joint gaze-correction and beautification of DIBR-synthesized human face via dual sparse codingabstractGaze mismatch is a common problem in video conferencing, where the viewpoint captured by a camera (usually located above or below a display monitor) is not aligned with the gaze direction of the human subject, who typically looks at his counterpart in the center of the screen. This means that the two parties cannot converse eye-to-eye, hampering the quality of visual communication. One conventional approach to the gaze mismatch problem is to synthesize a gaze-corrected face image as viewed from center of the screen via depth-image-based rendering (DIBR), assuming texture and depth maps are available at the camera-captured viewpoint(s). Due to self-occlusion, however, there will be missing pixels in the DIBR-synthesized view image that require satisfactory filling. In this paper, we propose to jointly solve the hole-filling problem and the face beautification problem (subtle modifications of facial features to enhance attractiveness of the rendered face) via a unified dual sparse coding framework. Specifically, we first train two dictionaries separately: one for face images of the intended conference subject, one for images of “beautiful” human faces. During synthesis, we simultaneously seek two code vectors - one is sparse in the first dictionary and explains the available DIBR-synthesized pixels, the other is sparse in the second dictionary and matches well with the first vector up to a restricted linear transform. This ensures a good match with the intended target face, while increasing proximity to “beautiful” facial features to improve attractiveness. Experimental results show naturally rendered human faces with noticeably improved attractiveness. Gene Cheung, Deming Zhai, Debin Zhao, Hiroshi Sankoh, Sei Naito |
ICIP | 6 |
| 2014 | An efficient method for human pointing estimation for robot interactionabstractIn this paper, we propose an efficient calibration method to estimate the pointing direction via a human pointing gesture to facilitate robot interaction. The ways in which pointing gestures are used by humans to indicate an object are individually diverse. In addition, people do not always point at the object carefully, which means there is a divergence between the line from the eye to the tip of the index finger and the line of sight. Hence, we focus on adapting to these individual ways of pointing to improve the accuracy of target object identification by means of an effective calibration process. We model these individual ways as two offsets, the horizontal offset and the vertical offset. After locating the head and fingertip positions, we learn these offsets for each individual through a training process with the person pointing at the camera. Experimental results show that our proposed method outperforms other conventional head-hand, head-fingertip, and eye-fingertip-based pointing recognition methods. Satoshi Ueno, Sei Naito, Tsuhan Chen |
ICIP | 2 |
| 2014 | Color Transfer based on Spatial Structure for TelepresenceabstractIn this paper, we propose a novel color transfer method based on spatial structure. This work considers an immersive telepresence system, in which distant users can feel as if they are present at a place other than their true location. In preceding studies, the region of an attendee in a remote room is extracted and synthesized on the display of a local room with a preset background image that is similar to the local room. However, the difference between the structure of preset background image and remote room image often degrades the quality of image in the synthesized video. For example, when a part of the human region is occluded by an object that does not exist in the preset background image, the deficient human region is shown despite the inexistence of the occluding object. To solve this problem, instead of using a preset background image, we propose a method that applies a color transfer technique to images of the remote room. Our proposed method can offer users an experience that allows them to feel as if they are in the same room as the other participant of telepresence by changing the colors of the remote room to match the colors of the local room. Furthermore, we have improved the similarity between the rooms based on spatial structure that was overlooked by the conventional color transfer methods. The experimental results show that the proposed method can provide the same-room experience for users in the telepresence system through color similarity of the remote and local room. Kentaro Yamada, Hiroshi Sankoh, Sei Naito |
ACM Multimedia | 3 |
| 2013 | Occlusion robust free-viewpoint video synthesis based on inter-camera/-frame interpolationabstractIn this paper, we propose a novel free-viewpoint video synthesis method that adaptively extracts the texture even from occluded areas. The conventional method based on object segmentation and inter-camera interpolation has two major problems. One is that textures of incorrectly segmented objects degrade image quality in the synthesized output of free-viewpoint video. For example, some object textures have missing regions and other object textures include unwanted regions. The other problem is that the inter-camera interpolation often causes inconsistency between the object appearance and corresponding moving direction. In order to overcome these problems, we propose a new texture acquisition scheme based on inter-frame interpolation to handle the case where object segmentation and inter-camera interpolation are both insufficient. In addition, the proposed method enables adaptive selection among three texture acquisition schemes, segmentation, inter-camera interpolation and inter-frame interpolation. This selection is optimally conducted considering the segmentation results and the direction of the virtual view point. The experimental results revealed that the proposed method can acquire an appropriate texture. Consequently, the subjective quality of generated free-viewpoint video is successfully improved while maintaining the original motion property even for occluded objects. Kentaro Yamada, Hiroshi Sankoh, Masaru Sugano, Sei Naito |
ICIP | 4 |
| 2013 | Robust foreground segmentation from sparsely arranged multi-view camerasabstractIn this paper, we propose robust segmentation of foreground objects, such as human regions from sparsely arranged multi-view cameras. This work is intended for an immersive telepresence system. The system can be realized by segmenting a conferee (i.e. the foreground) from captured video at each conference site and then synthesizing life-sized textures with the background of another space. However, segmentation is a very challenging problem where the background has a similar texture to the foreground or the illumination varies according to time changes. Actually, segmentation accuracies of conventional methods are not sufficient to realize the telepresence system. The proposed method achieves sufficient segmentation quality to realize the system by directly estimating the foreground regions in a three-dimensional space based on the object existence probability for an individual camera and the color similarity among multiple cameras. Experimental results showed the effectiveness of the proposed method regarding foreground segmentation accuracy. Furthermore, confirmation was made that the experience of motion parallax for head movement could be naturally realized. Hiroshi Sankoh, Masaru Sugano, Sei Naito |
MMSP | 3 |
| 2013 | In-loop colour-space-transform coding based on integered SVD for HEVC range extensionsabstractInter colour-component correlation is generally very high in RGB 4:4:4 chroma format. To improve the coding performance of the high efficiency video coding (HEVC) especially for such content, we propose the in-loop colour-space-transform. The colour space is dynamically transformed into un-correlated space by employing singular value decomposition (SVD) for each block at both the encoder and decoder. Signals in transformed colour space are coded with the existing intra / inter coding framework. We utilize the simplified SVD process implemented only by integer operations for the complexity reduction. Compared with HM10.0 as an anchor method, BD-bitrate gain reached 23.8% and 23.4% for the all intra case and the random access case, respectively, while a runtime of the decoder increase 4.8-9.8%. Kei Kawamura, Haruhisa Kato, Sei Naito |
PCS | 3 |
| 2012 | Interactive music video application for smartphones based on free-viewpoint video and audio renderingabstractThis paper presents a novel interactive music video application for smartphones based on free-viewpoint video technology in conjunction with three-dimensional positional audio technology. A user can enjoy a music video from a moving viewpoint that the user can manipulate by the touch screen, with the positional audio through the headphone. A user can even manipulate the positions of the performers on the stage as well as the viewpoint. The application, consisting of our audio rendering engine for multiple AAC ADTS files and our video rendering engine for multiple H.264 ES files, runs on a smartphone in stand-alone mode. The application has been released as official content from a music label for Android and iOS. Toshiharu Horiuchi, Hiroshi Sankoh, Tsuneo Kato, Sei Naito |
ACM Multimedia | 4 |
| 2012 | Dynamic camera calibration method for free-viewpoint experience in sport videosabstractIn this paper, we propose a dynamic camera calibration and object extraction method for sport videos captured with a moving pan-tilt-zoom camera. Such technology realizes an immersive free-viewpoint experience whereby audiences can see real sport scenes from any viewpoint. Camera calibration and object extraction are two of the most important processes for rendering free-viewpoint video, since a 3-dimensional model of each object needs to be reconstructed in every frame based on objects' textures and camera parameters. Most conventional rendering methods only use static cameras whose camera parameters and background models change little. However, since the cameras have to be set apart widely enough to capture the entire scene, the resolution of each object becomes low and is not sufficient for rendering high-quality free-viewpoint video. In order to obtain the texture of an object in high resolution from a moving pan-tilt-zoom camera, our proposed method estimates camera parameters by identifying reliable corresponding feature points between video frames, and also extracts the precise textures of objects using estimated camera parameters. Experimental results revealed that the proposed method successfully estimated precise camera parameters compared to the conventional method. Furthermore, by applying our proposed approach, the free-viewpoint video was rendered without visual defects. Hiroshi Sankoh, Masaru Sugano, Sei Naito |
ACM Multimedia | 3 |
| 2012 | Asymmetric partitioning with non-power-of-two transform for intra codingabstractHEVC (High Efficiency Video Coding) is an ongoing standardization target as the next generation of video compression technology. HEVC employs a coding tree block, which is a quad-tree structure of a coding unit. It also employs some unit types; coding unit, prediction unit, and transform unit. A coding unit can be divided into smaller units as prediction units. Though an asymmetric unit is used for inter coding, only symmetric units are permitted for intra coding. In this paper, we propose an asymmetric partitioning with a non-power-of-two transform as a prediction and transform unit. While conventional partitioning locates the cross-point of partitioning lines at the center of the coding unit, the proposed method locates the cross-point in places except center. The proposed method reduces 2.0% BD-bitrate compared with HM5.0 under all intra / high efficiency condition. The validity of the proposed method is confirmed by some experimental results. Kei Kawamura, Haruhisa Kato, Sei Naito |
PCS | 3 |
| 2011 | No reference metric of video coding quality based on parametric analysis of video bitstreamabstractIn this paper, we propose a novel method to measure the perceived picture quality of H.264 coded video based on parametric analysis of the coded bitstream. The parametric analysis means that the proposed method utilizes only bitstream parameters, while it does not have any access to the baseband signal (pixel level information) of the decoded video. The proposed method extracts quantiser-scale, slice type and transformed coefficients from each macroblock. These parameters are used to calculate spatiotemporal image features to reflect coding artifacts which have a strong relation to the subjective quality. A computer simulation shows that the proposed method can estimate the subjective quality at a correlation coefficient of 0.867 whereas the PSNR metric, which is referred to as a benchmark, correlates the subjective quality at a correlation coefficient of 0.773. Osamu Sugimoto, Sei Naito |
ICIP | 2 |
| 2011 | Adaptive loop filter technology based on analytical design considering local image characteristicsabstractBased on Wiener algorithm, the adaptive loop filter (ALF) technology is applied to minimize coding distortion over a frame. As the ALF is designed based on the Wiener algorithm, coding performance can be improved by spatially adaptive selection of filter coefficients dependent on a local image feature. However, since the conventional ALF approaches designed only a single set of frame basis filter coefficients, the filter coefficients are not necessarily optimal to achieve the significant coding gain. As the conventional work, the modified adaptive loop filter approach that employs the segmentation pattern has been proposed. This approach divides the current frame into a few segments based on the structure of image texture, and the filter coefficients are determined for each segment type. Although 14 candidates are defined as available segmentation modes, further improvement can be expected because such segmentation can not handle the complicated combination of various texture characteristics within a video frame. From this perspective, we propose an enhanced ALF approach based on spatially adaptive control of filter design dependent on the texture property of the local image region. Experimental results show that the bit reduction performance is improved by 1.6 points in the maximum case against the conventional scheme. Furthermore, the proposed scheme contributes to improving the subjective picture quality. Tomonobu Yoshino, Sei Naito |
ICIP | 2 |
| 2011 | Arbitrary product detection from advertisement video by using object independent featuresabstractThis research proposes a novel method to extract image regions of products from an advertisement video, by analyzing features which are completely independent from the target object. Namely, we focus on how each product is emphasized in the video production, and propose the utilization of low-level visual features which leverage the technical know-how of video producers. By using such features, our proposal can achieve highly accurate detection of temporal and spatial locations of the advertised products, regardless of the product domain. Evaluation of the proposed method has been conducted with the actual advertisement video, in which accuracy of 79.4% in F-measure has been achieved. Tomohiko Takahashi, Masaru Sugano, Keiichiro Hoashi, Sei Naito |
ICME | 4 |
| 2010 | Low-complexity scheme for adaptive interpolation filter based on amplitude characteristic analysisabstractMotion-compensated (MC) prediction is one of the most effective coding tools for high-compression video coding. In former works, adaptive interpolation filter (AIF) technology, which improves the coding performance of the fractional-pel MC, has been proposed. As for the real-time codec system, the implementation of AIF based on the Wiener filter algorithm is difficult because of its high computational complexity. To overcome this problem, in this paper, we propose a low-complexity AIF scheme, which can maintain the coding performance achieved by the conventional AIF. The key technology is the efficient filter design process based on the amplitude characteristic analysis with the precise estimation of MC error signal. The experimental result confirmed that almost the same coding performance was maintained at only 30% of the complexity compared to the conventional AIF. Tomonobu Yoshino, Sei Naito, Shigeyuki Sakazawa, Shuichi Matsumoto |
ICIP | 2 |
| 2010 | Robust background subtraction method based on 3D model projections with likelihoodabstractWe propose a robust background subtraction method for multi-view images, which is essential for realizing free viewpoint video where an accurate 3D model is required. Most of the conventional methods determine background using only visual information from a single camera image, and the precise silhouette cannot be obtained. Our method employs an approach of integrating multi-view images taken by multiple cameras, in which the background region is determined using a 3D model generated by multi-view images. We apply the likelihood of background to each pixel of camera images, and derive an integrated likelihood for each voxel in a 3D model. Then, the background region is determined based on the minimization of energy functions of the voxel likelihood. Furthermore, the proposed method also applies a robust refining process, where a foreground region obtained by a projection of a 3D model is improved according to geometric information as well as visual information. A 3D model is finally reconstructed using the improved foreground silhouettes. Experimental results show the effectiveness of the proposed method compared with conventional works. Hiroshi Sankoh, Akio Ishikawa, Sei Naito, Shigeyuki Sakazawa |
MMSP | 3 |
| 2010 | Efficient free viewpoint video-on-demand scheme realizing walk-through experienceabstractThis paper presents an efficient video-on-demand (VOD) scheme for free viewpoint television (FTV), and proposes a data format and its data generation method to provide a walkthrough experience. We employ a hybrid rendering approach to describe a 3D scene using 3D model data for objects and textures. However, conventional hybrid rendering methods such as multi-texturing include excessive redundancy in texture data and demand a great deal of bandwidth to transmit. In this paper we propose an efficient texture data format, which removes the redundancy due to occlusion of objects by employing an orthogonal projection image for each object. The additional advantage of the data format is great simplification at the server to choose the transmitted images that correspond to the requested viewpoint. Experiments using multiview real video sequences confirm that the proposed scheme can reduce the transmission of texture data by as much as 42% compared to the conventional scheme. Akio Ishikawa, Hiroshi Sankoh, Sei Naito, Shigeyuki Sakazawa |
PCS | 3 |
| 2010 | Complementary coding mode design based on R-D cost minimization for extending H.264 coding technologyabstractTo improve high resolution video coding efficiency under low bit-rate condition, an appropriate coding mode is required from an R-D optimization (RDO) perspective, although a coding mode defined within the H.264 standard is not always optimal for RDO criteria. With this in mind, we previously proposed extended SKIP modes with close-to-optimal R-D characteristics. However, the additional modes did not always satisfy the optimal R-D characteristics, especially for low bit-rate coding. In this paper, we propose an enhanced coding mode capable of providing a candidate corresponding to the minimum R-D cost by controlling the residual signal associated with the extended SKIP mode. The experimental result showed that the PSNR improvement against H.264 and our previous approach reached 0.42 dB and 0.24 dB in the maximum case, respectively. Tomonobu Yoshino, Sei Naito, Shigeyuki Sakazawa, Shuichi Matsumoto |
PCS | 2 |
| 2009 | Objective perceptual video quality measurement method based on hybrid no reference frameworkabstractIn this paper, the authors propose a novel method to measure the perceived picture quality of H.264 coded video based on hybrid no reference framework. The latter term means that the proposed model uses only receiver-side information for objective video quality assessment, but analyzes both the compressed bitstream and baseband signals of the decoded picture to improve the estimation accuracy of the subjective quality. The proposed method extracts quantizer-scale information from the bit-stream along with two spatiotemporal image features from the baseband signal, which are integrated to express the overall quality using the weighted Minkowski metric. A computer simulation shows the proposed method can estimate the subjective quality at a correlation coefficient of 0.909 whereas the PSNR metric, which is referred to as a benchmark, correlates the subjective quality at a correlation coefficient of 0.773. Osamu Sugimoto, Sei Naito, Shigeyuki Sakazawa, Atsushi Koike |
ICIP | 2 |
| 2008 | Free viewpoint video generation for walk-through experience using image-based renderingabstractThis paper presents a novel method to represent a real 3D world using IBR (Image-based Rendering) technology. The major achievement is realization of "walk-through" experience, in which audiences can see the scene of sport games as if they were the player, while the conventional IBR method cannot provide such an experience. The key idea is a newly introduced framework "locally divided ray space" to capture the ray inside the scene. Consequently, the proposed algorithm can generate a virtual image inside the scene using multiple real images captured from outside. In this paper, the outline of the proposed method as well as the possible demonstration at the conference venue is described. Akio Ishikawa, Mehrdad Panahpour Tehrani, Sei Naito, Shigeyuki Sakazawa, Atsushi Koike |
ACM Multimedia | 3 |
| 2006 | Real-time multi-scale brain data acquisition, assembly, and analysis using an end-to-end OptIPuter
Rajvikram Singh, Nicholas Schwarz, Nut Taesombut, Byungil Jeong, Luc Renambot, Abel W. Lin, Ruth West, Hiromu Otsuka, Sei Naito, Steven Peltier, Maryann E. Martone, Kazunori Nozaki, Jason Leigh, Mark H. Ellisman |
Future Gener. Comput. Syst. | 10 |
| 2005 | Effective rate control method for minimizing temporal fluctuations in picture quality applicable for MPEG-4 AVC/H.264 encodingabstractAppropriate rate control plays a very important role in encoding motion pictures under the constant bit-rate. One of the requirements for rate control is minimizing temporal fluctuations in picture quality, which may cause flicker artifacts. To satisfy the requirement, some methods for MPEG-1 and MPEG-2 have been proposed. However, these methods cannot be applied to the H.264 encoding algorithm based on R-D optimization since they require the DCT coefficients and the accurate output bit amount for each quantization parameter of a current picture in advance. To overcome the problem, we propose a rate control method to satisfy the requirement in the H.264 encoding algorithm by introducing approximations to estimate the bit amount and the distortion. Atsushi Matsumura, Sei Naito, Ryoichi Kawada, Atsushi Koike |
ICIP (1) | 2 |
| 2004 | Optimal JPEG2000 encoder mechanism for low delay and efficient distribution of HDTV programsabstractMotion JPEG2000 has been developed as a new video coding standard for motion pictures, utilizing the still image coding standard JPEG2000. Motion JPEG2000 specifies a normative bitstream syntax that the decoder must recognize and allows flexible selection of detailed coding parameters at the encoder. The compression performance largely depends, therefore, on the encoder design. However, the simple implementation of JPEG2000 encoder might cause picture quality degradation due specifically to wavelet image coding of interlaced TV signals and unsatisfactory bit assignments. That degradation might then become conspicuous in subjective picture quality especially during low bit-rate coding. To overcome this problem, we investigated an optimal design for a Motion JPEG2000 encoder that can be used for primary distribution applications of HDTV programs, assuming that the available bit-rate is approximately 50 Mbps. Our study introduces advanced key technologies not yet recognized officially. Coding experiments using those technologies confirmed that a significant coding gain was achieved versus conventional encoding schemes. Sei Naito, Atsushi Koike, Shuichi Matsumoto |
ICASSP (5) | 1 |
| 2004 | Effective interpolation for free viewpoint images using multi-layered dynamic background buffersabstractPresenting images from a free viewpoint is a promising interactive video application. Though reference images from one viewpoint and their depth maps are often used to render free viewpoint images, picture quality degradation may occur because of lack of information in background regions that are occluded by foreground regions. In this paper an interpolation method for free viewpoint images using multilayered dynamic background buffers is proposed. In the proposed method, the buffers, updated using the reference images divided in each frame, are used to store the background regions and the output images are interpolated by the background buffers. Since the background buffers are created and updated using only the reference images and their depth maps, additional information on the background buffers is not required for the interpolation. The effectiveness of the proposed method was evaluated by several simulation experiments. Atsushi Matsumura, Sei Naito, Ryoichi Kawada, Atsushi Koike, Shuichi Matsumoto |
ICIP | 2 |
| 2002 | Optimal MPEG-2 encoder design for low bit-rate HDTV digital broadcastingabstractMPEG-2 is an international standard specifying a normative bitstream syntax that decoders must recognize and that allows flexible selection of detailed coding parameters at the encoder. The compression performance is therefore largely dependent on the encoder design and the difference in coded picture quality is present even among encoders developed by different codec makers at identical bit-rates. Conventional MPEG-2 encoders were also designed for high picture quality transmission at an adequate bit-rate, and the overhead portion occupancy in a coded bitstream can be ignored in the final picture quality. However, the unexpected increase in overhead portion occupancy might cause fatal picture quality degradation at extremely low bitrates. To overcome this problem, we investigated an optimal design for an MPEG-2 encoder for use in HDTV digital terrestrial broadcasting where the available bit-rate is limited to less than 0.2 bit/pixel. In our study, advanced key technologies which have not been clearly recognized until now were introduced into possible optimization points, picture type selection macroblock coding mode selection and rate control, and a significant coding gain versus conventional encoding schemes was confirmed in coding experiments. Sei Naito, Atsushi Koike, Masahiro Wada, Shuichi Matsumoto |
ICIP (3) | 1 |
| 1999 | Advanced Rate Control Technologies for 3D-Hdtv Digital Coding Based on Mpeg-2 Multi-View ProfileabstractThis paper describes advanced rate control technologies suitable for the digital compression coding based on MPEG-2 multi-view profile for transmitting three dimensional stereo HDTV signals at bit-rates 45 Mbps for PDH networks or 52 Mbps for SDH networks. First of all, a common buffer control is introduced, which achieves statistical multiplexing effect while keeping the picture qualify of both the left and right pictures well balanced. Under the assumption of applying this common buffer control, we investigate the optimal setting of GOP structure which has a close relation with the rate control. And then, a discriminatory bit allocation is proposed which improves the left picture quality without any degradation of the right picture. From the results of coding experiment conducted to evaluate the coding picture achieved by this coding scheme, it is confined that our coding scheme gives satisfactory picture quality even at 45 Mbps including audio and FEC data. Sei Naito, Shuichi Matsumoto |
ICIP (1) | 1 |