VLDB 2026 Research / reviewers in the wild / expert
Daisuke Sugimura
dblp:56/9054
· DBLP profile ↗
27ranked-venue papers
7as first author
5since 2021 · last 2025
0000-0002-8305-9790ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 5 first-authorComputer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Theoretical and Numerical Analysis of Measurement Limits in Quantum Image Representations: Qubit Lattice, FRQI, and NEQRabstractWith the advancement of quantum computing, quantum image representations (QIRs) have been widely studied as platforms for quantum image processing. While previous studies have focused on theoretical qubit counts and circuit complexities, there has been limited quantitative evaluation of image reconstruction performance from quantum states. In this study, we clarify the theoretical upper bounds of restoration performance for three representative QIRs: qubit lattice (QL), flexible representation of quantum images (FRQI), and novel enhanced quantum representation (NEQR). Also, we validate them through large-scale GPU-based simulations using cuQuantum. Specifically, we (1) derive a prediction formula for PSNR in probabilistic schemes like QL and FRQI, independent of image content and size; (2) show that deterministic schemes like NEQR require a number of samples following a harmonic number for complete restoration; and (3) perform large-scale simulations to verify these findings. These results provide fundamental insights for designing QIRs and their applications on NISQ devices. Kento Okada, Akira Hasegawa, Yoshihiro Maeda, Daisuke Sugimura, Mikio Hasegawa, Norishige Fukushima |
VCIP | 4 |
| 2024 | Physiological Modeling With Multispectral Imaging for Heart Rate EstimationabstractHeart rate (HR) is a key parameter in evaluating the physiological and emotional states of a person. In this paper, we propose a novel video-based heart rate (HR) estimation method based on physiological modeling with multispectral imaging. To capture blood volume pulse (BVP) associated with a person’s heartbeat, we utilize a camera that records multispectral video consisting of red, green, blue, and near-infrared information. The novelty of the proposed method is the incorporation of a physiological BVP model into a multispectral HR estimation framework. The integration of a physiological model-based BVP signal extraction scheme into an adaptive multispectral framework enables the suppression of noise derived from ambient light and the accurate extraction of the BVP signal, thereby enhancing HR estimation performance. The experiments using RGB/NIR video datasets demonstrate the effectiveness of the proposed method. Kosuke Kurihara, Yoshihiro Maeda, Daisuke Sugimura, Takayuki Hamamoto |
ICIP | 3 |
| 2022 | Blood Volume Pulse Signal Extraction based on Spatio-Temporal Low-Rank Approximation for Heart Rate EstimationabstractWe propose a novel blood volume pulse (BVP) signal extraction method for heart rate estimation that incorporates the self-similarity properties of BVP in the spatial and temporal domains. The main novelty of the proposed method is the incorporation of the temporal self-similarity of BVP via low-rank approximation in the time-delay coordinate system for BVP signal extraction. To make a low-rank approximation of BVP in the time domain, we introduce knowledge of linear time-invariant systems, i.e., the autoregressive (AR) model lies in the low-rank subspace in the time-delay coordinate system. In the medical field, it is widely known that BVP has quasi-periodic temporal characteristics owing to the cardiac pulse and exhibits self-similarity properties in the temporal domain. Hence, we model the temporal behavior of BVP as an AR process, allowing for a low-rank approximation of BVP in the time-delay coordinate system. Low-rank approximation of BVP in the time and spatial domains enables reliable BVP signal extraction, resulting in accurate heart rate estimation. The experiments demonstrate the effectiveness of the proposed method. Kosuke Kurihara, Yoshihiro Maeda, Daisuke Sugimura, Takayuki Hamamoto |
VCIP | 3 |
| 2021 | Non-Contact Heart Rate Estimation via Adaptive RGB/NIR Signal FusionabstractWe propose a non-contact heart rate (HR) estimation method that is robust to various situations, such as bright, low-light, and varying illumination scenes. We utilize a camera that records red, green, and blue (RGB) and near-infrared (NIR) information to capture the subtle skin color changes induced by the cardiac pulse of a person. The key novelty of our method is the adaptive fusion of RGB and NIR signals for HR estimation based on the analysis of background illumination variations. RGB signals are suitable indicators for HR estimation in bright scenes. Conversely, NIR signals are more reliable than RGB signals in scenes with more complex illumination, as they can be captured independently of the changes in background illumination. By measuring the correlations between the lights reflected from the background and facial regions, we adaptively utilize RGB and NIR observations for HR estimation. The experiments demonstrate the effectiveness of the proposed method. Kosuke Kurihara, Daisuke Sugimura, Takayuki Hamamoto |
IEEE Trans. Image Process. | 2 |
| 2021 | Hierarchical Group-Level Emotion RecognitionabstractGroup-level emotion recognition is a technique for estimating the emotion of a group of people. In this paper, we propose a novel method for group-level emotion recognition. Our method lies in the two-fold contributions: (1) recognition of group-level emotion using a hierarchical classification approach; (2) incorporation of novel features to contribute to the description of the group-level emotion. We consider that the use of facial expressions of people will only be effective in differentiating images labeled as “Positive” because those labeled as “Neutral” or “Negative” are likely to include similar facial expressions. Therefore, we first perform binary classification based on facial expression recognition to distinguish “Positive” labels that include discriminative facial expressions (e.g., smile) from the others. We evaluate outcomes that are not classified as “Positive” during the first classification by exploiting scene features that describe what type of events (e.g., demonstration or funeral) are shown in the image. The other novelty of our method lies in two-fold. The first is the exploitation of visual attention for the first classification. It allows us to estimate which faces are the main subjects in the target image, thereby suppressing the influences of faces in the background that contribute less to group-level emotion. The second is the exploitation of object-wise semantic information (labels) for the second classification. This allows a more detailed description of the scene context in the image and enables performance enhancement in the second classification. We demonstrate the effectiveness of our method through experiments using public datasets. Katsuya Fujii, Daisuke Sugimura, Takayuki Hamamoto |
IEEE Trans. Multim. | 2 |
| 2020 | Network-Density-Controlled Decentralized Parallel Stochastic Gradient Descent in Wireless SystemsabstractThis paper proposes a communication strategy for decentralized learning on wireless systems. Our discussion is based on the decentralized parallel stochastic gradient descent (D-PSGD), which is one of the state-of-the-art algorithms for decentralized learning. The main contribution of this paper is to raise a novel open question for decentralized learning on wireless systems: there is a possibility that the density of a network topology significantly influences the runtime performance of DPSGD. In general, it is difficult to guarantee delay-free communications without any communication deterioration in real wireless network systems because of path loss and multi-path fading. These factors significantly degrade the runtime performance of D-PSGD. To alleviate such problems, we first analyze the runtime performance of D-PSGD by considering real wireless systems. This analysis yields the key insights that dense network topology (1) does not significantly gain the training accuracy of D-PSGD compared to sparse one, and (2) strongly degrades the runtime performance because this setting generally requires to utilize a low-rate transmission. Based on these findings, we propose a novel communication strategy, in which each node estimates optimal transmission rates such that communication time during the D-PSGD optimization is minimized under the constraint of network density, which is characterized by radio propagation property. The proposed strategy enables to improve the runtime performance of D-PSGD in wireless systems. Numerical simulations reveal that the proposed strategy is capable of enhancing the runtime performance of D-PSGD. Koya Sato, Yasuyuki Satoh, Daisuke Sugimura |
ICC | 3 |
| 2020 | Three-Dimensional Point Cloud Object Detection Using Scene Appearance Consistency Among Multi-View Projection DirectionsabstractThree-dimensional (3D) object detection in point clouds is an important technique for various high-level computer vision tasks. In this study, we propose a method for point-wise detection of regions of objects in a scene. We regard the 3D object detection problem as a series of optimal matching problems between object and scene images, which are obtained by projecting point clouds into multiple viewpoints. The main novelty of this study is treating the 3D object detection problem as the determination of optimal correspondence among image sets. Unlike the existing methods that directly employ individual correspondences between projected image pairs, the simultaneous matching of projected image sets allows the evaluation of the appearance consistency of the target object in multi-viewpoint scene images. The other novelty of the proposed method is using principal component analysis to estimate effective image-projection directions for object point clouds. By projecting object point clouds in directions orthogonal to the first principal component basis, the projected images can include plenty of point clouds information, thus providing highly discriminative features for image matching. We back-project reliable matching results retrieved from the image-set correspondence into 3D space to achieve point-wise object detection. Experiments using public datasets demonstrate the effectiveness and performance of the proposed method. Daisuke Sugimura, Tomoaki Yamazaki, Takayuki Hamamoto |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Hierarchical Group-level Emotion Recognition in the WildabstractWe propose a method for group-level emotion recognition in the wild. The main novelty of our method lies with the recognition of group-level emotions using a hierarchical classification approach. We consider that using the facial expressions of people will only be effective in differentiating images labeled as "Positive" because those labeled as "Neutral" or "Negative" are likely to include similar facial expressions (i.e., less discriminative). Therefore, we first perform binary classification based on facial expression recognition to distinguish "Positive" labels that include discriminative facial expressions (e.g., smile) from the others. We evaluate outcomes that are not classified as "Positive" at the first classification by exploiting scene features that describe what type of events (e.g., demonstration or funeral) are taking place in the image. Classification using scene features will not only be effective in differentiating "Negative" and "Neutral" labels but also in recognizing "Positive" labels, where facial expression features show less discriminative characteristics. The other novelty of the proposed method is to the exploitation of visual attention. Using visual attention allows us to estimate which faces are the main subjects in the target image, thereby suppressing the influences of faces in the background that contribute less to group-level emotion. We demonstrate the effectiveness of our proposed method through experiments using a public dataset. Katsuya Fujii, Daisuke Sugimura, Takayuki Hamamoto |
FG | 2 |
| 2019 | Adaptive Fusion of RGB/NIR Signals Based on Face/Background Cross-Spectral Analysis for Heart Rate EstimationabstractWe propose a method for heart rate (HR) estimation that is robust to various situations such as bright, low-light, and varying illumination scenes. We capture temporal variations in the pixel values owing to person's cardiac pulse by using a camera that records red, green, and blue (RGB) and near-infrared (NIR) information. The key novelty of our method is to introduce a scheme for adaptive fusion of RGB and NIR signals for HR estimation, by analyzing variations in the background illuminations. RGB signals will be a good cue for HR estimation under bright scenes. In contrast, NIR signals are more reliable in HR estimation than RGB ones in complex illumination scenes, because NIR signals can be captured independent to changes in the background illuminations. By measuring correlations of signals between background and face regions, we adaptively utilize RGB and NIR signals for HR estimation. Experiments demonstrate the effectiveness of our method. Kosuke Kurihara, Daisuke Sugimura, Takayuki Hamamoto |
ICIP | 2 |
| 2018 | Discovering Correspondence Among Image Sets with Projection View Preservation For 3D Object Detection in Point CloudsabstractWe propose a method for detecting objects that correspond to given three-dimensional (3D) point clouds in a scene. We regard the 3D object detection as a series of optimal matching of the object and scene images that are obtained by projecting point clouds into multiple viewpoints. The key novelty of the proposed method is to introduce a constraint imposed by the spatial relationship among the image-projection directions for the object point clouds, to discover the optimal matching of the projected image sets. This constraint allows to evaluate the appearance consistency of the object in multi-viewpoint scene images. Thus, image-projection directions can be effective cues to detect objects even in cluttered scenes, where previous methods are not effective. We estimate the image-projection directions for the object point clouds by applying principal component analysis to the object point clouds and hence include highly discriminative image features. Then, we back-project reliable matching results, which are retrieved from the image set correspondence, into 3D space to achieve a point-wise object detection. Experiments using public datasets demonstrate the effectiveness and performance of the proposed method. Tomoaki Yamazaki, Daisuke Sugimura, Takayuki Hamamoto |
ICASSP | 2 |
| 2018 | Low-Light Color Image Super-Resolution Using RGB/NIR SensorabstractWe propose a method for super-resolution (SR) of low-resolution (LR) color images taken in low-light scenes. Our method is based on multi-frame SR technique, which fuses multiple LR images taken at different camera positions to synthesize a high-resolution color image. Previous methods have implicitly assumed that LR images could be captured with less noise and blur. However, heavy noise and motion blur will be imposed on images taken in low-light scene. They make it difficult to super-resolve low-light images with high quality. To overcome this problems, we utilize a single sensor that captures red, green, blue (RGB) and near-infrared (NIR) information. Since NIR images taken using an NIR flash unit can be captured with less noise, they contribute to effective reduction of image artifacts in super-resolving LR images. Using multiple RGB/NIR raw images, we jointly perform deblurring, denoising and SR of LR images. Experiments using real raw data demonstrate the effectiveness of our method. Takayuki Honda, Takayuki Hamamoto, Daisuke Sugimura |
ICIP | 3 |
| 2018 | Online background subtraction with freely moving cameras using different motion boundaries
Daisuke Sugimura, Fumihiro Teshima, Takayuki Hamamoto |
Image Vis. Comput. | 1 |
| 2018 | Underwater Image Color Correction using Exposure-Bracketing ImagingabstractAbsorption and scattering of light in an underwater scene saliently attenuate red spectrum components. They cause heavy color distortions in the captured underwater images. In this letter, we propose a method for color-correcting underwater images, utilizing a framework of gray information estimation for color constancy. The key novelty of our method is to utilize exposure-bracketing imaging: a technique to capture multiple images with different exposure times for color correction. The long-exposure image is useful for sufficiently acquiring red spectrum information of underwater scenes. In contrast, pixel values in the green and blue channels in the short-exposure image are suitable because they are unlikely to attenuate more than the red ones. By selecting appropriate images (i.e., least over- and under-exposed images) for each color channel from those taken with exposure-bracketing imaging, we fuse an image that includes sufficient spectral information of underwater scenes. The fused image allows us to extract reliable gray information of scenes; thus, effective color corrections can be achieved. We perform color correction by linear regression of gray information estimated from the fused image. Experiments using real underwater images demonstrate the effectiveness of our method. Kohei Nomura, Daisuke Sugimura, Takayuki Hamamoto |
IEEE Signal Process. Lett. | 2 |
| 2017 | Disparity estimation in stereo videos using spatio-temporal disparity hyperplane modelsabstractWe propose a method for disparity estimation in stereo video. We address the problems associated with spatially-temporally-correlated disparity variations (STCDV). STCDV problems are caused by complex motions, e.g., yaw-rotation, pan-tilt-zoom camera movements, etc. The key novelty of this study is to introduce a spatio-temporal disparity hyperplane (STDH) model. The proposed STDH model represents a hyperplane defined in four-dimensional space spanned by disparity, image plane, and time coordinates. Our STDH model is represented by surface normals varying with the spatially-temporally-correlated changes in disparity. Thus, our STDH model is effective in estimating disparity in a stereo video including STCDVs. We estimate video disparity by incorporating our STDH model into the PatchMatch brief propagation framework. Our experiments demonstrate that the proposed method outperforms other methods. Hiroki Nakano, Daisuke Sugimura, Takayuki Hamamoto |
ICASSP | 2 |
| 2017 | RGB-NIR imaging with exposure bracketing for joint denoising and deblurring of low-light color imagesabstractColor images taken in low light scenes are deteriorated with noise and motion blur. The simultaneous reduction of noise and motion blur from the low-light color images is difficult because the imposed noise hinders accurate motion blur kernel estimation. To overcome this problem, we build a novel imaging system using a single sensor that captures red, green, blue (RGB) and near-infrared (NIR) images. Our imaging system captures low-light scenes with exposure bracketing, which is a technique to acquire multiple images with different exposure times. It thus allows us to obtain the short- and long-exposure RGB/NIR images. Both the short- and long-exposure NIR images taken using an NIR flash unit can be captured with less noise; thus they enable estimation of motion blur kernel accurately. Based on this fact, we perform joint denoising and deblurring of the low-light color image with the estimated motion blur kernel. Our experiments using real raw data captured by our imaging system demonstrate the effectiveness of our method. Hiroki Yamashita, Daisuke Sugimura, Takayuki Hamamoto |
ICASSP | 2 |
| 2017 | Color correction of underwater images based on multi-illuminant estimation with exposure bracketing imagingabstractWe propose a method for color correction of underwater images based on multi-illuminant estimation. We regard the color distortion of underwater images as the color cast that is illuminated by multiple light sources. In order to effectively remove the color distortions from the underwater image, we capture underwater scenes by exposure bracketing imaging. Using multiple images taken with different exposure times, we fuse an image where the attenuation differences in the spectra information of the incoming light are mitigated. We apply a multi-illuminant estimation to the fused image to reconstruct the underwater images so as to be those in a canonical (white) illumination environment. Our experiments demonstrate the effectiveness of our method. Kohei Nomura, Daisuke Sugimura, Takayuki Hamamoto |
ICIP | 2 |
| 2017 | Depth upsampling by depth predictionabstractWe propose a method for depth upsampling with the aid of high-resolution color image. The key novelty of our method is to exploit a spatio-temporal coherency between the color and depth image sequences. It allows us to perform a depth prediction using the responses from motion estimation in the color image sequence. The predicted depth image is able to estimate the scene boundary regions in the color image; it enables to suppress the influences of color image textures in depth upsampling. We synthesize high-resolution depth images with the help of the estimated scene boundary. Our experiments demonstrate the effectiveness of our method. Atsuhiko Tsuchiya, Daisuke Sugimura, Takayuki Hamamoto |
ICIP | 2 |
| 2016 | Two-layer light field imaging using an organic photoelectric conversion filmabstractIn this paper, we propose a novel method for light field imaging. Previous systems are difficult to obtain multi-viewpoint images (sub-images) at the high-resolution. In order to overcome this problem, we propose a two-layer light field imaging system by using an organic photoelectric conversion film (OPCF). Our imaging system places the OPCF having the green spectral sensitivity onto the micro-lens array of the conventional light field camera. It allows us to obtain the green spectrum information at the full-resolution of the image sensor. In contrast, the other spectra information (red and blue) are coded by the optical system of the light field camera, and are recorded by the image sensor. Using the captured images, we synthesize the full-resolution sub-images. Our experiments using synthetic images demonstrate that our method outperformed other previous methods. Suguru Kobayashi, Daisuke Sugimura, Takayuki Hamamoto |
ICIP | 2 |
| 2016 | Enhanced Cascading Classifier Using Multi-Scale HOG for Pedestrian Detection from Aerial ImagesabstractWe propose a method for pedestrian detection from aerial images captured by unmanned aerial vehicles (UAVs). Aerial images are captured at considerably low resolution, and they are often subject to heavy noise and blur as a result of atmospheric influences. Furthermore, significant changes to the appearance of pedestrians frequently occur because of UAV motion. In order to address these crucial problems, we propose a cascading classifier that concatenates a pre-trained classifier and an online learning-based classifier. We construct the first classifier using deep belief network (DBN) with an extended input layer. Unlike previous approaches that use raw images as the input layer of the DBN, we exploit multi-scale histogram of oriented gradients (MS-HOG) features. The MS-HOG enables us to supply better and richer information than low-resolution aerial images for constructing a reliable deep structure of DBN, because the dimensions of the input features can be expanded. Furthermore, the MS-HOG effectively extracts the necessary edge information while reducing trivial gradients and noise. The second classifier is based on online learning, and it uses predictions of the target appearance using UAV motions. Predicting the target appearance enables us to collect reliable training samples for the classifier’s online learning process. Experiments using aerial videos demonstrate the effectiveness of the proposed method. Daisuke Sugimura, Takayuki Fujimura, Takayuki Hamamoto |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2016 | Detecting flaws in golf swing using common movements of professional players
Daisuke Sugimura, Hiroharu Tsutsui, Takayuki Hamamoto |
Mach. Vis. Appl. | 1 |
| 2016 | Compressive multi-spectral imaging using self-correlations of images based on hierarchical joint sparsity models
Daisuke Sugimura, Masaru Tomabechi, Tadaaki Hosaka, Takayuki Hamamoto |
Mach. Vis. Appl. | 1 |
| 2015 | Enhancing low-light color images using an RGB-NIR single sensorabstractIn this paper, we propose a method to enhance the color image of a low-light scene by using a single sensor that simultaneously captures red, green, blue (RGB) and near-infrared (NIR) information. Typical image enhancement methods require two cameras to simultaneously capture color and NIR images. In such cases, meticulous calibration is required to adjust the pixel positions of the two cameras. By contrast, our proposed system is calibration free, but achieves accurate color image restoration. We divide the captured multi-spectral data into RGB and NIR information based on the spectral sensitivity of our imaging system. Using the NIR information for guidance, we reconstruct the corresponding clear color image based on a joint demosaicking and denoising technique. Our experiments show the effectiveness of our method using raw data captured by our imaging system. Hiroki Yamashita, Daisuke Sugimura, Takayuki Hamamoto |
VCIP | 2 |
| 2015 | Enhancing Color Images of Extremely Low Light Scenes Based on RGB/NIR Images Acquisition With Different Exposure TimesabstractWe propose a novel method to synthesize a noise- and blur-free color image sequence using near-infrared (NIR) images captured in extremely low light conditions. In extremely low light scenes, heavy noise and motion blur are simultaneously produced in the captured images. Our goal is to enhance the color image sequence of an extremely low light scene. In this paper, we augment the imaging system as well as enhancing the image synthesis scheme. We propose a novel imaging system that can simultaneously capture the red, green, blue (RGB) and the NIR images with different exposure times. An RGB image is taken with a long exposure time to acquire sufficient color information and mitigates the effects of heavy noise. By contrast, the NIR images are captured with a short exposure time to measure the structure of the scenes. Our imaging system using different exposure times allows us to ensure sufficient information to reconstruct a clear color image sequence. Using the captured image pairs, we reconstruct a latent color image sequence using an adaptive smoothness condition based on gradient and color correlations. Our experiments using both synthetic images and real image sequences show that our method outperforms other state-of-the-art methods. Daisuke Sugimura, Takuya Mikami, Hiroki Yamashita, Takayuki Hamamoto |
IEEE Trans. Image Process. | 1 |
| 2014 | Capturing color and near-infrared images with different exposure times for image enhancement under extremely low-light sceneabstractNoise and blur are annoying, common problems in low-light photography. In this paper, we propose a novel framework to reconstruct a blur and noise-free color image sequence using near-infrared (NIR) images. In extremely low light, previous works may fail in color image restoration because both heavy noise and motion blur are produced simultaneously. To overcome these essential problems, we augment both the imaging system and image synthesis method. Our imaging system captures the color and NIR images with different exposure times. Capturing color image with a long exposure time allows us to mitigate the heavy noise. In contrast, the NIR images are taken with a short exposure time to measure the structure of the scene. By leveraging the captured image pairs, we reconstruct a latent color image sequence using a scale map and color correlation. Our experiments on actual images show that our method outperforms other state-of-the-art methods under extremely low-light condition. Takuya Mikami, Daisuke Sugimura, Takayuki Hamamoto |
ICIP | 2 |
| 2013 | Head direction estimation from low resolution images with scene adaptation
Isarun Chamveha, Yusuke Sugano, Daisuke Sugimura, Teera Siriteerakul, Takahiro Okabe, Yoichi Sato 0001, Akihiro Sugimoto |
Comput. Vis. Image Underst. | 3 |
| 2009 | Using individuality to track individuals: Clustering individual trajectories in crowds using local appearance and frequency traitabstractIn this work, we propose a method for tracking individuals in crowds. Our method is based on a trajectory-based clustering approach that groups trajectories of image features that belong to the same person. The key novelty of our method is to make use of a person's individuality, that is, the gait features and the temporal consistency of local appearance to track each individual in a crowd. Gait features in the frequency domain have been shown to be an effective biometric cue in discriminating between individuals, and our method uses such features for tracking people in crowds for the first time. Unlike existing trajectory-based tracking methods, our method evaluates the dissimilarity of trajectories with respect to a group of three adjacent trajectories. In this way, we incorporate the temporal consistency of local patch appearance to differentiate trajectories of multiple people moving in close proximity. Our experiments show that the use of gait features and the temporal consistency of local appearance contributes to significant performance improvement in tracking people in crowded scenes. Daisuke Sugimura, Kris Makoto Kitani, Takahiro Okabe, Yoichi Sato 0001, Akihiro Sugimoto |
ICCV | 1 |
| 2006 | 3D Head Tracking using the Particle Filter with Cascaded ClassifiersabstractWe propose a method for real-time people tracking using multiple cameras. The particle filter framework is known to be effective for tracking people, but most of existing methods adopt only simple perceptual cues such as color histogram or contour similarity for hypothesis evaluation. To improve the robustness and accuracy of tracking more sophisticated hypothesis evaluation is indispensable. We therefore present a novel technique for human head tracking using cascaded classifiers based on AdaBoost and Haar-like features for hypothesis evaluation. In addition, we use multiple classifiers, each of which is trained respectively to detect one direction of a human head. During real-time tracking the most suitable classifier is adaptively selected by considering each hypothesis and known camera position. Our experimental results demonstrate the effectiveness and robustness of our method. 1 1 Yoshinori Kobayashi, Daisuke Sugimura, Yoichi Sato 0001, Kousuke Hirasawa, Naohiko Suzuki, Hiroshi Kage, Akihiro Sugimoto |
BMVC | 2 |