VLDB 2026 Research / reviewers in the wild / expert
Hitoshi Imaoka
dblp:09/3516
· DBLP profile ↗
16ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0003-3891-9410ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 3 since 2021Security and privacy · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning the Optimal Stopping for Early Classification within Finite Horizons via Sequential Probability Ratio TestabstractTime-sensitive machine learning benefits from Sequential Probability Ratio Test (SPRT), which provides an optimal stopping time for early classification of time series. However, in *finite horizon* scenarios, where input lengths are finite, determining the optimal stopping rule becomes computationally intensive due to the need for *backward induction*, limiting practical applicability. We thus introduce FIRMBOUND, an SPRT-based framework that efficiently estimates the solution to backward induction from training data, bridging the gap between optimal stopping theory and real-world deployment. It employs *density ratio estimation* and *convex function learning* to provide statistically consistent estimators for sufficient statistic and conditional expectation, both essential for solving backward induction; consequently, FIRMBOUND minimizes Bayes risk to reach optimality. Additionally, we present a faster alternative using Gaussian process regression, which significantly reduces training time while retaining low deployment overhead, albeit with potential compromise in statistical consistency. Experiments across independent and identically distributed (i.i.d.), non-i.i.d., binary, multiclass, synthetic, and real-world datasets show that FIRMBOUND achieves optimalities in the sense of Bayes risk and speed-accuracy tradeoff. Furthermore, it advances the tradeoff boundary toward optimality when possible and reduces decision-time variance, ensuring reliable decision-making. Code is included in the supplementary materials. Akinori F. Ebihara, Taiki Miyagawa, Kazuyuki Sakurai, Hitoshi Imaoka |
ICLR | 4 |
| 2025 | ComFace: Facial Representation Learning with Synthetic Data for Comparing FacesabstractDaily monitoring of intra-personal facial changes associated with health and emotional conditions has great potential to be useful for medical, healthcare, and emotion recognition fields. However, the approach for capturing intra-personal facial changes is relatively unexplored due to the difficulty of collecting temporally changing face images. In this paper, we propose a facial representation learning method using synthetic images for comparing faces, called ComFace, which is designed to capture intra-personal facial changes. For effective representation learning, ComFace aims to acquire two feature representations, i.e., inter-personal facial differences and intra-personal facial changes. The key point of our method is the use of synthetic face images to overcome the limitations of collecting real intra-personal face images. Facial representations learned by ComFace are transferred to three extensive downstream tasks for comparing faces: estimating facial expression changes, weight changes, and age changes from two face images of the same individual. Our Com-Face, trained using only synthetic data, achieves comparable to or better transfer performance than general pretraining and state-of-the-art representation learning methods trained using real images. Yusuke Akamatsu, Terumi Umematsu, Hitoshi Imaoka, Shizuko Gomi, Hideo Tsurushima |
WACV | 3 |
| 2024 | CalibrationPhys: Self-Supervised Video-Based Heart and Respiratory Rate Measurements by Calibrating Between Multiple CamerasabstractVideo-based heart and respiratory rate measurements using facial videos are more useful and user-friendly than traditional contact-based sensors. However, most of the current deep learning approaches require ground-truth pulse and respiratory waves for model training, which are expensive to collect. In this paper, we propose CalibrationPhys, a self-supervised video-based heart and respiratory rate measurement method that calibrates between multiple cameras. CalibrationPhys trains deep learning models without supervised labels by using facial videos captured simultaneously by multiple cameras. Contrastive learning is performed so that the pulse and respiratory waves predicted from the synchronized videos using multiple cameras are positive and those from different videos are negative. CalibrationPhys also improves the robustness of the models by means of a data augmentation technique and successfully leverages a pre-trained model for a particular camera. Experimental results utilizing two datasets demonstrate that CalibrationPhys outperforms state-of-the-art heart and respiratory rate measurement methods. Since we optimize camera-specific models using only videos from multiple cameras, our approach makes it easy to use arbitrary cameras for heart and respiratory rate measurements. Yusuke Akamatsu, Terumi Umematsu, Hitoshi Imaoka |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | Blood Oxygen Saturation Estimation from Facial Video Via DC and AC Components of Spatio-Temporal MapabstractPeripheral blood oxygen saturation (SpO2), an indicator of oxygen levels in the blood, is one of the most important physiological parameters. Although SpO2 is usually measured using a pulse oximeter, non-contact SpO2 estimation methods from facial or hand videos have been attracting attention in recent years. In this paper, we propose an SpO2 estimation method from facial videos based on convolutional neural networks (CNN). Our method constructs CNN models that consider the direct current (DC) and alternating current (AC) components extracted from the RGB signals of facial videos, which are important in the principle of SpO2 estimation. Specifically, we extract the DC and AC components from the spatio-temporal map using filtering processes and train CNN models to predict SpO2 from these components. We also propose an end-to-end model that predicts SpO2 directly from the spatio-temporal map by extracting the DC and AC components via convolutional layers. Experiments using facial videos and SpO2 data from 50 subjects demonstrate that the proposed method achieves a better estimation performance than current state-of-the-art SpO2 estimation methods. Yusuke Akamatsu, Yoshifumi Onishi, Hitoshi Imaoka |
ICASSP | 3 |
| 2023 | Toward Asymptotic Optimality: Sequential Unsupervised Regression of Density Ratio for Early ClassificationabstractTheoretically-inspired sequential density ratio estimation (SDRE) algorithms are proposed for the early classification of time series. Conventional SDRE algorithms can fail to estimate DRs precisely due to the internal overnormalization problem, which prevents the DR-based sequential algorithm, Sequential Probability Ratio Test (SPRT), from reaching its asymptotic Bayes optimality. Two novel SPRT-based algorithms, B2Bsqrt-TANDEM and TANDEMformer, are designed to avoid the overnormalization problem for precise unsupervised regression of SDRs. The two algorithms statistically significantly reduce DR estimation errors and classification errors on an artificial sequential Gaussian dataset and real datasets (SiW, UCF101, and HMDB51), respectively. The code is available at: https://github.com/Akinori-F-Ebihara/LLR_saturation_problem. Akinori F. Ebihara, Taiki Miyagawa, Kazuyuki Sakurai, Hitoshi Imaoka |
ICASSP | 4 |
| 2023 | Edema Estimation From Facial Images Taken Before and After Dialysis via Contrastive Multi-Patient Pre-TrainingabstractEdema is a common symptom of kidney disease, and quantitative measurement of edema is desired. This paper presents a method to estimate the degree of edema from facial images taken before and after dialysis of renal failure patients. As tasks to estimate the degree of edema, we perform pre- and post-dialysis classification and body weight prediction. We develop a multi-patient pre-training framework for acquiring knowledge of edema and transfer the pre-trained model to a model for each patient. For effective pre-training, we propose a novel contrastive representation learning, called weight-aware supervised momentum contrast (WeightSupMoCo). WeightSupMoCo aims to make feature representations of facial images closer in similarity of patient weight when the pre- and post-dialysis labels are the same. Experimental results show that our pre-training approach improves the accuracy of pre- and post-dialysis classification by 15.1% and reduces the mean absolute error of weight prediction by 0.243 kg compared with training from scratch. The proposed method accurately estimate the degree of edema from facial images; our edema estimation system could thus be beneficial to dialysis patients. Yusuke Akamatsu, Yoshifumi Onishi, Hitoshi Imaoka, Junko Kameyama, Hideo Tsurushima |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | Heart Rate and Oxygen Saturation Estimation from Facial Video with Multimodal Physiological Data GenerationabstractEfforts to estimate multiple physiological parameters such as heart rate and oxygen saturation from facial videos have been made. However, training robust machine learning models for the estimation is challenging without large multimodal physiological datasets containing multiple physiological parameters and facial videos. In this paper, we propose a method to estimate heart rate and oxygen saturation from facial videos with multimodal physiological data generation. To collect sufficient datasets, the proposed method generates multimodal physiological datasets from several datasets containing a part of physiological modalities. Furthermore, to accurately estimate physiological parameters for unseen subjects, i.e., not included in the training data, we generate a multimodal physiological dataset for unseen subjects by using short facial videos of unseen subjects. Experimental results using three public datasets show the effectiveness of our multimodal physiological data generation. Yusuke Akamatsu, Yoshifumi Onishi, Hitoshi Imaoka |
ICASSP | 3 |
| 2021 | Joint Feature Distribution Alignment Learning for NIR-VIS and VIS-VIS Face RecognitionabstractFace recognition for visible light (VIS) images achieve high accuracy thanks to the recent development of deep learning. However, heterogeneous face recognition (HFR), which is a face matching in different domains, is still a difficult task due to the domain discrepancy and lack of large HFR dataset. Several methods have attempted to reduce the domain discrepancy by means of fine-tuning, which causes significant degradation of the performance in the VIS domain because it loses the highly discriminative VIS representation. To overcome this problem, we propose joint feature distribution alignment learning (JFDAL) which is a joint learning approach utilizing knowledge distillation. It enables us to achieve high HFR performance with retaining the original performance for the VIS domain. Extensive experiments demonstrate that our proposed method delivers statistically significantly better performances compared with the conventional fine-tuning approach on a public HFR dataset Oulu-CASIA NIR&VIS and popular verification datasets in VIS domain such as FLW, CFP, AgeDB. Furthermore, comparative experiments with existing state-of-the-art HFR methods show that our method achieves a comparable HFR performance on the Oulu-CASIA NIR&VIS dataset with less degradation of VIS performance. Takaya Miyamoto, Hiroshi Hashimoto, Akihiro Hayasaka, Akinori F. Ebihara, Hitoshi Imaoka |
IJCB | 5 |
| 2021 | Sequential Density Ratio Estimation for Simultaneous Optimization of Speed and Accuracy
Akinori F. Ebihara, Taiki Miyagawa, Kazuyuki Sakurai, Hitoshi Imaoka |
ICLR | 4 |
| 2020 | Specular- and Diffuse-reflection-based Face Spoofing Detection for Mobile DevicesabstractIn light of the rising demand for biometric-authentication systems, preventing face spoofing attacks is a critical issue for the safe deployment of face recognition systems. Here, we propose an efficient face presentation attack detection (PAD) algorithm that requires minimal hardware and only a small database, making it suitable for resource-constrained devices such as mobile phones. Utilizing one monocular visible light camera, the proposed algorithm takes two facial photos, one taken with a flash, the other without a flash. The proposed SpecDiff descriptor is constructed by leveraging two types of reflection: (i) specular reflections from the iris region that have a specific intensity distribution depending on liveness, and (ii) diffuse reflections from the entire face region that represents the 3D structure of a subject's face. Classifiers trained with SpecDiff descriptor outperforms other flash-based PAD algorithms on both an in-house database and on publicly available NUAA, Replay-Attack, and SiW databases. Moreover, the proposed algorithm achieves statistically significantly better accuracy to that of an end-to-end, deep neural network classifier, while being approximately six-times faster execution speed. The code is publicly available at https://github.com/Akinori-F-Ebihara/SpecDiff-spoofing-detector. Akinori F. Ebihara, Kazuyuki Sakurai, Hitoshi Imaoka |
IJCB | 3 |
| 2017 | Fast k-Nearest Neighbor Search for Face Identification Using Bounds of Residual ScoreabstractA novel fast k-nearest neighbor (k-NN) search method is proposed for the face identification task. It is well suited for this task because (1) it works well with high dimensionality, (2) it can be used with various similarity scores such as inner product, Euclidean distance, and correlation coefficient, (3) it can achieve not only fast exact k-NN search but much faster approximate search, and (4) it does not require any training or special data structure, resulting in low maintenance cost for the target database. Similarity scores between query and target samples are aggregated sequentially along with their dimensions, and target samples with no possibility of being included in k-NNs are rejected. The possibility is evaluated on the basis of the upper and lower bounds of the score for residual dimensions. Experimental results for a face database demonstrated that the proposed method achieves equal or better accuracy than other methods and is ten times faster than an exhaustive search with no degradation in the rank-k identification rate. Masato Ishii, Hitoshi Imaoka, Atsushi Sato |
FG | 2 |
| 2016 | Fast and accurate scale estimation method for object trackingabstractMany of the existing tracking methods do not estimate the object scale (width, height), only the location (x, y). In this paper we present a method which can accurately estimate the object scale given the location. The proposed approach works by cascading two methods together; such that each method refines the estimate by removing the false scale samples. Our method does not depend on the tracking technique and can be applied with any tracking system. We apply our approach to an existing tracker and compare the performance on benchmark sequences. The proposed method outperforms the existing tracker, while hardly affecting the speed. Karan Rampal, Kazuyuki Sakurai, Hitoshi Imaoka |
ICPR | 3 |
| 2015 | Occlusion handling in feature point tracking using ranked parts based modelsabstractA method for feature point tracking with partial occlusion handling is proposed. Occlusion causes distortion of entire face shape and not just the occluded part. To address this multiple models are learnt using regression, each aligning some part of the complete feature point set. A ranking SVM is then used to select the best feature points from among the aligned parts. The proposed method gives improved results compared to state of the art methods. Karon Rampal, Kazuyuki Sakurai, Hitoshi Imaoka |
ICIP | 3 |
| 2011 | Real-time face recognition demonstrationabstractIn recent years there have been great expectations of biometric authentication in view of increasing vicious crimes and terrorist threats. Face recognition is expected to be applied widely not only to security applications but also to image indexing and natural user interfaces. Accuracy of face recognition has been improved steadily in these years, but further improvements are demanded to meet performance requirements of these applications. We participated in Multiple Biometric Evaluation Still test conducted by National Institute of Standards and Technology (NIST) in 2010. In this evaluation, our algorithm achieved the best performance among all participants, with the highest identification rate of 95% among 1.8 million enrolled population, the lowest false match rate of 0.3% at false non-match rate 0.1%. In this demonstration, we show a real-time face recognition system using the above algorithm. Hitoshi Imaoka, Yusuke Morishita, Akihiro Hayasaka |
FG | 1 |
| 2004 | An Algorithm for the Detection of Faces on the Basis of Gabor Features and Information MaximizationabstractWe propose an algorithm for the detection of facial regions within input images. The characteristics of this algorithm are (1) a vast number of Gabor-type features (196,800) in various orientations, and with various frequencies and central positions, which are used as feature candidates in representing the patterns of an image, and (2) an information maximization principle, which is used to select several hundred features that are suitable for the detection of faces from among these candidates. Using only the selected features in face detection leads to reduced computational cost and is also expected to reduce generalization error. We applied the system, after training, to 42 input images with complex backgrounds (Test Set A from the Carnegie Mellon University face data set). The result was a high detection rate of 87.0%, with only six false detections. We compared the result with other published face detection algorithms. Hitoshi Imaoka, Kenji Okajima |
Neural Comput. | 1 |
| 2001 | A Complex Cell-Like Receptive Field Obtained by Information MaximizationabstractThe energy model (Pollen & Ronner, 1983; Adelson & Bergen, 1985) for a complex cell in the visual cortex is investigated theoretically. The energy model describes the output of a complex cell as the squared sum of outputs of two linear operators. An information-maximization problem to determine the two linear operators is investigated assuming the low signal-to-noise ratio limit and a localization term in the objective function. As a result, two linear operators characterized by a quadrature pair of Gabor functions are obtained as solutions. The result agrees with the energy model, which well describes the shift-invariant and orientation-selective responses of actual complex cells, and thus suggests that complex cells are optimally designed from an information-theoretic viewpoint. Kenji Okajima, Hitoshi Imaoka |
Neural Comput. | 2 |