VLDB 2026 Research / reviewers in the wild / expert
Hidetsugu Uchida
dblp:158/4248
· DBLP profile ↗
11ranked-venue papers
4as first author
4since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 3 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CP-VLM: Causal Prompting for Human Intention Inference with Vision-Language Models
Kazuki Osamura, Hidetsugu Uchida, Narishige Abe |
FG | 2 |
| 2024 | A Human-Centered Risk Evaluation of Biometric Systems Using Conjoint AnalysisabstractBiometric recognition systems, known for their convenience, are widely adopted across various fields. However, their security faces risks depending on the authentication algorithm and deployment environment. Current risk assessment methods faces significant challenges in incorporating the crucial factor of attacker’s motivation, leading to incomplete evaluations. This paper presents a novel human-centered risk evaluation framework using conjoint analysis to quantify the impact of risk factors, such as surveillance cameras, on attacker’s motivation. Our framework calculates risk values incorporating the False Acceptance Rate (FAR) and attack probability, allowing comprehensive comparisons across use cases. A survey of 600 Japanese participants demonstrates our method’s effectiveness, showing how security measures influence attacker’s motivation. This approach helps decision-makers customize biometric systems to enhance security while maintaining usability. Tetsushi Ohki, Narishige Abe, Hidetsugu Uchida, Shigefumi Yamada |
IJCB | 3 |
| 2024 | Multi-Masked Prompt Learning For Clothing-Change Person Re-IdentificationabstractClothing-change person re-identification (CC-ReID) aims to match persons even if they change clothes. Thus, extracting clothing-independent features is the key challenge in CC-ReID. Recently, many researches have primarily focused on auxiliary information to realize discriminative feature learning including soft-biometrics features, such as body shape and gaits, and additional clothes labels. However, this method does not fully capture personal information hidden in RGB images. Owing to pretrained vision-language models like CLIP, textual information can describe a person in such a way that it includes all details to capturing personal characteristics, e.g., gender, age, hairstyle, glass, and clothes. Considering that certain personal textual information remains unchanged in CC-ReID, we propose a novel CC-ReID method called Multi-Masked Prompt Learning(MMPL), which takes full advantage of the ability of CLIP to extract clothing-independent features for CC-ReID. MMPL can extract clothing-independent textual and visual features via Prompt Mask and Visual Mask, respectively. This approach enables to discard clothing information from the trained model. The proposed MMPL outperforms other state-of-the-art approaches, improving the rank-1 accuracy to 1.5% and 1.0% on PRCC and LTCC datasets, respectively. Kazuki Osamura, Hidetsugu Uchida, Shijie Nie, Narishige Abe |
IJCB | 2 |
| 2022 | Sample-Level and Class-Level Adaptive Training for Face RecognitionabstractMarginal softmax loss function has been widely used for face recognition, where a universal angular margin is added between weight prototypes. However, this method neglects the differences between classes and samples. On class-level, the imbalanced real world training dataset requires different margin for the head and tail classes to equally squeeze each class's feature space. On the sample-level, it's also necessary to assign larger importance for the hard samples during training. In this paper, we address these two issues by combining two strategies: (1) explicitly assign the adaptive margin according to the image quantity so that the margin is enlarged for the tail classes; (2) semantically identify the ‘hard positive, samples and misclassified samples [1] to attach adaptive weights to increase the training emphasis on these samples. Extensive experiments on LFW/CFP/AGEDB and IJB-B/IJB-C show our method's effectiveness. Mengjiao Wang 0001, Rujie Liu, Narishige Abe, Tomoaki Matsunami, Hidetsugu Uchida, Lina Septiana |
ICME | 5 |
| 2018 | Discover the Effective Strategy for Face Recognition Model Compression by Improved Knowledge DistillationabstractFor the sake of better accuracy, the face recognition model is becoming larger and larger, which makes them difficult to be deployed on embedded systems. This work proposes an effective model compression method using knowledge distillation, where a fast student model is trained under the guidance of a complex teacher model. Firstly, different loss combinations and network architectures are analyzed through comprehensive experiments to find the most effective approach. To augment the performance, the feature layer is further normalized to make the optimization objective consistent with cosine similarity metric. Moreover, a teacher weighting strategy is proposed to address the issue when teacher provides wrong guidance. Experimental results show that the student model built by our approach can surpass the teacher model while achieving 3× acceleration. Mengjiao Wang 0001, Rujie Liu, Narishige Abe, Hidetsugu Uchida, Tomoaki Matsunami, Shigefumi Yamada |
ICIP | 4 |
| 2017 | Parallel-Data-Free Many-to-Many Voice Conversion Based on DNN Integrated with Eigenspace Using a Non-Parallel Speech Corpus
Tetsuya Hashimoto, Hidetsugu Uchida, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 2 |
| 2017 | Acoustic-to-Articulatory Mapping Based on Mixture of Probabilistic Canonical Correlation Analysis
Hidetsugu Uchida, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 1 |
| 2016 | Prediction of the Articulatory Movements of Unseen Phonemes of a Speaker Using the Speech Structure of Another Speaker
Hidetsugu Uchida, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 1 |
| 2016 | Voice Conversion Based on Matrix Variate Gaussian Mixture Model Using Multiple Frame Features
Hidetsugu Uchida, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 2 |
| 2015 | Statistical acoustic-to-articulatory mapping unified with speaker normalization based on voice conversionabstractThis paper proposes a model of speaker-normalized acoustic-toarticulatory mapping using statistical voice conversion. A mapping function from acoustic parameters to articulatory parameters is usually developed with a single speaker’s parallel data. Hence the constructed mapping model can work appropriately only for this specific speaker, and applying this model to other speakers degrades the performance of acoustic-to-articulatory mapping. In this paper, two models of speaker conversion and acoustic-to-articulatory mapping are implemented using Gaussian Mixture Models (GMM), and by integrating these two models, we propose two methods of speaker-normalized acoustic-to-articulatory mapping. One is concatenating these models sequentially, and the other integrates the two models into a unified model, where acoustic parameters of a speaker can be converted directly to articulatory parameters of another speaker. Experiments show that both methods can improve the mapping accuracy and that the latter method works better than the former method. Especially in the case of velar stop consonants, the mapping accuracy is higher by 0.6 mm. Index Terms: acoustic-to-articulatory mapping, Gaussian mixture model, voice conversion, speaker normalization Hidetsugu Uchida, Daisuke Saito, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 1 |
| 2014 | A study on the improvement of measurement accuracy of the three-dimensional electromagnetic articulography
Hidetsugu Uchida, Kohei Wakamiya, Tokihiko Kaburagi |
INTERSPEECH | 1 |