VLDB 2026 Research / reviewers in the wild / expert
Ko Watanabe 0001
dblp:148/6442-1
· DBLP profile ↗
7ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0003-0252-1785ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Expanding the Range of Pre-Trained Gaze Estimation via Lightweight On-Device Fine-TuningabstractAppearance-based gaze estimation models exhibit degraded accuracy when deployed on devices not present in the training data. For instance, iTracker, trained exclusively on the GazeCapture dataset collected from smartphones and tablets, shows increased error in the lower screen regions on laptops due to out-of-distribution downward gaze angles. We propose a lightweight fine-tuning approach that adapts a CoreML-format model using 13-point calibration data collected in approximately 30 seconds. In an evaluation with 11 participants, we compared four conditions: baseline, homography transformation, cross-user fine-tuning, and user-specific fine-tuning. User-specific fine-tuning reduced the mean estimation error from 6.17 cm to 2.91 cm (52.8% reduction), outperforming homography transformation (3.38 cm). These results demonstrate that fine-tuning only the final layer can effectively extend pre-trained gaze estimation models to unseen device configurations. Soki Kokado, Ko Watanabe 0001, Shoya Ishimaru |
ETRA | 2 |
| 2025 | TKG-DM: Training-free Chroma Key Content Generation Diffusion ModelabstractDiffusion models have enabled the generation of high-quality images with a strong focus on realism and textual fidelity. Yet, large-scale text-to-image models, such as Stable Diffusion, struggle to generate images where foreground objects are placed over a chroma key background, limiting their ability to separate foreground and background elements without fine-tuning. To address this limitation, we present a novel Training-Free Chroma Key Content Generation Diffusion Model (TKG-DM), which optimizes the initial random noise to produce images with foreground objects on a specifiable color background. Our proposed method is the first to explore the manipulation of the color aspects in initial noise for controlled background generation, enabling precise separation of foreground and background without fine-tuning. Extensive experiments demonstrate that our training-free method outperforms existing methods in both qualitative and quantitative evaluations, matching or surpassing fine-tuned models. Finally, we successfully extend it to other tasks (e.g., consistency models and text-to-video), highlighting its transformative potential across various generative applications where independent control of foreground and background is crucial. Ryugo Morita, Stanislav Frolov, Brian B. Moser, Takahiro Shirakawa, Ko Watanabe 0001, Andreas Dengel 0001, Jinjia Zhou |
CVPR | 5 |
| 2025 | PupilSense: A Novel Application for Webcam-Based Pupil Diameter Estimation
Vijul Shah, Ko Watanabe 0001, Brian B. Moser, Andreas Dengel 0001 |
ETRA | 2 |
| 2025 | Webcam-Based Pupil Diameter Prediction Benefits from Upscaling
Vijul Shah, Brian B. Moser, Ko Watanabe 0001, Andreas Dengel 0001 |
ICAART (2) | 3 |
| 2024 | Feature Estimation of Global Language Processing in EEG Using Attention Maps
Dai Shimizu, Ko Watanabe 0001, Andreas Dengel 0001 |
ACCV (2) | 2 |
| 2024 | Eye Movement in a Controlled Dialogue SettingabstractDesigning realistic eye movements for animated avatars poses a challenge, as gaze behavior is predominantly unconscious. Accurately modulating those movements is crucial to avoid the Uncanny Valley. The human gaze exhibits different characteristics in conversations, depending on speaking or listening. Albeit these distinctions are known, data for synthesizing eye movement models suitable for avatars is scarce. This research introduces a novel dataset involving human gaze behavior during remote screen conversations. The data are collected from 19 participants, offering 4 hours of gaze data labeled as Speaking and Listening. Our data analysis substantiates prior knowledge of gaze behavior while providing new insights through higher precision. Furthermore, we demonstrate the dataset’s suitability for machine learning algorithms by training a classifier, achieving 88.1% binary classification accuracy. David Dembinsky, Ko Watanabe 0001, Andreas Dengel 0001, Shoya Ishimaru |
ETRA | 2 |
| 2021 | A Privacy-Aware Browser Extension to Track User Search Behavior for Programming Course Supplement
Jihed Makhlouf, Yutaka Arakawa, Ko Watanabe 0001 |
MobiQuitous | 3 |