Kazuya Ueki

dblp:15/2806 · DBLP profile ↗
← Back
12ranked-venue papers
9as first author
7since 2021 · last 2026
0009-0005-1691-1858ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 8 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 5 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 U-Cker at VBS2026: A Web-Based Interactive Video Retrieval System with Multimodal Query Support
Kazuya Ueki, Ryo Mutou, Takuya Wada, Ryota Akaba
MMM (4)1
2025 U-Cker: Initial Development of an Interactive Video Retrieval System for Novice Users
abstract
We present U-Cker, an interactive video retrieval system designed for novice users, developed for the VR4B (Video Retrieval for Beginners) special session at the CBMI conference. The system builds on our previous developments for the TRECVID video retrieval benchmark, where we solely focused on search accuracy and did not consider real-time operation or usability. U-Cker addresses this limitation by enabling real-time interaction while maintaining high retrieval performance. Furthermore, the user interface has been significantly simplified to ensure accessibility for users with no prior experience in video retrieval. The system runs entirely in a web browser, which does not require additional software installation and provides easy access for all users.
Kazuya Ueki, Ryo Muto, Takuya Wada, Ryota Akaba, Genesis Faith Fernandez
CBMI1
2025 Optimization of CLIP Models for Domain-Specific Video Search
abstract
Pre-trained vision-language multimodal models, such as CLIP, have demonstrated effectiveness across a wide range of tasks, including image recognition, image retrieval, and video analysis. This study explores fine-tuning these models specifically for video retrieval using the large-scale video dataset provided by the TRECVID benchmark. Evaluations carried out on various models available in OpenCLIP showed significant improvements in retrieval accuracy in all models tested, confirming the effectiveness of the proposed approach.
Kazuya Ueki, Haruki Sato, Yuma Suzuki, Takayuki Hori, Hiroki Takushima, Takumi Takada, Hayato Tanoue, Aiswariya Manoj Kumar, Hiroki Nishihara, Yuki Shibata
ICMR1
2024 Direct Metric Optimization for Image Captioning through Reward-Weighted Augmented Data Utilization
abstract
Takumi Takada, Yuma Suzuki, Hiroki Takushima, Hayato Tanoue, Haruki Sato, Aiswariya Kumar, Hiroki Nishihara, Takayuki Hori, Kazuya Ueki. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Takumi Takada, Yuma Suzuki, Hiroki Takushima, Hayato Tanoue, Haruki Sato, Aiswariya Manoj Kumar, Hiroki Nishihara, Takayuki Hori, Kazuya Ueki
ACL (1)9
2024 Elevating Video Retrieval Capabilities: A Cross-Modal Approach Utilizing Text and Image Generative Models
abstract
In this study, we propose a new methodology to enhance the accuracy of video retrieval by integrating rapidly developing text and image generation technologies. The core of our study is to transform search queries into a more diverse range of search requirements using ChatGPT, and convert them into more specific visual features using stable diffusion, thereby accurately retrieving video content that is highly relevant to search queries. In our experiments, we used a large-scale video dataset from the TRECVID benchmark, which is considered the industry standard as concerns diversity and inclusiveness, to validate the effectiveness of our method. The combination of the advanced natural language understanding capabilities of ChatGPT and image generation prowess of stable diffusion enabled us to successfully retrieve highly relevant search results, even for complicated queries.
Kazuya Ueki, Yuma Suzuki, Haruki Sato, Takayuki Hori, Takumi Takada, Hiroki Takushima, Hayato Tanoue, Aiswariya Manoj Kumar, Hiroki Nishihara
CBMI1
2024 Waseda_Meisei_SoftBank at Video Browser Showdown 2024
Takayuki Hori, Kazuya Ueki, Yuma Suzuki, Hiroki Takushima, Hayato Tanoue, Haruki Sato, Takumi Takada, Aiswariya Manoj Kumar
MMM (4)2
2021 Survey of Visual-Semantic Embedding Methods for Zero-Shot Image Retrieval
abstract
Visual-semantic embedding is an interesting research topic because it is useful for various tasks, such as visual question answering (VQA), image-text retrieval, image captioning, and scene graph generation. In this paper, we focus on zero-shot image retrieval using sentences as queries and present a survey of the technological trends in this area. First, we provide a comprehensive overview of the history of the technology, starting with a discussion of the early studies of image-to-text matching and how the technology has evolved over time. In addition, a description of the datasets commonly used in experiments and a comparison of the evaluation results of each method are presented. We also introduce the implementation available on github for use in confirming the accuracy of experiments and for further improvement. We hope that this survey paper will encourage researchers to further develop their research on bridging images and languages.
Kazuya Ueki
ICMLA1
2020 Japanese Cursive Character Recognition for Efficient Transcription
Kazuya Ueki, Tomoka Kojima
ICPRAM1
2018 Fine-grained Video Retrieval using Query Phrases - Waseda_Meisei TRECVID 2017 AVS System -
abstract
In this paper, a joint team from Waseda University and Meisei University (team name: Waseda_Meisei) report their efforts on the ad-hoc video search (AVS) task for the TRECVID benchmark, which is conducted annually by the National Institute of Standards and Technology (NIST). For the AVS task, a system is required to perform a fine-grained search of target videos from a large-scale video database using a query phrase including multiple keywords, such as objects, persons, scenes, and actions. The system we submitted has the following two characteristics. First, to improve the coverage rate of classes corresponding to keywords in query phrases, we prepared a large number of classifiers that can detect objects, persons, scenes, and actions, which were trained using various image and video datasets. Second, when choosing a concept classifier corresponding to a keyword, we introduced a mechanism that allows us to select additional concept classifiers by incorporating natural language processing techniques. We submitted multiple systems with these characteristics to the TRECVID 2017 AVS task and one of our systems ranked the highest among all the submitted systems from 22 teams.
Kazuya Ueki, Koji Hirakawa, Kotaro Kikuchi, Tetsunori Kobayashi
ICPR1
2016 Improving semantic video indexing: Efforts in Waseda TRECVID 2015 SIN system
abstract
In this paper, we propose a method for improving the performance of semantic video indexing. Our approach involves extracting features from multiple convolutional neural networks (CNNs), creating multiple classifiers, and integrating them. We employed four measures to accomplish this: (1) utilizing multiple evidences observed in each video and effectively compressing them into a fixed-length vector; (2) introducing gradient and motion features to CNNs; (3) enriching variations of the training and the testing sets; and (4) extracting features from several CNNs trained with various large-scale datasets. Using the test dataset from TRECVID's 2014 evaluation benchmark, we evaluated the performance of the proposal in terms of the mean extended inferred average precision measure. On this measure, our system's performance was 35.7, outperforming the state-of-the-art TRECVID 2014 benchmark performance of 33.2. Based on this work, our submission at TRECVID 2015 was ranked second among all submissions.
Kazuya Ueki, Tetsunori Kobayashi
ICASSP1
2010 Perceived Age Estimation under Lighting Condition Change by Covariate Shift Adaptation
abstract
Over the recent years, a great deal of effort has been made to age estimation from face images. It has been reported that age can be accurately estimated under controlled environment such as frontal faces, no expression, and static lighting conditions. However, it is not straightforward to achieve the same accuracy level in real-world environment because of considerable variations in camera settings, facial poses, and illumination conditions. In this paper, we apply a recently-proposed machine learning technique called covariate shift adaptation to alleviating lighting condition change between laboratory and practical environment. Through real-world age estimation experiments, we demonstrate the usefulness of our proposed method.
Kazuya Ueki, Masashi Sugiyama, Yasuyuki Ihara
ICPR1
2000 On-Line Estimation of Hidden Markov Model Parameters
Jun Mizuno, Tasuya Watanabe, Kazuya Ueki, Kazuyuki Amano, Eiji Takimoto, Akira Maruoka
Discovery Science3