EDBT 2026 Demo / reviewers in the wild / expert
Keita Suzuki
dblp:139/4913
· DBLP profile ↗
19ranked-venue papers
8as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating the Feasibility of Superposed QR Code Attacks Using High Refresh Rate DisplaysabstractQuick Response (QR) codes are widely used for encoding information accessible via mobile devices. However, since QR codes appear as a complex pattern of black and white modules, it is difficult for users to detect tampering. In this study, we propose a camouflaged QR code attack method utilizing high refresh rate displays. This method rapidly alternates two QR codes above the critical flicker fusion frequency of human vision, so humans perceive a single fused image, whereas cameras may capture either code. Preliminary experiments using LCD displays confirmed that the attack is feasible, but revealed a constraint on module modification due to LCD characteristics. To overcome this limitation, we applied the method to off-the-shelf LED displays and confirmed that the LCD-related constraint was resolved, enabling more flexible code modification with QR codes. Keita Suzuki, Kentaro Fukuchi |
AVI | 1 |
| 2025 | ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of MindabstractExisting Theory of Mind (ToM) benchmarks diverge from real-world scenarios in three aspects: 1) they assess a limited range of mental states such as beliefs, 2) false beliefs are not comprehensively explored, and 3) the diverse personality traits of characters are overlooked. To address these challenges, we introduce ToMATO, a new ToM benchmark formulated as multiple-choice QA over conversations. ToMATO is generated via LLM-LLM conversations featuring information asymmetry. By employing a prompting method that requires role-playing LLMs to verbalize their thoughts before each utterance, we capture both first- and second-order mental states across five categories: belief, intention, desire, emotion, and knowledge. These verbalized thoughts serve as answers to questions designed to assess the mental states of characters within conversations. Furthermore, the information asymmetry introduced by hiding thoughts from others induces the generation of false beliefs about various mental states. Assigning distinct personality traits to LLMs further diversifies both utterances and thoughts. ToMATO consists of 5.4k questions, 753 conversations, and 15 personality trait patterns. Our analysis shows that this dataset construction approach frequently generates false beliefs due to the information asymmetry between role-playing LLMs, and effectively reflects diverse personalities. We evaluate nine LLMs on ToMATO and find that even GPT-4o mini lags behind human performance, especially in understanding false beliefs, and lacks robustness to various personality traits. Kazutoshi Shinoda, Nobukatsu Hojo, Kyosuke Nishida, Saki Mizuno, Keita Suzuki, Ryo Masumura, Hiroaki Sugiyama, Kuniko Saito |
AAAI | 5 |
| 2025 | Correntropy-Based Improper Likelihood Model for Robust Electrophysiological Source ImagingabstractBayesian learning provides a unified skeleton to solve the electrophysiological source imaging task. From this perspective, existing source imaging algorithms utilize the Gaussian assumption for the observation noise to build the likelihood function for Bayesian inference. However, the electromagnetic measurements of brain activity are usually affected by miscellaneous artifacts, leading to a potentially non-Gaussian distribution for the observation noise. Hence the conventional Gaussian likelihood model is a suboptimal choice for the real-world source imaging task. In this study, we aim to solve this problem by proposing a new likelihood model which is robust with respect to non-Gaussian noises. Motivated by the robust maximum correntropy criterion, we propose a new improper distribution model concerning the noise assumption. This new noise distribution is leveraged to structure a robust likelihood function and integrated with hierarchical prior distributions to estimate source activities by variational inference. In particular, the score matching is adopted to determine the hyperparameters for the improper likelihood model. A comprehensive performance evaluation is performed to compare the proposed noise assumption to the conventional Gaussian model. Simulation results show that, the proposed method can realize more precise source reconstruction by designing known ground-truth. The real-world dataset also demonstrates the superiority of our new method with the visual perception task. This study provides a new backbone for Bayesian source imaging, which would facilitate its application using real-world noisy brain signal. Yuanhao Li 0004, Badong Chen, Zhongxu Hu, Keita Suzuki, Wenjun Bai, Yasuharu Koike, Okito Yamashita |
IEEE Trans. Medical Imaging | 4 |
| 2024 | Talking Face Generation for Impression Conversion Considering Speech SemanticsabstractThis study investigates the talking face generation method to convert a speaker’s video to give a target impression, such as “favorable” or “considerate”. Such an impression conversion method needs to consider the input speech semantics because they affect the impression of a speaker’s video along with the facial expression. Conventional emotional talking face generation methods utilize speech information to synchronize the lip and speech of the output video. However, they cannot consider speech semantics because the speech representations contain only phonetic information. To solve this problem, we propose a facial expression conversion model that uses a semantic vector obtained from BERT embeddings of speech recognition results of input speech. We first constructed an audio-visual dataset with impression labels assigned to each utterance. The evaluation results based on the dataset showed that the proposed method could improve the estimation accuracy of the facial expressions of the target video. Saki Mizuno, Nobukatsu Hojo, Kazutoshi Shinoda, Keita Suzuki, Mana Ihori, Tomohiro Tanaka, Naotaka Kawata, Satoshi Kobashikawa, Ryo Masumura |
ICASSP | 4 |
| 2024 | Optimal criterion for feature learning of two-layer linear neural network in high dimensional interpolation regimeabstractDeep neural networks with feature learning have shown surprising generalization performance in high dimensional settings, but it has not been fully understood how and when they enjoy the benefit of feature learning. In this paper, we theoretically analyze the statistical properties of the benefits from feature learning in a two-layer linear neural network with multiple outputs in a high-dimensional setting. For that purpose, we propose a new criterion that allows feature learning of a two-layer linear neural network in a high-dimensional setting. Interestingly, we can show that models with smaller values of the criterion generalize even in situations where normal ridge regression fails to generalize. This is because the proposed criterion contains a proper regularization for the feature mapping and acts as an upper bound on the predictive risk. As an important characterization of the criterion, the two-layer linear neural network that minimizes this criterion can achieve the optimal Bayes risk that is determined by the distribution of the true signals across the multiple outputs. To the best of our knowledge, this is the first study to specifically identify the conditions under which a model obtained by proper feature learning can outperform normal ridge regression in a high-dimensional multiple-output linear regression problem. Keita Suzuki, Taiji Suzuki |
ICLR | 1 |
| 2024 | Unified Multi-Talker ASR with and without Target-speaker Enrollment
Ryo Masumura, Naoki Makishima, Tomohiro Tanaka, Mana Ihori, Naotaka Kawata, Shota Orihashi, Kazutoshi Shinoda, Taiga Yamane, Saki Mizuno, Keita Suzuki, Nobukatsu Hojo, Takafumi Moriya, Atsushi Ando |
INTERSPEECH | 10 |
| 2024 | Learning from Multiple Annotator Biased Labels in Multimodal Conversation
Kazutoshi Shinoda, Nobukatsu Hojo, Saki Mizuno, Keita Suzuki, Satoshi Kobashikawa, Ryo Masumura |
INTERSPEECH | 4 |
| 2024 | Participant-Pair-Wise Bottleneck Transformer for Engagement Estimation from Video Conversation
Keita Suzuki, Nobukatsu Hojo, Kazutoshi Shinoda, Saki Mizuno, Ryo Masumura |
INTERSPEECH | 1 |
| 2024 | Balancing Analysis Time and Bug Detection: Daily Development-friendly Bug Detection in Linux
Keita Suzuki, Kenta Ishiguro, Kenji Kono |
USENIX ATC | 1 |
| 2023 | The effect of implicit theories on help-seeking behavior: Focusing on anticipated evaluation and perceived implicit theories of the peer member
Keita Suzuki, Tomoki Fujiwara, Yukiko Muramoto |
CogSci | 1 |
| 2023 | OnDA-DETR: Online Domain Adaptation for Detection Transformers with Self-Training FrameworkabstractThis paper presents a novel method for online domain adaptation (OnDA) for DEtection TRansformer (DETR)-based object detection models called OnDA-DETR. OnDA is a domain adaptation paradigm that adapts a model trained on the source domain data to perform well on the target domain in an online manner during testing, using only the unlabeled test data from the target domain. Due to challenging and realistic problem settings, OnDA has garnered significant attention. However, OnDA methods for DETR-based models, which have demonstrated excellent performance in object detection research fields, had not been developed. OnDA-DETR is the first OnDA method specifically designed for DETR-based models. OnDA-DETR incorporates a self-training framework that generates pseudo-labels for the unlabeled target domain data. To effectively incorporate the self-training framework into DETR-based models, we leverage recall-aware pseudo-labeling and quality-aware training in OnDA-DETR. Experimental results indicate that OnDA-DETR improves the performance of the source-trained model by about 3.0 % points through OnDA. Taiga Yamane, Naoki Makishima, Keita Suzuki, Atsushi Ando, Ryo Masumura |
ICIP | 4 |
| 2023 | Joint Autoregressive Modeling of End-to-End Multi-Talker Overlapped Speech Recognition and Utterance-level Timestamp Prediction
Naoki Makishima, Keita Suzuki, Atsushi Ando, Ryo Masumura |
INTERSPEECH | 2 |
| 2023 | End-to-End Joint Target and Non-Target Speakers ASR
Ryo Masumura, Naoki Makishima, Taiga Yamane, Yoshihiko Yamazaki, Saki Mizuno, Mana Ihori, Mihiro Uchida, Keita Suzuki, Hiroshi Sato 0002, Tomohiro Tanaka, Akihiko Takashima, Takafumi Moriya, Nobukatsu Hojo, Atsushi Ando |
INTERSPEECH | 8 |
| 2023 | Multi-region CNN-Transformer for Micro-gesture Recognition in Face and Upper BodyabstractThis paper presents a novel task that recognizes from a video unintentional micro-gestures (UMGs), which are movements made by people unconsciously and unintentionally. Recognizing UMGs is crucial because they reveal a person’s underlying psychological state. Since a UMG is composed of subtle sequential movements, the recognition model must be able to capture accurate information in both the spatial and temporal directions. Therefore, we utilize a convolutional neural network (CNN) to capture information in the spatial direction and a Transformer to merge the features extracted by the CNN in the temporal direction. However, this model often misrecognizes UMGs because it is not possible to capture slight differences in movements, such as in the face and mouth regions. To address this issue, we propose a novel model for UMG recognition, the Multi-Region CNN-Transformer model, that inputs cropped videos from multiple upper body regions simultaneously. The key advance of our method is to capture subtle changes in regions such as the upper body, face, head, and mouth for recognizing UMGs. We demonstrate the effectiveness of the proposed method through experiments using our newly created UMG dataset for this task. Keita Suzuki, Ryo Masumura, Atsushi Ando, Naoki Makishima |
MMAsia | 1 |
| 2022 | On the Use of Modality-Specific Large-Scale Pre-Trained Encoders for Multimodal Sentiment AnalysisabstractThis paper investigates the effectiveness and implementation of modality-specific large-scale pre-trained encoders for multimodal sentiment analysis (MSA). Although the effectiveness of pre-trained encoders in various fields has been reported, conventional MSA methods employ them for only linguistic modality, and their application has not been investigated. This paper compares the features yielded by large-scale pre-trained encoders with conventional heuristic features. One each of the largest pre-trained encoders publicly available for each modality are used; CLIP-ViT, WavLM, and BERT for visual, acoustic, and linguistic modalities, respectively. Experiments on two datasets reveal that methods with domain-specific pre-trained encoders attain better performance than those with conventional features in both unimodal and multimodal scenarios. We also find it better to use the outputs of the intermediate layers of the encoders than those of the output layer. The codes are available at https://github.com/ando-hub/MSA_Pretrain. Atsushi Ando, Ryo Masumura, Akihiko Takashima, Naoki Makishima, Keita Suzuki, Takafumi Moriya, Takanori Ashihara, Hiroshi Sato 0002 |
SLT | 6 |
| 2020 | LSTM Neural Network for Fine-Granularity Estimation on Baseline Load of Fast Demand Response
Shun Matsukawa, Keita Suzuki, Chuzo Ninagawa, Junji Morikawa, Seiji Kondo |
EANN | 2 |
| 2019 | Detecting and Analyzing Year 2038 Problem Bugs in User-Level ApplicationsabstractThe year 2038 problem is a well-known year problem that might cause severe damage to many existing software systems. However, no current tool can detect the bugs since it requires the understandings of the problem unique encoding semantics. In this paper, we analyze real-world applications and raise the alarm over the fact that the Year 2038 problem is a real threat. We target all of the C based projects uploaded on GitHub in the years 2012 to 2018 (32,921 in total), between the dates July 1 to July 10. Our analysis shows that 7.35% of the compiled projects have bugs. Some of the bugs trigger undefined behavior and are dangerous enough to crash the software systems. Our bug fixing patches sent to six projects have been confirmed and approved, including large-scale, real-world projects such as the Amazon Web Service support tools and the Linux Test Project. Keita Suzuki, Takafumi Kubota, Kenji Kono |
PRDC | 1 |
| 2017 | FaceShare: Mirroring with Pseudo-Smile Enriches Video Chat Communicationsabstract"Mirroring" refers to the unconscious mimicry of another person's behaviors, such as their facial expressions. Mirroring has many positive effects, such as enhancing closeness and improving the flow of a conversation, which enriches the quality of communication. Our study set out to devise a means of evoking these positive effects in a video chat without any conscious effort of participants. We constructed a videophone system, called FaceShare, which can deform the user's face into a smile in response to their partner's smiling. That is, our system generates mirroring by producing a pseudo-smile through image processing. We conducted an experiment in which pairs of participants had brief conversations via FaceShare. The results implied that mirroring using the pseudo-smile lets the mimicker, whose face is deformed according to the expressions of their partner, feel a closeness, and improves the flow of the conversation for both the mimicker and the mimickee, who sees the mimicker's deformed face. Keita Suzuki, Masanori Yokoyama, Shigeo Yoshida, Takayoshi Mochizuki, Tomohiro Yamada, Takuji Narumi, Tomohiro Tanikawa, Michitaka Hirose |
CHI | 1 |
| 2016 | Need and impressions of communication robots for seniors with slight physical and cognitive disabilities: Evaluation using system usability scaleabstractVarious types of communication robots have emerged for supporting elderly people with impairments or frailties. However, the need of the robots for healthy elderly persons with decreasing physical and cognitive abilities remains unclear. Knowledge about these needs can help determine the guidelines or implications of the introduction of communication robots for healthy seniors. The objective of this study is to clarify the needs and the impressions of various communication robots for seniors with slight physical or cognitive problems. The scores on the system usability scale (SUS) and interview results indicated that, with increasing participant age, mammal-like communication robots were more preferable than other types. In addition, elderly persons with minor physical problems tended not to prefer manually operated robots. Takahiro Miura, Taichi Goto, Kazuki Kaneko, Yuka Sumikawa, Ayako Ishii, Mio Doke, Keita Suzuki, Taiyu Okatani, Akihiro Kubota, Mingzhen Zhang, Yuki Kinoshita, Hazuki Yoshinaga, Masahiro Tsuruta, Yuri Kominami, Misato Nihei, Takenobu Inoue, Minoru Kamata, Junichiro Okata |
SMC | 7 |