EDBT 2026 Demo / reviewers in the wild / expert
Bowen Wu 0002
dblp:156/8281-2
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SignFlow: End-to-End Sign Language Generation for One-to-Many Modeling using Conditional Flow Matching
Khan Nabeela Khanum, Bowen Wu 0002, Sihan Tan, Carlos Toshinori Ishi, Kazuhiro Nakadai |
ICMI | 2 |
| 2025 | MultiGAU: Real Time Sign Language Generation Using Multimodal Gated Attention
Khan Nabeela Khanum, Bowen Wu 0002, Carlos Toshinori Ishi, Kazuhiro Nakadai |
IEA/AIE (1) | 2 |
| 2025 | HAM-GNN: A hierarchical attention-based multi-dimensional edge graph neural network for dialogue act classification
Changzeng Fu, Yikai Su, Kaifeng Su, Yinghao Liu, Bowen Wu 0002, Carlos Toshinori Ishi, Hiroshi Ishiguro |
Expert Syst. Appl. | 6 |
| 2024 | Retargeting Human Facial Expression to Human-like Robotic Face through Neural Network Surrogate-based OptimizationabstractFacial mimicry is crucial for human-like robots in human-robot interaction. The challenge is that the high diversity of facial expressions proposes difficulties in programming a robotic face to mimic human facial expressions using traditional methods. In this paper, we present a data-driven method to retarget human facial expressions to robotic faces without human effort. Our data collection is fully automatic, where only a robotic face and Apple ARKit are involved to sample actuator commands and record the resulting facial blendshape values. We trained a neural network that predicts blendshape values from commands, which is then used as a surrogate model to optimize command values to resemble given facial expressions. Experiments show that the proposed method has achieved lower error in terms of facial blendshape values than baselines. Moreover, the response time can be reduced to 0.2 seconds via TCP/IP through WiFi, offering great potential for real-time application. Our method is a novel framework for retargeting facial expressions to robotic faces, which can be incorporated into various human-robot interaction systems. Bowen Wu 0002, Carlos Toshinori Ishi, Takashi Minato, Hiroshi Ishiguro |
IROS | 1 |
| 2024 | Speech-Driven Gesture Generation Using Transformer-Based Denoising Diffusion Probabilistic ModelsabstractWhile it is crucial for human-like avatars to perform co-speech gestures, existing approaches struggle to generate natural and realistic movements. In the present study, a novel transformer-based denoising diffusion model is proposed to generate co-speech gestures. Moreover, we introduce a practical sampling trick for diffusion models to maintain the continuity between the generated motion segments while improving the within-segment motion likelihood and naturalness. Our model can be used for online generation since it generates gestures for a short segment of speech, e.g., 2 s. We evaluate our model on two large-scale speech-gesture datasets with finger movements using objective measurements and a user study, showing that our model outperforms all other baselines. Our user study is based on the Metahuman platform in the Unreal Engine, a popular tool for creating human-like avatars and motions. Bowen Wu 0002, Carlos Toshinori Ishi, Hiroshi Ishiguro |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2023 | HAG: Hierarchical Attention with Graph Network for Dialogue Act Classification in ConversationabstractThe prediction of dialogue acts (DA) labels on utterance-level in conversations can be treated as a sequence labeling problem, which requires context- and speaker-aware semantic comprehension, especially for Japanese. In this study, we pro-posed a hierarchical attention with the graph neural network (HAG) to consider the contextual interconnections as well as the semantics carried by the sentence itself. Concretely, the model use long-short term memory networks (LSTMs) to perform a context-aware encoding within a dialogue window. Then, we construct the context graph by aggregating the neighboring utterances. Subsequently, a speaker feature transformation is executed with a graph attention network (GAT) to calculate the interconnections, while a context-level feature selection is performed with a gated graph convolutional network (GatedGCN) to select the salient utterances that contribute to the DA classification. Finally, we merge the representations of different levels and conduct a classification with two dense layers. We evaluate the proposed model on Japanese dialogue act dataset (JPS-DA). The experimental results show that our method outperforms the baselines. Changzeng Fu, Zhenghan Chen, Bowen Wu 0002, Carlos Toshinori Ishi, Hiroshi Ishiguro |
ICASSP | 4 |
| 2023 | Recognizing Real-World Intentions using A Multimodal Deep Learning Approach with Spatial-Temporal Graph Convolutional NetworksabstractIdentifying intentions is a critical task for comprehending the actions of others, anticipating their future behavior, and making informed decisions. However, it is challenging to recognize intentions due to the uncertainty of future human activities and the complex influence factors. In this work, we explore the method of recognizing intentions alluded under human behaviors in the real world, aiming to boost intelligent systems' ability to recognize potential intentions and understand human behaviors. We collect data containing real-world human behaviors before using a hand dispenser and a temperature scanner at the building entrance. These data are processed and labeled into intention categories. A questionnaire is conducted to survey the human ability in inferring the intentions of others. Skeleton data and image features are extracted inspired by the answer to the questionnaire. For skeleton-based intention recognition, we propose a spatial-temporal graph convolutional network that performs graph convolutions on both part-based graphs and adaptive graphs, which achieves the best performance compared with baseline models in the same task. A deep-learning-based method using multimodal features is proposed to automatically infer intentions, which is demonstrated to accurately predict intentions based on past behaviors in the experiment, significantly outperforming humans. Carlos Toshinori Ishi, Bowen Wu 0002, Hiroshi Ishiguro |
IROS | 4 |
| 2022 | Controlling the Impression of Robots via GAN-based Gesture GenerationabstractAs a type of body language, gestures can largely affect the impressions of human-like robots perceived by users. Recent data-driven approaches to the generation of co-speech gestures have successfully promoted the naturalness of produced gestures. These approaches also possess greater generalizability to work under various contexts than rule-based methods. However, most have no direct control over the human impressions of robots. The main obstacle is that creating a dataset that covers various impression labels is not trivial. In this study, based on previous findings in cognitive science on robot impressions, we present a heuristic method to control them without manual labeling, and demonstrate its effectiveness on a virtual agent and partially on a humanoid robot through subjective experiments with 50 participants. Bowen Wu 0002, Carlos Toshinori Ishi, Hiroshi Ishiguro |
IROS | 1 |