Hiromi Wakaki

dblp:07/3863 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Generative modeling · 29% Knowledge representation and reasoning · 21% Trustworthy machine learning · 16%
Computer graphics and multimedia
1 paper
Audio and music processing · 100%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
2.332025
VinaBench: Benchmark for Faithful and Consistent Visual Narratives · CVPR 2025
DiffuCOMET: Contextual Commonsense Knowledge Diffusion · ACL (1) 2024
PeaCoK: Persona Commonsense Knowledge for Consistent and Engaging Narratives · ACL (1) 2023
Machine learning › Generative modeling
diffusion model
1.122025
Distillation of Discrete Diffusion through Dimensional Correlations · ICML 2025
DiffuCOMET: Contextual Commonsense Knowledge Diffusion · ACL (1) 2024
Machine learning › Trustworthy machine learning › fairness › social bias
cultural bias
0.912025
CARE: Multilingual Human Preference Learning for Cultural Awareness · EMNLP 2025
Machine learning › Generative modeling › diffusion model
diffusion distillation
0.912025
Distillation of Discrete Diffusion through Dimensional Correlations · ICML 2025
Machine learning › Generative modeling › diffusion model
discrete diffusion model
0.912025
Distillation of Discrete Diffusion through Dimensional Correlations · ICML 2025
Machine learning › Trustworthy machine learning
fairness
0.912025
CARE: Multilingual Human Preference Learning for Cultural Awareness · EMNLP 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
Distillation of Discrete Diffusion through Dimensional Correlations · ICML 2025
Machine learning › Reinforcement learning
preference learning
0.912025
CARE: Multilingual Human Preference Learning for Cultural Awareness · EMNLP 2025
Audio and music processing › music information retrieval
music understanding
0.912025
DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuning · EMNLP 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.312025
VinaBench: Benchmark for Faithful and Consistent Visual Narratives · CVPR 2025
Natural language and speech › Language models and text generation › text generation
story generation
0.212023
PeaCoK: Persona Commonsense Knowledge for Consistent and Engaging Narratives · ACL (1) 2023

Methods — techniques the papers use, named apart from their topics

multimodal fusion · 1.7instruction tuning · 1.7imagebind embeddings · 1.7reinforcement learning from human feedback · 0.9mixture model · 0.9human preference learning · 0.9evaluation metrics · 0.9distillation · 0.9benchmark · 0.9diffusion model · 0.8
YearPublicationVenuePosition
2025 VinaBench: Benchmark for Faithful and Consistent Visual Narratives
abstract
Visual narrative generation transforms textual narratives into sequences of images illustrating the content of the text. However, generating visual narratives that are faithful to the input text and self-consistent across generated images remains an open challenge, due to the lack of knowledge constraints used for planning the stories. In this work, we propose a new benchmark, VinaBench, to address this challenge. Our benchmark annotates the underlying commonsense and discourse constraints in visual narrative samples, offering systematic scaffolds for learning the implicit strategies of visual storytelling. Based on the incorporated narrative constraints, we further propose novel metrics to closely evaluate the consistency of generated narrative images and the alignment of generations with the input textual narrative. Our results across three generative vision models demonstrate that learning with VinaBench’s knowledge constraints effectively improves the faithfulness and cohesion of generated visual narratives.1
Silin Gao, Sheryl Mathew, Li Mi, Sepideh Mamooler, Hiromi Wakaki, Yuki Mitsufuji, Syrielle Montariol, Antoine Bosselut
CVPR6
2025 CARE: Multilingual Human Preference Learning for Cultural Awareness
abstract
Language Models (LMs) are typically tuned with human preferences to produce helpful responses, but the impact of preference tuning on the ability to handle culturally diverse queries remains understudied.In this paper, we systematically analyze how native human cultural preferences can be incorporated into the preference learning process to train more culturally aware LMs.We introduce CARE, a multilingual resource containing 3,490 culturally specific questions and 31.7kresponses with human judgments.We demonstrate how a modest amount of high-quality native preferences improves cultural awareness across various LMs, outperforming larger generic preference data.Our analyses reveal that models with stronger initial cultural performance benefit more from alignment, leading to gaps among models developed in different regions with varying access to culturally relevant data.CARE is publicly available at https://github.com/Guochry/ CARE.
Geyang Guo, Tarek Naous, Hiromi Wakaki, Yukiko Nishimura, Yuki Mitsufuji, Alan Ritter, Wei Xu 0004
EMNLP3
2025 DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuning
abstract
Recent advancements in music large language models (LLMs) have significantly improved music understanding tasks, which involve the model's ability to analyze and interpret various musical elements.These improvements primarily focused on integrating both music and text inputs.However, the potential of incorporating additional modalities such as images, videos and textual music features to enhance music understanding remains unexplored.To bridge this gap, we propose DeepResonance, a multimodal music understanding LLM fine-tuned via multiway instruction tuning with multi-way aligned music, text, image, and video data.To this end, we construct Music4way-MI2T, Music4way-MV2T, and Music4way-Any2T, three 4-way training and evaluation datasets designed to enable DeepResonance to integrate both visual and textual music feature content.We also introduce multi-sampled ImageBind embeddings and a pre-LLM fusion Transformer to enhance modality fusion prior to input into text LLMs, tailoring for multi-way instruction tuning.Our model achieves state-of-the-art performances across six music understanding tasks, highlighting the benefits of the auxiliary modalities and the structural superiority of DeepResonance.We open-source the codes, models and datasets we constructed: https: //github.com/sony/DeepResonance.
Zhuoyuan Mao, Qiyu Wu 0001, Hiromi Wakaki, Yuki Mitsufuji
EMNLP4
2025 Distillation of Discrete Diffusion through Dimensional Correlations
abstract
Diffusion models have demonstrated exceptional performances in various fields of generative modeling, but suffer from slow sampling speed due to their iterative nature. While this issue is being addressed in continuous domains, discrete diffusion models face unique challenges, particularly in capturing dependencies between elements (e.g., pixel relationships in image, sequential dependencies in language) mainly due to the computational cost of processing high-dimensional joint distributions. In this paper, (i) we propose "mixture" models for discrete diffusion that are capable of treating dimensional correlations while remaining scalable, and (ii) we provide a set of loss functions for distilling the iterations of existing models. Two primary theoretical insights underpin our approach: First, conventional models with element-wise independence can well approximate the data distribution, but essentially require *many sampling steps*. Second, our loss functions enable the mixture models to distill such many-step conventional models into just a few steps by learning the dimensional correlations. Our experimental results show the effectiveness of the proposed method in distilling pretrained discrete diffusion models across image and language domains. The code used in the paper is available at https://github.com/sony/di4c.
Satoshi Hayakawa, Yuhta Takida, Masaaki Imaizumi, Hiromi Wakaki, Yuki Mitsufuji
ICML4
2024 DiffuCOMET: Contextual Commonsense Knowledge Diffusion
abstract
Silin Gao, Mete Ismayilzada, Mengjie Zhao, Hiromi Wakaki, Yuki Mitsufuji, Antoine Bosselut. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Silin Gao, Mete Ismayilzada, Hiromi Wakaki, Yuki Mitsufuji, Antoine Bosselut
ACL (1)4
2023 PeaCoK: Persona Commonsense Knowledge for Consistent and Engaging Narratives
abstract
Silin Gao, Beatriz Borges, Soyoung Oh, Deniz Bayazit, Saya Kanno, Hiromi Wakaki, Yuki Mitsufuji, Antoine Bosselut. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Silin Gao, Beatriz Borges, Soyoung Oh, Deniz Bayazit, Saya Kanno, Hiromi Wakaki, Yuki Mitsufuji, Antoine Bosselut
ACL (1)6
2021 Fundamental Exploration of Evaluation Metrics for Persona Characteristics of Text Utterances
abstract
To maintain utterance quality of a personaaware dialog system, inappropriate utterances for the persona should be thoroughly filtered.When evaluating the appropriateness of a large number of arbitrary utterances to be registered in the utterance database of a retrieval-based dialog system, evaluation metrics that require a reference (or a "correct" utterance) for each evaluation target cannot be used.In addition, practical utterance filtering requires the ability to select utterances based on the intensity of persona characteristics.Therefore, we are developing metrics that can be used to capture the intensity of persona characteristics and can be computed without references tailored to the evaluation targets.To this end, we explore existing metrics and propose two new metrics: persona speaker probability and persona term salience.Experimental results show that our proposed metrics show weak to moderate correlations between scores of persona characteristics based on human judgments and outperform other metrics overall in filtering inappropriate utterances for particular personas.
Chiaki Miyazaki, Saya Kanno, Makoto Yoda, Junya Ono, Hiromi Wakaki
SIGDIAL5
2006 A New Measure for Query Disambiguation Using Term Co-occurrences
Hiromi Wakaki, Tomonari Masada, Atsuhiro Takasu, Jun Adachi
IDEAL1
2003 AVICE: Evolving Avatar's Movernent
Hiromi Wakaki, Hitoshi Iba
GECCO1
2002 3D-CG Avatar Motion Design by means of Interactive Evolutionary Computation
Hitoshi Iba, N. Tokui, Hiromi Wakaki
HIS3