VLDB 2026 Research / reviewers in the wild / expert
Zhuoyuan Mao
dblp:256/9496
· DBLP profile ↗
8ranked-venue papers
6as first author
7since 2021 · last 2025
0000-0001-5273-2738ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 6 first-author · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Language models and text generation · 42% Representation and self-supervised learning · 29% Information extraction and text analysis · 29% | |
| Computer graphics and multimedia
1 paper |
Audio and music processing · 100% |
Topics — the 9 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis
relation extraction |
1.2 | 2 | 2023 | GPT-RE: In-context Learning for Relation Extraction using Large Language Models · EMNLP 2023 Rescue Implicit and Long-tail Cases: Nearest Neighbor Relation Extraction · EMNLP 2022 |
Audio and music processing › music information retrieval
music understanding |
0.9 | 1 | 2025 | DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuning · EMNLP 2025 |
Machine learning › Representation and self-supervised learning › text embedding › text representation learning
cross-lingual representation learning |
0.8 | 1 | 2024 | EMS: Efficient and Effective Massively Multilingual Sentence Embedding Learning · IEEE ACM Trans. Audio Speech Lang. Process. 2024 |
Machine learning › Representation and self-supervised learning › word representation
multilingual word embedding |
0.8 | 1 | 2024 | EMS: Efficient and Effective Massively Multilingual Sentence Embedding Learning · IEEE ACM Trans. Audio Speech Lang. Process. 2024 |
Natural language and speech › Language models and text generation › in-context learning
demonstration retrieval |
0.7 | 1 | 2023 | GPT-RE: In-context Learning for Relation Extraction using Large Language Models · EMNLP 2023 |
Natural language and speech › Language models and text generation
in-context learning |
0.7 | 1 | 2023 | GPT-RE: In-context Learning for Relation Extraction using Large Language Models · EMNLP 2023 |
Natural language and speech › Language models and text generation
prompting |
0.7 | 1 | 2023 | GPT-RE: In-context Learning for Relation Extraction using Large Language Models · EMNLP 2023 |
Natural language and speech › Information extraction and text analysis › relation extraction › low-resource relation extraction
long-tail relation extraction |
0.6 | 1 | 2022 | Rescue Implicit and Long-tail Cases: Nearest Neighbor Relation Extraction · EMNLP 2022 |
Natural language and speech › Information extraction and text analysis › relation extraction
distant supervision |
0.2 | 1 | 2022 | Rescue Implicit and Long-tail Cases: Nearest Neighbor Relation Extraction · EMNLP 2022 |
Methods — techniques the papers use, named apart from their topics
multimodal fusion · 1.7instruction tuning · 1.7imagebind embeddings · 1.7sentence-level contrastive learning · 0.8cross-lingual token-level reconstruction · 0.8reasoning logic · 0.7in-context learning · 0.7demonstration retrieval · 0.7nearest neighbor search · 0.6k-nearest neighbors · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction TuningabstractRecent advancements in music large language models (LLMs) have significantly improved music understanding tasks, which involve the model's ability to analyze and interpret various musical elements.These improvements primarily focused on integrating both music and text inputs.However, the potential of incorporating additional modalities such as images, videos and textual music features to enhance music understanding remains unexplored.To bridge this gap, we propose DeepResonance, a multimodal music understanding LLM fine-tuned via multiway instruction tuning with multi-way aligned music, text, image, and video data.To this end, we construct Music4way-MI2T, Music4way-MV2T, and Music4way-Any2T, three 4-way training and evaluation datasets designed to enable DeepResonance to integrate both visual and textual music feature content.We also introduce multi-sampled ImageBind embeddings and a pre-LLM fusion Transformer to enhance modality fusion prior to input into text LLMs, tailoring for multi-way instruction tuning.Our model achieves state-of-the-art performances across six music understanding tasks, highlighting the benefits of the auxiliary modalities and the structural superiority of DeepResonance.We open-source the codes, models and datasets we constructed: https: //github.com/sony/DeepResonance. Zhuoyuan Mao, Qiyu Wu 0001, Hiromi Wakaki, Yuki Mitsufuji |
EMNLP | 1 |
| 2024 | EMS: Efficient and Effective Massively Multilingual Sentence Embedding LearningabstractMassively multilingual sentence representation models, e.g., LASER, SBERT-distill, and LaBSE, help significantly improve cross-lingual downstream tasks. However, the use of a large amount of data or inefficient model architectures results in heavy computation to train a new model according to our preferred languages and domains. To resolve this issue, we introduce efficient and effective massively multilingual sentence embedding (EMS), using cross-lingual token-level reconstruction (XTR) and sentence-level contrastive learning as training objectives. Compared with related studies, the proposed model can be efficiently trained using significantly fewer parallel sentences and GPU computation resources. Empirical results showed that the proposed model significantly yields better or comparable results with regard to cross-lingual sentence retrieval, zero-shot cross-lingual genre classification, and sentiment classification. Ablative analyses demonstrated the efficiency and effectiveness of each component of the proposed model. We release the codes for model training and the EMS pre-trained sentence embedding model, which supports 62 languages (https://github.com/Mao-KU/EMS). Zhuoyuan Mao, Chenhui Chu, Sadao Kurohashi |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | LEALLA: Learning Lightweight Language-agnostic Sentence Embeddings with Knowledge DistillationabstractLarge-scale language-agnostic sentence embedding models such as LaBSE (Feng et al., 2022) obtain state-of-the-art performance for parallel sentence alignment.However, these large-scale models can suffer from inference speed and computation overhead.This study systematically explores learning language-agnostic sentence embeddings with lightweight models.We demonstrate that a thin-deep encoder can construct robust low-dimensional sentence embeddings for 109 languages.With our proposed distillation methods, we achieve further improvements by incorporating knowledge from a teacher model.Empirical results on Tatoeba, United Nations, and BUCC show the effectiveness of our lightweight models.We release our lightweight language-agnostic sentence embedding models LEALLA on Tensor-Flow Hub. 1 * Currently at Kurohashi-Chu-Murawaki Lab., Zhuoyuan Mao, Tetsuji Nakagawa |
EACL | 1 |
| 2023 | GPT-RE: In-context Learning for Relation Extraction using Large Language ModelsabstractIn spite of the potential for ground-breaking achievements offered by large language models (LLMs) (e.g., GPT-3) via in-context learning (ICL), they still lag significantly behind fullysupervised baselines (e.g., fine-tuned BERT) in relation extraction (RE).This is due to the two major shortcomings of ICL for RE: (1) low relevance regarding entity and relation in existing sentence-level demonstration retrieval approaches for ICL; and (2) the lack of explaining input-label mappings of demonstrations leading to poor ICL effectiveness.In this paper, we propose GPT-RE to successfully address the aforementioned issues by (1) incorporating task-aware representations in demonstration retrieval; and (2) enriching the demonstrations with gold label-induced reasoning logic.We evaluate GPT-RE on four widely-used RE datasets and observe that GPT-RE achieves improvements over not only existing GPT-3 baselines, but also fully-supervised baselines as in Figure 1.Specifically, GPT-RE achieves SOTA performances on the Semeval and SciERC datasets, and competitive performances on the TACRED and ACE05 datasets.Additionally, a critical issue of LLMs revealed by previous work, the strong inclination to wrongly classify NULL examples into other predefined labels, is substantially alleviated by our method.We show an empirical analysis.1 Fei Cheng 0002, Zhuoyuan Mao, Qianying Liu, Haiyue Song, Jiwei Li 0001, Sadao Kurohashi |
EMNLP | 3 |
| 2022 | Rescue Implicit and Long-tail Cases: Nearest Neighbor Relation ExtractionabstractRelation extraction (RE) has achieved remarkable progress with the help of pre-trained language models.However, existing RE models are usually incapable of handling two situations: implicit expressions and long-tail relation types, caused by language complexity and data sparsity.In this paper, we introduce a simple enhancement of RE using k nearest neighbors (kNN-RE).kNN-RE allows the model to consult training relations at test time through a nearest-neighbor search and provides a simple yet effective means to tackle the two issues above.Additionally, we observe that kNN-RE serves as an effective way to leverage distant supervision (DS) data for RE.Experimental results show that the proposed kNN-RE achieves state-of-the-art performances on a variety of supervised RE datasets, i.e., ACE05, SciERC, and Wiki80, along with outperforming the best model to date on the i2b2 and Wiki80 datasets in the setting of allowing using DS.Our code and models are available at: https://github.com/YukinoWan/kNN-RE. Qianying Liu, Zhuoyuan Mao, Fei Cheng 0002, Sadao Kurohashi, Jiwei Li 0001 |
EMNLP | 3 |
| 2022 | Linguistically Driven Multi-Task Pre-Training for Low-Resource Neural Machine TranslationabstractIn the present study, we propose novel sequence-to-sequence pre-training objectives for low-resource machine translation (NMT): Japanese-specific sequence to sequence (JASS) for language pairs involving Japanese as the source or target language, and English-specific sequence to sequence (ENSS) for language pairs involving English. JASS focuses on masking and reordering Japanese linguistic units known as bunsetsu, whereas ENSS is proposed based on phrase structure masking and reordering tasks. Experiments on ASPEC Japanese–English & Japanese–Chinese, Wikipedia Japanese–Chinese, News English–Korean corpora demonstrate that JASS and ENSS outperform MASS and other existing language-agnostic pre-training methods by up to +2.9 BLEU points for the Japanese–English tasks, up to +7.0 BLEU points for the Japanese–Chinese tasks and up to +1.3 BLEU points for English–Korean tasks. Empirical analysis, which focuses on the relationship between individual parts in JASS and ENSS, reveals the complementary nature of the subtasks of JASS and ENSS. Adequacy evaluation using LASER, human evaluation, and case studies reveals that our proposed methods significantly outperform pre-training methods without injected linguistic knowledge and they have a larger positive impact on the adequacy as compared to the fluency. Zhuoyuan Mao, Chenhui Chu, Sadao Kurohashi |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2021 | Lightweight Cross-Lingual Sentence Representation LearningabstractZhuoyuan Mao, Prakhar Gupta, Chenhui Chu, Martin Jaggi, Sadao Kurohashi. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zhuoyuan Mao, Prakhar Gupta, Chenhui Chu, Martin Jaggi, Sadao Kurohashi |
ACL/IJCNLP (1) | 1 |
| 2020 | JASS: Japanese-specific Sequence to Sequence Pre-training for Neural Machine TranslationabstractNeural machine translation (NMT) needs large parallel corpora for state-of-the-art translation quality. Low-resource NMT is typically addressed by transfer learning which leverages large monolingual or parallel corpora for pre-training. Monolingual pre-training approaches such as MASS (MAsked Sequence to Sequence) are extremely effective in boosting NMT quality for languages with small parallel corpora. However, they do not account for linguistic information obtained using syntactic analyzers which is known to be invaluable for several Natural Language Processing (NLP) tasks. To this end, we propose JASS, Japanese-specific Sequence to Sequence, as a novel pre-training alternative to MASS for NMT involving Japanese as the source or target language. JASS is joint BMASS (Bunsetsu MASS) and BRSS (Bunsetsu Reordering Sequence to Sequence) pre-training which focuses on Japanese linguistic units called bunsetsus. In our experiments on ASPEC Japanese–English and News Commentary Japanese–Russian translation we show that JASS can give results that are competitive with if not better than those given by MASS. Furthermore, we show for the first time that joint MASS and JASS pre-training gives results that significantly surpass the individual methods indicating their complementary nature. We will release our code, pre-trained models and bunsetsu annotated data as resources for researchers to use in their own NLP tasks. Zhuoyuan Mao, Fabien Cromières, Raj Dabre, Haiyue Song, Sadao Kurohashi |
LREC | 1 |