Qiyu Wu 0001

dblp:197/6506-1 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0002-6013-8533ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Representation and self-supervised learning · 43% Machine translation · 22% Language models and text generation · 15%
Databases, data mining, and information retrieval
1 paper
Data integration and cleaning · 100%
Computer graphics and multimedia
1 paper
Audio and music processing · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Smart cities and intelligent transportation · 100%

Topics — the 15 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Machine translation › statistical machine translation
word alignment
1.422024
Word Alignment as Preference for Machine Translation · EMNLP 2024
WSPAlign: Word Alignment Pre-training via Large-Scale Weakly Supervised Span Prediction · ACL (1) 2023
Machine learning › Representation and self-supervised learning
contrastive learning
1.122022
Stable Contrastive Learning for Self-Supervised Sentence Embeddings With Pseudo-Siamese Mutual Learning · IEEE ACM Trans. Audio Speech Lang. Process. 2022
PCL: Peer-Contrastive Learning with Diverse Augmentations for Unsupervised Sentence Embeddings · EMNLP 2022
Machine learning › Representation and self-supervised learning › text embedding
sentence embedding
1.122022
Stable Contrastive Learning for Self-Supervised Sentence Embeddings With Pseudo-Siamese Mutual Learning · IEEE ACM Trans. Audio Speech Lang. Process. 2022
PCL: Peer-Contrastive Learning with Diverse Augmentations for Unsupervised Sentence Embeddings · EMNLP 2022
Audio and music processing › music information retrieval
music understanding
0.912025
DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuning · EMNLP 2025
Machine learning › Generative modeling
diffusion model
0.812024
Unified Generation, Reconstruction, and Representation: Generalized Diffusion with Adaptive Latent Encoding-Decoding · ICML 2024
Machine learning › Generative modeling
variational autoencoder
0.812024
Unified Generation, Reconstruction, and Representation: Generalized Diffusion with Adaptive Latent Encoding-Decoding · ICML 2024
Machine learning › Representation and self-supervised learning
pre-training
0.722023
Taking Notes on the Fly Helps Language Pre-Training · ICLR 2021
WSPAlign: Word Alignment Pre-training via Large-Scale Weakly Supervised Span Prediction · ACL (1) 2023
Data integration and cleaning
entity resolution
0.712023
Blocker and Matcher Can Mutually Benefit: A Co-Learning Framework for Low-Resource Entity Resolution · Proc. VLDB Endow. 2023
Data integration and cleaning › entity resolution
low-resource entity resolution
0.712023
Blocker and Matcher Can Mutually Benefit: A Co-Learning Framework for Low-Resource Entity Resolution · Proc. VLDB Endow. 2023
Machine learning › Representation and self-supervised learning › contrastive learning
self-supervised contrastive learning
0.612022
PCL: Peer-Contrastive Learning with Diverse Augmentations for Unsupervised Sentence Embeddings · EMNLP 2022
Machine learning › Representation and self-supervised learning › text embedding › sentence embedding
unsupervised sentence embeddings
0.612022
PCL: Peer-Contrastive Learning with Diverse Augmentations for Unsupervised Sentence Embeddings · EMNLP 2022
Natural language and speech › Language models and text generation › large language model training
language model pretraining
0.512021
Taking Notes on the Fly Helps Language Pre-Training · ICLR 2021
Machine learning › Graph learning › graph neural network › dynamic graph neural network
spatio-temporal graph neural network
0.512021
Community-Aware Multi-Task Transportation Demand Prediction · AAAI 2021
Smart cities and intelligent transportation
demand prediction
0.512021
Community-Aware Multi-Task Transportation Demand Prediction · AAAI 2021
Natural language and speech › Language models and text generation
pre-trained language model
0.212022
Stable Contrastive Learning for Self-Supervised Sentence Embeddings With Pseudo-Siamese Mutual Learning · IEEE ACM Trans. Audio Speech Lang. Process. 2022

Methods — techniques the papers use, named apart from their topics

multimodal fusion · 1.7instruction tuning · 1.7imagebind embeddings · 1.7contrastive learning · 1.1parameterized encoding-decoding · 0.8large language model · 0.8joint training · 0.8direct preference optimization · 0.8GPT-4 · 0.8span prediction · 0.7pseudo-labeling · 0.7deep learning · 0.7co-training · 0.7multi-task learning · 0.5graph neural network · 0.5adaptive task grouping · 0.5
YearPublicationVenuePosition
2025 DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuning
abstract
Recent advancements in music large language models (LLMs) have significantly improved music understanding tasks, which involve the model's ability to analyze and interpret various musical elements.These improvements primarily focused on integrating both music and text inputs.However, the potential of incorporating additional modalities such as images, videos and textual music features to enhance music understanding remains unexplored.To bridge this gap, we propose DeepResonance, a multimodal music understanding LLM fine-tuned via multiway instruction tuning with multi-way aligned music, text, image, and video data.To this end, we construct Music4way-MI2T, Music4way-MV2T, and Music4way-Any2T, three 4-way training and evaluation datasets designed to enable DeepResonance to integrate both visual and textual music feature content.We also introduce multi-sampled ImageBind embeddings and a pre-LLM fusion Transformer to enhance modality fusion prior to input into text LLMs, tailoring for multi-way instruction tuning.Our model achieves state-of-the-art performances across six music understanding tasks, highlighting the benefits of the auxiliary modalities and the structural superiority of DeepResonance.We open-source the codes, models and datasets we constructed: https: //github.com/sony/DeepResonance.
Zhuoyuan Mao, Qiyu Wu 0001, Hiromi Wakaki, Yuki Mitsufuji
EMNLP3
2024 Leveraging Multi-lingual Positive Instances in Contrastive Learning to Improve Sentence Embedding
abstract
Learning multilingual sentence embeddings is a fundamental task in natural language processing.Recent trends in learning both monolingual and multilingual sentence embeddings are mainly based on contrastive learning (CL) among an anchor, one positive, and multiple negative instances.In this work, we argue that leveraging multiple positives should be considered for multilingual sentence embeddings because (1) positives in a diverse set of languages can benefit cross-lingual learning, and (2) transitive similarity across multiple positives can provide reliable structural information for learning.In order to investigate the impact of multiple positives in CL, we propose a novel approach, named MPCL, to effectively utilize multiple positive instances to improve the learning of multilingual sentence embeddings.Experimental results on various backbone models and downstream tasks demonstrate that MPCL leads to better retrieval, semantic similarity, and classification performance compared to conventional CL.We also observe that in unseen languages, sentence embedding models trained on multiple positives show better cross-lingual transfer performance than models trained on a single positive instance.
Kaiyan Zhao, Qiyu Wu 0001, Xin-Qiang Cai, Yoshimasa Tsuruoka
EACL (1)2
2024 Word Alignment as Preference for Machine Translation
abstract
The problem of hallucination and omission, a long-standing problem in machine translation (MT), is more pronounced when a large language model (LLM) is used in MT because an LLM itself is susceptible to these phenomena.In this work, we mitigate the problem in an LLM-based MT model by guiding it to better word alignment.We first study the correlation between word alignment and the phenomena of hallucination and omission in MT.Then we propose to utilize word alignment as preference to optimize the LLM-based MT model.The preference data are constructed by selecting chosen and rejected translations from multiple MT tools.Subsequently, direct preference optimization is used to optimize the LLM-based model towards the preference signal.Given the absence of evaluators specifically designed for hallucination and omission in MT, we further propose selecting hard instances and utilizing GPT-4 to directly evaluate the performance of the models in mitigating these issues.We verify the rationality of these designed evaluation methods by experiments, followed by extensive results demonstrating the effectiveness of word alignment-based preference optimization to mitigate hallucination and omission.On the other hand, although it shows promise in mitigating hallucination and omission, the overall performance of MT in different language directions remains mixed, with slight increases in BLEU and decreases in COMET.
Qiyu Wu 0001, Masaaki Nagata, Zhongtao Miao, Yoshimasa Tsuruoka
EMNLP1
2024 Unified Generation, Reconstruction, and Representation: Generalized Diffusion with Adaptive Latent Encoding-Decoding
abstract
The vast applications of deep generative models are anchored in three core capabilities---*generating* new instances, *reconstructing* inputs, and learning compact *representations*---across various data types, such as discrete text/protein sequences and continuous images. Existing model families, like variational autoencoders (VAEs), generative adversarial networks (GANs), autoregressive models, and (latent) diffusion models, generally excel in specific capabilities and data types but fall short in others. We introduce *Generalized* ***E****ncoding*-***D****ecoding ****D****iffusion ****P****robabilistic ****M****odels* (EDDPMs) which integrate the core capabilities for broad applicability and enhanced performance. EDDPMs generalize the Gaussian noising-denoising in standard diffusion by introducing parameterized encoding-decoding. Crucially, EDDPMs are compatible with the well-established diffusion model objective and training recipes, allowing effective learning of the encoder-decoder parameters *jointly* with diffusion. By choosing appropriate encoder/decoder (e.g., large language models), EDDPMs naturally apply to different data types. Extensive experiments on text, proteins, and images demonstrate the flexibility to handle diverse data and tasks and the strong improvement over various existing models. Code is available at https://github.com/guangyliu/EDDPM .
Guangyi Liu 0005, Yu Wang 0170, Zeyu Feng, Qiyu Wu 0001, Zhen Li 0026, Shuguang Cui, Julian J. McAuley, Eric P. Xing, Zhiting Hu
ICML4
2023 WSPAlign: Word Alignment Pre-training via Large-Scale Weakly Supervised Span Prediction
abstract
Most existing word alignment methods rely on manual alignment datasets or parallel corpora, which limits their usefulness.Here, to mitigate the dependence on manual data, we broaden the source of supervision by relaxing the requirement for correct, fully-aligned, and parallel sentences.Specifically, we make noisy, partially aligned, and non-parallel paragraphs.We then use such a large-scale weakly-supervised dataset for word alignment pre-training via span prediction.Extensive experiments with various settings empirically demonstrate that our approach, which is named WSPAlign, is an effective and scalable way to pre-train word aligners without manual data.When fine-tuned on standard benchmarks, WSPAlign has set a new state of the art by improving upon the best supervised baseline by 3.3~6.1 points in F1 and 1.5~6.1 points in AER .Furthermore, WSPAlign also achieves competitive performance compared with the corresponding baselines in few-shot, zero-shot and cross-lingual tests, which demonstrates that WSPAlign is potentially more practical for low-resource languages than existing methods. 1(1) Data Collection and Annotation (2) Pre-training for word alignment Transformer Encoder
Qiyu Wu 0001, Masaaki Nagata, Yoshimasa Tsuruoka
ACL (1)1
2023 Blocker and Matcher Can Mutually Benefit: A Co-Learning Framework for Low-Resource Entity Resolution
abstract
Entity resolution (ER) approaches typically consist of a blocker and a matcher. They share the same goal and cooperate in different roles: the blocker first quickly removes obvious non-matches, and the matcher subsequently determines whether the remaining pairs refer to the same real-world entity. Despite the state-of-the-art performance achieved by deep learning methods in ER, these techniques often rely on a large amount of labeled data for training, which can be challenging or costly to obtain. Thus, there is a need to develop effective ER systems under low-resource settings. In this work, we propose an end-to-end iterative Co-learning framework for ER, aimed at jointly training the blocker and the matcher by leveraging their cooperative relationship. In particular, we let the blocker and the matcher share their learned knowledge with each other via iteratively updated pseudo labels, which broaden the supervision signals. To mitigate the impact of noise in pseudo labels, we develop optimization techniques from three aspects: label generation, label selection and model training. Through extensive experiments on benchmark datasets, we demonstrate that our proposed framework outperforms baselines by an average of 9.13--51.55%. Furthermore, our analysis confirms that our framework achieves mutual benefits between the blocker and the matcher.
Shiwen Wu, Qiyu Wu 0001, Honghua Dong, Wen Hua, Xiaofang Zhou 0001
Proc. VLDB Endow.2
2022 PCL: Peer-Contrastive Learning with Diverse Augmentations for Unsupervised Sentence Embeddings
abstract
Learning sentence embeddings in an unsupervised manner is fundamental in natural language processing.Recent common practice is to couple pre-trained language models with unsupervised contrastive learning, whose success relies on augmenting a sentence with a semantically-close positive instance to construct contrastive pairs.Nonetheless, existing approaches usually depend on a monoaugmenting strategy, which causes learning shortcuts towards the augmenting biases and thus corrupts the quality of sentence embeddings.A straightforward solution is resorting to more diverse positives from a multiaugmenting strategy, while an open question remains about how to unsupervisedly learn from the diverse positives but with uneven augmenting qualities in the text field.As one answer, we propose a novel Peer-Contrastive Learning (PCL) with diverse augmentations.PCL constructs diverse contrastive positives and negatives at the group level for unsupervised sentence embeddings.PCL performs peer-positive contrast as well as peer-network cooperation, which offers an inherent anti-bias ability and an effective way to learn from diverse augmentations.Experiments on STS benchmarks verify the effectiveness of PCL against its competitors in unsupervised sentence embeddings. 1
Qiyu Wu 0001, Chongyang Tao, Tao Shen 0001, Can Xu 0002, Xiubo Geng, Daxin Jiang
EMNLP1
2022 Stable Contrastive Learning for Self-Supervised Sentence Embeddings With Pseudo-Siamese Mutual Learning
abstract
Learning semantic sentence embeddings is beneficial to a variety of natural language processing tasks. Recently, methods using the contrastive learning framework to fine-tune pre-trained language models have been proposed and have achieved significant performance on sentence embeddings. However, sentence embeddings are easy to “overfit” to the contrastive learning goal. With the training of contrastive learning, the gap between contrastive learning and test tasks leads to unstable even declining performance on test tasks. For this reason, existing methods rely on the labeled development set to frequently evaluate the performance on test tasks and get the best checkpoints. In such a way, models are limited when the labeled data is unavailable or extremely scarce. To address this problem, we proposePseudo-Siamese networkMutualLearning (PSML) for self-supervised sentence embeddings to reduce the gap between contrastive learning and test tasks. Consisting of the main encoder and the auxiliary encoder, PSML utilizes mutual learning as the basic framework. Between the two encoders, two mutual learning losses are constructed to share learning signals. The proposed model framework and losses of PSML help the model be optimized more stably and generalize better to test tasks, such as semantic textual similarity. Extensive experiments on seven public semantic textual similarity datasets show that PSML performs better than previous unsupervised contrastive methods for sentence embeddings. Besides, PSML also gives a stable performance curve on test tasks with training and is able to get the comparative performance without frequent evaluation on the labeled development set.
Qiyu Wu 0001, Wei Chen 0056, Tengjiao Wang 0003
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 Community-Aware Multi-Task Transportation Demand Prediction
abstract
Transportation demand prediction is of great importance to urban governance and has become an essential function in many online applications. While many efforts have been made for regional transportation demand prediction, predicting the diversified transportation demand for different communities (e.g., the aged, the juveniles) remains an unexplored problem. However, this task is challenging because of the joint influence of spatio-temporal correlation among regions and implicit correlation among different communities. To this end, in this paper, we propose the Multi-task Spatio-Temporal Network with Mutually-supervised Adaptive task grouping (Ada-MSTNet) for community-aware transportation demand prediction. Specifically, we first construct a sequence of multi-view graphs from both spatial and community perspectives, and devise a spatio-temporal neural network to simultaneously capture the sophisticated correlations between regions and communities, respectively. Then, we propose an adaptively clustered multi-task learning module, where the prediction of each region-community specific transportation demand is regarded as distinct task. Moreover, a mutually supervised adaptive task grouping strategy is introduced to softly cluster each task into different task groups, by leveraging the supervision signal from one another graph view. In such a way, Ada-MSTNet is not only able to share common knowledge among highly related communities and regions, but also shield the noise from unrelated tasks in an end-to-end fashion. Finally, extensive experiments on two real-world datasets demonstrate the effectiveness of our approach compared with seven baselines.
Hao Liu 0026, Qiyu Wu 0001, Fuzhen Zhuang, Xinjiang Lu, Dejing Dou, Hui Xiong 0001
AAAI2
2021 Taking Notes on the Fly Helps Language Pre-Training
Qiyu Wu 0001, Chen Xing, Yatao Li, Guolin Ke, Di He 0001, Tie-Yan Liu
ICLR1
2020 Stance Detection with Stance-Wise Convolution Network
Dechuan Yang, Qiyu Wu 0001, Wei Chen 0021, Tengjiao Wang 0003, Yingbao Cui
NLPCC (1)2