VLDB 2026 Research / reviewers in the wild / expert
Yaobo Liang
dblp:245/8600
· DBLP profile ↗
19ranked-venue papers
2as first author
13since 2021 · last 2025
0000-0002-6595-5145ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | UniGraspTransformer: Simplified Policy Distillation for Scalable Dexterous Robotic GraspingabstractWe introduce UniGraspTransformer, a universal Transformer-based network for dexterous robotic grasping that simplifies training while enhancing scalability and performance. Unlike prior methods such as UniDexGrasp++, which require complex, multi-step training pipelines, UniGraspTransformer follows a streamlined process: first, dedicated policy networks are trained for individual objects using reinforcement learning to generate successful grasp trajectories; then, these trajectories are distilled into a single, universal network. Our approach enables UniGraspTransformer to scale effectively, incorporating up to 12 self-attention blocks for handling thousands of objects with diverse poses. Additionally, it generalizes well to both idealized and real-world inputs, evaluated in state-based and vision-based settings. Notably, UniGraspTransformer generates a broader range of grasping poses for objects in various shapes and orientations, resulting in more diverse grasp strategies. Experimental results demonstrate significant improvements over state-of-the-art, UniDexGrasp++, across various object categories, achieving success rate gains of 3.5%, 7.7%, and 10.1% on seen objects, unseen objects within seen categories, and completely unseen objects, respectively, in the vision-based setting. Project page: https://dexhand.github.io/UniGraspTransformer/. Fangyun Wei, Xiaohan Yi, Yaobo Liang, Chang Xu 0002, Yan Lu 0001, Jiaolong Yang, Baining Guo |
CVPR | 8 |
| 2025 | Direct Preference Optimization for LLM-Enhanced Recommendation SystemsabstractLarge Language Models (LLMs) have exhibited remarkable performance across a wide range of domains, motivating research into their potential for recommendation systems. Early efforts have leveraged LLMs’ rich knowledge and strong generalization capabilities via in-context learning, where recommendation tasks are framed as prompts. However, LLM performance in recommendation scenarios remains limited due to the mismatch between their pretraining objectives and recommendation tasks, as well as the lack of recommendation-specific data during pretraining. To address these challenges, we propose DPO4Rec, a novel framework that integrates Direct Preference Optimization (DPO) into LLM-enhanced recommendation systems. First, we prompt the LLM to infer user preferences from historical interactions, which are then used to augment traditional ID-based sequential recommendation models. Next, we train a reward model based on knowledge-augmented recommendation architectures to assess the quality of LLM-generated reasoning. Using this, we select the highest- and lowest-ranked responses from N samples to construct a dataset for LLM fine-tuning. Finally, we apply a structure alignment strategy via DPO to align the LLM’s outputs with desirable recommendation behavior. Extensive experiments show that DPO4Rec significantly improves re-ranking performance over strong baselines, demonstrating enhanced instruction-following capabilities of LLMs in recommendation tasks. Yaobo Liang, Yaming Yang 0001, Shilin Xu 0001, Tianmeng Yang, Yunhai Tong |
ICME | 2 |
| 2025 | VideoVLA: Video Generators Can Be Generalizable Robot ManipulatorsabstractGeneralization in robot manipulation is essential for deploying robots in open-world environments and advancing toward artificial general intelligence. While recent Vision-Language-Action (VLA) models leverage large pre-trained understanding models for perception and instruction following, their ability to generalize to novel tasks, objects, and settings remains limited. In this work, we present VideoVLA, a simple approach that explores the potential of transforming large video generation models into robotic VLA manipulators. Given a language instruction and an image, VideoVLA predicts an action sequence as well as the future visual outcomes. Built on a multi-modal Diffusion Transformer, VideoVLA jointly models video, language, and action modalities, using pre-trained video generative models for joint visual and action forecasting. Our experiments show that high-quality imagined futures correlate with reliable action predictions and task success, highlighting the importance of visual imagination in manipulation. VideoVLA demonstrates strong generalization, including imitating other embodiments' skills and handling novel objects. This dual-prediction strategy—forecasting both actions and their visual consequences—explores a paradigm shift in robot learning and unlocks generalization capabilities in manipulation systems. Yichao Shen 0001, Fangyun Wei, Zhiying Du, Yaobo Liang, Yan Lu 0001, Jiaolong Yang, Nanning Zheng 0001, Baining Guo |
NeurIPS | 4 |
| 2024 | Machine-Created Universal Language for Cross-Lingual TransferabstractThere are two primary approaches to addressing cross-lingual transfer: multilingual pre-training, which implicitly aligns the hidden representations of various languages, and translate-test, which explicitly translates different languages into an intermediate language, such as English. Translate-test offers better interpretability compared to multilingual pre-training. However, it has lower performance than multilingual pre-training and struggles with word-level tasks due to translation altering word order. As a result, we propose a new Machine-created Universal Language (MUL) as an alternative intermediate language. MUL comprises a set of discrete symbols forming a universal vocabulary and a natural language to MUL translator for converting multiple natural languages to MUL. MUL unifies shared concepts from various languages into a single universal word, enhancing cross-language transfer. Additionally, MUL retains language-specific words and word order, allowing the model to be easily applied to word-level tasks. Our experiments demonstrate that translating into MUL yields improved performance compared to multilingual pre-training, and our analysis indicates that MUL possesses strong interpretability. The code is at: https://github.com/microsoft/Unicoder/tree/master/MCUL. Yaobo Liang, Quanzhi Zhu, Junhe Zhao |
AAAI | 1 |
| 2023 | Analyzing and Reducing the Performance Gap in Cross-Lingual Transfer with Fine-tuning Slow and FastabstractExisting research has shown that a multilingual pre-trained language model fine-tuned with one (source) language also performs well on downstream tasks for non-source languages, even though no fine-tuning is done on these languages.However, there is a clear gap between the performance of the source language and that of the non-source languages.This paper analyzes the fine-tuning process, discovers when the performance gap changes and identifies which network weights affect the overall performance most.Additionally, the paper seeks to answer to what extent the gap can be reduced by reducing forgetting.Based on the analysis results, a method named Fine-tuning slow and fast with four training policies is proposed to address these issues.Experimental results show the proposed method outperforms baselines by a clear margin. Yiduo Guo, Yaobo Liang, Dongyan Zhao 0001, Bing Liu 0001, Nan Duan 0001 |
ACL (1) | 2 |
| 2023 | Modeling Sequential Sentence Relation to Improve Cross-lingual Dense Retrieval
Shunyu Zhang, Yaobo Liang, Ming Gong 0001, Daxin Jiang, Nan Duan 0001 |
ICLR | 2 |
| 2022 | XLM-K: Improving Cross-Lingual Language Model Pre-training with Multilingual KnowledgeabstractCross-lingual pre-training has achieved great successes using monolingual and bilingual plain text corpora. However, most pre-trained models neglect multilingual knowledge, which is language agnostic but comprises abundant cross-lingual structure alignment. In this paper, we propose XLM-K, a cross-lingual language model incorporating multilingual knowledge in pre-training. XLM-K augments existing multilingual pre-training with two knowledge tasks, namely Masked Entity Prediction Task and Object Entailment Task. We evaluate XLM-K on MLQA, NER and XNLI. Experimental results clearly demonstrate significant improvements over existing multilingual language models. The results on MLQA and NER exhibit the superiority of XLM-K in knowledge related tasks. The success in XNLI shows a better cross-lingual transferability obtained in XLM-K. What is more, we provide a detailed probing analysis to confirm the desired knowledge captured in our pre-training regimen. The code is available at https://github.com/microsoft/Unicoder/tree/master/pretraining/xlmk. Xiaoze Jiang, Yaobo Liang, Weizhu Chen |
AAAI | 2 |
| 2022 | Cross-Lingual Ability of Multilingual Masked Language Models: A Study of Language StructureabstractMultilingual pre-trained language models, such as mBERT and XLM-R, have shown impressive cross-lingual ability. Surprisingly, both of them use multilingual masked language model (MLM) without any cross-lingual supervision or aligned data. Despite the encouraging results, we still lack a clear understanding of why cross-lingual ability could emerge from multilingual MLM. In our work, we argue that cross-language ability comes from the commonality between languages. Specifically, we study three language properties: constituent order, composition and word co-occurrence. First, we create an artificial language by modifying property in source language. Then we study the contribution of modified property through the change of cross-language transfer results on target language. We conduct experiments on six languages and two cross-lingual NLP tasks (textual entailment, sentence retrieval). Our main conclusion is that the contribution of constituent order and word co-occurrence is limited, while the composition is more crucial to the success of cross-linguistic transfer. Yaobo Liang |
ACL (1) | 2 |
| 2022 | Multi-View Document Representation Learning for Open-Domain Dense RetrievalabstractDense retrieval has achieved impressive advances in first-stage retrieval from a largescale document collection, which is built on bi-encoder architecture to produce single vector representation of query and document.However, a document can usually answer multiple potential queries from different views.So the single vector representation of a document is hard to match with multi-view queries, and faces a semantic mismatch problem.This paper proposes a multi-view document representation learning framework, aiming to produce multiview embeddings to represent documents and enforce them to align with different queries.First, we propose a simple yet effective method of generating multiple embeddings through viewers.Second, to prevent multi-view embeddings from collapsing to the same one, we further propose a global-local loss with annealed temperature to encourage the multiple viewers to better align with different potential queries.Experiments show our method outperforms recent works and achieves state-of-the-art results. * Work done during internship at Microsoft Research Asia.Q1: Where can people using iPods on planes view the device's interface?A1: Individual seat-back displays.Q2: What are two airlines that considered implementing iPod connections but did not join the 2007 agreement?A2: KLM and Air France. Shunyu Zhang, Yaobo Liang, Ming Gong 0001, Daxin Jiang, Nan Duan 0001 |
ACL (1) | 2 |
| 2022 | Unsupervised Context Aware Sentence Representation Pretraining for Multi-lingual Dense RetrievalabstractRecent research demonstrates the effectiveness of using pretrained language models (PLM) to improve dense retrieval and multilingual dense retrieval. In this work, we present a simple but effective monolingual pretraining task called contrastive context prediction (CCP) to learn sentence representation by modeling sentence level contextual relation. By pushing the embedding of sentences in a local context closer and pushing random negative samples away, different languages could form isomorphic structure, then sentence pairs in two different languages will be automatically aligned. Our experiments show that model collapse and information leakage are very easy to happen during contrastive training of language model, but language-specific memory bank and asymmetric batch normalization operation play an essential role in preventing collapsing and information leakage, respectively. Besides, a post-processing for sentence embedding is also very effective to achieve better retrieval performance. On the multilingual sentence retrieval task Tatoeba, our model achieves new SOTA results among methods without using bilingual data. Our model also shows larger gain on Tatoeba when transferring between non-English pairs. On two multi-lingual query-passage retrieval tasks, XOR Retrieve and Mr.TYDI, our model even achieves two SOTA results in both zero-shot and supervised setting among all pretraining models using bilingual data. Ning Wu 0013, Yaobo Liang, Houxing Ren, Linjun Shou, Nan Duan 0001, Ming Gong 0001, Daxin Jiang |
IJCAI | 2 |
| 2022 | Less-forgetting Multi-lingual Fine-tuningabstractMulti-lingual fine-tuning (MLF), which fine-tunes a multi-lingual language model (MLLM) with multiple source languages, aims to gain good zero-shot performance on target languages. In MLF, the fine-tuned model tends to fit the source languages while forgetting its cross-lingual knowledge obtained from the pre-training stage. This forgetting phenomenon degenerates the zero-shot performance of MLF, which remains under-explored. To fill this gap, this paper proposes a multi-lingual fine-tuning method, dubbed Less-forgetting Multi-lingual Fine-tuning (LF-MLF). In LF-MLF, we cast multi-lingual fine-tuning as a constrained optimization problem, where the optimization objective is to minimize forgetting, and constraints are reducing the fine-tuning loss. The proposed method has superior zero-shot performance; furthermore, it can achieve the Pareto stationarity. Extensive experiments on Named Entity Recognition, Question Answering and Natural Language Inference back up our theoretical analysis and validate the superiority of our proposals. Yuren Mao, Yaobo Liang, Haobo Wang 0001, Kai Wang 0037, Lu Chen 0001, Yunjun Gao |
NeurIPS | 2 |
| 2021 | Simpson's Bias in NLP TrainingabstractIn most machine learning tasks, we evaluate a model M on a given data population S by measuring a population-level metric F(S;M). Examples of such evaluation metric F include precision/recall for (binary) recognition, the F1 score for multi-class classification, and the BLEU metric for language generation. On the other hand, the model M is trained by optimizing a sample-level loss G(S_t; M) at each learning step t, where S_t is a subset of S (a.k.a. the mini-batch). Popular choices of G include cross-entropy loss, the Dice loss, and sentence-level BLEU scores. A fundamental assumption behind this paradigm is that the mean value of the sample-level loss G, if averaged over all possible samples, should effectively represent the population-level metric F of the task, such as, that E[ G(S_t; M) ] ~ F(S; M). In this paper, we systematically investigate the above assumption in several NLP tasks. We show, both theoretically and experimentally, that some popular designs of the sample-level loss G may be inconsistent with the true population-level metric F of the task, so that models trained to optimize the former can be substantially sub-optimal to the latter, a phenomenon we call it, Simpson's bias, due to its deep connections with the classic paradox known as Simpson's reversal paradox in statistics and social sciences. Longtu Zhang, Bojun Huang, Yaobo Liang |
AAAI | 4 |
| 2021 | GLOW : Global Weighted Self-Attention Network for Web SearchabstractDeep matching models aim to facilitate search engines retrieving more relevant documents by mapping queries and documents into semantic vectors in the first-stage retrieval. When leveraging BERT as the deep matching model, the attention score across two words are solely built upon local contextualized word embeddings. It lacks prior global knowledge to distinguish the importance of different words, which has been proved to play a critical role in information retrieval tasks. In addition to this, BERT only performs attention across sub-words tokens which weakens whole word attention representation. We propose a novel Global Weighted Self-Attention (GLOW) network for web document search. GLOW fuses global corpus statistics into the deep matching model. By adding prior weights into attention generation from global information, like BM25, GLOW successfully learns weighted attention scores jointly with query matrix Q and key matrix K. We also present an efficient whole word weight sharing solution to bring prior whole word knowledge into sub-words level attention. It aids Transformer to learn whole word level attention. To make our models applicable to complicated web search scenarios, we introduce combined fields representation to accommodate documents with multiple fields even with variable number of instances. We demonstrate GLOW is more efficient to capture the topical and semantic representation both in queries and documents. Intrinsic evaluation and experiments conducted on public data sets reveal GLOW to be a general framework for document retrieve task. It significantly outperforms BERT and other competitive baselines by a large margin while retaining the same model complexity with BERT. The source code is available at https://github.com/GLOW-deep/GLOW. Xuan Shan, Chuanjie Liu, Yiqian Xia, Qi Chen 0009, Kaize Ding, Yaobo Liang, Angen Luo, Yuxiang Luo |
IEEE BigData | 7 |
| 2020 | Enhancing Answer Boundary Detection for Multilingual Machine Reading ComprehensionabstractMultilingual pre-trained models could leverage the training data from a rich source language (such as English) to improve the performance on low resource languages.However, the transfer effectiveness on the multilingual Machine Reading Comprehension (MRC) task is substantially poorer than that for sentence classification tasks, mainly due to the requirement of MRC to detect the word level answer boundary.In this paper, we propose two auxiliary tasks to introduce additional phrase boundary supervision in the fine-tuning stage:(1) a mixed MRC task, which translates the question or passage to other languages and builds cross-lingual question-passage pairs; and (2) a language-agnostic knowledge masking task by leveraging knowledge phrases mined from the Web.Extensive experiments on two cross-lingual MRC datasets show the effectiveness of our proposed approach.† Random N-gram Masking shows gains in English SQuAD. Fei Yuan 0010, Linjun Shou, Xuanyu Bai, Ming Gong 0001, Yaobo Liang, Nan Duan 0001, Daxin Jiang |
ACL | 5 |
| 2020 | Document Modeling with Graph Attention Networks for Multi-grained Machine Reading ComprehensionabstractNatural Questions is a new challenging machine reading comprehension benchmark with two-grained answers, which are a long answer (typically a paragraph) and a short answer (one or more entities inside the long answer).Despite the effectiveness of existing methods on this benchmark, they treat these two sub-tasks individually during training while ignoring their dependencies.To address this issue, we present a novel multi-grained machine reading comprehension framework that focuses on modeling documents at their hierarchical nature, which are different levels of granularity: documents, paragraphs, sentences, and tokens.We utilize graph attention networks to obtain different levels of representations so that they can be learned simultaneously.The long and short answers can be extracted from paragraphlevel representation and token-level representation, respectively.In this way, we can model the dependencies between the two-grained answers to provide evidence for each other.We jointly train the two sub-tasks, and our experiments show that our approach significantly outperforms previous systems at both long and short answer criteria. Bo Zheng 0010, Haoyang Wen, Yaobo Liang, Nan Duan 0001, Wanxiang Che, Daxin Jiang, Ming Zhou 0001, Ting Liu 0001 |
ACL | 3 |
| 2020 | Cross-Lingual Transfer Learning for Medical Named Entity Recognition
Pengjie Ding, Yaobo Liang, Wei Lu 0015, Buzhou Tang, Jun Yan 0010 |
DASFAA (1) | 3 |
| 2020 | XGLUE: A New Benchmark Datasetfor Cross-lingual Pre-training, Understanding and GenerationabstractYaobo Liang, Nan Duan, Yeyun Gong, Ning Wu, Fenfei Guo, Weizhen Qi, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Cui, Sining Wei, Taroon Bharti, Ying Qiao, Jiun-Hung Chen, Winnie Wu, Shuguang Liu, Fan Yang, Daniel Campos, Rangan Majumder, Ming Zhou. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Yaobo Liang, Nan Duan 0001, Yeyun Gong, Ning Wu 0013, Fenfei Guo, Weizhen Qi, Ming Gong 0001, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Dong Bo Cui, Sining Wei, Taroon Bharti, Jiun-Hung Chen, Winnie Wu, Fan Yang 0024, Daniel Campos, Rangan Majumder, Ming Zhou 0001 |
EMNLP (1) | 1 |
| 2019 | Dense Procedure Captioning in Narrated Instructional VideosabstractUnderstanding narrated instructional videos is important for both research and real-world web applications.Motivated by video dense captioning, we propose a model to generate procedure captions from narrated instructional videos which are a sequence of stepwise clips with description.Previous works on video dense captioning learn video segments and generate captions without considering transcripts.We argue that transcripts in narrated instructional videos can enhance video representation by providing fine-grained complimentary and semantic textual information.In this paper, we introduce a framework to ( 1) extract procedures by a cross-modality module, which fuses video content with the entire transcript; and (2) generate captions by encoding video frames as well as a snippet of transcripts within each extracted procedure.Experiments show that our model can achieve state-of-the-art performance in procedure extraction and captioning, and the ablation studies demonstrate that both the video frames and the transcripts are important for the task. Botian Shi, Lei Ji 0001, Yaobo Liang, Nan Duan 0001, Peng Chen 0029, Zhendong Niu, Ming Zhou 0001 |
ACL (1) | 3 |
| 2019 | Unicoder: A Universal Language Encoder by Pre-training with Multiple Cross-lingual TasksabstractHaoyang Huang, Yaobo Liang, Nan Duan, Ming Gong, Linjun Shou, Daxin Jiang, Ming Zhou. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Haoyang Huang, Yaobo Liang, Nan Duan 0001, Ming Gong 0001, Linjun Shou, Daxin Jiang, Ming Zhou 0001 |
EMNLP/IJCNLP (1) | 2 |