VLDB 2026 Research / reviewers in the wild / expert
Shuo Ren 0002
dblp:147/6063-2
· DBLP profile ↗
20ranked-venue papers
7as first author
12since 2021 · last 2025
0009-0003-6948-9270ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 7 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Teaching Vision-Language Models to Ask: Resolving Ambiguity in Visual QuestionsabstractIn visual question answering (VQA) context, users often pose ambiguous questions to visual language models (VLMs) due to varying expression habits.Existing research addresses such ambiguities primarily by rephrasing questions.These approaches neglect the inherently interactive nature of user interactions with VLMs, where ambiguities can be clarified through user feedback.However, research on interactive clarification faces two major challenges: (1) Benchmarks are absent to assess VLMs' capacity for resolving ambiguities through interaction; (2) VLMs are trained to prefer answering rather than asking, preventing them from seeking clarification.To overcome these challenges, we introduce ClearVQA benchmark 1 , which targets three common categories of ambiguity in VQA context, and encompasses various VQA scenarios.Furthermore, we propose an automated pipeline to generate ambiguity-clarification question pairs.Experimental results demonstrate that training based on the automated generated data enables VLMs to ask reasonable clarification questions, thereby generating more accurate and specific answers based on user feedback. Pu Jian, Donglei Yu, Shuo Ren 0002, Jiajun Zhang 0001 |
ACL (1) | 4 |
| 2025 | Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language ModelsabstractRecent advances in text-only "slow-thinking" reasoning have prompted efforts to transfer this capability to vision-language models (VLMs), for training visual reasoning models (VRMs).However, such transfer faces critical challenges: Effective "slow thinking" in VRMs requires visual reflection, the ability to check the reasoning process based on visual information.Through quantitative analysis, we observe that current VRMs exhibit limited visual reflection, as their attention to visual information diminishes rapidly with longer generated responses.To address this challenge, we propose a new VRM Reflection-V 1 , which enhances visual reflection based on reasoning data construction for cold-start and reward design for reinforcement learning (RL).Firstly, we construct vision-centered reasoning data by leveraging an agent that interacts between VLMs and reasoning LLMs, enabling cold-start learning of visual reflection patterns.Secondly, a visual attention based reward model is employed during RL to encourage reasoning based on visual information.Therefore, Reflection-V demonstrates significant improvements across multiple visual reasoning benchmarks.Furthermore, Reflection-V maintains a stronger and more consistent reliance on visual information during visual reasoning, indicating effective enhancement in visual reflection capabilities. Pu Jian, Junhong Wu, Shuo Ren 0002, Jiajun Zhang 0001 |
EMNLP | 5 |
| 2025 | Collaborative Beam Search: Enhancing LLM Reasoning via Collective ConsensusabstractComplex multi-step reasoning remains challenging for large language models (LLMs).While parallel inference-time scaling methods, such as step-level beam search, offer a promising solution, existing approaches typically depend on either domain-specific external verifiers, or self-evaluation which is brittle and prompt-sensitive.To address these issues, we propose Collaborative Beam Search (CBS), an iterative framework that harnesses the collective intelligence of multiple LLMs across both generation and verification stages.For generation, CBS leverages multiple LLMs to explore a broader search space, resulting in more diverse candidate steps.For verifications, CBS employs a perplexity-based collective consensus among these models, eliminating reliance on an external verifier or complex prompts.Between iterations, CBS leverages a dynamic quota allocation strategy that reassigns generation budget based on each model's past performance, striking a balance between candidate diversity and quality.Experimental results on six tasks across arithmetic, logical, and commonsense reasoning show that CBS outperforms single-model scaling and multi-model ensemble baselines by over 4 percentage points in average accuracy, demonstrating its effectiveness and general applicability. Yangyifan Xu, Shuo Ren 0002, Jiajun Zhang 0001 |
EMNLP | 2 |
| 2025 | KTAE: A Model-Free Algorithm to Key-Tokens Advantage Estimation in Mathematical ReasoningabstractRecent advances have demonstrated that integrating reinforcement learning with rule-based rewards can significantly enhance the reasoning capabilities of large language models (LLMs), even without supervised fine-tuning (SFT). However, prevalent reinforcement learning algorithms such as GRPO and its variants like DAPO, suffer from a coarse granularity issue when computing the advantage. Specifically, they compute rollout-level advantages that assign identical values to every token within a sequence, failing to capture token-specific contributions. To address this limitation, we propose Key-token Advantage Estimation (KTAE)—a novel algorithm that estimates fine-grained, token-level advantages without introducing additional models. KTAE leverages the correctness of sampled rollouts and applies statistical analysis to quantify the importance of individual tokens within a sequence to the final outcome. This quantified token-level importance is then combined with the rollout-level advantage to obtain a more fine-grained token-level advantage estimation. Empirical results show that models trained with GRPO+KTAE and DAPO+KTAE outperform baseline methods across five mathematical reasoning benchmarks. Notably, they achieve higher accuracy with shorter responses and even surpass R1-Distill-Qwen-1.5B using the same base model. Pu Jian, Qianlong Du, Fuwei Cui, Shuo Ren 0002, Jiajun Zhang 0001 |
NeurIPS | 6 |
| 2025 | Beyond One-Size-Fits-All: Adaptive Fine-Tuning for LLMs Based on Data Inherent Heterogeneity
Wanyue Zhang, Yangyifan Xu, Shuo Ren 0002, Jiajun Zhang 0001 |
NLPCC (1) | 3 |
| 2024 | SpeechLM: Enhanced Speech Pre-Training With Unpaired Textual DataabstractHow to boost speech pre-training with textual data is an unsolved problem due to the fact that speech and text are very different modalities with distinct characteristics. In this paper, we propose a cross-modalSpeechandLanguageModel (SpeechLM) to explicitly align speech and text pre-training with a pre-defined unified discrete representation. Specifically, we introduce two alternative discrete tokenizers to bridge the speech and text modalities, including phoneme-unit and hidden-unit tokenizers, which can be trained using unpaired speech or a small amount of paired speech-text data. Based on the trained tokenizers, we convert the unlabeled speech and text data into tokens of phoneme units or hidden units. The pre-training objective is designed to unify the speech and the text into the same discrete semantic space with a unified Transformer network. We evaluate SpeechLM on various spoken language processing tasks including speech recognition, speech translation, and universal representation evaluation framework SUPERB, demonstrating significant improvements on content-related tasks. Code and models are available athttps://aka.ms/SpeechLM. Sanyuan Chen, Yu Wu 0012, Shuo Ren 0002, Shujie Liu 0001, Zhuoyuan Yao, Xun Gong 0005, Li-Rong Dai 0001, Jinyu Li 0001, Furu Wei |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2022 | SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language ProcessingabstractJunyi Ao, Rui Wang, Long Zhou, Chengyi Wang, Shuo Ren, Yu Wu, Shujie Liu, Tom Ko, Qing Li, Yu Zhang, Zhihua Wei, Yao Qian, Jinyu Li, Furu Wei. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Junyi Ao, Rui Wang 0073, Chengyi Wang 0002, Shuo Ren 0002, Yu Wu 0012, Shujie Liu 0001, Tom Ko, Qing Li 0001, Yu Zhang 0006, Zhihua Wei 0001, Yao Qian, Jinyu Li 0001, Furu Wei |
ACL (1) | 5 |
| 2022 | RAPO: An Adaptive Ranking Paradigm for Bilingual Lexicon InductionabstractZhoujin Tian, Chaozhuo Li, Shuo Ren, Zhiqiang Zuo, Zengxuan Wen, Xinyue Hu, Xiao Han, Haizhen Huang, Denvy Deng, Qi Zhang, Xing Xie. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Zhoujin Tian, Chaozhuo Li, Shuo Ren 0002, Zhiqiang Zuo 0004, Zengxuan Wen, Xinyue Hu 0003, Haizhen Huang, Denvy Deng, Qi Zhang 0066, Xing Xie 0001 |
EMNLP | 3 |
| 2022 | Optimizing Alignment of Speech and Language Latent Spaces for End-To-End Speech Recognition and UnderstandingabstractThe advances in attention-based encoder-decoder (AED) networks have brought great progress to end-to-end (E2E) automatic speech recognition (ASR). One way to further improve the performance of AED-based E2E ASR is to introduce an extra text encoder for leveraging extensive text data and thus capture more context-aware linguistic information. However, this approach brings a mismatch problem between the speech encoder and the text encoder due to the different units used for modeling. In this paper, we propose an embedding aligner and modality switch training to better align the speech and text latent spaces. The embedding aligner is a shared linear projection between text encoder and speech encoder trained by masked language modeling (MLM) loss and connectionist temporal classification (CTC), respectively. The modality switch training randomly swaps speech and text embeddings based on the forced alignment result to learn a joint representation space. Experimental results show that our proposed approach achieves a relative 14% to 19% word error rate (WER) reduction on Librispeech ASR task. We further verify its effectiveness on spoken language understanding (SLU), i.e., an absolute 2.5% to 2.8% F1 score improvement on SNIPS slot filling task. Wei Wang 0010, Shuo Ren 0002, Yao Qian, Shujie Liu 0001, Yu Shi 0001, Yanmin Qian, Michael Zeng 0001 |
ICASSP | 2 |
| 2022 | Speech Pre-training with Acoustic PieceabstractPrevious speech pre-training methods, such as wav2vec2.0 and HuBERT, pre-train a Transformer encoder to learn deep representations from audio data, with objectives predicting either elements from latent vector quantized space or pre-generated labels (known as target codes) with offline clustering. However, those training signals (quantized elements or codes) are independent across different tokens without considering their relations. According to our observation and analysis, the target codes share obvious patterns aligned with phonemized text data. Based on that, we propose to leverage those patterns to better pre-train the model considering the relations among the codes. The patterns we extracted, called "acoustic piece"s, are from the sentence piece result of HuBERT codes. With the acoustic piece as the training signal, we can implicitly bridge the input audio and natural language, which benefits audio-to-text tasks, such as automatic speech recognition (ASR). Simple but effective, our method "HuBERT-AP" significantly outperforms strong baselines on the LibriSpeech ASR task. Shuo Ren 0002, Shujie Liu 0001, Yu Wu 0012, Furu Wei |
INTERSPEECH | 1 |
| 2021 | SemFace: Pre-training Encoder and Decoder with a Semantic Interface for Neural Machine TranslationabstractShuo Ren, Long Zhou, Shujie Liu, Furu Wei, Ming Zhou, Shuai Ma. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Shuo Ren 0002, Shujie Liu 0001, Furu Wei, Ming Zhou 0001, Shuai Ma 0001 |
ACL/IJCNLP (1) | 1 |
| 2021 | GraphCodeBERT: Pre-training Code Representations with Data Flow
Daya Guo, Shuo Ren 0002, Zhangyin Feng, Duyu Tang, Shujie Liu 0001, Nan Duan 0001, Alexey Svyatkovskiy, Shengyu Fu, Michele Tufano, Shao Kun Deng, Colin B. Clement, Dawn Drain, Neel Sundaresan, Jian Yin 0001, Daxin Jiang, Ming Zhou 0001 |
ICLR | 2 |
| 2020 | A Graph-based Coarse-to-fine Method for Unsupervised Bilingual Lexicon InductionabstractUnsupervised bilingual lexicon induction is the task of inducing word translations from monolingual corpora of two languages.Recent methods are mostly based on unsupervised cross-lingual word embeddings, the key to which is to find initial solutions of word translations, followed by the learning and refinement of mappings between the embedding spaces of two languages.However, previous methods find initial solutions just based on word-level information, which may be (1) limited and inaccurate, and (2) prone to contain some noise introduced by the insufficiently pre-trained embeddings of some words.To deal with those issues, in this paper, we propose a novel graph-based paradigm to induce bilingual lexicons in a coarse-to-fine way.We first build a graph for each language with its vertices representing different words.Then we extract word cliques from the graphs and map the cliques of two languages.Based on that, we induce the initial word translation solution with the central words of the aligned cliques.This coarse-to-fine approach not only leverages clique-level information, which is richer and more accurate, but also effectively reduces the bad effect of the noise in the pre-trained embeddings.Finally, we take the initial solution as the seed to learn cross-lingual embeddings, from which we induce bilingual lexicons.Experiments show that our approach improves the performance of bilingual lexicon induction compared with previous methods. Shuo Ren 0002, Shujie Liu 0001, Ming Zhou 0001, Shuai Ma 0001 |
ACL | 1 |
| 2020 | A Retrieve-and-Rewrite Initialization Method for Unsupervised Machine TranslationabstractThe commonly used framework for unsupervised machine translation builds initial translation models of both translation directions, and then performs iterative back-translation to jointly boost their translation performance.The initialization stage is very important since bad initialization may wrongly squeeze the search space, and too much noise introduced in this stage may hurt the final performance.In this paper, we propose a novel retrieval and rewriting based method to better initialize unsupervised translation models.We first retrieve semantically comparable sentences from monolingual corpora of two languages and then rewrite the target side to minimize the semantic gap between the source and retrieved targets with a designed rewriting model.The rewritten sentence pairs are used to initialize SMT models which are used to generate pseudo data for two NMT models, followed by the iterative back-translation.Experiments show that our method can build better initial unsupervised translation models and improve the final translation performance by over 4 BLEU scores. Shuo Ren 0002, Yu Wu 0012, Shujie Liu 0001, Ming Zhou 0001, Shuai Ma 0001 |
ACL | 1 |
| 2020 | Semantic Mask for Transformer Based End-to-End Speech RecognitionabstractAttention-based encoder-decoder model has achieved impressive results for both automatic speech recognition (ASR) and text-to-speech (TTS) tasks.This approach takes advantage of the memorization capacity of neural networks to learn the mapping from the input sequence to the output sequence from scratch, without the assumption of prior knowledge such as the alignments.However, this model is prone to overfitting, especially when the amount of training data is limited.Inspired by SpecAugment and BERT, in this paper, we propose a semantic mask based regularization for training such kind of end-toend (E2E) model.The idea is to mask the input features corresponding to a particular output token, e.g., a word or a wordpiece, in order to encourage the model to fill the token based on the contextual information.While this approach is applicable to the encoder-decoder framework with any type of neural network architecture, we study the transformer-based model for ASR in this work.We perform experiments on Librispeech 960h and TedLium2 data sets, and achieve the state-of-the-art performance on the test set in the scope of E2E models. Chengyi Wang 0002, Yu Wu 0012, Yujiao Du, Jinyu Li 0001, Shujie Liu 0001, Liang Lu 0001, Shuo Ren 0002, Guoli Ye, Sheng Zhao 0002, Ming Zhou 0001 |
INTERSPEECH | 7 |
| 2019 | Unsupervised Neural Machine Translation with SMT as Posterior RegularizationabstractWithout real bilingual corpus available, unsupervised Neural Machine Translation (NMT) typically requires pseudo parallel data generated with the back-translation method for the model training. However, due to weak supervision, the pseudo data inevitably contain noises and errors that will be accumulated and reinforced in the subsequent training process, leading to bad translation performance. To address this issue, we introduce phrase based Statistic Machine Translation (SMT) models which are robust to noisy data, as posterior regularizations to guide the training of unsupervised NMT models in the iterative back-translation process. Our method starts from SMT models built with pre-trained language models and word-level translation tables inferred from cross-lingual embeddings. Then SMT and NMT models are optimized jointly and boost each other incrementally in a unified EM framework. In this way, (1) the negative effect caused by errors in the iterative back-translation process can be alleviated timely by SMT filtering noises from its phrase tables; meanwhile, (2) NMT can compensate for the deficiency of fluency inherent in SMT. Experiments conducted on en-fr and en-de translation tasks show that our method outperforms the strong baseline and achieves new state-of-the-art unsupervised machine translation performance. Shuo Ren 0002, Zhirui Zhang, Shujie Liu 0001, Ming Zhou 0001, Shuai Ma 0001 |
AAAI | 1 |
| 2019 | Explicit Cross-lingual Pre-training for Unsupervised Machine TranslationabstractShuo Ren, Yu Wu, Shujie Liu, Ming Zhou, Shuai Ma. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Shuo Ren 0002, Yu Wu 0012, Shujie Liu 0001, Ming Zhou 0001, Shuai Ma 0001 |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Triangular Architecture for Rare Language TranslationabstractNeural Machine Translation (NMT) performs poor on the low-resource language pair (X, Z), especially when Z is a rare language.By introducing another rich language Y , we propose a novel triangular training architecture (TA-NMT) to leverage bilingual data (Y, Z) (may be small) and (X, Y ) (can be rich) to improve the translation performance of lowresource pairs.In this triangular architecture, Z is taken as the intermediate latent variable, and translation models of Z are jointly optimized with a unified bidirectional EM algorithm under the goal of maximizing the translation likelihood of (X, Y ).Empirical results demonstrate that our method significantly improves the translation quality of rare languages on MultiUN and IWSLT2012 datasets, and achieves even better performance combining back-translation methods. Shuo Ren 0002, Wenhu Chen, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Shuai Ma 0001 |
ACL (1) | 1 |
| 2018 | Generative Bridging Network for Neural Sequence PredictionabstractWenhu Chen, Guanlin Li, Shuo Ren, Shujie Liu, Zhirui Zhang, Mu Li, Ming Zhou. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Wenhu Chen, Shuo Ren 0002, Shujie Liu 0001, Zhirui Zhang, Mu Li 0001, Ming Zhou 0001 |
NAACL-HLT | 3 |
| 2016 | Knowledge-Based Semantic Embedding for Machine TranslationabstractIn this paper, with the help of knowledge base, we build and formulate a semantic space to connect the source and target languages, and apply it to the sequence-to-sequence framework to propose a Knowledge-Based Semantic Embedding (KBSE) method.In our KB-SE method, the source sentence is firstly mapped into a knowledge based semantic space, and the target sentence is generated using a recurrent neural network with the internal meaning preserved.Experiments are conducted on two translation tasks, the electric business data and movie data, and the results show that our proposed method can achieve outstanding performance, compared with both the traditional SMT methods and the existing encoder-decoder models. Shujie Liu 0001, Shuo Ren 0002, Mu Li 0001, Ming Zhou 0001, Xu Sun 0001, Houfeng Wang |
ACL (1) | 3 |