VLDB 2026 Research / reviewers in the wild / expert
Raphael Tang
dblp:207/7684
· DBLP profile ↗
21ranked-venue papers
9as first author
14since 2021 · last 2026
0009-0007-2873-892XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 11 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Role of Mixed-Language Documents for Multilingual Large Language Model PretrainingabstractJiandong Shao, Raphael Tang, Crystina Zhang, Karin Sevegnani, Pontus Stenetorp, Jianfei Yang, Yao Lu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiandong Shao, Raphael Tang, Xinyu Zhang 0018, Karin Sevegnani, Pontus Stenetorp |
ACL (1) | 2 |
| 2025 | Rank-Without-GPT: Building GPT-Independent Listwise Rerankers on Open-Source Large Language Models
Xinyu Zhang 0018, Sebastian Hofstätter, Patrick Lewis 0002, Raphael Tang, Jimmy Lin |
ECIR (2) | 4 |
| 2025 | Multilingual Language Model Pretraining using Machine-translated DataabstractJiayi Wang, Yao Lu, Maurice Weber, Max Ryabinin, David Ifeoluwa Adelani, Yihong Chen, Raphael Tang, Pontus Stenetorp. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Jiayi Wang 0010, Maurice Weber, Max Ryabinin, David Ifeoluwa Adelani, Raphael Tang, Pontus Stenetorp |
EMNLP | 7 |
| 2024 | Understanding Retrieval Robustness for Retrieval-augmented Image CaptioningabstractRecent advances in retrieval-augmented models for image captioning highlight the benefit of retrieving related captions for efficient, lightweight models with strong domain-transfer capabilities.While these models demonstrate the success of retrieval augmentation, retrieval models are still far from perfect in practice: the retrieved information can sometimes mislead the model, resulting in incorrect generation and worse performance.In this paper, we analyze the robustness of a retrieval-augmented captioning model SMALLCAP.Our analysis shows that the model is sensitive to tokens that appear in the majority of the retrieved captions, and the input attribution shows that those tokens are likely copied into the generated output.Given these findings, we propose to train the model by sampling retrieved captions from more diverse sets.This decreases the chance that the model learns to copy majority tokens, and improves both in-domain and cross-domain performance. Wenyan Li 0001, Jiaang Li 0002, Rita Ramos, Raphael Tang, Desmond Elliott |
ACL (1) | 4 |
| 2024 | FoodieQA: A Multimodal Dataset for Fine-Grained Understanding of Chinese Food CultureabstractWenyan Li, Crystina Zhang, Jiaang Li, Qiwei Peng, Raphael Tang, Li Zhou, Weijia Zhang, Guimin Hu, Yifei Yuan, Anders Søgaard, Daniel Hershcovich, Desmond Elliott. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Wenyan Li 0001, Xinyu Zhang 0018, Jiaang Li 0002, Qiwei Peng 0003, Raphael Tang, Li Zhou 0010, Weijia Zhang 0004, Guimin Hu, Yifei Yuan 0002, Anders Søgaard, Daniel Hershcovich, Desmond Elliott |
EMNLP | 5 |
| 2024 | Words Worth a Thousand Pictures: Measuring and Understanding Perceptual Variability in Text-to-Image GenerationabstractDiffusion models are the state of the art in textto-image generation, but their perceptual variability remains understudied.In this paper, we examine how prompts affect image variability in black-box diffusion-based models.We propose W1KP, a human-calibrated measure of variability in a set of images, bootstrapped from existing image-pair perceptual distances.Current datasets do not cover recent diffusion models, thus we curate three test sets for evaluation.Our best perceptual distance outperforms nine baselines by up to 18 points in accuracy, and our calibration matches graded human judgements 78% of the time.Using W1KP, we study prompt reusability and show that Imagen prompts can be reused for 10-50 random seeds before new images become too similar to already generated images, while Stable Diffusion XL and DALL-E 3 can be reused 50-200 times.Lastly, we analyze 56 linguistic features of real prompts, finding that the prompt's length, CLIP embedding norm, concreteness, and word senses influence variability most.As far as we are aware, we are the first to analyze diffusion variability from a visuolinguistic perspective.Our project page is at http://w1kp.com. Raphael Tang, Xinyu Zhang 0018, Lixinyu Xu, Wenyan Li 0001, Pontus Stenetorp, Jimmy Lin, Ferhan Ture |
EMNLP | 1 |
| 2024 | Strings from the Library of Babel: Random Sampling as a Strong Baseline for Prompt OptimisationabstractYao Lu, Jiayi Wang, Raphael Tang, Sebastian Riedel, Pontus Stenetorp. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jiayi Wang 0010, Raphael Tang, Sebastian Riedel 0001, Pontus Stenetorp |
NAACL-HLT | 3 |
| 2024 | Found in the Middle: Permutation Self-Consistency Improves Listwise Ranking in Large Language ModelsabstractRaphael Tang, Crystina Zhang, Xueguang Ma, Jimmy Lin, Ferhan Ture. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Raphael Tang, Xinyu Zhang 0018, Xueguang Ma, Jimmy Lin, Ferhan Ture |
NAACL-HLT | 1 |
| 2024 | "Ask Me Anything": How Comcast Uses LLMs to Assist Agents in Real TimeabstractCustomer service is how companies interface with their customers. It can contribute heavily towards the overall customer satisfaction. However, high-quality service can become expensive, creating an incentive to make it as cost efficient as possible and prompting most companies to utilize AI-powered assistants, or "chat bots". On the other hand, human-to-human interaction is still desired by customers, especially when it comes to complex scenarios such as disputes and sensitive topics like bill payment. Scott Rome, Tianwen Chen, Raphael Tang, Luwei Zhou, Ferhan Ture |
SIGIR | 3 |
| 2023 | What the DAAM: Interpreting Stable Diffusion Using Cross AttentionabstractRaphael Tang, Linqing Liu, Akshat Pandey, Zhiying Jiang, Gefei Yang, Karun Kumar, Pontus Stenetorp, Jimmy Lin, Ferhan Ture. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Raphael Tang, Linqing Liu, Akshat Pandey, Zhiying Jiang, Gefei Yang, Karun Kumar, Pontus Stenetorp, Jimmy Lin, Ferhan Ture |
ACL (1) | 1 |
| 2022 | Temporal Early Exiting for Streaming Speech Commands RecognitionabstractLimited-vocabulary speech commands recognition is the task of classifying a short utterance as one of several speech commands, for which neural networks obtain state-of-the-art results. In particular, recurrent neural networks represent a common approach for streaming commands recognition systems. In this paper, we explore resource-efficient methods to short-circuit such systems in the time domain when the model is confident in its prediction. We propose applying a frame-level labeling objective to further improve the efficiency–accuracy trade-off. On two datasets in limited-vocabulary commands recognition, our best method achieves an average time savings of 45% of the utterance without reducing the absolute accuracy by more than 0.6 points. We show that the per-instance savings depend on the length of the unique prefix in the phonemes across a dataset. Raphael Tang, Karun Kumar, Ji Xin, Piyush Vyas, Wenyan Li 0001, Gefei Yang, Yajie Mao, G. Craig Murray, Jimmy Lin |
ICASSP | 1 |
| 2021 | The Art of Abstention: Selective Prediction and Error Regularization for Natural Language ProcessingabstractJi Xin, Raphael Tang, Yaoliang Yu, Jimmy Lin. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ji Xin, Raphael Tang, Yaoliang Yu, Jimmy Lin |
ACL/IJCNLP (1) | 2 |
| 2021 | BERxiT: Early Exiting for BERT with Better Fine-Tuning and Extension to RegressionabstractThe slow speed of BERT has motivated much research on accelerating its inference, and the early exiting idea has been proposed to make trade-offs between model quality and efficiency.This paper aims to address two weaknesses of previous work: (1) existing fine-tuning strategies for early exiting models fail to take full advantage of BERT; (2) methods to make exiting decisions are limited to classification tasks.We propose a more advanced fine-tuning strategy and a learning-toexit module that extends early exiting to tasks other than classification.Experiments demonstrate improved early exiting for BERT, with better trade-offs obtained by the proposed finetuning strategy, successful application to regression tasks, and the possibility to combine it with other acceleration methods. Ji Xin, Raphael Tang, Yaoliang Yu, Jimmy Lin |
EACL | 2 |
| 2021 | Voice Query Auto CompletionabstractQuery auto completion (QAC) is the task of predicting a search engine user’s final query from their intermediate, incomplete query. In this paper, we extend QAC to the streaming voice search setting, where automatic speech recognition systems produce intermediate transcriptions as users speak. Naively applying existing methods fails because the intermediate transcriptions often don’t form prefixes or even substrings of the final transcription. To address this issue, we propose to condition QAC approaches on intermediate transcriptions to complete voice queries. We evaluate our models on a speech-enabled smart television with real-life voice search traffic, finding that this ASR-aware conditioning improves the completion quality. Our best method obtains an 18% relative improvement in mean reciprocal rank over previous methods. Raphael Tang, Karun Kumar, Kendra Chalkley, Ji Xin, Wenyan Li 0001, Gefei Yang, Yajie Mao, Junho Shin, G. Craig Murray, Jimmy Lin |
EMNLP (1) | 1 |
| 2020 | Showing Your Work Doesn't Always WorkabstractIn natural language processing, a recently popular line of work explores how to best report the experimental results of neural networks.One exemplar publication, titled "Show Your Work: Improved Reporting of Experimental Results" (Dodge et al., 2019), advocates for reporting the expected validation effectiveness of the best-tuned model, with respect to the computational budget.In the present work, we critically examine this paper.As far as statistical generalizability is concerned, we find unspoken pitfalls and caveats with this approach.We analytically show that their estimator is biased and uses error-prone assumptions.We find that the estimator favors negative errors and yields poor bootstrapped confidence intervals.We derive an unbiased alternative and bolster our claims with empirical evidence from statistical simulation.Our codebase is at https://github.com/ castorini/meanmax. Raphael Tang, Ji Xin, Yaoliang Yu, Jimmy Lin |
ACL | 1 |
| 2020 | DeeBERT: Dynamic Early Exiting for Accelerating BERT InferenceabstractLarge-scale pre-trained language models such as BERT have brought significant improvements to NLP applications.However, they are also notorious for being slow in inference, which makes them difficult to deploy in realtime applications.We propose a simple but effective method, DeeBERT, to accelerate BERT inference.Our approach allows samples to exit earlier without passing through the entire model.Experiments show that DeeBERT is able to save up to ∼40% inference time with minimal degradation in model quality.Further analyses show different behaviors in the BERT transformer layers and also reveal their redundancy.Our work provides new ideas to efficiently apply deep transformer-based models to downstream tasks. Ji Xin, Raphael Tang, Yaoliang Yu, Jimmy Lin |
ACL | 2 |
| 2019 | Incorporating Contextual and Syntactic Structures Improves Semantic Similarity ModelingabstractLinqing Liu, Wei Yang, Jinfeng Rao, Raphael Tang, Jimmy Lin. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Linqing Liu, Wei Yang 0017, Jinfeng Rao, Raphael Tang, Jimmy Lin |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Yelling at Your TV: An Analysis of Speech Recognition Errors and Subsequent User Behavior on Entertainment SystemsabstractMillions of consumers issue voice queries through television-based entertainment systems such as the Comcast X1, the Amazon Fire TV, and Roku TV. Automatic speech recognition (ASR) systems are responsible for transcribing these voice queries into text to feed downstream natural language understanding modules. However, ASR is far from perfect, often producing incorrect transcriptions and forcing users to take corrective action. To better understand their impact on sessions, this paper characterizes speech recognition errors as well as subsequent user responses. We provide both quantitative and qualitative analyses, examining the acoustic as well as lexical attributes of the utterances. This work represents, to our knowledge, the first analysis of speech recognition errors from real users on a widely-deployed entertainment system. Raphael Tang, Ferhan Ture, Jimmy Lin |
SIGIR | 1 |
| 2019 | Challenges and Opportunities in Understanding Spoken Queries Directed at Modern Entertainment PlatformsabstractModern in-home entertainment platforms---representing the evolution of the humble television of yesteryear---are packed with features and content: they offer a dizzying array of programs spanning hundreds of channels as well as a catalog of on-demand programs offering tens of thousands of options. Furthermore, the entertainment platform may serve as an in-home hub, providing capabilities ranging from playing music to controlling the home security system. At a high level, our goal is to provide natural speech-based access to these myriad features as an alternative to physical button entry on a remote control. Ferhan Ture, Jinfeng Rao, Raphael Tang, Jimmy Lin |
SIGIR | 3 |
| 2018 | Deep Residual Learning for Small-Footprint Keyword SpottingabstractWe explore the application of deep residual learning and dilated convolutions to the keyword spotting task, using the recently-released Google Speech Commands Dataset as our benchmark. Our best residual network (ResNet) implementation significantly outperforms Google's previous convolutional neural networks in terms of accuracy. By varying model depth and width, we can achieve compact models that also outperform previous small-footprint variants. To our knowledge, we are the first to examine these approaches for keyword spotting, and our results establish an open-source state-of-the-art reference to support the development of future speech-based interfaces. Raphael Tang, Jimmy Lin |
ICASSP | 1 |
| 2018 | An Experimental Analysis of the Power Consumption of Convolutional Neural Networks for Keyword SpottingabstractNearly all previous work on small-footprint keyword spotting with neural networks quantify model footprint in terms of the number of parameters and multiply operations for a feedforward inference pass. These values are, however, proxy measures since empirical performance in actual deployments is determined by many factors. In this paper, we study the power consumption of a family of convolutional neural networks for keyword spotting on a Raspberry Pi. We find that both proxies are good predictors of energy usage, although the number of multiplies is more predictive than the number of model parameters. We also confirm that models with the highest accuracies are, unsurprisingly, the most power hungry. Raphael Tang, Zhucheng Tu, Jimmy Lin |
ICASSP | 1 |