EDBT 2026 Demo / reviewers in the wild / expert
Yuan Cao 0007
dblp:52/4472-7
· DBLP profile ↗
28ranked-venue papers
1as first author
16since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 1 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | IG Captioner: Information Gain Captioners Are Strong Zero-Shot Classifiers
Siyuan Qiao, Yuan Cao 0007, Yu Zhang 0033, Tao Zhu 0005, Alan L. Yuille |
ECCV (64) | 3 |
| 2024 | Retrieval Augmented End-to-End Spoken Dialog ModelsabstractWe recently developed a joint speech and language model (SLM [1]) which fuses a pretrained foundational speech model and a large language model (LLM), while preserving the in-context learning capability intrinsic to the pretrained LLM. In this paper, we apply SLM to dialog applications where the dialog states are inferred directly from the audio signal.Task-oriented dialogs often contain domain-specific entities, i.e., restaurants, hotels, train stations, and city names, which are difficult to recognize, however, critical for the downstream applications. Inspired by the RAG (retrieval-augmented generation) models, we propose a retrieval augmented SLM (ReSLM) that overcomes this weakness. We first train a retriever to retrieve text entities given audio inputs. The retrieved entities are then added as text inputs to the underlying LLM to bias model predictions. We evaluated ReSLM on speech MultiWoz task (DSTC-11 Challenge), and found that the retrieval augmentation boosts model performance, achieving joint goal accuracy (38.6% vs 32.7%), slot error rate (20.6% vs 24.8%) and ASR word error rate (5.5% vs 6.7%). While demonstrated on dialog state tracking, our approach is broadly applicable to speech tasks requiring custom contextual information or domain-specific entities. Mingqiu Wang, Izhak Shafran, Hagen Soltau, Wei Han 0002, Yuan Cao 0007, Laurent El Shafey |
ICASSP | 5 |
| 2024 | RoboVQA: Multimodal Long-Horizon Reasoning for RoboticsabstractWe present a scalable, bottom-up and intrinsically diverse data collection scheme that can be used for high-level reasoning with long and medium horizons and that has 2.2x higher throughput compared to traditional narrow top-down step-by-step collection. We collect realistic data by performing any user requests within the entirety of 3 office buildings and using multiple embodiments (robot, human, human with grasping tool). With this data, we show that models trained on all embodiments perform better than ones trained on the robot data only, even when evaluated solely on robot episodes. We explore the economics of collection costs and find that for a fixed budget it is beneficial to take advantage of the cheaper human collection along with robot collection. We release a large and highly diverse (29,520 unique instructions) dataset dubbed RoboVQA containing 829,502 (video, text) pairs for robotics-focused visual question answering. We also demonstrate how evaluating real robot experiments with an intervention mechanism enables performing tasks to completion, making it deployable with human oversight even if imperfect while also providing a single performance metric. We demonstrate a single video-conditioned model named RoboVQA-VideoCoCa trained on our dataset that is capable of performing a variety of grounded high-level reasoning tasks in broad realistic settings with a cognitive intervention rate 46% lower than the zeroshot state of the art visual language model (VLM) baseline and is able to guide real robots through long-horizon tasks. The performance gap with zero-shot state-of-the-art models indicates that a lot of grounded data remains to be collected for real-world deployment, emphasizing the critical need for scalable data collection approaches. Finally, we show that video VLMs significantly outperform single-image VLMs with an average error rate reduction of 19% across all VQA tasks. Thanks to video conditioning and dataset diversity, the model can be used as general video value functions (e.g. success and affordance) in situations where actions needs to be recognized rather than states, expanding capabilities and environment understanding for robots. Data and videos are available at robovqa.github.io Pierre Sermanet, Tianli Ding, Jeffrey Zhao, Fei Xia 0002, Debidatta Dwibedi, Keerthana Gopalakrishnan, Christine Chan, Gabriel Dulac-Arnold, Sharath Maddineni, Nikhil J. Joshi, Peter R. Florence, Wei Han 0002, Robert Baruch, Yao Lu 0006, Suvir Mirchandani, Peng Xu 0010, Pannag R. Sanketi, Karol Hausman, Izhak Shafran, Brian Ichter, Yuan Cao 0007 |
ICRA | 21 |
| 2023 | SLM: Bridge the Thin Gap Between Speech and Text Foundation ModelsabstractWe present a joint Speech and Language Model (SLM), a multitask, multilingual, and dual-modal model that takes advantage of pretrained foundational speech and language models. SLM freezes the pretrained foundation models to maximally preserves their capabilities, and only trains a simple adapter with just 1% (156M) of the foundation models’ parameters. This adaptation not only leads SLM to achieve strong performance on conventional tasks such as automatic speech recognition (ASR) and automatic speech translation (AST), but also unlocks the novel capability of zero-shot instruction-following for more diverse tasks. Given a speech input and a text instruction, SLM is able to perform unseen generation tasks including contextual biasing ASR using real-time context, dialog generation, speech continuation, and question answering. Our approach demonstrates that the representational gap between pretrained speech and language models is narrower than one would expect, and can be bridged by a simple adaptation mechanism. As a result, SLM is not only efficient to train, but also inherits strong capabilities already present in foundation models of different modalities. Mingqiu Wang, Wei Han 0002, Izhak Shafran, Zelin Wu, Chung-Cheng Chiu, Yuan Cao 0007, Nanxin Chen, Yu Zhang 0033, Hagen Soltau, Paul K. Rubenstein, Lukas Zilka, Golan Pundak, Nikhil Siddhartha, Johan Schalkwyk |
ASRU | 6 |
| 2023 | AnyTOD: A Programmable Task-Oriented Dialog SystemabstractJeffrey Zhao, Yuan Cao, Raghav Gupta, Harrison Lee, Abhinav Rastogi, Mingqiu Wang, Hagen Soltau, Izhak Shafran, Yonghui Wu. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Jeffrey Zhao, Yuan Cao 0007, Harrison Lee 0001, Abhinav Rastogi, Mingqiu Wang, Hagen Soltau, Izhak Shafran |
EMNLP | 2 |
| 2023 | ReAct: Synergizing Reasoning and Acting in Language Models
Shunyu Yao 0006, Jeffrey Zhao, Nan Du 0002, Izhak Shafran, Karthik Narasimhan, Yuan Cao 0007 |
ICLR | 7 |
| 2023 | Speech Aware Dialog System Technology Challenge (DSTC11)
Hagen Soltau, Izhak Shafran, Mingqiu Wang, Abhinav Rastogi, Jeffrey Zhao, Ye Jia, Wei Han 0002, Yuan Cao 0007, Aramys Miranda |
INTERSPEECH | 8 |
| 2023 | Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsabstractLanguage models are increasingly being deployed for general problem solving across a wide range of tasks, but are still confined to token-level, left-to-right decision-making processes during inference. This means they can fall short in tasks that require exploration, strategic lookahead, or where initial decisions play a pivotal role. To surmount these challenges, we introduce a new framework for language model inference, Tree of Thoughts (ToT), which generalizes over the popular Chain of Thought approach to prompting language models, and enables exploration over coherent units of text (thoughts) that serve as intermediate steps toward problem solving. ToT allows LMs to perform deliberate decision making by considering multiple different reasoning paths and self-evaluating choices to decide the next course of action, as well as looking ahead or backtracking when necessary to make global choices.
Our experiments show that ToT significantly enhances language models’ problem-solving abilities on three novel tasks requiring non-trivial planning or search: Game of 24, Creative Writing, and Mini Crosswords. For instance, in Game of 24, while GPT-4 with chain-of-thought prompting only solved 4\% of tasks, our method achieved a success rate of 74\%. Code repo with all prompts: https://github.com/princeton-nlp/tree-of-thought-llm. Shunyu Yao 0006, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths 0001, Yuan Cao 0007, Karthik Narasimhan |
NeurIPS | 6 |
| 2022 | SGD-X: A Benchmark for Robust Generalization in Schema-Guided Dialogue SystemsabstractZero/few-shot transfer to unseen services is a critical challenge in task-oriented dialogue research. The Schema-Guided Dialogue (SGD) dataset introduced a paradigm for enabling models to support any service in zero-shot through schemas, which describe service APIs to models in natural language. We explore the robustness of dialogue systems to linguistic variations in schemas by designing SGD-X - a benchmark extending SGD with semantically similar yet stylistically diverse variants for every schema. We observe that two top state tracking models fail to generalize well across schema variants, measured by joint goal accuracy and a novel metric for measuring schema sensitivity. Additionally, we present a simple model-agnostic data augmentation method to improve schema robustness. Harrison Lee 0001, Abhinav Rastogi, Yuan Cao 0007 |
AAAI | 4 |
| 2022 | Multilingual Mix: Example Interpolation Improves Multilingual Neural Machine TranslationabstractMultilingual neural machine translation models are trained to maximize the likelihood of a mix of examples drawn from multiple language pairs.The dominant inductive bias applied to these models is a shared vocabulary and a shared set of parameters across languages; the inputs and labels corresponding to examples drawn from different language pairs might still reside in distinct subspaces.In this paper, we introduce multilingual crossover encoder-decoder (mXEncDec) to fuse language pairs at an instance level.Our approach interpolates instances from different language pairs into joint 'crossover examples' in order to encourage sharing input and output spaces across languages.To ensure better fusion of examples in multilingual settings, we propose several techniques to improve example interpolation across dissimilar languages under heavy data imbalance.Experiments on a large-scale WMT multilingual dataset demonstrate that our approach significantly improves quality on English-to-Many, Many-to-English and zero-shot translation tasks (from +0.5 BLEU up to +5.5 BLEU points).Results on code-switching sets demonstrate the capability of our approach to improve model generalization to out-of-distribution multilingual examples.We also conduct qualitative and quantitative representation comparisons to analyze the advantages of our approach at the representation level. Yong Cheng 0003, Ankur Bapna, Orhan Firat, Yuan Cao 0007, Pidong Wang, Wolfgang Macherey |
ACL (1) | 4 |
| 2022 | SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, Yuan Cao 0007 |
ICLR | 6 |
| 2022 | Show, Don't Tell: Demonstrations Outperform Descriptions for Schema-Guided Task-Oriented DialogueabstractRaghav Gupta, Harrison Lee, Jeffrey Zhao, Yuan Cao, Abhinav Rastogi, Yonghui Wu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Harrison Lee 0001, Jeffrey Zhao, Yuan Cao 0007, Abhinav Rastogi |
NAACL-HLT | 4 |
| 2022 | Unsupervised Slot Schema Induction for Task-oriented DialogabstractDian Yu, Mingqiu Wang, Yuan Cao, Izhak Shafran, Laurent Shafey, Hagen Soltau. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Mingqiu Wang, Yuan Cao 0007, Izhak Shafran, Laurent El Shafey, Hagen Soltau |
NAACL-HLT | 3 |
| 2021 | Effective Sequence-to-Sequence Dialogue State TrackingabstractSequence-to-sequence models have been applied to a wide variety of NLP tasks, but how to properly use them for dialogue state tracking has not been systematically investigated.In this paper, we study this problem from the perspectives of pre-training objectives as well as the formats of context representations.We demonstrate that the choice of pre-training objective makes a significant difference to the state tracking quality.In particular, we find that masked span prediction is more effective than auto-regressive language modeling.We also explore using Pegasus, a span prediction-based pre-training objective for text summarization, for the state tracking model.We found that pre-training for the seemingly distant summarization task works surprisingly well for dialogue state tracking.In addition, we found that while recurrent state context representation works also reasonably well, the model may have a hard time recovering from earlier mistakes.We conducted experiments on the MultiWOZ 2.1-2.4,WOZ 2.0, and DSTC2 datasets with consistent observations. Jeffrey Zhao, Mahdis Mahdieh, Yuan Cao 0007 |
EMNLP (1) | 4 |
| 2021 | Echo State Speech RecognitionabstractWe propose automatic speech recognition (ASR) models inspired by echo state network (ESN) [1], in which a subset of recurrent neural networks (RNN) layers in the models are randomly initialized and untrained. Our study focuses on RNN-T and Conformer models, and we show that model quality does not drop even when the decoder is fully randomized. Furthermore, such models can be trained more efficiently as the decoders do not require to be updated. By contrast, randomizing encoders hurts model quality, indicating that optimizing encoders and learn proper representations for acoustic inputs are more vital for speech recognition. Overall, we challenge the common practice of training ASR models for all components, and demonstrate that ESN-based models can perform equally well but enable more efficient training and storage than fully-trainable counterparts. Harsh Shrivastava 0001, Ankush Garg, Yuan Cao 0007, Yu Zhang 0033, Tara N. Sainath |
ICASSP | 3 |
| 2021 | Gradient Vaccine: Investigating and Improving Multi-task Optimization in Massively Multilingual Models
Yulia Tsvetkov, Orhan Firat, Yuan Cao 0007 |
ICLR | 4 |
| 2020 | Leveraging Monolingual Data with Self-Supervision for Multilingual Neural Machine TranslationabstractAditya Siddhant, Ankur Bapna, Yuan Cao, Orhan Firat, Mia Chen, Sneha Kudugunta, Naveen Arivazhagan, Yonghui Wu. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Aditya Siddhant, Ankur Bapna, Yuan Cao 0007, Orhan Firat, Mia Xu Chen, Sneha Reddy Kudugunta, Naveen Arivazhagan |
ACL | 3 |
| 2020 | Generating Diverse and Natural Text-to-Speech Samples Using a Quantized Fine-Grained VAE and Autoregressive Prosody PriorabstractRecent neural text-to-speech (TTS) models with fine-grained latent features enable precise control of the prosody of synthesized speech. Such models typically incorporate a fine-grained variational autoencoder (VAE) structure, extracting latent features at each input token (e.g., phonemes). However, generating samples with the standard VAE prior often results in unnatural and discontinuous speech, with dramatic prosodic variation between tokens. This paper proposes a sequential prior in a discrete latent space which can generate more naturally sounding samples. This is accomplished by discretizing the latent features using vector quantization (VQ), and separately training an autoregressive (AR) prior model over the result. We evaluate the approach using listening tests, objective metrics of automatic speech recognition (ASR) performance, and measurements of prosody attributes. Experimental results show that the proposed model significantly improves the naturalness in random sample generation. Furthermore, initial experiments demonstrate that randomly sampling from the proposed model can be used as data augmentation to improve the ASR performance. Guangzhi Sun, Yu Zhang 0033, Ron J. Weiss, Yuan Cao 0007, Heiga Zen, Andrew Rosenberg, Bhuvana Ramabhadran |
ICASSP | 4 |
| 2020 | Fully-Hierarchical Fine-Grained Prosody Modeling For Interpretable Speech SynthesisabstractThis paper proposes a hierarchical, fine-grained and interpretable latent variable model for prosody based on the Tacotron 2 text-to-speech model. It achieves multi-resolution modeling of prosody by conditioning finer level representations on coarser level ones. Additionally, it imposes hierarchical conditioning across all latent dimensions using a conditional variational auto-encoder (VAE) with an auto-regressive structure. Evaluation of reconstruction performance illustrates that the new structure does not degrade the model while allowing better interpretability. Interpretations of prosody attributes are provided together with the comparison between word-level and phone-level prosody representations. Moreover, both qualitative and quantitative evaluations are used to demonstrate the improvement in the disentanglement of the latent dimensions. Guangzhi Sun, Yu Zhang 0033, Ron J. Weiss, Yuan Cao 0007, Heiga Zen |
ICASSP | 4 |
| 2019 | Leveraging Weakly Supervised Data to Improve End-to-end Speech-to-text TranslationabstractEnd-to-end Speech Translation (ST) models have many potential advantages when compared to the cascade of Automatic Speech Recognition (ASR) and text Machine Translation (MT) models, including lowered inference latency and the avoidance of error compounding. However, the quality of end-to-end ST is often limited by a paucity of training data, since it is difficult to collect large parallel corpora of speech and translated transcript pairs. Previous studies have proposed the use of pre-trained components and multi-task learning in order to benefit from weakly supervised training data, such as speech-to-transcript or text-to-foreign-text pairs. In this paper, we demonstrate that using pre-trained MT or text-to-speech (TTS) synthesis models to convert weakly supervised data into speech-to-translation pairs for ST training can be more effective than multi-task learning. Furthermore, we demonstrate that a high quality end-to-end ST model can be trained using only weakly supervised datasets, and that synthetic data sourced from unlabeled monolingual text or speech can be used to improve performance. Finally, we discuss methods for avoiding overfitting to synthetic speech with a quantitative ablation study. Ye Jia, Melvin Johnson, Wolfgang Macherey, Ron J. Weiss, Yuan Cao 0007, Chung-Cheng Chiu, Naveen Ari, Stella Laurenzo |
ICASSP | 5 |
| 2019 | Hierarchical Generative Modeling for Controllable Speech Synthesis
Wei-Ning Hsu, Yu Zhang 0033, Ron J. Weiss, Heiga Zen, Yuxuan Wang 0002, Yuan Cao 0007, Ye Jia, Jonathan Shen, Patrick Nguyen, Ruoming Pang |
ICLR (Poster) | 7 |
| 2019 | Gmail Smart Compose: Real-Time Assisted WritingabstractIn this paper, we present Smart Compose, a novel system for generating interactive, real-time suggestions in Gmail that assists users in writing mails by reducing repetitive typing. In the design and deployment of such a large-scale and complicated system, we faced several challenges including model selection, performance evaluation, serving and other practical issues. At the core of Smart Compose is a large-scale neural language model. We leveraged state-of-the-art machine learning techniques for language model training which enabled high-quality suggestion prediction, and constructed novel serving infrastructure for high-throughput and real-time inference. Experimental results show the effectiveness of our proposed system design and deployment approach. This system is currently being served in Gmail. Mia Xu Chen, Benjamin N. Lee, Gagan Bansal, Yuan Cao 0007, Shuyuan Zhang 0002, Justin Lu, Jackie Tsay, Andrew M. Dai, Timothy Sohn |
KDD | 4 |
| 2018 | Training Deeper Neural Machine Translation Models with Transparent AttentionabstractWhile current state-of-the-art NMT models, such as RNN seq2seq and Transformers, possess a large number of parameters, they are still shallow in comparison to convolutional models used for both text and vision applications.In this work we attempt to train significantly (2-3x) deeper Transformer and Bi-RNN encoders for machine translation.We propose a simple modification to the attention mechanism that eases the optimization of deeper models, and results in consistent gains of 0.7-1.1 BLEU on the benchmark WMT'14 English-German and WMT'15 Czech-English tasks for both architectures. Ankur Bapna, Mia Xu Chen, Orhan Firat, Yuan Cao 0007 |
EMNLP | 4 |
| 2014 | Online Learning in Tensor SpaceabstractWe propose an online learning algorithm based on tensor-space models. A tensor-space model represents data in a compact way, and via rank-1 approximation the weight tensor can be made highly struc-tured, resulting in a significantly smaller number of free parameters to be estimated than in comparable vector-space models. This regularizes the model complexity and makes the tensor model highly effective in situations where a large feature set is de-fined but very limited resources are avail-able for training. We apply with the pro-posed algorithm to a parsing task, and show that even with very little training data the learning algorithm based on a ten-sor model performs well, and gives signif-icantly better results than standard learn-ing algorithms based on traditional vector-space models. 1 Yuan Cao 0007, Sanjeev Khudanpur |
ACL (1) | 1 |
| 2012 | Semi-supervised discriminative language modeling for Turkish ASRabstractWe present our work on semi-supervised learning of discriminative language models where the negative examples for sentences in a text corpus are generated using confusion models for Turkish at various granularities, specifically, word, sub-word, syllable and phone levels. We experiment with different language models and various sampling strategies to select competing hypotheses for training with a variant of the perceptron algorithm. We find that morph-based confusion models with a sample selection strategy aiming to match the error distribution of the baseline ASR system gives the best performance. We also observe that substituting half of the supervised training examples with those obtained in a semi-supervised manner gives similar results. Arda Çelebi, Hasim Sak, Erinç Dikici, Murat Saraclar, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Damianos Karakos, Sanjeev Khudanpur, Brian Roark, Kenji Sagae, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
ICASSP | 16 |
| 2012 | Hallucinated n-best lists for discriminative language modelingabstractThis paper investigates semi-supervised methods for discriminative language modeling, whereby n-best lists are “hallucinated” for given reference text and are then used for training n-gram language models using the perceptron algorithm. We perform controlled experiments on a very strong baseline English CTS system, comparing three methods for simulating ASR output, and compare the results with training with “real” n-best list output from the baseline recognizer. We find that methods based on extracting phrasal cohorts - similar to methods from machine translation for extracting phrase tables - yielded the largest gains of our three methods, achieving over half of the WER reduction of the fully supervised methods. Kenji Sagae, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Damianos Karakos, Sanjeev Khudanpur, Brian Roark, Murat Saraclar, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
ICASSP | 13 |
| 2012 | Continuous space discriminative language modelingabstractDiscriminative language modeling is a structured classification problem. Log-linear models have been previously used to address this problem. In this paper, the standard dot-product feature representation used in log-linear models is replaced by a non-linear function parameterized by a neural network. Embeddings are learned for each word and features are extracted automatically through the use of convolutional layers. Experimental results show that as a stand-alone model the continuous space model yields significantly lower word error rate (1% absolute), while having a much more compact parameterization (60%-90% smaller). If the baseline scores are combined, our approach performs equally well. Puyang Xu, Sanjeev Khudanpur, Maider Lehr, Emily Tucker Prud'hommeaux, Nathan Glenn, Damianos Karakos, Brian Roark, Kenji Sagae, Murat Saraclar, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
ICASSP | 13 |
| 2012 | Deriving conversation-based features from unlabeled speech for discriminative language modeling
Damianos Karakos, Brian Roark, Izhak Shafran, Kenji Sagae, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Sanjeev Khudanpur, Murat Saraclar, Dan Bikel, Mark Dredze, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
INTERSPEECH | 14 |