Yingbo Zhou 0002

dblp:72/8614-2 · DBLP profile ↗
← Back
43ranked-venue papers
5as first author
24since 2021 · last 2025
0000-0001-6034-9667ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 2 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 1 since 2021Security and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
abstract
Embedding models play a crucial role in a variety of downstream tasks, including semantic similarity, information retrieval, and clustering. While there has been a surge of interest in developing universal text embedding models that generalize across tasks (e.g., MTEB), progress in learning universal multimodal embedding models has been comparatively slow, despite their importance and practical applications. In this work, we explore the potential of building universal multimodal embeddings capable of handling a broad range of downstream tasks. Our contributions are twofold: (1) we propose MMEB (Massive Multimodal Embedding Benchmark), which covers four meta-tasks (classification, visual question answering, multimodal retrieval, and visual grounding) and 36 datasets, including 20 training datasets and 16 evaluation datasets spanning both in-distribution and out-of-distribution tasks, and (2) VLM2Vec (Vision-Language Model → Vector), a contrastive training framework that transforms any vision-language model into an embedding model through contrastive training on MMEB. Unlike previous models such as CLIP and BLIP, which encode text and images independently without task-specific guidance, VLM2Vec can process any combination of images and text while incorporating task instructions to generate a fixed-dimensional vector. We develop a series of VLM2Vec models based on state-of-the-art VLMs, including Phi-3.5-V, LLaVA-1.6, and Qwen2-VL, and evaluate them on MMEB’s benchmark. With LoRA tuning, VLM2Vec achieves a 10% to 20% improvement over existing multimodal embedding models on MMEB’s evaluation sets. Our findings reveal that VLMs are surprisingly strong embedding models.
Ziyan Jiang, Xinyi Yang 0002, Semih Yavuz, Yingbo Zhou 0002, Wenhu Chen
ICLR5
2025 Diversity Empowers Intelligence: Integrating Expertise of Software Engineering Agents
abstract
Large language model (LLM) agents have shown great potential in solving real-world software engineering (SWE) problems. The most advanced open-source SWE agent can resolve over 27% of real GitHub issues in SWE-Bench Lite. However, these sophisticated agent frameworks exhibit varying strengths, excelling in certain tasks while underperforming in others. To fully harness the diversity of these agents, we propose DEI (Diversity Empowered Intelligence), a framework that leverages their unique expertise. DEI functions as a meta-module atop existing SWE agent frameworks, managing agent collectives for enhanced problem-solving. Experimental results show that a DEI-guided committee of agents is able to surpass the best individual agent's performance by a large margin. For instance, a group of open-source SWE agents, with a maximum individual resolve rate of 27.3% on SWE-Bench Lite, can achieve a 34.3% resolve rate with DEI, making a 25% improvement and beating most closed-source solutions. Our best-performing group excels with a 55% resolve rate, securing the highest ranking on SWE-Bench Lite. Our findings contribute to the growing body of research on collaborative AI systems and their potential to solve complex software engineering challenges.
Kexun Zhang, Weiran Yao, Zuxin Liu, Yihao Feng, Zhiwei Liu 0001, Rithesh R. N., Tian Lan 0006, Lei Li 0005, Renze Lou, Jiacheng Xu 0001, Bo Pang 0004, Yingbo Zhou 0002, Shelby Heinecke, Silvio Savarese, Huan Wang 0016, Caiming Xiong
ICLR12
2025 CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models
abstract
Jierui Li, Hung Le, Yingbo Zhou, Caiming Xiong, Silvio Savarese, Doyen Sahoo. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Jierui Li, Hung Le 0003, Yingbo Zhou 0002, Caiming Xiong, Silvio Savarese, Doyen Sahoo
NAACL (Long Papers)3
2025 Breaking the Batch Barrier (B3) of Contrastive Learning via Smart Batch Mining
abstract
Contrastive learning (CL) is a prevalent technique for training embedding models, which pulls semantically similar examples (positives) closer in the representation space while pushing dissimilar ones (negatives) further apart. A key source of negatives are "in-batch" examples, i.e., positives from other examples in the batch. Effectiveness of such models is hence strongly influenced by the size and quality of training batches. In this work, we propose *Breaking the Batch Barrier* (B3), a novel batch construction strategy designed to curate high-quality batches for CL. Our approach begins by using a pretrained teacher embedding model to rank all examples in the dataset, from which a sparse similarity graph is constructed. A community detection algorithm is then applied to this graph to identify clusters of examples that serve as strong negatives for one another. The clusters are then used to construct batches that are rich in in-batch negatives. Empirical results on the MMEB multimodal embedding benchmark (36 tasks) demonstrate that our method sets a new state of the art, outperforming previous best methods by +1.3 and +2.9 points at the 7B and 2B model scales, respectively. Notably, models trained with B3 surpass existing state-of-the-art results even with a batch size as small as 64, which is 4–16× smaller than that required by other methods. Moreover, experiments show that B3 generalizes well across domains and tasks, maintaining strong performance even when trained with considerably weaker teachers.
Raghuveer Thirukovalluru, Ye Liu 0006, Karthikeyan K, Mingyi Su, Ping Nie, Semih Yavuz, Yingbo Zhou 0002, Wenhu Chen, Bhuwan Dhingra
NeurIPS8
2024 FOLIO: Natural Language Reasoning with First-Order Logic
abstract
Simeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi, Martin Riddell, Wenfei Zhou, James Coady, David Peng, Yujie Qiao, Luke Benson, Lucy Sun, Alexander Wardle-Solano, Hannah Szabó, Ekaterina Zubova, Matthew Burtell, Jonathan Fan, Yixin Liu, Brian Wong, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu, Rui Zhang, Alexander Fabbri, Wojciech Maciej Kryscinski, Semih Yavuz, Ye Liu, Xi Victoria Lin, Shafiq Joty, Yingbo Zhou, Caiming Xiong, Rex Ying, Arman Cohan, Dragomir Radev. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Simeng Han, Hailey Schoelkopf, Yilun Zhao 0001, Zhenting Qi, Martin Riddell, Wenfei Zhou, James Coady, David Peng, Yujie Qiao, Luke Benson, Lucy Sun, Alexander Wardle-Solano, Hannah Szabó, Ekaterina Zubova, Matthew Burtell, Jonathan Fan 0001, Yixin Liu 0003, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu 0009, Rui Zhang 0037, Alexander R. Fabbri, Wojciech Kryscinski, Semih Yavuz, Ye Liu 0006, Xi Victoria Lin, Shafiq R. Joty, Yingbo Zhou 0002, Caiming Xiong, Rex Ying, Arman Cohan, Dragomir R. Radev
EMNLP31
2024 Unlocking Anticipatory Text Generation: A Constrained Approach for Large Language Models Decoding
abstract
Large Language Models (LLMs) have demonstrated a powerful ability for text generation.However, achieving optimal results with a given prompt or instruction can be challenging, especially for billion-sized models.Additionally, undesired behaviors such as toxicity or hallucinations can manifest.While much larger models (e.g., ChatGPT) may demonstrate strength in mitigating these issues, there is still no guarantee of complete prevention.In this work, we propose formalizing text generation as a future-constrained generation problem to minimize undesirable behaviors and enforce faithfulness to instructions.The estimation of future constraint satisfaction, accomplished using LLMs, guides the text generation process.Our extensive experiments demonstrate the effectiveness of the proposed approach across three distinct text generation tasks: keywordconstrained generation (Lin et al., 2020), toxicity reduction (Gehman et al., 2020), and factual correctness in question-answering (Gao et al., 2023). 1
Lifu Tu, Semih Yavuz, Jin Qu, Jiacheng Xu 0001, Caiming Xiong, Yingbo Zhou 0002
EMNLP7
2024 ARM: Alignment with Residual Energy-Based Model
abstract
Bo Pang, Caiming Xiong, Yingbo Zhou. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Bo Pang 0004, Caiming Xiong, Yingbo Zhou 0002
NAACL-HLT3
2024 INDICT: Code Generation with Internal Dialogues of Critiques for Both Security and Helpfulness
abstract
Large language models (LLMs) for code are typically trained to align with natural language instructions to closely follow their intentions and requirements. However, in many practical scenarios, it becomes increasingly challenging for these models to navigate the intricate boundary between helpfulness and safety, especially against highly complex yet potentially malicious instructions. In this work, we introduce INDICT: a new framework that empowers LLMs with Internal Dialogues of Critiques for both safety and helpfulness guidance. The internal dialogue is a dual cooperative system between a safety-driven critic and a helpfulness-driven critic. Each critic provides analysis against the given task and corresponding generated response, equipped with external knowledge queried through relevant code snippets and tools like web search and code interpreter. We engage the dual critic system in both code generation stage as well as code execution stage, providing preemptive and post-hoc guidance respectively to LLMs. We evaluated INDICT on 8 diverse tasks across 8 programming languages from 5 benchmarks, using LLMs from 7B to 70B parameters. We observed that our approach can provide an advanced level of critiques of both safety and helpfulness analysis, significantly improving the quality of output codes (+10% absolute improvements in all models).
Hung Le 0003, Doyen Sahoo, Yingbo Zhou 0002, Caiming Xiong, Silvio Savarese
NeurIPS3
2024 L2CEval: Evaluating Language-to-Code Generation Capabilities of Large Language Models
abstract
Abstract Recently, large language models (LLMs), especially those that are pretrained on code, have demonstrated strong capabilities in generating programs from natural language inputs. Despite promising results, there is a notable lack of a comprehensive evaluation of these models’ language-to-code generation capabilities. Existing studies often focus on specific tasks, model architectures, or learning paradigms, leading to a fragmented understanding of the overall landscape. In this work, we present L2CEval, a systematic evaluation of the language-to-code generation capabilities of LLMs on 7 tasks across the domain spectrum of semantic parsing, math reasoning, and Python programming, analyzing the factors that potentially affect their performance, such as model size, pretraining data, instruction tuning, and different prompting methods. In addition, we assess confidence calibration, and conduct human evaluations to identify typical failures across different tasks and models. L2CEval offers a comprehensive understanding of the capabilities and limitations of LLMs in language-to-code generation. We release the evaluation framework1 and all model outputs, hoping to lay the groundwork for further future research. All future evaluations (e.g., LLaMA-3, StarCoder2, etc) will be updated on the project website: https://l2c-eval.github.io/.
Ansong Ni, Yilun Zhao 0001, Martin Riddell, Troy Feng, Stephen Yin, Ye Liu 0006, Semih Yavuz, Caiming Xiong, Shafiq R. Joty, Yingbo Zhou 0002, Dragomir R. Radev, Arman Cohan
Trans. Assoc. Comput. Linguistics12
2023 Best-k Search Algorithm for Neural Text Generation
abstract
Modern natural language generation paradigms require a decoding strategy to obtain quality sequences out of the model.Beam search yields high-quality but low diversity outputs; stochastic approaches suffer from high variance and sometimes low quality.In this work, we propose a deterministic search algorithm balancing both quality and diversity.We first investigate the vanilla best-first search (BFS) algorithm and then propose the best-k search algorithm.Inspired by BFS, we greedily expand the top k nodes, instead of the first node, to boost efficiency and diversity.Upweighting recently discovered nodes accompanied by heap pruning ensures the completeness of the search procedure.Experiments on four NLG tasks show that best-k search yields more diverse and natural outputs compared to strong baselines, while our approach maintains high text quality.The proposed algorithm is parameter-free, lightweight, efficient, and easy-to-use.1
Jiacheng Xu 0001, Caiming Xiong, Silvio Savarese, Yingbo Zhou 0002
ACL (1)4
2023 CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis
Erik Nijkamp, Bo Pang 0004, Hiroaki Hayashi, Lifu Tu, Huan Wang 0016, Yingbo Zhou 0002, Silvio Savarese, Caiming Xiong
ICLR6
2023 UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild
abstract
Achieving machine autonomy and human control often represent divergent objectives in the design of interactive AI systems. Visual generative foundation models such as Stable Diffusion show promise in navigating these goals, especially when prompted with arbitrary languages. However, they often fall short in generating images with spatial, structural, or geometric controls. The integration of such controls, which can accommodate various visual conditions in a single unified model, remains an unaddressed challenge. In response, we introduce UniControl, a new generative foundation model that consolidates a wide array of controllable condition-to-image (C2I) tasks within a singular framework, while still allowing for arbitrary language prompts. UniControl enables pixel-level-precise image generation, where visual conditions primarily influence the generated structures and language prompts guide the style and context. To equip UniControl with the capacity to handle diverse visual conditions, we augment pretrained text-to-image diffusion models and introduce a task-aware HyperNet to modulate the diffusion models, enabling the adaptation to different C2I tasks simultaneously. Trained on nine unique C2I tasks, UniControl demonstrates impressive zero-shot generation abilities with unseen visual conditions. Experimental results show that UniControl often surpasses the performance of single-task-controlled methods of comparable model sizes. This control versatility positions UniControl as a significant advancement in the realm of controllable visual generation.
Can Qin, Shu Zhang 0007, Ning Yu 0006, Yihao Feng, Xinyi Yang 0002, Yingbo Zhou 0002, Huan Wang 0016, Juan Carlos Niebles, Caiming Xiong, Silvio Savarese, Stefano Ermon, Yun Fu 0001, Ran Xu 0001
NeurIPS6
2023 Merlion: End-to-End Machine Learning for Time Series
abstract
We introduce Merlion, an open-source machine learning library for time series. It features a unified interface for many commonly used models and datasets for forecasting and anomaly detection on both univariate and multivariate time series, along with standard pre/post-processing layers. It has several modules to improve ease-of-use, including a no-code visual dashboard, anomaly score calibration to improve interpetability, AutoML for hyperparameter tuning and model selection, and model ensembling. Merlion also provides an evaluation framework that simulates the live deployment of a model in production, and a distributed computing backend to run time series models at industrial scale. This library aims to provide engineers and researchers a one-stop solution to rapidly develop models for their specific time series needs and benchmark them across multiple datasets.
Aadyot Bhatnagar, Paul Kassianik, Tian Lan 0006, Wenzhuo Yang, Rowan Cassius, Doyen Sahoo, Devansh Arpit, Sri Subramanian, Gerald Woo, Amrita Saha, Arun Kumar Jagota, Gokulakrishnan Gopalakrishnan, K. C. Krithika, Sukumar Maddineni, Dae-ki Cho, Bo Zong, Yingbo Zhou 0002, Caiming Xiong, Silvio Savarese, Steven C. H. Hoi, Huan Wang 0016
J. Mach. Learn. Res.19
2022 Modeling Multi-hop Question Answering as Single Sequence Prediction
abstract
Fusion-in-decoder (FID) (Izacard and Grave, 2021) is a generative question answering (QA) model that leverages passage retrieval with a pre-trained transformer and pushed the state of the art on single-hop QA.However, the complexity of multi-hop QA hinders the effectiveness of the generative QA approach.In this work, we propose a simple generative approach (PATHFID) that extends the task beyond just answer generation by explicitly modeling the reasoning process to resolve the answer for multihop questions.By linearizing the hierarchical reasoning path of supporting passages, their key sentences, and finally the factoid answer, we cast the problem as a single sequence prediction task.To facilitate complex reasoning with multiple clues, we further extend the unified flat representation of multiple input documents by encoding cross-passage interactions.Our extensive experiments demonstrate that PATHFID leads to strong performance gains on two multihop QA datasets: HotpotQA and IIRC.Besides the performance gains, PATHFID is more interpretable, which in turn yields answers that are more faithfully grounded to the supporting passages and facts compared to the baseline FID model.
Semih Yavuz, Kazuma Hashimoto, Yingbo Zhou 0002, Nitish Shirish Keskar, Caiming Xiong
ACL (1)3
2022 RNG-KBQA: Generation Augmented Iterative Ranking for Knowledge Base Question Answering
abstract
Existing KBQA approaches, despite achieving strong performance on i.i.d.test data, often struggle in generalizing to questions involving unseen KB schema items.Prior rankingbased approaches have shown some success in generalization, but suffer from the coverage issue.We present RnG-KBQA, a Rank-and-Generate approach for KBQA, which remedies the coverage issue with a generation model while preserving a strong generalization capability.Our approach first uses a contrastive ranker to rank a set of candidate logical forms obtained by searching over the knowledge graph.It then introduces a tailored generation model conditioned on the question and the top-ranked candidates to compose the final logical form.We achieve new state-ofthe-art results on GRAILQA and WEBQSP datasets.In particular, our method surpasses the prior state-of-the-art by a large margin on the GRAILQA leaderboard.In addition, RnG-KBQA outperforms all prior approaches on the popular WEBQSP benchmark, even including the ones that use the oracle entity linking.The experimental results demonstrate the effectiveness of the interplay between ranking and generation, which leads to the superior performance of our proposed approach across all settings with especially strong improvements in zero-shot generalization. 1 * Work done during internship at Salesforce Research. 1 Code available at https://github.com/salesforce/rng-kbqa.
Xi Ye 0003, Semih Yavuz, Kazuma Hashimoto, Yingbo Zhou 0002, Caiming Xiong
ACL (1)4
2022 Uni-Parser: Unified Semantic Parser for Question Answering on Knowledge Base and Database
abstract
Parsing natural language questions into executable logical forms is a useful and interpretable way to perform question answering on structured data such as knowledge bases (KB) or databases (DB).However, existing approaches on semantic parsing cannot adapt to both modalities, as they suffer from the exponential growth of the logical form candidates and can hardly generalize to unseen data.In this work, we propose Uni-Parser, a unified semantic parser for question answering (QA) on both KB and DB.We introduce the primitive (relation and entity in KB, and table name, column name and cell value in DB) as an essential element in our framework.The number of primitives grows linearly with the number of retrieved relations in KB and DB, preventing us from dealing with exponential logic form candidates.We leverage the generator to predict final logical forms by altering and composing topranked primitives with different operations (e.g.select, where, count).With sufficiently pruned search space by a contrastive primitive ranker, the generator is empowered to capture the composition of primitives enhancing its generalization ability.We achieve competitive results on multiple KB and DB QA benchmarks more efficiently, especially in the compositional and zero-shot settings.
Ye Liu 0006, Semih Yavuz, Dragomir R. Radev, Caiming Xiong, Yingbo Zhou 0002
EMNLP6
2022 Efficient and Differentiable Conformal Prediction with General Function Classes
Yu Bai 0017, Song Mei, Huan Wang 0016, Yingbo Zhou 0002, Caiming Xiong
ICLR4
2022 Ensemble of Averages: Improving Model Selection and Boosting Performance in Domain Generalization
abstract
In Domain Generalization (DG) settings, models trained independently on a given set of training domains have notoriously chaotic performance on distribution shifted test domains, and stochasticity in optimization (e.g. seed) plays a big role. This makes deep learning models unreliable in real world settings. We first show that this chaotic behavior exists even along the training optimization trajectory of a single model, and propose a simple model averaging protocol that both significantly boosts domain generalization and diminishes the impact of stochasticity by improving the rank correlation between the in-domain validation accuracy and out-domain test accuracy, which is crucial for reliable early stopping. Taking advantage of our observation, we show that instead of ensembling unaveraged models (that is typical in practice), ensembling moving average models (EoA) from independent runs further boosts performance. We theoretically explain the boost in performance of ensembling and model averaging by adapting the well known Bias-Variance trade-off to the domain generalization setting. On the DomainBed benchmark, when using a pre-trained ResNet-50, this ensemble of averages achieves an average of $68.0\%$, beating vanilla ERM (w/o averaging/ensembling) by $\sim 4\%$, and when using a pre-trained RegNetY-16GF, achieves an average of $76.6\%$, beating vanilla ERM by $\sim 6\%$.
Devansh Arpit, Huan Wang 0016, Yingbo Zhou 0002, Caiming Xiong
NeurIPS3
2021 WOAD: Weakly Supervised Online Action Detection in Untrimmed Videos
abstract
Online action detection in untrimmed videos aims to identify an action as it happens, which makes it very important for real-time applications. Previous methods rely on tedious annotations of temporal action boundaries for training, which hinders the scalability of online action detection systems. We propose WOAD, a weakly supervised framework that can be trained using only video-class labels. WOAD contains two jointly-trained modules, i.e., temporal proposal generator (TPG) and online action recognizer (OAR). Supervised by video-class labels, TPG works offline and targets at accurately mining pseudo frame-level labels for OAR. With the supervisory signals from TPG, OAR learns to conduct action detection in an online fashion. Experimental results on THUMOS’14, ActivityNet1.2 and ActivityNet1.3 show that our weakly-supervised method largely outperforms weakly-supervised baselines and achieves comparable performance to the previous strongly-supervised methods. Beyond that, WOAD is flexible to leverage strong supervision when it is available. When strongly supervised, our method obtains the state-of-the-art results in the tasks of both online per-frame action recognition and online detection of action start.
Mingfei Gao, Yingbo Zhou 0002, Ran Xu 0001, Richard Socher, Caiming Xiong
CVPR2
2021 Unsupervised Paraphrasing with Pretrained Language Models
abstract
Paraphrase generation has benefited extensively from recent progress in the designing of training objectives and model architectures.However, previous explorations have largely focused on supervised methods, which require a large amount of labeled data that is costly to collect.To address this drawback, we adopt a transfer learning approach and propose a training pipeline that enables pre-trained language models to generate high-quality paraphrases in an unsupervised setting.Our recipe consists of task-adaptation, self-supervision, and a novel decoding algorithm named Dynamic Blocking (DB).To enforce a surface form dissimilar from the input, whenever the language model emits a token contained in the source sequence, DB prevents the model from outputting the subsequent source token for the next generation step.We show with automatic and human evaluations that our approach achieves state-of-the-art performance on both the Quora Question Pair (QQP) and the ParaNMT datasets and is robust to domain shift between the two datasets of distinct distributions.We also demonstrate that our model transfers to paraphrasing in other languages without any additional finetuning.
Semih Yavuz, Yingbo Zhou 0002, Nitish Shirish Keskar, Huan Wang 0016, Caiming Xiong
EMNLP (1)3
2021 Representation Learning for Sequence Data with Deep Autoencoding Predictive Components
Junwen Bai, Yingbo Zhou 0002, Caiming Xiong
ICLR3
2021 CoCo: Controllable Counterfactuals for Evaluating Dialogue State Trackers
Semih Yavuz, Kazuma Hashimoto, Jia Li 0015, Nazneen Fatema Rajani, Xifeng Yan, Yingbo Zhou 0002, Caiming Xiong
ICLR8
2021 Focused Attention Improves Document-Grounded Generation
abstract
Shrimai Prabhumoye, Kazuma Hashimoto, Yingbo Zhou, Alan W Black, Ruslan Salakhutdinov. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Shrimai Prabhumoye, Kazuma Hashimoto, Yingbo Zhou 0002, Alan W. Black, Ruslan Salakhutdinov
NAACL-HLT3
2021 QueryBlazer: Efficient Query Autocompletion Framework
abstract
Query autocompletion is an essential feature in search engines that predicts and suggests query completions to a user's incomplete prefix input, a critical feature to enhance the user experience. While a generic lookup-based system can provide completions with great efficiency, it is unable to address prefixes not seen in the past. On the other hand, a generative system can complete unseen queries with superior accuracy but requires substantial computational overhead at runtime, making it costly for a large-scale system. Here, we present an efficient, fully-generative query autocompletion framework. Our framework employs an n-gram language model at a subword-level and exploits the n-gram model's inherent data structure to precompute completions prior to runtime. Evaluation results on public dataset show that our framework is not only as effective as previous systems with neural language models, but also reduces computational overhead at runtime, expediting the speed by more than two orders of magnitude. The goal of this work is to showcase a generative query completion system that is an attractive choice for large-scale deployments.
Young Mo Kang, Wenhao Liu 0003, Yingbo Zhou 0002
WSDM3
2020 An Investigation of Phone-Based Subword Units for End-to-End Speech Recognition
abstract
Phones and their context-dependent variants have been the standard modeling units for conventional speech recognition systems, while characters and subwords have demonstrated their effectiveness for end-to-end recognition systems.We investigate the use of phone-based subwords, in particular, byte pair encoder (BPE), as modeling units for end-to-end speech recognition.In addition, we also developed multi-level language model-based decoding algorithms based on a pronunciation dictionary.Besides the use of the lexicon, which is easily available, our system avoids the need of additional expert knowledge or processing steps from conventional systems.Experimental results show that phone-based BPEs tend to yield more accurate recognition systems than the character-based counterpart.In addition, further improvement can be obtained with a novel one-pass joint beam search decoder, which efficiently combines phone-and character-based BPE systems.For Switchboard, our phone-based BPE system achieves 6.8%/14.4% word error rate (WER) on the Switchboard/CallHome portion of the test set while joint decoding achieves 6.3%/13.3%WER.On Fisher + Switchboard, joint decoding leads to 4.9%/9.5% WER, setting new milestones for telephony speech recognition.
Guangsen Wang, Aadyot Bhatnagar, Yingbo Zhou 0002, Caiming Xiong, Richard Socher
INTERSPEECH4
2020 Online Structured Meta-learning
abstract
Learning quickly is of great importance for machine intelligence deployed in online platforms. With the capability of transferring knowledge from learned tasks, meta-learning has shown its effectiveness in online scenarios by continuously updating the model with the learned prior. However, current online meta-learning algorithms are limited to learn a globally-shared meta-learner, which may lead to sub-optimal results when the tasks contain heterogeneous information that are difficult to share. We overcome this limitation by proposing an online structured meta-learning (OSML) framework. Inspired by the knowledge organization of human and hierarchical feature representation, OSML explicitly disentangles the meta-learner as a meta-hierarchical graph with different knowledge blocks. When a new task is encountered, it constructs a meta-knowledge pathway by either utilizing the most relevant knowledge blocks or exploring new blocks. Through the meta-knowledge pathway, the model is able to quickly adapt to the new task. In addition, new knowledge is further incorporated into the selected blocks. Experiments on three datasets empirically demonstrate the effectiveness and interpretability of our proposed framework, not only under heterogeneous tasks but also under homogeneous settings.
Huaxiu Yao, Yingbo Zhou 0002, Mehrdad Mahdavi, Zhenhui Li, Richard Socher, Caiming Xiong
NeurIPS2
2019 Augmented Cyclic Adversarial Learning for Low Resource Domain Adaptation
Ehsan Hosseini-Asl, Yingbo Zhou 0002, Caiming Xiong, Richard Socher
ICLR (Poster)2
2019 Learn to Grow: A Continual Structure Learning Framework for Overcoming Catastrophic Forgetting
abstract
Addressing catastrophic forgetting is one of the key challenges in continual learning where machine learning systems are trained with sequential or streaming tasks. Despite recent remarkable progress in state-of-the-art deep learning, deep neural networks (DNNs) are still plagued with the catastrophic forgetting problem. This paper presents a conceptually simple yet general and effective framework for handling catastrophic forgetting in continual learning with DNNs. The proposed method consists of two components: a neural structure optimization component and a parameter learning and/or fine-tuning component. By separating the explicit neural structure learning and the parameter estimation, not only is the proposed method capable of evolving neural structures in an intuitively meaningful way, but also shows strong capabilities of alleviating catastrophic forgetting in experiments. Furthermore, the proposed method outperforms all other baselines on the permuted MNIST dataset, the split CIFAR100 dataset and the Visual Domain Decathlon dataset in continual learning setting.
Xilai Li, Yingbo Zhou 0002, Tianfu Wu 0001, Richard Socher, Caiming Xiong
ICML2
2018 End-to-End Dense Video Captioning With Masked Transformer
abstract
Dense video captioning aims to generate text descriptions for all events in an untrimmed video. This involves both detecting and describing events. Therefore, all previous methods on dense video captioning tackle this problem by building two models, i.e. an event proposal and a captioning model, for these two sub-problems. The models are either trained separately or in alternation. This prevents direct influence of the language description to the event proposal, which is important for generating accurate descriptions. To address this problem, we propose an end-to-end transformer model for dense video captioning. The encoder encodes the video into appropriate representations. The proposal decoder decodes from the encoding with different anchors to form video event proposals. The captioning decoder employs a masking network to restrict its attention to the proposal event over the encoding feature. This masking network converts the event proposal to a differentiable mask, which ensures the consistency between the proposal and captioning during training. In addition, our model employs a self-attention mechanism, which enables the use of efficient non-recurrent structure during encoding and leads to performance improvements. We demonstrate the effectiveness of this end-to-end model on ActivityNet Captions and YouCookII datasets, where we achieved 10.12 and 6.58 METEOR score, respectively.
Luowei Zhou, Yingbo Zhou 0002, Jason J. Corso, Richard Socher, Caiming Xiong
CVPR2
2018 Improving End-to-End Speech Recognition with Policy Learning
abstract
Connectionist temporal classification (CTC) is widely used for maximum likelihood learning in end-to-end speech recognition models. However, there is usually a disparity between the negative maximum likelihood and the performance metric used in speech recognition, e.g., word error rate (WER). This results in a mismatch between the objective function and metric during training. We show that the above problem can be mitigated by jointly training with maximum likelihood and policy gradient. In particular, with policy learning we are able to directly optimize on the (otherwise non-differentiable) performance metric. We show that joint training improves relative performance by 4% to 13% for our end-to-end model as compared to the same model learned through maximum likelihood. The model achieves 5.53% WER on Wall Street Journal dataset, and 5.42% and 14.70% on Librispeech test-clean and test-other set, respectively.
Yingbo Zhou 0002, Caiming Xiong, Richard Socher
ICASSP1
2018 A Multi-Discriminator CycleGAN for Unsupervised Non-Parallel Speech Domain Adaptation
abstract
Domain adaptation plays an important role for speech recognition models, in particular, for domains that have low resources. We propose a novel generative model based on cyclic-consistent generative adversarial network (CycleGAN) for unsupervised non-parallel speech domain adaptation. The proposed model employs multiple independent discriminators on the power spectrogram, each in charge of different frequency bands. As a result we have 1) better discriminators that focus on fine-grained details of the frequency features, and 2) a generator that is capable of generating more realistic domain-adapted spectrogram. We demonstrate the effectiveness of our method on speech recognition with gender adaptation, where the model only has access to supervised data from one gender during training, but is evaluated on the other at test time. Our model is able to achieve an average of $7.41\%$ on phoneme error rate, and $11.10\%$ word error rate relative performance improvement as compared to the baseline, on TIMIT and WSJ dataset, respectively. Qualitatively, our model also generates more natural sounding speech, when conditioned on data from the other domain.
Ehsan Hosseini-Asl, Yingbo Zhou 0002, Caiming Xiong, Richard Socher
INTERSPEECH2
2016 Normalization Propagation: A Parametric Technique for Removing Internal Covariate Shift in Deep Networks
abstract
While the authors of Batch Normalization (BN) identify and address an important problem involved in training deep networks– \textitInternal Covariate Shift– the current solution has certain drawbacks. For instance, BN depends on batch statistics for layerwise input normalization during training which makes the estimates of mean and standard deviation of input (distribution) to hidden layers inaccurate due to shifting parameter values (especially during initial training epochs). Another fundamental problem with BN is that it cannot be used with batch-size 1 during training. We address these drawbacks of BN by proposing a non-adaptive normalization technique for removing covariate shift, that we call \textitNormalization Propagation. Our approach does not depend on batch statistics, but rather uses a data-independent parametric estimate of mean and standard-deviation in every layer thus being computationally faster compared with BN. We exploit the observation that the pre-activation before Rectified Linear Units follow Gaussian distribution in deep networks, and that once the first and second order statistics of any given dataset are normalized, we can forward propagate this normalization without the need for recalculating the approximate statistics for hidden layers.
Devansh Arpit, Yingbo Zhou 0002, Bhargava Urala Kota, Venu Govindaraju
ICML2
2016 Why Regularized Auto-Encoders learn Sparse Representation?
abstract
Sparse distributed representation is the key to learning useful features in deep learning algorithms, because not only it is an efficient mode of data representation, but also – more importantly – it captures the generation process of most real world data. While a number of regularized auto-encoders (AE) enforce sparsity explicitly in their learned representation and others don’t, there has been little formal analysis on what encourages sparsity in these models in general. Our objective is to formally study this general problem for regularized auto-encoders. We provide sufficient conditions on both regularization and activation functions that encourage sparsity. We show that multiple popular models (de-noising and contractive auto encoders, e.g.) and activations (rectified linear and sigmoid, e.g.) satisfy these conditions; thus, our conditions help explain sparsity in their learned representation. Thus our theoretical and empirical analysis together shed light on the properties of regularization/activation that are conductive to sparsity and unify a number of existing auto-encoder models and activation functions under the same analytical framework.
Devansh Arpit, Yingbo Zhou 0002, Hung Q. Ngo 0001, Venu Govindaraju
ICML2
2015 Challenges in representation learning: A report on three machine learning contests
Ian J. Goodfellow, Dumitru Erhan, Pierre Luc Carrier, Aaron C. Courville, Mehdi Mirza, Benjamin Hamner, Will Cukierski, Yichuan Tang, Dave Thaler, Yingbo Zhou 0002, Chetan Ramaiah, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006, Dimitris Athanasakis, John Shawe-Taylor, Maxim Milakov, John Park, Radu Tudor Ionescu, Marius Popescu, Cristian Grozea, James Bergstra, Jingjing Xie, Lukasz Romaszko, Yoshua Bengio
Neural Networks11
2014 Parallel Feature Selection Inspired by Group Testing
Yingbo Zhou 0002, Utkarsh Porwal, Ce Zhang 0001, Hung Q. Ngo 0001, XuanLong Nguyen, Christopher Ré, Venu Govindaraju
NIPS1
2014 Shared features for multiple face-based biometrics
abstract
People often make instant judgments about the age, health, mood, personality and character of others based on their facial features. It is not clear from a cognitive aspect whether these different traits require different sets of features or a shared feature set. Till date, much of the computational face image analysis work such as face recognition, face-based deceit detection, age estimation, gender estimation, etc, have been developed on datasets and features specific only to the problem-at-hand. In this paper, we explore an approach for performing face image analysis using a shared set of features for different tasks. By performing unsupervised learning on a large collection of face images, we learn the parameters of a probabilistic generative face model, and by projecting a new face image into this probabilistic space, we obtain a set of face features not created for any specific face analysis tasks. We investigate the use of such shared features and successfully predict the level of attractiveness, whether or not a face is made-up, the facial expression, and the gender of a person, given any arbitrary, near-frontal face image.
Ifeoma Nwogu, Yingbo Zhou 0002
SMC2
2013 Challenges in Representation Learning: A Report on Three Machine Learning Contests
Ian J. Goodfellow, Dumitru Erhan, Pierre Luc Carrier, Aaron C. Courville, Mehdi Mirza, Benjamin Hamner, Will Cukierski, Yichuan Tang, Dave Thaler, Yingbo Zhou 0002, Chetan Ramaiah, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006, Dimitris Athanasakis, John Shawe-Taylor, Maxim Milakov, John Park, Radu Tudor Ionescu, Marius Popescu, Cristian Grozea, James Bergstra, Jingjing Xie, Lukasz Romaszko, Yoshua Bengio
ICONIP (3)11
2013 Labeling Spain With Stanford
abstract
We present an end-to-end framework for outdoor scene region decomposition, learned on a small set of randomly selected images that generalizes well to multiple data sets containing images from around the world. We discuss the different aspects of the framework especially a generalized variational inference method with better approximations to the true marginals of a graphical model. Experimentally, we explain why the framework is robust and performs competitively on many diverse scene data sets, including several unseen scene types. We have obtained high pixel-level accuracies (≈ 80%) in three of the four data sets, which include a benchmark data set known as the Stanford background data set. Our model obtained over 70% accuracy on the fourth data set, which contained a number of indoor and close-up images that are significantly different from our training examples.
Yingbo Zhou 0002, Ifeoma Nwogu, Venu Govindaraju
IEEE Trans. Image Process.1
2012 Handwritten Arabic text recognition using Deep Belief Networks
Utkarsh Porwal, Yingbo Zhou 0002, Venu Govindaraju
ICPR2
2012 Human Identification Using Finger Images
abstract
This paper presents a new approach to improve the performance of finger-vein identification systems presented in the literature. The proposed system simultaneously acquires the finger-vein and low-resolution fingerprint images and combines these two evidences using a novel score-level combination strategy. We examine the previously proposed finger-vein identification approaches and develop a new approach that illustrates it superiority over prior published efforts. The utility of low-resolution fingerprint images acquired from a webcam is examined to ascertain the matching performance from such images. We develop and investigate two new score-level combinations, i.e., holistic and nonlinear fusion, and comparatively evaluate them with more popular score-level fusion approaches to ascertain their effectiveness in the proposed system. The rigorous experimental results presented on the database of 6264 images from 156 subjects illustrate significant improvement in the performance, i.e., both from the authentication and recognition experiments.
Ajay Kumar 0001, Yingbo Zhou 0002
IEEE Trans. Image Process.2
2011 DISCO: Describing Images Using Scene Contexts and Objects
abstract
In this paper, we propose a bottom-up approach to generating short descriptive sentences from images, to enhance scene understanding. We demonstrate automatic methods for mapping the visual content in an image to natural spoken or written language. We also introduce a human-in-the-loop evaluation strategy that quantitatively captures the meaningfulness of the generated sentences. We recorded a correctness rate of 60.34% when human users were asked to judge the meaningfulness of the sentences generated from relatively challenging images. Also, our automatic methods compared well with the state-of-the-art techniques for the related computer vision tasks.
Ifeoma Nwogu, Yingbo Zhou 0002, Christopher Brown 0001
AAAI2
2011 Human Identification Using Palm-Vein Images
abstract
This paper presents two new approaches to improve the performance of palm-vein-based identification systems presented in the literature. The proposed approach attempts to more effectively accommodate the potential deformations, rotational and translational changes by encoding the orientation preserving features and utilizing a novel region-based matching scheme. We systematically compare the previously proposed palm-vein identification approaches with our proposed ones on two different databases that are acquired with the contactless and touch-based imaging setup. We evaluate the performance improvement in both verification and recognition scenarios and analyze the influence of enrollment size on the performance. In this context, the proposed approaches are also compared for its superiority using single image enrollment on two different databases. The rigorous experimental results presented in this paper, on the databases of 100 and 250 subjects, consistently conforms the superiority of the proposed approach in both the verification and recognition scenario.
Yingbo Zhou 0002, Ajay Kumar 0001
IEEE Trans. Inf. Forensics Secur.1
2010 Personal Identification from Iris Images Using Localized Radon Transform
abstract
Personal identification using iris images has invited lots of attention in the literature and offered higher accuracy. However, the computational complexity in the feature extraction from the normalized iris images is still of key concern and further efforts are required to develop efficient feature extraction approaches. In this paper, we investigate a new approach for the efficient and effective extraction of iris features using localized Radon transforms. The feature extraction process exploits the orientation information from the local iris texture features using finite Radon transform. The dominant orientation from these Radon transform features is used to generate a binarized/compact feature representation. The similarity between two feature vectors is computed from the minimum matching distance that can account for the variations resulting from translation and rotation of the images. The feasibility of this approach is rigorously evaluated on two publically available iris image databases, i.e. IITD iris image database v1 and CASIA v3 iris image database. We also investigate the multi-scale analysis of iris images to enhance the performance. The experimental results presented in this paper are highly promising and suggest the computationally attractive alternative for the online iris identification.
Yingbo Zhou 0002, Ajay Kumar 0001
ICPR1