EDBT 2026 Demo / reviewers in the wild / expert
Yuntian Deng
dblp:166/1720
· DBLP profile ↗
30ranked-venue papers
9as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 8 first-author · 16 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorSystems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TokDrift: When LLM Speaks in Subwords but Code Speaks in GrammarabstractLarge language models (LLMs) for code rely on subword tokenizers, such as byte-pair encoding (BPE), learned from mixed natural language text and programming language code but driven by statistics rather than grammar.As a result, semantically identical code snippets can be tokenized differently depending on superficial factors such as whitespace or identifier naming.To measure the impact of this misalignment, we introduce TOKDRIFT, a framework that applies semantic-preserving rewrite rules to create code variants differing only in tokenization.Across nine code LLMs, including large ones with over 30B parameters, even minor formatting changes can cause substantial shifts in model behavior.Layer-wise analysis shows that the issue originates in early embeddings, where subword segmentation fails to capture grammar token boundaries.Our findings identify misaligned tokenization as a hidden obstacle to reliable code understanding and generation, highlighting the need for grammaraware tokenization for future code LLMs.Benchmark Source Task Input PL Output PL # Samples HumanEval-Fix-py HumanEvalPack (Muennighoff et al., 2023) bug fixing Python Python 164 HumanEval-Fix-java Java Java 164 HumanEval-Explain-py code summarization Python Python 164 HumanEval-Explain-java Java Java 164 Avatar-py2java Avatar (Ahmad et al., 2023; Pan et al., 2024) code translation Python Java 244 Avatar-java2py Java Python 246 CodeNet-py2java CodeNet (Puri et al., 2021; Pan et al., 2024) Python Java 200 CodeNet-java2py Java Python 200 Yinxi Li, Yuntian Deng, Pengyu Nie 0001 |
ACL (1) | 2 |
| 2025 | From Chat Logs to Collective Insights: Aggregative Question AnsweringabstractConversational agents powered by large language models (LLMs) are rapidly becoming integral to our daily interactions, generating unprecedented amounts of conversational data.Such datasets offer a powerful lens into societal interests, trending topics, and collective concerns.Yet, existing approaches typically treat these interactions as independent and miss critical insights that could emerge from aggregating and reasoning across large-scale conversation logs.In this paper, we introduce Aggregative Question Answering, a novel task requiring models to reason over thousands of user-chatbot interactions to answer aggregative queries, such as identifying emerging concerns among specific demographics.To enable research in this direction, we constructed WildChat-AQA, a benchmark comprising 6,027 aggregative questions derived from 182,330 real-world chatbot conversations.Experiments show that existing methods either struggle to reason effectively or incur prohibitive computational costs, underscoring the need for new approaches capable of extracting collective insights from large-scale conversational data.Figure 2: Overview of the WildChat-AQA dataset creation process. Woojeong Kim, Yuntian Deng |
EMNLP | 3 |
| 2025 | WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the WildabstractWe introduce WildBench, an automated evaluation framework designed to benchmark large language models (LLMs) using challenging, real-world user queries. WildBench consists of 1,024 tasks carefully selected from over one million human-chatbot conversation logs. For automated evaluation with WildBench, we have developed two metrics, WB-Reward and WB-Score, which are computable using advanced LLMs such as GPT-4-turbo. WildBench evaluation uses task-specific checklists to evaluate model outputs systematically and provides structured explanations that justify the scores and comparisons, resulting in more reliable and interpretable automatic judgments. WB-Reward employs fine-grained pairwise comparisons between model responses, generating five potential outcomes: much better, slightly better, slightly worse, much worse, or a tie. Unlike previous evaluations that employed a single baseline model, we selected three baseline models at varying performance levels to ensure a comprehensive pairwise evaluation. Additionally, we propose a simple method to mitigate length bias, by converting outcomes of “slightly better/worse” to “tie” if the winner response exceeds the loser one by more than K characters. WB-Score evaluates the quality of model outputs individually, making it a fast and cost-efficient evaluation metric. WildBench results demonstrate a strong correlation with the human-voted Elo ratings from Chatbot Arena on hard tasks. Specifically, WB-Reward achieves a Pearson correlation of 0.98 with top-ranking models. Additionally, WB-Score reaches 0.95, surpassing both ArenaHard’s 0.91 and AlpacaEval2.0’s 0.89 for length-controlled win rates, as well as the 0.87 for regular win rates. Bill Y. Lin, Yuntian Deng, Khyathi Raghavi Chandu, Abhilasha Ravichander, Valentina Pyatkin, Nouha Dziri, Ronan Le Bras 0001, Yejin Choi 0001 |
ICLR | 2 |
| 2025 | MixEval-X: Any-to-any Evaluations from Real-world Data MixtureabstractPerceiving and generating diverse modalities are crucial for AI models to effectively learn from and engage with real-world signals, necessitating reliable evaluations for their development. We identify two major issues in current evaluations: (1) inconsistent standards, shaped by different communities with varying protocols and maturity levels; and (2) significant query, grading, and generalization biases. To address these, we introduce MixEval-X, the first any-to-any, real-world benchmark designed to optimize and standardize evaluations across diverse input and output modalities. We propose multi-modal benchmark mixture and adaptation-rectification pipelines to reconstruct real-world task distributions, ensuring evaluations generalize effectively to real-world use cases. Extensive meta-evaluations show our approach effectively aligns benchmark samples with real-world task distributions. Meanwhile, MixEval-X's model rankings correlate strongly with that of crowd-sourced real-world evaluations (up to 0.98) while being much more efficient. We provide comprehensive leaderboards to rerank existing models and organizations and offer insights to enhance understanding of multi-modal evaluations and inform future research. Jinjie Ni, Deepanway Ghosal, Bo Li 0080, Junhao Zhang 0001, Xiang Yue, Fuzhao Xue, Yuntian Deng, Zian Zheng 0001, Kaichen Zhang, Mahir Shah, Kabir Jain, Yang You 0001, Michael Shieh |
ICLR | 8 |
| 2025 | Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with NothingabstractHigh-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent existing open-source data creation methods from scaling effectively, potentially limiting the diversity and quality of public alignment datasets. Is it possible to synthesize high-quality instruction data at scale by extracting it directly from an aligned LLM? We present a self-synthesis method for generating large-scale alignment data named Magpie. Our key observation is that aligned LLMs like Llama-3-Instruct can generate a user query when we input only the pre-query templates up to the position reserved for user messages, thanks to their auto-regressive nature. We use this method to prompt Llama-3-Instruct and generate 4 million instructions along with their corresponding responses. We further introduce extensions of Magpie for filtering, generating multi-turn, preference optimization, domain-specific and multilingual datasets. We perform a comprehensive analysis of the Magpie-generated data. To compare Magpie-generated data with other public instruction datasets (e.g., ShareGPT, WildChat, Evol-Instruct, UltraChat, OpenHermes, Tulu-V2-Mix, GenQA), we fine-tune Llama-3-8B-Base with each dataset and evaluate the performance of the fine-tuned models. Our results indicate that using Magpie for supervised fine-tuning (SFT) solely can surpass the performance of previous public datasets utilized for both SFT and preference optimization, such as direct preference optimization with UltraFeedback. We also show that in some tasks, models supervised fine-tuned with Magpie perform comparably to the official Llama-3-8B-Instruct, despite the latter being enhanced with 10 million data points through SFT and subsequent preference optimization. This advantage is evident on alignment benchmarks such as AlpacaEval, ArenaHard, and WildBench. Zhangchen Xu, Fengqing Jiang, Luyao Niu, Yuntian Deng, Radha Poovendran, Yejin Choi 0001, Bill Y. Lin |
ICLR | 4 |
| 2025 | The Leaderboard IllusionabstractMeasuring progress is fundamental to the advancement of any scientific field. As benchmarks play an increasingly central role, they also grow more susceptible to distortion.Chatbot Arena has emerged as the go-to leaderboard for ranking the most capable AI systems. Yet, in this work we identify systematic issues that have resulted in a distorted playing field. We find that undisclosed private testing practices benefit a handful of providers who are able to test multiple variants before public release and retract scores if desired. We establish that the ability of these providers to choose the best score leads to biased Arena scores due to selective disclosure of performance results. At an extreme, we found one provider testing 27 private variants before making one model public at the second position on the leaderboard. We also establish that proprietary closed models are sampled at higher rates (number of battles) and have fewer models removed from the arena than open-weight and open-source alternatives. Both these policies lead to large data access asymmetries over time. The top two providers have individually received an estimated 19.2% and 20.4% of all data on the arena. In contrast, a combined 83 open-weight models have only received an estimated 29.7% of the total data. With conservative estimates, we show that access to Chatbot Arena data yields substantial benefits; even limited additional data can result in relative performance gains of up to 112% on ArenaHard, a test set from the arena distribution.Together, these dynamics result in overfitting to Arena-specific dynamics rather than general model quality. The Arena builds on the substantial efforts of both the organizers and an open community that maintains this valuable evaluation platform. We offer actionable recommendations to reform the Chatbot Arena's evaluation framework and promote fairer, more transparent benchmarking for the field. Shivalika Singh, Yiyang Nan, Daniel D'souza, Sayash Kapoor, Ahmet Üstün, Oluwasanmi Koyejo, Yuntian Deng, Shayne Longpre, Noah A. Smith, Beyza Ermis, Marzieh Fadaee, Sara Hooker |
NeurIPS | 8 |
| 2024 | WildChat: 1M ChatGPT Interaction Logs in the WildabstractChatbots such as GPT-4 and ChatGPT are now serving millions of users. Despite their widespread use, there remains a lack of public datasets showcasing how these tools are used by a population of users in practice. To bridge this gap, we offered free access to ChatGPT for online users in exchange for their affirmative, consensual opt-in to anonymously collect their chat transcripts and request headers. From this, we compiled WildChat, a corpus of 1 million user-ChatGPT conversations, which consists of over 2.5 million interaction turns. We compare WildChat with other popular user-chatbot interaction datasets, and find that our dataset offers the most diverse user prompts, contains the largest number of languages, and presents the richest variety of potentially toxic use-cases for researchers to study. In addition to timestamped chat transcripts, we enrich the dataset with demographic data, including state, country, and hashed IP addresses, alongside request headers. This augmentation allows for more detailed analysis of user behaviors across different geographical regions and temporal dimensions. Finally, because it captures a broad range of use cases, we demonstrate the dataset’s potential utility in fine-tuning instruction-following models. WildChat is released at https://wildchat.allen.ai under AI2 ImpACT Licenses. Xiang Ren 0001, Jack Hessel, Claire Cardie, Yejin Choi 0001, Yuntian Deng |
ICLR | 6 |
| 2024 | MixEval: Deriving Wisdom of the Crowd from LLM Benchmark MixturesabstractEvaluating large language models (LLMs) is challenging. Traditional ground-truth- based benchmarks fail to capture the comprehensiveness and nuance of real-world queries, while LLM-as-judge benchmarks suffer from grading biases and limited query quantity. Both of them may also become contaminated over time. User- facing evaluation, such as Chatbot Arena, provides reliable signals but is costly and slow. In this work, we propose MixEval, a new paradigm for establishing efficient, gold-standard LLM evaluation by strategically mixing off-the-shelf bench- marks. It bridges (1) comprehensive and well-distributed real-world user queries and (2) efficient and fairly-graded ground-truth-based benchmarks, by matching queries mined from the web with similar queries from existing benchmarks. Based on MixEval, we further build MixEval-Hard, which offers more room for model improvement. Our benchmarks’ advantages lie in (1) a 0.96 model ranking correlation with Chatbot Arena arising from the highly impartial query distribution and grading mechanism, (2) fast, cheap, and reproducible execution (6% of the time and cost of MMLU), and (3) dynamic evaluation enabled by the rapid and stable data update pipeline. We provide extensive meta-evaluation and analysis for our and existing LLM benchmarks to deepen the community’s understanding of LLM evaluation and guide future research directions. Jinjie Ni, Fuzhao Xue, Xiang Yue, Yuntian Deng, Mahir Shah, Kabir Jain, Graham Neubig, Yang You 0001 |
NeurIPS | 4 |
| 2023 | Tree Prompting: Efficient Task Adaptation without Fine-TuningabstractPrompting language models (LMs) is the main interface for applying them to new tasks.However, for smaller LMs, prompting provides low accuracy compared to gradient-based finetuning.Tree Prompting is an approach to prompting which builds a decision tree of prompts, linking multiple LM calls together to solve a task.At inference time, each call to the LM is determined by efficiently routing the outcome of the previous call using the tree.Experiments on classification datasets show that Tree Prompting improves accuracy over competing methods and is competitive with fine-tuning.We also show that variants of Tree Prompting allow inspection of a model's decision-making process. 1 Chandan Singh, John X. Morris, Alexander M. Rush, Jianfeng Gao 0001, Yuntian Deng |
EMNLP | 5 |
| 2023 | Markup-to-Image Diffusion Models with Scheduled Sampling
Yuntian Deng, Noriyuki Kojima, Alexander M. Rush |
ICLR | 1 |
| 2023 | Semi-Parametric Inducing Point Networks and Neural Processes
Richa Rastogi, Yair Schiff, Alon Hacohen, Zhaozhi Li, Ian Lee, Yuntian Deng, Mert R. Sabuncu, Volodymyr Kuleshov |
ICLR | 6 |
| 2022 | Weighted Gaussian Process Bandits for Non-stationary EnvironmentsabstractIn this paper, we consider the Gaussian process (GP) bandit optimization problem in a non-stationary environment. To capture external changes, the black-box function is allowed to be time-varying within a reproducing kernel Hilbert space (RKHS). To this end, we develop WGP-UCB, a novel UCB-type algorithm based on weighted Gaussian process regression. A key challenge is how to cope with infinite-dimensional feature maps. To that end, we leverage kernel approximation techniques to prove a sublinear regret bound, which is the first (frequentist) sublinear regret guarantee on weighted time-varying bandits with general nonlinear rewards. This result generalizes both non-stationary linear bandits and standard GP-UCB algorithms. Further, a novel concentration inequality is achieved for weighted Gaussian process regression with general weights. We also provide universal upper bounds and weight-dependent upper bounds for weighted maximum information gains. These results are of independent interest for applications such as news ranking and adaptive pricing, where weights can be adopted to capture the importance or quality of data. Finally, we conduct experiments to highlight the favorable gains of the proposed algorithm in many cases when compared to existing methods. Yuntian Deng, Xingyu Zhou 0001, Baekjin Kim, Ambuj Tewari, Abhishek Gupta 0002, Ness Shroff |
AISTATS | 1 |
| 2022 | Model Criticism for Long-Form Text GenerationabstractLanguage models have demonstrated the ability to generate highly fluent text; however, it remains unclear whether their output retains coherent high-level structure (e.g., story progression).Here, we propose to apply a statistical tool, model criticism in latent space, to evaluate the high-level structure of the generated text.Model criticism compares the distributions between real and generated data in a latent space obtained according to an assumptive generative process.Different generative processes identify specific failure modes of the underlying model.We perform experiments on three representative aspects of highlevel discourse-coherence, coreference, and topicality-and find that transformer-based language models are able to capture topical structures but have a harder time maintaining structural coherence or modeling coreference. Yuntian Deng, Volodymyr Kuleshov, Alexander M. Rush |
EMNLP | 1 |
| 2022 | Interference Constrained Beam Alignment for Time-Varying Channels via Kernelized BanditsabstractTo fully utilize the abundant spectrum resources in millimeter wave (mmWave), Beam Alignment (BA) is necessary for large antenna arrays to achieve large array gains. In practical dynamic wireless environments, channel modeling is challenging due to time-varying and multipath effects. In this paper, we formulate the beam alignment problem as a nonstationary online learning problem with the objective to maximize the received signal strength under interference constraint. In particular, we employ the non-stationary kernelized bandit to leverage the correlation among beams and model the complex beamforming and multipath channel functions. Furthermore, to mitigate interference to other user equipment, we leverage the primal-dual method to design a constrained UCB-type kernelized bandit algorithm. Our theoretical analysis indicates that the proposed algorithm can adaptively adjust the beam in time-varying environments, such that both the cumulative regret of the received signal and constraint violations have sublinear bounds with respect to time. This result is of independent interest for applications such as adaptive pricing and news ranking. In addition, the algorithm assumes the channel is a black-box function and does not require any prior knowledge for dynamic channel modeling, and thus is applicable in a variety of scenarios. We further show that if the information about the channel variation is known, the algorithm will have better theoretical guarantees and performance. Finally, we conduct simulations to highlight the effectiveness of the proposed algorithm. Yuntian Deng, Xingyu Zhou 0001, Arnob Ghosh, Abhishek Gupta 0002, Ness Shroff |
WiOpt | 1 |
| 2021 | Rationales for Sequential PredictionsabstractSequence models are a critical component of modern NLP systems, but their predictions are difficult to explain.We consider model explanations though rationales, subsets of context that can explain individual model predictions.We find sequential rationales by solving a combinatorial optimization: the best rationale is the smallest subset of input tokens that would predict the same output as the full sequence.Enumerating all subsets is intractable, so we propose an efficient greedy algorithm to approximate this objective.The algorithm, which is called greedy rationalization, applies to any model.For this approach to be effective, the model should form compatible conditional distributions when making predictions on incomplete subsets of the context.This condition can be enforced with a short finetuning step.We study greedy rationalization on language modeling and machine translation.Compared to existing baselines, greedy rationalization is best at optimizing the sequential objective and provides the most faithful rationales.On a new dataset of annotated sequential rationales, greedy rationales are most similar to human rationales. Keyon Vafa, Yuntian Deng, David M. Blei, Alexander M. Rush |
EMNLP (1) | 2 |
| 2021 | Low-Rank Constraints for Fast Inference in Structured ModelsabstractStructured distributions, i.e. distributions over combinatorial spaces, are commonly used to learn latent probabilistic representations from observed data. However, scaling these models is bottlenecked by the high computational and memory complexity with respect to the size of the latent representations. Common models such as Hidden Markov Models (HMMs) and Probabilistic Context-Free Grammars (PCFGs) require time and space quadratic and cubic in the number of hidden states respectively. This work demonstrates a simple approach to reduce the computational and memory complexity of a large class of structured models. We show that by viewing the central inference step as a matrix-vector product and using a low-rank constraint, we can trade off model expressivity and speed via the rank. Experiments with neural parameterized structured models for language modeling, polyphonic music modeling, unsupervised grammar induction, and video modeling show that our approach matches the accuracy of standard models at large state spaces while providing practical speedups. Justin T. Chiu, Yuntian Deng, Alexander M. Rush |
NeurIPS | 2 |
| 2021 | Residual Energy-Based Models for TextabstractCurrent large-scale auto-regressive language models display impressive fluency and can generate convincing text. In this work we start by asking the question: Can the generations of these models be reliably distinguished from real text by statistical discriminators? We find experimentally that the answer is affirmative when we have access to the training data for the model, and guardedly affirmative even if we do not. This suggests that the auto-regressive models can be improved by incorporating the (globally normalized) discriminators into the generative process. We give a formalism for this using the Energy-Based Model framework, and show that it indeed improves the results of the generative models, measured both in terms of perplexity and in terms of human evaluation. Anton Bakhtin, Yuntian Deng, Sam Gross, Myle Ott, Marc'Aurelio Ranzato, Arthur Szlam |
J. Mach. Learn. Res. | 2 |
| 2020 | Algorithm-Hardware Co-Design of Adaptive Floating-Point Encodings for Resilient Deep Learning InferenceabstractConventional hardware-friendly quantization methods, such as fixed-point or integer, tend to perform poorly at very low precision as their shrunken dynamic ranges cannot adequately capture the wide data distributions commonly seen in sequence transduction models. We present an algorithm-hardware co-design centered around a novel floating-point inspired number format, AdaptivFloat, that dynamically maximizes and optimally clips its available dynamic range, at a layer granularity, in order to create faithful encodings of neural network parameters. AdaptivFloat consistently produces higher inference accuracies compared to block floating-point, uniform, IEEE-like float or posit encodings at low bit precision (≤8-bit) across a diverse set of state-of-the-art neural networks, exhibiting narrow to wide weight distribution. Notably, at 4-bit weight precision, only a 2.1 degradation in BLEU score is observed on the AdaptivFloat-quantized Transformer network compared to total accuracy loss when encoded in the above-mentioned prominent datatypes. Furthermore, experimental results on a deep neural network (DNN) processing element (PE), exploiting AdaptivFloat logic in its computational datapath, demonstrate per-operation energy and area that is 0.9× and 1.14×, width, respectively that of an equivalent bit NVDLA-like integer-based PE. Thierry Tambe, En-Yu Yang, Zishen Wan, Yuntian Deng, Vijay Janapa Reddi, Alexander M. Rush, David Brooks 0001, Gu-Yeon Wei |
DAC | 4 |
| 2020 | Residual Energy-Based Models for Text Generation
Yuntian Deng, Anton Bakhtin, Myle Ott, Arthur Szlam, Marc'Aurelio Ranzato |
ICLR | 1 |
| 2020 | Cascaded Text Generation with Markov TransformersabstractThe two dominant approaches to neural text generation are fully autoregressive models, using serial beam search decoding, and non-autoregressive models, using parallel decoding with no output dependencies. This work proposes an autoregressive model with sub-linear parallel time generation. Noting that conditional random fields with bounded context can be decoded in parallel, we propose an efficient cascaded decoding approach for generating high-quality output. To parameterize this cascade, we introduce a Markov transformer, a variant of the popular fully autoregressive model that allows us to simultaneously decode with specific autoregressive context cutoffs. This approach requires only a small modification from standard autoregressive training, while showing competitive accuracy/speed tradeoff compared to existing methods on five machine translation datasets. Yuntian Deng, Alexander M. Rush |
NeurIPS | 1 |
| 2019 | Neural Linguistic SteganographyabstractZachary Ziegler, Yuntian Deng, Alexander Rush. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Zachary M. Ziegler, Yuntian Deng, Alexander M. Rush |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Challenges in End-to-End Neural Scientific Table RecognitionabstractIn recent years, end-to-end trained neural models have been applied successfully to various optical character recognition (OCR) tasks. However, the same level of success has not yet been achieved in end-to-end neural scientific table recognition, which involves multi-row/multi-column structures and math formulas in cells. In this paper, we take a step forward to full end-to-end scientific table recognition by constructing a large dataset consisting of 450K table images paired with corresponding LaTeX sources. We apply a popular attentional encoder-decoder model to this dataset and show the potential of end-to-end trained neural systems, as well as associated challenges. Yuntian Deng, David S. Rosenberg, Gideon Mann |
ICDAR | 1 |
| 2018 | Bottom-Up Abstractive SummarizationabstractNeural network-based methods for abstractive summarization produce outputs that are more fluent than other techniques, but which can be poor at content selection.This work proposes a simple technique for addressing this issue: use a data-efficient content selector to overdetermine phrases in a source document that should be part of the summary.We use this selector as a bottom-up attention step to constrain the model to likely phrases.We show that this approach improves the ability to compress text, while still generating fluent summaries.This two-step process is both simpler and higher performing than other end-to-end content selection models, leading to significant improvements on ROUGE for both the CNN-DM and NYT corpus.Furthermore, the content selector can be trained with as little as 1,000 sentences, making it easy to transfer a trained summarizer to a new domain. Sebastian Gehrmann, Yuntian Deng, Alexander M. Rush |
EMNLP | 2 |
| 2018 | Latent Alignment and Variational AttentionabstractNeural attention has become central to many state-of-the-art models in natural language processing and related domains. Attention networks are an easy-to-train and effective method for softly simulating alignment; however, the approach does not marginalize over latent alignments in a probabilistic sense. This property makes it difficult to compare attention to other alignment approaches, to compose it with probabilistic models, and to perform posterior inference conditioned on observed data. A related latent approach, hard attention, fixes these issues, but is generally harder to train and less accurate. This work considers variational attention networks, alternatives to soft and hard attention for learning latent variable alignment models, with tighter approximation bounds based on amortized variational inference. We further propose methods for reducing the variance of gradients to make these approaches computationally feasible. Experiments show that for machine translation and visual question answering, inefficient exact latent variable models outperform standard neural attention, but these gains go away when using hard attention based training. On the other hand, variational attention retains most of the performance gain but with training speed comparable to neural attention. Yuntian Deng, Justin T. Chiu, Demi Guo, Alexander M. Rush |
NeurIPS | 1 |
| 2017 | Dropout with Expectation-linear Regularization
Xuezhe Ma, Yingkai Gao, Zhiting Hu, Yaoliang Yu, Yuntian Deng, Eduard H. Hovy |
ICLR (Poster) | 5 |
| 2017 | Image-to-Markup Generation with Coarse-to-Fine AttentionabstractWe present a neural encoder-decoder model to convert images into presentational markup based on a scalable coarse-to-fine attention mechanism. Our method is evaluated in the context of image-to-LaTeX generation, and we introduce a new dataset of real-world rendered mathematical expressions paired with LaTeX markup. We show that unlike neural OCR techniques using CTC-based models, attention-based approaches can tackle this non-standard OCR task. Our approach outperforms classical mathematical OCR systems by a large margin on in-domain rendered data, and, with pretraining, also performs well on out-of-domain handwritten data. To reduce the inference complexity associated with the attention-based approaches, we introduce a new coarse-to-fine attention layer that selects a support region before applying attention. Yuntian Deng, Anssi Kanervisto, Jeffrey Ling, Alexander M. Rush |
ICML | 1 |
| 2017 | Learning Latent Space Models with Angular ConstraintsabstractThe large model capacity of latent space models (LSMs) enables them to achieve great performance on various applications, but meanwhile renders LSMs to be prone to overfitting. Several recent studies investigate a new type of regularization approach, which encourages components in LSMs to be diverse, for the sake of alleviating overfitting. While they have shown promising empirical effectiveness, in theory why larger “diversity” results in less overfitting is still unclear. To bridge this gap, we propose a new diversity-promoting approach that is both theoretically analyzable and empirically effective. Specifically, we use near-orthogonality to characterize “diversity” and impose angular constraints (ACs) on the components of LSMs to promote diversity. A generalization error analysis shows that larger diversity results in smaller estimation error and larger approximation error. An efficient ADMM algorithm is developed to solve the constrained LSM problems. Experiments demonstrate that ACs improve generalization performance of LSMs and outperform other diversity-promoting approaches. Pengtao Xie, Yuntian Deng, Yi Zhou 0017, Abhimanu Kumar, Yaoliang Yu, James Zou 0001, Eric P. Xing |
ICML | 2 |
| 2016 | Learning Concept Taxonomies from Multi-modal DataabstractWe study the problem of automatically building hypernym taxonomies from textual and visual data.Previous works in taxonomy induction generally ignore the increasingly prominent visual data, which encode important perceptual semantics.Instead, we propose a probabilistic model for taxonomy induction by jointly leveraging text and images.To avoid hand-crafted feature engineering, we design end-to-end features based on distributed representations of images and words.The model is discriminatively trained given a small set of existing ontologies and is capable of building full taxonomies from scratch for a collection of unseen conceptual label items with associated images.We evaluate our model and features on the WordNet hierarchies, where our system outperforms previous approaches by a large gap. Hao Zhang 0025, Zhiting Hu, Yuntian Deng, Mrinmaya Sachan, Zhicheng Yan 0001, Eric P. Xing |
ACL (1) | 3 |
| 2015 | Entity Hierarchy EmbeddingabstractZhiting Hu, Poyao Huang, Yuntian Deng, Yingkai Gao, Eric Xing. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Zhiting Hu, Po-Yao Huang 0001, Yuntian Deng, Yingkai Gao, Eric P. Xing |
ACL (1) | 3 |
| 2015 | Diversifying Restricted Boltzmann Machine for Document ModelingabstractRestricted Boltzmann Machine (RBM) has shown great effectiveness in document modeling. It utilizes hidden units to discover the latent topics and can learn compact semantic representations for documents which greatly facilitate document retrieval, clustering and classification. The popularity (or frequency) of topics in text corpora usually follow a power-law distribution where a few dominant topics occur very frequently while most topics (in the long-tail region) have low probabilities. Due to this imbalance, RBM tends to learn multiple redundant hidden units to best represent dominant topics and ignore those in the long-tail region, which renders the learned representations to be redundant and non-informative. To solve this problem, we propose Diversified RBM (DRBM) which diversifies the hidden units, to make them cover not only the dominant topics, but also those in the long-tail region. We define a diversity metric and use it as a regularizer to encourage the hidden units to be diverse. Since the diversity metric is hard to optimize directly, we instead optimize its lower bound and prove that maximizing the lower bound with projected gradient ascent can increase this diversity metric. Experiments on document retrieval and clustering demonstrate that with diversification, the document modeling power of DRBM can be greatly improved. Pengtao Xie, Yuntian Deng, Eric P. Xing |
KDD | 2 |