VLDB 2026 Research / reviewers in the wild / expert
Dzmitry Bahdanau
dblp:151/6504
· DBLP profile ↗
19ranked-venue papers
4as first author
10since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Evaluating In-Context Learning of Libraries for Code GenerationabstractArkil Patel, Siva Reddy, Dzmitry Bahdanau, Pradeep Dasigi. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Arkil Patel, Siva Reddy, Dzmitry Bahdanau, Pradeep Dasigi |
NAACL-HLT | 3 |
| 2023 | MAGNIFICo: Evaluating the In-Context Learning Ability of Large Language Models to Generalize to Novel InterpretationsabstractHumans possess a remarkable ability to assign novel interpretations to linguistic expressions, enabling them to learn new words and understand community-specific connotations.However, Large Language Models (LLMs) have a knowledge cutoff and are costly to finetune repeatedly.Therefore, it is crucial for LLMs to learn novel interpretations in-context.In this paper, we systematically analyse the ability of LLMs to acquire novel interpretations using incontext learning.To facilitate our study, we introduce MAGNIFICO, an evaluation suite implemented within a text-to-SQL semantic parsing framework that incorporates diverse tokens and prompt settings to simulate real-world complexity.Experimental results on MAGNIFICO demonstrate that LLMs exhibit a surprisingly robust capacity for comprehending novel interpretations from natural language descriptions as well as from discussions within long conversations.Nevertheless, our findings also highlight the need for further improvements, particularly when interpreting unfamiliar words or when composing multiple novel interpretations simultaneously in the same example.Additionally, our analysis uncovers the semantic predispositions in LLMs and reveals the impact of recency bias for information presented in long contexts. Arkil Patel, Satwik Bhattamishra, Siva Reddy, Dzmitry Bahdanau |
EMNLP | 4 |
| 2023 | PromptMix: A Class Boundary Augmentation Method for Large Language Model DistillationabstractData augmentation is a widely used technique to address the problem of text classification when there is a limited amount of training data.Recent work often tackles this problem using large language models (LLMs) like GPT3 that can generate new examples given already available ones.In this work, we propose a method to generate more helpful augmented data by utilizing the LLM's abilities to follow instructions and perform few-shot classifications.Our specific PromptMix method consists of two steps: 1) generate challenging text augmentations near class boundaries; however, generating borderline examples increases the risk of false positives in the dataset, so we 2) relabel the text augmentations using a prompting-based LLM classifier to enhance the correctness of labels in the generated data.We evaluate the proposed method in challenging 2-shot and zero-shot settings on four text classification datasets: Bank-ing77, TREC6, Subjectivity (SUBJ), and Twitter Complaints.Our experiments show that generating and, crucially, relabeling borderline examples facilitates the transfer of knowledge of a massive LLM like GPT3.5-turbo into smaller and cheaper classifiers like DistilBERT base and BERT base .Furthermore, 2-shot Prompt-Mix outperforms multiple 5-shot data augmentation methods on the four datasets. Gaurav Sahu, Olga Vechtomova, Dzmitry Bahdanau, Issam H. Laradji |
EMNLP | 3 |
| 2022 | Compositional Generalization in Dependency ParsingabstractCompositionality-the ability to combine familiar units like words into novel phrases and sentences-has been the focus of intense interest in artificial intelligence in recent years.To test compositional generalization in semantic parsing, Keysers et al. (2020) introduced Compositional Freebase Queries (CFQ).This dataset maximizes the similarity between the test and train distributions over primitive units, like words, while maximizing the compound divergence-the dissimilarity between test and train distributions over larger structures, like phrases.Dependency parsing, however, lacks a compositional generalization benchmark.In this work, we introduce a gold-standard set of dependency parses for CFQ, and use this to analyze the behavior of a state-of-the art dependency parser (Qi et al., 2020) on the CFQ dataset.We find that increasing compound divergence degrades dependency parsing performance, although not as dramatically as semantic parsing performance.Additionally, we find the performance of the dependency parser does not uniformly degrade relative to compound divergence, and the parser performs differently on different splits with the same compound divergence.We explore a number of hypotheses for what causes the non-uniform degradation in dependency parsing performance, and identify a number of syntactic structures that drive the dependency parser's lower performance on the most challenging splits. Emily Goodwin, Siva Reddy, Timothy J. O'Donnell, Dzmitry Bahdanau |
ACL (1) | 4 |
| 2022 | LAGr: Label Aligned Graphs for Better Systematic Generalization in Semantic ParsingabstractSemantic parsing is the task of producing structured meaning representations for natural language sentences.Recent research has pointed out that the commonly-used sequenceto-sequence (seq2seq) semantic parsers struggle to generalize systematically, i.e. to handle examples that require recombining known knowledge in novel settings.In this work, we show that better systematic generalization can be achieved by producing the meaning representation directly as a graph and not as a sequence.To this end we propose LAGr (Label Aligned Graphs), a general framework to produce semantic parses by independently predicting node and edge labels for a complete multi-layer input-aligned graph.The strongly-supervised LAGr algorithm requires aligned graphs as inputs, whereas weaklysupervised LAGr infers alignments for originally unaligned target graphs using approximate maximum-a-posteriori inference.Experiments demonstrate that LAGr achieves significant improvements in systematic generalization upon the baseline seq2seq parsers in both strongly-and weakly-supervised settings. Dora Jambor, Dzmitry Bahdanau |
ACL (1) | 2 |
| 2021 | PICARD: Parsing Incrementally for Constrained Auto-Regressive Decoding from Language ModelsabstractLarge pre-trained language models for textual data have an unconstrained output space; at each decoding step, they can produce any of 10,000s of sub-word tokens.When fine-tuned to target constrained formal languages like SQL, these models often generate invalid code, rendering it unusable.We propose PICARD 1 , a method for constraining auto-regressive decoders of language models through incremental parsing.PICARD helps to find valid output sequences by rejecting inadmissible tokens at each decoding step.On the challenging Spider and CoSQL text-to-SQL translation tasks, we show that PICARD transforms fine-tuned T5 models with passable performance into stateof-the-art solutions. Torsten Scholak, Nathan Schucher, Dzmitry Bahdanau |
EMNLP (1) | 3 |
| 2021 | Combating False Negatives in Adversarial Imitation LearningabstractIn adversarial imitation learning, a discriminator is trained to differentiate agent episodes from expert demonstrations representing the desired behavior. However, as the trained policy learns to be more successful, the negative examples (the ones produced by the agent) become increasingly similar to expert ones. Despite the fact that the task is successfully accomplished in some of the agent's trajectories, the discriminator is trained to output low values for them. We hypothesize that this inconsistent training signal for the discriminator can impede its learning, and consequently leads to worse overall performance of the agent. We show experimental evidence for this hypothesis and that the ‘False Negatives’ (i.e. successful agent episodes) significantly hinder adversarial imitation learning, which is the first contribution of this paper. Then, we propose a method to alleviate the impact of false negatives and test it on the BabyAI environment. This method consistently improves sample efficiency over the baselines by at least an order of magnitude. Konrad Zolna, Chitwan Saharia, Léonard Boussioux, David Yu-Tung Hui, Maxime Chevalier-Boisvert, Dzmitry Bahdanau, Yoshua Bengio |
IJCNN | 6 |
| 2021 | Understanding by Understanding Not: Modeling Negation in Language ModelsabstractArian Hosseini, Siva Reddy, Dzmitry Bahdanau, R Devon Hjelm, Alessandro Sordoni, Aaron Courville. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Seyed Arian Hosseini, Siva Reddy, Dzmitry Bahdanau, R. Devon Hjelm, Alessandro Sordoni, Aaron C. Courville |
NAACL-HLT | 3 |
| 2021 | DuoRAT: Towards Simpler Text-to-SQL ModelsabstractTorsten Scholak, Raymond Li, Dzmitry Bahdanau, Harm de Vries, Chris Pal. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Torsten Scholak, Raymond Li, Dzmitry Bahdanau, Harm de Vries, Christopher Joseph Pal |
NAACL-HLT | 3 |
| 2021 | Systematic Generalization with Edge TransformersabstractRecent research suggests that systematic generalization in natural language understanding remains a challenge for state-of-the-art neural models such as Transformers and Graph Neural Networks. To tackle this challenge, we propose Edge Transformer, a new model that combines inspiration from Transformers and rule-based symbolic AI. The first key idea in Edge Transformers is to associate vector states with every edge, that is, with every pair of input nodes---as opposed to just every node, as it is done in the Transformer model. The second major innovation is a triangular attention mechanism that updates edge representations in a way that is inspired by unification from logic programming. We evaluate Edge Transformer on compositional generalization benchmarks in relational reasoning, semantic parsing, and dependency parsing. In all three settings, the Edge Transformer outperforms Relation-aware, Universal and classical Transformer baselines. Leon Bergen, Timothy J. O'Donnell, Dzmitry Bahdanau |
NeurIPS | 3 |
| 2020 | Combating False Negatives in Adversarial Imitation Learning (Student Abstract)abstractWe define the False Negatives problem and show that it is a significant limitation in adversarial imitation learning. We propose a method that solves the problem by leveraging the nature of goal-conditioned tasks. The method, dubbed Fake Conditioning, is tested on instruction following tasks in BabyAI environments, where it improves sample efficiency over the baselines by at least an order of magnitude. Konrad Zolna, Chitwan Saharia, Léonard Boussioux, David Yu-Tung Hui, Maxime Chevalier-Boisvert, Dzmitry Bahdanau, Yoshua Bengio |
AAAI | 6 |
| 2019 | Learning to Understand Goal Specifications by Modelling Reward
Dzmitry Bahdanau, Felix Hill, Jan Leike, Edward Hughes 0001, Seyed Arian Hosseini, Pushmeet Kohli, Edward Grefenstette |
ICLR (Poster) | 1 |
| 2019 | Systematic Generalization: What Is Required and Can It Be Learned?
Dzmitry Bahdanau, Shikhar Murty, Michael Noukhovitch, Thien Huu Nguyen, Harm de Vries, Aaron C. Courville |
ICLR (Poster) | 1 |
| 2019 | BabyAI: A Platform to Study the Sample Efficiency of Grounded Language Learning
Maxime Chevalier-Boisvert, Dzmitry Bahdanau, Salem Lahlou, Lucas Willems, Chitwan Saharia, Thien Huu Nguyen, Yoshua Bengio |
ICLR (Poster) | 2 |
| 2017 | An Actor-Critic Algorithm for Sequence Prediction
Dzmitry Bahdanau, Philemon Brakel, Kelvin Xu, Anirudh Goyal, Ryan Lowe, Joelle Pineau, Aaron C. Courville, Yoshua Bengio |
ICLR (Poster) | 1 |
| 2017 | Sequence Tutor: Conservative Fine-Tuning of Sequence Generation Models with KL-controlabstractThis paper proposes a general method for improving the structure and quality of sequences generated by a recurrent neural network (RNN), while maintaining information originally learned from data, as well as sample diversity. An RNN is first pre-trained on data using maximum likelihood estimation (MLE), and the probability distribution over the next token in the sequence learned by this model is treated as a prior policy. Another RNN is then trained using reinforcement learning (RL) to generate higher-quality outputs that account for domain-specific incentives while retaining proximity to the prior policy of the MLE RNN. To formalize this objective, we derive novel off-policy RL methods for RNNs from KL-control. The effectiveness of the approach is demonstrated on two applications; 1) generating novel musical melodies, and 2) computational molecular generation. For both problems, we show that the proposed method improves the desired properties and structure of the generated sequences, while maintaining information learned from data. Natasha Jaques, Shixiang Gu, Dzmitry Bahdanau, José Miguel Hernández-Lobato, Richard E. Turner, Douglas Eck |
ICML | 3 |
| 2016 | End-to-end attention-based large vocabulary speech recognitionabstractMany state-of-the-art Large Vocabulary Continuous Speech Recognition (LVCSR) Systems are hybrids of neural networks and Hidden Markov Models (HMMs). Recently, more direct end-to-end methods have been investigated, in which neural architectures were trained to model sequences of characters [1,2]. To our knowledge, all these approaches relied on Connectionist Temporal Classification [3] modules. We investigate an alternative method for sequence modelling based on an attention mechanism that allows a Recurrent Neural Network (RNN) to learn alignments between sequences of input frames and output labels. We show how this setup can be applied to LVCSR by integrating the decoding RNN with an n-gram language model and by speeding up its operation by constraining selections made by the attention mechanism and by reducing the source sequence lengths by pooling information over time. Recognition accuracies similar to other HMM-free RNN-based approaches are reported for the Wall Street Journal corpus. Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, Yoshua Bengio |
ICASSP | 1 |
| 2015 | Attention-Based Models for Speech RecognitionabstractRecurrent sequence generators conditioned on input data through an attention mechanism have recently shown very good performance on a range of tasks including machine translation, handwriting synthesis and image caption generation. We extend the attention-mechanism with features needed for speech recognition. We show that while an adaptation of the model used for machine translation reaches a competitive 18.6\% phoneme error rate (PER) on the TIMIT phoneme recognition task, it can only be applied to utterances which are roughly as long as the ones it was trained on. We offer a qualitative explanation of this failure and propose a novel and generic method of adding location-awareness to the attention mechanism to alleviate this issue. The new method yields a model that is robust to long inputs and achieves 18\% PER in single utterances and 20\% in 10-times longer (repeated) utterances. Finally, we propose a change to the attention mechanism that prevents it from concentrating too much on single frames, which further reduces PER to 17.6\% level. Jan Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, Yoshua Bengio |
NIPS | 2 |
| 2014 | Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine TranslationabstractKyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, Yoshua Bengio. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2014. Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, Yoshua Bengio |
EMNLP | 4 |