EDBT 2026 Demo / reviewers in the wild / expert
Andrea Madotto
dblp:174/2905
· DBLP profile ↗
33ranked-venue papers
4as first author
19since 2021 · last 2025
0000-0002-8672-715XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 4 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Proactive Assistant Dialogue Generation from Streaming Egocentric VideosabstractYichi Zhang, Xin Luna Dong, Zhaojiang Lin, Andrea Madotto, Anuj Kumar, Babak Damavandi, Joyce Chai, Seungwhan Moon. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yichi Zhang 0001, Xin Dong 0001, Zhaojiang Lin, Andrea Madotto, Babak Damavandi, Joyce Y. Chai, Seungwhan Moon |
EMNLP | 4 |
| 2025 | Perception Encoder: The best visual embeddings are not at the output of the networkabstractWe introduce Perception Encoder (PE), a family of state-of-the-art vision encoders for image and video understanding. Traditionally, vision encoders have relied on a variety of pretraining objectives, each excelling at different downstream tasks. Surprisingly, after scaling a carefully tuned image pretraining recipe and refining with a robust video data engine, we find that contrastive vision-language training alone can produce strong, general embeddings for all of these downstream tasks. There is only one caveat: these embeddings are hidden within the intermediate layers of the network. To draw them out, we introduce two alignment methods: language alignment for multimodal language modeling, and spatial alignment for dense prediction. Together, our PE family of models achieves state-of-the-art results on a wide variety of tasks, including zero-shot image and video classification and retrieval; document, image, and video Q&A; and spatial tasks such as detection, tracking, and depth estimation. We release our models, code, and novel dataset of synthetically and human-annotated videos: https://github.com/facebookresearch/perception_models Daniel Bolya, Po-Yao Huang 0001, Peize Sun, Jang Hyun Cho, Andrea Madotto, Chen Wei 0005, Tengyu Ma 0005, Jiale Zhi, Jathushan Rajasegaran, Hanoona Rasheed, Marco Monteiro, Hu Xu 0001, Shiyu Dong, Nikhila Ravi, Shang-Wen Li 0001, Piotr Dollár, Christoph Feichtenhofer |
NeurIPS | 5 |
| 2025 | PerceptionLM: Open-Access Data and Models for Detailed Visual UnderstandingabstractVision-language models are integral to computer vision research, yet many high-performing models remain closed-source, obscuring their data, design and training recipe. The research community has responded by using distillation from black-box models to label training data, achieving strong benchmark results, at the cost of measurable scientific progress. However, without knowing the details of the teacher model and its data sources, scientific progress remains difficult to measure. In this paper, we study building a Perception Language Model (PLM) in a fully open and reproducible framework for transparent research in image and video understanding. We analyze standard training pipelines without distillation from proprietary models and explore large-scale synthetic data to identify critical data gaps, particularly in detailed video understanding. To bridge these gaps, we release 2.8M human-labeled instances of fine-grained video question-answer pairs and spatio-temporally grounded video captions. Additionally, we introduce PLM–VideoBench, a suite for evaluating challenging video understanding tasks focusing on the ability to reason about ''what'', ''where'', ''when'', and ''how'' of a video. We make our work fully reproducible by providing data, training recipes, code & models. Jang Hyun Cho, Andrea Madotto, Effrosyni Mavroudi, Triantafyllos Afouras, Tushar Nagarajan, Muhammad Maaz 0001, Yale Song, Tengyu Ma 0005, Shuming Hu, Suyog Dutt Jain, Hanoona Abdul Rasheed, Peize Sun, Po-Yao Huang 0001, Daniel Bolya, Nikhila Ravi, Shashank Jain, Tammy Stark, Seungwhan Moon, Babak Damavandi, Vivian Lee, Andrew Westbury, Salman Khan 0001, Philipp Krähenbühl, Piotr Dollár, Lorenzo Torresani, Kristen Grauman, Christoph Feichtenhofer |
NeurIPS | 2 |
| 2025 | Corgi: Cached Memory Guided Video GenerationabstractText-to-Video generation has achieved remarkable progress with the rise of diffusion models. In this work, we introduce Cached Memory-Guided Video Generation (Corgi), aiming to generate multi-scene videos with arbi-trary number of video clips, conditioned on input images and instruction prompts. This is a challenging task, as tra-ditional T2V methods often struggle to maintain the quality of longer videos due to the difficulties in preserving visual context from earlier scenes. We address this by introducing a cached memory mechanism that stores the key frames. Our multi-scene video generation process is explicitly con-ditioned on the cached memories to avoid forgetting the vi-sual appearance of target subjects. Corgi shows significant improvement in multi-scene video generation compared to the prior art, with up to 59.2% in long-term consistency and 7.6% in diversity. Xindi Wu, Uriel Singer, Zhaojiang Lin, Andrea Madotto, Xide Xia, Paul A. Crook, Xin Dong 0001, Seungwhan Moon |
WACV | 4 |
| 2024 | Fine-Tuned Language Models Generate Stable Inorganic Materials as TextabstractWe propose fine-tuning large language models for generation of stable materials. While unorthodox, fine-tuning large language models on text-encoded atomistic data is simple to implement yet reliable, with around 90\% of sampled structures obeying physical constraints on atom positions and charges. Using energy above hull calculations from both learned ML potentials and gold-standard DFT calculations, we show that our strongest model (fine-tuned LLaMA-2 70B) can generate materials predicted to be metastable at about twice the rate (49\% vs 28\%) of CDVAE, a competing diffusion model. Because of text prompting's inherent flexibility, our models can simultaneously be used for unconditional generation of stable material, infilling of partial structures and text-conditional generation. Finally, we show that language models' ability to capture key symmetries of crystal structures improves with model scale, suggesting that the biases of pretrained LLMs are surprisingly well-suited for atomistic data. Nate Gruver, Anuroop Sriram, Andrea Madotto, Andrew Gordon Wilson, C. Lawrence Zitnick, Zachary W. Ulissi |
ICLR | 3 |
| 2023 | Training Models to Generate, Recognize, and Reframe Unhelpful ThoughtsabstractMounica Maddela, Megan Ung, Jing Xu, Andrea Madotto, Heather Foran, Y-Lan Boureau. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Mounica Maddela, Megan Ung, Jing Xu 0014, Andrea Madotto, Heather Foran, Y-Lan Boureau |
ACL (1) | 4 |
| 2023 | SIMMC-VR: A Task-oriented Multimodal Dialog Dataset with Situated and Immersive VR StreamsabstractTe-Lin Wu, Satwik Kottur, Andrea Madotto, Mahmoud Azab, Pedro Rodriguez, Babak Damavandi, Nanyun Peng, Seungwhan Moon. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Te-Lin Wu, Satwik Kottur, Andrea Madotto, Mahmoud Azab, Pedro Rodríguez 0001, Babak Damavandi, Nanyun Peng 0001, Seungwhan Moon |
ACL (1) | 3 |
| 2023 | Continual Dialogue State Tracking via Example-Guided Question AnsweringabstractHyundong Cho, Andrea Madotto, Zhaojiang Lin, Khyathi Chandu, Satwik Kottur, Jing Xu, Jonathan May, Chinnadhurai Sankar. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Hyundong Cho, Andrea Madotto, Zhaojiang Lin, Khyathi Raghavi Chandu, Satwik Kottur, Jonathan May, Chinnadhurai Sankar |
EMNLP | 2 |
| 2022 | QAConv: Question Answering on Informative ConversationsabstractThis paper introduces QAConv, 1 , a new question answering (QA) dataset that uses conversations as a knowledge source.We focus on informative conversations, including business emails, panel discussions, and work channels.Unlike open-domain and task-oriented dialogues, these conversations are usually long, complex, asynchronous, and involve strong domain knowledge.In total, we collect 34,608 QA pairs from 10,259 selected conversations with both human-written and machinegenerated questions.We use a question generator and a dialogue summarizer as auxiliary tools to collect and recommend questions.The dataset has two testing scenarios: chunk mode and full mode, depending on whether the grounded partial conversation is provided or retrieved.Experimental results show that stateof-the-art pretrained QA systems have limited zero-shot performance and tend to predict our questions as unanswerable.Our dataset provides a new training and evaluation testbed to facilitate QA on conversations research. Chien-Sheng Wu, Andrea Madotto, Wenhao Liu 0003, Pascale Fung, Caiming Xiong |
ACL (1) | 2 |
| 2022 | NeuS: Neutral Multi-News Summarization for Mitigating Framing BiasabstractNayeon Lee, Yejin Bang, Tiezheng Yu, Andrea Madotto, Pascale Fung. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Nayeon Lee, Yejin Bang, Tiezheng Yu, Andrea Madotto, Pascale Fung |
NAACL-HLT | 4 |
| 2021 | The Adapter-Bot: All-In-One Controllable Conversational ModelabstractIn this paper, we present the Adapter-Bot, a generative chat-bot that uses a fixed backbone conversational model such as DialGPT (Zhang et al. 2019) and triggers on-demand dialogue skills via different adapters (Houlsby et al. 2019). Each adapter can be trained independently, thus allowing a continual integration of skills without retraining the entire model. Depending on the skills, the model is able to process multiple knowledge types, such as text, tables, and graphs, in a seamless manner. The dialogue skills can be triggered automatically via a dialogue manager, or manually, thus allowing high-level control of the generated responses. At the current stage, we have implemented 12 response styles (e.g., positive, negative etc.), 6 goal-oriented skills (e.g. weather information, movie recommendation, etc.), and personalized and emphatic responses. Zhaojiang Lin, Andrea Madotto, Yejin Bang, Pascale Fung |
AAAI | 2 |
| 2021 | CrossNER: Evaluating Cross-Domain Named Entity RecognitionabstractCross-domain named entity recognition (NER) models are able to cope with the scarcity issue of NER samples in target domains. However, most of the existing NER benchmarks lack domain-specialized entity types or do not focus on a certain domain, leading to a less effective cross-domain evaluation. To address these obstacles, we introduce a cross-domain NER dataset (CrossNER), a fully-labeled collection of NER data spanning over five diverse domains with specialized entity categories for different domains. Additionally, we also provide a domain-related corpus since using it to continue pre-training language models (domain-adaptive pre-training) is effective for the domain adaptation. We then conduct comprehensive experiments to explore the effectiveness of leveraging different levels of the domain corpus and pre-training strategies to do domain-adaptive pre-training for the cross-domain task. Results show that focusing on the fractional corpus containing domain-specialized entities and utilizing a more challenging pre-training strategy in domain-adaptive pre-training are beneficial for the NER domain adaptation, and our proposed method can consistently outperform existing cross-domain NER baselines. Nevertheless, experiments also illustrate the challenge of this cross-domain NER task. We hope that our dataset and baselines will catalyze research in the NER domain adaptation area. The code and data are available at https://github.com/zliucr/CrossNER. Zihan Liu 0001, Yan Xu 0012, Tiezheng Yu, Wenliang Dai, Ziwei Ji 0001, Samuel Cahyawijaya, Andrea Madotto, Pascale Fung |
AAAI | 7 |
| 2021 | On the Importance of Word Order Information in Cross-lingual Sequence LabelingabstractCross-lingual models trained on source language tasks possess the capability to directly transfer to target languages. However, since word order variances generally exist in different languages, cross-lingual models that overfit into the word order of the source language could have sub-optimal performance in target languages. In this paper, we hypothesize that reducing the word order information fitted into the models can improve the adaptation performance in target languages. To verify this hypothesis, we introduce several methods to make models encode less word order information of the source language and test them based on cross-lingual word embeddings and the pre-trained multilingual model. Experimental results on three sequence labeling tasks (i.e., part-of-speech tagging, named entity recognition and slot filling tasks) show that reducing word order information injected into the model can achieve better zero-shot cross-lingual performance. Further analysis illustrates that fitting excessive or insufficient word order information into the model results in inferior cross-lingual performance. Moreover, our proposed methods can also be applied to strong cross-lingual models and further improve their performance. Zihan Liu 0001, Genta Indra Winata, Samuel Cahyawijaya, Andrea Madotto, Zhaojiang Lin, Pascale Fung |
AAAI | 4 |
| 2021 | Neural Path Hunter: Reducing Hallucination in Dialogue Systems via Path GroundingabstractDialogue systems powered by large pretrained language models exhibit an innate ability to deliver fluent and natural-sounding responses.Despite their impressive performance, these models are fitful and can often generate factually incorrect statements impeding their widespread adoption.In this paper, we focus on the task of improving faithfulness and reducing hallucination of neural dialogue systems to known facts supplied by a Knowledge Graph (KG).We propose NEU-RAL PATH HUNTER which follows a generatethen-refine strategy whereby a generated response is amended using the KG.NEURAL PATH HUNTER leverages a separate tokenlevel fact critic to identify plausible sources of hallucination followed by a refinement stage that retrieves correct entities by crafting a query signal that is propagated over a k-hop subgraph.We empirically validate our proposed approach on the OpenDialKG dataset (Moon et al., 2019) against a suite of metrics and report a relative improvement of faithfulness over dialogue responses by 20.35% based on FeQA (Durmus et al., 2020).The code is available at https://github.com/ nouhadziri/Neural-Path-Hunter. Nouha Dziri, Andrea Madotto, Osmar R. Zaïane, Joey Bose |
EMNLP (1) | 2 |
| 2021 | Zero-Shot Dialogue State Tracking via Cross-Task TransferabstractZhaojiang Lin, Bing Liu, Andrea Madotto, Seungwhan Moon, Zhenpeng Zhou, Paul Crook, Zhiguang Wang, Zhou Yu, Eunjoon Cho, Rajen Subba, Pascale Fung. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Zhaojiang Lin, Andrea Madotto, Seungwhan Moon, Zhenpeng Zhou, Paul A. Crook, Zhiguang Wang, Zhou Yu 0005, Eunjoon Cho, Rajen Subba, Pascale Fung |
EMNLP (1) | 3 |
| 2021 | Continual Learning in Task-Oriented Dialogue SystemsabstractAndrea Madotto, Zhaojiang Lin, Zhenpeng Zhou, Seungwhan Moon, Paul Crook, Bing Liu, Zhou Yu, Eunjoon Cho, Pascale Fung, Zhiguang Wang. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Andrea Madotto, Zhaojiang Lin, Zhenpeng Zhou, Seungwhan Moon, Paul A. Crook, Zhou Yu 0005, Eunjoon Cho, Pascale Fung, Zhiguang Wang |
EMNLP (1) | 1 |
| 2021 | Towards Few-shot Fact-Checking via PerplexityabstractNayeon Lee, Yejin Bang, Andrea Madotto, Pascale Fung. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Nayeon Lee, Yejin Bang, Andrea Madotto, Pascale Fung |
NAACL-HLT | 3 |
| 2021 | Leveraging Slot Descriptions for Zero-Shot Cross-Domain Dialogue StateTrackingabstractZhaojiang Lin, Bing Liu, Seungwhan Moon, Paul Crook, Zhenpeng Zhou, Zhiguang Wang, Zhou Yu, Andrea Madotto, Eunjoon Cho, Rajen Subba. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Zhaojiang Lin, Seungwhan Moon, Paul A. Crook, Zhenpeng Zhou, Zhiguang Wang, Andrea Madotto, Eunjoon Cho, Rajen Subba |
NAACL-HLT | 8 |
| 2021 | Assessing Political Prudence of Open-domain ChatbotsabstractPolitically sensitive topics are still a challenge for open-domain chatbots.However, dealing with politically sensitive content in a responsible, non-partisan, and safe behavior way is integral for these chatbots.Currently, the main approach to handling political sensitivity is by simply changing such a topic when it is detected.This is safe but evasive and results in a chatbot that is less engaging.In this work, as a first step towards a politically safe chatbot, we propose a group of metrics for assessing their political prudence.We then conduct political prudence analysis of various chatbots and discuss their behavior from multiple angles through our automatic metric and human evaluation metrics.The testsets and codebase are released to promote research in this area.1 Yejin Bang, Nayeon Lee, Etsuko Ishii, Andrea Madotto, Pascale Fung |
SIGDIAL | 4 |
| 2020 | MinTL: Minimalist Transfer Learning for Task-Oriented Dialogue SystemsabstractIn this paper, we propose Minimalist Transfer Learning (MinTL) to simplify the system design process of task-oriented dialogue systems and alleviate the over-dependency on annotated data.MinTL is a simple yet effective transfer learning framework, which allows us to plug-and-play pre-trained seq2seq models, and jointly learn dialogue state tracking and dialogue response generation.Unlike previous approaches, which use a copy mechanism to "carryover" the old dialogue states to the new one, we introduce Levenshtein belief spans (Lev), that allows efficient dialogue state tracking with a minimal generation length.We instantiate our learning framework with two pretrained backbones: T5 (Raffel et al., 2019) and BART (Lewis et al., 2019), and evaluate them on MultiWOZ.Extensive experiments demonstrate that: 1) our systems establish new state-of-the-art results on end-to-end response generation, 2) MinTL-based systems are more robust than baseline methods in the low resource setting, and they achieve competitive results with only 20% training data, and 3) Lev greatly improves the inference efficiency 1 . Zhaojiang Lin, Andrea Madotto, Genta Indra Winata, Pascale Fung |
EMNLP (1) | 2 |
| 2020 | Generating Empathetic Responses by Looking Ahead the User's SentimentabstractAn important aspect of human conversation difficult for machines is conversing with empathy, which is to understand the user's emotion and respond appropriately. Recent neural conversation models that attempted to generate empathetic responses either focused on conditioning the output to a given emotion, or incorporating the current user emotional state. However, these approaches do not factor in how the user would feel towards the generated response. Hence, in this paper, we propose Sentiment Look-ahead, which is a novel perspective for empathy that models the future user emotional state. In short, Sentiment Look-ahead is a reward function under a reinforcement learning framework that provides a higher reward to the generative model when the generated utterance improves the user's sentiment. We implement and evaluate three different possible implementations of sentiment look-ahead and empirically show that our proposed approach can generate significantly more empathetic, relevant, and fluent responses than other competitive baselines such as multitask learning. Jamin Shin, Peng Xu 0008, Andrea Madotto, Pascale Fung |
ICASSP | 3 |
| 2020 | Plug and Play Language Models: A Simple Approach to Controlled Text Generation
Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, Rosanne Liu |
ICLR | 2 |
| 2020 | Exploration Based Language Learning for Text-Based GamesabstractThis work presents an exploration and imitation-learning-based agent capable of state-of-the-art performance in playing text-based computer games. These games are of interest as they can be seen as a testbed for language understanding, problem-solving, and language generation by artificial agents. Moreover, they provide a learning setting in which these skills can be acquired through interactions with an environment rather than using fixed corpora. One aspect that makes these games particularly challenging for learning agents is the combinatorially large action space. Existing methods for solving text-based games are limited to games that are either very simple or have an action space restricted to a predetermined set of admissible actions. In this work, we propose to use the exploration approach of Go-Explore for solving text-based games. More specifically, in an initial exploration phase, we first extract trajectories with high rewards, after which we train a policy to solve the game by imitating these trajectories. Our experiments show that this approach outperforms existing solutions in solving text-based games, and it is more sample efficient in terms of the number of interactions with the environment. Moreover, we show that the learned policy can generalize better than existing solutions to unseen games without using any restriction on the action space. Andrea Madotto, Mahdi Namazifar, Joost Huizinga, Piero Molino, Adrien Ecoffet, Huaixiu Zheng, Alexandros Papangelis, Dian Yu 0002, Chandra Khatri, Gökhan Tür |
IJCAI | 1 |
| 2020 | Learning Fast Adaptation on Cross-Accented Speech RecognitionabstractLocal dialects influence people to pronounce words of the same language differently from each other. The great variability and complex characteristics of accents create a major challenge for training a robust and accent-agnostic automatic speech recognition (ASR) system. In this paper, we introduce a cross-accented English speech recognition task as a benchmark for measuring the ability of the model to adapt to unseen accents using the existing CommonVoice corpus. We also propose an accent-agnostic approach that extends the model-agnostic meta-learning (MAML) algorithm for fast adaptation to unseen accents. Our approach significantly outperforms joint training in both zero-shot, few-shot, and all-shot in the mixed-region and cross-region settings in terms of word error rate. Copyright © 2020 ISCA Genta Indra Winata, Samuel Cahyawijaya, Zihan Liu 0001, Zhaojiang Lin, Andrea Madotto, Peng Xu 0008, Pascale Fung |
INTERSPEECH | 5 |
| 2020 | Getting To Know You: User Attribute Extraction from DialoguesabstractUser attributes provide rich and useful information for user understanding, yet structured and easy-to-use attributes are often sparsely populated. In this paper, we leverage dialogues with conversational agents, which contain strong suggestions of user information, to automatically extract user attributes. Since no existing dataset is available for this purpose, we apply distant supervision to train our proposed two-stage attribute extractor, which surpasses several retrieval and generation baselines on human evaluation. Meanwhile, we discuss potential applications (e.g., personalized recommendation and dialogue systems) of such extracted user attributes, and point out current limitations to cast light on future work. Chien-Sheng Wu, Andrea Madotto, Zhaojiang Lin, Peng Xu 0008, Pascale Fung |
LREC | 2 |
| 2019 | Personalizing Dialogue Agents via Meta-LearningabstractExisting personalized dialogue models use human designed persona descriptions to improve dialogue consistency.Collecting such descriptions from existing dialogues is expensive and requires hand-crafted feature designs.In this paper, we propose to extend Model-Agnostic Meta-Learning (MAML) (Finn et al., 2017) to personalized dialogue learning without using any persona descriptions.Our model learns to quickly adapt to new personas by leveraging only a few dialogue samples collected from the same user, which is fundamentally different from conditioning the response on the persona descriptions.Empirical results on Persona-chat dataset (Zhang et al., 2018) indicate that our solution outperforms non-metalearning baselines using automatic evaluation metrics, and in terms of human-evaluated fluency and consistency. Andrea Madotto, Zhaojiang Lin, Chien-Sheng Wu, Pascale Fung |
ACL (1) | 1 |
| 2019 | Transferable Multi-Domain State Generator for Task-Oriented Dialogue SystemsabstractOver-dependence on domain ontology and lack of knowledge sharing across domains are two practical and yet less studied problems of dialogue state tracking.Existing approaches generally fall short in tracking unknown slot values during inference and often have difficulties in adapting to new domains.In this paper, we propose a TRAnsferable Dialogue statE generator (TRADE) that generates dialogue states from utterances using a copy mechanism, facilitating knowledge transfer when predicting (domain, slot, value) triplets not encountered during training.Our model is composed of an utterance encoder, a slot gate, and a state generator, which are shared across domains.Empirical results demonstrate that TRADE achieves state-of-the-art joint goal accuracy of 48.62% for the five domains of Mul-tiWOZ, a human-human dialogue dataset.In addition, we show its transferring ability by simulating zero-shot and few-shot dialogue state tracking for unseen domains.TRADE achieves 60.58% joint goal accuracy in one of the zero-shot domains, and is able to adapt to few-shot cases without forgetting already trained domains.* Work partially done while the first author was an intern at Salesforce Research.Usr: I am looking for a cheap restaurant in the centre of the Chien-Sheng Wu, Andrea Madotto, Ehsan Hosseini-Asl, Caiming Xiong, Richard Socher, Pascale Fung |
ACL (1) | 2 |
| 2019 | Code-Switched Language Models Using Neural Based Synthetic Data from Parallel SentencesabstractTraining code-switched language models is difficult due to lack of data and complexity in the grammatical structure.Linguistic constraint theories have been used for decades to generate artificial code-switching sentences to cope with this issue.However, this require external word alignments or constituency parsers that create erroneous results on distant languages.We propose a sequence-to-sequence model using a copy mechanism to generate code-switching data by leveraging parallel monolingual translations from a limited source of code-switching data.The model learns how to combine words from parallel sentences and identifies when to switch one language to the other.Moreover, it captures code-switching constraints by attending and aligning the words in inputs, without requiring any external knowledge.Based on experimental results, the language model trained with the generated sentences achieves state-of-theart performance and improves end-to-end automatic speech recognition. Genta Indra Winata, Andrea Madotto, Chien-Sheng Wu, Pascale Fung |
CoNLL | 2 |
| 2019 | MoEL: Mixture of Empathetic ListenersabstractZhaojiang Lin, Andrea Madotto, Jamin Shin, Peng Xu, Pascale Fung. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Zhaojiang Lin, Andrea Madotto, Jamin Shin, Peng Xu 0008, Pascale Fung |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Zero-shot Cross-lingual Dialogue Systems with Transferable Latent VariablesabstractZihan Liu, Jamin Shin, Yan Xu, Genta Indra Winata, Peng Xu, Andrea Madotto, Pascale Fung. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Zihan Liu 0001, Jamin Shin, Yan Xu 0012, Genta Indra Winata, Peng Xu 0008, Andrea Madotto, Pascale Fung |
EMNLP/IJCNLP (1) | 6 |
| 2019 | Clickbait? Sensational Headline Generation with Auto-tuned Reinforcement LearningabstractPeng Xu, Chien-Sheng Wu, Andrea Madotto, Pascale Fung. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Peng Xu 0008, Chien-Sheng Wu, Andrea Madotto, Pascale Fung |
EMNLP/IJCNLP (1) | 3 |
| 2018 | Mem2Seq: Effectively Incorporating Knowledge Bases into End-to-End Task-Oriented Dialog SystemsabstractEnd-to-end task-oriented dialog systems usually suffer from the challenge of incorporating knowledge bases.In this paper, we propose a novel yet simple end-toend differentiable model called memoryto-sequence (Mem2Seq) to address this issue.Mem2Seq is the first neural generative model that combines the multihop attention over memories with the idea of pointer network.We empirically show how Mem2Seq controls each generation step, and how its multi-hop attention mechanism helps in learning correlations between memories.In addition, our model is quite general without complicated taskspecific designs.As a result, we show that Mem2Seq can be trained faster and attain the state-of-the-art performance on three different task-oriented dialog datasets. Andrea Madotto, Chien-Sheng Wu, Pascale Fung |
ACL (1) | 1 |
| 2018 | End-to-End Dynamic Query Memory Network for Entity-Value Independent Task-Oriented DialogabstractIn this paper, we propose an end-to-end Dynamic Query Memory Network (DQMemNN) with a delexicalization mechanism for task-oriented dialog systems. The added dynamic component enables memory networks to capture the dialog's sequential dependencies by using a context-based query. Besides, the delexicalization mechanism reduces learning complexity and it alleviates the out-of-vocabulary entity problems. Experiments show that DQMemNN outperforms original end-to-end memory network models on bAbI full-dialog task by 3.1 % per-response and 39.3% per-dialog accuracy. In addition, the proposed framework achieves a promising average per-response accuracy of 99.7% and per-dialog accuracy of 97.8% without hand-crafted rules and features. Chien-Sheng Wu, Andrea Madotto, Genta Indra Winata, Pascale Fung |
ICASSP | 2 |