EDBT 2026 Demo / reviewers in the wild / expert
Antoine Bosselut
dblp:184/3742
· DBLP profile ↗
66ranked-venue papers
5as first author
53since 2021 · last 2026
0000-0001-8968-9649ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 62 · 5 first-author · 49 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM PretrainingabstractLarge language models (LLMs) learn nontrivial abstractions during pretraining, such as detecting irregular plural noun subjects.However, because traditional evaluation methods (e.g., benchmarking) fail to reveal how models acquire these concepts and capabilities, it is not well understood when and how these specific linguistic abilities emerge.To bridge this gap and better understand model training at the concept level, we use sparse crosscoders to discover and align features across model checkpoints.Using this approach, we track the evolution of linguistic features during pretraining.We train crosscoders between opensourced checkpoint triplets with significant performance and representation shifts, and introduce a novel metric, Relative Indirect Effects (RELIE), to trace training stages at which individual features become causally important for task performance.We show that crosscoders can detect feature emergence, maintenance, and discontinuation during pretraining.Our approach is architecture-agnostic and scalable, offering a promising path toward more interpretable and fine-grained analysis of representation learning throughout pretraining.1 Deniz Bayazit, Aaron Mueller, Antoine Bosselut |
ACL (1) | 3 |
| 2026 | Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in TokenizationabstractNegar Foroutan, Clara Meister, Debjit Paul, Joel Niklaus, Sina Ahmadi, Antoine Bosselut, Rico Sennrich. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Negar Foroutan Eghlidi, Clara Meister, Debjit Paul, Joel Niklaus, Sina Ahmadi, Antoine Bosselut, Rico Sennrich |
ACL (1) | 6 |
| 2026 | Apertus: Democratizing Open and Compliant LLMs for Global Language EnvironmentsabstractAlejandro Hernández-Cano, Alexander Hägele, Allen Hao Huang, Angelika Romanou, Antoni-Joan Solergibert, Barna Pásztor, Bettina Messmer, Dhia Garbaya, Eduard Frank Ďurech, Ido Hakimi, Juan Garcia Giraldo, Mete Ismayilzada, Negar Foroutan, Skander Moalla, Tiancheng Chen, Vinko Sabolčec, Yixuan Xu, Michael Aerni, Badr AlKhamissi, Inés Altemir Marinas, Mohammad Hossein Amani, Matin Ansaripour, Ilia Badanin, Harold Benoit, Emanuela Boros, Nicholas John Browning, Fabian Bösch, Maximilian Böther, Niklas Canova, Camille Challier, Clément Charmillot, Jonathan Coles, Jan Milan Deriu, Arnout Devos, Lukas Drescher, Daniil Dzenhaliou, Maud Ehrmann, Dongyang Fan, Simin Fan, Silin Gao, Miguel Gila, María Grandury, Diba Hashemi, Alexander Miserlis Hoyle, Jiaming Jiang, Mark Klein, Andrei Kucharavy, Anastasiia Kucherenko, Frederike Lübeck, Roman Machacek, Theofilos Ioannis Manitaras, Andreas Marfurt, Kyle Matoba, Simon Matrenok, Henrique Mendonça, Fawzi Roberto Mohamed, Syrielle Montariol, Luca Mouchel, Sven Najem-Meyer, Jingwei Ni, Gennaro Oliva, Matteo Pagliardini, Elia Palme, Andrei Panferov, Léo Paoletti, Marco Passerini, Ivan Pavlov, Auguste Poiroux, Kaustubh Ponkshe, Nathan Ranchin, Javier Rando, Mathieu Sauser, Jakhongir Saydaliev, Mukhammadali Sayfiddinov, Marian Schneider, Stefano Schuppli, Marco Scialanga, Andrei Semenov, Kumar Shridhar, Raghav Singhal, Anna Sotnikova, Alexander Sternfeld, Ayush Kumar Tarun, Paul Teiletche, Jannis Vamvas, Xiaozhe Yao, Hao Zhao, Alexander Ilic, Ana Klimovic, Andreas Krause, Caglar Gulcehre, David Rosenthal, Elliott Ash, Florian Tramèr, Joost VandeVondele, Livio Veraldi, Martin Rajman, Thomas C. Schulthess, Torsten Hoefler, Antoine Bosselut, Martin Jaggi, Imanol Schlag. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Alejandro Hernández-Cano, Alexander Hägele, Allen Hao Huang, Angelika Romanou, Antoni-Joan Solergibert i Llaquet, Barna Pásztor, Bettina Messmer, Dhia Garbaya, Eduard Durech, Ido Hakimi, Juan Garcia Giraldo, Mete Ismayilzada, Negar Foroutan Eghlidi, Skander Moalla, Tiancheng Chen, Vinko Sabolcec, Yixuan Even Xu, Michael Aerni, Badr AlKhamissi, Ines Altemir Marinas, Mohammad Hossein Amani, Matin Ansaripour, Ilia Badanin, Harold Benoit, Emanuela Boros, Nicholas John Browning, Fabian Bösch, Maximilian Böther, Niklas Canova, Camille Challier, Clément Charmillot, Jonathan Coles, Jan Deriu, Arnout Devos, Lukas Drescher, Daniil Dzenhaliou, Maud Ehrmann, Dongyang Fan, Simin Fan, Silin Gao, Miguel Gila, María Grandury, Diba Hashemi, Alexander Miserlis Hoyle, Jiaming Jiang, Mark Klein 0002, Andrei Kucharavy, Anastasiia Kucherenko, Frederike Lübeck, Roman Machacek, Theofilos Ioannis Manitaras, Andreas Marfurt, Kyle Matoba, Simon Matrenok, Henrique Mendonça, Fawzi Roberto Mohamed, Syrielle Montariol, Luca Mouchel, Sven Najem-Meyer, Jingwei Ni, Gennaro Oliva, Matteo Pagliardini, Elia Palme, Andrei Panferov, Léo Paoletti, Marco Passerini, Ivan Pavlov, Auguste Poiroux, Kaustubh Ponkshe, Nathan Ranchin, Javier Rando, Mathieu Sauser, Jakhongir Saydaliev, Mukhammadali Sayfiddinov, Marian Schneider, Stefano Schuppli, Marco Scialanga, Andrei Semenov, Kumar Shridhar, Raghav Singhal, Anna Sotnikova, Alexander Sternfeld, Ayush K. Tarun, Paul Teiletche, Jannis Vamvas, Xiaozhe Yao, Alexander Ilic, Ana Klimovic, Andreas Krause 0001, Caglar Gulcehre, David Rosenthal, Elliott Ash, Florian Tramèr, Joost VandeVondele, Livio Veraldi, Martin Rajman, Thomas C. Schulthess, Torsten Hoefler, Antoine Bosselut, Martin Jaggi, Imanol Schlag |
ACL (1) | 100 |
| 2026 | AI meets Mathematics Education: Supporting Instructors in Large Mathematics Classes with Context-Aware AIabstractLarge-enrollment university courses face persistent challenges in providing timely and scalable instructional support. While generative AI holds promise, its effective use depends on reliability and pedagogical alignment. We present a human-centered case study of AI-assisted support in a Calculus I course, implemented in close collaboration with the course instructor. We developed a system to answer students’ questions on a discussion forum, fine-tuning a lightweight language model on 2,588 historical student–instructor interactions. The model achieved 75.3% accuracy on a benchmark of 150 representative questions annotated by five instructors, and in 36% of cases, its responses were rated equal to or better than instructor answers. Post-deployment student survey (N = 105) indicated that students valued the alignment of the responses with the course materials and their immediate availability, while still relying on the instructor verification for trust. We highlight the importance of hybrid human–AI workflows for safe and effective course support. Jérémy Valentin Barghorn, Anna Sotnikova, Sacha Friedli, Antoine Bosselut |
CHI | 4 |
| 2025 | Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual EvaluationabstractShivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, Raymond Ng, Shayne Longpre, Sebastian Ruder, Wei-Yin Ko, Antoine Bosselut, Alice Oh, Andre Martins, Leshem Choshen, Daphne Ippolito, Enzo Ferrante, Marzieh Fadaee, Beyza Ermis, Sara Hooker. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, Raymond Ng, Shayne Longpre, Sebastian Ruder, Wei-Yin Ko, Antoine Bosselut, Alice Oh, André F. T. Martins, Leshem Choshen, Daphne Ippolito, Enzo Ferrante, Marzieh Fadaee, Beyza Ermis, Sara Hooker |
ACL (1) | 15 |
| 2025 | Efficient Tool Use with Chain-of-Abstraction ReasoningabstractTo achieve faithful reasoning that aligns with human expectations, large language models (LLMs) need to ground their reasoning to real-world knowledge (e.g., web facts, math and physical rules). Tools help LLMs access this external knowledge, but there remains challenges for fine-tuning LLM agents (e.g., Toolformer) to invoke tools in multi-step reasoning problems, where inter-connected tool calls require holistic and efficient tool usage planning. In this work, we propose a new method for LLMs to better leverage tools in multi-step reasoning. Our method, Chain-of-Abstraction (CoA), trains LLMs to first decode reasoning chains with abstract placeholders, and then call domain tools to reify each reasoning chain by filling in specific knowledge. This planning with abstract chains enables LLMs to learn more general reasoning strategies, which are robust to shifts of domain knowledge (e.g., math results) relevant to different reasoning questions. It also allows LLMs to perform decoding and calling of external tools in parallel, which avoids the inference delay caused by waiting for tool responses. In mathematical reasoning and Wiki QA domains, we show that our method consistently outperforms previous chain-of-thought and tool-augmented baselines on both in-distribution and out-of-distribution test sets, with an average ~6% absolute QA accuracy improvement. LLM agents trained with our method also show more efficient tool use, with inference speed being on average ~1.4x faster than baseline tool-augmented LLMs. Silin Gao, Jane Dwivedi-Yu, Xiaoqing Ellen Tan, Ramakanth Pasunuru, Olga Golovneva, Koustuv Sinha, Asli Celikyilmaz, Antoine Bosselut |
COLING | 9 |
| 2025 | VinaBench: Benchmark for Faithful and Consistent Visual NarrativesabstractVisual narrative generation transforms textual narratives into sequences of images illustrating the content of the text. However, generating visual narratives that are faithful to the input text and self-consistent across generated images remains an open challenge, due to the lack of knowledge constraints used for planning the stories. In this work, we propose a new benchmark, VinaBench, to address this challenge. Our benchmark annotates the underlying commonsense and discourse constraints in visual narrative samples, offering systematic scaffolds for learning the implicit strategies of visual storytelling. Based on the incorporated narrative constraints, we further propose novel metrics to closely evaluate the consistency of generated narrative images and the alignment of generations with the input textual narrative. Our results across three generative vision models demonstrate that learning with VinaBench’s knowledge constraints effectively improves the faithfulness and cohesion of generated visual narratives.1 Silin Gao, Sheryl Mathew, Li Mi, Sepideh Mamooler, Hiromi Wakaki, Yuki Mitsufuji, Syrielle Montariol, Antoine Bosselut |
CVPR | 9 |
| 2025 | From Language to Cognition: How LLMs Outgrow the Human Language NetworkabstractLarge language models (LLMs) exhibit remarkable similarity to neural activity in the human language network. However, the key properties of language underlying this alignment—and how brain-like representations emerge and change across training—remain unclear. We here benchmark 34 training checkpoints spanning 300B tokens across 8 different model sizes to analyze how brain alignment relates to linguistic competence. Specifically, we find that brain alignment tracks the development of formal linguistic competence—i.e., knowledge of linguistic rules—more closely than functional linguistic competence. While functional competence, which involves world knowledge and reasoning, continues to develop throughout training, its relationship with brain alignment is weaker, suggesting that the human language network primarily encodes formal linguistic structure rather than broader cognitive functions. Notably, we find that the correlation between next-word prediction, behavioral alignment, and brain alignment fades once models surpass human language proficiency. We further show that model size is not a reliable predictor of brain alignment when controlling for the number of features. Finally, using the largest set of rigorous neural language benchmarks to date, we show that language brain alignment benchmarks remain unsaturated, highlighting opportunities for improving future models. Taken together, our findings suggest that the human language network is best modeled by formal, rather than functional, aspects of language. Badr AlKhamissi, Greta Tuckute, Yingtian Tang, Taha Binhuraib, Antoine Bosselut, Martin Schrimpf |
EMNLP | 5 |
| 2025 | CAVE : Detecting and Explaining Commonsense Anomalies in Visual EnvironmentsabstractCAVE: Commonsense Anomalies in Visual Environment 🏠 Project Page📄 Paper (EMNLP 2025)💻 Code Dataset Details Dataset Description CAVE is the first benchmark of real-world visual anomalies for evaluating Vision-Language Models (VLMs). It is curated from images captured in real-life settings (photographs and screenshots taken by individuals), sourced from Reddit. The benchmark is grounded in cognitive science literature on how humans detect and resolve anomalies. Each image is annotated with rich, multi-task annotations that support three open-ended tasks (anomaly description, explanation, and justification), one visual grounding task (anomaly localization via bounding boxes), and classification along four dimensions (anomaly category, severity, surprisal, and complexity) that characterize the anomaly. CAVE reveals that state-of-the-art VLMs struggle substantially with visual anomaly perception and commonsense reasoning: the best model (GPT-4o) achieves only ~57% F1-score on anomaly detection even with advanced prompting strategies. Curated by: Rishika Bhagwatkar, Syrielle Montariol, Angelika Romanou, Beatriz Borges, Irina Rish, Antoine Bosselut Affiliations: EPFL, MILA Language: English License: CC-BY-4.0 Published at: EMNLP 2025 (Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing) Dataset Sources Project Page: https://smontariol.github.io/cave-visual-anomalies/ Paper: CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments Contact: [email protected], [email protected] Uses Intended Uses CAVE is designed to evaluate VLMs on their ability to: Detect real-world commonsense anomalies in images (anomaly description). Explain why a detected situation is anomalous (anomaly explanation). Justify how an anomaly might have occurred (anomaly justification). Localize anomalies within images via bounding boxes (anomaly localization). Classify anomalies by their visual manifestation type and numerical features (severity, surprisal, complexity). It also serves as a resource for studying the alignment between human and machine processing of visual anomalies, and for developing improved prompting strategies or fine-tuning approaches for anomaly-related tasks. Out-of-Scope Uses CAVE is a benchmark for evaluation purposes. Its small size (361 images) makes it unsuitable as a training set. It should not be used to deploy anomaly detection systems in safety-critical settings without additional validation. Dataset Structure Overview CAVE consists of 361 images: 309 anomalous and 52 normal (non-anomalous) images. Anomalous images contain up to 3 anomalies each, totaling 334 annotated anomalies. Each anomaly is paired with a unique bounding box. Annotation Fields Each sample includes the following fields: Field Description image The image (photograph or screenshot) image_description Short description of the image content (without describing the anomaly) anomaly_description Textual description of what is anomalous in the image anomaly_explanation Explanation of why the situation is anomalous (commonsense reasoning) anomaly_justification Plausible explanation of how the anomaly might have occurred anomaly_category Category of the anomaly's visual manifestation (see taxonomy below) bounding_box Coordinates of the bounding box demarcating the anomalous region severity 1–5 score: does the anomaly require immediate action? surprisal 1–5 score: how much does the situation deviate from expectations? complexity 1–5 score: how hard is the anomaly to detect? Anomaly Category Taxonomy Anomalies are categorized by how they visually manifest, inspired by MMBench's taxonomy of visual reasoning types: Category Description Example Entity Presence An object is present when it shouldn't be A black bear in an industrial building Entity Absence An expected object is missing A person using a cutter without protective gear Entity Attribute An object has an anomalous attribute (color, shape, label, orientation, usage) A snack packet opened from the wrong side Spatial Relation An object is incorrectly positioned relative to another Furniture blocking an emergency button Uniformity Breach A disruption in an expected uniform/symmetrical pattern One tile with a different orientation Textual Anomaly Text in the image conveys an unexpected or contradictory message A "KEEP RIGHT" sign with an arrow pointing left Dataset Creation Images were collected from four Reddit subreddits that specialize in content featuring unusual or uncommon situations: r/ocdtriggers r/mildlyconfusing r/mildlyinfuriating r/OSHA The top 1,000 posts from each subreddit were downloaded using the PRAW library. Images were filtered through both automatic and manual processes to remove: Unclear or ambiguous content Non-realistic images NSFW or sensitive content Images with text annotations, circles, or other overlaid marks Images below icon resolution Annotation proceeded in two rounds, with Amazon Mechanical Turk followed by Expert Verification & Consolidation, with 3 independent raters per anomaly for severity, surprisal, and complexity scores. Citation @inproceedings{bhagwatkar-etal-2025-cave, title = "{CAVE} : Detecting and Explaining Commonsense Anomalies in Visual Environments", author = "Bhagwatkar, Rishika and Montariol, Syrielle and Romanou, Angelika and Borges, Beatriz and Rish, Irina and Bosselut, Antoine", booktitle = "Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing", month = nov, year = "2025", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2025.emnlp-main.1379/", doi = "10.18653/v1/2025.emnlp-main.1379", pages = "27110--27151", } Acknowledgements The authors acknowledge support from Canada CIFAR AI Chair Program, Canada Excellence Research Chairs Program, Swiss National Science Foundation (No. 215390), Innosuisse (PFFS-21-29), EPFL Center for Imaging, Sony Group Corporation, and a Meta LLM Evaluation Research Grant. Computational resources were provided by MILA - Quebec AI Institute. Rishika Bhagwatkar, Syrielle Montariol, Angelika Romanou, Beatriz Borges, Irina Rish, Antoine Bosselut |
EMNLP | 6 |
| 2025 | Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document EmbeddingsabstractA limitation of modern document retrieval embedding methods is that they typically encode passages (chunks) from the same documents independently, often overlooking crucial contextual information from the rest of the document that could greatly improve individual chunk representations.In this work, we introduce ConTEB (Contextaware Text Embedding Benchmark), a benchmark designed to evaluate retrieval models on their ability to leverage document-wide context.Our results show that state-of-the-art embedding models struggle in retrieval scenarios where context is required.To address this limitation, we propose InSeNT (In-sequence Negative Training), a novel contrastive posttraining approach which combined with late chunking pooling enhances contextual representation learning while preserving computational efficiency.Our method significantly improves retrieval quality on ConTEB without sacrificing base model performance.We further find chunks embedded with our method are more robust to suboptimal chunking strategies and larger retrieval corpus sizes.We opensource all artifacts at https://github.com/ illuin-tech/contextual-embeddings. Max Conti, Manuel Faysse, Gautier Viaud, Antoine Bosselut, Céline Hudelot, Pierre Colombo |
EMNLP | 4 |
| 2025 | Reliable Evaluation and Benchmarks for Statement AutoformalizationabstractEvaluating statement autoformalization, translating natural language mathematics into formal languages like Lean 4, remains a significant challenge, with few metrics, datasets, and standards to robustly measure progress.In this work, we present a comprehensive approach combining improved metrics, robust benchmarks, and systematic evaluation, to fill this gap.First, we introduce BEq+, an automated metric that correlates strongly with human judgment, along with ProofNetVerif, a new dataset for assessing the quality of evaluation metrics, containing 3,752 annotated examples.Second, we develop two new autoformalization benchmarks: ProofNet#, a corrected version of ProofNet, and RLM25, with 619 new pairs of research-level mathematics from six formalization projects.Through systematic experimentation across these benchmarks, we find that current techniques can achieve up to 45.1% accuracy on undergraduate mathematics but struggle with research-level content without proper context.Our work establishes a reliable foundation for evaluating and advancing autoformalization systems. Auguste Poiroux, Gail Weiss, Viktor Kuncak, Antoine Bosselut |
EMNLP | 4 |
| 2025 | GeoExplorer: Active Geo-Localization with Curiosity-Driven ExplorationabstractActive Geo-localization (AGL) is the task of localizing a goal, represented in various modalities (e.g., aerial images, ground-level images, or text), within a predefined search area. Current methods approach AGL as a goal-reaching reinforcement learning (RL) problem with a distance-based reward. They localize the goal by implicitly learning to minimize the relative distance from it. However, when distance estimation becomes challenging or when encountering unseen targets and environments, the agent exhibits reduced robustness and generalization ability due to the less reliable exploration strategy learned during training. In this paper, we propose GeoExplorer, an AGL agent that incorporates curiosity-driven exploration through intrinsic rewards. Unlike distance-based rewards, our curiosity-driven reward is goal-agnostic, enabling robust, diverse, and contextually relevant exploration based on effective environment modeling. These capabilities have been proven through extensive experiments across four AGL benchmarks, demonstrating the effectiveness and generalization ability of GeoExplorer in diverse settings, particularly in localizing unfamiliar targets and environments. Li Mi, Manon Béchaz, Zeming Chen 0001, Antoine Bosselut, Devis Tuia |
ICCV | 4 |
| 2025 | INCLUDE: Evaluating Multilingual Language Understanding with Regional KnowledgeabstractThe performance differential of large language models (LLM) between languages hinders their effective deployment in many regions, inhibiting the potential economic and societal value of generative AI tools in many communities. However, the development of functional LLMs in many languages (i.e., multilingual LLMs) is bottlenecked by the lack of high-quality evaluation resources in languages other than English. Moreover, current practices in multilingual benchmark construction often translate English resources, ignoring the regional and cultural knowledge of the environments in which multilingual systems would be used. In this work, we construct an evaluation suite of 197,243 QA pairs from local exam sources to measure the capabilities of multilingual LLMs in a variety of regional contexts.
Our novel resource, INCLUDE, is a comprehensive knowledge- and reasoning-centric benchmark across 44 written languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed. Angelika Romanou, Negar Foroutan Eghlidi, Anna Sotnikova, Zeming Chen 0001, Sree Harsha Nelaturu, Shivalika Singh, Rishabh Maheshwary, Micol Altomare, Mohamed A. Haggag, Imanol Schlag, Marzieh Fadaee, Sara Hooker, Antoine Bosselut, Snegha A, Alfonso Amayuelas, Azril Hafizi Amirudin, Viraat Aryabumi, Danylo Boiko, Jenny Chim, Gal Cohen, Aditya Kumar Dalmia, Abraham Diress, Sharad Duwal, Daniil Dzenhaliou, Daniel Fernando Erazo Florez, Fabian Farestam, Joseph Marvin Imperial, Shayekh Bin Islam, Perttu Isotalo, Maral Jabbarishiviari, Börje Karlsson 0001, Eldar Khalilov, Christopher Klamm, Fajri Koto, Dominik Krzeminski, Gabriel Adriano de Melo, Syrielle Montariol, Yiyang Nan, Joel Niklaus, Jekaterina Novikova, Johan S. Obando-Ceron, Debjit Paul, Esther Ploeger, Jebish Purbey, Swati Rajwal, Selvan Sunitha Ravi, Sara Rydell, Roshan Santhosh, Drishti Sharma, Marjana Prifti Skenduli, Arshia Soltani Moakhar, Bardia Soltani Moakhar, Ran Tamir, Ayush K. Tarun, Azmine Toushik Wasi, Thenuka Ovin Weerasinghe, Serhan Yilmaz, Mike Zhang |
ICLR | 13 |
| 2025 | The LLM Language Network: A Neuroscientific Approach for Identifying Causally Task-Relevant UnitsabstractBadr AlKhamissi, Greta Tuckute, Antoine Bosselut, Martin Schrimpf. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Badr AlKhamissi, Greta Tuckute, Antoine Bosselut, Martin Schrimpf |
NAACL (Long Papers) | 3 |
| 2025 | Evaluating Morphological Compositional Generalization in Large Language ModelsabstractMete Ismayilzada, Defne Circi, Jonne Sälevä, Hale Sirin, Abdullatif Köksal, Bhuwan Dhingra, Antoine Bosselut, Duygu Ataman, Lonneke Van Der Plas. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Mete Ismayilzada, Defne Circi, Jonne Sälevä, Hale Sirin, Abdullatif Köksal, Bhuwan Dhingra, Antoine Bosselut, Duygu Ataman, Lonneke van der Plas |
NAACL (Long Papers) | 7 |
| 2025 | PICLe: Pseudo-annotations for In-Context Learning in Low-Resource Named Entity DetectionabstractSepideh Mamooler, Syrielle Montariol, Alexander Mathis, Antoine Bosselut. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Sepideh Mamooler, Syrielle Montariol, Alexander Mathis, Antoine Bosselut |
NAACL (Long Papers) | 4 |
| 2025 | A Logical Fallacy-Informed Framework for Argument GenerationabstractLuca Mouchel, Debjit Paul, Shaobo Cui, Robert West, Antoine Bosselut, Boi Faltings. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Luca Mouchel, Debjit Paul, Shaobo Cui 0006, Robert West 0001, Antoine Bosselut, Boi Faltings |
NAACL (Long Papers) | 5 |
| 2025 | Measuring what Matters: Construct Validity in Large Language Model BenchmarksabstractEvaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstract and complex phenomena such as safety' androbustness' requires strong construct validity, that is, having measures that represent what matters to the phenomenon. With a team of 29 expert reviewers, we conduct a systematic review of 445 LLM benchmarks from leading conferences in natural language processing and machine learning. Across the reviewed articles, we find patterns related to the measured phenomena, tasks, and scoring metrics which undermine the validity of the resulting claims. To address these shortcomings, we provide eight key recommendations and detailed actionable guidance to researchers and practitioners in developing LLM benchmarks. Andrew M. Bean 0001, Ryan Othniel Kearns, Angelika Romanou, Franziska Sofia Hafner, Harry Mayne, Jan Batzner, Negar Foroutan Eghlidi, Chris Schmitz, Karolina Korgul, Hunar Batra, Oishi Deb, Emma Beharry, Cornelius Emde, Thomas Foster, Anna Gausen, María Grandury, Simeng Han, Valentin Hofmann, Lujain Ibrahim, Hazel Kim, Hannah Kirk, Fangru Lin, Gabrielle K. Liu, Lennart Luettgau, Jabez Magomere, Jonathan Rystrøm, Anna Sotnikova, Yilun Zhao 0001, Adel Bibi, Antoine Bosselut, Ronald Clark, Arman Cohan, Jakob N. Foerster, Yarin Gal, Scott A. Hale, Inioluwa Deborah Raji, Christopher Summerfield, Philip Torr 0001, Cozmin Ududec, Luc Rocher, Adam Mahdi |
NeurIPS | 31 |
| 2025 | For Better or for Worse, Transformers Seek Patterns for MemorizationabstractMemorization in language models is a critical yet poorly understood phenomenon. In this work, we investigate memorization in transformer-based language models by analyzing their memorization dynamics during training over multiple epochs. We find that memorization is neither a constant accumulation of sequences nor simply dictated by the recency of exposure to these sequences. Instead, much like generalization, memorization appears to be driven by pattern recognition. Tracking memorization dynamics in mixed datasets, we observe that models memorize different sub-datasets in distinct bursts, suggesting that each subset is associated with unique underlying patterns, and that the model prefers to learn these patterns in a consistent order. We also find that easily learnable patterns tend to support generalization on unseen data, while more complex patterns do not. Furthermore, in datasets with weak or absent patterns, larger models may delay memorization relative to smaller ones, a behavior we term $\textit{overthinking}$. Our results show that the subset of sequences memorized by a model over time is not arbitrary, and give insights into the internal processes a model goes through during training. Our code is available at: https://github.com/mdrpanwar/memorization-patterns. Madhur Panwar, Gail Weiss, Navin Goyal, Antoine Bosselut |
NeurIPS | 4 |
| 2025 | Positional Fragility in LLMs: How Offset Effects Reshape Our Understanding of Memorization RisksabstractLarge language models are known to memorize parts of their training data, posing risk of copyright violations. To systematically examine this risk, we pretrain language models (1B/3B/8B) from scratch on 83B tokens, mixing web-scale data with public domain books used to simulate copyrighted content at controlled frequencies at lengths at least ten times longer than prior work. We thereby identified the offset effect, a phenomenon characterized by two key findings: (1) verbatim memorization is most strongly triggered by short prefixes drawn from the beginning of the context window, with memorization decreasing counterintuitively as prefix length increases; and (2) a sharp decline in verbatim recall when prefix begins offset from the initial tokens of the context window. We attribute this to positional fragility: models rely disproportionately on the earliest tokens in their context window as retrieval anchors, making them sensitive to even slight shifts. We further observe that when the model fails to retrieve memorized content, it often produces degenerated text. Leveraging these findings, we show that shifting sensitive data deeper into the context window suppresses both extractable memorization and degeneration. Our results suggest that positional offset is a critical and previously overlooked axis for evaluating memorization risks, since prior work implicitly assumed uniformity by probing only from the beginning of documents or training sequences. Yixuan Even Xu, Antoine Bosselut, Imanol Schlag |
NeurIPS | 2 |
| 2024 | ConVQG: Contrastive Visual Question Generation with Multimodal GuidanceabstractAsking questions about visual environments is a crucial way for intelligent agents to understand rich multi-faceted scenes, raising the importance of Visual Question Generation (VQG) systems. Apart from being grounded to the image, existing VQG systems can use textual constraints, such as expected answers or knowledge triplets, to generate focused questions. These constraints allow VQG systems to specify the question content or leverage external commonsense knowledge that can not be obtained from the image content only. However, generating focused questions using textual constraints while enforcing a high relevance to the image content remains a challenge, as VQG systems often ignore one or both forms of grounding. In this work, we propose Contrastive Visual Question Generation (ConVQG), a method using a dual contrastive objective to discriminate questions generated using both modalities from those based on a single one. Experiments on both knowledge-aware and standard VQG benchmarks demonstrate that ConVQG outperforms the state-of-the-art methods and generates image-grounded, text-guided, and knowledge-rich questions. Our human evaluation results also show preference for ConVQG questions compared to non-contrastive baselines. Li Mi, Syrielle Montariol, Javiera Castillo-Navarro, Xianjie Dai, Antoine Bosselut, Devis Tuia |
AAAI | 5 |
| 2024 | Complex Reasoning over Logical Queries on Commonsense Knowledge GraphsabstractEvent commonsense reasoning requires the ability to reason about the relationship between events, as well as infer implicit context underlying that relationship.However, data scarcity makes it challenging for language models to learn to generate commonsense inferences for contexts and questions involving interactions between complex events.To address this demand, we present COM 2 (COMplex COMmonsense), a new dataset created by sampling multi-hop logical queries (e.g., the joint effect or cause of both event A and B, or the effect of the effect of event C) from an existing commonsense knowledge graph (CSKG), and verbalizing them using handcrafted rules and large language models into multiple-choice and text generation questions.Our experiments show that language models trained on COM 2 exhibit significant improvements in complex reasoning ability, resulting in enhanced zero-shot performance in both indomain and out-of-domain tasks for question answering and generative commonsense reasoning, without expensive human annotations. 1* Work done during internship at EPFL. 1 Code and data are available at https://github.com/ tqfang/complex-commonsense-reasoning Tianqing Fang, Zeming Chen 0001, Yangqiu Song, Antoine Bosselut |
ACL (1) | 4 |
| 2024 | DiffuCOMET: Contextual Commonsense Knowledge DiffusionabstractSilin Gao, Mete Ismayilzada, Mengjie Zhao, Hiromi Wakaki, Yuki Mitsufuji, Antoine Bosselut. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Silin Gao, Mete Ismayilzada, Hiromi Wakaki, Yuki Mitsufuji, Antoine Bosselut |
ACL (1) | 6 |
| 2024 | A Design Space for Intelligent and Interactive Writing AssistantsabstractIn our era of rapid technological advancement, the research landscape for writing assistants has become increasingly fragmented across various research communities. We seek to address this challenge by proposing a design space as a structured way to examine and explore the multidimensional space of intelligent and interactive writing assistants. Through community collaboration, we explore five aspects of writing assistants: task, user, technology, interaction, and ecosystem. Within each aspect, we define dimensions and codes by systematically reviewing 115 papers, while leveraging the expertise of researchers in various disciplines. Our design space aims to offer researchers and designers a practical tool to navigate, comprehend, and compare the various possibilities of writing assistants, and aid in the design of new writing assistants. Mina Lee 0002, Katy Ilonka Gero, John Joon Young Chung, Simon Buckingham Shum, Vipul Raheja, Hua Shen 0005, Subhashini Venugopalan, Thiemo Wambsganss, David Zhou, Emad A. Alghamdi, Tal August, Avinash Bhat, Madiha Zahrah Choksi, Senjuti Dutta, Jin L. C. Guo, Md. Naimul Hoque, Simon Knight 0001, Seyed Parsa Neshaei, Antonette Shibani, Disha Shrivastava, Lila Shroff, Agnia Sergeyuk, Jessi Stark, Sarah Sterman, Sitong Wang 0001, Antoine Bosselut, Daniel Buschek, Joseph Chee Chang, Sherol Chen, Max Kreminski, Joonsuk Park, Roy D. Pea, Eugenia Ha Rim Rho, Shannon Shen 0001, Pao Siangliulue |
CHI | 27 |
| 2024 | REFINER: Reasoning Feedback on Intermediate RepresentationsabstractDebjit Paul, Mete Ismayilzada, Maxime Peyrard, Beatriz Borges, Antoine Bosselut, Robert West, Boi Faltings. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Debjit Paul, Mete Ismayilzada, Maxime Peyrard, Beatriz Borges, Antoine Bosselut, Robert West 0001, Boi Faltings |
EACL (1) | 5 |
| 2024 | ConGeo: Robust Cross-View Geo-Localization Across Ground View Variations
Li Mi, Chang Xu 0027, Javiera Castillo-Navarro, Syrielle Montariol, Wen Yang 0001, Antoine Bosselut, Devis Tuia |
ECCV (14) | 6 |
| 2024 | Discovering Knowledge-Critical Subnetworks in Pretrained Language ModelsabstractPretrained language models (LMs) encode implicit representations of knowledge in their parameters.However, localizing these representations and disentangling them from each other remains an open problem.In this work, we investigate whether pretrained language models contain various knowledge-critical subnetworks: particular sparse computational subgraphs that can, if removed, precisely suppress specific knowledge the model has memorized.We propose a multi-objective differentiable masking scheme that can be applied to both weights and neurons to discover such subnetworks and show that we can use them to precisely remove specific knowledge from models while minimizing adverse effects on the behavior of the original model.We demonstrate our method on multiple GPT2 variants, uncovering highly sparse subnetworks (98%+ sparsity) that are critical for expressing specific collections of relational knowledge.When these subnetworks are removed, the remaining network maintains most of its initial abilities but struggles to represent the suppressed knowledge.1 Deniz Bayazit, Negar Foroutan Eghlidi, Zeming Chen 0001, Gail Weiss, Antoine Bosselut |
EMNLP | 5 |
| 2024 | Let Me Teach You: Pedagogical Foundations of Feedback for Language ModelsabstractNatural Language Feedback (NLF) is an increasingly popular mechanism for aligning Large Language Models (LLMs) to human preferences.Despite the diversity of the information it can convey, NLF methods are often handdesigned and arbitrary, with little systematic grounding.At the same time, research in learning sciences has long established several effective feedback models.In this opinion piece, we compile ideas from pedagogy to introduce FELT, a feedback framework for LLMs that outlines various characteristics of the feedback space, and a feedback content taxonomy based on these variables, providing a general mapping of the feedback space.In addition to streamlining NLF designs, FELT also brings out new, unexplored directions for research in NLF.We make our taxonomy available to the community, providing guides and examples for mapping our categorizations to future research. Beatriz Borges, Niket Tandon, Tanja Käser, Antoine Bosselut |
EMNLP | 4 |
| 2024 | "Flex Tape Can't Fix That": Bias and Misinformation in Edited Language ModelsabstractWeight-based model editing methods update the parametric knowledge of language models post-training.However, these methods can unintentionally alter unrelated parametric knowledge representations, potentially increasing the risk of harm.In this work, we investigate how weight editing methods unexpectedly amplify model biases after edits.We introduce a novel benchmark dataset, SEESAW-CF, for measuring bias amplification of model editing methods for demographic traits such as race, geographic origin, and gender.We use SEESAW-CF to examine the impact of model editing on bias in five large language models.Our results demonstrate that edited models exhibit, to various degrees, more biased behavior for certain demographic groups than before they were edited, specifically becoming less confident in properties for Asian and African subjects.Additionally, editing facts about place of birth, country of citizenship, or gender has particularly negative effects on the model's knowledge about unrelated properties, such as field of work, a pattern observed across multiple models. Karina Halevy, Anna Sotnikova, Badr AlKhamissi, Syrielle Montariol, Antoine Bosselut |
EMNLP | 5 |
| 2024 | Training Visual Language Models with Object Detection: Grounded Change Descriptions in Satellite ImagesabstractRecently, generalist Vision Language Models (VLMs) have shown exceptional progress in tasks previously dominated by specialized computer vision models. This becomes more prevalent when visual grounding capabilities, such as the ability to reason over input text and image to generate bounding boxes around objects, are required. However, how these capabilities transfer to specialized domains such as remote sensing remains understudied, despite the recent increase in specialized models for Earth observation. In this work, we evaluate how grounding visual entities – by generating bounding-box coordinates – affects VLM performance in satellite imagery. To this end, we create two instruction-following tasks sourced from the xBD dataset, describing changes due to natural disasters observed in satellite images. We fine-tune several instances of MiniGPTv2, an open-source VLM with grounding capabilities, and evaluate their performance under the "grounded" vs. "not grounded" settings. We find that generating bounding boxes to refer to visual entities increases performance in tasks related to objects in the image, but only when the number of entities in the image is limited. João Luis Prado, Syrielle Montariol, Javiera Castillo-Navarro, Devis Tuia, Antoine Bosselut |
IGARSS | 5 |
| 2024 | Course Recommender Systems Need to Consider the Job MarketabstractCurrent course recommender systems primarily leverage learner-course interactions, course content, learner preferences, and supplementary course details like instructor, institution, ratings, and reviews, to make their recommendation. However, these systems often overlook a critical aspect: the evolving skill demand of the job market. This paper focuses on the perspective of academic researchers, working in collaboration with the industry, aiming to develop a course recommender system that incorporates job market skill demands. In light of the job market's rapid changes and the current state of research in course recommender systems, we outline essential properties for course recommender systems to address these demands effectively, including explainable, sequential, unsupervised, and aligned with the job market and user's goals. Our discussion extends to the challenges and research questions this objective entails, including unsupervised skill extraction from job listings, course descriptions, and resumes, as well as predicting recommendations that align with learner objectives and the job market and designing metrics to evaluate this alignment. Furthermore, we introduce an initial system that addresses some existing limitations of course recommender systems using large Language Models (LLMs) for skill extraction and Reinforcement Learning (RL) for alignment with the job market. We provide empirical results using open-source data to demonstrate its effectiveness. Jibril Frej, Anna Dai, Syrielle Montariol, Antoine Bosselut, Tanja Käser |
SIGIR | 4 |
| 2023 | DISCO: Distilling Counterfactuals with Large Language ModelsabstractModels trained with counterfactually augmented data learn representations of the causal structure of tasks, enabling robust generalization.However, high-quality counterfactual data is scarce for most tasks and not easily generated at scale.When crowdsourced, such data is typically limited in scale and diversity; when generated using supervised methods, it is computationally expensive to extend to new counterfactual dimensions.In this work, we introduce DISCO (DIStilled COunterfactual Data), a new method for automatically generating highquality counterfactual data at scale.DISCO engineers prompts to generate phrasal perturbations with a large general language model.Then, a task-specific teacher model filters these generations to distill high-quality counterfactual data.While task-agnostic, we apply our pipeline to the task of natural language inference (NLI) and find that on challenging evaluations such as the NLI stress test, comparatively smaller student models trained with DISCOgenerated counterfactuals are more robust (6% absolute) and generalize better across distributions (2%) compared to models trained without data augmentation.Furthermore, DISCOaugmented models are 10% more consistent between counterfactual pairs on three evaluation sets, demonstrating that DISCO augmentation enables models to more reliably learn causal representations.Our Zeming Chen 0001, Qiyue Gao, Antoine Bosselut, Ashish Sabharwal, Kyle Richardson 0001 |
ACL (1) | 3 |
| 2023 | Mitigating Label Biases for In-context LearningabstractVarious design settings for in-context learning (ICL), such as the choice and order of the incontext examples, can bias the model's predictions.While many studies discuss these design choices, there have been few systematic investigations into categorizing them and mitigating their impact.In this work, we define a typology for three types of label biases in ICL for text classification: vanilla-label bias, contextlabel bias, and domain-label bias (which we conceptualize and detect for the first time).Our analysis demonstrates that prior label bias calibration methods fall short of addressing all three types of biases.Specifically, domainlabel bias restricts LLMs to random-level performance on many tasks regardless of the choice of in-context examples.To mitigate the effect of these biases, we propose a simple bias calibration method that estimates a language model's label bias using random indomain words from the task corpus.After controlling for this estimated bias when making predictions, our novel domain-context calibration significantly improves the ICL performance of GPT-J and GPT-3 on a wide range of tasks.The gain is substantial on tasks with large domain-label bias (up to 37% in Macro-F1).Furthermore, our results generalize to models with different scales, pretraining methods, and manually-designed task instructions, showing the prevalence of label biases in ICL. Zeming Chen 0001, Antoine Bosselut |
ACL (1) | 4 |
| 2023 | PeaCoK: Persona Commonsense Knowledge for Consistent and Engaging NarrativesabstractSilin Gao, Beatriz Borges, Soyoung Oh, Deniz Bayazit, Saya Kanno, Hiromi Wakaki, Yuki Mitsufuji, Antoine Bosselut. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Silin Gao, Beatriz Borges, Soyoung Oh, Deniz Bayazit, Saya Kanno, Hiromi Wakaki, Yuki Mitsufuji, Antoine Bosselut |
ACL (1) | 8 |
| 2023 | Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language ModelsabstractYifan Hou, Jiaoda Li, Yu Fei, Alessandro Stolfo, Wangchunshu Zhou, Guangtao Zeng, Antoine Bosselut, Mrinmaya Sachan. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Jiaoda Li, Alessandro Stolfo, Wangchunshu Zhou, Guangtao Zeng, Antoine Bosselut, Mrinmaya Sachan |
EMNLP | 7 |
| 2023 | CRoW: Benchmarking Commonsense Reasoning in Real-World TasksabstractRecent efforts in natural language processing (NLP) commonsense reasoning research have yielded a considerable number of new datasets and benchmarks.However, most of these datasets formulate commonsense reasoning challenges in artificial scenarios that are not reflective of the tasks which real-world NLP systems are designed to solve.In this work, we present CROW, a manually-curated, multitask benchmark that evaluates the ability of models to apply commonsense reasoning in the context of six real-world NLP tasks.CROW is constructed using a multi-stage data collection pipeline that rewrites examples from existing datasets using commonsense-violating perturbations.We use CROWto study how NLP systems perform across different dimensions of commonsense knowledge, such as physical, temporal, and social reasoning.We find a significant performance gap when NLP systems are evaluated on CROWcompared to humans, showcasing that commonsense reasoning is far from being solved in real-world task settings.We make our dataset and leaderboard available to the research community.1 * Equal contribution 1 https://github.com/mismayil/crowDialogue Agent: Hi, would you like some free candies?Human: Sure.What are you handing these out for?Agent: Well, we're trying to gather some Mete Ismayilzada, Debjit Paul, Syrielle Montariol, Mor Geva, Antoine Bosselut |
EMNLP | 5 |
| 2023 | CRAB: Assessing the Strength of Causal Relationships Between Real-world EventsabstractUnderstanding narratives requires reasoning about the cause-and-effect relationships between events mentioned in the text.While existing foundation models yield impressive results in many NLP tasks requiring reasoning, it is unclear whether they understand the complexity of the underlying network of causal relationships of events in narratives.In this work, we present CRAB, a new Causal Reasoning Assessment Benchmark designed to evaluate causal understanding of events in real-world narratives.CRAB contains fine-grained, contextual causality annotations for ∼ 2.7K pairs of real-world events that describe various newsworthy event timelines (e.g., the acquisition of Twitter by Elon Musk).Using CRAB, we measure the performance of several large language models, demonstrating that most systems achieve poor performance on the task.Motivated by classical causal principles, we also analyze the causal structures of groups of events in CRAB, and find that models perform worse on causal reasoning when events are derived from complex causal structures compared to simple linear causal chains.We make our dataset and code available to the research community. Angelika Romanou, Syrielle Montariol, Debjit Paul, Léo Laugier, Karl Aberer, Antoine Bosselut |
EMNLP | 6 |
| 2023 | RECKONING: Reasoning through Dynamic Knowledge EncodingabstractRecent studies on transformer-based language models show that they can answer questions by reasoning over knowledge provided as part of the context (i.e., in-context reasoning). However, since the available knowledge is often not filtered for a particular question, in-context reasoning can be sensitive to distractor facts, additional content that is irrelevant to a question but that may be relevant for a different question (i.e., not necessarily random noise). In these situations, the model fails to
distinguish the necessary knowledge to answer the question, leading to spurious reasoning and degraded performance. This reasoning failure contrasts with the model’s apparent ability to distinguish its contextual knowledge from all the knowledge it has memorized during pre-training. Following this observation, we propose teaching the model to reason more robustly by folding the provided contextual knowledge into the model’s parameters before presenting it with a question. Our method, RECKONING, is a bi-level learning algorithm that teaches language models to reason by updating their parametric knowledge through back-propagation, allowing them to answer questions using the updated parameters. During training, the inner loop rapidly adapts a copy of the model weights to encode contextual knowledge into its parameters. In the outer loop, the model learns to use the updated weights to reproduce and answer reasoning questions about the memorized knowledge. Our experiments on three diverse multi-hop reasoning datasets show that RECKONING’s performance improves over the in-context reasoning baseline (by up to 4.5%). We also find that compared to in-context reasoning, RECKONING generalizes better to longer reasoning chains unseen during training, is more robust to distractors in the context, and is computationally more efficient when multiple questions are asked about the same knowledge. Zeming Chen 0001, Gail Weiss, Eric Mitchell, Asli Celikyilmaz, Antoine Bosselut |
NeurIPS | 5 |
| 2022 | Synthetic Disinformation Attacks on Automated Fact Verification SystemsabstractAutomated fact-checking is a needed technology to curtail the spread of online misinformation. One current framework for such solutions proposes to verify claims by retrieving supporting or refuting evidence from related textual sources. However, the realistic use cases for fact-checkers will require verifying claims against evidence sources that could be affected by the same misinformation. Furthermore, the development of modern NLP tools that can produce coherent, fabricated content would allow malicious actors to systematically generate adversarial disinformation for fact-checkers. In this work, we explore the sensitivity of automated fact-checkers to synthetic adversarial evidence in two simulated settings: ADVERSARIAL ADDITION, where we fabricate documents and add them to the evidence repository available to the fact-checking system, and ADVERSARIAL MODIFICATION, where existing evidence source documents in the repository are automatically altered. Our study across multiple models on three benchmarks demonstrates that these systems suffer significant performance drops against these attacks. Finally, we discuss the growing threat of modern NLG systems as generators of disinformation in the context of the challenges they pose to automated fact-checkers. Yibing Du, Antoine Bosselut, Christopher D. Manning |
AAAI | 2 |
| 2022 | Discovering Language-neutral Sub-networks in Multilingual Language ModelsabstractMultilingual pre-trained language models transfer remarkably well on cross-lingual downstream tasks.However, the extent to which they learn language-neutral representations (i.e., shared representations that encode similar phenomena across languages), and the effect of such representations on cross-lingual transfer performance, remain open questions.In this work, we conceptualize language neutrality of multilingual models as a function of the overlap between language-encoding subnetworks of these models.We employ the lottery ticket hypothesis to discover sub-networks that are individually optimized for various languages and tasks.Our evaluation across three distinct tasks and eleven typologically-diverse languages demonstrates that sub-networks for different languages are topologically similar (i.e., language-neutral), making them effective initializations for cross-lingual transfer with limited performance degradation. 1 Negar Foroutan Eghlidi, Mohammadreza Banaei, Rémi Lebret, Antoine Bosselut, Karl Aberer |
EMNLP | 4 |
| 2022 | Conditional set generation using Seq2seq modelsabstractConditional set generation learns a mapping from an input sequence of tokens to a set.Several NLP tasks, such as entity typing and dialogue emotion tagging, are instances of set generation.SEQ2SEQ models, a popular choice for set generation, treat a set as a sequence and do not fully leverage its key properties, namely order-invariance and cardinality.We propose a novel algorithm for effectively sampling informative orders over the combinatorial space of label orders.We jointly model the set cardinality and output by prepending the set size and taking advantage of the autoregressive factorization used by SEQ2SEQ models.Our method is a model-independent data augmentation approach that endows any SEQ2SEQ model with the signals of order-invariance and cardinality.Training a SEQ2SEQ model on this augmented data (without any additional annotations) gets an average relative improvement of 20% on four benchmark datasets across various models: BART-base, T5-11B, and GPT3-175B. 1 Aman Madaan, Dheeraj Rajagopal, Niket Tandon, Yiming Yang 0002, Antoine Bosselut |
EMNLP | 5 |
| 2022 | GreaseLM: Graph REASoning Enhanced Language Models
Xikun Zhang 0001, Antoine Bosselut, Michihiro Yasunaga, Hongyu Ren, Percy Liang, Christopher D. Manning, Jure Leskovec |
ICLR | 2 |
| 2022 | Fast Model Editing at Scale
Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, Christopher D. Manning |
ICLR | 3 |
| 2022 | Memory-Based Model Editing at ScaleabstractEven the largest neural networks make errors, and once-correct predictions can become invalid as the world changes. Model editors make local updates to the behavior of base (pre-trained) models to inject updated knowledge or correct undesirable behaviors. Existing model editors have shown promise, but also suffer from insufficient expressiveness: they struggle to accurately model an edit’s intended scope (examples affected by the edit), leading to inaccurate predictions for test inputs loosely related to the edit, and they often fail altogether after many edits. As a higher-capacity alternative, we propose Semi-Parametric Editing with a Retrieval-Augmented Counterfactual Model (SERAC), which stores edits in an explicit memory and learns to reason over them to modulate the base model’s predictions as needed. To enable more rigorous evaluation of model editors, we introduce three challenging language model editing problems based on question answering, fact-checking, and dialogue generation. We find that only SERAC achieves high performance on all three problems, consistently outperforming existing approaches to model editing by a significant margin. Code, data, and additional project information will be made available at https://sites.google.com/view/serac-editing. Eric Mitchell, Charles Lin, Antoine Bosselut, Christopher D. Manning, Chelsea Finn |
ICML | 3 |
| 2022 | Deep Bidirectional Language-Knowledge Graph PretrainingabstractPretraining a language model (LM) on text has been shown to help various downstream NLP tasks. Recent works show that a knowledge graph (KG) can complement text data, offering structured background knowledge that provides a useful scaffold for reasoning. However, these works are not pretrained to learn a deep fusion of the two modalities at scale, limiting the potential to acquire fully joint representations of text and KG. Here we propose DRAGON (Deep Bidirectional Language-Knowledge Graph Pretraining), a self-supervised approach to pretraining a deeply joint language-knowledge foundation model from text and KG at scale. Specifically, our model takes pairs of text segments and relevant KG subgraphs as input and bidirectionally fuses information from both modalities. We pretrain this model by unifying two self-supervised reasoning tasks, masked language modeling and KG link prediction. DRAGON outperforms existing LM and LM+KG models on diverse downstream tasks including question answering across general and biomedical domains, with +5% absolute gain on average. In particular, DRAGON achieves notable performance on complex reasoning about language and knowledge (+10% on questions involving long contexts or multi-step reasoning) and low-resource QA (+8% on OBQA and RiddleSense), and new state-of-the-art results on various BioNLP tasks. Our code and trained models are available at https://github.com/michiyasunaga/dragon. Michihiro Yasunaga, Antoine Bosselut, Hongyu Ren, Xikun Zhang 0001, Christopher D. Manning, Percy Liang, Jure Leskovec |
NeurIPS | 2 |
| 2022 | End-to-End Task-Oriented Dialog Modeling With Semi-Structured Knowledge ManagementabstractCurrent task-oriented dialog (TOD) systems mostly manage structured knowledge (e.g. databases and tables) to guide the goal-oriented conversations. However, they fall short of handling dialogs which also involve unstructured knowledge (e.g. reviews and documents). In this paper, we formulate a task of modeling TOD grounded on a fusion of structured and unstructured knowledge. To address this task, we propose a TOD system with semi-structured knowledge management, SeKnow, which extends the belief state to manage knowledge with both structured and unstructured contents. Furthermore, we introduce two implementations of SeKnow based on a non-pretrained sequence-to-sequence model and a pretrained language model, respectively. Both implementations use the end-to-end manner to jointly optimize dialog modeling grounded on structured and unstructured knowledge. We conduct experiments on a modified version of MultiWOZ 2.1 dataset, Mod-MultiWOZ 2.1, where dialogs are processed to involve semi-structured knowledge. Experimental results show that SeKnow has strong performances in both end-to-end dialog and intermediate knowledge management, compared to existing TOD systems and their extensions with pipeline knowledge management schemes. Silin Gao, Ryuichi Takanobu, Antoine Bosselut, Minlie Huang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Dynamic Neuro-Symbolic Knowledge Graph Construction for Zero-shot Commonsense Question AnsweringabstractUnderstanding narratives requires reasoning about implicit world knowledge related to the causes, effects, and states of situations described in text. At the core of this challenge is how to access contextually relevant knowledge on demand and reason over it. In this paper, we present initial studies toward zero-shot commonsense question answering by formulating the task as inference over dynamically generated commonsense knowledge graphs. In contrast to previous studies for knowledge integration that rely on retrieval of existing knowledge from static knowledge graphs, our study requires commonsense knowledge integration where contextually relevant knowledge is often not present in existing knowledge bases. Therefore, we present a novel approach that generates contextually-relevant symbolic knowledge structures on demand using generative neural commonsense knowledge models. Empirical results on two datasets demonstrate the efficacy of our neuro-symbolic approach for dynamically constructing knowledge graphs for reasoning. Our approach achieves significant performance boosts over pretrained language models and vanilla knowledge models, all while providing interpretable reasoning paths for its predictions. Antoine Bosselut, Ronan Le Bras 0001, Yejin Choi 0001 |
AAAI | 1 |
| 2021 | (Comet-) Atomic 2020: On Symbolic and Neural Commonsense Knowledge GraphsabstractRecent years have brought about a renewed interest in commonsense representation and reasoning in the field of natural language understanding. The development of new commonsense knowledge graphs (CSKG) has been central to these advances as their diverse facts can be used and referenced by machine learning models for tackling new and challenging tasks. At the same time, there remain questions about the quality and coverage of these resources due to the massive scale required to comprehensively encompass general commonsense knowledge. In this work, we posit that manually constructed CSKGs will never achieve the coverage necessary to be applicable in all situations encountered by NLP agents. Therefore, we propose a new evaluation framework for testing the utility of KGs based on how effectively implicit knowledge representations can be learned from them. With this new goal, we propose Atomic 2020, a new CSKG of general-purpose commonsense knowledge containing knowledge that is not readily available in pretrained language models. We evaluate its properties in comparison with other leading CSKGs, performing the first large-scale pairwise study of commonsense knowledge resources. Next, we show that Atomic 2020 is better suited for training knowledge models that can generate accurate, representative knowledge for new, unseen entities and events. Finally, through human evaluation, we show that the few-shot performance of GPT-3 (175B parameters), while impressive, remains ~12 absolute points lower than a BART-based knowledge model trained on Atomic 2020 despite using over 430x fewer parameters. Jena D. Hwang, Chandra Bhagavatula, Ronan Le Bras 0001, Jeff Da, Keisuke Sakaguchi, Antoine Bosselut, Yejin Choi 0001 |
AAAI | 6 |
| 2021 | Edited Media Understanding Frames: Reasoning About the Intent and Implications of Visual MisinformationabstractJeff Da, Maxwell Forbes, Rowan Zellers, Anthony Zheng, Jena D. Hwang, Antoine Bosselut, Yejin Choi. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jeff Da, Maxwell Forbes, Rowan Zellers, Anthony Zheng, Jena D. Hwang, Antoine Bosselut, Yejin Choi 0001 |
ACL/IJCNLP (1) | 6 |
| 2021 | Discourse Understanding and Factual Consistency in Abstractive SummarizationabstractSaadia Gabriel, Antoine Bosselut, Jeff Da, Ari Holtzman, Jan Buys, Kyle Lo, Asli Celikyilmaz, Yejin Choi. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Saadia Gabriel, Antoine Bosselut, Jeff Da, Ari Holtzman, Jan Buys, Kyle Lo, Asli Celikyilmaz, Yejin Choi 0001 |
EACL | 2 |
| 2021 | Conversational Multi-Hop Reasoning with Neural Commonsense Knowledge and Symbolic Logic RulesabstractOne of the challenges faced by conversational agents is their inability to identify unstated presumptions of their users' commands, a task trivial for humans due to their common sense.In this paper, we propose a zeroshot commonsense reasoning system for conversational agents in an attempt to achieve this.Our reasoner uncovers unstated presumptions from user commands satisfying a general template of if-(state ), then-(action ), because-(goal ).Our reasoner uses a state-ofthe-art transformer-based generative commonsense knowledge base (KB) as its source of background knowledge for reasoning.We propose a novel and iterative knowledge query mechanism to extract multi-hop reasoning chains from the neural KB which uses symbolic logic rules to significantly reduce the search space.Similar to any KBs gathered to date, our commonsense KB is prone to missing knowledge.Therefore, we propose to conversationally elicit the missing knowledge from human users with our novel dynamic question generation strategy, which generates and presents contextualized queries to human users.We evaluate the model with a user study with human users that achieves a 35% higher success rate compared to SOTA. Forough Arabshahi, Jennifer Lee, Antoine Bosselut, Yejin Choi 0001, Tom M. Mitchell |
EMNLP (1) | 3 |
| 2021 | "I'm Not Mad": Commonsense Implications of Negation and ContradictionabstractLiwei Jiang, Antoine Bosselut, Chandra Bhagavatula, Yejin Choi. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Antoine Bosselut, Chandra Bhagavatula, Yejin Choi 0001 |
NAACL-HLT | 2 |
| 2021 | QA-GNN: Reasoning with Language Models and Knowledge Graphs for Question AnsweringabstractMichihiro Yasunaga, Hongyu Ren, Antoine Bosselut, Percy Liang, Jure Leskovec. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Michihiro Yasunaga, Hongyu Ren, Antoine Bosselut, Percy Liang, Jure Leskovec |
NAACL-HLT | 3 |
| 2020 | Commonsense Knowledge Base Completion with Structural and Semantic ContextabstractAutomatic KB completion for commonsense knowledge graphs (e.g., ATOMIC and ConceptNet) poses unique challenges compared to the much studied conventional knowledge bases (e.g., Freebase). Commonsense knowledge graphs use free-form text to represent nodes, resulting in orders of magnitude more nodes compared to conventional KBs ( ∼18x more nodes in ATOMIC compared to Freebase (FB15K-237)). Importantly, this implies significantly sparser graph structures — a major challenge for existing KB completion methods that assume densely connected graphs over a relatively smaller set of nodes.In this paper, we present novel KB completion models that can address these challenges by exploiting the structural and semantic context of nodes. Specifically, we investigate two key ideas: (1) learning from local graph structure, using graph convolutional networks and automatic graph densification and (2) transfer learning from pre-trained language models to knowledge graphs for enhanced contextual representation of knowledge. We describe our method to incorporate information from both these sources in a joint model and provide the first empirical results for KB completion on ATOMIC and evaluation with ranking metrics on ConceptNet. Our results demonstrate the effectiveness of language model representations in boosting link prediction performance and the advantages of learning from local graph structure (+1.5 points in MRR for ConceptNet) when training on subgraphs for computational efficiency. Further analysis on model predictions shines light on the types of commonsense knowledge that language models capture well. Chaitanya Malaviya, Chandra Bhagavatula, Antoine Bosselut, Yejin Choi 0001 |
AAAI | 3 |
| 2020 | Back to the Future: Unsupervised Backprop-based Decoding for Counterfactual and Abductive Commonsense ReasoningabstractLianhui Qin, Vered Shwartz, Peter West, Chandra Bhagavatula, Jena D. Hwang, Ronan Le Bras, Antoine Bosselut, Yejin Choi. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Lianhui Qin, Vered Shwartz, Peter West, Chandra Bhagavatula, Jena D. Hwang, Ronan Le Bras 0001, Antoine Bosselut, Yejin Choi 0001 |
EMNLP (1) | 7 |
| 2019 | COMET: Commonsense Transformers for Automatic Knowledge Graph ConstructionabstractWe present the first comprehensive study on automatic knowledge base construction for two prevalent commonsense knowledge graphs: ATOMIC (Sap et al., 2019) and Con-ceptNet (Speer et al., 2017).Contrary to many conventional KBs that store knowledge with canonical templates, commonsense KBs only store loosely structured open-text descriptions of knowledge.We posit that an important step toward automatic commonsense completion is the development of generative models of commonsense knowledge, and propose COMmonsEnse Transformers (COMET ) that learn to generate rich and diverse commonsense descriptions in natural language.Despite the challenges of commonsense modeling, our investigation reveals promising results when implicit knowledge from deep pre-trained language models is transferred to generate explicit knowledge in commonsense knowledge graphs.Empirical results demonstrate that COMET is able to generate novel knowledge that humans rate as high quality, with up to 77.5% (ATOMIC) and 91.7% (ConceptNet) precision at top 1, which approaches human performance for these resources.Our findings suggest that using generative commonsense models for automatic commonsense KB completion could soon be a plausible alternative to extractive methods. Antoine Bosselut, Hannah Rashkin, Maarten Sap, Chaitanya Malaviya, Asli Celikyilmaz, Yejin Choi 0001 |
ACL (1) | 1 |
| 2019 | Everything Happens for a Reason: Discovering the Purpose of Actions in Procedural TextabstractBhavana Dalvi, Niket Tandon, Antoine Bosselut, Wen-tau Yih, Peter Clark. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Bhavana Dalvi, Niket Tandon, Antoine Bosselut, Scott Yih, Peter Clark |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Counterfactual Story Reasoning and GenerationabstractLianhui Qin, Antoine Bosselut, Ari Holtzman, Chandra Bhagavatula, Elizabeth Clark, Yejin Choi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Lianhui Qin, Antoine Bosselut, Ari Holtzman, Chandra Bhagavatula, Elizabeth Clark, Yejin Choi 0001 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | WIQA: A dataset for "What if..." reasoning over procedural textabstractNiket Tandon, Bhavana Dalvi, Keisuke Sakaguchi, Peter Clark, Antoine Bosselut. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Niket Tandon, Bhavana Dalvi, Keisuke Sakaguchi, Peter Clark, Antoine Bosselut |
EMNLP/IJCNLP (1) | 5 |
| 2018 | Learning to Write with Cooperative DiscriminatorsabstractAri Holtzman, Jan Buys, Maxwell Forbes, Antoine Bosselut, David Golub, Yejin Choi. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Ari Holtzman, Jan Buys, Maxwell Forbes, Antoine Bosselut, David Golub, Yejin Choi 0001 |
ACL (1) | 4 |
| 2018 | Modeling Naive Psychology of Characters in Simple Commonsense StoriesabstractUnderstanding a narrative requires reading between the lines and reasoning about the unspoken but obvious implications about events and people's mental states -a capability that is trivial for humans but remarkably hard for machines.To facilitate research addressing this challenge, we introduce a new annotation framework to explain naive psychology of story characters as fully-specified chains of mental states with respect to motivations and emotional reactions.Our work presents a new largescale dataset with rich low-level annotations and establishes baseline performance on several new tasks, suggesting avenues for future research. Hannah Rashkin, Antoine Bosselut, Maarten Sap, Kevin Knight, Yejin Choi 0001 |
ACL (1) | 2 |
| 2018 | Reasoning about Actions and State Changes by Injecting Commonsense KnowledgeabstractComprehending procedural text, e.g., a paragraph describing photosynthesis, requires modeling actions and the state changes they produce, so that questions about entities at different timepoints can be answered.Although several recent systems have shown impressive progress in this task, their predictions can be globally inconsistent or highly improbable.In this paper, we show how the predicted effects of actions in the context of a paragraph can be improved in two ways: (1) by incorporating global, commonsense constraints (e.g., a non-existent entity cannot be destroyed), and (2) by biasing reading with preferences from large-scale corpora (e.g., trees rarely move).Unlike earlier methods, we treat the problem as a neural structured prediction task, allowing hard and soft constraints to steer the model away from unlikely predictions.We show that the new model significantly outperforms earlier systems on a benchmark dataset for procedural text comprehension (+8% relative gain), and that it also avoids some of the nonsensical predictions that earlier systems make. Niket Tandon, Bhavana Dalvi, Joel Grus, Scott Yih, Antoine Bosselut, Peter Clark |
EMNLP | 5 |
| 2018 | Simulating Action Dynamics with Neural Process Networks
Antoine Bosselut, Omer Levy, Ari Holtzman, Corin Ennis, Dieter Fox, Yejin Choi 0001 |
ICLR (Poster) | 1 |
| 2018 | Discourse-Aware Neural Rewards for Coherent Text GenerationabstractAntoine Bosselut, Asli Celikyilmaz, Xiaodong He, Jianfeng Gao, Po-Sen Huang, Yejin Choi. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Antoine Bosselut, Asli Celikyilmaz, Xiaodong He 0001, Jianfeng Gao 0001, Po-Sen Huang, Yejin Choi 0001 |
NAACL-HLT | 1 |
| 2018 | Deep Communicating Agents for Abstractive SummarizationabstractAsli Celikyilmaz, Antoine Bosselut, Xiaodong He, Yejin Choi. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Asli Celikyilmaz, Antoine Bosselut, Xiaodong He 0001, Yejin Choi 0001 |
NAACL-HLT | 2 |
| 2016 | Learning Prototypical Event Structure from Photo AlbumsabstractActivities and events in our lives are structural, be it a vacation, a camping trip, or a wedding.While individual details vary, there are characteristic patterns that are specific to each of these scenarios.For example, a wedding typically consists of a sequence of events such as walking down the aisle, exchanging vows, and dancing.In this paper, we present a data-driven approach to learning event knowledge from a large collection of photo albums.We formulate the task as constrained optimization to induce the prototypical temporal structure of an event, integrating both visual and textual cues.Comprehensive evaluation demonstrates that it is possible to learn multimodal knowledge of event structure from noisy web content. Antoine Bosselut, Jianfu Chen, David Scott Warren, Hannaneh Hajishirzi, Yejin Choi 0001 |
ACL (1) | 1 |