Syrielle Montariol

dblp:245/2618 · DBLP profile ↗
← Back
18ranked-venue papers
2as first author
18since 2021 · last 2026
0000-0003-1355-8778ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 2 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
abstract
Alejandro Hernández-Cano, Alexander Hägele, Allen Hao Huang, Angelika Romanou, Antoni-Joan Solergibert, Barna Pásztor, Bettina Messmer, Dhia Garbaya, Eduard Frank Ďurech, Ido Hakimi, Juan Garcia Giraldo, Mete Ismayilzada, Negar Foroutan, Skander Moalla, Tiancheng Chen, Vinko Sabolčec, Yixuan Xu, Michael Aerni, Badr AlKhamissi, Inés Altemir Marinas, Mohammad Hossein Amani, Matin Ansaripour, Ilia Badanin, Harold Benoit, Emanuela Boros, Nicholas John Browning, Fabian Bösch, Maximilian Böther, Niklas Canova, Camille Challier, Clément Charmillot, Jonathan Coles, Jan Milan Deriu, Arnout Devos, Lukas Drescher, Daniil Dzenhaliou, Maud Ehrmann, Dongyang Fan, Simin Fan, Silin Gao, Miguel Gila, María Grandury, Diba Hashemi, Alexander Miserlis Hoyle, Jiaming Jiang, Mark Klein, Andrei Kucharavy, Anastasiia Kucherenko, Frederike Lübeck, Roman Machacek, Theofilos Ioannis Manitaras, Andreas Marfurt, Kyle Matoba, Simon Matrenok, Henrique Mendonça, Fawzi Roberto Mohamed, Syrielle Montariol, Luca Mouchel, Sven Najem-Meyer, Jingwei Ni, Gennaro Oliva, Matteo Pagliardini, Elia Palme, Andrei Panferov, Léo Paoletti, Marco Passerini, Ivan Pavlov, Auguste Poiroux, Kaustubh Ponkshe, Nathan Ranchin, Javier Rando, Mathieu Sauser, Jakhongir Saydaliev, Mukhammadali Sayfiddinov, Marian Schneider, Stefano Schuppli, Marco Scialanga, Andrei Semenov, Kumar Shridhar, Raghav Singhal, Anna Sotnikova, Alexander Sternfeld, Ayush Kumar Tarun, Paul Teiletche, Jannis Vamvas, Xiaozhe Yao, Hao Zhao, Alexander Ilic, Ana Klimovic, Andreas Krause, Caglar Gulcehre, David Rosenthal, Elliott Ash, Florian Tramèr, Joost VandeVondele, Livio Veraldi, Martin Rajman, Thomas C. Schulthess, Torsten Hoefler, Antoine Bosselut, Martin Jaggi, Imanol Schlag. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Alejandro Hernández-Cano, Alexander Hägele, Allen Hao Huang, Angelika Romanou, Antoni-Joan Solergibert i Llaquet, Barna Pásztor, Bettina Messmer, Dhia Garbaya, Eduard Durech, Ido Hakimi, Juan Garcia Giraldo, Mete Ismayilzada, Negar Foroutan Eghlidi, Skander Moalla, Tiancheng Chen, Vinko Sabolcec, Yixuan Even Xu, Michael Aerni, Badr AlKhamissi, Ines Altemir Marinas, Mohammad Hossein Amani, Matin Ansaripour, Ilia Badanin, Harold Benoit, Emanuela Boros, Nicholas John Browning, Fabian Bösch, Maximilian Böther, Niklas Canova, Camille Challier, Clément Charmillot, Jonathan Coles, Jan Deriu, Arnout Devos, Lukas Drescher, Daniil Dzenhaliou, Maud Ehrmann, Dongyang Fan, Simin Fan, Silin Gao, Miguel Gila, María Grandury, Diba Hashemi, Alexander Miserlis Hoyle, Jiaming Jiang, Mark Klein 0002, Andrei Kucharavy, Anastasiia Kucherenko, Frederike Lübeck, Roman Machacek, Theofilos Ioannis Manitaras, Andreas Marfurt, Kyle Matoba, Simon Matrenok, Henrique Mendonça, Fawzi Roberto Mohamed, Syrielle Montariol, Luca Mouchel, Sven Najem-Meyer, Jingwei Ni, Gennaro Oliva, Matteo Pagliardini, Elia Palme, Andrei Panferov, Léo Paoletti, Marco Passerini, Ivan Pavlov, Auguste Poiroux, Kaustubh Ponkshe, Nathan Ranchin, Javier Rando, Mathieu Sauser, Jakhongir Saydaliev, Mukhammadali Sayfiddinov, Marian Schneider, Stefano Schuppli, Marco Scialanga, Andrei Semenov, Kumar Shridhar, Raghav Singhal, Anna Sotnikova, Alexander Sternfeld, Ayush K. Tarun, Paul Teiletche, Jannis Vamvas, Xiaozhe Yao, Alexander Ilic, Ana Klimovic, Andreas Krause 0001, Caglar Gulcehre, David Rosenthal, Elliott Ash, Florian Tramèr, Joost VandeVondele, Livio Veraldi, Martin Rajman, Thomas C. Schulthess, Torsten Hoefler, Antoine Bosselut, Martin Jaggi, Imanol Schlag
ACL (1)57
2026 Checkmate: Interpretable and Explainable RSVQA is the Endgame
Lucrezia Tosato, Christel Tartini-Chappuis, Syrielle Montariol, Flora Weissgerber, Sylvain Lobry, Devis Tuia
ICPR (7)3
2025 VinaBench: Benchmark for Faithful and Consistent Visual Narratives
abstract
Visual narrative generation transforms textual narratives into sequences of images illustrating the content of the text. However, generating visual narratives that are faithful to the input text and self-consistent across generated images remains an open challenge, due to the lack of knowledge constraints used for planning the stories. In this work, we propose a new benchmark, VinaBench, to address this challenge. Our benchmark annotates the underlying commonsense and discourse constraints in visual narrative samples, offering systematic scaffolds for learning the implicit strategies of visual storytelling. Based on the incorporated narrative constraints, we further propose novel metrics to closely evaluate the consistency of generated narrative images and the alignment of generations with the input textual narrative. Our results across three generative vision models demonstrate that learning with VinaBench’s knowledge constraints effectively improves the faithfulness and cohesion of generated visual narratives.1
Silin Gao, Sheryl Mathew, Li Mi, Sepideh Mamooler, Hiromi Wakaki, Yuki Mitsufuji, Syrielle Montariol, Antoine Bosselut
CVPR8
2025 CAVE : Detecting and Explaining Commonsense Anomalies in Visual Environments
abstract
CAVE: Commonsense Anomalies in Visual Environment 🏠 Project Page📄 Paper (EMNLP 2025)💻 Code Dataset Details Dataset Description CAVE is the first benchmark of real-world visual anomalies for evaluating Vision-Language Models (VLMs). It is curated from images captured in real-life settings (photographs and screenshots taken by individuals), sourced from Reddit. The benchmark is grounded in cognitive science literature on how humans detect and resolve anomalies. Each image is annotated with rich, multi-task annotations that support three open-ended tasks (anomaly description, explanation, and justification), one visual grounding task (anomaly localization via bounding boxes), and classification along four dimensions (anomaly category, severity, surprisal, and complexity) that characterize the anomaly. CAVE reveals that state-of-the-art VLMs struggle substantially with visual anomaly perception and commonsense reasoning: the best model (GPT-4o) achieves only ~57% F1-score on anomaly detection even with advanced prompting strategies. Curated by: Rishika Bhagwatkar, Syrielle Montariol, Angelika Romanou, Beatriz Borges, Irina Rish, Antoine Bosselut Affiliations: EPFL, MILA Language: English License: CC-BY-4.0 Published at: EMNLP 2025 (Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing) Dataset Sources Project Page: https://smontariol.github.io/cave-visual-anomalies/ Paper: CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments Contact: [email protected], [email protected] Uses Intended Uses CAVE is designed to evaluate VLMs on their ability to: Detect real-world commonsense anomalies in images (anomaly description). Explain why a detected situation is anomalous (anomaly explanation). Justify how an anomaly might have occurred (anomaly justification). Localize anomalies within images via bounding boxes (anomaly localization). Classify anomalies by their visual manifestation type and numerical features (severity, surprisal, complexity). It also serves as a resource for studying the alignment between human and machine processing of visual anomalies, and for developing improved prompting strategies or fine-tuning approaches for anomaly-related tasks. Out-of-Scope Uses CAVE is a benchmark for evaluation purposes. Its small size (361 images) makes it unsuitable as a training set. It should not be used to deploy anomaly detection systems in safety-critical settings without additional validation. Dataset Structure Overview CAVE consists of 361 images: 309 anomalous and 52 normal (non-anomalous) images. Anomalous images contain up to 3 anomalies each, totaling 334 annotated anomalies. Each anomaly is paired with a unique bounding box. Annotation Fields Each sample includes the following fields: Field Description image The image (photograph or screenshot) image_description Short description of the image content (without describing the anomaly) anomaly_description Textual description of what is anomalous in the image anomaly_explanation Explanation of why the situation is anomalous (commonsense reasoning) anomaly_justification Plausible explanation of how the anomaly might have occurred anomaly_category Category of the anomaly's visual manifestation (see taxonomy below) bounding_box Coordinates of the bounding box demarcating the anomalous region severity 1–5 score: does the anomaly require immediate action? surprisal 1–5 score: how much does the situation deviate from expectations? complexity 1–5 score: how hard is the anomaly to detect? Anomaly Category Taxonomy Anomalies are categorized by how they visually manifest, inspired by MMBench's taxonomy of visual reasoning types: Category Description Example Entity Presence An object is present when it shouldn't be A black bear in an industrial building Entity Absence An expected object is missing A person using a cutter without protective gear Entity Attribute An object has an anomalous attribute (color, shape, label, orientation, usage) A snack packet opened from the wrong side Spatial Relation An object is incorrectly positioned relative to another Furniture blocking an emergency button Uniformity Breach A disruption in an expected uniform/symmetrical pattern One tile with a different orientation Textual Anomaly Text in the image conveys an unexpected or contradictory message A "KEEP RIGHT" sign with an arrow pointing left Dataset Creation Images were collected from four Reddit subreddits that specialize in content featuring unusual or uncommon situations: r/ocdtriggers r/mildlyconfusing r/mildlyinfuriating r/OSHA The top 1,000 posts from each subreddit were downloaded using the PRAW library. Images were filtered through both automatic and manual processes to remove: Unclear or ambiguous content Non-realistic images NSFW or sensitive content Images with text annotations, circles, or other overlaid marks Images below icon resolution Annotation proceeded in two rounds, with Amazon Mechanical Turk followed by Expert Verification & Consolidation, with 3 independent raters per anomaly for severity, surprisal, and complexity scores. Citation @inproceedings{bhagwatkar-etal-2025-cave, title = "{CAVE} : Detecting and Explaining Commonsense Anomalies in Visual Environments", author = "Bhagwatkar, Rishika and Montariol, Syrielle and Romanou, Angelika and Borges, Beatriz and Rish, Irina and Bosselut, Antoine", booktitle = "Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing", month = nov, year = "2025", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2025.emnlp-main.1379/", doi = "10.18653/v1/2025.emnlp-main.1379", pages = "27110--27151", } Acknowledgements The authors acknowledge support from Canada CIFAR AI Chair Program, Canada Excellence Research Chairs Program, Swiss National Science Foundation (No. 215390), Innosuisse (PFFS-21-29), EPFL Center for Imaging, Sony Group Corporation, and a Meta LLM Evaluation Research Grant. Computational resources were provided by MILA - Quebec AI Institute.
Rishika Bhagwatkar, Syrielle Montariol, Angelika Romanou, Beatriz Borges, Irina Rish, Antoine Bosselut
EMNLP2
2025 INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge
abstract
The performance differential of large language models (LLM) between languages hinders their effective deployment in many regions, inhibiting the potential economic and societal value of generative AI tools in many communities. However, the development of functional LLMs in many languages (i.e., multilingual LLMs) is bottlenecked by the lack of high-quality evaluation resources in languages other than English. Moreover, current practices in multilingual benchmark construction often translate English resources, ignoring the regional and cultural knowledge of the environments in which multilingual systems would be used. In this work, we construct an evaluation suite of 197,243 QA pairs from local exam sources to measure the capabilities of multilingual LLMs in a variety of regional contexts. Our novel resource, INCLUDE, is a comprehensive knowledge- and reasoning-centric benchmark across 44 written languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed.
Angelika Romanou, Negar Foroutan Eghlidi, Anna Sotnikova, Zeming Chen 0001, Sree Harsha Nelaturu, Shivalika Singh, Rishabh Maheshwary, Micol Altomare, Mohamed A. Haggag, Imanol Schlag, Marzieh Fadaee, Sara Hooker, Antoine Bosselut, Snegha A, Alfonso Amayuelas, Azril Hafizi Amirudin, Viraat Aryabumi, Danylo Boiko, Jenny Chim, Gal Cohen, Aditya Kumar Dalmia, Abraham Diress, Sharad Duwal, Daniil Dzenhaliou, Daniel Fernando Erazo Florez, Fabian Farestam, Joseph Marvin Imperial, Shayekh Bin Islam, Perttu Isotalo, Maral Jabbarishiviari, Börje Karlsson 0001, Eldar Khalilov, Christopher Klamm, Fajri Koto, Dominik Krzeminski, Gabriel Adriano de Melo, Syrielle Montariol, Yiyang Nan, Joel Niklaus, Jekaterina Novikova, Johan S. Obando-Ceron, Debjit Paul, Esther Ploeger, Jebish Purbey, Swati Rajwal, Selvan Sunitha Ravi, Sara Rydell, Roshan Santhosh, Drishti Sharma, Marjana Prifti Skenduli, Arshia Soltani Moakhar, Bardia Soltani Moakhar, Ran Tamir, Ayush K. Tarun, Azmine Toushik Wasi, Thenuka Ovin Weerasinghe, Serhan Yilmaz, Mike Zhang
ICLR38
2025 Intrinsic User-Centric Interpretability through Global Mixture of Experts
abstract
In human-centric settings like education or healthcare, model accuracy and model explainability are key factors for user adoption. Towards these two goals, intrinsically interpretable deep learning models have gained popularity, focusing on accurate predictions alongside faithful explanations. However, there exists a gap in the human-centeredness of these approaches, which often produce nuanced and complex explanations that are not easily actionable for downstream users. We present InterpretCC (interpretable conditional computation), a family of intrinsically interpretable neural networks at a unique point in the design space that optimizes for ease of human understanding and explanation faithfulness, while maintaining comparable performance to state-of-the-art models. InterpretCC achieves this through adaptive sparse activation of features before prediction, allowing the model to use a different, minimal set of features for each instance. We extend this idea into an interpretable, global mixture-of-experts (MoE) model that allows users to specify topics of interest, discretely separates the feature space for each data point into topical subnetworks, and adaptively and sparsely activates these topical subnetworks for prediction. We apply InterpretCC for text, time series and tabular data across several real-world datasets, demonstrating comparable performance with non-interpretable baselines and outperforming intrinsically interpretable baselines. Through a user study involving 56 teachers, InterpretCC explanations are found to have higher actionability and usefulness over other intrinsically interpretable approaches.
Vinitra Swamy, Syrielle Montariol, Julian Blackwell, Jibril Frej, Martin Jaggi, Tanja Käser
ICLR2
2025 PICLe: Pseudo-annotations for In-Context Learning in Low-Resource Named Entity Detection
abstract
Sepideh Mamooler, Syrielle Montariol, Alexander Mathis, Antoine Bosselut. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Sepideh Mamooler, Syrielle Montariol, Alexander Mathis, Antoine Bosselut
NAACL (Long Papers)2
2024 ConVQG: Contrastive Visual Question Generation with Multimodal Guidance
abstract
Asking questions about visual environments is a crucial way for intelligent agents to understand rich multi-faceted scenes, raising the importance of Visual Question Generation (VQG) systems. Apart from being grounded to the image, existing VQG systems can use textual constraints, such as expected answers or knowledge triplets, to generate focused questions. These constraints allow VQG systems to specify the question content or leverage external commonsense knowledge that can not be obtained from the image content only. However, generating focused questions using textual constraints while enforcing a high relevance to the image content remains a challenge, as VQG systems often ignore one or both forms of grounding. In this work, we propose Contrastive Visual Question Generation (ConVQG), a method using a dual contrastive objective to discriminate questions generated using both modalities from those based on a single one. Experiments on both knowledge-aware and standard VQG benchmarks demonstrate that ConVQG outperforms the state-of-the-art methods and generates image-grounded, text-guided, and knowledge-rich questions. Our human evaluation results also show preference for ConVQG questions compared to non-contrastive baselines.
Li Mi, Syrielle Montariol, Javiera Castillo-Navarro, Xianjie Dai, Antoine Bosselut, Devis Tuia
AAAI2
2024 ConGeo: Robust Cross-View Geo-Localization Across Ground View Variations
Li Mi, Chang Xu 0027, Javiera Castillo-Navarro, Syrielle Montariol, Wen Yang 0001, Antoine Bosselut, Devis Tuia
ECCV (14)4
2024 "Flex Tape Can't Fix That": Bias and Misinformation in Edited Language Models
abstract
Weight-based model editing methods update the parametric knowledge of language models post-training.However, these methods can unintentionally alter unrelated parametric knowledge representations, potentially increasing the risk of harm.In this work, we investigate how weight editing methods unexpectedly amplify model biases after edits.We introduce a novel benchmark dataset, SEESAW-CF, for measuring bias amplification of model editing methods for demographic traits such as race, geographic origin, and gender.We use SEESAW-CF to examine the impact of model editing on bias in five large language models.Our results demonstrate that edited models exhibit, to various degrees, more biased behavior for certain demographic groups than before they were edited, specifically becoming less confident in properties for Asian and African subjects.Additionally, editing facts about place of birth, country of citizenship, or gender has particularly negative effects on the model's knowledge about unrelated properties, such as field of work, a pattern observed across multiple models.
Karina Halevy, Anna Sotnikova, Badr AlKhamissi, Syrielle Montariol, Antoine Bosselut
EMNLP4
2024 Training Visual Language Models with Object Detection: Grounded Change Descriptions in Satellite Images
abstract
Recently, generalist Vision Language Models (VLMs) have shown exceptional progress in tasks previously dominated by specialized computer vision models. This becomes more prevalent when visual grounding capabilities, such as the ability to reason over input text and image to generate bounding boxes around objects, are required. However, how these capabilities transfer to specialized domains such as remote sensing remains understudied, despite the recent increase in specialized models for Earth observation. In this work, we evaluate how grounding visual entities – by generating bounding-box coordinates – affects VLM performance in satellite imagery. To this end, we create two instruction-following tasks sourced from the xBD dataset, describing changes due to natural disasters observed in satellite images. We fine-tune several instances of MiniGPTv2, an open-source VLM with grounding capabilities, and evaluate their performance under the "grounded" vs. "not grounded" settings. We find that generating bounding boxes to refer to visual entities increases performance in tasks related to objects in the image, but only when the number of entities in the image is limited.
João Luis Prado, Syrielle Montariol, Javiera Castillo-Navarro, Devis Tuia, Antoine Bosselut
IGARSS2
2024 Course Recommender Systems Need to Consider the Job Market
abstract
Current course recommender systems primarily leverage learner-course interactions, course content, learner preferences, and supplementary course details like instructor, institution, ratings, and reviews, to make their recommendation. However, these systems often overlook a critical aspect: the evolving skill demand of the job market. This paper focuses on the perspective of academic researchers, working in collaboration with the industry, aiming to develop a course recommender system that incorporates job market skill demands. In light of the job market's rapid changes and the current state of research in course recommender systems, we outline essential properties for course recommender systems to address these demands effectively, including explainable, sequential, unsupervised, and aligned with the job market and user's goals. Our discussion extends to the challenges and research questions this objective entails, including unsupervised skill extraction from job listings, course descriptions, and resumes, as well as predicting recommendations that align with learner objectives and the job market and designing metrics to evaluate this alignment. Furthermore, we introduce an initial system that addresses some existing limitations of course recommender systems using large Language Models (LLMs) for skill extraction and Reinforcement Learning (RL) for alignment with the job market. We provide empirical results using open-source data to demonstrate its effectiveness.
Jibril Frej, Anna Dai, Syrielle Montariol, Antoine Bosselut, Tanja Käser
SIGIR3
2023 CRoW: Benchmarking Commonsense Reasoning in Real-World Tasks
abstract
Recent efforts in natural language processing (NLP) commonsense reasoning research have yielded a considerable number of new datasets and benchmarks.However, most of these datasets formulate commonsense reasoning challenges in artificial scenarios that are not reflective of the tasks which real-world NLP systems are designed to solve.In this work, we present CROW, a manually-curated, multitask benchmark that evaluates the ability of models to apply commonsense reasoning in the context of six real-world NLP tasks.CROW is constructed using a multi-stage data collection pipeline that rewrites examples from existing datasets using commonsense-violating perturbations.We use CROWto study how NLP systems perform across different dimensions of commonsense knowledge, such as physical, temporal, and social reasoning.We find a significant performance gap when NLP systems are evaluated on CROWcompared to humans, showcasing that commonsense reasoning is far from being solved in real-world task settings.We make our dataset and leaderboard available to the research community.1 * Equal contribution 1 https://github.com/mismayil/crowDialogue Agent: Hi, would you like some free candies?Human: Sure.What are you handing these out for?Agent: Well, we're trying to gather some
Mete Ismayilzada, Debjit Paul, Syrielle Montariol, Mor Geva, Antoine Bosselut
EMNLP3
2023 CRAB: Assessing the Strength of Causal Relationships Between Real-world Events
abstract
Understanding narratives requires reasoning about the cause-and-effect relationships between events mentioned in the text.While existing foundation models yield impressive results in many NLP tasks requiring reasoning, it is unclear whether they understand the complexity of the underlying network of causal relationships of events in narratives.In this work, we present CRAB, a new Causal Reasoning Assessment Benchmark designed to evaluate causal understanding of events in real-world narratives.CRAB contains fine-grained, contextual causality annotations for ∼ 2.7K pairs of real-world events that describe various newsworthy event timelines (e.g., the acquisition of Twitter by Elon Musk).Using CRAB, we measure the performance of several large language models, demonstrating that most systems achieve poor performance on the task.Motivated by classical causal principles, we also analyze the causal structures of groups of events in CRAB, and find that models perform worse on causal reasoning when events are derived from complex causal structures compared to simple linear causal chains.We make our dataset and code available to the research community.
Angelika Romanou, Syrielle Montariol, Debjit Paul, Léo Laugier, Karl Aberer, Antoine Bosselut
EMNLP2
2023 Don't Start Your Data Labeling from Scratch: OpSaLa - Optimized Data Sampling Before Labeling
Andraz Pelicon, Syrielle Montariol, Petra Kralj Novak
IDA2
2022 Effectiveness of Data Augmentation and Pretraining for Improving Neural Headline Generation in Low-Resource Settings
abstract
We tackle the problem of neural headline generation in a low-resource setting, where only limited amount of data is available to train a model. We compare the ideal high-resource scenario on English with results obtained on a smaller subset of the same data and also run experiments on two small news corpora covering low-resource languages, Croatian and Estonian. Two options for headline generation in a multilingual low-resource scenario are investigated: a pretrained multilingual encoder-decoder model and a combination of two pretrained language models, one used as an encoder and the other as a decoder, connected with a cross-attention layer that needs to be trained from scratch. The results show that the first approach outperforms the second one by a large margin. We explore several data augmentation and pretraining strategies in order to improve the performance of both models and show that while we can drastically improve the second approach using these strategies, they have little to no effect on the performance of the pretrained encoder-decoder model. Finally, we propose two new measures for evaluating the performance of the models besides the classic ROUGE scores.
Matej Martinc, Syrielle Montariol, Lidia Pivovarova, Elaine Zosa
LREC2
2021 Measure and Evaluation of Semantic Divergence across Two Languages
abstract
Syrielle Montariol, Alexandre Allauzen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Syrielle Montariol, Alexandre Allauzen
ACL/IJCNLP (1)1
2021 Scalable and Interpretable Semantic Change Detection
abstract
Syrielle Montariol, Matej Martinc, Lidia Pivovarova. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Syrielle Montariol, Matej Martinc, Lidia Pivovarova
NAACL-HLT1