Antonios Anastasopoulos

dblp:148/9479 · also Antonis Anastasopoulos · DBLP profile ↗
← Back
69ranked-venue papers
6as first author
44since 2021 · last 2026
0000-0002-8544-246XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 62 · 6 first-author · 38 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 VIGNETTE: Socially Grounded Bias Evaluation for Vision-Language Models
abstract
While bias in large language models (LLMs) is well-studied, similar concerns in vision-language models (VLMs) have received comparatively less attention. Existing VLM bias studies often focus on portrait-style images and gender-occupation associations, overlooking broader and more complex social stereotypes and their implied harm. This work introduces VIGNETTE, a large-scale VQA benchmark with 30M+ images for evaluating bias in VLMs through a question-answering framework spanning four directions: factuality, perception, stereotyping, and decision making. Beyond narrowly-centered studies, we assess how VLMs interpret identities in contextualized settings, revealing how models make trait and capability assumptions and exhibit patterns of discrimination. Drawing from social psychology, we examine how VLMs connect visual identity cues to trait and role-based inferences, encoding social hierarchies, through biased selections. Our findings uncover subtle, multifaceted, and surprising stereotypical patterns, offering insights into how VLMs construct social meaning from inputs.
Chahat Raj, Bowen Wei, Aylin Caliskan, Antonios Anastasopoulos, Ziwei Zhu 0001
ACL (1)4
2026 TigerCoder: A Novel Suite of LLMs for Code Generation in Bangla
Antonios Anastasopoulos, Marcos Zampieri
LREC2
2025 Tracing L1 Interference in English Learner Writing: A Longitudinal Corpus with Error Annotations
abstract
Language transfer is an important topic of research in second language acquisition and computational linguistics.The availability of suitable learner corpora is paramount for the study of second language acquisition (SLA) and language transfer.However, curating learner corpora is a challenging endeavor as high quality learner data is rarely publicly available.This results in only a few such corpora available to the community.To address this important gap, in this paper we present LENS, a novel English learner corpus with longitudinal data which enables researchers to investigate language learning over time.LENS contains 687 instances written by speakers of 15 different L1s.We use LENS two perform two important tasks at the intersection of SLA and Computational Linguistics: (1) Native Language Identification (NLI); and (2) an evaluation of large language models as a tool for high-precision, semi-automated annotation of L1 interference features.1
Poorvi Acharya, J. Elizabeth Liebl, Dhiman Goswami, Kai North, Marcos Zampieri, Antonios Anastasopoulos
EMNLP6
2025 Graph Enhanced Trajectory Anomaly Detection
abstract
Trajectory anomaly detection is essential for identifying unusual and unexpected movement patterns in applications ranging from intelligent transportation systems to urban safety and fraud prevention. Existing methods only consider limited aspects of the trajectory nature and its movement space by treating trajectories as sequences of sampled locations, with sampling determined by positioning technology, e.g., GPS, or by high-level abstractions such as staypoints. Trajectories are analyzed in Euclidean space, neglecting the constraints and connectivity information of the underlying movement network, e.g., road or transit networks. The proposed Graph Enhanced Trajectory Anomaly Detection (GETAD) framework tightly integrates road network topology, segment semantics, and historical travel patterns to model trajectory data. GETAD uses a Graph Attention Network to learn road-aware embeddings that capture both physical attributes and transition behavior, and augments these with graph-based positional encodings that reflect the spatial layout of the road network. A Transformer-based decoder models sequential movement, while a multiobjective loss function combining autoregressive prediction and supervised link prediction ensures realistic and structurally coherent representations. To improve the robustness of anomaly detection, we introduce Confidence Weighted Negative Log Likelihood (CW NLL), an anomaly scoring function that emphasizes high-confidence deviations. Experiments on real-world and synthetic datasets demonstrate that GETAD achieves consistent improvements over existing methods, particularly in detecting subtle anomalies in road-constrained environments. These results highlight the benefits of incorporating graph structure and contextual semantics into trajectory modeling, enabling more precise and context-aware anomaly detection.
Jonathan Mbuya, Dieter Pfoser, Antonios Anastasopoulos
SIGSPATIAL/GIS3
2025 The ML-SUPERB 2.0 Challenge: Towards Inclusive ASR Benchmarking for All Language Varieties
Chutong Meng, Jiatong Shi, Martijn Bartelds, Shih-Heng Wang, Hsiu-Hsuan Wang, Rafael Mosquera, Sara Hincapie, Daniel Jurafsky, Antonios Anastasopoulos, Hung-yi Lee, Karen Livescu, Shinji Watanabe 0001
INTERSPEECH10
2025 Script-Agnosticism and its Impact on Language Identification for Dravidian Languages
abstract
Milind Agarwal, Joshua Otten, Antonios Anastasopoulos. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Milind Agarwal, Joshua Otten, Antonios Anastasopoulos
NAACL (Long Papers)3
2025 Follow the Beaten Path: The Role of Route Patterns on Vision-Language Navigation Agents Generalization Abilities
abstract
Kourosh T Baghaei, Dieter Pfoser, Antonios Anastasopoulos. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Kourosh T. Baghaei, Dieter Pfoser, Antonios Anastasopoulos
NAACL (Long Papers)3
2025 mHumanEval - A Multilingual Benchmark to Evaluate Large Language Models for Code Generation
abstract
Nishat Raihan, Antonios Anastasopoulos, Marcos Zampieri. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Antonios Anastasopoulos, Marcos Zampieri
NAACL (Long Papers)2
2025 Crossroads of Continents: Automated Artifact Extraction for Cultural Adaptation with Large Multimodal Models
Anjishnu Mukherjee, Ziwei Zhu 0001, Antonios Anastasopoulos
WACV3
2024 DIALECTBENCH: An NLP Benchmark for Dialects, Varieties, and Closely-Related Languages
abstract
Fahim Faisal, Orevaoghene Ahia, Aarohi Srivastava, Kabir Ahuja, David Chiang, Yulia Tsvetkov, Antonios Anastasopoulos. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Fahim Faisal, Orevaoghene Ahia, Aarohi Srivastava, Kabir Ahuja, David Chiang 0001, Yulia Tsvetkov, Antonios Anastasopoulos
ACL (1)7
2024 Breaking Bias, Building Bridges: Evaluation and Mitigation of Social Biases in LLMs via Contact Hypothesis
abstract
Large Language Models (LLMs) perpetuate social biases, reflecting prejudices in their training data and reinforcing societal stereotypes and inequalities. Our work explores the potential of the Contact Hypothesis, a concept from social psychology for debiasing LLMs. We simulate various forms of social contact through LLM prompting to measure their influence on the model’s biases, mirroring how intergroup interactions can reduce prejudices in social contexts. We create a dataset of 108,000 prompts following a principled approach replicating social contact to measure biases in three LLMs (LLaMA 2, Tulu, and NousHermes) across 13 social bias dimensions. We propose a unique debiasing technique, Social Contact Debiasing (SCD), that instruction-tunes these models with unbiased responses to prompts. Our research demonstrates that LLM responses exhibit social biases when subject to contact probing, but more importantly, these biases can be significantly reduced by up to 40% in 1 epoch of instruction tuning LLaMA 2 following our SCD strategy.
Chahat Raj, Anjishnu Mukherjee, Aylin Caliskan, Antonios Anastasopoulos, Ziwei Zhu 0001
AIES (1)4
2024 Language and Speech Technology for Central Kurdish Varieties
abstract
Kurdish, an Indo-European language spoken by over 30 million speakers, is considered a dialect continuum and known for its diversity in language varieties. Previous studies addressing language and speech technology for Kurdish handle it in a monolithic way as a macro-language, resulting in disparities for dialects and varieties for which there are few resources and tools available. In this paper, we take a step towards developing resources for language and speech technology for varieties of Central Kurdish, creating a corpus by transcribing movies and TV series as an alternative to fieldwork. Additionally, we report the performance of machine translation, automatic speech recognition, and language identification as downstream tasks evaluated on Central Kurdish subdialects. Data and models are publicly available under an open license at https://github.com/sinaahmadi/CORDI.
Sina Ahmadi, Daban Q. Jaff, Md Mahfuz Ibn Alam, Antonios Anastasopoulos
LREC/COLING4
2024 SALSA: Salience-Based Switching Attack for Adversarial Perturbations in Fake News Detection Models
Chahat Raj, Anjishnu Mukherjee, Hemant Purohit, Antonios Anastasopoulos, Ziwei Zhu 0001
ECIR (5)4
2024 Birdie: Advancing State Space Language Modeling with Dynamic Mixtures of Training Objectives
abstract
Efficient state space models (SSMs), including linear recurrent neural networks and linear attention variants, have emerged as potential alternative language models to Transformers.While efficient, SSMs struggle with tasks requiring in-context retrieval, such as text copying and associative recall, limiting their usefulness in practical settings.Prior work on how to meet this challenge has focused on the internal model architecture and not investigated the role of the training procedure.This paper proposes a new training procedure that improve the performance of SSMs on retrieval-intensive tasks.This novel pre-training procedure combines a bidirectional processing of the input with dynamic mixtures of pre-training objectives to improve the utilization of the SSM's fixed-size state.Our experimental evaluations show that this procedure significantly improves performance on retrieval-intensive tasks that challenge current SSMs, such as phone book lookup, long paragraph question-answering, and infilling tasks.Our findings offer insights into a new direction to advance the training of SSMs to close the performance gap with Transformers.
Sam Blouir, Jimmy T. H. Smith, Antonios Anastasopoulos, Amarda Shehu
EMNLP3
2024 The LLM Effect: Are Humans Truly Using LLMs, or Are They Being Influenced By Them Instead?
abstract
Large Language Models (LLMs) have shown capabilities close to human performance in various analytical tasks, leading researchers to use them for time and labor-intensive analyses.However, their capability to handle highly specialized and open-ended tasks in domains like policy studies remains in question.This paper investigates the efficiency and accuracy of LLMs in specialized tasks through a structured user study focusing on Human-LLM partnership.The study, conducted in two stages-Topic Discovery and Topic Assignment-integrates LLMs with expert annotators to observe the impact of LLM suggestions on what is usually human-only analysis.Results indicate that LLM-generated topic lists have significant overlap with human generated topic lists, with minor hiccups in missing document-specific topics.However, LLM suggestions may significantly improve task completion speed, but at the same time introduce anchoring bias, potentially affecting the depth and nuance of the analysis, raising a critical question about the trade-off between increased efficiency and the risk of biased analysis.1 * Equal contribution.
Alexander S. Choi, Syeda Sabrina Akter, JP Singh, Antonios Anastasopoulos
EMNLP4
2024 Back to School: Translation Using Grammar Books
abstract
Machine translation systems for high resource languages perform exceptionally well and produce high quality translations.Unfortunately, the vast majority of languages lack the quantity of parallel sentences needed to train such systems.These under-represented languages are not entirely without resources, as bilingual dictionaries and grammar books may be available as linguistic reference material.With current large language models (LLMs) supporting near book-length contexts, we can use the available material to ensure advancements are shared among all of the world's languages.In this paper, we use dictionaries and grammar books to improve machine translation.We evaluate on 16 typologically diverse low-resource languages, showing encouraging improvements.1
Jonathan Hus, Antonios Anastasopoulos
EMNLP2
2024 Urban Mobility Assessment Using LLMs
abstract
In urban science, understanding mobility patterns and analyzing how people move around cities helps improve the overall quality of life and supports the development of more livable, efficient, and sustainable urban areas. A challenging aspect of this work is the collection of mobility data through user tracking or travel surveys, given the associated privacy concerns, noncompliance, and high cost. This work proposes an innovative AI-based approach for synthesizing travel surveys by prompting large language models (LLMs), aiming to leverage their vast amount of relevant background knowledge and text generation capabilities. Our study evaluates the effectiveness of this approach across various U.S. metropolitan areas by comparing the results against existing survey data at different granularity levels. These levels include (i) pattern level, which compares aggregated metrics such as the average number of locations traveled and travel time, (ii) trip level, which focuses on comparing trips as whole units using transition probabilities, and (iii) activity chain level, which examines the sequence of locations visited by individuals. Our work covers several proprietary and open-source LLMs, revealing that open-source base models like Llama-2, when fine-tuned on even a limited amount of actual data, can generate synthetic data that closely mimics the actual travel survey data and, as such, provides an argument for using such data in mobility studies.
Prabin Bhandari, Antonios Anastasopoulos, Dieter Pfoser
SIGSPATIAL/GIS2
2024 Trajectory Anomaly Detection with Language Models
abstract
This paper presents a novel approach for trajectory anomaly detection using an autoregressive causal-attention model, termed LM-TAD. This method leverages the similarities between language statements and trajectories, both of which consist of ordered elements requiring coherence through external rules and contextual variations. By treating trajectories as sequences of tokens, our model learns the probability distributions over trajectories, enabling the identification of anomalous locations with high precision. We incorporate user-specific tokens to account for individual behavior patterns, enhancing anomaly detection tailored to user context. Our experiments demonstrate the effectiveness of LM-TAD on both synthetic and real-world datasets. In particular, the model outperforms existing methods on the Pattern of Life (PoL) dataset by detecting user-contextual anomalies and achieves competitive results on the Porto taxi dataset, highlighting its adaptability and robustness. Additionally, we introduce the use of perplexity and surprisal rate metrics for detecting outliers and pinpointing specific anomalous locations within trajectories. The LM-TAD framework supports various trajectory representations, including GPS coordinates, staypoints, and activity types, proving its versatility in handling diverse trajectory data. Moreover, our approach is well-suited for online trajectory anomaly detection, significantly reducing computational latency by caching key-value states of the attention mechanism, thereby avoiding repeated computations. The code to reproduce experiments in this paper can be found at the following link: https://github.com/jonathankabala/LMTAD.
Jonathan Mbuya, Dieter Pfoser, Antonios Anastasopoulos
SIGSPATIAL/GIS3
2024 Enhancing End-to-End Conversational Speech Translation Through Target Language Context Utilization
abstract
Incorporating longer context has been shown to benefit machine translation, but the inclusion of context in end-to-end speech translation (E2E-ST) remains under-studied. To bridge this gap, we introduce target language context in E2E-ST, enhancing coherence and overcoming memory constraints of extended audio segments. Additionally, we propose context dropout to ensure robustness to the absence of context, and further improve performance by adding speaker information. Our proposed contextual E2E-ST outperforms the isolated utterance-based E2E-ST approach. Lastly, we demonstrate that in conversational speech, contextual information primarily contributes to capturing context style, as well as resolving anaphora and named entities.
Amir Hussein, Brian Yan, Antonios Anastasopoulos, Shinji Watanabe 0001, Sanjeev Khudanpur
ICASSP3
2024 Cross-Modal Multi-Tasking for Speech-to-Text Translation via Hard Parameter Sharing
abstract
Recent works in end-to-end speech-to-text translation (ST) have proposed multi-tasking methods with soft parameter sharing which leverage machine translation (MT) data via secondary encoders that map text inputs to an eventual cross-modal representation. In this work, we instead propose a ST/MT multi-tasking framework with hard parameter sharing in which all model parameters are shared cross-modally. Our method reduces the speech-text modality gap via a pre-processing stage which converts speech and text inputs into two discrete token sequences of similar length – this allows models to indiscriminately process both modalities simply using a joint vocabulary. With experiments on MuST-C, we demonstrate that our multi-tasking framework improves attentional encoder-decoder, Connectionist Temporal Classification (CTC), transducer, and joint CTC/attention models by an average of +0.5 BLEU without any external MT data. Further, we show that this framework incorporates external MT data, yielding +0.8 BLEU, and also improves transfer learning from pre-trained textual models, yielding +1.8 BLEU.1
Brian Yan, Xuankai Chang, Antonios Anastasopoulos, Yuya Fujita, Shinji Watanabe 0001
ICASSP3
2024 Speech Recognition for Greek Dialects: A Challenging Benchmark
Socrates Vakirtzian, Chara Tsoukala, Stavros Bompolas, Katerina Mouzou, Vivian Stamou, Georgios Paraskevopoulos, Antonios Dimakis, Stella Markantonatou, Angela Ralli, Antonios Anastasopoulos
INTERSPEECH10
2024 Global Gallery: The Fine Art of Painting Culture Portraits through Multilingual Instruction Tuning
abstract
Anjishnu Mukherjee, Aylin Caliskan, Ziwei Zhu, Antonios Anastasopoulos. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Anjishnu Mukherjee, Aylin Caliskan, Ziwei Zhu 0001, Antonios Anastasopoulos
NAACL-HLT4
2024 Clinical risk prediction using language models: benefits and considerations
abstract
OBJECTIVE: The use of electronic health records (EHRs) for clinical risk prediction is on the rise. However, in many practical settings, the limited availability of task-specific EHR data can restrict the application of standard machine learning pipelines. In this study, we investigate the potential of leveraging language models (LMs) as a means to incorporate supplementary domain knowledge for improving the performance of various EHR-based risk prediction tasks. METHODS: We propose two novel LM-based methods, namely "LLaMA2-EHR" and "Sent-e-Med." Our focus is on utilizing the textual descriptions within structured EHRs to make risk predictions about future diagnoses. We conduct a comprehensive comparison with previous approaches across various data types and sizes. RESULTS: Experiments across 6 different methods and 3 separate risk prediction tasks reveal that employing LMs to represent structured EHRs, such as diagnostic histories, results in significant performance improvements when evaluated using standard metrics such as area under the receiver operating characteristic (ROC) curve and precision-recall (PR) curve. Additionally, they offer benefits such as few-shot learning, the ability to handle previously unseen medical concepts, and adaptability to various medical vocabularies. However, it is noteworthy that outcomes may exhibit sensitivity to a specific prompt. CONCLUSION: LMs encompass extensive embedded knowledge, making them valuable for the analysis of EHRs in the context of risk prediction. Nevertheless, it is important to exercise caution in their application, as ongoing safety concerns related to LMs persist and require continuous consideration.
Angeela Acharya, Sulabh Shrestha, Anyi Chen, Joseph Conte, Sanja Avramovic, Siddhartha Sikdar, Antonios Anastasopoulos, Sanmay Das
J. Am. Medical Informatics Assoc.7
2023 Script Normalization for Unconventional Writing of Under-Resourced Languages in Bilingual Communities
abstract
The wide accessibility of social media has provided linguistically under-represented communities with an extraordinary opportunity to create content in their native languages.This, however, comes with certain challenges in script normalization, particularly where the speakers of a language in a bilingual community rely on another script or orthography to write their native language.This paper addresses the problem of script normalization for several such languages that are mainly written in a Perso-Arabic script.Using synthetic data with various levels of noise and a transformerbased model, we demonstrate that the problem can be effectively remediated.We conduct a small-scale evaluation of real data as well.Our experiments indicate that script normalization is also beneficial to improve the performance of downstream tasks such as machine translation and language identification. 1 Language Unconventional script Unconventional writing Conventional writing Gilaki Persian ‫زﻧﻦ‬ ‫ﮔﺐ‬ ‫ﺟﯽ‬ ‫اون‬ ‫ﮔﻴﻠﮑﻦ‬ ‫ﮔﻪ‬ ‫ﻫﻴﺴﻪ‬ ‫ﻧﻢ‬ ‫زون‬ ‫ﯾﺘﻪ‬ ‫زﻧﻦ‬ ‫ﮔﺐ‬ ‫ﺟﻲ‬ ٚ ‫اۊن‬ ‫ﮔﻴﻠﮑﺆن‬ ‫ﮔﻪ‬ ‫ﻫﻴﺴﻪ‬ ‫ﻧﺆم‬ ٚ ‫زوؤن‬ ‫ﯾﺘﻪ‬ Kashmiri Urdu ‫#"۔‬ $% & ' ( % ) * + ,"-.% / , .0 % 1 "-2( % 3 ‫َر۔‬ ‫جانو‬ ‫ُرٲس4‬ ‫و‬ ‫َکھ‬ ‫ا‬ ُ ‫چھ‬ ‫ٛور‬ ‫بر‬ Kurmanji Arabic ‫دا‬ ‫دهوك‬ ‫پارزكار‬ ‫بةرثوا‬ ‫اامدي‬ ‫قايمقام‬ ‫دا‬ ‫دهۆکێ‬ ‫پارێزگارێ‬ ‫بەرسڤا‬ ‫ئامێدیێ‬ ‫قایمقامێ‬ Sorani Arabic ‫دةويت‬ ‫فهديان‬ ‫ديارة‬ ‫شانؤوة‬ ‫يةكةم‬ ‫لة‬ ‫هةر‬ ‫دەوێت‬ ‫فەهەدیان‬ ‫دیارە‬ ‫شانۆوە‬ ‫یەکەم‬ ‫لە‬ ‫هەر‬ Sindhi Urdu 5 6ٔ 8 % 9 $ :% ; < =% > '% ? 5 @",-A B % C D 5 6 E F G H $% I G JG % K -G LA( M % N < = O
Sina Ahmadi, Antonios Anastasopoulos
ACL (1)2
2023 BIG-C: a Multimodal Multi-Purpose Dataset for Bemba
abstract
We present BIG-C (Bemba Image Grounded Conversations), a large multimodal dataset for Bemba.While Bemba is the most populous language of Zambia, it exhibits a dearth of resources which render the development of language technologies or language processing research almost impossible.The dataset is comprised of multi-turn dialogues between Bemba speakers based on images, transcribed and translated into English.There are more than 92,000 utterances/sentences, amounting to more than 180 hours of audio data with corresponding transcriptions and English translations.We also provide baselines on speech recognition (ASR), machine translation (MT) and speech translation (ST) tasks, and sketch out other potential future multimodal uses of our dataset.We hope that by making the dataset available to the research community, 1 this work will foster research and encourage collaboration across the language, speech, and vision communities especially for languages outside the "traditionally" used high-resourced ones.
Claytone Sikasote, Eunice Mukonde, Md Mahfuz Ibn Alam, Antonios Anastasopoulos
ACL (1)4
2023 Language Generation Models Can Cause Harm: So What Can We Do About It? An Actionable Survey
abstract
Sachin Kumar, Vidhisha Balachandran, Lucille Njoo, Antonios Anastasopoulos, Yulia Tsvetkov. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Sachin Kumar 0009, Vidhisha Balachandran, Lucille Njoo, Antonios Anastasopoulos, Yulia Tsvetkov
EACL4
2023 LIMIT: Language Identification, Misidentification, and Translation using Hierarchical Models in 350+ Languages
abstract
Knowing the language of an input text/audio is a necessary first step for using almost every NLP tool such as taggers, parsers, or translation systems.Language identification is a wellstudied problem, sometimes even considered solved; in reality, due to lack of data and computational challenges, current systems cannot accurately identify most of the world's 7000 languages.To tackle this bottleneck, we first compile a corpus, MCS-350, of 50K multilingual and parallel children's stories in 350+ languages.MCS-350 can serve as a benchmark for language identification of short texts and for 1400+ new translation directions in lowresource Indian and African languages.Second, we propose a novel misprediction-resolution hierarchical model, LIMIT, for language identification that reduces error by 55% (from 0.71 to 0.32) on our compiled children's stories dataset and by 40% (from 0.23 to 0.14) on the FLORES-200 benchmark.Our method can expand language identification coverage into low-resource languages by relying solely on systemic misprediction patterns, bypassing the need to retrain large models from scratch. 1
Milind Agarwal, Md Mahfuz Ibn Alam, Antonios Anastasopoulos
EMNLP3
2023 Global Voices, Local Biases: Socio-Cultural Prejudices across Languages
abstract
Human biases are ubiquitous but not uniform: disparities exist across linguistic, cultural, and societal borders.As large amounts of recent literature suggest, language models (LMs) trained on human data can reflect and often amplify the effects of these social biases.However, the vast majority of existing studies on bias are heavily skewed towards Western and European languages.In this work, we scale the Word Embedding Association Test (WEAT) to 24 languages, enabling broader studies and yielding interesting findings about LM bias.We additionally enhance this data with culturally relevant information for each language, capturing local contexts on a global scale.Further, to encompass more widely prevalent societal biases, we examine new bias dimensions across toxicity, ableism, and more.Moreover, we delve deeper into the Indian linguistic landscape, conducting a comprehensive regional bias analysis across six prevalent Indian languages.Finally, we highlight the significance of these social biases and the new dimensions through an extensive comparison of embedding methods, reinforcing the need to address them in pursuit of more equitable language models. 1
Anjishnu Mukherjee, Chahat Raj, Ziwei Zhu 0001, Antonios Anastasopoulos
EMNLP4
2023 GlobalBench: A Benchmark for Global Progress in Natural Language Processing
abstract
Yueqi Song, Simran Khanuja, Pengfei Liu, Fahim Faisal, Alissa Ostapenko, Genta Winata, Alham Aji, Samuel Cahyawijaya, Yulia Tsvetkov, Antonios Anastasopoulos, Graham Neubig. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Yueqi Song, Simran Khanuja, Pengfei Liu 0003, Fahim Faisal, Alissa Ostapenko, Genta Indra Winata, Alham Fikri Aji, Samuel Cahyawijaya, Yulia Tsvetkov, Antonios Anastasopoulos, Graham Neubig
EMNLP10
2023 Are Large Language Models Geospatially Knowledgeable?
abstract
Despite the impressive performance of Large Language Models (LLM) for various natural language processing tasks, little is known about their comprehension of geographic data and related ability to facilitate informed geospatial decision-making. This paper investigates the extent of geospatial knowledge, awareness, and reasoning abilities encoded within such pretrained LLMs. With a focus on autoregressive language models, we devise experimental approaches related to (i) probing LLMs for geo-coordinates to assess geospatial knowledge, (ii) using geospatial and non-geospatial prepositions to gauge their geospatial awareness, and (iii) utilizing a multidimensional scaling (MDS) experiment to assess the models' geospatial reasoning capabilities and to determine locations of cities based on prompting. Our results confirm that it does not only take larger but also more sophisticated LLMs to synthesize geospatial knowledge from textual information. As such, this research contributes to understanding the potential and limitations of LLMs in dealing with geospatial information.
Prabin Bhandari, Antonios Anastasopoulos, Dieter Pfoser
SIGSPATIAL/GIS2
2023 Towards a Universal Python: Translating the Natural Modality of Python into Other Human Languages
abstract
The Python programming language plays a large role in computer science today, both in industry and education. While the pseudo-code nature of its keywords and built-in functions/modules makes programming easy to learn for English speakers, non-English speakers do not have this advantage. Our goal is to further the democratization of computer science, allowing anyone to code in their native language, anywhere. This paper describes our vision for realizing this goal by automatically translating Python (keywords, error messages, identifiers) into other human languages, leveraging recent developments in machine translation and language technologies in general. As a first step, we introduce a preliminary multi-lingual Python tool that enables a user to code, translate, and execute Python in 5 additional languages, as well as a roadmap for the future development of our automated framework.
Joshua Otten, Antonios Anastasopoulos, Kevin Moran
ICSME2
2023 Zambezi Voice: A Multilingual Speech Corpus for Zambian Languages
abstract
This work introduces ZambeziVoice, an open-source multilingual speech resource for Zambian languages.It contains two collections of datasets: unlabelled audio recordings of radio news and talk shows programs (160 hours) and labelled data (over 80 hours) consisting of read speech recorded from text sourced from publicly available literature books.The dataset is created for speech recognition but can be extended to multilingual speech processing research for both supervised and unsupervised learning approaches.To our knowledge, this is the first multilingual speech dataset created for Zambian languages.We exploit pretraining and cross-lingual transfer learning by finetuning the Wav2Vec2.0large-scale multilingual pretrained model to build end-to-end (E2E) speech recognition models for our baseline models.The dataset is released publicly under a Creative
Claytone Sikasote, Kalinda Siaminwe, Stanly Mwape, Bangiwe Zulu, Mofya Phiri, Martin Phiri, David Zulu, Mayumbo Nyirenda, Antonios Anastasopoulos
INTERSPEECH9
2022 Systematic Inequalities in Language Technology Performance across the World's Languages
abstract
Natural language processing (NLP) systems have become a central technology in communication, education, medicine, artificial intelligence, and many other domains of research and development.While the performance of NLP methods has grown enormously over the last decade, this progress has been restricted to a minuscule subset of the world's ≈6,500 languages.We introduce a framework for estimating the global utility of language technologies as revealed in a comprehensive snapshot of recent publications in NLP.Our analyses involve the field at large, but also more in-depth studies on both user-facing technologies (machine translation, language understanding, question answering, text-to-speech synthesis) as well as foundational NLP tasks (dependency parsing, morphological inflection).In the process, we (1) quantify disparities in the current state of NLP research, (2) explore some of its associated societal and academic factors, and (3) produce tailored recommendations for evidencebased policy making aimed at promoting more global and equitable language technologies.1
Damián E. Blasi, Antonios Anastasopoulos, Graham Neubig
ACL (1)2
2022 Dataset Geography: Mapping Language Data to Language Users
abstract
As language technologies become more ubiquitous, there are increasing efforts towards expanding the language diversity and coverage of natural language processing (NLP) systems.Arguably, the most important factor influencing the quality of modern NLP systems is data availability.In this work, we study the geographical representativeness of NLP datasets, aiming to quantify if and by how much do NLP datasets match the expected needs of the language speakers.In doing so, we use entity recognition and linking systems, presenting an approach for good-enough entity linking without entity recognition first.Last, we explore some geographical and economic factors that may explain the observed dataset distributions. 1
Fahim Faisal, Yinkai Wang, Antonios Anastasopoulos
ACL (1)3
2022 Cross-Lingual Text Classification of Transliterated Hindi and Malayalam
abstract
Transliteration is very common on social media, but transliterated text is not adequately handled by modern neural models for various NLP tasks. In this work, we combine data augmentation approaches with a Teacher-Student training scheme to address this issue in a cross-lingual transfer setting for fine-tuning state-of-the-art pre-trained multilingual language models such as mBERT and XLM-R. We evaluate our method on transliterated Hindi and Malayalam, also introducing new datasets for benchmarking on real-world scenarios: one on sentiment classification in transliterated Malayalam, and another on crisis tweet classification in transliterated Hindi and Malayalam (related to the 2013 North India and 2018 Kerala floods). Our method yielded an average improvement of +5.6% on mBERT and +4.7% on XLM-R in F1 scores over their strong baselines.1
Jitin Krishnan, Antonios Anastasopoulos, Hemant Purohit, Huzefa Rangwala
IEEE Big Data2
2022 PROBER: A System for Real-time Propaganda Behavior Analytics on Social Media and Web Data Streams
abstract
Social media and online platforms provide a public space for many people to share opinions. Social media has numerous benefits to society; however, previous research has identified that individuals use social media for propagandizing purposes which can be detrimental to society, especially during humanitarian crises. Therefore, communities must look into this content to understand and effectively mitigate propaganda, especially when social media messages contain targeted hate or fake/disinformation. In this paper, we propose a human-centered system called PROBER, for propaganda behavior analytics in social and web data streams, which the relevant authorities, such as government institutions, could use for operational decision support and informing policy analysts for crisis management.
Yasas Senarath, Antonios Anastasopoulos, Tonya Thornton, Hemant Purohit
IEEE Big Data2
2022 UniMorph 4.0: Universal Morphology
abstract
The Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological inflection tables for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation, and a type-level resource of annotated data in diverse languages realizing that schema. This paper presents the expansions and improvements on several fronts that were made in the last couple of years (since McCarthy et al. (2020)). Collaborative efforts by numerous linguists have added 66 new languages, including 24 endangered languages. We have implemented several improvements to the extraction pipeline to tackle some issues, e.g., missing gender and macrons information. We have amended the schema to use a hierarchical structure that is needed for morphological phenomena like multiple-argument agreement and case stacking, while adding some missing morphological features to make the schema more inclusive. In light of the last UniMorph release, we also augmented the database with morpheme segmentation for 16 languages. Lastly, this new release makes a push towards inclusion of derivational morphology in UniMorph by enriching the data and annotation schema with instances representing derivational processes from MorphyNet.
Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieras, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina J. Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Lane 0002, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóga, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer C. White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo Maria Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar 0002, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Tucker Prud'hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova
LREC34
2022 BembaSpeech: A Speech Recognition Corpus for the Bemba Language
abstract
We present a preprocessed, ready-to-use automatic speech recognition corpus, BembaSpeech, consisting over 24 hours of read speech in the Bemba language, a written but low-resourced language spoken by over 30% of the population in Zambia. To assess its usefulness for training and testing ASR systems for Bemba, we explored different approaches; supervised pre-training (training from scratch), cross-lingual transfer learning from a monolingual English pre-trained model using DeepSpeech on the portion of the dataset and fine-tuning large scale self-supervised Wav2Vec2.0 based multilingual pre-trained models on the complete BembaSpeech corpus. From our experiments, the 1 billion XLS-R parameter model gives the best results. The model achieves a word error rate (WER) of 32.91%, results demonstrating that model capacity significantly improves performance and that multilingual pre-trained models transfers cross-lingual acoustic representation better than monolingual pre-trained English model on the BembaSpeech for the Bemba ASR. Lastly, results also show that the corpus can be used for building ASR systems for Bemba language.
Claytone Sikasote, Antonios Anastasopoulos
LREC2
2021 When is Wall a Pared and when a Muro?: Extracting Rules Governing Lexical Selection
abstract
Learning fine-grained distinctions between vocabulary items is a key challenge in learning a new language.For example, the noun "wall" has different lexical manifestations in Spanish -"pared" refers to an indoor wall while "muro" refers to an outside wall.However, this variety of lexical distinction may not be obvious to non-native learners unless the distinction is explained in such a way.In this work, we present a method for automatically identifying fine-grained lexical distinctions, and extracting concise descriptions explaining these distinctions in a human-and machine-readable format.We confirm the quality of these extracted descriptions in a language learning setup for two languages, Spanish and Greek, where we use them to teach non-native speakers when to translate a given ambiguous word into its different possible translations.Code and data are publicly released here.1
Aditi Chaudhary, Kayo Yin, Antonios Anastasopoulos, Graham Neubig
EMNLP (1)3
2021 Evaluating the Morphosyntactic Well-formedness of Generated Texts
abstract
Adithya Pratapa, Antonios Anastasopoulos, Shruti Rijhwani, Aditi Chaudhary, David R. Mortensen, Graham Neubig, Yulia Tsvetkov. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Adithya Pratapa, Antonios Anastasopoulos, Shruti Rijhwani, Aditi Chaudhary, David R. Mortensen, Graham Neubig, Yulia Tsvetkov
EMNLP (1)2
2021 Phoneme Recognition Through Fine Tuning of Phonetic Representations: A Case Study on Luhya Language Varieties
abstract
Models pre-trained on multiple languages have shown significant promise for improving speech recognition, particularly for low-resource languages. In this work, we focus on phoneme recognition using Allosaurus, a method for multilingual recognition based on phonetic annotation, which incorporates phonological knowledge through a language-dependent allophone layer that associates a universal narrow phone-set with the phonemes that appear in each language. To evaluate in a challenging real-world scenario, we curate phone recognition datasets for Bukusu and Saamia, two varieties of the Luhya language cluster of western Kenya and eastern Uganda. To our knowledge, these datasets are the first of their kind. We carry out similar experiments on the dataset of an endangered Tangkhulic language, East Tusom, a Tibeto-Burman language variety spoken mostly in India. We explore both zero-shot and few-shot recognition by fine-tuning using datasets of varying sizes (10 to 1000 utterances). We find that fine-tuning of Allosaurus, even with just 100 utterances, leads to significant improvements in phone error rates.
Kathleen Siminyu, Antonios Anastasopoulos, David R. Mortensen, Michael R. Marlo, Graham Neubig
Interspeech3
2021 When Being Unseen from mBERT is just the Beginning: Handling New Languages With Multilingual Language Models
abstract
Benjamin Muller, Antonios Anastasopoulos, Benoît Sagot, Djamé Seddah. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Benjamin Muller, Antonios Anastasopoulos, Benoît Sagot, Djamé Seddah
NAACL-HLT2
2021 Reducing Confusion in Active Learning for Part-Of-Speech Tagging
Aditi Chaudhary, Zaid Sheikh, Antonios Anastasopoulos, Graham Neubig
Trans. Assoc. Comput. Linguistics3
2021 Lexically Aware Semi-Supervised Learning for OCR Post-Correction
abstract
Abstract Much of the existing linguistic data in many languages of the world is locked away in non- digitized books and documents. Optical character recognition (OCR) can be used to produce digitized text, and previous work has demonstrated the utility of neural post-correction methods that improve the results of general- purpose OCR systems on recognition of less- well-resourced languages. However, these methods rely on manually curated post- correction data, which are relatively scarce compared to the non-annotated raw images that need to be digitized. In this paper, we present a semi-supervised learning method that makes it possible to utilize these raw images to improve performance, specifically through the use of self-training, a technique where a model is iteratively trained on its own outputs. In addition, to enforce consistency in the recognized vocabulary, we introduce a lexically aware decoding method that augments the neural post-correction model with a count-based language model constructed from the recognized texts, implemented using weighted finite-state automata (WFSA) for efficient and effective decoding. Results on four endangered languages demonstrate the utility of the proposed method, with relative error reductions of 15%–29%, where we find the combination of self-training and lexically aware decoding essential for achieving consistent improvements.1
Shruti Rijhwani, Daisy Rosenblum, Antonios Anastasopoulos, Graham Neubig
Trans. Assoc. Comput. Linguistics3
2020 Towards Minimal Supervision BERT-Based Grammar Error Correction (Student Abstract)
abstract
Current grammatical error correction (GEC) models typically consider the task as sequence generation, which requires large amounts of annotated data and limit the applications in data-limited settings. We try to incorporate contextual information from pre-trained language model to leverage annotation and benefit multilingual scenarios. Results show strong potential of Bidirectional Encoder Representations from Transformers (BERT) in grammatical error correction task.
Yiyuan Li, Antonios Anastasopoulos, Alan W. Black
AAAI2
2020 Should All Cross-Lingual Embeddings Speak English?
abstract
Most of recent work in cross-lingual word embeddings is severely Anglocentric.The vast majority of lexicon induction evaluation dictionaries are between English and another language, and the English embedding space is selected by default as the hub when learning in a multilingual setting.With this work, however, we challenge these practices.First, we show that the choice of hub language can significantly impact downstream lexicon induction and zero-shot POS tagging performance.Second, we both expand a standard Englishcentered evaluation dictionary collection to include all language pairs using triangulation, and create new dictionaries for under-represented languages.1 Evaluating established methods over all these language pairs sheds light into their suitability for aligning embeddings from distant languages and presents new challenges for the field.Finally, in our analysis we identify general guidelines for strong cross-lingual embedding baselines, that extend to language pairs that do not include English.
Antonios Anastasopoulos, Graham Neubig
ACL1
2020 It's Easier to Translate out of English than into it: Measuring Neural Translation Difficulty by Cross-Mutual Information
abstract
The performance of neural machine translation systems is commonly evaluated in terms of BLEU. However, due to its reliance on target language properties and generation, the BLEU metric does not allow an assessment of which translation directions are more difficult to model. In this paper, we propose cross-mutual information (XMI): an asymmetric information-theoretic metric of machine translation difficulty that exploits the probabilistic nature of most neural machine translation models. XMI allows us to better evaluate the difficulty of translating text into the target language while controlling for the difficulty of the target-side generation component independent of the translation task. We then present the first systematic and controlled study of cross-lingual translation difficulties using modern neural translation systems. Code for replicating our experiments is available online at https://github.com/e-bug/nmt-difficulty.
Emanuele Bugliarello, Sabrina J. Mielke, Antonios Anastasopoulos, Ryan Cotterell, Naoaki Okazaki
ACL3
2020 Predicting Performance for Natural Language Processing Tasks
abstract
Given the complexity of combinations of tasks, languages, and domains in natural language processing (NLP) research, it is computationally prohibitive to exhaustively test newly proposed models on each possible experimental setting.In this work, we attempt to explore the possibility of gaining plausible judgments of how well an NLP model can perform under an experimental setting, without actually training or testing the model.To do so, we build regression models to predict the evaluation score of an NLP experiment given the experimental settings as input.Experimenting on 9 different NLP tasks, we find that our predictors can produce meaningful predictions over unseen languages and different modeling architectures, outperforming reasonable baselines as well as human experts.Going further, we outline how our predictor can be used to find a small subset of representative experiments that should be run in order to obtain plausible predictions for all other experimental settings.1
Mengzhou Xia, Antonios Anastasopoulos, Ruochen Xu, Yiming Yang 0002, Graham Neubig
ACL2
2020 Automatic Interlinear Glossing for Under-Resourced Languages Leveraging Translations
abstract
Interlinear Glossed Text (IGT) is a widely used format for encoding linguistic information in language documentation projects and scholarly papers.Manual production of IGT takes time and requires linguistic expertise.We attempt to address this issue by creating automatic glossing models, using modern multi-source neural models that additionally leverage easy-to-collect translations.We further explore cross-lingual transfer and a simple output length control mechanism, further refining our models.Evaluated on three challenging low-resource scenarios, our approach significantly outperforms a recent, state-of-the-art baseline, particularly improving on overall accuracy as well as lemma and tag recall.
Xingyuan Zhao, Satoru Ozaki, Antonios Anastasopoulos, Graham Neubig, Lori S. Levin
COLING3
2020 Automatic Extraction of Rules Governing Morphological Agreement
abstract
Aditi Chaudhary, Antonios Anastasopoulos, Adithya Pratapa, David R. Mortensen, Zaid Sheikh, Yulia Tsvetkov, Graham Neubig. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Aditi Chaudhary, Antonios Anastasopoulos, Adithya Pratapa, David R. Mortensen, Zaid Sheikh, Yulia Tsvetkov, Graham Neubig
EMNLP (1)2
2020 Dynamic Data Selection and Weighting for Iterative Back-Translation
abstract
Back-translation has proven to be an effective method to utilize monolingual data in neural machine translation (NMT), and iteratively conducting back-translation can further improve the model performance.Selecting which monolingual data to back-translate is crucial, as we require that the resulting synthetic data are of high quality and reflect the target domain.To achieve these two goals, data selection and weighting strategies have been proposed, with a common practice being to select samples close to the target domain but also dissimilar to the average general-domain text.In this paper, we provide insights into this commonly used approach and generalize it to a dynamic curriculum learning strategy, which is applied to iterative back-translation models.In addition, we propose weighting strategies based on both the current quality of the sentence and its improvement over the previous iteration.We evaluate our models on domain adaptation, low-resource, and high-resource MT settings and on two language pairs.Experimental results demonstrate that our methods achieve improvements of up to 1.8 BLEU points over competitive baselines.1
Zi-Yi Dou, Antonios Anastasopoulos, Graham Neubig
EMNLP (1)2
2020 X-FACTR: Multilingual Factual Knowledge Retrieval from Pretrained Language Models
abstract
Language models (LMs) have proven surprisingly successful at capturing factual knowledge by completing cloze-style fill-in-theblank questions such as "Punta Cana is located in _."However, while knowledge is both written and queried in many languages, studies on LMs' factual representation ability have almost invariably been performed on English.To assess factual knowledge retrieval in LMs in different languages, we create a multilingual benchmark of cloze-style probes for 23 typologically diverse languages.To properly handle language variations, we expand probing methods from single-to multi-word entities, and develop several decoding algorithms to generate multi-token predictions.Extensive experimental results provide insights about how well (or poorly) current state-of-theart LMs perform at this task in languages with more or fewer available resources.We further propose a code-switching-based method to improve the ability of multilingual LMs to access knowledge, and verify its effectiveness on several benchmark languages.Benchmark data and code have be released at https: //x-factr.github.io.
Zhengbao Jiang, Antonios Anastasopoulos, Jun Araki, Haibo Ding, Graham Neubig
EMNLP (1)2
2020 OCR Post Correction for Endangered Language Texts
abstract
There is little to no data available to build natural language processing models for most endangered languages.However, textual data in these languages often exists in formats that are not machine-readable, such as paper books and scanned images.In this work, we address the task of extracting text from these resources.We create a benchmark dataset of transcriptions for scanned books in three critically endangered languages and present a systematic analysis of how general-purpose OCR tools are not robust to the data-scarce setting of endangered languages.We develop an OCR postcorrection method tailored to ease training in this data-scarce setting, reducing the recognition error rate by 34% on average across the three languages.1
Shruti Rijhwani, Antonios Anastasopoulos, Graham Neubig
EMNLP (1)2
2020 Universal Phone Recognition with a Multilingual Allophone System
abstract
Multilingual models can improve language processing, particularly for low resource situations, by sharing parameters across languages. Multilingual acoustic models, however, generally ignore the difference between phonemes (sounds that can support lexical contrasts in a particular language) and their corresponding phones (the sounds that are actually spoken, which are language independent). This can lead to performance degradation when combining a variety of training languages, as identically annotated phonemes can actually correspond to several different underlying phonetic realizations. In this work, we propose a joint model of both language-independent phone and language-dependent phoneme distributions. In multilingual ASR experiments over 11 languages, we find that this model improves testing performance by 2% phoneme error rate absolute in low-resource conditions. Additionally, because we are explicitly modeling language-independent phones, we can build a (nearly-)universal phone recognizer that, when combined with the PHOIBLE [1] large, manually curated database of phone inventories, can be customized into 2,000 language dependent recognizers. Experiments on two low-resourced indigenous languages, Inuktitut and Tusom, show that our recognizer achieves phone accuracy improvements of more than 17%, moving a step closer to speech recognition for all languages in the world.1
Siddharth Dalmia, Juncheng Li 0001, Matthew Lee 0012, Patrick Littell, Jiali Yao, Antonios Anastasopoulos, David R. Mortensen, Graham Neubig, Alan W. Black, Florian Metze
ICASSP7
2020 Optimizing Data Usage via Differentiable Rewards
abstract
To acquire a new skill, humans learn better and faster if a tutor, based on their current knowledge level, informs them of how much attention they should pay to particular content or practice problems. Similarly, a machine learning model could potentially be trained better with a scorer that “adapts” to its current learning state and estimates the importance of each training data instance. Training such an adaptive scorer efficiently is a challenging problem; in order to precisely quantify the effect of a data instance at a given time during the training, it is typically necessary to first complete the entire training process. To efficiently optimize data usage, we propose a reinforcement learning approach called Differentiable Data Selection (DDS). In DDS, we formulate a scorer network as a learnable function of the training data, which can be efficiently updated along with the main model being trained. Specifically, DDS updates the scorer with an intuitive reward signal: it should up-weigh the data that has a similar gradient with a dev set upon which we would finally like to perform well. Without significant computing overhead, DDS delivers strong and consistent improvements over several strong baselines on two very different tasks of machine translation and image classification.
Xinyi Wang 0001, Paul Michel, Antonios Anastasopoulos, Jaime G. Carbonell, Graham Neubig
ICML4
2020 A Resource for Studying Chatino Verbal Morphology
abstract
We present the first resource focusing on the verbal inflectional morphology of San Juan Quiahije Chatino, a tonal mesoamerican language spoken in Mexico. We provide a collection of complete inflection tables of 198 lemmata, with morphological tags based on the UniMorph schema. We also provide baseline results on three core NLP tasks: morphological analysis, lemmatization, and morphological inflection.
Hilaria Cruz, Antonios Anastasopoulos, Gregory Stump
LREC2
2020 A Resource for Computational Experiments on Mapudungun
abstract
We present a resource for computational experiments on Mapudungun, a polysynthetic indigenous language spoken in Chile with upwards of 200 thousand speakers. We provide 142 hours of culturally significant conversations in the domain of medical treatment. The conversations are fully transcribed and translated into Spanish. The transcriptions also include annotations for code-switching and non-standard pronunciations. We also provide baseline results on three core NLP tasks: speech recognition, speech synthesis, and machine translation between Spanish and Mapudungun. We further explore other applications for which the corpus will be suitable, including the study of code-switching, historical orthography change, linguistic structure, and sociological and anthropological studies.
Mingjun Duan, Carlos Fasola, Sai Krishna Rallabandi, Rodolfo Vega, Antonios Anastasopoulos, Lori S. Levin, Alan W. Black
LREC5
2020 AlloVera: A Multilingual Allophone Database
abstract
We introduce a new resource, AlloVera, which provides mappings from 218 allophones to phonemes for 14 languages. Phonemes are contrastive phonological units, and allophones are their various concrete realizations, which are predictable from phonological context. While phonemic representations are language specific, phonetic representations (stated in terms of (allo)phones) are much closer to a universal (language-independent) transcription. AlloVera allows the training of speech recognition models that output phonetic transcriptions in the International Phonetic Alphabet (IPA), regardless of the input language. We show that a “universal” allophone model, Allosaurus, built with AlloVera, outperforms “universal” phonemic models and language-specific models on a speech-transcription task. We explore the implications of this technology (and related technologies) for the documentation of endangered and minority languages. We further explore other applications for which AlloVera will be suitable as it grows, including phonological typology.
David R. Mortensen, Patrick Littell, Alexis Michaud, Shruti Rijhwani, Antonios Anastasopoulos, Alan W. Black, Florian Metze, Graham Neubig
LREC6
2019 Choosing Transfer Languages for Cross-Lingual Learning
abstract
Yu-Hsiang Lin, Chian-Yu Chen, Jean Lee, Zirui Li, Yuyan Zhang, Mengzhou Xia, Shruti Rijhwani, Junxian He, Zhisong Zhang, Xuezhe Ma, Antonios Anastasopoulos, Patrick Littell, Graham Neubig. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Yu-Hsiang Lin, Chian-Yu Chen, Jean Lee, Mengzhou Xia, Shruti Rijhwani, Junxian He, Zhisong Zhang, Xuezhe Ma, Antonios Anastasopoulos, Patrick Littell, Graham Neubig
ACL (1)11
2019 Generalized Data Augmentation for Low-Resource Translation
abstract
Translation to or from low-resource languages (LRLs) poses challenges for machine translation in terms of both adequacy and fluency.Data augmentation utilizing large amounts of monolingual data is regarded as an effective way to alleviate these problems.In this paper, we propose a general framework for data augmentation in low-resource machine translation that not only uses target-side monolingual data, but also pivots through a related highresource language (HRL).Specifically, we experiment with a two-step pivoting method to convert high-resource data to the LRL, making use of available resources to better approximate the true data distribution of the LRL.First, we inject LRL words into HRL sentences through an induced bilingual dictionary.Second, we further edit these modified sentences using a modified unsupervised machine translation framework.Extensive experiments on four low-resource datasets show that under extreme low-resource settings, our data augmentation techniques improve translation quality by up to 1.5 to 8 BLEU points compared to supervised back-translation baselines.1
Mengzhou Xia, Xiang Kong, Antonios Anastasopoulos, Graham Neubig
ACL (1)3
2019 Pushing the Limits of Low-Resource Morphological Inflection
abstract
Antonios Anastasopoulos, Graham Neubig. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Antonios Anastasopoulos, Graham Neubig
EMNLP/IJCNLP (1)1
2019 Unsupervised Domain Adaptation for Neural Machine Translation with Domain-Aware Feature Embeddings
abstract
Zi-Yi Dou, Junjie Hu, Antonios Anastasopoulos, Graham Neubig. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Zi-Yi Dou, Junjie Hu 0001, Antonios Anastasopoulos, Graham Neubig
EMNLP/IJCNLP (1)3
2019 Investigating Meta-Learning Algorithms for Low-Resource Natural Language Understanding Tasks
abstract
Zi-Yi Dou, Keyi Yu, Antonios Anastasopoulos. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Zi-Yi Dou, Keyi Yu, Antonios Anastasopoulos
EMNLP/IJCNLP (1)3
2018 Part-of-Speech Tagging on an Endangered Language: a Parallel Griko-Italian Resource
abstract
Most work on part-of-speech (POS) tagging is focused on high resource languages, or examines low-resource and active learning settings through simulated studies. We evaluate POS tagging techniques on an actual endangered language, Griko. We present a resource that contains 114 narratives in Griko, along with sentence-level translations in Italian, and provides gold annotations for the test set. Based on a previously collected small corpus, we investigate several traditional methods, as well as methods that take advantage of monolingual data or project cross-lingual POS tags. We show that the combination of a semi-supervised method with cross-lingual transfer is more appropriate for this extremely challenging setting, with the best tagger achieving an accuracy of 72.9%. With an applied active learning scheme, which we use to collect sentence-level annotations over the test set, we achieve improvements of more than 21 percentage points.
Antonios Anastasopoulos, Marika Lekakou, Josep Quer, Eleni Zimianiti, Justin DeBenedetto, David Chiang 0001
COLING1
2018 Leveraging Translations for Speech Transcription in Low-resource Settings
abstract
Recently proposed data collection frameworks for endangered language documentation aim not only to collect speech in the language of interest, but also to collect translations into a high-resource language that will render the collected resource interpretable. We focus on this scenario and explore whether we can improve transcription quality under these extremely low-resource settings with the assistance of text translations. We present a neural multi-source model and evaluate several variations of it on three low-resource datasets. We find that our multi-source model with shared attention outperforms the baselines, reducing transcription character error rate by up to 12.3%.
Antonios Anastasopoulos, David Chiang 0001
INTERSPEECH1
2018 Tied Multitask Learning for Neural Speech Translation
abstract
Antonios Anastasopoulos, David Chiang. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Antonios Anastasopoulos, David Chiang 0001
NAACL-HLT1
2016 An Unsupervised Probability Model for Speech-to-Translation Alignment of Low-Resource Languages
abstract
For many low-resource languages, spoken language resources are more likely to be annotated with translations than with transcriptions. Translated speech data is potentially valuable for documenting endangered languages or for training speech translation systems. A first step towards making use of such data would be to automatically align spoken words with their translations. We present a model that combines Dyer et al.'s reparameterization of IBM Model 2 (fast-align) and k-means clustering using Dynamic Time Warping as a distance metric. The two components are trained jointly using expectation-maximization. In an extremely low-resource scenario, our model performs significantly better than both a neural model and a strong baseline.
Antonios Anastasopoulos, David Chiang 0001, Long Duong
EMNLP1
2016 An Attentional Model for Speech Translation Without Transcription
abstract
Long Duong, Antonios Anastasopoulos, David Chiang, Steven Bird, Trevor Cohn. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Long Duong, Antonios Anastasopoulos, David Chiang 0001, Steven Bird, Trevor Cohn
HLT-NAACL2
2014 Adaptive Quality Estimation for Machine Translation
abstract
The automatic estimation of machine translation (MT) output quality is a hard task in which the selection of the appropriate algorithm and the most predictive features over reasonably sized training sets plays a crucial role.When moving from controlled lab evaluations to real-life scenarios the task becomes even harder.For current MT quality estimation (QE) systems, additional complexity comes from the difficulty to model user and domain changes.Indeed, the instability of the systems with respect to data coming from different distributions calls for adaptive solutions that react to new operating conditions.To tackle this issue we propose an online framework for adaptive QE that targets reactivity and robustness to user and domain changes.Contrastive experiments in different testing conditions involving user and domain changes demonstrate the effectiveness of our approach.
Marco Turchi, Antonios Anastasopoulos, José Guilherme Camargo de Souza, Matteo Negri
ACL (1)2