Sina Ahmadi

dblp:213/1031 · also Mohammad Sina Ahmadi · DBLP profile ↗
← Back
19ranked-venue papers
8as first author
15since 2021 · last 2026
0000-0001-7904-6551ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 8 first-author · 14 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization
abstract
Negar Foroutan, Clara Meister, Debjit Paul, Joel Niklaus, Sina Ahmadi, Antoine Bosselut, Rico Sennrich. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Negar Foroutan Eghlidi, Clara Meister, Debjit Paul, Joel Niklaus, Sina Ahmadi, Antoine Bosselut, Rico Sennrich
ACL (1)5
2026 CommonMorph: Participatory Morphological Documentation Platform
Aso Mahmudi, Sina Ahmadi, Kemal Kurniawan, Rico Sennrich, Eduard H. Hovy, Ekaterina Vylomova
LREC2
2026 Are Language Models Borrowing-Blind? A Multilingual Evaluation of Loanword Identification across 10 Languages
Mérilin Sousa Silva, Sina Ahmadi
LREC2
2025 ConLoan: A Contrastive Multilingual Dataset for Evaluating Loanwords
abstract
Sina Ahmadi, Micha David Hess, Elena Álvarez-Mellado, Alessia Battisti, Cui Ding, Anne Göhring, Yingqiang Gao, Zifan Jiang, Andrianos Michail, Peshmerge Morad, Joel Niklaus, Maria Christina Panagiotopoulou, Stefano Perrella, Juri Opitz, Anastassia Shaitarova, Rico Sennrich. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Sina Ahmadi, Micha David Hess, Elena Álvarez Mellado, Alessia Battisti, Cui Ding, Anne Göhring, Yingqiang Gao, Zifan Jiang, Andrianos Michail, Peshmerge Morad, Joel Niklaus, Maria Christina Panagiotopoulou, Stefano Perrella, Juri Opitz, Anastassia Shaitarova, Rico Sennrich
ACL (1)1
2025 PARME: Parallel Corpora for Low-Resourced Middle Eastern Languages
abstract
Sina Ahmadi, Rico Sennrich, Erfan Karami, Ako Marani, Parviz Fekrazad, Gholamreza Akbarzadeh Baghban, Hanah Hadi, Semko Heidari, Mahîr Dogan, Pedram Asadi, Dashne Bashir, Mohammad Amin Ghodrati, Kourosh Amini, Zeynab Ashourinezhad, Mana Baladi, Farshid Ezzati, Alireza Ghasemifar, Daryoush Hosseinpour, Behrooz Abbaszadeh, Amin Hassanpour, Bahaddin Jalal Hamaamin, Saya Kamal Hama, Ardeshir Mousavi, Sarko Nazir Hussein, Isar Nejadgholi, Mehmet Ölmez, Horam Osmanpour, Rashid Roshan Ramezani, Aryan Sediq Aziz, Ali Salehi, Mohammadreza Yadegari, Kewyar Yadegari, Sedighe Zamani Roodsari. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Sina Ahmadi, Rico Sennrich, Erfan Karami, Ako Marani, Parviz Fekrazad, Gholamreza Akbarzadeh Baghban, Hanah Hadi, Semko Heidari, Mahîr Dogan, Pedram Asadi, Dashne Bashir, Mohammad Amin Ghodrati, Kourosh Amini, Zeynab Ashourinezhad, Mana Baladi, Farshid Ezzati, Alireza Ghasemifar, Daryoush Hosseinpour, Behrooz Abbaszadeh, Amin Hassanpour, Bahaddin Jalal Hamaamin, Saya Kamal Hama, Ardeshir Mousavi, Sarko Nazir Hussein, Isar Nejadgholi, Mehmet Ölmez, Horam Osmanpour, Rashid Roshan Ramezani, Aryan Sediq Aziz, Ali Salehi, Mohammadreza Yadegari, Kewyar Yadegari, Sedighe Zamani Roodsari
ACL (1)1
2025 SwiLTra-Bench: The Swiss Legal Translation Benchmark
abstract
In Switzerland legal translation is uniquely important due to the country’s four official languages and requirements for multilingual legal documentation. However, this process traditionally relies on professionals who must be both legal experts and skilled translators—creating bottlenecks and impacting effective access to justice. To address this challenge, we introduce SwiLTra-Bench, a comprehensive multilingual benchmark of over 180K aligned Swiss legal translation pairs comprising laws, headnotes, and press releases across all Swiss languages along with English, designed to evaluate LLM-based translation systems. Our systematic evaluation reveals that frontier models achieve superior translation performance across all document types, while specialized translation systems excel specifically in laws but under-perform in headnotes. Through rigorous testing and human expert validation, we demonstrate that while fine-tuning open SLMs significantly improves their translation quality, they still lag behind the best zero-shot prompted frontier models such as Claude-3.5-Sonnet. Additionally, we present SwiLTra-Judge, a specialized LLM evaluation system that aligns best with human expert assessments.
Joel Niklaus, Jakob Merane, Luka Nenadic, Sina Ahmadi, Yingqiang Gao, Cyrill A. H. Chevalley, Claude Humbel, Christophe Gösken, Lorenzo Tanzi, Thomas Lüthi, Stefan Palombo, Spencer Poff, Boling Yang, Matthew Guillod, Robin Mamié, Daniel Brunner, Julio Pereyra, Niko Grupen
ACL (1)4
2025 Automatic Speech Recognition for Low-Resourced Middle Eastern Languages
Razhan Hameed, Sina Ahmadi, Hanah Hadi, Rico Sennrich
INTERSPEECH2
2025 Conversational Lexicography: Querying Lexicographic Data on Knowledge Graphs with SPARQL through Natural Language
abstract
Knowledge graphs offer an excellent solution for representing the lexical-semantic structures of lexicographic data. However, working with the SPARQL query language represents a considerable hurdle for many non-expert users who could benefit from the advantages of this technology. This paper addresses the challenge of creating natural language interfaces for lexicographic data retrieval on knowledge graphs such as Wikidata. We develop a multidimensional taxonomy capturing the complexity of Wikidata’s lexicographic data ontology module through four dimensions and create a template-based dataset with over 1.2 million mappings from natural language utterances to SPARQL queries. Our experiments with GPT-2 (124M), Phi-1.5 (1.3B), and GPT-3.5-Turbo reveal significant differences in model capabilities. While all models perform well on familiar patterns, only GPT-3.5-Turbo demonstrates meaningful generalization capabilities, suggesting that model size and diverse pre-training are crucial for adaptability in this domain. However, significant challenges remain in achieving robust generalization, handling diverse linguistic data, and developing scalable solutions that can accommodate the full complexity of lexicographic knowledge representation.
Kilian Sennrich, Sina Ahmadi
LDK2
2025 ELICA: Efficient and Load Balanced I/O Cache Architecture for Hyperconverged Infrastructures
abstract
Hyperconverged Infrastructures(HCIs) combine processing and storage elements to meet the requirements of data-intensive applications in performance, scalability, and quality of service. As an emerging paradigm, HCI should couple with a variety of traditional performance improvement approaches such as I/O caching in virtualized platforms. Contemporary I/O caching schemes are optimized for traditional single-node storage architectures and suffer from two major shortcomings for multi-node architectures: a) imbalanced cache space requirement and b) imbalanced I/O traffic and load. This makes existing schemes inefficient in distributing cache resources over an array of separate physical nodes. In this paper, we propose anEfficient andLoad BalancedI/OCacheArchitecture(ELICA), managing thesolid-state drive(SSD) cache resources across HCI nodes to enhance I/O performance. ELICA dynamically reconfigures and distributes the SSD cache resources throughout the array of HCI nodes and also balances the network traffic and I/O cache load by dynamic reallocation of cache resources. To maximize the performance, we further present an optimization problem defined byInteger Linear Programmingto efficiently distribute cache resources and balance the network traffic and I/O cache relocations. Our experimental results on a real platform show that ELICA improves quality of service in terms of average and worst-case latency in HCIs by 3.1× and 23%, respectively, compared to the state-of-the-art.
Mostafa Kishani, Sina Ahmadi, Saba Ahmadian, Reza Salkhordeh, Zdenek Becvar, Onur Mutlu, André Brinkmann, Hossein Asadi 0001
IEEE Trans. Parallel Distributed Syst.2
2024 Language and Speech Technology for Central Kurdish Varieties
abstract
Kurdish, an Indo-European language spoken by over 30 million speakers, is considered a dialect continuum and known for its diversity in language varieties. Previous studies addressing language and speech technology for Kurdish handle it in a monolithic way as a macro-language, resulting in disparities for dialects and varieties for which there are few resources and tools available. In this paper, we take a step towards developing resources for language and speech technology for varieties of Central Kurdish, creating a corpus by transcribing movies and TV series as an alternative to fieldwork. Additionally, we report the performance of machine translation, automatic speech recognition, and language identification as downstream tasks evaluated on Central Kurdish subdialects. Data and models are publicly available under an open license at https://github.com/sinaahmadi/CORDI.
Sina Ahmadi, Daban Q. Jaff, Md Mahfuz Ibn Alam, Antonios Anastasopoulos
LREC/COLING1
2023 Script Normalization for Unconventional Writing of Under-Resourced Languages in Bilingual Communities
abstract
The wide accessibility of social media has provided linguistically under-represented communities with an extraordinary opportunity to create content in their native languages.This, however, comes with certain challenges in script normalization, particularly where the speakers of a language in a bilingual community rely on another script or orthography to write their native language.This paper addresses the problem of script normalization for several such languages that are mainly written in a Perso-Arabic script.Using synthetic data with various levels of noise and a transformerbased model, we demonstrate that the problem can be effectively remediated.We conduct a small-scale evaluation of real data as well.Our experiments indicate that script normalization is also beneficial to improve the performance of downstream tasks such as machine translation and language identification. 1 Language Unconventional script Unconventional writing Conventional writing Gilaki Persian ‫زﻧﻦ‬ ‫ﮔﺐ‬ ‫ﺟﯽ‬ ‫اون‬ ‫ﮔﻴﻠﮑﻦ‬ ‫ﮔﻪ‬ ‫ﻫﻴﺴﻪ‬ ‫ﻧﻢ‬ ‫زون‬ ‫ﯾﺘﻪ‬ ‫زﻧﻦ‬ ‫ﮔﺐ‬ ‫ﺟﻲ‬ ٚ ‫اۊن‬ ‫ﮔﻴﻠﮑﺆن‬ ‫ﮔﻪ‬ ‫ﻫﻴﺴﻪ‬ ‫ﻧﺆم‬ ٚ ‫زوؤن‬ ‫ﯾﺘﻪ‬ Kashmiri Urdu ‫#"۔‬ $% & ' ( % ) * + ,"-.% / , .0 % 1 "-2( % 3 ‫َر۔‬ ‫جانو‬ ‫ُرٲس4‬ ‫و‬ ‫َکھ‬ ‫ا‬ ُ ‫چھ‬ ‫ٛور‬ ‫بر‬ Kurmanji Arabic ‫دا‬ ‫دهوك‬ ‫پارزكار‬ ‫بةرثوا‬ ‫اامدي‬ ‫قايمقام‬ ‫دا‬ ‫دهۆکێ‬ ‫پارێزگارێ‬ ‫بەرسڤا‬ ‫ئامێدیێ‬ ‫قایمقامێ‬ Sorani Arabic ‫دةويت‬ ‫فهديان‬ ‫ديارة‬ ‫شانؤوة‬ ‫يةكةم‬ ‫لة‬ ‫هةر‬ ‫دەوێت‬ ‫فەهەدیان‬ ‫دیارە‬ ‫شانۆوە‬ ‫یەکەم‬ ‫لە‬ ‫هەر‬ Sindhi Urdu 5 6ٔ 8 % 9 $ :% ; < =% > '% ? 5 @",-A B % C D 5 6 E F G H $% I G JG % K -G LA( M % N < = O
Sina Ahmadi, Antonios Anastasopoulos
ACL (1)1
2022 Cross-Lingual Link Discovery for Under-Resourced Languages
abstract
In this paper, we provide an overview of current technologies for cross-lingual link discovery, and we discuss challenges, experiences and prospects of their application to under-resourced languages. We rst introduce the goals of cross-lingual linking and associated technologies, and in particular, the role that the Linked Data paradigm (Bizer et al., 2011) applied to language data can play in this context. We de ne under-resourced languages with a speci c focus on languages actively used on the internet, i.e., languages with a digitally versatile speaker community, but limited support in terms of language technology. We argue that languages for which considerable amounts of textual data and (at least) a bilingual word list are available, techniques for cross-lingual linking can be readily applied, and that these enable the implementation of downstream applications for under-resourced languages via the localisation and adaptation of existing technologies and resources.
Mike Rosner, Sina Ahmadi, Elena Apostol, Julia Bosque-Gil, Christian Chiarcos, Milan Dojchinovski, Katerina Gkirtzou, Jorge Gracia, Dagmar Gromann, Chaya Liebeskind, Giedre Valunaite Oleskeviciene, Gilles Sérasset, Ciprian-Octavian Truica
LREC2
2022 CoFiF Plus: A French Financial Narrative Summarisation Corpus
abstract
Natural Language Processing is increasingly being applied in the finance and business industry to analyse the text of many different types of financial documents. Given the increasing growth of firms around the world, the volume of financial disclosures and financial texts in different languages and forms is increasing sharply and therefore the study of language technology methods that automatically summarise content has grown rapidly into a major research area. Corpora for financial narrative summarisation exists in English, but there is a significant lack of financial text resources in the French language. To remedy this, we present CoFiF Plus, the first French financial narrative summarisation dataset providing a comprehensive set of financial text written in French. The dataset has been extracted from French financial reports published in PDF file format. It is composed of 1,703 reports from the most capitalised companies in France (Euronext Paris) covering a time frame from 1995 to 2021. This paper describes the collection, annotation and validation of the financial reports and their summaries. It also describes the dataset and gives the results of some baseline summarisers. Our datasets will be openly available upon the acceptance of the paper.
Nadhem Zmandar, Tobias Daudert, Sina Ahmadi, Mahmoud El-Haj, Paul Rayson
LREC3
2022 Leveraging Multilingual News Websites for Building a Kurdish Parallel Corpus
abstract
Machine translation has been a major motivation of development in natural language processing. Despite the burgeoning achievements in creating more efficient machine translation systems, thanks to deep learning methods, parallel corpora have remained indispensable for progress in the field. In an attempt to create parallel corpora for the Kurdish language, in this article, we describe our approach in retrieving potentially alignable news articles from multi-language websites and manually align them across dialects and languages based on lexical similarity and transliteration of scripts. We present a corpus containing 12,327 translation pairs in the two major dialects of Kurdish, Sorani and Kurmanji. We also provide 1,797 and 650 translation pairs in English-Kurmanji and English-Sorani. The corpus is publicly available under the CC BY-NC-SA 4.0 license. 1
Sina Ahmadi, Hossein Hassani 0001, Daban Q. Jaff
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2021 Monolingual Word Sense Alignment as a Classification Problem
abstract
Words are defined based on their meanings in various ways in different resources.Aligning word senses across monolingual lexicographic resources increases domain coverage and enables integration and incorporation of data.In this paper, we explore the application of classification methods using manually-extracted features along with representation learning techniques in the task of word sense alignment and semantic relationship detection.We demonstrate that the performance of classification methods dramatically varies based on the type of semantic relationships due to the nature of the task but outperforms the previous experiments.
Sina Ahmadi, John P. McCrae
GWC1
2020 A Multilingual Evaluation Dataset for Monolingual Word Sense Alignment
abstract
Aligning senses across resources and languages is a challenging task with beneficial applications in the field of natural language processing and electronic lexicography. In this paper, we describe our efforts in manually aligning monolingual dictionaries. The alignment is carried out at sense-level for various resources in 15 languages. Moreover, senses are annotated with possible semantic relationships such as broadness, narrowness, relatedness, and equivalence. In comparison to previous datasets for this task, this dataset covers a wide range of languages and resources and focuses on the more challenging task of linking general-purpose language. We believe that our data will pave the way for further advances in alignment and evaluation of word senses by creating new solutions, particularly those notoriously requiring data such as neural networks. Our resources are publicly available at https://github.com/elexis-eu/MWSA.
Sina Ahmadi, John P. McCrae, Sanni Nimb, Anas Fahad Khan, Monica Monachini, Bolette S. Pedersen, Thierry Declerck, Tanja Wissik, Andrea Bellandi, Irene Pisani, Thomas Troelsgård, Sussi Olsen, Simon Krek, Veronika Lipp, Tamás Váradi, László Simon, András Gyorffy, Carole Tiberius, Tanneke Schoonheim, Yifat Ben Moshe, Maya Rudich, Raya Abu Ahmad, Dorielle Lonke, Kira Kovalenko, Margit Langemets, Jelena Kallas, Oksana Dereza, Theodorus Fransen, David Cillessen, David Lindemann, Mikel Alonso, Ana Salgado, José-Luis Sancho-Gómez, Rafael-J. Ureña-Ruiz, Jordi Porta-Zamorano, Kiril Ivanov Simov, Petya Osenova, Zara Kancheva, Ivaylo Radev, Ranka Stankovic, Andrej Perdih, Dejan Gabrovsek
LREC1
2020 Defying Wikidata: Validation of Terminological Relations in the Web of Data
abstract
In this paper we present an approach to validate terminological data retrieved from open encyclopaedic knowledge bases. This need arises from the enrichment of automatically extracted terms with information from existing resources in theLinguistic Linked Open Data cloud. Specifically, the resource employed for this enrichment is WIKIDATA, since it is one of the biggest knowledge bases freely available within the Semantic Web. During the experiment, we noticed that certain RDF properties in the Knowledge Base did not contain the data they are intended to represent, but a different type of information. In this paper we propose an approach to validate the retrieved data based on four axioms that rely on two linguistic theories: the x-bar theory and the multidimensional theory of terminology. The validation process is supported by a second knowledge base specialised in linguistic data; in this case, CONCEPTNET. In our experiment, we validate terms from the legal domain in four languages: Dutch, English, German and Spanish. The final aim is to generate a set of sound and reliable terminological resources in RDF to contribute to the population of the Linguistic Linked Open Data cloud.
Patricia Martín-Chozas, Sina Ahmadi, Elena Montiel-Ponsoda
LREC2
2019 A Rule-Based Kurdish Text Transliteration System
abstract
In this article, we present a rule-based approach for transliterating two of the most used orthographies in Sorani Kurdish. Our work consists of detecting a character in a word by removing the possible ambiguities and mapping it into the target orthography. We describe different challenges in Kurdish text mining and propose novel ideas concerning the transliteration task for Sorani Kurdish. Our transliteration system, named Wergor , achieves 82.79% overall precision and more than 99% in detecting the double-usage characters. We also present a manually transliterated corpus for Kurdish.
Sina Ahmadi
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2014 Towards Building KurdNet, the Kurdish WordNet
abstract
In this paper we highlight the main challenges in building a lexical database for Kurdish, a resource-scarce and diverse language.We also report on our effort in building the first prototype of KurdNetthe Kurdish WordNet-along with a preliminary evaluation of its impact on Kurdish information retrieval.
Purya Aliabadi, Sina Ahmadi, Shahin Salavati, Kyumars Sheykh Esmaili
GWC2