Constantine Lignos

dblp:78/4390 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0001-6410-2848ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 3 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2026 Comparing Approaches to Automatic Summarization in Less-Resourced Languages
Chester Palen-Michel, Constantine Lignos
LREC2
2025 OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ Languages
abstract
We present OpenNER 1.0, a standardized collection of openly-available named entity recognition (NER) datasets.OpenNER contains 36 NER corpora that span 52 languages, humanannotated in varying named entity ontologies.We correct annotation format issues, standardize the original datasets into a uniform representation with consistent entity type names across corpora, and provide the collection in a structure that enables research in multilingual and multi-ontology NER.We provide baseline results using three pretrained multilingual language models and two large language models to compare the performance of recent models and facilitate future research in NER.We find that no single model is best in all languages and that significant work remains to obtain high performance from LLMs on the NER task.OpenNER is released at https://github.com/bltlab/open-ner.
Chester Palen-Michel, Maxwell Pickering, Maya Kruse, Jonne Sälevä, Constantine Lignos
EMNLP5
2024 QueryNER: Segmentation of E-commerce Queries
abstract
We present QueryNER, a manually-annotated dataset and accompanying model for e-commerce query segmentation. Prior work in sequence labeling for e-commerce has largely addressed aspect-value extraction which focuses on extracting portions of a product title or query for narrowly defined aspects. Our work instead focuses on the goal of dividing a query into meaningful chunks with broadly applicable types. We report baseline tagging results and conduct experiments comparing token and entity dropping for null and low recall query recovery. Challenging test sets are created using automatic transformations and show how simple data augmentation techniques can make the models more robust to noise. We make the QueryNER dataset publicly available.
Chester Palen-Michel, Lizzie Liang, Constantine Lignos
LREC/COLING4
2024 CoNLL#: Fine-grained Error Analysis and a Corrected Test Set for CoNLL-03 English
abstract
Modern named entity recognition systems have steadily improved performance in the age of larger and more powerful neural models. However, over the past several years, the state-of-the-art has seemingly hit another plateau on the benchmark CoNLL-03 English dataset. In this paper, we perform a deep dive into the test outputs of the highest-performing NER models, conducting a fine-grained evaluation of their performance by introducing new document-level annotations on the test set. We go beyond F1 scores by categorizing errors in order to interpret the true state of the art for NER and guide future work. We review previous attempts at correcting the various flaws of the test set and introduce CoNLL#, a new corrected version of the test set that addresses its systematic and most prevalent errors, allowing for low-noise, interpretable error analysis.
Andrew Rueda, Elena Álvarez Mellado, Constantine Lignos
LREC/COLING3
2024 ParaNames 1.0: Creating an Entity Name Corpus for 400+ Languages Using Wikidata
abstract
We introduce ParaNames, a massively multilingual parallel name resource consisting of 140 million names spanning over 400 languages. Names are provided for 16.8 million entities, and each entity is mapped from a complex type hierarchy to a standard type (PER/LOC/ORG). Using Wikidata as a source, we create the largest resource of this type to date. We describe our approach to filtering and standardizing the data to provide the best quality possible. ParaNames is useful for multilingual language processing, both in defining tasks for name translation/transliteration and as supplementary data for tasks such as named entity recognition and linking. We demonstrate the usefulness of ParaNames on two tasks. First, we perform canonical name translation between English and 17 other languages. Second, we use it as a gazetteer for multilingual named entity recognition, obtaining performance improvements on all 10 languages evaluated.
Jonne Sälevä, Constantine Lignos
LREC/COLING2
2022 Detecting Unassimilated Borrowings in Spanish: An Annotated Corpus and Approaches to Modeling
abstract
This work presents a new resource for borrowing identification and analyzes the performance and errors of several models on this task.We introduce a new annotated corpus of Spanish newswire rich in unassimilated lexical borrowings-words from one language that are introduced into another without orthographic adaptation-and use it to evaluate how several sequence labeling models (CRF, BiLSTM-CRF, and Transformer-based models) perform.The corpus contains 370,000 tokens and is larger, more borrowing-dense, OOV-rich, and topic-varied than previous corpora available for this task.Our results show that a BiLSTM-CRF model fed with subword embeddings along with either Transformerbased embeddings pretrained on codeswitched data or a combination of contextualized word embeddings outperforms results obtained by a multilingual BERT-based model.
Elena Álvarez Mellado, Constantine Lignos
ACL (1)2
2022 MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity Recognition
abstract
David Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani, Michael Beukman, Chester Palen-Michel, Constantine Lignos, Jesujoba Alabi, Shamsuddeen Muhammad, Peter Nabende, Cheikh M. Bamba Dione, Andiswa Bukula, Rooweither Mabuya, Bonaventure F. P. Dossou, Blessing Sibanda, Happy Buzaaba, Jonathan Mukiibi, Godson Kalipe, Derguene Mbaye, Amelia Taylor, Fatoumata Kabore, Chris Chinenye Emezue, Anuoluwapo Aremu, Perez Ogayo, Catherine Gitau, Edwin Munkoh-Buabeng, Victoire Memdjokam Koagne, Allahsera Auguste Tapo, Tebogo Macucwa, Vukosi Marivate, Mboning Tchiaze Elvis, Tajuddeen Gwadabe, Tosin Adewumi, Orevaoghene Ahia, Joyce Nakatumba-Nabende, Neo Lerato Mokono, Ignatius Ezeani, Chiamaka Chukwuneke, Mofetoluwa Oluwaseun Adeyemi, Gilles Quentin Hacheme, Idris Abdulmumin, Odunayo Ogundepo, Oreen Yousuf, Tatiana Moteu, Dietrich Klakow. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
David Ifeoluwa Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani, Michael Beukman, Chester Palen-Michel, Constantine Lignos, Jesujoba O. Alabi, Shamsuddeen Hassan Muhammad, Peter Nabende, Cheikh M. Bamba Dione, Andiswa Bukula, Rooweither Mabuya, Bonaventure F. P. Dossou, Blessing K. Sibanda, Happy Buzaaba, Jonathan Mukiibi, Godson Kalipe, Derguene Mbaye, Amelia V. Taylor, Fatoumata Ouoba Kabore, Chris C. Emezue, Aremu Anuoluwapo, Perez Ogayo, Catherine Gitau, Edwin Munkoh-Buabeng, Victoire Memdjokam Koagne, Allahsera Tapo, Tebogo Macucwa, Vukosi Marivate, Elvis Mboning, Tajuddeen Rabiu Gwadabe, Tosin P. Adewumi, Orevaoghene Ahia, Joyce Nakatumba-Nabende, Neo L. Mokono, Ignatius Ezeani, Chiamaka Ijeoma Chukwuneke, Mofe Adeyemi, Gilles Hacheme, Idris Abdulmumin, Odunayo Ogundepo, Oreen Yousuf, Tatiana Moteu Ngoli, Dietrich Klakow
EMNLP7
2022 Borrowing or Codeswitching? Annotating for Finer-Grained Distinctions in Language Mixing
abstract
We present a new corpus of Twitter data annotated for codeswitching and borrowing between Spanish and English. The corpus contains 9,500 tweets annotated at the token level with codeswitches, borrowings, and named entities. This corpus differs from prior corpora of codeswitching in that we attempt to clearly define and annotate the boundary between codeswitching and borrowing and do not treat common “internet-speak” (lol, etc.) as codeswitching when used in an otherwise monolingual context. The result is a corpus that enables the study and modeling of Spanish-English borrowing and codeswitching on Twitter in one dataset. We present baseline scores for modeling the labels of this corpus using Transformer-based language models. The annotation itself is released with a CC BY 4.0 license, while the text it applies to is distributed in compliance with the Twitter terms of service.
Elena Álvarez Mellado, Constantine Lignos
LREC2
2022 Multilingual Open Text Release 1: Public Domain News in 44 Languages
abstract
We present a Multilingual Open Text (MOT), a new multilingual corpus containing text in 44 languages, many of which have limited existing text resources for natural language processing. The first release of the corpus contains over 2.8 million news articles and an additional 1 million short snippets (photo captions, video descriptions, etc.) published between 2001–2022 and collected from Voice of America’s news websites. We describe our process for collecting, filtering, and processing the data. The source material is in the public domain, our collection is licensed using a creative commons license (CC BY 4.0), and all software used to create the corpus is released under the MIT License. The corpus will be regularly updated as additional documents are published.
Chester Palen-Michel, June Kim, Constantine Lignos
LREC3
2021 Macro-Average: Rare Types Are Important Too
abstract
Thamme Gowda, Weiqiu You, Constantine Lignos, Jonathan May. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Thamme Gowda, Weiqiu You, Constantine Lignos, Jonathan May
NAACL-HLT3
2021 MasakhaNER: Named Entity Recognition for African Languages
abstract
Abstract We take a step towards addressing the under- representation of the African continent in NLP research by bringing together different stakeholders to create the first large, publicly available, high-quality dataset for named entity recognition (NER) in ten African languages. We detail the characteristics of these languages to help researchers and practitioners better understand the challenges they pose for NER tasks. We analyze our datasets and conduct an extensive empirical evaluation of state- of-the-art methods across both supervised and transfer learning settings. Finally, we release the data, code, and models to inspire future research on African NLP.1
David Ifeoluwa Adelani, Jade Z. Abbott, Graham Neubig, Daniel D'souza, Julia Kreutzer, Constantine Lignos, Chester Palen-Michel, Happy Buzaaba, Shruti Rijhwani, Sebastian Ruder, Stephen Mayhew 0002, Israel Abebe Azime, Shamsuddeen Hassan Muhammad, Chris C. Emezue, Joyce Nakatumba-Nabende, Perez Ogayo, Aremu Anuoluwapo, Catherine Gitau, Derguene Mbaye, Jesujoba O. Alabi, Seid Muhie Yimam, Tajuddeen Rabiu Gwadabe, Ignatius Ezeani, Rubungo Andre Niyongabo, Jonathan Mukiibi, Verrah Otiende, Iroro Orife, Davis David, Samba Ngom, Tosin P. Adewumi, Paul Rayson, Mofe Adeyemi, Gerald Muriuki, Emmanuel Anebi, Chiamaka Ijeoma Chukwuneke, Nkiruka Odu, Eric Peter Wairagala, Samuel Oyerinde, Clemencia Siro, Tobius Saul Bateesa, Temilola Oloyede, Yvonne Wambui, Victor Akinode, Deborah Nabagereka, Maurice Katusiime, Ayodele Awokoya, Mouhamadane Mboup, Dibora Gebreyohannes, Henok Tilaye, Kelechi Nwaike, Degaga Wolde, Abdoulaye Faye, Blessing K. Sibanda, Orevaoghene Ahia, Bonaventure F. P. Dossou, Kelechi Ogueji, Thierno Ibrahima Diop, Abdoulaye Diallo, Adewale Akinfaderin, Tendai Marengereke, Salomey Osei
Trans. Assoc. Comput. Linguistics6
2019 The Challenges of Optimizing Machine Translation for Low Resource Cross-Language Information Retrieval
abstract
Constantine Lignos, Daniel Cohen, Yen-Chieh Lien, Pratik Mehta, W. Bruce Croft, Scott Miller. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Constantine Lignos, Yen-Chieh Lien, Pratik Mehta, W. Bruce Croft
EMNLP/IJCNLP (1)1
2018 Combining rule-based and statistical mechanisms for low-resource named entity recognition
Ryan Gabbard, Jay DeYoung, Constantine Lignos, Marjorie Freedman, Ralph M. Weischedel
Mach. Transl.3
2012 Situation understanding bot through language and environment
abstract
This video shows a demonstration of a fully autonomous robot, an iRobot ATRV-JR, which can be given commands using natural language. Users type commands to the robot on a tablet computer, which are then parsed and processed using semantic analysis. This information is used to build a plan representing the high level autonomous behaviors the robot should perform [2][1]. The robot can be given commands to be executed immediately (e.g., "Search the floor for hostages.") as well as standing orders for use over the entire run (e.g., "Let me know if you see any bombs.").
Daniel J. Brooks, Constantine Lignos, Mikhail S. Medvedev, Ian Perera, Cameron Finucane, Vasumathi Raman, Abraham Shultz, Sean McSheehy, Adam Norton, Hadas Kress-Gazit, Mitchell P. Marcus, Holly A. Yanco
HRI2
2011 Modeling Infant Word Segmentation
Constantine Lignos
CoNLL1
2010 Recession Segmentation: Simpler Online Word Segmentation Using Limited Resources
Constantine Lignos, Charles Yang 0001
CoNLL1
2006 Effects of head movement on perceptions of humanoid robot behavior
abstract
This paper examines human perceptions of humanoid robot behavior, specifically how perception is affected by variations in head tracking behavior under constant gestural behavior. Subjects were invited to the lab to "play with Nico," an upper-torso humanoid robot. The follow-up survey asked subjects to rate and write about the experience. A coding scheme originally created to gauge human intentionality was applied to written responses to measure the level of intentionality that subjects perceived in the robot. Subjects were presented with one of four variations of head movement: a motionless head, a smooth tracking head, a tracking head without smoothed movements, and an avoidance behavior, while a pre-scripted wave and beckon sequence was carried out in all cases. Surprisingly, subjects rated the interaction as most enjoyable and Nico as possessing more intentionality when avoidance and unsmooth tracking were used. These data suggest that naïve users of robots may prefer caricatured and exaggerated behaviors to more natural ones. Also, correlations between ratings across modes suggest that simple features of robot behavior reliably evoke notable changes in many perception scales.
Qian (Emily) Wang, Constantine Lignos, Ashish Vatsal, Brian Scassellati
HRI2