EDBT 2026 Demo / reviewers in the wild / expert
Kai North
dblp:293/7192
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0002-9970-2402ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Information extraction and text analysis · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational social science and digital humanities · 100% |
Topics — the 3 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › text mining › authorship analysis
native language identification |
0.9 | 1 | 2025 | Tracing L1 Interference in English Learner Writing: A Longitudinal Corpus with Error Annotations · EMNLP 2025 |
Computational social science and digital humanities
second language acquisition |
0.9 | 1 | 2025 | Tracing L1 Interference in English Learner Writing: A Longitudinal Corpus with Error Annotations · EMNLP 2025 |
Natural language and speech › Information extraction and text analysis › data annotation
LLM-based annotation |
0.3 | 1 | 2025 | Tracing L1 Interference in English Learner Writing: A Longitudinal Corpus with Error Annotations · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
large language model annotation · 1.7error annotation · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Tracing L1 Interference in English Learner Writing: A Longitudinal Corpus with Error AnnotationsabstractLanguage transfer is an important topic of research in second language acquisition and computational linguistics.The availability of suitable learner corpora is paramount for the study of second language acquisition (SLA) and language transfer.However, curating learner corpora is a challenging endeavor as high quality learner data is rarely publicly available.This results in only a few such corpora available to the community.To address this important gap, in this paper we present LENS, a novel English learner corpus with longitudinal data which enables researchers to investigate language learning over time.LENS contains 687 instances written by speakers of 15 different L1s.We use LENS two perform two important tasks at the intersection of SLA and Computational Linguistics: (1) Native Language Identification (NLI); and (2) an evaluation of large language models as a tool for high-precision, semi-automated annotation of L1 interference features.1 Poorvi Acharya, J. Elizabeth Liebl, Dhiman Goswami, Kai North, Marcos Zampieri, Antonios Anastasopoulos |
EMNLP | 4 |
| 2025 | Deep learning approaches to lexical simplification: A surveyabstractAbstract Lexical Simplification (LS) is the task of substituting complex words within a sentence for simpler alternatives while maintaining the sentence’s original meaning. LS is the lexical component of Text Simplification (TS) systems with the aim of improving accessibility to various target populations such as individuals with low literacy or reading disabilities. Prior surveys have been published several years before the introduction of transformers, transformer-based large language models (LLMs), and prompt learning that have drastically changed the field of NLP. The high performance of these models has sparked renewed interest in LS. To reflect these recent advances, we present a comprehensive survey of papers published since 2017 on LS and its sub-tasks focusing on deep learning. Finally, we describe available benchmark datasets for the future development of LS systems. Kai North, Tharindu Ranasinghe, Matthew Shardlow, Marcos Zampieri |
J. Intell. Inf. Syst. | 1 |
| 2024 | Language Variety Identification with True LabelsabstractLanguage identification is an important first step in many NLP applications. Most publicly available language identification datasets, however, are compiled under the assumption that the gold label of each instance is determined by where texts are retrieved from. Research has shown that this is a problematic assumption, particularly in the case of very similar languages (e.g., Croatian and Serbian) and national language varieties (e.g., Brazilian and European Portuguese), where texts may contain no distinctive marker of the particular language or variety. To overcome this important limitation, this paper presents DSL True Labels (DSL-TL), the first human-annotated multilingual dataset for language variety identification. DSL-TL contains a total of 12,900 instances in Portuguese, split between European Portuguese and Brazilian Portuguese; Spanish, split between Argentine Spanish and Castilian Spanish; and English, split between American English and British English. We trained multiple models to discriminate between these language varieties, and we present the results in detail. The data and models presented in this paper provide a reliable benchmark toward the development of robust and fairer language variety identification systems. We make DSL-TL freely available to the research community. Marcos Zampieri, Kai North, Tommi Jauhiainen, Mariano Felice, Neha Kumari 0010, Nishant Nair, Yash Bangera |
LREC/COLING | 2 |
| 2024 | Native Language Identification in Texts: A SurveyabstractDhiman Goswami, Sharanya Thilagan, Kai North, Shervin Malmasi, Marcos Zampieri. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Dhiman Goswami, Sharanya Thilagan, Kai North, Shervin Malmasi, Marcos Zampieri |
NAACL-HLT | 3 |
| 2024 | Health text simplification: An annotated corpus for digestive cancer education and novel strategies for reinforcement learning
Md. Mushfiqur Rahman, Mohammad Sabik Irbaz, Kai North, Michelle S. Williams, Marcos Zampieri, Kevin Lybarger |
J. Biomed. Informatics | 3 |
| 2022 | ALEXSIS-PT: A New Resource for Portuguese Lexical SimplificationabstractLexical simplification (LS) is the task of automatically replacing complex words for easier ones making texts more accessible to various target populations (e.g. individuals with low literacy, individuals with learning disabilities, second language learners). To train and test models, LS systems usually require corpora that feature complex words in context along with their potential substitutions. To continue improving the performance of LS systems we introduce ALEXSIS-PT, a novel multi-candidate dataset for Brazilian Portuguese LS containing 9,605 candidate substitutions for 387 complex words. ALEXSIS-PT has been compiled following the ALEXSIS-ES protocol for Spanish opening exciting new avenues for cross-lingual models. ALEXSIS-PT is the first LS multi-candidate dataset that contains Brazilian newspaper articles. We evaluated three models for substitute generation on this dataset, namely mBERT, XLM-R, and BERTimbau. The latter achieved the highest performance across all evaluation metrics. Kai North, Marcos Zampieri, Tharindu Ranasinghe |
COLING | 1 |