Kai North

dblp:293/7192 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0002-9970-2402ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Information extraction and text analysis · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational social science and digital humanities · 100%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › text mining › authorship analysis
native language identification
0.912025
Tracing L1 Interference in English Learner Writing: A Longitudinal Corpus with Error Annotations · EMNLP 2025
Computational social science and digital humanities
second language acquisition
0.912025
Tracing L1 Interference in English Learner Writing: A Longitudinal Corpus with Error Annotations · EMNLP 2025
Natural language and speech › Information extraction and text analysis › data annotation
LLM-based annotation
0.312025
Tracing L1 Interference in English Learner Writing: A Longitudinal Corpus with Error Annotations · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

large language model annotation · 1.7error annotation · 1.7
YearPublicationVenuePosition
2025 Tracing L1 Interference in English Learner Writing: A Longitudinal Corpus with Error Annotations
abstract
Language transfer is an important topic of research in second language acquisition and computational linguistics.The availability of suitable learner corpora is paramount for the study of second language acquisition (SLA) and language transfer.However, curating learner corpora is a challenging endeavor as high quality learner data is rarely publicly available.This results in only a few such corpora available to the community.To address this important gap, in this paper we present LENS, a novel English learner corpus with longitudinal data which enables researchers to investigate language learning over time.LENS contains 687 instances written by speakers of 15 different L1s.We use LENS two perform two important tasks at the intersection of SLA and Computational Linguistics: (1) Native Language Identification (NLI); and (2) an evaluation of large language models as a tool for high-precision, semi-automated annotation of L1 interference features.1
Poorvi Acharya, J. Elizabeth Liebl, Dhiman Goswami, Kai North, Marcos Zampieri, Antonios Anastasopoulos
EMNLP4
2025 Deep learning approaches to lexical simplification: A survey
abstract
Abstract Lexical Simplification (LS) is the task of substituting complex words within a sentence for simpler alternatives while maintaining the sentence’s original meaning. LS is the lexical component of Text Simplification (TS) systems with the aim of improving accessibility to various target populations such as individuals with low literacy or reading disabilities. Prior surveys have been published several years before the introduction of transformers, transformer-based large language models (LLMs), and prompt learning that have drastically changed the field of NLP. The high performance of these models has sparked renewed interest in LS. To reflect these recent advances, we present a comprehensive survey of papers published since 2017 on LS and its sub-tasks focusing on deep learning. Finally, we describe available benchmark datasets for the future development of LS systems.
Kai North, Tharindu Ranasinghe, Matthew Shardlow, Marcos Zampieri
J. Intell. Inf. Syst.1
2024 Language Variety Identification with True Labels
abstract
Language identification is an important first step in many NLP applications. Most publicly available language identification datasets, however, are compiled under the assumption that the gold label of each instance is determined by where texts are retrieved from. Research has shown that this is a problematic assumption, particularly in the case of very similar languages (e.g., Croatian and Serbian) and national language varieties (e.g., Brazilian and European Portuguese), where texts may contain no distinctive marker of the particular language or variety. To overcome this important limitation, this paper presents DSL True Labels (DSL-TL), the first human-annotated multilingual dataset for language variety identification. DSL-TL contains a total of 12,900 instances in Portuguese, split between European Portuguese and Brazilian Portuguese; Spanish, split between Argentine Spanish and Castilian Spanish; and English, split between American English and British English. We trained multiple models to discriminate between these language varieties, and we present the results in detail. The data and models presented in this paper provide a reliable benchmark toward the development of robust and fairer language variety identification systems. We make DSL-TL freely available to the research community.
Marcos Zampieri, Kai North, Tommi Jauhiainen, Mariano Felice, Neha Kumari 0010, Nishant Nair, Yash Bangera
LREC/COLING2
2024 Native Language Identification in Texts: A Survey
abstract
Dhiman Goswami, Sharanya Thilagan, Kai North, Shervin Malmasi, Marcos Zampieri. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Dhiman Goswami, Sharanya Thilagan, Kai North, Shervin Malmasi, Marcos Zampieri
NAACL-HLT3
2024 Health text simplification: An annotated corpus for digestive cancer education and novel strategies for reinforcement learning
Md. Mushfiqur Rahman, Mohammad Sabik Irbaz, Kai North, Michelle S. Williams, Marcos Zampieri, Kevin Lybarger
J. Biomed. Informatics3
2022 ALEXSIS-PT: A New Resource for Portuguese Lexical Simplification
abstract
Lexical simplification (LS) is the task of automatically replacing complex words for easier ones making texts more accessible to various target populations (e.g. individuals with low literacy, individuals with learning disabilities, second language learners). To train and test models, LS systems usually require corpora that feature complex words in context along with their potential substitutions. To continue improving the performance of LS systems we introduce ALEXSIS-PT, a novel multi-candidate dataset for Brazilian Portuguese LS containing 9,605 candidate substitutions for 387 complex words. ALEXSIS-PT has been compiled following the ALEXSIS-ES protocol for Spanish opening exciting new avenues for cross-lingual models. ALEXSIS-PT is the first LS multi-candidate dataset that contains Brazilian newspaper articles. We evaluated three models for substitute generation on this dataset, namely mBERT, XLM-R, and BERTimbau. The latter achieved the highest performance across all evaluation metrics.
Kai North, Marcos Zampieri, Tharindu Ranasinghe
COLING1