Erik F. Tjong Kim Sang

dblp:07/2399 · also Erik Tjong Kim Sang · DBLP profile ↗
← Back
25ranked-venue papers
12as first author
4since 2021 · last 2025
0000-0002-8431-081XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 9 first-authorSoftware engineering, systems software and programming languages · 6 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2025 OnToxKG: An Ontology-Based Knowledge Graph of Toxic Symbols and Their Manifestations
Delfina Sol Martinez Pandiani, Erik F. Tjong Kim Sang, Davide Ceolin
ICWE2
2025 Evaluating Locally Run Large Language Models on Toxic Meme Analysis
Erik F. Tjong Kim Sang, Delfina Sol Martinez Pandiani, Davide Ceolin
ICWE1
2024 Investigating the Usefulness of Product Reviews Through Bipolar Argumentation Frameworks
Atefeh Keshavarzi Zafarghandi, Laura Hollink, Erik F. Tjong Kim Sang, Davide Ceolin
ICWE4
2021 Extracting Stances on Pandemic Measures from Social Media Data
abstract
Support for national measures against the COVID-19 pandemic can be measured by collecting responses to questionnaires. In this paper we explore a less costly and less time-consuming method: by analyzing social media data. We compare stances regarding anti-pandemic measures extracted from tweets with the questionnaire results provided by the Dutch national health institute RIVM. We find similarities and differences and discuss the results.
Erik F. Tjong Kim Sang, Shihan Wang 0001, Marijn Schraagen, Mehdi Dastani
e-Science1
2018 Utilizing a Transparency-Driven Environment Toward Trusted Automatic Genre Classification: A Case Study in Journalism History
abstract
With the growing abundance of unlabeled data in real-world tasks, researchers have to rely on the predictions given by black-boxed computational models. However, it is an often neglected fact that these models may be scoring high on accuracy for the wrong reasons. In this paper, we present a practical impact analysis of enabling model transparency by various presentation forms. For this purpose, we developed an environment that empowers non-computer scientists to become practicing data scientists in their own research field. We demonstrate the gradually increasing understanding of journalism historians through a real-world use case study on automatic genre classification of newspaper articles. This study is a first step towards trusted usage of machine learning pipelines in a responsible way.
Aysenur Bilgin, Erik F. Tjong Kim Sang, Kim Smeenk, Laura Hollink, Jacco van Ossenbruggen, Frank Harbers, Marcel Broersma
eScience2
2017 Determining the Function of Political Tweets
abstract
We study the discursive practices of politicians and journalists on social media. For this we need more annotated data than we currently have but the annotation process is time-consuming and costly. In this paper we examine machine learning methods for automatically annotating unseen tweetsbased on a small set of manually annotated tweets. Forimproving the performance of the learner, we focus onmethods related to training data expansion, like artificialtraining data, active learning and incorporating languagemodels developed from unannotated text.
Erik F. Tjong Kim Sang, Herbert Teun Kruitbosch, Marcel Broersma, Marc Esteve Del Valle
eScience1
2016 Nederlab: Towards a Single Portal and Research Environment for Diachronic Dutch Text Corpora
Hennie Brugman, Martin Reynaert, Nicoline van der Sijs, René van Stipriaan, Erik F. Tjong Kim Sang, Antal van den Bosch
LREC5
2010 GikiCLEF: Crosscultural Issues in Multilingual Information Access
Diana Santos, Luís Miguel Cabral, Corina Forascu, Pamela Forner, Fredric C. Gey, Katrin Lamm, Thomas Mandl 0001, Petya Osenova, Anselmo Peñas, Álvaro Rodrigo, Julia Maria Struß, Yvonne Skalban, Erik F. Tjong Kim Sang
LREC13
2009 Lexical Patterns or Dependency Patterns: Which Is Better for Hypernym Extraction?
Erik F. Tjong Kim Sang, Katja Hofmann
CoNLL1
2007 Extracting Hypernym Pairs from the Web
Erik F. Tjong Kim Sang
ACL1
2007 An Experiment in Automatic Classification of Pathological Reports
Janneke M. van der Zwaan, Erik F. Tjong Kim Sang, Maarten de Rijke
AIME2
2007 A Constraint Satisfaction Approach to Dependency Parsing
Sander Canisius, Erik F. Tjong Kim Sang
EMNLP-CoNLL2
2007 Automatic extension of non-english wordnets
abstract
Lexical taxonomies, such as WordNet, are important resources underlying natural language processing techniques such as machine translation and word-sense disambiguation. However, creating and maintaining such taxonomies manually is a tedious and time-consuming task. This has led to a great deal of interest in automatic methods for retrieving taxonomic relations. Efforts for both manual development of taxonomies and automatic acquisition methods have largely focused on English-language resources. Although WordNets exist for many languages, these are usually much smaller than Princeton WordNet (PWN) [1], the major semantic resource for English. For English, an excellent method has been developed for automatically extending lexical taxonomies. Snow et al. [4] show that it is possible to predict new and precise hypernymhyponym relations from a parsed corpus. Unfortunately, the method cannot be easily applied to other languages, as it relies on existing lexical resources. First, the authors use the large amount of data already contained in the PWN to extract a pattern lexicon and train their hypernym classifier. Second, the authors apply word-sense disambiguation based on an existing sense-tagged corpus. Problems in transferring the method to another language include the small size of non-English WordNets and the lack of sense-tagged corpora in other languages. In this paper we apply the method of [4] to a non-English language (Dutch) for which only a basic WordNet and a parser are available. We find that, without an additional sense-tagged corpus, this approach is highly susceptible to noise due to word sense ambiguity. We propose and evaluate two methods to address this problem.
Katja Hofmann, Erik F. Tjong Kim Sang
SIGIR2
2006 Dependency Parsing by Inference over High-recall Dependency Predictions
Sander Canisius, Toine Bogers, Antal van den Bosch, Jeroen Geertzen, Erik F. Tjong Kim Sang
CoNLL5
2005 Applying Spelling Error Correction Techniques for Improving Semantic Role Labelling
Erik F. Tjong Kim Sang, Sander Canisius, Antal van den Bosch, Toine Bogers
CoNLL1
2004 Memory-based semantic role labeling: Optimizing features, algorithm, and output
Antal van den Bosch, Sander Canisius, Walter Daelemans, Iris Hendrickx, Erik F. Tjong Kim Sang
CoNLL5
2004 Automatic Sentence Simplification for Subtitling in Dutch and English
Walter Daelemans, Anja Höthker, Erik F. Tjong Kim Sang
LREC3
2004 Using a Parallel Transcript/Subtitle Corpus for Sentence Compression
Vincent Vandeghinste, Erik F. Tjong Kim Sang
LREC2
2003 Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition
Erik F. Tjong Kim Sang, Fien De Meulder
CoNLL1
2002 Introduction to the CoNLL-2002 Shared Task: Language-Independent Named Entity Recognition
Erik F. Tjong Kim Sang
CoNLL1
2002 Memory-Based Named Entity Recognition
Erik F. Tjong Kim Sang
CoNLL1
2002 Memory-Based Shallow Parsing
Erik F. Tjong Kim Sang
J. Mach. Learn. Res.1
2000 Applying System Combination to Base Noun Phrase Identification
Erik F. Tjong Kim Sang, Walter Daelemans, Hervé Déjean, Rob Koeling, Yuval Krymolowski, Vasin Punyakanok, Dan Roth 0001
COLING1
2000 Meta-Learning for Phonemic Annotation of Corpora
Véronique Hoste, Walter Daelemans, Erik F. Tjong Kim Sang, Steven Gillis
ICML3
1999 Representing Text Chunks
Erik F. Tjong Kim Sang, Jorn Veenstra
EACL1