Gerold Schneider

dblp:13/4960 · DBLP profile ↗
← Back
25ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0002-1905-6237ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Information extraction and text analysis · 90% Trustworthy machine learning · 9% Transfer learning and domain adaptation · 1%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 6 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › text mining › authorship analysis
native language identification
0.912025
Robust Native Language Identification through Agentic Decomposition · EMNLP 2025
Natural language and speech › Information extraction and text analysis
named entity recognition
0.812024
NeuroTrialNER: An Annotated Corpus for Neurological Diseases and Therapies in Clinical Trial Registries · EMNLP 2024
Machine learning › Trustworthy machine learning › robustness › spurious correlation
spurious cue robustness
0.312025
Robust Native Language Identification through Agentic Decomposition · EMNLP 2025
Natural language and speech › Information extraction and text analysis › syntactic parsing
dependency parsing
0.112007
Pro3Gres Parser in the CoNLL Domain Adaptation Shared Task · EMNLP-CoNLL 2007
Natural language and speech › Information extraction and text analysis
syntactic parsing
0.112007
Pro3Gres Parser in the CoNLL Domain Adaptation Shared Task · EMNLP-CoNLL 2007
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.012007
Pro3Gres Parser in the CoNLL Domain Adaptation Shared Task · EMNLP-CoNLL 2007

Methods — techniques the papers use, named apart from their topics

corpus annotation · 1.5prompting · 0.9large language model · 0.9agentic decomposition · 0.9
YearPublicationVenuePosition
2025 Linguistic Features Extracted by GPT-4 Improve Alzheimer's Disease Detection based on Spontaneous Speech
abstract
Alzheimer’s Disease (AD) is a significant and growing public health concern. Investigating alterations in speech and language patterns offers a promising path towards cost-effective and non-invasive early detection of AD on a large scale. Large language models (LLMs), such as GPT, have enabled powerful new possibilities for semantic text analysis. In this study, we leverage GPT-4 to extract five semantic features from transcripts of spontaneous patient speech. The features capture known symptoms of AD, but they are difficult to quantify effectively using traditional methods of computational linguistics. We demonstrate the clinical significance of these features and further validate one of them (“Word-Finding Difficulties”) against a proxy measure and human raters. When combined with established linguistic features and a Random Forest classifier, the GPT-derived features significantly improve the detection of AD. Our approach proves effective for both manually transcribed and automatically generated transcripts, representing a novel and impactful use of recent advancements in LLMs for AD speech analysis.
Jonathan Heitz, Gerold Schneider, Nicolas Langer
COLING2
2025 Robust Native Language Identification through Agentic Decomposition
abstract
Large language models (LLMs) often achieve high performance in native language identification (NLI) benchmarks by leveraging superficial contextual clues such as names, locations, and cultural stereotypes, rather than the underlying linguistic patterns indicative of native language (L1) influence.To improve robustness, previous work has instructed LLMs to disregard such clues.In this work, we demonstrate that such a strategy is unreliable and model predictions can be easily altered by misleading hints.To address this problem, we introduce an agentic NLI pipeline inspired by forensic linguistics, where specialized agents accumulate and categorize diverse linguistic evidence before an independent final overall assessment.In this final assessment, a goal-aware coordinating agent synthesizes all evidence to make the NLI prediction.On two benchmark datasets, our approach significantly enhances NLI robustness against misleading contextual clues and performance consistency compared to standard prompting methods. 1
Ahmet Yavuz Uluslu, Tannon Kew, Tilia Ellendorff, Gerold Schneider, Rico Sennrich
EMNLP4
2024 The Influence of Automatic Speech Recognition on Linguistic Features and Automatic Alzheimer's Disease Detection from Spontaneous Speech
abstract
Alzheimer’s disease (AD) represents a major problem for society and a heavy burden for those affected. The study of changes in speech offers a potential means for large-scale AD screening that is non-invasive and inexpensive. Automatic Speech Recognition (ASR) is necessary for a fully automated system. We compare different ASR systems in terms of Word Error Rate (WER) using a publicly available benchmark dataset of speech recordings of AD patients and controls. Furthermore, this study is the first to quantify how popular linguistic features change when replacing manual transcriptions with ASR output. This contributes to the understanding of linguistic features in the context of AD detection. Moreover, we investigate how ASR affects AD classification performance by implementing two popular approaches: A fine-tuned BERT model, and Random Forest on popular linguistic features. Our results show best classification performance when using manual transcripts, but the degradation when using ASR is not dramatic. Performance stays strong, achieving an AUROC of 0.87. Our BERT-based approach is affected more strongly by ASR transcription errors than the simpler and more explainable approach based on linguistic features.
Jonathan Heitz, Gerold Schneider, Nicolas Langer
LREC/COLING2
2024 NeuroTrialNER: An Annotated Corpus for Neurological Diseases and Therapies in Clinical Trial Registries
abstract
Simona Emilova Doneva, Tilia Ellendorff, Beate Sick, Jean-Philippe Goldman, Amelia Elaine Cannon, Gerold Schneider, Benjamin Victor Ineichen. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Simona Doneva, Tilia Ellendorff, Beate Sick, Jean-Philippe Goldman, Amelia Cannon, Gerold Schneider, Benjamin Ineichen
EMNLP6
2024 Improving Adversarial Data Collection by Supporting Annotators: Lessons from GAHD, a German Hate Speech Dataset
abstract
Janis Goldzycher, Paul Röttger, Gerold Schneider. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Janis Goldzycher, Paul Röttger, Gerold Schneider
NAACL-HLT3
2023 Replicable semi-supervised approaches to state-of-the-art stance detection of tweets
abstract
Stance is defined as the expression of a speaker's standpoint towards a given target or entity. To date, the most reliable method for measuring stance is opinion surveys. However, people's increased reliance on social media makes these online platforms an essential source of complementary information about public opinion. Our study contributes to the discussion surrounding replicable methods through which to conduct reliable stance detection by establishing a rule-based model, which we replicated for several targets independently. To test our model, we relied on a widely used dataset of annotated tweets - the SemEval Task 6A dataset, which contains 5 targets with 4,163 manually labelled tweets. We relied on “off-the-shelf” sentiment lexica to expand the scope of our custom dictionaries, while also integrating linguistic markers and using word-pairs dependency information to conduct stance classification. While positive and negative evaluative words are the clearest markers of expression of stance, we demonstrate the added value of linguistic markers to identify the direction of the stance more precisely. Our model achieves an average classification accuracy of 75% (ranging from 67% to 89% across targets). This study is concluded by discussing practical implications and outlooks for future research, while highlighting that each target poses specific challenges to stance detection.
Maud Reveilhac, Gerold Schneider
Inf. Process. Manag.2
2020 Using Multilingual Resources to Evaluate CEFRLex for Learner Applications
abstract
The Common European Framework of Reference for Languages (CEFR) defines six levels of learner proficiency, and links them to particular communicative abilities. The CEFRLex project aims at compiling lexical resources that link single words and multi-word expressions to particular CEFR levels. The resources are thought to reflect second language learner needs as they are compiled from CEFR-graded textbooks and other learner-directed texts. In this work, we investigate the applicability of CEFRLex resources for building language learning applications. Our main concerns were that vocabulary in language learning materials might be sparse, i.e. that not all vocabulary items that belong to a particular level would also occur in materials for that level, and, on the other hand, that vocabulary items might be used on lower-level materials if required by the topic (e.g. with a simpler paraphrasing or translation). Our results indicate that the English CEFRLex resource is in accordance with external resources that we jointly employ as gold standard. Together with other values obtained from monolingual and parallel corpora, we can indicate which entries need to be adjusted to obtain values that are even more in line with this gold standard. We expect that this finding also holds for the other languages
Johannes Graën, David Alfter, Gerold Schneider
LREC3
2019 Cognitive Aging Effects on Language Use in Real-Life Contexts: A Naturalistic Observation Study
Minxia Luo, Gerold Schneider, Mike Martin, Burcu Demiray
CogSci2
2017 Measuring Encoding Efficiency in Swedish and English Language Learner Speech Production
abstract
We use n-gram language models to investigate how far language approximates an optimal code for human communication in terms of Information Theory, and what differences there are between Learner proficiency levels. Although the language of lower level learners is simpler, it is less optimal in terms of information theory, and as a consequence more difficult to process.
Gintare Grigonyte, Gerold Schneider
INTERSPEECH2
2012 Dependency parsing for interaction detection in pharmacogenomics
Gerold Schneider, Fabio Rinaldi 0001, Simon Clematide
LREC1
2012 Relation mining experiments in the pharmacogenomics domain
Fabio Rinaldi 0001, Gerold Schneider, Simon Clematide
J. Biomed. Informatics2
2011 The Protein-Protein Interaction tasks of BioCreative III: classification/ranking of articles and linking bio-ontology concepts to full text
abstract
BACKGROUND: Determining usefulness of biomedical text mining systems requires realistic task definition and data selection criteria without artificial constraints, measuring performance aspects that go beyond traditional metrics. The BioCreative III Protein-Protein Interaction (PPI) tasks were motivated by such considerations, trying to address aspects including how the end user would oversee the generated output, for instance by providing ranked results, textual evidence for human interpretation or measuring time savings by using automated systems. Detecting articles describing complex biological events like PPIs was addressed in the Article Classification Task (ACT), where participants were asked to implement tools for detecting PPI-describing abstracts. Therefore the BCIII-ACT corpus was provided, which includes a training, development and test set of over 12,000 PPI relevant and non-relevant PubMed abstracts labeled manually by domain experts and recording also the human classification times. The Interaction Method Task (IMT) went beyond abstracts and required mining for associations between more than 3,500 full text articles and interaction detection method ontology concepts that had been applied to detect the PPIs reported in them. RESULTS: A total of 11 teams participated in at least one of the two PPI tasks (10 in ACT and 8 in the IMT) and a total of 62 persons were involved either as participants or in preparing data sets/evaluating these tasks. Per task, each team was allowed to submit five runs offline and another five online via the BioCreative Meta-Server. From the 52 runs submitted for the ACT, the highest Matthew's Correlation Coefficient (MCC) score measured was 0.55 at an accuracy of 89% and the best AUC iP/R was 68%. Most ACT teams explored machine learning methods, some of them also used lexical resources like MeSH terms, PSI-MI concepts or particular lists of verbs and nouns, some integrated NER approaches. For the IMT, a total of 42 runs were evaluated by comparing systems against manually generated annotations done by curators from the BioGRID and MINT databases. The highest AUC iP/R achieved by any run was 53%, the best MCC score 0.55. In case of competitive systems with an acceptable recall (above 35%) the macro-averaged precision ranged between 50% and 80%, with a maximum F-Score of 55%. CONCLUSIONS: The results of the ACT task of BioCreative III indicate that classification of large unbalanced article collections reflecting the real class imbalance is still challenging. Nevertheless, text-mining tools that report ranked lists of relevant articles for manual selection can potentially reduce the time needed to identify half of the relevant articles to less than 1/4 of the time when compared to unranked results. Detecting associations between full text articles and interaction detection method PSI-MI terms (IMT) is more difficult than might be anticipated. This is due to the variability of method term mentions, errors resulting from pre-processing of articles provided as PDF files, and the heterogeneity and different granularity of method term concepts encountered in the ontology. However, combining the sophisticated techniques developed by the participants with supporting evidence strings derived from the articles for human interpretation could result in practical modules for biological annotation workflows.
Martin Krallinger, Miguel Vázquez, Florian Leitner, David Salgado, Andrew Chatr-aryamontri, Andrew G. Winter, Livia Perfetto, Leonardo Briganti, Luana Licata, Marta Iannuccelli, Luisa Castagnoli, Gianni Cesareni, Mike Tyers, Gerold Schneider, Fabio Rinaldi 0001, Robert Leaman, Graciela Gonzalez-Hernandez, Sérgio Matos, Sun Kim, W. John Wilbur, Luis M. Rocha, Hagit Shatkay, Ashish V. Tendulkar, Shashank Agarwal, Xinglong Wang, Rafal Rak, Keith Noto, Charles Elkan, Zhiyong Lu
BMC Bioinform.14
2011 Detection of interaction articles and experimental methods in biomedical literature
abstract
BACKGROUND: This article describes the approaches taken by the OntoGene group at the University of Zurich in dealing with two tasks of the BioCreative III competition: classification of articles which contain curatable protein-protein interactions (PPI-ACT) and extraction of experimental methods (PPI-IMT). RESULTS: Two main achievements are described in this paper: (a) a system for document classification which crucially relies on the results of an advanced pipeline of natural language processing tools; (b) a system which is capable of detecting all experimental methods mentioned in scientific literature, and listing them with a competitive ranking (AUC iP/R > 0.5). CONCLUSIONS: The results of the BioCreative III shared evaluation clearly demonstrate that significant progress has been achieved in the domain of biomedical text mining in the past few years. Our own contribution, together with the results of other participants, provides evidence that natural language processing techniques have become by now an integral part of advanced text mining approaches.
Gerold Schneider, Simon Clematide, Fabio Rinaldi 0001
BMC Bioinform.1
2010 OntoGene in BioCreative II.5
abstract
We describe a system for the detection of mentions of protein-protein interactions in the biomedical scientific literature. The original system was developed as a part of the OntoGene project, which focuses on using advanced computational linguistic techniques for text mining applications in the biomedical domain. In this paper, we focus in particular on the participation to the BioCreative II.5 challenge, where the OntoGene system achieved best-ranked results. Additionally, we describe a feature-analysis experiment performed after the challenge, which shows the unexpected result that one single feature alone performs better than the combination of features used in the challenge.
Fabio Rinaldi 0001, Gerold Schneider, Kaarel Kaljurand, Simon Clematide, Thérèse Vachon, Martin Romacker
IEEE ACM Trans. Comput. Biol. Bioinform.2
2009 Using Existing Biomedical Resources to Detect and Ground Terms in Biomedical Literature
Kaarel Kaljurand, Fabio Rinaldi 0001, Thomas Kappeler, Gerold Schneider
AIME4
2009 Detecting Protein-Protein Interactions in Biomedical Texts Using a Parser and Linguistic Resources
Gerold Schneider, Kaarel Kaljurand, Fabio Rinaldi 0001
CICLing1
2008 Dependency-Based Relation Mining for Biomedical Literature
Fabio Rinaldi 0001, Gerold Schneider, Kaarel Kaljurand, Michael Hess 0001
LREC2
2007 Pro3Gres Parser in the CoNLL Domain Adaptation Shared Task
Gerold Schneider, Kaarel Kaljurand, Fabio Rinaldi 0001, Tobias Kuhn
EMNLP-CoNLL1
2007 Mining of relations between proteins over biomedical scientific literature using a deep-linguistic approach
Fabio Rinaldi 0001, Gerold Schneider, Kaarel Kaljurand, Michael Hess 0001, Christos Andronis, Ourania Konstanti, Andreas Persidis
Artif. Intell. Medicine2
2006 Tools for Text Mining over Biomedical Literature
Fabio Rinaldi 0001, Gerold Schneider, Kaarel Kaljurand, Michael Hess 0001
ECAI2
2006 An environment for relation mining over richly annotated corpora: the case of GENIA
abstract
BACKGROUND: The biomedical domain is witnessing a rapid growth of the amount of published scientific results, which makes it increasingly difficult to filter the core information. There is a real need for support tools that 'digest' the published results and extract the most important information. RESULTS: We describe and evaluate an environment supporting the extraction of domain-specific relations, such as protein-protein interactions, from a richly-annotated corpus. We use full, deep-linguistic parsing and manually created, versatile patterns, expressing a large set of syntactic alternations, plus semantic ontology information. CONCLUSION: The experiments show that our approach described is capable of delivering high-precision results, while maintaining sufficient levels of recall. The high level of abstraction of the rules used by the system, which are considerably more powerful and versatile than finite-state approaches, allows speedy interactive development and validation.
Fabio Rinaldi 0001, Gerold Schneider, Kaarel Kaljurand, Michael Hess 0001, Martin Romacker
BMC Bioinform.2
2005 Relation Mining over a Corpus of Scientific Literature
Fabio Rinaldi 0001, Gerold Schneider, Kaarel Kaljurand, Michael Hess 0001, Christos Andronis, Andreas Persidis, Ourania Konstanti
AIME2
2003 A low-complexity, broad-coverage probabilistic Dependency Parser for English
Gerold Schneider
HLT-NAACL1
2002 Answer Extraction in Technical Domains
Fabio Rinaldi 0001, Michael Hess 0001, Diego Mollá Aliod, Rolf Schwitter, James Dowdall, Gerold Schneider, Rachel Fournier
CICLing6
2002 Using Syntactic Analysis to Increase Efficiency in Visualizing Text Collections
James Henderson 0001, Paola Merlo, Ivan Petroff, Gerold Schneider
COLING4