VLDB 2026 Research / reviewers in the wild / expert
Gerold Schneider
dblp:13/4960
· DBLP profile ↗
25ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0002-1905-6237ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Information extraction and text analysis · 90% Trustworthy machine learning · 9% Transfer learning and domain adaptation · 1% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% |
Topics — the 6 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › text mining › authorship analysis
native language identification |
0.9 | 1 | 2025 | Robust Native Language Identification through Agentic Decomposition · EMNLP 2025 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.8 | 1 | 2024 | NeuroTrialNER: An Annotated Corpus for Neurological Diseases and Therapies in Clinical Trial Registries · EMNLP 2024 |
Machine learning › Trustworthy machine learning › robustness › spurious correlation
spurious cue robustness |
0.3 | 1 | 2025 | Robust Native Language Identification through Agentic Decomposition · EMNLP 2025 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
dependency parsing |
0.1 | 1 | 2007 | Pro3Gres Parser in the CoNLL Domain Adaptation Shared Task · EMNLP-CoNLL 2007 |
Natural language and speech › Information extraction and text analysis
syntactic parsing |
0.1 | 1 | 2007 | Pro3Gres Parser in the CoNLL Domain Adaptation Shared Task · EMNLP-CoNLL 2007 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
0.0 | 1 | 2007 | Pro3Gres Parser in the CoNLL Domain Adaptation Shared Task · EMNLP-CoNLL 2007 |
Methods — techniques the papers use, named apart from their topics
corpus annotation · 1.5prompting · 0.9large language model · 0.9agentic decomposition · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Linguistic Features Extracted by GPT-4 Improve Alzheimer's Disease Detection based on Spontaneous SpeechabstractAlzheimer’s Disease (AD) is a significant and growing public health concern. Investigating alterations in speech and language patterns offers a promising path towards cost-effective and non-invasive early detection of AD on a large scale. Large language models (LLMs), such as GPT, have enabled powerful new possibilities for semantic text analysis. In this study, we leverage GPT-4 to extract five semantic features from transcripts of spontaneous patient speech. The features capture known symptoms of AD, but they are difficult to quantify effectively using traditional methods of computational linguistics. We demonstrate the clinical significance of these features and further validate one of them (“Word-Finding Difficulties”) against a proxy measure and human raters. When combined with established linguistic features and a Random Forest classifier, the GPT-derived features significantly improve the detection of AD. Our approach proves effective for both manually transcribed and automatically generated transcripts, representing a novel and impactful use of recent advancements in LLMs for AD speech analysis. Jonathan Heitz, Gerold Schneider, Nicolas Langer |
COLING | 2 |
| 2025 | Robust Native Language Identification through Agentic DecompositionabstractLarge language models (LLMs) often achieve high performance in native language identification (NLI) benchmarks by leveraging superficial contextual clues such as names, locations, and cultural stereotypes, rather than the underlying linguistic patterns indicative of native language (L1) influence.To improve robustness, previous work has instructed LLMs to disregard such clues.In this work, we demonstrate that such a strategy is unreliable and model predictions can be easily altered by misleading hints.To address this problem, we introduce an agentic NLI pipeline inspired by forensic linguistics, where specialized agents accumulate and categorize diverse linguistic evidence before an independent final overall assessment.In this final assessment, a goal-aware coordinating agent synthesizes all evidence to make the NLI prediction.On two benchmark datasets, our approach significantly enhances NLI robustness against misleading contextual clues and performance consistency compared to standard prompting methods. 1 Ahmet Yavuz Uluslu, Tannon Kew, Tilia Ellendorff, Gerold Schneider, Rico Sennrich |
EMNLP | 4 |
| 2024 | The Influence of Automatic Speech Recognition on Linguistic Features and Automatic Alzheimer's Disease Detection from Spontaneous SpeechabstractAlzheimer’s disease (AD) represents a major problem for society and a heavy burden for those affected. The study of changes in speech offers a potential means for large-scale AD screening that is non-invasive and inexpensive. Automatic Speech Recognition (ASR) is necessary for a fully automated system. We compare different ASR systems in terms of Word Error Rate (WER) using a publicly available benchmark dataset of speech recordings of AD patients and controls. Furthermore, this study is the first to quantify how popular linguistic features change when replacing manual transcriptions with ASR output. This contributes to the understanding of linguistic features in the context of AD detection. Moreover, we investigate how ASR affects AD classification performance by implementing two popular approaches: A fine-tuned BERT model, and Random Forest on popular linguistic features. Our results show best classification performance when using manual transcripts, but the degradation when using ASR is not dramatic. Performance stays strong, achieving an AUROC of 0.87. Our BERT-based approach is affected more strongly by ASR transcription errors than the simpler and more explainable approach based on linguistic features. Jonathan Heitz, Gerold Schneider, Nicolas Langer |
LREC/COLING | 2 |
| 2024 | NeuroTrialNER: An Annotated Corpus for Neurological Diseases and Therapies in Clinical Trial RegistriesabstractSimona Emilova Doneva, Tilia Ellendorff, Beate Sick, Jean-Philippe Goldman, Amelia Elaine Cannon, Gerold Schneider, Benjamin Victor Ineichen. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Simona Doneva, Tilia Ellendorff, Beate Sick, Jean-Philippe Goldman, Amelia Cannon, Gerold Schneider, Benjamin Ineichen |
EMNLP | 6 |
| 2024 | Improving Adversarial Data Collection by Supporting Annotators: Lessons from GAHD, a German Hate Speech DatasetabstractJanis Goldzycher, Paul Röttger, Gerold Schneider. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Janis Goldzycher, Paul Röttger, Gerold Schneider |
NAACL-HLT | 3 |
| 2023 | Replicable semi-supervised approaches to state-of-the-art stance detection of tweetsabstractStance is defined as the expression of a speaker's standpoint towards a given target or entity. To date, the most reliable method for measuring stance is opinion surveys. However, people's increased reliance on social media makes these online platforms an essential source of complementary information about public opinion. Our study contributes to the discussion surrounding replicable methods through which to conduct reliable stance detection by establishing a rule-based model, which we replicated for several targets independently. To test our model, we relied on a widely used dataset of annotated tweets - the SemEval Task 6A dataset, which contains 5 targets with 4,163 manually labelled tweets. We relied on “off-the-shelf” sentiment lexica to expand the scope of our custom dictionaries, while also integrating linguistic markers and using word-pairs dependency information to conduct stance classification. While positive and negative evaluative words are the clearest markers of expression of stance, we demonstrate the added value of linguistic markers to identify the direction of the stance more precisely. Our model achieves an average classification accuracy of 75% (ranging from 67% to 89% across targets). This study is concluded by discussing practical implications and outlooks for future research, while highlighting that each target poses specific challenges to stance detection. Maud Reveilhac, Gerold Schneider |
Inf. Process. Manag. | 2 |
| 2020 | Using Multilingual Resources to Evaluate CEFRLex for Learner ApplicationsabstractThe Common European Framework of Reference for Languages (CEFR) defines six levels of learner proficiency, and links them to particular communicative abilities. The CEFRLex project aims at compiling lexical resources that link single words and multi-word expressions to particular CEFR levels. The resources are thought to reflect second language learner needs as they are compiled from CEFR-graded textbooks and other learner-directed texts. In this work, we investigate the applicability of CEFRLex resources for building language learning applications. Our main concerns were that vocabulary in language learning materials might be sparse, i.e. that not all vocabulary items that belong to a particular level would also occur in materials for that level, and, on the other hand, that vocabulary items might be used on lower-level materials if required by the topic (e.g. with a simpler paraphrasing or translation). Our results indicate that the English CEFRLex resource is in accordance with external resources that we jointly employ as gold standard. Together with other values obtained from monolingual and parallel corpora, we can indicate which entries need to be adjusted to obtain values that are even more in line with this gold standard. We expect that this finding also holds for the other languages Johannes Graën, David Alfter, Gerold Schneider |
LREC | 3 |
| 2019 | Cognitive Aging Effects on Language Use in Real-Life Contexts: A Naturalistic Observation Study
Minxia Luo, Gerold Schneider, Mike Martin, Burcu Demiray |
CogSci | 2 |
| 2017 | Measuring Encoding Efficiency in Swedish and English Language Learner Speech ProductionabstractWe use n-gram language models to investigate how far language approximates an optimal code for human communication in terms of Information Theory, and what differences there are between Learner proficiency levels. Although the language of lower level learners is simpler, it is less optimal in terms of information theory, and as a consequence more difficult to process. Gintare Grigonyte, Gerold Schneider |
INTERSPEECH | 2 |
| 2012 | Dependency parsing for interaction detection in pharmacogenomics
Gerold Schneider, Fabio Rinaldi 0001, Simon Clematide |
LREC | 1 |
| 2012 | Relation mining experiments in the pharmacogenomics domain
Fabio Rinaldi 0001, Gerold Schneider, Simon Clematide |
J. Biomed. Informatics | 2 |
| 2011 | The Protein-Protein Interaction tasks of BioCreative III: classification/ranking of articles and linking bio-ontology concepts to full textabstractBACKGROUND: Determining usefulness of biomedical text mining systems requires realistic task definition and data selection criteria without artificial constraints, measuring performance aspects that go beyond traditional metrics. The BioCreative III Protein-Protein Interaction (PPI) tasks were motivated by such considerations, trying to address aspects including how the end user would oversee the generated output, for instance by providing ranked results, textual evidence for human interpretation or measuring time savings by using automated systems. Detecting articles describing complex biological events like PPIs was addressed in the Article Classification Task (ACT), where participants were asked to implement tools for detecting PPI-describing abstracts. Therefore the BCIII-ACT corpus was provided, which includes a training, development and test set of over 12,000 PPI relevant and non-relevant PubMed abstracts labeled manually by domain experts and recording also the human classification times. The Interaction Method Task (IMT) went beyond abstracts and required mining for associations between more than 3,500 full text articles and interaction detection method ontology concepts that had been applied to detect the PPIs reported in them. RESULTS: A total of 11 teams participated in at least one of the two PPI tasks (10 in ACT and 8 in the IMT) and a total of 62 persons were involved either as participants or in preparing data sets/evaluating these tasks. Per task, each team was allowed to submit five runs offline and another five online via the BioCreative Meta-Server. From the 52 runs submitted for the ACT, the highest Matthew's Correlation Coefficient (MCC) score measured was 0.55 at an accuracy of 89% and the best AUC iP/R was 68%. Most ACT teams explored machine learning methods, some of them also used lexical resources like MeSH terms, PSI-MI concepts or particular lists of verbs and nouns, some integrated NER approaches. For the IMT, a total of 42 runs were evaluated by comparing systems against manually generated annotations done by curators from the BioGRID and MINT databases. The highest AUC iP/R achieved by any run was 53%, the best MCC score 0.55. In case of competitive systems with an acceptable recall (above 35%) the macro-averaged precision ranged between 50% and 80%, with a maximum F-Score of 55%. CONCLUSIONS: The results of the ACT task of BioCreative III indicate that classification of large unbalanced article collections reflecting the real class imbalance is still challenging. Nevertheless, text-mining tools that report ranked lists of relevant articles for manual selection can potentially reduce the time needed to identify half of the relevant articles to less than 1/4 of the time when compared to unranked results. Detecting associations between full text articles and interaction detection method PSI-MI terms (IMT) is more difficult than might be anticipated. This is due to the variability of method term mentions, errors resulting from pre-processing of articles provided as PDF files, and the heterogeneity and different granularity of method term concepts encountered in the ontology. However, combining the sophisticated techniques developed by the participants with supporting evidence strings derived from the articles for human interpretation could result in practical modules for biological annotation workflows. Martin Krallinger, Miguel Vázquez, Florian Leitner, David Salgado, Andrew Chatr-aryamontri, Andrew G. Winter, Livia Perfetto, Leonardo Briganti, Luana Licata, Marta Iannuccelli, Luisa Castagnoli, Gianni Cesareni, Mike Tyers, Gerold Schneider, Fabio Rinaldi 0001, Robert Leaman, Graciela Gonzalez-Hernandez, Sérgio Matos, Sun Kim, W. John Wilbur, Luis M. Rocha, Hagit Shatkay, Ashish V. Tendulkar, Shashank Agarwal, Xinglong Wang, Rafal Rak, Keith Noto, Charles Elkan, Zhiyong Lu |
BMC Bioinform. | 14 |
| 2011 | Detection of interaction articles and experimental methods in biomedical literatureabstractBACKGROUND: This article describes the approaches taken by the OntoGene group at the University of Zurich in dealing with two tasks of the BioCreative III competition: classification of articles which contain curatable protein-protein interactions (PPI-ACT) and extraction of experimental methods (PPI-IMT). RESULTS: Two main achievements are described in this paper: (a) a system for document classification which crucially relies on the results of an advanced pipeline of natural language processing tools; (b) a system which is capable of detecting all experimental methods mentioned in scientific literature, and listing them with a competitive ranking (AUC iP/R > 0.5). CONCLUSIONS: The results of the BioCreative III shared evaluation clearly demonstrate that significant progress has been achieved in the domain of biomedical text mining in the past few years. Our own contribution, together with the results of other participants, provides evidence that natural language processing techniques have become by now an integral part of advanced text mining approaches. Gerold Schneider, Simon Clematide, Fabio Rinaldi 0001 |
BMC Bioinform. | 1 |
| 2010 | OntoGene in BioCreative II.5abstractWe describe a system for the detection of mentions of protein-protein interactions in the biomedical scientific literature. The original system was developed as a part of the OntoGene project, which focuses on using advanced computational linguistic techniques for text mining applications in the biomedical domain. In this paper, we focus in particular on the participation to the BioCreative II.5 challenge, where the OntoGene system achieved best-ranked results. Additionally, we describe a feature-analysis experiment performed after the challenge, which shows the unexpected result that one single feature alone performs better than the combination of features used in the challenge. Fabio Rinaldi 0001, Gerold Schneider, Kaarel Kaljurand, Simon Clematide, Thérèse Vachon, Martin Romacker |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2009 | Using Existing Biomedical Resources to Detect and Ground Terms in Biomedical Literature
Kaarel Kaljurand, Fabio Rinaldi 0001, Thomas Kappeler, Gerold Schneider |
AIME | 4 |
| 2009 | Detecting Protein-Protein Interactions in Biomedical Texts Using a Parser and Linguistic Resources
Gerold Schneider, Kaarel Kaljurand, Fabio Rinaldi 0001 |
CICLing | 1 |
| 2008 | Dependency-Based Relation Mining for Biomedical Literature
Fabio Rinaldi 0001, Gerold Schneider, Kaarel Kaljurand, Michael Hess 0001 |
LREC | 2 |
| 2007 | Pro3Gres Parser in the CoNLL Domain Adaptation Shared Task
Gerold Schneider, Kaarel Kaljurand, Fabio Rinaldi 0001, Tobias Kuhn |
EMNLP-CoNLL | 1 |
| 2007 | Mining of relations between proteins over biomedical scientific literature using a deep-linguistic approach
Fabio Rinaldi 0001, Gerold Schneider, Kaarel Kaljurand, Michael Hess 0001, Christos Andronis, Ourania Konstanti, Andreas Persidis |
Artif. Intell. Medicine | 2 |
| 2006 | Tools for Text Mining over Biomedical Literature
Fabio Rinaldi 0001, Gerold Schneider, Kaarel Kaljurand, Michael Hess 0001 |
ECAI | 2 |
| 2006 | An environment for relation mining over richly annotated corpora: the case of GENIAabstractBACKGROUND: The biomedical domain is witnessing a rapid growth of the amount of published scientific results, which makes it increasingly difficult to filter the core information. There is a real need for support tools that 'digest' the published results and extract the most important information. RESULTS: We describe and evaluate an environment supporting the extraction of domain-specific relations, such as protein-protein interactions, from a richly-annotated corpus. We use full, deep-linguistic parsing and manually created, versatile patterns, expressing a large set of syntactic alternations, plus semantic ontology information. CONCLUSION: The experiments show that our approach described is capable of delivering high-precision results, while maintaining sufficient levels of recall. The high level of abstraction of the rules used by the system, which are considerably more powerful and versatile than finite-state approaches, allows speedy interactive development and validation. Fabio Rinaldi 0001, Gerold Schneider, Kaarel Kaljurand, Michael Hess 0001, Martin Romacker |
BMC Bioinform. | 2 |
| 2005 | Relation Mining over a Corpus of Scientific Literature
Fabio Rinaldi 0001, Gerold Schneider, Kaarel Kaljurand, Michael Hess 0001, Christos Andronis, Andreas Persidis, Ourania Konstanti |
AIME | 2 |
| 2003 | A low-complexity, broad-coverage probabilistic Dependency Parser for English
Gerold Schneider |
HLT-NAACL | 1 |
| 2002 | Answer Extraction in Technical Domains
Fabio Rinaldi 0001, Michael Hess 0001, Diego Mollá Aliod, Rolf Schwitter, James Dowdall, Gerold Schneider, Rachel Fournier |
CICLing | 6 |
| 2002 | Using Syntactic Analysis to Increase Efficiency in Visualizing Text Collections
James Henderson 0001, Paola Merlo, Ivan Petroff, Gerold Schneider |
COLING | 4 |