VLDB 2026 Research / reviewers in the wild / expert
Felix A. Gers
dblp:79/2607 · also Felix Alexander Gers
· DBLP profile ↗
28ranked-venue papers
8as first author
10since 2021 · last 2026
0000-0002-2258-2983ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 8 first-author · 9 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking the Idiomaticity Decomposability Hypothesis: Evidence from Distributional LearningabstractMaggie Mi, Golzar Atefi, Atsuki Yamaguchi, Felix Gers, Aline Villavicencio, Nafise Sadat Moosavi. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Maggie Mi, Golzar Atefi, Atsuki Yamaguchi, Felix A. Gers, Aline Villavicencio, Nafise Sadat Moosavi |
ACL (1) | 4 |
| 2026 | DeepICD-R1: Medical Reasoning through Hierarchical Rewards and Unsupervised Distillation
Tom Röhr, Thomas Steffek, Roman Teucher, Keno Bressem, Alexei Figueroa Rosero, Paul Grundmann, Peter Tröger, Felix A. Gers, Alexander Löser |
LREC | 8 |
| 2025 | "Where does it hurt?" - Dataset and Study on Physician Intent Trajectories in Doctor Patient DialoguesabstractIn a doctor-patient dialogue, the primary objective of physicians is to diagnose patients and propose a treatment plan. Medical doctors guide these conversations through targeted questioning to efficiently gather the information required to provide the best possible outcomes for patients. To the best of our knowledge, this is the first work that studies physician intent trajectories in doctor-patient dialogues. We use the ‘Ambient Clinical Intelligence Benchmark’ (Aci-bench) dataset for our study. We collaborate with medical professionals to develop a fine-grained taxonomy of physician intents based on the SOAP framework (Subjective, Objective, Assessment, and Plan). We then conduct a large-scale annotation effort to label over 5000 doctor-patient turns with the help of a large number of medical experts recruited using Prolific, a popular crowd-sourcing platform. This large labeled dataset is an important resource contribution that we use for benchmarking the state-of-the-art generative and encoder models for medical intent classification tasks. Our findings show that our models understand the general structure of medical dialogues with high accuracy, but often fail to identify transitions between SOAP categories. We also report for the first time common trajectories in medical dialogue structures that provide valuable insights for designing ‘differential diagnosis’ systems. Finally, we extensively study the impact of intent filtering for medical dialogue summarization and observe a significant boost in performance. We make the codes and data, including annotation guidelines, publicly available at https://github.com/DATEXIS/medical-intent-classification. Tom Röhr, Soumyadeep Roy, Fares Al Mohamad, Jens-Michalis Papaioannou, Wolfgang Nejdl, Felix A. Gers, Alexander Löser |
ECAI | 6 |
| 2024 | DDxGym: Online Transformer Policies in a Knowledge Graph Based Natural Language EnvironmentabstractDifferential diagnosis (DDx) is vital for physicians and challenging due to the existence of numerous diseases and their complex symptoms. Model training for this task is generally hindered by limited data access due to privacy concerns. To address this, we present DDxGym, a specialized OpenAI Gym environment for clinical differential diagnosis. DDxGym formulates DDx as a natural-language-based reinforcement learning (RL) problem, where agents emulate medical professionals, selecting examinations and treatments for patients with randomly sampled diseases. This RL environment utilizes data labeled from online resources, evaluated by medical professionals for accuracy. Transformers, while effective for encoding text in DDxGym, are unstable in online RL. For that reason we propose a novel training method using an auxiliary masked language modeling objective for policy optimization, resulting in model stabilization and significant performance improvement over strong baselines. Following this approach, our agent effectively navigates large action spaces and identifies universally applicable actions. All data, environment details, and implementation, including experiment reproduction code, are made publicly available. Benjamin Winter, Alexei Figueroa Rosero, Alexander Löser, Felix A. Gers, Nancy Katerina Figueroa Rosero, Ralf Krestel |
LREC/COLING | 4 |
| 2024 | Boosting Long-Tail Data Classification with Sparse Prototypical Networks
Alexei Figueroa Rosero, Jens-Michalis Papaioannou, Conor Fallon, Alexandra Bekiaridou, Keno Bressem, Stavros Zanos, Felix A. Gers, Wolfgang Nejdl, Alexander Löser |
ECML/PKDD (7) | 7 |
| 2022 | Attention Networks for Augmenting Clinical Text with Support Sets for Diagnosis PredictionabstractDiagnosis prediction on admission notes is a core clinical task. However, these notes may incompletely describe the patient. Also, clinical language models may suffer from idiosyncratic language or imbalanced vocabulary for describing diseases or symptoms. We tackle the task of diagnosis prediction, which consists of predicting future patient diagnoses from clinical texts at the time of admission. We improve the performance on this task by introducing an additional signal from support sets of diagnostic codes from prior admissions or as they emerge during differential diagnosis. To enhance the robustness of diagnosis prediction methods, we propose to augment clinical text with potentially complementary set data from diagnosis codes from previous patient visits or from codes that emerge from the current admission as they become available through diagnostics. We discuss novel attention network architectures and augmentation strategies to solve this problem. Our experiments reveal that support sets improve the performance drastically to predict less common diagnosis codes. Our approach clearly outperforms the previous state-of-the-art PubMedBERT baseline by up 3% points. Furthermore, we find that support sets drastically improve the performance for pregnancy- and gynecology-related diagnoses up to 32.9% points compared to the baseline. Paul Grundmann, Tom Oberhauser, Felix A. Gers, Alexander Löser |
COLING | 3 |
| 2022 | Cross-Lingual Knowledge Transfer for Clinical PhenotypingabstractClinical phenotyping enables the automatic extraction of clinical conditions from patient records, which can be beneficial to doctors and clinics worldwide. However, current state-of-the-art models are mostly applicable to clinical notes written in English. We therefore investigate cross-lingual knowledge transfer strategies to execute this task for clinics that do not use the English language and have a small amount of in-domain data available. Our results reveal two strategies that outperform the state-of-the-art: Translation-based methods in combination with domain-specific encoders and cross-lingual encoders plus adapters. We find that these strategies perform especially well for classifying rare phenotypes and we advise on which method to prefer in which situation. Our results show that using multilingual data overall improves clinical phenotyping models and can compensate for data sparseness. Jens-Michalis Papaioannou, Paul Grundmann, Betty van Aken, Athanasios Samaras, Ilias Kyparissidis, George Giannakoulas, Felix A. Gers, Alexander Löser |
LREC | 7 |
| 2022 | KIMERA: Injecting Domain Knowledge into Vacant Transformer HeadsabstractTraining transformer language models requires vast amounts of text and computational resources. This drastically limits the usage of these models in niche domains for which they are not optimized, or where domain-specific training data is scarce. We focus here on the clinical domain because of its limited access to training data in common tasks, while structured ontological data is often readily available. Recent observations in model compression of transformer models show optimization potential in improving the representation capacity of attention heads. We propose KIMERA (Knowledge Injection via Mask Enforced Retraining of Attention) for detecting, retraining and instilling attention heads with complementary structured domain knowledge. Our novel multi-task training scheme effectively identifies and targets individual attention heads that are least useful for a given downstream task and optimizes their representation with information from structured data. KIMERA generalizes well, thereby building the basis for an efficient fine-tuning. KIMERA achieves significant performance boosts on seven datasets in the medical domain in Information Retrieval and Clinical Outcome Prediction settings. We apply KIMERA to BERT-base to evaluate the extent of the domain transfer and also improve on the already strong results of BioBERT in the clinical domain. Benjamin Winter, Alexei Figueroa Rosero, Alexander Löser, Felix A. Gers, Amy Siu |
LREC | 4 |
| 2021 | Clinical Outcome Prediction from Admission Notes using Self-Supervised Knowledge IntegrationabstractBetty van Aken, Jens-Michalis Papaioannou, Manuel Mayrdorfer, Klemens Budde, Felix Gers, Alexander Loeser. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Betty van Aken, Jens-Michalis Papaioannou, Manuel Mayrdorfer, Klemens Budde, Felix A. Gers, Alexander Löser |
EACL | 5 |
| 2021 | Aspect-Based Passage Retrieval with Contextualized Discourse Vectors
Jens-Michalis Papaioannou, Manuel Mayrdorfer, Sebastian Arnold 0001, Felix A. Gers, Klemens Budde, Alexander Löser |
ECIR (2) | 4 |
| 2020 | Is Language Modeling Enough? Evaluating Effective Embedding CombinationsabstractUniversal embeddings, such as BERT or ELMo, are useful for a broad set of natural language processing tasks like text classification or sentiment analysis. Moreover, specialized embeddings also exist for tasks like topic modeling or named entity disambiguation. We study if we can complement these universal embeddings with specialized embeddings. We conduct an in-depth evaluation of nine well known natural language understanding tasks with SentEval. Also, we extend SentEval with two additional tasks to the medical domain. We present PubMedSection, a novel topic classification dataset focussed on the biomedical domain. Our comprehensive analysis covers 11 tasks and combinations of six embeddings. We report that combined embeddings outperform state of the art universal embeddings without any embedding fine-tuning. We observe that adding topic model based embeddings helps for most tasks and that differing pre-training tasks encode complementary features. Moreover, we present new state of the art results on the MPQA and SUBJ tasks in SentEval. Rudolf Schneider 0001, Tom Oberhauser, Paul Grundmann, Felix A. Gers, Alexander Löser, Steffen Staab |
LREC | 4 |
| 2020 | Learning Contextualized Document Representations for Healthcare Answer RetrievalabstractWe present Contextual Discourse Vectors (CDV), a distributed document representation for efficient answer retrieval from long healthcare documents. Our approach is based on structured query tuples of entities and aspects from free text and medical taxonomies. Our model leverages a dual encoder architecture with hierarchical LSTM layers and multi-task training to encode the position of clinical entities and aspects alongside the document discourse. We use our continuous representations to resolve queries with short latency using approximate nearest neighbor search on sentence level. We apply the CDV model for retrieving coherent answer passages from nine English public health resources from the Web, addressing both patients and medical professionals. Because there is no end-to-end training data available for all application scenarios, we train our model with self-supervised data from Wikipedia. We show that our generalized model significantly outperforms several state-of-the-art baselines for healthcare passage ranking and is able to adapt to heterogeneous domains without additional fine-tuning. Sebastian Arnold 0001, Betty van Aken, Paul Grundmann, Felix A. Gers, Alexander Löser |
WWW | 4 |
| 2019 | How Does BERT Answer Questions?: A Layer-Wise Analysis of Transformer RepresentationsabstractBidirectional Encoder Representations from Transformers (BERT) reach state-of-the-art results in a variety of Natural Language Processing tasks. However, understanding of their internal functioning is still insufficient and unsatisfactory. In order to better understand BERT and other Transformer-based models, we present a layer-wise analysis of BERT's hidden states. Unlike previous research, which mainly focuses on explaining Transformer models by their attention weights, we argue that hidden states contain equally valuable information. Specifically, our analysis focuses on models fine-tuned on the task of Question Answering (QA) as an example of a complex downstream task. We inspect how QA models transform token vectors in order to find the correct answer. To this end, we apply a set of general and QA-specific probing tasks that reveal the information stored in each representation layer. Our qualitative analysis of hidden state visualizations provides additional insights into BERT's reasoning process. Our results show that the transformations within BERT go through phases that are related to traditional pipeline tasks. The system can therefore implicitly incorporate task-specific information into its token representations. Furthermore, our analysis reveals that fine-tuning has little impact on the models' semantic abilities and that prediction errors can be recognized in the vector representations of even early layers. Betty van Aken, Benjamin Winter, Alexander Löser, Felix A. Gers |
CIKM | 4 |
| 2019 | SECTOR: A Neural Model for Coherent Topic Segmentation and ClassificationabstractWhen searching for information, a human reader first glances over a document, spots relevant sections, and then focuses on a few sentences for resolving her intention. However, the high variance of document structure complicates the identification of the salient topic of a given section at a glance. To tackle this challenge, we present SECTOR, a model to support machine reading systems by segmenting documents into coherent sections and assigning topic labels to each section. Our deep neural network architecture learns a latent topic embedding over the course of a document. This can be leveraged to classify local topics from plain text and segment a document at topic shifts. In addition, we contribute WikiSection, a publicly available data set with 242k labeled sections in English and German from two distinct domains: diseases and cities. From our extensive evaluation of 20 architectures, we report a highest score of 71.6% F1 for the segmentation and classification of 30 topics from the English city domain, scored by our SECTOR long short-term memory model with Bloom filter embeddings and bidirectional segmentation. This is a significant improvement of 29.5 points F1 over state-of-the-art CNN classifiers with baseline segmentation. Sebastian Arnold 0001, Rudolf Schneider 0001, Philippe Cudré-Mauroux, Felix A. Gers, Alexander Löser |
Trans. Assoc. Comput. Linguistics | 4 |
| 2003 | Kalman filters improve LSTM network performance in problems unsolvable by traditional recurrent nets
Juan Antonio Pérez-Ortiz, Felix A. Gers, Douglas Eck, Jürgen Schmidhuber |
Neural Networks | 2 |
| 2002 | DEKF-LSTM
Felix A. Gers, Juan Antonio Pérez-Ortiz, Douglas Eck, Jürgen Schmidhuber |
ESANN | 1 |
| 2002 | Learning Context Sensitive Languages with LSTM Trained with Kalman Filters
Felix A. Gers, Juan Antonio Pérez-Ortiz, Douglas Eck, Jürgen Schmidhuber |
ICANN | 1 |
| 2002 | Improving Long-Term Online Prediction with Decoupled Extended Kalman Filters
Juan Antonio Pérez-Ortiz, Jürgen Schmidhuber, Felix A. Gers, Douglas Eck |
ICANN | 3 |
| 2002 | Learning Precise Timing with LSTM Recurrent Networks
Felix A. Gers, Nicol N. Schraudolph, Jürgen Schmidhuber |
J. Mach. Learn. Res. | 1 |
| 2002 | Learning Nonregular Languages: A Comparison of Simple Recurrent Networks and LSTMabstractIn response to Rodriguez's recent article (2001), we compare the performance of simple recurrent nets and long short-term memory recurrent nets on context-free and context-sensitive languages. Jürgen Schmidhuber, Felix A. Gers, Douglas Eck |
Neural Comput. | 2 |
| 2001 | Applying LSTM to Time Series Predictable through Time-Window Approaches
Felix A. Gers, Douglas Eck, Jürgen Schmidhuber |
ICANN | 1 |
| 2001 | LSTM recurrent networks learn simple context-free and context-sensitive languagesabstractPrevious work on learning regular languages from exemplary training sequences showed that long short-term memory (LSTM) outperforms traditional recurrent neural networks (RNNs). We demonstrate LSTMs superior performance on context-free language benchmarks for RNNs, and show that it works even better than previous hardwired or highly specialized architectures. To the best of our knowledge, LSTM variants are also the first RNNs to learn a simple context-sensitive language, namely a(n)b(n)c(n). Felix A. Gers, Jürgen Schmidhuber |
IEEE Trans. Neural Networks | 1 |
| 2000 | Recurrent Nets that Time and CountabstractThe size of the time intervals between events conveys information essential for numerous sequential tasks such as motor control and rhythm detection. While hidden Markov models tend to ignore this information, recurrent neural networks (RNN) can in principle learn to make use of it. We focus on long short-term memory (LSTM) because it usually outperforms other RNN. Surprisingly, LSTM augmented by "peephole connections" from its internal cells to its multiplicative gates can learn the fine distinction between sequences of spikes separated by either 50 or 49 discrete time steps, without the help of any short training exemplars. Without external resets or teacher forcing or loss of performance on tasks reported earlier, our LSTM variant also learns to generate very stable sequences of highly nonlinear, precisely timed spikes. This makes LSTM a promising approach for real-world tasks that require to time and count. Felix A. Gers, Jürgen Schmidhuber |
IJCNN (3) | 1 |
| 2000 | Neural Processing of Complex Continual Input StreamsabstractLong short-term memory (LSTM) can learn algorithms for temporal pattern processing not learnable by alternative recurrent neural networks or other methods such as hidden Markov models and symbolic grammar learning. Here, we present tasks involving arithmetic operations on continual input streams that even LSTM cannot solve. However, an LSTM variant based on "forget gates," has superior arithmetic capabilities and does solve the tasks. Felix A. Gers, Jürgen Schmidhuber |
IJCNN (4) | 1 |
| 2000 | Learning to Forget: Continual Prediction with LSTMabstractLong short-term memory (LSTM; Hochreiter & Schmidhuber, 1997) can solve numerous tasks not solvable by previous learning algorithms for recurrent neural networks (RNNs). We identify a weakness of LSTM networks processing continual input streams that are not a priori segmented into subsequences with explicitly marked ends at which the network's internal state could be reset. Without resets, the state may grow indefinitely and eventually cause the network to break down. Our remedy is a novel, adaptive "forget gate" that enables an LSTM cell to learn to reset itself at appropriate times, thus releasing internal resources. We review illustrative benchmark problems on which standard LSTM outperforms other RNN algorithms. All algorithms (including LSTM) fail to solve continual versions of these problems. LSTM with forget gates, however, easily solves them, and in an elegant way. Felix A. Gers, Jürgen Schmidhuber, Fred A. Cummins |
Neural Comput. | 1 |
| 1999 | ATR's artificial brain (CAM-brain) project: a sample of what individual CoDi-1Bit model evolved neural net modules can doabstractThe paper presents a sample of what evolved neural net circuit modules using the so called "CoDi-1Bit" neural net model can do. This work is part of an 8 year research project at ATR which aims to build an artificial brain containing a billion neurons by the year 2001, that will be used to control the behaviors of a kitten robot "Robokoneko". It looks as though the figure is more likely to be 40 million, but the numbers are not of great concern. What is more important is the issue of evolvability of the cellular automata (CA) based neural net circuits which grow and evolve in special FPGA (field programmable gate array) hardware, at hardware speeds, e.g. updating 150 billion CA cells per second, and performing a complete run of a genetic algorithm, i.e. tens of thousands of circuit growths and fitness evaluations, to evolve the elite neural net circuit in about 1 second. The specialized hardware which performs this evolution is labeled the CAM-Brain Machine (CBM). It implements the CoDi-1Bit model, and will be delivered to ATR probably in January 1999. The CBM should make practical the assemblage of 10000s of evolved neural net modules into humanly defined artificial brains. For the past few months, the latest hardware version of the CBM has been simulated in software to see just how evolvable and functional individual evolved modules can be. The paper reports on some of the results of these simulations. Hugo de Garis, Michael Korkin, Felix A. Gers, Michael Hough |
CEC | 3 |
| 1999 | ATR's Artificial Brain ("CAM-Brain") Project: A Sample of What Individual "CoDi-1Bit" Model Evolved Neural Net Modules Can Do with Digital and Analog I/O
Hugo de Garis, Andrzej Buller, Michael Korkin, Felix A. Gers, Norberto Eiji Nawa, Michael Hough |
GECCO | 4 |
| 1999 | Language identification from prosody without explicit featuresabstractMost current language identification (LID) systems make little or no use of prosodic information, despite the importance of prosody in LID by humans. The greatest obstacle has been that of finding an appropriate feature set which captures linguistically relevant prosodic information. The only system to attempt LID entirely on the basis of prosodic variables uses a set of over 200 features which are selected and combined in a task-specific manner [12]. We apply a novel recurrent neural network model to the task of pairwise discrimination among languages. Network inputs are limited to delta-F 0 and the first difference of the band limited amplitude envelope. Initial results are based on all pairwise combinations of English, German, Japanese, Mandarin and Spanish, with 90 speakers per language. Keywords: Language identification, Recurrent neural networks, prosody 1. PROSODY AND LANGUAGE IDENTIFICATION Most current approaches to automatic language identification use some form of segment re... Fred A. Cummins, Felix A. Gers, Jürgen Schmidhuber |
EUROSPEECH | 2 |