VLDB 2026 Research / reviewers in the wild / expert
Nils Reimers 0001
dblp:129/5173 · also Nils Fabian Reimers
· DBLP profile ↗
17ranked-venue papers
7as first author
8since 2021 · last 2024
0000-0002-8663-1131ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 7 first-author · 8 since 2021Security and privacy · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Representation and self-supervised learning · 32% Information extraction and text analysis · 32% Efficient and distributed learning · 18% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 86% Machine learning and data management · 14% |
Topics — the 23 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › text embedding
sentence embedding |
1.1 | 2 | 2024 | Triple-Encoders: Representations That Fire Together, Wire Together · ACL (1) 2024 Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks · EMNLP/IJCNLP (1) 2019 |
Information retrieval › retrieval models
neural retrieval |
0.8 | 1 | 2024 | DAPR: A Benchmark on Document-Aware Passage Retrieval · ACL (1) 2024 |
Information retrieval › document retrieval
passage retrieval |
0.8 | 1 | 2024 | DAPR: A Benchmark on Document-Aware Passage Retrieval · ACL (1) 2024 |
Information retrieval › reranking
document re-ranking |
0.6 | 1 | 2022 | Incorporating Relevance Feedback for Information-Seeking Retrieval using Few-Shot Document Re-Ranking · EMNLP 2022 |
Machine learning and data management › transfer learning
few-shot learning |
0.6 | 1 | 2022 | Incorporating Relevance Feedback for Information-Seeking Retrieval using Few-Shot Document Re-Ranking · EMNLP 2022 |
Information retrieval › reranking
neural re-ranking |
0.6 | 1 | 2022 | Incorporating Relevance Feedback for Information-Seeking Retrieval using Few-Shot Document Re-Ranking · EMNLP 2022 |
Information retrieval
relevance feedback |
0.6 | 1 | 2022 | Incorporating Relevance Feedback for Information-Seeking Retrieval using Few-Shot Document Re-Ranking · EMNLP 2022 |
Machine learning › Efficient and distributed learning
model compression |
0.5 | 1 | 2021 | AdapterDrop: On the Efficiency of Adapters in Transformers · EMNLP (1) 2021 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
0.5 | 1 | 2021 | AdapterDrop: On the Efficiency of Adapters in Transformers · EMNLP (1) 2021 |
Machine learning › Representation and self-supervised learning › word representation
multilingual word embedding |
0.4 | 1 | 2020 | Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation · EMNLP (1) 2020 |
Natural language and speech › Information extraction and text analysis › argument mining
argument classification |
0.4 | 1 | 2019 | Classification and Clustering of Arguments with Contextualized Word Embeddings · ACL (1) 2019 |
Natural language and speech › Information extraction and text analysis
argument mining |
0.4 | 1 | 2019 | Classification and Clustering of Arguments with Contextualized Word Embeddings · ACL (1) 2019 |
Natural language and speech › Information extraction and text analysis
coreference resolution |
0.4 | 1 | 2019 | Revisiting Joint Modeling of Cross-document Entity and Event Coreference Resolution · ACL (1) 2019 |
Machine learning › Deep learning architectures and training
siamese network |
0.4 | 1 | 2019 | Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks · EMNLP/IJCNLP (1) 2019 |
Machine learning › Deep learning architectures and training › recurrent neural network
LSTM |
0.3 | 1 | 2017 | Reporting Score Distributions Makes a Difference: Performance Study of LSTM-networks for Sequence Tagging · EMNLP 2017 |
Natural language and speech › Information extraction and text analysis
sequence labeling |
0.3 | 1 | 2017 | Reporting Score Distributions Makes a Difference: Performance Study of LSTM-networks for Sequence Tagging · EMNLP 2017 |
Natural language and speech › Information extraction and text analysis
temporal information extraction |
0.2 | 1 | 2016 | Temporal Anchoring of Events for the TimeBank Corpus · ACL (1) 2016 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.2 | 1 | 2024 | Triple-Encoders: Representations That Fire Together, Wire Together · ACL (1) 2024 |
Natural language and speech › Language models and text generation
multi-task inference |
0.1 | 1 | 2021 | AdapterDrop: On the Efficiency of Adapters in Transformers · EMNLP (1) 2021 |
Machine learning › Deep learning architectures and training
transformer |
0.1 | 1 | 2021 | AdapterDrop: On the Efficiency of Adapters in Transformers · EMNLP (1) 2021 |
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer |
0.1 | 1 | 2020 | Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation · EMNLP (1) 2020 |
Natural language and speech › Information extraction and text analysis › coreference resolution
event coreference resolution |
0.1 | 1 | 2019 | Revisiting Joint Modeling of Cross-document Entity and Event Coreference Resolution · ACL (1) 2019 |
Performance modeling and evaluation
benchmarking |
0.1 | 1 | 2017 | Reporting Score Distributions Makes a Difference: Performance Study of LSTM-networks for Sequence Tagging · EMNLP 2017 |
Methods — techniques the papers use, named apart from their topics
contextualized passage representation · 0.8BM25 · 0.8BERT · 0.8meta-learning · 0.6kNN · 0.6fine-tuning · 0.6cross-encoder · 0.6pruning · 0.5adapter fusion · 0.5adapter · 0.5machine translation · 0.4knowledge distillation · 0.4predicate-argument structure · 0.4neural architecture · 0.4contextualized word embeddings · 0.4ELMo · 0.4score distribution analysis · 0.3hyperparameter evaluation · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Triple-Encoders: Representations That Fire Together, Wire TogetherabstractJustus-Jonas Erker, Florian Mai, Nils Reimers, Gerasimos Spanakis, Iryna Gurevych. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Justus-Jonas Erker, Florian Mai, Nils Reimers 0001, Gerasimos Spanakis, Iryna Gurevych |
ACL (1) | 3 |
| 2024 | DAPR: A Benchmark on Document-Aware Passage RetrievalabstractThe work of neural retrieval so far focuses on ranking short texts and is challenged with long documents.There are many cases where the users want to find a relevant passage within a long document from a huge corpus, e.g.Wikipedia articles, research papers, etc.We propose and name this task Document-Aware Passage Retrieval (DAPR).While analyzing the errors of the State-of-The-Art (SoTA) passage retrievers, we find the major errors (53.5%) are due to missing document context.This drives us to build a benchmark for this task including multiple datasets from heterogeneous domains.In the experiments, we extend the SoTA passage retrievers with document context via (1) hybrid retrieval with BM25 and (2) contextualized passage representations, which inform the passage representation with document context.We find despite that hybrid retrieval performs the strongest on the mixture of the easy and the hard queries, it completely fails on the hard queries that require document-context understanding.On the other hand, contextualized passage representations (e.g.prepending document titles) achieve good improvement on these hard queries, but overall they also perform rather poorly.Our created benchmark enables future research on developing and comparing retrieval systems for the new task.The code and the data are available 1 . Nils Reimers 0001, Iryna Gurevych |
ACL (1) | 2 |
| 2022 | Incorporating Relevance Feedback for Information-Seeking Retrieval using Few-Shot Document Re-RankingabstractPairing a lexical retriever with a neural reranking model has set state-of-the-art performance on large-scale information retrieval datasets.This pipeline covers scenarios like question answering or navigational queries, however, for information-seeking scenarios, users often provide information on whether a document is relevant to their query in form of clicks or explicit feedback.Therefore, in this work, we explore how relevance feedback can be directly integrated into neural re-ranking models by adopting few-shot and parameterefficient learning techniques.Specifically, we introduce a kNN approach that re-ranks documents based on their similarity with the query and the documents the user considers relevant.Further, we explore Cross-Encoder models that we pre-train using meta-learning and subsequently fine-tune for each query, training only on the feedback documents.To evaluate our different integration strategies, we transform four existing information retrieval datasets into the relevance feedback scenario.Extensive experiments demonstrate that integrating relevance feedback directly in neural re-ranking models improves their performance, and fusing lexical ranking with our best performing neural reranker outperforms all other methods by 5.2% nDCG@20. Tim Baumgärtner, Leonardo F. R. Ribeiro, Nils Reimers 0001, Iryna Gurevych |
EMNLP | 3 |
| 2022 | GPL: Generative Pseudo Labeling for Unsupervised Domain Adaptation of Dense RetrievalabstractKexin Wang, Nandan Thakur, Nils Reimers, Iryna Gurevych. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Nandan Thakur, Nils Reimers 0001, Iryna Gurevych |
NAACL-HLT | 3 |
| 2022 | Retrieve Fast, Rerank Smart: Cooperative and Joint Approaches for Improved Cross-Modal RetrievalabstractAbstract Current state-of-the-art approaches to cross- modal retrieval process text and visual input jointly, relying on Transformer-based architectures with cross-attention mechanisms that attend over all words and objects in an image. While offering unmatched retrieval performance, such models: 1) are typically pretrained from scratch and thus less scalable, 2) suffer from huge retrieval latency and inefficiency issues, which makes them impractical in realistic applications. To address these crucial gaps towards both improved and efficient cross- modal retrieval, we propose a novel fine-tuning framework that turns any pretrained text-image multi-modal model into an efficient retrieval model. The framework is based on a cooperative retrieve-and-rerank approach that combines: 1) twin networks (i.e., a bi-encoder) to separately encode all items of a corpus, enabling efficient initial retrieval, and 2) a cross-encoder component for a more nuanced (i.e., smarter) ranking of the retrieved small set of items. We also propose to jointly fine- tune the two components with shared weights, yielding a more parameter-efficient model. Our experiments on a series of standard cross-modal retrieval benchmarks in monolingual, multilingual, and zero-shot setups, demonstrate improved accuracy and huge efficiency benefits over the state-of-the-art cross- encoders.1 Gregor Geigle, Jonas Pfeiffer, Nils Reimers 0001, Ivan Vulic, Iryna Gurevych |
Trans. Assoc. Comput. Linguistics | 3 |
| 2021 | AdapterDrop: On the Efficiency of Adapters in TransformersabstractTransformer models are expensive to fine-tune, slow for inference, and have large storage requirements.Recent approaches tackle these shortcomings by training smaller models, dynamically reducing the model size, and by training light-weight adapters.In this paper, we propose AdapterDrop, removing adapters from lower transformer layers during training and inference, which incorporates concepts from all three directions.We show that Adap-terDrop can dynamically reduce the computational overhead when performing inference over multiple tasks simultaneously, with minimal decrease in task performances.We further prune adapters from AdapterFusion, which improves the inference efficiency while maintaining the task performances entirely. Andreas Rücklé, Gregor Geigle, Max Glockner, Tilman Beck, Jonas Pfeiffer, Nils Reimers 0001, Iryna Gurevych |
EMNLP (1) | 6 |
| 2021 | Augmented SBERT: Data Augmentation Method for Improving Bi-Encoders for Pairwise Sentence Scoring TasksabstractNandan Thakur, Nils Reimers, Johannes Daxenberger, Iryna Gurevych. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Nandan Thakur, Nils Reimers 0001, Johannes Daxenberger, Iryna Gurevych |
NAACL-HLT | 2 |
| 2021 | Generalizing Cross-Document Event Coreference Resolution Across Multiple CorporaabstractCross-document event coreference resolution (CDCR) is an NLP task in which mentions of events need to be identified and clustered throughout a collection of documents. CDCR aims to benefit downstream multidocument applications, but despite recent progress on corpora and system development, downstream improvements from applying CDCR have not been shown yet. We make the observation that every CDCR system to date was developed, trained, and tested only on a single respective corpus. This raises strong concerns on their generalizability—a must-have for downstream applications where the magnitude of domains or event mentions is likely to exceed those found in a curated corpus. To investigate this assumption, we define a uniform evaluation setup involving three CDCR corpora: ECB+, the Gun Violence Corpus, and the Football Coreference Corpus (which we reannotate on token level to make our analysis possible). We compare a corpus-independent, feature-based system against a recent neural system developed for ECB+. Although being inferior in absolute numbers, the feature-based system shows more consistent performance across all corpora whereas the neural system is hit-or-miss. Via model introspection, we find that the importance of event actions, event time, and so forth, for resolving coreference in practice varies greatly between the corpora. Additional analysis shows that several systems overfit on the structure of the ECB+ corpus. We conclude with recommendations on how to achieve generally applicable CDCR systems in the future—the most important being that evaluation on multiple CDCR corpora is strongly necessary. To facilitate future research, we release our dataset, annotation guidelines, and system implementation to the public.1 Michael Bugert, Nils Reimers 0001, Iryna Gurevych |
Comput. Linguistics | 2 |
| 2020 | Making Monolingual Sentence Embeddings Multilingual using Knowledge DistillationabstractWe present an easy and efficient method to extend existing sentence embedding models to new languages.This allows to create multilingual versions from previously monolingual models.The training is based on the idea that a translated sentence should be mapped to the same location in the vector space as the original sentence.We use the original (monolingual) model to generate sentence embeddings for the source language and then train a new system on translated sentences to mimic the original model.Compared to other methods for training multilingual sentence embeddings, this approach has several advantages: It is easy to extend existing models with relatively few samples to new languages, it is easier to ensure desired properties for the vector space, and the hardware requirements for training are lower.We demonstrate the effectiveness of our approach for 50+ languages from various language families.Code to extend sentence embeddings models to more than 400 languages is publicly available.1 Nils Reimers 0001, Iryna Gurevych |
EMNLP (1) | 1 |
| 2019 | Revisiting Joint Modeling of Cross-document Entity and Event Coreference ResolutionabstractRecognizing coreferring events and entities across multiple texts is crucial for many NLP applications.Despite the task's importance, research focus was given mostly to withindocument entity coreference, with rather little attention to the other variants.We propose a neural architecture for cross-document coreference resolution.Inspired by Lee et al. (2012), we jointly model entity and event coreference.We represent an event (entity) mention using its lexical span, surrounding context, and relation to entity (event) mentions via predicate-arguments structures.Our model outperforms the previous state-of-the-art event coreference model on ECB+, while providing the first entity coreference results on this corpus.Our analysis confirms that all our representation elements, including the mention span itself, its context, and the relation to other mentions contribute to the model's success. Shany Barhom, Vered Shwartz, Alon Eirew, Michael Bugert, Nils Reimers 0001, Ido Dagan |
ACL (1) | 5 |
| 2019 | Classification and Clustering of Arguments with Contextualized Word EmbeddingsabstractWe experiment with two recent contextualized word embedding methods (ELMo and BERT) in the context of open-domain argument search.For the first time, we show how to leverage the power of contextualized word embeddings to classify and cluster topic-dependent arguments, achieving impressive results on both tasks and across multiple datasets.For argument classification, we improve the state-of-the-art for the UKP Sentential Argument Mining Corpus by 20.8 percentage points and for the IBM Debater -Evidence Sentences dataset by 7.4 percentage points.For the understudied task of argument clustering, we propose a pre-training step which improves by 7.8 percentage points over strong baselines on a novel dataset, and by 12.3 percentage points for the Argument Facet Similarity (AFS) Corpus. 1 Nils Reimers 0001, Benjamin Schiller, Tilman Beck, Johannes Daxenberger, Christian Stab, Iryna Gurevych |
ACL (1) | 1 |
| 2019 | Sentence-BERT: Sentence Embeddings using Siamese BERT-NetworksabstractNils Reimers, Iryna Gurevych. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Nils Reimers 0001, Iryna Gurevych |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Event Time Extraction with a Decision Tree of Neural ClassifiersabstractExtracting the information from text when an event happened is challenging. Documents do not only report on current events, but also on past events as well as on future events. Often, the relevant time information for an event is scattered across the document. In this paper we present a novel method to automatically anchor events in time. To our knowledge it is the first approach that takes temporal information from the complete document into account. We created a decision tree that applies neural network based classifiers at its nodes. We use this tree to incrementally infer, in a stepwise manner, at which time frame an event happened. We evaluate the approach on the TimeBank-EventTime Corpus (Reimers et al., 2016) achieving an accuracy of 42.0% compared to an inter-annotator agreement (IAA) of 56.7%. For events that span over a single day we observe an accuracy improvement of 33.1 points compared to the state-of-the-art CAEVO system (Chambers et al., 2014). Without retraining, we apply this model to the SemEval-2015 Task 4 on automatic timeline generation and achieve an improvement of 4.01 points F1-score compared to the state-of-the-art. Our code is publically available. Nils Reimers 0001, Nazanin Dehghani, Iryna Gurevych |
Trans. Assoc. Comput. Linguistics | 1 |
| 2017 | Reporting Score Distributions Makes a Difference: Performance Study of LSTM-networks for Sequence TaggingabstractIn this paper we show that reporting a single performance score is insufficient to compare non-deterministic approaches.We demonstrate for common sequence tagging tasks that the seed value for the random number generator can result in statistically significant (p < 10 -4 ) differences for state-of-the-art systems.For two recent systems for NER, we observe an absolute difference of one percentage point F 1 -score depending on the selected seed value, making these systems perceived either as state-of-the-art or mediocre.Instead of publishing and reporting single performance scores, we propose to compare score distributions based on multiple executions.Based on the evaluation of 50.000LSTMnetworks for five sequence tagging tasks, we present network architectures that produce both superior performance as well as are more stable with respect to the remaining hyperparameters.The full experimental results are published in (Reimers and Gurevych, 2017). 1 The implementation of our network is publicly available.2 Nils Reimers 0001, Iryna Gurevych |
EMNLP | 1 |
| 2016 | Temporal Anchoring of Events for the TimeBank CorpusabstractToday's extraction of temporal information for events heavily depends on annotated temporal links.These so called TLINKs capture the relation between pairs of event mentions and time expressions.One problem is that the number of possible TLINKs grows quadratic with the number of event mentions, therefore most annotation studies concentrate on links for mentions in the same or in adjacent sentences.However, as our annotation study shows, this restriction results for 58% of the event mentions in a less precise information when the event took place.This paper proposes a new annotation scheme to anchor events in time.Not only is the annotation effort much lower as it scales linear with the number of events, it also gives a more precise anchoring when the events have happened as the complete document can be taken into account.Using this scheme, we annotated a subset of the TimeBank Corpus and compare our results to other annotation schemes.Additionally, we present some baseline experiments to automatically anchor events in time.Our annotation scheme, the automated system and the annotated corpus are publicly available. Nils Reimers 0001, Nazanin Dehghani, Iryna Gurevych |
ACL (1) | 1 |
| 2016 | Task-Oriented Intrinsic Evaluation of Semantic Textual SimilarityabstractSemantic Textual Similarity (STS) is a foundational NLP task and can be used in a wide range of tasks. To determine the STS of two texts, hundreds of different STS systems exist, however, for an NLP system designer, it is hard to decide which system is the best one. To answer this question, an intrinsic evaluation of the STS systems is conducted by comparing the output of the system to human judgments on semantic similarity. The comparison is usually done using Pearson correlation. In this work, we show that relying on intrinsic evaluations with Pearson correlation can be misleading. In three common STS based tasks we could observe that the Pearson correlation was especially ill-suited to detect the best STS system for the task and other evaluation measures were much better suited. In this work we define how the validity of an intrinsic evaluation can be assessed and compare different intrinsic evaluation methods. Understanding of the properties of the targeted task is crucial and we propose a framework for conducting the intrinsic evaluation which takes the properties of the targeted task into account. Nils Reimers 0001, Philip Beyer, Iryna Gurevych |
COLING | 1 |
| 2013 | Computing on Authenticated Data for Adjustable Predicates
Björn Deiseroth, Victoria Fehr, Marc Fischlin, Manuel Maasz, Nils Reimers 0001, Richard Stein |
ACNS | 5 |