VLDB 2026 Research / reviewers in the wild / expert
George Zerveas
dblp:232/1820
· DBLP profile ↗
8ranked-venue papers
4as first author
7since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Parameter-efficient Modularised Bias Mitigation via AdapterFusionabstractDeepak Kumar, Oleg Lesota, George Zerveas, Daniel Cohen, Carsten Eickhoff, Markus Schedl, Navid Rekabsaz. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Deepak Kumar 0015, Oleg Lesota, George Zerveas, Carsten Eickhoff, Markus Schedl, Navid Rekabsaz |
EACL | 3 |
| 2023 | Enhancing the Ranking Context of Dense Retrieval through Reciprocal Nearest NeighborsabstractSparse annotation poses persistent challenges to training dense retrieval models, for example by distorting the training signal when unlabeled relevant documents are used spuriously as negatives in contrastive learning.To alleviate this problem, we introduce evidence-based label smoothing, a novel, computationally efficient method that prevents penalizing the model for assigning high relevance to false negatives.To compute the target relevance distribution over candidate documents within the ranking context of a given query, those candidates most similar to the ground truth are assigned a nonzero relevance probability based on the degree of their similarity to the ground-truth document(s).To estimate relevance we leverage an improved similarity metric based on reciprocal nearest neighbors, which can also be used independently to rerank candidates in postprocessing.Through extensive experiments on two large-scale ad hoc text retrieval datasets, we demonstrate that reciprocal nearest neighbors can improve the ranking effectiveness of dense retrieval models, both when used for label smoothing, as well as for reranking.This indicates that by considering relationships between documents and queries beyond simple geometric distance we can effectively enhance the ranking context. 1 1 Our code and other resources are available at: https://github.com/gzerveas/CODER 2 E.g., on average 1. George Zerveas, Navid Rekabsaz, Carsten Eickhoff |
EMNLP | 1 |
| 2022 | CODER: An efficient framework for improving retrieval through COntextual Document Embedding RerankingabstractContrastive learning has been the dominant approach to training dense retrieval models.In this work, we investigate the impact of ranking context -an often overlooked aspect of learning dense retrieval models.In particular, we examine the effect of its constituent parts: jointly scoring a large number of negatives per query, using retrieved (query-specific) instead of random negatives, and a fully list-wise loss.To incorporate these factors into training, we introduce Contextual Document Embedding Reranking (CODER), a highly efficient retrieval framework.When reranking, it incurs only a negligible computational overhead on top of a firststage method at run time (∼ 5 ms delay per query), allowing it to be easily combined with any state-of-the-art dual encoder method.Models trained through CODER can also be used as stand-alone retrievers.Evaluating CODER in a large set of experiments on the MS MARCO and TripClick collections, we show that the contextual reranking of precomputed document embeddings leads to a significant improvement in retrieval performance.This improvement becomes even more pronounced when more relevance information per query is available, shown in the TripClick collection, where we establish new state-of-the-art results by a large margin. George Zerveas, Navid Rekabsaz, Carsten Eickhoff |
EMNLP | 1 |
| 2022 | Unsupervised Multivariate Time-Series Transformers for Seizure Identification on EEGabstractEpilepsy is one of the most common neurological disorders, typically observed via seizure episodes. Epileptic seizures are commonly monitored through electroencephalogram (EEG) recordings due to their routine and low expense collection. The stochastic nature of EEG makes seizure identification via manual inspections performed by highly-trained experts a tedious endeavor, motivating the use of automated identification. The literature on automated identification focuses mostly on supervised learning methods requiring expert labels of EEG segments that contain seizures, which are difficult to obtain. Motivated by these observations, we pose seizure identification as an unsupervised anomaly detection problem. To this end, we employ the first unsupervised transformer-based model for seizure identification on raw EEG. We train an autoencoder involving a transformer encoder via an unsupervised loss function, incorporating a novel masking strategy uniquely designed for multivariate time-series data such as EEG. Training employs EEG recordings that do not contain any seizures, while seizures are identified with respect to reconstruction errors at inference time. We evaluate our method on three publicly available benchmark EEG datasets for distinguishing seizure vs. non-seizure windows. Our method leads to significantly better seizure identification performance than supervised learning counterparts, by up to 16% recall, 9% accuracy, and 9% Area under the Receiver Operating Characteristics Curve (AUC), establishing particular benefits on highly imbalanced data. Through accurate seizure identification, our method could facilitate widely accessible and early detection of epilepsy development, without needing expensive label collection or manual feature extraction. Ilkay Yildiz, George Zerveas, Carsten Eickhoff, Dominique Duncan |
ICMLA | 2 |
| 2022 | Mitigating Bias in Search Results Through Contextual Document Reranking and Neutrality RegularizationabstractSocietal biases can influence Information Retrieval system results, and conversely, search results can potentially reinforce existing societal biases. Recent research has therefore focused on developing methods for quantifying and mitigating bias in search results and applied them to contemporary retrieval systems that leverage transformer-based language models. In the present work, we expand this direction of research by considering bias mitigation within a framework for contextual document embedding reranking. In this framework, the transformer-based query encoder is optimized for relevance ranking through a list-wise objective, by jointly scoring for the same query a large set of candidate document embeddings in the context of one another, instead of in isolation. At the same time, we impose a regularization loss which penalizes highly scoring documents that deviate from neutrality with respect to a protected attribute (e.g., gender). Our approach for bias mitigation is end-to-end differentiable and efficient. Compared to the existing alternatives for deep neural retrieval architectures, which are based on adversarial training, we demonstrate that it can attain much stronger bias mitigation/fairness. At the same time, for the same amount of bias mitigation, it offers significantly better relevance performance (utility). Crucially, our method allows for a more finely controllable and predictable intensity of bias mitigation, which is essential for practical deployment in production systems. George Zerveas, Navid Rekabsaz, Carsten Eickhoff |
SIGIR | 1 |
| 2022 | CATS: Customizable Abstractive Topic-based SummarizationabstractNeural sequence-to-sequence models are the state-of-the-art approach used in abstractive summarization of textual documents, useful for producing condensed versions of source text narratives without being restricted to using only words from the original text. Despite the advances in abstractive summarization, custom generation of summaries (e.g., towards a user’s preference) remains unexplored. In this article, we present CATS, an abstractive neural summarization model that summarizes content in a sequence-to-sequence fashion while also introducing a new mechanism to control the underlying latent topic distribution of the produced summaries. We empirically illustrate the efficacy of our model in producing customized summaries and present findings that facilitate the design of such systems. We use the well-known CNN/DailyMail dataset to evaluate our model. Furthermore, we present a transfer-learning method and demonstrate the effectiveness of our approach in a low resource setting, i.e., abstractive summarization of meetings minutes, where combining the main available meetings’ transcripts datasets, AMI and International Computer Science Institute(ICSI) , results in merely a few hundred training documents. Seyed Ali Bahrainian, George Zerveas, Fabio Crestani, Carsten Eickhoff |
ACM Trans. Inf. Syst. | 2 |
| 2021 | A Transformer-based Framework for Multivariate Time Series Representation LearningabstractWe present a novel framework for multivariate time series representation learning based on the transformer encoder architecture. The framework includes an unsupervised pre-training scheme, which can offer substantial performance benefits over fully supervised learning on downstream tasks, both with but even without leveraging additional unlabeled data, i.e., by reusing the existing data samples. Evaluating our framework on several public multivariate time series datasets from various domains and with diverse characteristics, we demonstrate that it performs significantly better than the best currently available methods for regression and classification, even for datasets which consist of only a few hundred training samples. Given the pronounced interest in unsupervised learning for nearly all domains in the sciences and in industry, these findings represent an important landmark, presenting the first unsupervised method shown to push the limits of state-of-the-art performance for multivariate time series regression and classification. George Zerveas, Srideepika Jayaraman, Dhaval Patel 0002, Anuradha Bhamidipaty, Carsten Eickhoff |
KDD | 1 |
| 2020 | Extracting Angina Symptoms from Clinical Notes Using Pre-Trained Transformer Architectures
Aaron S. Eisman, Nishant R. Shah, Carsten Eickhoff, George Zerveas, Elizabeth S. Chen, Wen-Chih Wu, Indra Neil Sarkar |
AMIA | 4 |