EDBT 2026 Demo / reviewers in the wild / expert
Arnold Overwijk
dblp:16/7404
· DBLP profile ↗
11ranked-venue papers
2as first author
9since 2021 · last 2025
0009-0000-5915-2541ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
7 papers |
Information retrieval · 81% Recommender systems · 19% | |
| Artificial intelligence
5 papers |
Language models and text generation · 25% Efficient and distributed learning · 25% Trustworthy machine learning · 21% |
Topics — the 15 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › retrieval models › neural retrieval
dense retrieval |
2.8 | 5 | 2023 | Augmenting Zero-Shot Dense Retrievers with Plug-in Mixture-of-Memories · EMNLP 2023 COCO-DR: Combating the Distribution Shift in Zero-Shot Dense Retrieval with Contrastive and Distributionally Robust Learning · EMNLP 2022 Reduce Catastrophic Forgetting of Dense Retrieval Training with Teleportation Negatives · EMNLP 2022 |
Information retrieval › document retrieval
zero-shot retrieval |
1.2 | 2 | 2023 | Augmenting Zero-Shot Dense Retrievers with Plug-in Mixture-of-Memories · EMNLP 2023 COCO-DR: Combating the Distribution Shift in Zero-Shot Dense Retrieval with Contrastive and Distributionally Robust Learning · EMNLP 2022 |
Natural language and speech › Language models and text generation › large language model training
pretraining data selection |
0.9 | 1 | 2025 | Group-Level Data Selection for Efficient Pretraining · NeurIPS 2025 |
Information retrieval › evaluation
benchmark |
0.9 | 1 | 2025 | ORBIT - Open Recommendation Benchmark for Reproducible Research with Hidden Tests · NeurIPS 2025 |
Recommender systems
recommender system evaluation |
0.9 | 1 | 2025 | ORBIT - Open Recommendation Benchmark for Reproducible Research with Hidden Tests · NeurIPS 2025 |
Recommender systems › content recommendation
web page recommendation |
0.9 | 1 | 2025 | ORBIT - Open Recommendation Benchmark for Reproducible Research with Hidden Tests · NeurIPS 2025 |
Information retrieval
retrieval augmentation |
0.7 | 1 | 2023 | Augmenting Zero-Shot Dense Retrievers with Plug-in Mixture-of-Memories · EMNLP 2023 |
Machine learning › Trustworthy machine learning › robustness
distributionally robust optimization |
0.6 | 1 | 2022 | COCO-DR: Combating the Distribution Shift in Zero-Shot Dense Retrieval with Contrastive and Distributionally Robust Learning · EMNLP 2022 |
Information retrieval › retrieval models › neural retrieval › dense retrieval
dense retriever training |
0.6 | 1 | 2022 | Reduce Catastrophic Forgetting of Dense Retrieval Training with Teleportation Negatives · EMNLP 2022 |
Information retrieval
retrieval models |
0.6 | 1 | 2022 | Reduce Catastrophic Forgetting of Dense Retrieval Training with Teleportation Negatives · EMNLP 2022 |
Information retrieval
web corpus |
0.6 | 1 | 2022 | ClueWeb22: 10 Billion Web Documents with Rich Information · SIGIR 2022 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.5 | 1 | 2021 | Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval · ICLR 2021 |
Natural language and speech › Information extraction and text analysis
keyphrase extraction |
0.4 | 1 | 2019 | Open Domain Web Keyphrase Extraction Beyond Language Modeling · EMNLP/IJCNLP (1) 2019 |
Machine learning › Trustworthy machine learning › robustness
distribution shift |
0.2 | 1 | 2022 | COCO-DR: Combating the Distribution Shift in Zero-Shot Dense Retrieval with Contrastive and Distributionally Robust Learning · EMNLP 2022 |
Algorithms and data structures › similarity search › nearest neighbor search
approximate nearest neighbor search |
0.1 | 1 | 2021 | Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval · ICLR 2021 |
Methods — techniques the papers use, named apart from their topics
contrastive learning · 3.6approximate nearest neighbor search · 1.5distributionally robust optimization · 1.1weak decoder · 1.0prompted LLM baseline · 0.9data influence model · 0.9clustering · 0.9t5 · 0.7joint learning · 0.7hard negative mining · 0.7negative sampling · 0.6momentum contrastive learning · 0.6language modeling · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ORBIT - Open Recommendation Benchmark for Reproducible Research with Hidden TestsabstractRecommender systems are among the most impactful AI applications, interacting with billions of users every day, guiding them to relevant products, services, or information tailored to their preferences.However, the research and development of recommender systems are hindered by existing datasets that fail to capture realistic user behaviors and inconsistent evaluation settings that lead to ambiguous conclusions.This paper introduces the \textbf{O}pen \textbf{R}ecommendation \textbf{B}enchmark for Reproducible Research with H\textbf{I}dden \textbf{T}ests (\textbf{ORBIT}), a unified benchmark for consistent and realistic evaluation of recommendation models. ORBIT offers a standardized evaluation framework of public datasets with reproducible splits and transparent settings for its public leaderboard. Additionally, ORBIT introduces a new webpage recommendation task, ClueWeb-Reco, featuring web browsing sequences from 87 million public, high-quality webpages. ClueWeb-Reco is a synthetic dataset derived from real, user-consented, and privacy-guaranteed browsing data. It aligns with modern recommendation scenarios and is reserved as the hidden test part of our leaderboard to challenge recommendation models' generalization ability. ORBIT measures 12 representative recommendation models on its public benchmark and introduces a prompted LLM baseline on the ClueWeb-Reco hidden test.Our benchmark results reflect general improvements of recommender systems on the public datasets, with variable individual performances.The results on the hidden test reveal the limitations of existing approaches in large-scale webpage recommendation and highlight the potential for improvements with LLM integrations.ORBIT benchmark, leaderboard, and codebase are available at \url{https://www.open-reco-bench.ai}. Jingyuan He, Jiongnan Liu 0001, Vishan Vishesh Oberoi, Bolin Wu, Mahima Jagadeesh Patel, Kangrui Mao, Chuning Shi, I-Ta Lee, Arnold Overwijk, Chenyan Xiong |
NeurIPS | 9 |
| 2025 | Group-Level Data Selection for Efficient PretrainingabstractThe efficiency and quality of language model pretraining are largely determined by the way pretraining data are selected. In this paper, we introduce *Group-MATES*, an efficient group-level data selection approach to optimize the speed-quality frontier of language model pretraining. Specifically, Group-MATES parameterizes costly group-level selection with a relational data influence model. To train this model, we sample training trajectories of the language model and collect oracle data influences alongside. The relational data influence model approximates the oracle data influence by weighting individual influence with relationships among training data. To enable efficient selection with our relational data influence model, we partition the dataset into small clusters using relationship weights and select data within each cluster independently. Experiments on DCLM 400M-4x, 1B-1x, and 3B-1x show that Group-MATES achieves 3.5\%-9.4\% relative performance gains over random selection across 22 downstream tasks, nearly doubling the improvements achieved by state-of-the-art individual data selection baselines. Furthermore, Group-MATES reduces the number of tokens required to reach a certain downstream performance by up to 1.75x, substantially elevating the speed-quality frontier. Further analyses highlight the critical role of relationship weights in the relational data influence model and the effectiveness of our cluster-based inference. Our code is open-sourced at https://github.com/facebookresearch/Group-MATES. Zichun Yu, Arnold Overwijk, Scott Yih, Chenyan Xiong |
NeurIPS | 4 |
| 2023 | Augmenting Zero-Shot Dense Retrievers with Plug-in Mixture-of-MemoriesabstractIn this paper we improve the zero-shot generalization ability of language models via Mixture-Of-Memory Augmentation (MoMA), a mechanism that retrieves augmentation documents from multiple information corpora ("external memories"), with the option to "plug in" unseen memory at inference time.We develop a joint learning mechanism that trains the augmentation component with latent labels derived from the end retrieval task, paired with hard negatives from the memory mixture.We instantiate the model in a zero-shot dense retrieval setting by augmenting strong T5-based retrievers with MoMA.With only T5-base, our model obtains strong zero-shot retrieval accuracy on the eighteen tasks included in the standard BEIR benchmark, outperforming some systems with larger model sizes.As a plug-inplay model, our model can efficiently generalize to any unseen corpus, meanwhile achieving comparable or even better performance than methods relying on target-specific pretraining.Our analysis further illustrates the necessity of augmenting with mixture-of-memory for robust generalization, the benefits of augmentation learning, and how MoMA utilizes the plugin memory at inference time without changing its parameters. Suyu Ge, Chenyan Xiong, Corby Rosset, Arnold Overwijk, Jiawei Han 0001, Paul N. Bennett |
EMNLP | 4 |
| 2023 | Improving Multitask Retrieval by Promoting Task SpecializationabstractAbstract In multitask retrieval, a single retriever is trained to retrieve relevant contexts for multiple tasks. Despite its practical appeal, naive multitask retrieval lags behind task-specific retrieval, in which a separate retriever is trained for each task. We show that it is possible to train a multitask retriever that outperforms task-specific retrievers by promoting task specialization. The main ingredients are: (1) a better choice of pretrained model—one that is explicitly optimized for multitasking—along with compatible prompting, and (2) a novel adaptive learning method that encourages each parameter to specialize in a particular task. The resulting multitask retriever is highly performant on the KILT benchmark. Upon analysis, we find that the model indeed learns parameters that are more task-specialized compared to naive multitasking without prompting or adaptive learning.1 Wenzheng Zhang 0003, Chenyan Xiong, Karl Stratos, Arnold Overwijk |
Trans. Assoc. Comput. Linguistics | 4 |
| 2022 | Reduce Catastrophic Forgetting of Dense Retrieval Training with Teleportation NegativesabstractIn this paper, we investigate the instability in the standard dense retrieval training, which iterates between model training and hard negative selection using the being-trained model.We show the catastrophic forgetting phenomena behind the training instability, where models learn and forget different negative groups during training iterations.We then propose ANCE-Tele, which accumulates momentum negatives from past iterations and approximates future iterations using lookahead negatives, as "teleportations" along the time axis to smooth the learning process.On web search and OpenQA, ANCE-Tele outperforms previous state-of-theart systems of similar size, eliminates the dependency on sparse retrieval negatives, and is competitive among systems using significantly more (50x) parameters.Our analysis demonstrates that teleportation negatives reduce catastrophic forgetting and improve convergence speed for dense retrieval training. Si Sun, Chenyan Xiong, Arnold Overwijk, Zhiyuan Liu 0001 |
EMNLP | 4 |
| 2022 | COCO-DR: Combating the Distribution Shift in Zero-Shot Dense Retrieval with Contrastive and Distributionally Robust LearningabstractWe present a new zero-shot dense retrieval (Ze-roDR) method, COCO-DR, to improve the generalization ability of dense retrieval by combating the distribution shifts between source training tasks and target scenarios.To mitigate the impact of document differences, COCO-DR continues pretraining the language model on the target corpora to adapt the model to target distributions via COtinuous COtrastive learning.To prepare for unseen target queries, COCO-DR leverages implicit Distributionally Robust Optimization (iDRO) to reweight samples from different source query clusters for improving model robustness over rare queries during fine-tuning.COCO-DR achieves superior average performance on BEIR, the zero-shot retrieval benchmark.At BERT Base scale, COCO-DR Base outperforms other ZeroDR models with 60× larger size.At BERT Large scale, COCO-DR Large outperforms the giant GPT-3 embedding model which has 500× more parameters.Our analysis show the correlation of COCO-DR's effectiveness in combating distribution shifts and improving zero-shot accuracy.Our code and model can be found at https://github.com/OpenMatch/COCO-DR. Yue Yu 0001, Chenyan Xiong, Si Sun, Chao Zhang 0014, Arnold Overwijk |
EMNLP | 5 |
| 2022 | ClueWeb22: 10 Billion Web Documents with Rich InformationabstractClueWeb22, the newest iteration of the ClueWeb line of datasets, is the result of more than a year of collaboration between industry and academia. Its design is influenced by the research needs of the academic community and the real-world needs of large-scale industry systems. Compared with earlier ClueWeb datasets, the ClueWeb22 corpus is larger, more varied, and has higher-quality documents. Its core is raw HTML, but it includes clean text versions of documents to lower the barrier to entry. Several aspects of ClueWeb22 are available to the research community for the first time at this scale, for example, visual representations of rendered web pages, parsed structured information from the HTML document, and the alignment of document distributions (domains, languages, and topics) to commercial web search. Arnold Overwijk, Chenyan Xiong, Jamie Callan |
SIGIR | 1 |
| 2021 | Less is More: Pretrain a Strong Siamese Encoder for Dense Text Retrieval Using a Weak DecoderabstractShuqi Lu, Di He, Chenyan Xiong, Guolin Ke, Waleed Malik, Zhicheng Dou, Paul Bennett, Tie-Yan Liu, Arnold Overwijk. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Shuqi Lu, Di He 0001, Chenyan Xiong, Guolin Ke, Waleed Malik, Zhicheng Dou, Paul N. Bennett, Tie-Yan Liu, Arnold Overwijk |
EMNLP (1) | 9 |
| 2021 | Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval
Lee Xiong, Chenyan Xiong, Kwok-Fung Tang, Paul N. Bennett, Junaid Ahmed, Arnold Overwijk |
ICLR | 8 |
| 2019 | Open Domain Web Keyphrase Extraction Beyond Language ModelingabstractLee Xiong, Chuan Hu, Chenyan Xiong, Daniel Campos, Arnold Overwijk. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Lee Xiong, Chenyan Xiong, Daniel Campos, Arnold Overwijk |
EMNLP/IJCNLP (1) | 5 |
| 2011 | A Local Search Algorithm for Branchwidth
Arnold Overwijk, Eelko Penninkx, Hans L. Bodlaender |
SOFSEM | 1 |