VLDB 2026 Research / reviewers in the wild / expert
Nadav Borenstein
dblp:294/4730
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Information extraction and text analysis · 49% Language models and text generation · 32% Learning theory · 15% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% | |
| Theoretical computer science
2 papers |
Automata and formal languages · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational social science and digital humanities · 100% |
Topics — the 8 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
search engines |
0.8 | 1 | 2024 | Imitation of Life: A Search Engine for Biologically Inspired Design · AAAI 2024 |
Natural language and speech › Information extraction and text analysis
event extraction |
0.7 | 1 | 2023 | Multilingual Event Extraction from Historical Newspaper Adverts · ACL (1) 2023 |
Natural language and speech › Information extraction and text analysis › event extraction
multilingual event extraction |
0.7 | 1 | 2023 | Multilingual Event Extraction from Historical Newspaper Adverts · ACL (1) 2023 |
Natural language and speech › Language models and text generation › multimodal language model
pixel-based language models |
0.7 | 1 | 2023 | PHD: Pixel-Based Language Modeling of Historical Documents · EMNLP 2023 |
Natural language and speech › Information extraction and text analysis › document analysis
scholarly text analysis |
0.5 | 1 | 2021 | How Did This Get Funded?! Automatically Identifying Quirky Scientific Achievements · ACL/IJCNLP (1) 2021 |
Natural language and speech › Language models and text generation
natural language understanding |
0.2 | 1 | 2024 | Imitation of Life: A Search Engine for Biologically Inspired Design · AAAI 2024 |
Automata and formal languages
formal language representation |
0.2 | 1 | 2024 | Can Transformers Learn n-gram Language Models? · EMNLP 2024 |
Computer vision › Vision and language › multimodal understanding
multimodal document understanding |
0.2 | 1 | 2023 | PHD: Pixel-Based Language Modeling of Historical Documents · EMNLP 2023 |
Methods — techniques the papers use, named apart from their topics
natural language understanding · 1.5data programming · 1.5machine translation · 1.3extractive question answering · 1.3n-gram estimation · 0.8add-λ smoothing · 0.8synthetic scan generation · 0.7masked patch reconstruction · 0.7text classification · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Investigating Human Values in Online CommunitiesabstractNadav Borenstein, Arnav Arora, Lucie-Aimée Kaffee, Isabelle Augenstein. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Nadav Borenstein, Arnav Arora, Lucie-Aimée Kaffee, Isabelle Augenstein |
NAACL (Long Papers) | 1 |
| 2024 | Imitation of Life: A Search Engine for Biologically Inspired DesignabstractBiologically Inspired Design (BID), or Biomimicry, is a problem-solving methodology that applies analogies from nature to solve engineering challenges. For example, Speedo engineers designed swimsuits based on shark skin. Finding relevant biological solutions for real-world problems poses significant challenges, both due to the limited biological knowledge engineers and designers typically possess and to the limited BID resources. Existing BID datasets are hand-curated and small, and scaling them up requires costly human annotations. In this paper, we introduce BARcode (Biological Analogy Retriever), a search engine for automatically mining bio-inspirations from the web at scale. Using advances in natural language understanding and data programming, BARcode identifies potential inspirations for engineering challenges. Our experiments demonstrate that BARcode can retrieve inspirations that are valuable to engineers and designers tackling real-world problems, as well as recover famous historical BID examples. We release data and code; we view BARcode as a step towards addressing the challenges that have historically hindered the practical application of BID to engineering innovation. Hen Emuna, Nadav Borenstein, Hyeonsu B. Kang, Joel Chan, Aniket Kittur, Dafna Shahaf |
AAAI | 2 |
| 2024 | What Languages are Easy to Language-Model? A Perspective from Learning Probabilistic Regular LanguagesabstractNadav Borenstein, Anej Svete, Robin Chan, Josef Valvoda, Franz Nowak, Isabelle Augenstein, Eleanor Chodroff, Ryan Cotterell. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Nadav Borenstein, Anej Svete, Robin Shing Moon Chan, Josef Valvoda, Franz Nowak, Isabelle Augenstein, Eleanor Chodroff, Ryan Cotterell |
ACL (1) | 1 |
| 2024 | Can Transformers Learn n-gram Language Models?abstractMuch theoretical work has described the ability of transformers to represent formal languages.However, linking theoretical results to empirical performance is not straightforward due to the complex interplay between the architecture, the learning algorithm, and training data.To test whether theoretical lower bounds imply learnability of formal languages, we turn to recent work relating transformers to n-gram language models (LMs).We study transformers' ability to learn random n-gram LMs of two kinds: ones with arbitrary next-symbol probabilities and ones where those are defined with shared parameters.We find that classic estimation techniques for n-gram LMs such as add-λ smoothing outperform transformers on the former, while transformers perform better on the latter, outperforming methods specifically designed to learn n-gram LMs.github.com/rycolab/learning-ngrams Anej Svete, Nadav Borenstein, Mike Zhou, Isabelle Augenstein, Ryan Cotterell |
EMNLP | 2 |
| 2023 | Multilingual Event Extraction from Historical Newspaper AdvertsabstractNLP methods can aid historians in analyzing textual materials in greater volumes than manually feasible.Developing such methods poses substantial challenges though.First, acquiring large, annotated historical datasets is difficult, as only domain experts can reliably label them.Second, most available off-the-shelf NLP models are trained on modern language texts, rendering them significantly less effective when applied to historical corpora.This is particularly problematic for less well studied tasks, and for languages other than English.This paper addresses these challenges while focusing on the under-explored task of event extraction from a novel domain of historical texts.We introduce a new multilingual dataset in English, French, and Dutch composed of newspaper ads from the early modern colonial period reporting on enslaved people who liberated themselves from enslavement.We find that: 1) even with scarce annotated data, it is possible to achieve surprisingly good results by formulating the problem as an extractive QA task and leveraging existing datasets and models for modern languages; and 2) cross-lingual low-resource learning for historical languages is highly challenging, and machine translation of the historical datasets to the considered target languages is, in practice, often the best-performing solution. Nadav Borenstein, Natalia da Silva Perez, Isabelle Augenstein |
ACL (1) | 1 |
| 2023 | PHD: Pixel-Based Language Modeling of Historical DocumentsabstractThe digitisation of historical documents has provided historians with unprecedented research opportunities. Yet, the conventional approach to analysing historical documents involves converting them from images to text using OCR, a process that overlooks the potential benefits of treating them as images and introduces high levels of noise. To bridge this gap, we take advantage of recent advancements in pixel-based language models trained to reconstruct masked patches of pixels instead of predicting token distributions. Due to the scarcity of real historical scans, we propose a novel method for generating synthetic scans to resemble real historical documents. We then pre-train our model, PHD, on a combination of synthetic scans and real historical newspapers from the 1700-1900 period. Through our experiments, we demonstrate that PHD exhibits high proficiency in reconstructing masked image patches and provide evidence of our model's noteworthy language understanding capabilities. Notably, we successfully apply our model to a historical QA task, highlighting its usefulness in this domain. Nadav Borenstein, Phillip Rust, Desmond Elliott, Isabelle Augenstein |
EMNLP | 1 |
| 2021 | How Did This Get Funded?! Automatically Identifying Quirky Scientific AchievementsabstractChen Shani, Nadav Borenstein, Dafna Shahaf. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Chen Shani, Nadav Borenstein, Dafna Shahaf |
ACL/IJCNLP (1) | 2 |