Nadav Borenstein

dblp:294/4730 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Information extraction and text analysis · 49% Language models and text generation · 32% Learning theory · 15%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Theoretical computer science
2 papers
Automata and formal languages · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational social science and digital humanities · 100%

Topics — the 8 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
search engines
0.812024
Imitation of Life: A Search Engine for Biologically Inspired Design · AAAI 2024
Natural language and speech › Information extraction and text analysis
event extraction
0.712023
Multilingual Event Extraction from Historical Newspaper Adverts · ACL (1) 2023
Natural language and speech › Information extraction and text analysis › event extraction
multilingual event extraction
0.712023
Multilingual Event Extraction from Historical Newspaper Adverts · ACL (1) 2023
Natural language and speech › Language models and text generation › multimodal language model
pixel-based language models
0.712023
PHD: Pixel-Based Language Modeling of Historical Documents · EMNLP 2023
Natural language and speech › Information extraction and text analysis › document analysis
scholarly text analysis
0.512021
How Did This Get Funded?! Automatically Identifying Quirky Scientific Achievements · ACL/IJCNLP (1) 2021
Natural language and speech › Language models and text generation
natural language understanding
0.212024
Imitation of Life: A Search Engine for Biologically Inspired Design · AAAI 2024
Automata and formal languages
formal language representation
0.212024
Can Transformers Learn n-gram Language Models? · EMNLP 2024
Computer vision › Vision and language › multimodal understanding
multimodal document understanding
0.212023
PHD: Pixel-Based Language Modeling of Historical Documents · EMNLP 2023

Methods — techniques the papers use, named apart from their topics

natural language understanding · 1.5data programming · 1.5machine translation · 1.3extractive question answering · 1.3n-gram estimation · 0.8add-λ smoothing · 0.8synthetic scan generation · 0.7masked patch reconstruction · 0.7text classification · 0.5
YearPublicationVenuePosition
2025 Investigating Human Values in Online Communities
abstract
Nadav Borenstein, Arnav Arora, Lucie-Aimée Kaffee, Isabelle Augenstein. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Nadav Borenstein, Arnav Arora, Lucie-Aimée Kaffee, Isabelle Augenstein
NAACL (Long Papers)1
2024 Imitation of Life: A Search Engine for Biologically Inspired Design
abstract
Biologically Inspired Design (BID), or Biomimicry, is a problem-solving methodology that applies analogies from nature to solve engineering challenges. For example, Speedo engineers designed swimsuits based on shark skin. Finding relevant biological solutions for real-world problems poses significant challenges, both due to the limited biological knowledge engineers and designers typically possess and to the limited BID resources. Existing BID datasets are hand-curated and small, and scaling them up requires costly human annotations. In this paper, we introduce BARcode (Biological Analogy Retriever), a search engine for automatically mining bio-inspirations from the web at scale. Using advances in natural language understanding and data programming, BARcode identifies potential inspirations for engineering challenges. Our experiments demonstrate that BARcode can retrieve inspirations that are valuable to engineers and designers tackling real-world problems, as well as recover famous historical BID examples. We release data and code; we view BARcode as a step towards addressing the challenges that have historically hindered the practical application of BID to engineering innovation.
Hen Emuna, Nadav Borenstein, Hyeonsu B. Kang, Joel Chan, Aniket Kittur, Dafna Shahaf
AAAI2
2024 What Languages are Easy to Language-Model? A Perspective from Learning Probabilistic Regular Languages
abstract
Nadav Borenstein, Anej Svete, Robin Chan, Josef Valvoda, Franz Nowak, Isabelle Augenstein, Eleanor Chodroff, Ryan Cotterell. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Nadav Borenstein, Anej Svete, Robin Shing Moon Chan, Josef Valvoda, Franz Nowak, Isabelle Augenstein, Eleanor Chodroff, Ryan Cotterell
ACL (1)1
2024 Can Transformers Learn n-gram Language Models?
abstract
Much theoretical work has described the ability of transformers to represent formal languages.However, linking theoretical results to empirical performance is not straightforward due to the complex interplay between the architecture, the learning algorithm, and training data.To test whether theoretical lower bounds imply learnability of formal languages, we turn to recent work relating transformers to n-gram language models (LMs).We study transformers' ability to learn random n-gram LMs of two kinds: ones with arbitrary next-symbol probabilities and ones where those are defined with shared parameters.We find that classic estimation techniques for n-gram LMs such as add-λ smoothing outperform transformers on the former, while transformers perform better on the latter, outperforming methods specifically designed to learn n-gram LMs.github.com/rycolab/learning-ngrams
Anej Svete, Nadav Borenstein, Mike Zhou, Isabelle Augenstein, Ryan Cotterell
EMNLP2
2023 Multilingual Event Extraction from Historical Newspaper Adverts
abstract
NLP methods can aid historians in analyzing textual materials in greater volumes than manually feasible.Developing such methods poses substantial challenges though.First, acquiring large, annotated historical datasets is difficult, as only domain experts can reliably label them.Second, most available off-the-shelf NLP models are trained on modern language texts, rendering them significantly less effective when applied to historical corpora.This is particularly problematic for less well studied tasks, and for languages other than English.This paper addresses these challenges while focusing on the under-explored task of event extraction from a novel domain of historical texts.We introduce a new multilingual dataset in English, French, and Dutch composed of newspaper ads from the early modern colonial period reporting on enslaved people who liberated themselves from enslavement.We find that: 1) even with scarce annotated data, it is possible to achieve surprisingly good results by formulating the problem as an extractive QA task and leveraging existing datasets and models for modern languages; and 2) cross-lingual low-resource learning for historical languages is highly challenging, and machine translation of the historical datasets to the considered target languages is, in practice, often the best-performing solution.
Nadav Borenstein, Natalia da Silva Perez, Isabelle Augenstein
ACL (1)1
2023 PHD: Pixel-Based Language Modeling of Historical Documents
abstract
The digitisation of historical documents has provided historians with unprecedented research opportunities. Yet, the conventional approach to analysing historical documents involves converting them from images to text using OCR, a process that overlooks the potential benefits of treating them as images and introduces high levels of noise. To bridge this gap, we take advantage of recent advancements in pixel-based language models trained to reconstruct masked patches of pixels instead of predicting token distributions. Due to the scarcity of real historical scans, we propose a novel method for generating synthetic scans to resemble real historical documents. We then pre-train our model, PHD, on a combination of synthetic scans and real historical newspapers from the 1700-1900 period. Through our experiments, we demonstrate that PHD exhibits high proficiency in reconstructing masked image patches and provide evidence of our model's noteworthy language understanding capabilities. Notably, we successfully apply our model to a historical QA task, highlighting its usefulness in this domain.
Nadav Borenstein, Phillip Rust, Desmond Elliott, Isabelle Augenstein
EMNLP1
2021 How Did This Get Funded?! Automatically Identifying Quirky Scientific Achievements
abstract
Chen Shani, Nadav Borenstein, Dafna Shahaf. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Chen Shani, Nadav Borenstein, Dafna Shahaf
ACL/IJCNLP (1)2