VLDB 2026 Research / reviewers in the wild / expert
Noy Sternlicht
dblp:408/3580
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Language models and text generation · 57% Information extraction and text analysis · 38% 3D vision · 6% | |
| Databases, data mining, and information retrieval
1 paper |
Knowledge graphs · 100% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › document analysis › scholarly text analysis
citation analysis |
1.0 | 1 | 2026 | In-depth Research Impact Summarization through Fine-Grained Temporal Citation Analysis · ACL (1) 2026 |
Natural language and speech › Information extraction and text analysis › document analysis › scholarly text analysis
citation intent classification |
1.0 | 1 | 2026 | In-depth Research Impact Summarization through Fine-Grained Temporal Citation Analysis · ACL (1) 2026 |
Natural language and speech › Language models and text generation
text summarization |
1.0 | 1 | 2026 | In-depth Research Impact Summarization through Fine-Grained Temporal Citation Analysis · ACL (1) 2026 |
Knowledge graphs
knowledge graph construction |
1.0 | 1 | 2026 | CHIMERA: A Knowledge Base of Scientific Idea Recombinations for Research Analysis and Ideation · ACL (1) 2026 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.9 | 1 | 2025 | Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation · EMNLP 2025 |
Natural language and speech › Language models and text generation › large language model evaluation
LLM judge |
0.9 | 1 | 2025 | Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation · EMNLP 2025 |
Computer vision › 3D vision › geometric estimation › geometric model fitting
hypothesis generation |
0.3 | 1 | 2026 | CHIMERA: A Knowledge Base of Scientific Idea Recombinations for Research Analysis and Ideation · ACL (1) 2026 |
Methods — techniques the papers use, named apart from their topics
fine-tuning · 2.0LLM-based extraction · 2.0temporal analysis · 1.0human annotation · 0.9benchmarking · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | In-depth Research Impact Summarization through Fine-Grained Temporal Citation AnalysisabstractUnderstanding the impact of scientific publications is crucial for identifying breakthroughs and guiding future research.Traditional metrics based on citation counts often miss the nuanced ways a paper contributes to its field.In this work, we propose a new task: generating nuanced, expressive, and time-aware impact summaries that capture both praise (confirmation citations) and critique (correction citations) through the evolution of fine-grained citation intents.We introduce an evaluation framework tailored to this task, showing moderate to strong human correlation on subjective metrics such as insightfulness.Expert feedback from professors reveals a strong interest in these summaries and suggests future improvements.Data and code are made available.1 Hiba Arnaout, Noy Sternlicht, Tom Hope, Iryna Gurevych |
ACL (1) | 2 |
| 2026 | CHIMERA: A Knowledge Base of Scientific Idea Recombinations for Research Analysis and IdeationabstractA hallmark of human innovation is recombination-the creation of novel ideas by integrating elements from existing concepts and mechanisms.In this work, we introduce CHIMERA, the first large-scale Knowledge Base (KB) of recombination examples automatically mined from the scientific literature.CHIMERA enables empirical analysis of how scientists recombine concepts and draw inspiration from different areas, and enables training models that propose cross-disciplinary research directions.To construct this KB, we define a new information extraction task: identifying recombination instances in papers.We curate an expertannotated dataset and use it to fine-tune an LLM-based extraction model, which we apply to a broad corpus of AI papers.We also demonstrate generalization to a biological domain.We showcase the utility of CHIMERA through two applications.First, we analyze patterns of recombination across AI subfields.Second, we train a scientific hypothesis generation model using the KB, showing that it can propose directions that researchers rate as inspiring. Noy Sternlicht, Tom Hope |
ACL (1) | 1 |
| 2025 | Debatable Intelligence: Benchmarking LLM Judges via Debate Speech EvaluationabstractWe introduce Debate Speech Evaluation as a novel and challenging benchmark for assessing LLM judges.Evaluating debate speeches requires a deep understanding of the speech at multiple levels, including argument strength and relevance, the coherence and organization of the speech, the appropriateness of its style and tone, and so on.This task involves a unique set of cognitive abilities that previously received limited attention in systematic LLM benchmarking.To explore such skills, we leverage a dataset of over 600 meticulously annotated debate speeches and present the first in-depth analysis of how state-of-the-art LLMs compare to human judges on this task.Our findings reveal a nuanced picture: while larger models can approximate individual human judgments in some respects, they differ substantially in their overall judgment behavior.We also investigate the ability of frontier LLMs to generate persuasive, opinionated speeches, showing that models may perform at a human level on this task. Noy Sternlicht, Ariel Gera, Roy Bar-Haim, Tom Hope, Noam Slonim |
EMNLP | 1 |