Aviv Slobodkin

dblp:290/2100 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021
YearPublicationVenuePosition
2026 PrefixNLI: Detecting Factual Inconsistencies as Soon as They Arise
abstract
Natural Language Inference (NLI) models have been used in various ways to improve the factuality of LLM outputs.This is typically done by applying an NLI model to judge whether the model output is entailed from the supposed evidence, triggering some corrective actions, such as beam reranking at inference time or RL rewards during training.While NLI models are trained to detect factual inconsistencies over complete sentences, decisions in the common autoregressive generation architecture are made for each evolving text prefix, during decoding.Addressing this setting, we generalize the entailment detection task to apply over arbitrary text prefixes, and suggest its utility for improving generation faithfulness.Providing suitable evaluation and training datasets for this task, we train MiniTruePrefixes, a novel specialized model that better detects factual inconsistencies over text prefixes, outperforming comparable baseline NLI models by 5-14 F1 points in prefix-level entailment.We further demonstrate that integrating MiniTruePrefixes into a controlled decoding framework substantially improves factual consistency in abstractive summarization.When guided by Mini-TruePrefixes, LLaMA-3.2-3B-Instructmatches the faithfulness and runtime of the 8B model from the same model family, while using only half the memory.
Sapir Harary, Eran Hirsch, Aviv Slobodkin, David Wan, Mohit Bansal, Ido Dagan
ACL (1)3
2025 LAQuer: Localized Attribution Queries in Content-grounded Generation
abstract
Eran Hirsch, Aviv Slobodkin, David Wan, Elias Stengel-Eskin, Mohit Bansal, Ido Dagan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Eran Hirsch, Aviv Slobodkin, David Wan, Elias Stengel-Eskin, Mohit Bansal, Ido Dagan
ACL (1)2
2025 A Unifying Scheme for Extractive Content Selection Tasks
abstract
Abstract A broad range of NLP tasks involve selecting relevant text spans from given source texts. Despite this shared objective, such content selection tasks have traditionally been studied in isolation, each with its own modeling approaches, datasets, and evaluation metrics. In this work, we propose instruction-guided content selection (IGCS) as a beneficial unified framework for such settings, where the task definition and any instance-specific request are encapsulated as instructions to a language model. To promote this framework, we introduce IGCS-Bench, the first unified benchmark covering diverse content selection tasks. Further, we create a large generic synthetic dataset that can be leveraged for diverse content selection tasks, and show that transfer learning with these datasets often boosts performance, whether dedicated training for the targeted task is available or not. Finally, we address generic inference time issues that arise in LLM-based modeling of content selection, assess a generic evaluation metric, and overall propose the utility of our resources and methods for future content selection models.1
Shmuel Amar, Ori Shapira, Aviv Slobodkin, Ido Dagan
Trans. Assoc. Comput. Linguistics3
2024 Explicating the Implicit: Argument Detection Beyond Sentence Boundaries
abstract
Paul Roit, Aviv Slobodkin, Eran Hirsch, Arie Cattan, Ayal Klein, Valentina Pyatkin, Ido Dagan. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Paul Roit, Aviv Slobodkin, Eran Hirsch, Arie Cattan, Ayal Klein, Valentina Pyatkin, Ido Dagan
ACL (1)2
2024 Attribute First, then Generate: Locally-attributable Grounded Text Generation
abstract
Recent efforts to address hallucinations in Large Language Models (LLMs) have focused on attributed text generation, which supplements generated texts with citations of supporting sources for post-generation fact-checking and corrections.Yet, these citations often point to entire documents or paragraphs, burdening users with extensive verification work.In this paper, we introduce a locally-attributable text generation approach, prioritizing concise attributions.Our method, named "Attribute First, then Generate", breaks down the conventional end-to-end generation process into three intuitive steps: content selection, sentence planning, and sequential sentence generation.By initially identifying relevant source segments ("select first") and then conditioning the generation process on them ("then generate"), we ensure these segments also act as the output's fine-grained attributions ("select" becomes "attribute").Tested on Multi-document Summarization and Long-form Question-answering, our method not only yields more concise citations than the baselines but also maintains-and in some cases enhances-both generation quality and attribution accuracy.Furthermore, it significantly reduces the time required for fact verification by human assessors.
Aviv Slobodkin, Eran Hirsch, Arie Cattan, Tal Schuster, Ido Dagan
ACL (1)1
2024 Is It Really Long Context if All You Need Is Retrieval? Towards Genuinely Difficult Long Context NLP
abstract
Improvements in language models' capabilities have pushed their applications towards longer contexts, making long-context evaluation and development an active research area.However, many disparate use cases are grouped together under the umbrella term of "long-context", defined simply by the total length of the model's input, including -for example -Needle-in-a-Haystack tasks, book summarization, and information aggregation.Given their varied difficulty, in this position paper we argue that conflating different tasks by their context length is unproductive.As a community, we require a more precise vocabulary to understand what makes long-context tasks similar or different.We propose to unpack the taxonomy of longcontext based on the properties that make them more difficult with longer contexts.We propose two orthogonal axes of difficulty: (I) Dispersion: How hard is it to find the necessary information in the context?(II) Scope: How much necessary information is there to find?We survey the literature on long context, provide justification for this taxonomy as an informative descriptor, and situate the literature with respect to it.We conclude that the most difficult and interesting settings, whose necessary information is very long and highly dispersed within the input, is severely under-explored.By using a descriptive vocabulary and discussing the relevant properties of difficulty in long context, we can implement more informed research in this area.We call for a careful design of tasks and benchmarks with distinctly long context, taking into account the characteristics that make it qualitatively different from shorter context.
Omer Goldman, Alon Jacovi, Aviv Slobodkin, Aviya Maimon, Ido Dagan, Reut Tsarfaty
EMNLP3
2024 Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models
abstract
Imagine observing someone scratching their arm; to understand why, additional context would be necessary. However, spotting a mosquito nearby would immediately offer a likely explanation for the person’s discomfort, thereby alleviating the need for further information. This example illustrates how subtle visual cues can challenge our cognitive skills and demonstrates the complexity of interpreting visual scenarios. To study these skills, we present Visual Riddles, a benchmark aimed to test vision and language models on visual riddles requiring commonsense and world knowledge. The benchmark comprises 400 visual riddles, each featuring a unique image created by a variety of text-to-image models, question, ground-truth answer, textual hint, and attribution. Human evaluation reveals that existing models lag significantly behind human performance, which is at 82% accuracy, with Gemini-Pro-1.5 leading with 40% accuracy. Our benchmark comes with automatic evaluation tasks to make assessment scalable. These findings underscore the potential of Visual Riddles as a valuable resource for enhancing vision and language models’ capabilities in interpreting complex visual scenarios. Data, code, and leaderboard are available at https://visual-riddles.github.io/.
Nitzan Guetta, Aviv Slobodkin, Aviya Maimon, Eliya Habba, Royi Rassin, Yonatan Bitton, Idan Szpektor, Amir Globerson, Yuval Elovici
NeurIPS2
2023 The Curious Case of Hallucinatory (Un)answerability: Finding Truths in the Hidden States of Over-Confident Large Language Models
abstract
Large language models (LLMs) have been shown to possess impressive capabilities, while also raising crucial concerns about the faithfulness of their responses.A primary issue arising in this context is the management of (un)answerable queries by LLMs, which often results in hallucinatory behavior due to overconfidence.In this paper, we explore the behavior of LLMs when presented with (un)answerable queries.We ask: do models represent the fact that the question is (un)answerable when generating a hallucinatory answer?Our results show strong indications that such models encode the answerability of an input query, with the representation of the first decoded token often being a strong indicator.These findings shed new light on the spatial organization within the latent representations of LLMs, unveiling previously unexplored facets of these models.Moreover, they pave the way for the development of improved decoding techniques with better adherence to factual generation, particularly in scenarios where query (un)answerability is a concern. 1
Aviv Slobodkin, Omer Goldman, Avi Caciularu, Ido Dagan, Shauli Ravfogel
EMNLP1
2022 Controlled Text Reduction
abstract
Producing a reduced version of a source text, as in generic or focused summarization, inherently involves two distinct subtasks: deciding on targeted content and generating a coherent text conveying it.While some popular approaches address summarization as a single end-to-end task, prominent works support decomposed modeling for individual subtasks.Further, semi-automated text reduction is also very appealing, where users may identify targeted content while models would generate a corresponding coherent summary.In this paper, we focus on the second subtask, of generating coherent text given pre-selected content.Concretely, we formalize Controlled Text Reduction as a standalone task, whose input is a source text with marked spans of targeted content ("highlighting").A model then needs to generate a coherent text that includes all and only the target information.We advocate the potential of such models, both for modular fully-automatic summarization, as well as for semi-automated human-in-the-loop use cases.Facilitating proper research, we crowdsource high-quality dev and test datasets for the task.Further, we automatically generate a larger "silver" training dataset from available summarization benchmarks, leveraging a pretrained summary-source alignment model.Finally, employing these datasets, we present a supervised baseline model, showing promising results and insightful analyses.1
Aviv Slobodkin, Paul Roit, Eran Hirsch, Ori Ernst, Ido Dagan
EMNLP1
2021 Mediators in Determining what Processing BERT Performs First
abstract
Probing neural models for the ability to perform downstream tasks using their activation patterns is often used to localize what parts of the network specialize in performing what tasks.However, little work addressed potential mediating factors in such comparisons.As a test-case mediating factor, we consider the prediction's context length, namely the length of the span whose processing is minimally required to perform the prediction.We show that not controlling for context length may lead to contradictory conclusions as to the localization patterns of the network, depending on the distribution of the probing dataset.Indeed, when probing BERT with seven tasks, we find that it is possible to get 196 different rankings between them when manipulating the distribution of context lengths in the probing dataset.We conclude by presenting best practices for conducting such comparisons in the future. 1
Aviv Slobodkin, Leshem Choshen, Omri Abend
NAACL-HLT1