EDBT 2026 Demo / reviewers in the wild / expert
Jonathan Herzig
dblp:133/3687
· DBLP profile ↗
22ranked-venue papers
8as first author
13since 2021 · last 2024
0009-0000-7227-6557ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 7 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning ChainsabstractAlon Jacovi, Yonatan Bitton, Bernd Bohnet, Jonathan Herzig, Or Honovich, Michael Tseng, Michael Collins, Roee Aharoni, Mor Geva. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Alon Jacovi, Yonatan Bitton, Bernd Bohnet, Jonathan Herzig, Or Honovich, Michael Tseng, Michael Collins 0001, Roee Aharoni, Mor Geva |
ACL (1) | 4 |
| 2024 | Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?abstractWhen large language models are aligned via supervised fine-tuning, they may encounter new factual information that was not acquired through pre-training.It is often conjectured that this can teach the model the behavior of hallucinating factually incorrect responses, as the model is trained to generate facts that are not grounded in its pre-existing knowledge.In this work, we study the impact of such exposure to new knowledge on the capability of the fine-tuned model to utilize its pre-existing knowledge.To this end, we design a controlled setup, focused on closedbook QA, where we vary the proportion of the fine-tuning examples that introduce new knowledge.We demonstrate that large language models struggle to acquire new factual knowledge through fine-tuning, as fine-tuning examples that introduce new knowledge are learned significantly slower than those consistent with the model's knowledge.However, we also find that as the examples with new knowledge are eventually learned, they linearly increase the model's tendency to hallucinate.Taken together, our results highlight the risk in introducing new factual knowledge through fine-tuning, and support the view that large language models mostly acquire factual knowledge through pre-training, whereas finetuning teaches them to use it more efficiently. Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal, Amir Feder, Roi Reichart, Jonathan Herzig |
EMNLP | 7 |
| 2024 | Representation Surgery: Theory and Practice of Affine SteeringabstractLanguage models often exhibit undesirable behavior, e.g., generating toxic or gender-biased text. In the case of neural language models, an encoding of the undesirable behavior is often present in the model’s representations. Thus, one natural (and common) approach to prevent the model from exhibiting undesirable behavior is to steer the model’s representations in a manner that reduces the probability of it generating undesirable text. This paper investigates the formal and empirical properties of steering functions, i.e., transformation of the neural language model’s representations that alter its behavior. First, we derive two optimal, in the least-squares sense, affine steering functions under different constraints. Our theory provides justification for existing approaches and offers a novel, improved steering approach. Second, we offer a series of experiments that demonstrate the empirical effectiveness of the methods in mitigating bias and reducing toxic generation. Shashwat Singh 0001, Shauli Ravfogel, Jonathan Herzig, Roee Aharoni, Ryan Cotterell, Ponnurangam Kumaraguru |
ICML | 3 |
| 2024 | TACT: Advancing Complex Aggregative Reasoning with Information Extraction ToolsabstractLarge Language Models (LLMs) often do not perform well on queries that require the aggregation of information across texts. To better evaluate this setting and facilitate modeling efforts, we introduce TACT - Text And Calculations through Tables, a dataset crafted to evaluate LLMs' reasoning and computational abilities using complex instructions. TACT contains challenging instructions that demand stitching information scattered across one or more texts, and performing complex integration on this information to generate the answer. We construct this dataset by leveraging an existing dataset of texts and their associated tables. For each such tables, we formulate new queries, and gather their respective answers. We demonstrate that all contemporary LLMs perform poorly on this dataset, achieving an accuracy below 38%. To pinpoint the difficulties and thoroughly dissect the problem, we analyze model performance across three components: table-generation, Pandas command-generation, and execution. Unexpectedly, we discover that each component presents substantial challenges for current LLMs. These insights lead us to propose a focused modeling framework, which we refer to as IE as a tool. Specifically, we propose to add "tools" for each of the above steps, and implement each such tool with few-shot prompting. This approach shows an improvement over existing prompting techniques, offering a promising direction for enhancing model capabilities in these tasks. Avi Caciularu, Alon Jacovi, Eyal Ben-David, Sasha Goldshtein, Tal Schuster, Jonathan Herzig, Gal Elidan, Amir Globerson |
NeurIPS | 6 |
| 2023 | TrueTeacher: Learning Factual Consistency Evaluation with Large Language ModelsabstractFactual consistency evaluation is often conducted using Natural Language Inference (NLI) models, yet these models exhibit limited success in evaluating summaries.Previous work improved such models with synthetic training data.However, the data is typically based on perturbed human-written summaries, which often differ in their characteristics from real model-generated summaries and have limited coverage of possible factual errors.Alternatively, large language models (LLMs) have recently shown promising results in directly evaluating generative tasks, but are too computationally expensive for practical use.Motivated by these limitations, we introduce TrueTeacher, a method for generating synthetic data by annotating diverse model-generated summaries using a LLM.Unlike prior work, TrueTeacher does not rely on human-written summaries, and is multilingual by nature.Experiments on the TRUE benchmark show that a student model trained using our data, substantially outperforms both the state-of-the-art model with similar capacity, and the LLM teacher.In a systematic study, we compare TrueTeacher to existing synthetic data generation methods and demonstrate its superiority and robustness to domain-shift.We also show that our method generalizes to multilingual scenarios.Lastly, we release our largescale synthetic dataset (1.4M examples), generated using TrueTeacher, and a checkpoint trained on this data.1 Zorik Gekhman, Jonathan Herzig, Roee Aharoni, Chen Elkind, Idan Szpektor |
EMNLP | 2 |
| 2023 | Evaluating and Modeling Attribution for Cross-Lingual Question AnsweringabstractBenjamin Muller, John Wieting, Jonathan Clark, Tom Kwiatkowski, Sebastian Ruder, Livio Soares, Roee Aharoni, Jonathan Herzig, Xinyi Wang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Benjamin Muller, John Wieting, Jonathan H. Clark, Tom Kwiatkowski, Sebastian Ruder, Livio B. Soares, Roee Aharoni, Jonathan Herzig, Xinyi Wang 0001 |
EMNLP | 8 |
| 2023 | What You See is What You Read? Improving Text-Image Alignment EvaluationabstractAutomatically determining whether a text and a corresponding image are semantically aligned is a significant challenge for vision-language models, with applications in generative text-to-image and image-to-text tasks. In this work, we study methods for automatic text-image alignment evaluation. We first introduce SeeTRUE: a comprehensive evaluation set, spanning multiple datasets from both text-to-image and image-to-text generation tasks, with human judgements for whether a given text-image pair is semantically aligned. We then describe two automatic methods to determine alignment: the first involving a pipeline based on question generation and visual question answering models, and the second employing an end-to-end classification approach by finetuning multimodal pretrained models. Both methods surpass prior approaches in various text-image alignment tasks, with significant improvements in challenging cases that involve complex composition or unnatural images. Finally, we demonstrate how our approaches can localize specific misalignments between an image and a given text, and how they can be used to automatically re-rank candidates in text-to-image generation. Michal Yarom, Yonatan Bitton, Soravit Changpinyo, Roee Aharoni, Jonathan Herzig, Oran Lang, Eran Ofek, Idan Szpektor |
NeurIPS | 5 |
| 2022 | Evaluating the Impact of Model Scale for Compositional Generalization in Semantic ParsingabstractLinlu Qiu, Peter Shaw, Panupong Pasupat, Tianze Shi, Jonathan Herzig, Emily Pitler, Fei Sha, Kristina Toutanova. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Linlu Qiu, Peter Shaw 0004, Panupong Pasupat, Tianze Shi, Jonathan Herzig, Emily Pitler, Fei Sha, Kristina Toutanova |
EMNLP | 5 |
| 2022 | TRUE: Re-evaluating Factual Consistency EvaluationabstractOr Honovich, Roee Aharoni, Jonathan Herzig, Hagai Taitelbaum, Doron Kukliansy, Vered Cohen, Thomas Scialom, Idan Szpektor, Avinatan Hassidim, Yossi Matias. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Or Honovich, Roee Aharoni, Jonathan Herzig, Hagai Taitelbaum, Doron Kukliansky, Vered Cohen, Thomas Scialom, Idan Szpektor, Avinatan Hassidim, Yossi Matias |
NAACL-HLT | 3 |
| 2022 | Learning To Retrieve Prompts for In-Context LearningabstractIn-context learning is a recent paradigm in natural language understanding, where a large pretrained language model (LM) observes a test instance and a few training examples as its input, and directly decodes the output without any update to its parameters.However, performance has been shown to strongly depend on the selected training examples (termed prompts).In this work, we propose an efficient method for retrieving prompts for in-context learning using annotated data and an LM.Given an inputoutput pair, we estimate the probability of the output given the input and a candidate training example as the prompt, and label training examples as positive or negative based on this probability.We then train an efficient dense retriever from this data, which is used to retrieve training examples as prompts at test time.We evaluate our approach on three sequence-tosequence tasks where language utterances are mapped to meaning representations, and find that it substantially outperforms prior work and multiple baselines across the board. Ohad Rubin, Jonathan Herzig, Jonathan Berant |
NAACL-HLT | 2 |
| 2021 | Span-based Semantic Parsing for Compositional GeneralizationabstractJonathan Herzig, Jonathan Berant. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jonathan Herzig, Jonathan Berant |
ACL/IJCNLP (1) | 1 |
| 2021 | Finding needles in a haystack: Sampling Structurally-diverse Training Sets from Synthetic Data for Compositional GeneralizationabstractModern semantic parsers suffer from two principal limitations.First, training requires expensive collection of utterance-program pairs.Second, semantic parsers fail to generalize at test time to new compositions/structures that have not been observed during training.Recent research has shown that automatic generation of synthetic utterance-program pairs can alleviate the first problem, but its potential for the second has thus far been under-explored.In this work, we investigate automatic generation of synthetic utterance-program pairs for improving compositional generalization in semantic parsing.Given a small training set of annotated examples and an "infinite" pool of synthetic examples, we select a subset of synthetic examples that are structurally-diverse and use them to improve compositional generalization.We evaluate our approach on a new split of the schema2QA dataset, and show that it leads to dramatic improvements in compositional generalization as well as moderate improvements in the traditional i.i.d setup.Moreover, structurally-diverse sampling achieves these improvements with as few as 5K examples, compared to 1M examples when sampling uniformly at random -a 200x improvement in data efficiency. Inbar Oren, Jonathan Herzig, Jonathan Berant |
EMNLP (1) | 2 |
| 2021 | Open Domain Question Answering over Tables via Dense RetrievalabstractJonathan Herzig, Thomas Müller, Syrine Krichene, Julian Eisenschlos. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Jonathan Herzig, Thomas Müller 0009, Syrine Krichene, Julian Martin Eisenschlos |
NAACL-HLT | 1 |
| 2020 | TaPas: Weakly Supervised Table Parsing via Pre-trainingabstractAnswering natural language questions over tables is usually seen as a semantic parsing task.To alleviate the collection cost of full logical forms, one popular approach focuses on weak supervision consisting of denotations instead of logical forms.However, training semantic parsers from weak supervision poses difficulties, and in addition, the generated logical forms are only used as an intermediate step prior to retrieving the denotation.In this paper, we present TAPAS, an approach to question answering over tables without generating logical forms.TAPAS trains from weak supervision, and predicts the denotation by selecting table cells and optionally applying a corresponding aggregation operator to such selection.TAPAS extends BERT's architecture to encode tables as input, initializes from an effective joint pre-training of text segments and tables crawled from Wikipedia, and is trained end-to-end.We experiment with three different semantic parsing datasets, and find that TAPAS outperforms or rivals semantic parsing models by improving state-of-the-art accuracy on SQA from 55.1 to 67.2 and performing on par with the state-of-the-art on WIKISQL and WIKITQ, but with a simpler model architecture.We additionally find that transfer learning, which is trivial in our setting, from WIK-ISQL to WIKITQ, yields 48.7 accuracy, 4.2 points above the state-of-the-art. Jonathan Herzig, Pawel Krzysztof Nowak, Thomas Müller 0009, Francesco Piccinno, Julian Martin Eisenschlos |
ACL | 1 |
| 2019 | TalkSumm: A Dataset and Scalable Annotation Method for Scientific Paper Summarization Based on Conference TalksabstractCurrently, no large-scale training data is available for the task of scientific paper summarization.In this paper, we propose a novel method that automatically generates summaries for scientific papers, by utilizing videos of talks at scientific conferences.We hypothesize that such talks constitute a coherent and concise description of the papers' content, and can form the basis for good summaries.We collected 1716 papers and their corresponding videos, and created a dataset of paper summaries.A model trained on this dataset achieves similar performance as models trained on a dataset of summaries created manually.In addition, we validated the quality of our summaries by human experts. Guy Lev, Michal Shmueli-Scheuer, Jonathan Herzig, Achiya Jerbi, David Konopnicki |
ACL (1) | 3 |
| 2019 | Don't paraphrase, detect! Rapid and Effective Data Collection for Semantic ParsingabstractJonathan Herzig, Jonathan Berant. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jonathan Herzig, Jonathan Berant |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Detecting Persuasive Arguments based on Author-Reader Personality Traits and their InteractionabstractPersuasion is one of the most frequent, albeit challenging, tasks in human interaction. In a textual argument, one party (author) aims to change the view of the other party (reader). In this paper, we propose to detect persuasive textual arguments while considering the parties personality traits. We find that we can substantially improve accuracy by introducing features that capture author-reader personality traits and their interaction. Our model improves performance of state-of-the-art baselines from 66% to 71% on a new dataset of more than 19K arguments we collected. Michal Shmueli-Scheuer, Jonathan Herzig, David Konopnicki, Tommy Sandbank |
UMAP | 2 |
| 2018 | Decoupling Structure and Lexicon for Zero-Shot Semantic ParsingabstractBuilding a semantic parser quickly in a new domain is a fundamental challenge for conversational interfaces, as current semantic parsers require expensive supervision and lack the ability to generalize to new domains.In this paper, we introduce a zero-shot approach to semantic parsing that can parse utterances in unseen domains while only being trained on examples in other source domains.First, we map an utterance to an abstract, domainindependent, logical form that represents the structure of the logical form, but contains slots instead of KB constants.Then, we replace slots with KB constants via lexical alignment scores and global inference.Our model reaches an average accuracy of 53.4% on 7 domains in the OVERNIGHT dataset, substantially better than other zero-shot baselines, and performs as good as a parser trained on over 30% of the target domain examples. Jonathan Herzig, Jonathan Berant |
EMNLP | 1 |
| 2018 | Detecting Egregious Conversations between Customers and Virtual AgentsabstractTommy Sandbank, Michal Shmueli-Scheuer, Jonathan Herzig, David Konopnicki, John Richards, David Piorkowski. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Tommy Sandbank, Michal Shmueli-Scheuer, Jonathan Herzig, David Konopnicki, John T. Richards, David Piorkowski |
NAACL-HLT | 3 |
| 2017 | Neural Response Generation for Customer Service based on Personality TraitsabstractWe present a neural response generation model that generates responses conditioned on a target personality.The model learns high level features based on the target personality, and uses them to update its hidden state.Our model achieves performance improvements in both perplexity and BLEU scores over a baseline sequence-to-sequence model, and is validated by human judges. Jonathan Herzig, Michal Shmueli-Scheuer, Tommy Sandbank, David Konopnicki |
INLG | 1 |
| 2016 | Classifying Emotions in Customer Support Dialogues in Social MediaabstractJonathan Herzig, Guy Feigenblat, Michal Shmueli-Scheuer, David Konopnicki, Anat Rafaeli, Daniel Altman, David Spivak. Proceedings of the 17th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2016. Jonathan Herzig, Guy Feigenblat, Michal Shmueli-Scheuer, David Konopnicki, Anat Rafaeli, Daniel Altman, David Spivak |
SIGDIAL Conference | 1 |
| 2016 | Predicting Customer Satisfaction in Customer Support Conversations in Social Media Using Affective FeaturesabstractProviding customer support through social media channels is gaining popularity. In such a context, predicting customer satisfaction in an early stage of a service conversation is important. Such an analysis can help personalize agent assignment to maximize customer satisfaction, and prioritize conversations. In this paper, we show that affective features such as customer's and agent's personality traits and emotion expression improve prediction of customer satisfaction when added to more typical text based features. We only utilize information extracted from the first customer conversation turn and previous customer and agent social network activity. Thus, our customer satisfaction classifier outputs its prediction in an early stage of the conversation, before any interaction has taken place between the customer and an agent. Our model was trained and tested on a Twitter conversations dataset of two customer support services, and shows an improvement of 30% in F1-score for predicting dissatisfaction. Jonathan Herzig, Guy Feigenblat, Michal Shmueli-Scheuer, David Konopnicki, Anat Rafaeli |
UMAP | 1 |