Yftah Ziser

dblp:188/6096 · DBLP profile ↗
← Back
21ranked-venue papers
5as first author
16since 2021 · last 2026
0009-0002-6228-9471ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 4 first-author · 16 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Beyond Next Token Probabilities: Learnable, Fast Detection of Hallucinations and Data Contamination on LLM Output Distributions
abstract
The automated detection of hallucinations and training data contamination is pivotal to the safe deployment of Large Language Models (LLMs). These tasks are particularly challenging in settings where no access to model internals is available. Current approaches in this setup typically leverage only the probabilities of actual tokens in the text, relying on simple task-specific heuristics. Crucially, they overlook the information contained in the full sequence of next-token probability distributions. We propose to go beyond hand-crafted decision rules by learning directly from the complete observable output of LLMs — consisting not only of next-token probabilities, but also the full sequence of next-token distributions. We refer to this as the LLM Output Signature (LOS), and treat it as a reference data type for detecting hallucinations and data contamination. To that end, we introduce LOS-Net, a lightweight attention-based architecture trained on an efficient encoding of the LOS, which can provably approximate a broad class of existing techniques for both tasks. Empirically, LOS-Net achieves superior performance across diverse benchmarks and LLMs, while maintaining extremely low detection latency. Furthermore, it demonstrates promising transfer capabilities across datasets and LLMs.
Guy Bar-Shalom, Fabrizio Frasca, Derek Lim, Yoav Gelberg, Yftah Ziser, Ran El-Yaniv, Gal Chechik, Haggai Maron
AAAI5
2025 A Simple Yet Effective Method for Non-Refusing Context Relevant Fine-grained Safety Steering in LLMs
abstract
Fine-tuning large language models (LLMs) to meet evolving safety policies is costly and impractical.Mechanistic interpretability enables inference-time control through latent activation steering, but its potential for precise, customizable safety adjustments remains underexplored.We propose SAFESTEER, a simple and effective method to guide LLM outputs by (i) leveraging category-specific steering vectors for fine-grained control, (ii) applying a gradient-free, unsupervised approach that enhances safety while preserving text quality and topic relevance without forcing explicit refusals, and (iii) eliminating the need for contrastive safe data.Across multiple LLMs, datasets, and risk categories, SAFESTEER provides precise control, avoids blanket refusals, and directs models to generate safe, relevant content, aligning with recent findings that simple activation-steering techniques often outperform more complex alternatives.Content Warning: This paper contains examples of critically harmful language.
Shaona Ghosh, Amrita Bhattacharjee, Yftah Ziser, Christopher Parisien
EMNLP3
2025 Iterative Multilingual Spectral Attribute Erasure
abstract
Multilingual representations embed words with similar meanings to share a common semantic space across languages, creating opportunities to transfer debiasing effects between languages.However, existing methods for debiasing are unable to exploit this opportunity because they operate on individual languages.We present Iterative Multilingual Spectral Attribute Erasure (IMSAE), which identifies and mitigates joint bias subspaces across multiple languages through iterative SVD-based truncation.Evaluating IMSAE across eight languages and five demographic dimensions, we demonstrate its effectiveness in both standard and zero-shot settings, where target language data is unavailable, but linguistically similar languages can be used for debiasing.Our comprehensive experiments across diverse language models (BERT, Llama, Mistral) show that IMSAE outperforms traditional monolingual and cross-lingual approaches while maintaining model utility. 1
Shun Shao, Yftah Ziser, Zheng Zhao 0005, Yifu Qiu, Shay B. Cohen, Anna Korhonen
EMNLP2
2025 TSPRank: Bridging Pairwise and Listwise Methods with a Bilinear Travelling Salesman Model
abstract
Traditional Learning-To-Rank (LETOR) approaches, including pairwise methods like RankNet and LambdaMART, often fall short by solely focusing on pairwise comparisons, leading to sub-optimal global rankings. Conversely, deep learning based listwise methods, while aiming to optimise entire lists, require complex tuning and yield only marginal improvements over robust pairwise models. To overcome these limitations, we introduce Travelling Salesman Problem Rank (TSPRank), a hybrid pairwise-listwise ranking method. TSPRank reframes the ranking problem as a Travelling Salesman Problem (TSP), a well-known combinatorial optimisation challenge that has been extensively studied for its numerous solution algorithms and applications. This approach enables the modelling of pairwise relationships and leverages combinatorial optimisation to determine the listwise ranking. TSPRank can be directly integrated as an additional component into embeddings generated by existing backbone models to enhance ranking performance. Our extensive experiments across three backbone models on diverse tasks, including stock ranking, information retrieval, and historical events ordering, demonstrate that TSPRank significantly outperforms both pure pairwise and listwise methods. Our qualitative analysis reveals that TSPRank's main advantage over existing methods is its ability to harness global information better while ranking. TSPRank's robustness and superior performance across different domains highlight its potential as a versatile and effective LETOR solution.
Weixian Waylon Li, Yftah Ziser, Yifei Xie 0002, Shay B. Cohen, Tiejun Ma
KDD (1)2
2025 Beyond Token Probes: Hallucination Detection via Activation Tensors with ACT-ViT
abstract
Detecting hallucinations in Large Language Model-generated text is crucial for their safe deployment. While probing classifiers show promise, they operate on isolated layer–token pairs and are LLM-specific, limiting their effectiveness and hindering cross-LLM applications. In this paper, we introduce a novel approach to address these shortcomings. We build on the natural sequential structure of activation data in both axes (layers $\times$ tokens) and advocate treating full activation tensors akin to images. We design ACT-ViT, a Vision Transformer-inspired model that can be effectively and efficiently applied to activation tensors and supports training on data from multiple LLMs simultaneously. Through comprehensive experiments encompassing diverse LLMs and datasets, we demonstrate that ACT-ViT consistently outperforms traditional probing techniques while remaining extremely efficient for deployment. In particular, we show that our architecture benefits substantially from multi-LLM training, achieves strong zero-shot performance on unseen datasets, and can be transferred effectively to new LLMs through fine-tuning.
Guy Bar-Shalom, Fabrizio Frasca, Yaniv Galron, Yftah Ziser, Haggai Maron
NeurIPS4
2025 Policy Optimized Text-to-Image Pipeline Design
abstract
Text-to-image generation has evolved beyond single monolithic models to complex multi-component pipelines that combine various enhancement tools. While these pipelines significantly improve image quality, their effective design requires substantial expertise. Recent approaches automating this process through large language models (LLMs) have shown promise but suffer from two critical limitations: extensive computational requirements from generating images with hundreds of predefined pipelines, and poor generalization beyond memorized training examples. We introduce a novel reinforcement learning-based framework that addresses these inefficiencies. Our approach first trains an ensemble of reward models capable of predicting image quality scores directly from prompt-workflow combinations, eliminating the need for costly image generation during training. We then implement a two-phase training strategy: initial workflow prediction training followed by GRPO-based optimization that guides the model toward higher-performing regions of the workflow space. Additionally, we incorporate a classifier-free guidance based enhancement technique that extrapolates along the path between the initial and GRPO-tuned models, further improving output quality. We validate our approach through a set of comparisons, showing that it can successfully create new flows with greater diversity and lead to superior image quality compared to existing baselines.
Uri Gadot, Rinon Gal, Yftah Ziser, Gal Chechik, Shie Mannor
NeurIPS3
2024 Layer by Layer: Uncovering Where Multi-Task Learning Happens in Instruction-Tuned Large Language Models
abstract
Fine-tuning pre-trained large language models (LLMs) on a diverse array of tasks has become a common approach for building models that can solve various natural language processing (NLP) tasks. However, where and to what extent these models retain task-specific knowledge remains largely unexplored. This study investigates the task-specific information encoded in pre-trained LLMs and the effects of instruction tuning on their representations across a diverse set of over 60 NLP tasks. We use a set of matrix analysis tools to examine the differences between the way pre-trained and instruction-tuned LLMs store task-specific information. Our findings reveal that while some tasks are already encoded within the pre-trained LLMs, others greatly benefit from instruction tuning. Additionally, we pinpointed the layers in which the model transitions from high-level general representations to more task-oriented representations. This finding extends our understanding of the governing mechanisms of LLMs and facilitates future research in the fields of parameter-efficient transfer learning and multi-task learning. Our code is available at: https://github.com/zsquaredz/layerbylayer/
Zheng Zhao 0005, Yftah Ziser, Shay B. Cohen
EMNLP2
2024 Are Large Language Model Temporally Grounded?
abstract
Yifu Qiu, Zheng Zhao, Yftah Ziser, Anna Korhonen, Edoardo Ponti, Shay Cohen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yifu Qiu, Zheng Zhao 0005, Yftah Ziser, Anna Korhonen, Edoardo Maria Ponti, Shay B. Cohen
NAACL-HLT3
2024 Spectral Editing of Activations for Large Language Model Alignment
abstract
Large language models (LLMs) often exhibit undesirable behaviours, such as generating untruthful or biased content. Editing their internal representations has been shown to be effective in mitigating such behaviours on top of the existing alignment methods. We propose a novel inference-time editing method, namely spectral editing of activations (SEA), to project the input representations into directions with maximal covariance with the positive demonstrations (e.g., truthful) while minimising covariance with the negative demonstrations (e.g., hallucinated). We also extend our method to non-linear editing using feature functions. We run extensive experiments on benchmarks concerning truthfulness and bias with six open-source LLMs of different sizes and model families. The results demonstrate the superiority of SEA in effectiveness, generalisation to similar tasks, as well as computation and data efficiency. We also show that SEA editing only has a limited negative impact on other model capabilities.
Yifu Qiu, Zheng Zhao 0005, Yftah Ziser, Anna Korhonen, Edoardo Maria Ponti, Shay B. Cohen
NeurIPS3
2023 BERT Is Not The Count: Learning to Match Mathematical Statements with Proofs
abstract
We introduce a task consisting in matching a proof to a given mathematical statement.The task fits well within current research on Mathematical Information Retrieval and, more generally, mathematical article analysis (Mathematical Sciences, 2014).We present a dataset for the task (the MATCH dataset) consisting of over 180k statement-proof pairs extracted from modern mathematical research articles.1 We find this dataset highly representative of our task, as it consists of relatively new findings useful to mathematicians.We propose a bilinear similarity model and two decoding methods to match statements to proofs effectively.While the first decoding method matches a proof to a statement without being aware of other statements or proofs, the second method treats the task as a global matching problem.Through a symbol replacement procedure, we analyze the "insights" that pre-trained language models have in such mathematical article analysis and show that while these models perform well on this task with the best performing mean reciprocal rank of 73.7, they follow a relatively shallow symbolic analysis and matching to achieve that performance.2 * Work mostly done at the University of Edinburgh. 1 Our dataset and code are available at https:// github.com/waylonli/MATcH.2 Like Bert, The Count (or Count von Count; ) is a character from the television show Sesame Street.The Count likes counting, and his main role in the show is to teach this skill to children.
Weixian Waylon Li, Yftah Ziser, Maximin Coavoux, Shay B. Cohen
EACL2
2023 Gold Doesn't Always Glitter: Spectral Removal of Linear and Nonlinear Guarded Attribute Information
abstract
We describe a simple and effective method (Spectral Attribute removaL; SAL) to remove private or guarded information from neural representations.Our method uses matrix decomposition to project the input representations into directions with reduced covariance with the guarded information rather than maximal covariance as factorization methods normally use.We begin with linear information removal and proceed to generalize our algorithm to the case of nonlinear information removal using kernels.Our experiments demonstrate that our algorithm retains better main task performance after removing the guarded information compared to previous work.In addition, our experiments demonstrate that we need a relatively small amount of guarded attribute data to remove information about these attributes, which lowers the exposure to sensitive data and is more suitable for low-resource scenarios.1
Shun Shao, Yftah Ziser, Shay B. Cohen
EACL2
2023 Detecting and Mitigating Hallucinations in Multilingual Summarisation
abstract
Hallucinations pose a significant challenge to the reliability of neural models for abstractive summarisation.While automatically generated summaries may be fluent, they often lack faithfulness to the original document.This issue becomes even more pronounced in lowresource languages, where summarisation requires cross-lingual transfer.With the existing faithful metrics focusing on English, even measuring the extent of this phenomenon in crosslingual settings is hard.To address this, we first develop a novel metric, mFACT, evaluating the faithfulness of non-English summaries, leveraging translation-based transfer from multiple English faithfulness metrics.Through extensive experiments in multiple languages, we demonstrate that mFACT is best suited to detect hallucinations compared to alternative metrics.With mFACT, we assess a broad range of multilingual large language models, and find that they all tend to hallucinate often in languages different from English.We then propose a simple but effective method to reduce hallucinations in cross-lingual transfer, which weighs the loss of each training example by its faithfulness score.This method drastically increases both performance and faithfulness according to both automatic and human evaluation when compared to strong baselines for cross-lingual transfer such as MAD-X.
Yifu Qiu, Yftah Ziser, Anna Korhonen, Edoardo Maria Ponti, Shay B. Cohen
EMNLP2
2023 Erasure of Unaligned Attributes from Neural Representations
abstract
Abstract We present the Assignment-Maximization Spectral Attribute removaL (AMSAL) algorithm, which erases information from neural representations when the information to be erased is implicit rather than directly being aligned to each input example. Our algorithm works by alternating between two steps. In one, it finds an assignment of the input representations to the information to be erased, and in the other, it creates projections of both the input representations and the information to be erased into a joint latent space. We test our algorithm on an extensive array of datasets, including a Twitter dataset with multiple guarded attributes, the BiasBios dataset, and the BiasBench benchmark. The latter benchmark includes four datasets with various types of protected attributes. Our results demonstrate that bias can often be removed in our setup. We also discuss the limitations of our approach when there is a strong entanglement between the main task and the information to be erased.1
Shun Shao, Yftah Ziser, Shay B. Cohen
Trans. Assoc. Comput. Linguistics2
2022 Factorizing Content and Budget Decisions in Abstractive Summarization of Long Documents
abstract
We argue that disentangling content selection from the budget used to cover salient content improves the performance and applicability of abstractive summarizers.Our method, FAC-TORSUM 1 , does this disentanglement by factorizing summarization into two steps through an energy function: (1) generation of abstractive summary views covering salient information in subsets of the input document (document views) ; (2) combination of these views into a final summary, following a budget and content guidance.This guidance may come from different sources, including from an advisor model such as BART or BigBird, or in oracle modefrom the reference.This factorization achieves significantly higher ROUGE scores on multiple benchmarks for long document summarization, namely PubMed, arXiv, and GovReport.Notably, our model is effective for domain adaptation.When trained only on PubMed, it achieves a 46.29 ROUGE-1 score on arXiv, outperforming PEGASUS trained in domain by a large margin.Our experimental results indicate that the performance gains are due to more flexible budget adaptation and processing of shorter contexts provided by partial document views.
Marcio Fonseca, Yftah Ziser, Shay B. Cohen
EMNLP2
2021 DILBERT: Customized Pre-Training for Domain Adaptation with Category Shift, with an Application to Aspect Extraction
abstract
The rise of pre-trained language models has yielded substantial progress in the vast majority of Natural Language Processing (NLP) tasks.However, a generic approach towards the pre-training procedure can naturally be sub-optimal in some cases.Particularly, finetuning a pre-trained language model on a source domain and then applying it to a different target domain, results in a sharp performance decline of the eventual classifier for many source-target domain pairs.Moreover, in some NLP tasks, the output categories substantially differ between domains, making adaptation even more challenging.This, for example, happens in the task of aspect extraction, where the aspects of interest of reviews of, e.g., restaurants or electronic devices may be very different.This paper presents a new fine-tuning scheme for BERT, which aims to address the above challenges.We name this scheme DILBERT: Domain Invariant Learning with BERT, and customize it for aspect extraction in the unsupervised domain adaptation setting.DILBERT harnesses the categorical information of both the source and the target domains to guide the pre-training process towards a more domain and category invariant representation, thus closing the gap between the domains.We show that DILBERT yields substantial improvements over state-ofthe-art baselines while using a fraction of the unlabeled data, particularly in more challenging domain adaptation setups. 1
Entony Lekhtman, Yftah Ziser, Roi Reichart
EMNLP (1)2
2021 Answering Product-Questions by Utilizing Questions from Other Contextually Similar Products
abstract
Ohad Rozen, David Carmel, Avihai Mejer, Vitaly Mirkis, Yftah Ziser. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Ohad Rozen, David Carmel, Avihai Mejer, Vitaly Mirkis, Yftah Ziser
NAACL-HLT5
2020 Humor Detection in Product Question Answering Systems
abstract
Community question-answering (CQA) has been established as a prominent web service enabling users to post questions and get answers from the community. Product Question Answering (PQA) is a special CQA framework where questions are asked (and are answered) in the context of a specific product. Naturally, humorous questions are integral part of such platforms, especially as some products attract humor due to their unreasonable price, their peculiar functionality, or in cases that users emphasize their critical point-of-view through humor. Detecting humorous questions in such systems is important for sellers, to better understand user engagement with their products. It is also important to signal users about flippancy of humorous questions, and that answers for such questions should be taken with a grain of salt.
Yftah Ziser, Elad Kravi, David Carmel
SIGIR1
2019 Task Refinement Learning for Improved Accuracy and Stability of Unsupervised Domain Adaptation
abstract
Pivot Based Language Modeling (PBLM) (Ziser and Reichart, 2018a), combining LSTMs with pivot-based methods, has yielded significant progress in unsupervised domain adaptation. However, this approach is still challenged by the large pivot detection problem that should be solved, and by the inherent instability of LSTMs. In this paper we propose a Task Refinement Learning (TRL) approach, in order to solve these problems. Our algorithms iteratively train the PBLM model, gradually increasing the information exposed about each pivot. TRL-PBLM achieves stateof- the-art accuracy in six domain adaptation setups for sentiment classification. Moreover, it is much more stable than plain PBLM across model configurations, making the model much better fitted for practical use.
Yftah Ziser, Roi Reichart
ACL (1)1
2018 Deep Pivot-Based Modeling for Cross-language Cross-domain Transfer with Minimal Guidance
abstract
While cross-domain and cross-language transfer have long been prominent topics in NLP research, their combination has hardly been explored.In this work we consider this problem, and propose a framework that builds on pivotbased learning, structure-aware Deep Neural Networks (particularly LSTMs and CNNs) and bilingual word embeddings, with the goal of training a model on labeled data from one (language, domain) pair so that it can be effectively applied to another (language, domain) pair.We consider two setups, differing with respect to the unlabeled data available for model training.In the full setup the model has access to unlabeled data from both pairs, while in the lazy setup, which is more realistic for truly resource-poor languages, unlabeled data is available for both domains but only for the source language.We design our model for the lazy setup so that for a given target domain, it can train once on the source language and then be applied to any target language without re-training.In experiments with nine English-German and nine English-French domain pairs our best model substantially outperforms previous models even when it is trained in the lazy setup and previous models are trained in the full setup.1
Yftah Ziser, Roi Reichart
EMNLP1
2018 Pivot Based Language Modeling for Improved Neural Domain Adaptation
abstract
Representation learning with pivot-based methods and with Neural Networks (NNs) have lead to significant progress in domain adaptation for Natural Language Processing.However, most previous work that follows these approaches does not explicitly exploit the structure of the input text, and its output is most often a single representation vector for the entire text.In this paper we present the Pivot Based Language Model (PBLM), a representation learning model that marries together pivot-based and NN modeling in a structure aware manner.Particularly, our model processes the information in the text with a sequential NN (LSTM) and its output consists of a context-dependent representation vector for every input word.Unlike most previous representation learning models in domain adaptation, PBLM can naturally feed structure aware text classifiers such as LSTM and CNN.We experiment with the task of cross-domain sentiment classification on 20 domain pairs and show substantial improvements over strong baselines.1
Yftah Ziser, Roi Reichart
NAACL-HLT1
2017 Neural Structural Correspondence Learning for Domain Adaptation
abstract
We introduce a neural network model that marries together ideas from two prominent strands of research on domain adaptation through representation learning: structural correspondence learning (SCL, (Blitzer et al., 2006)) and autoencoder neural networks (NNs).Our model is a three-layer NN that learns to encode the non-pivot features of an input example into a lowdimensional representation, so that the existence of pivot features (features that are prominent in both domains and convey useful information for the NLP task) in the example can be decoded from that representation.The low-dimensional representation is then employed in a learning algorithm for the task.Moreover, we show how to inject pre-trained word embeddings into our model in order to improve generalization across examples with similar pivot features.We experiment with the task of cross-domain sentiment classification on 16 domain pairs and show substantial improvements over strong baselines.1
Yftah Ziser, Roi Reichart
CoNLL1