VLDB 2026 Research / reviewers in the wild / expert
Shauli Ravfogel
dblp:227/2231
· DBLP profile ↗
24ranked-venue papers
7as first author
23since 2021 · last 2026
0000-0001-8442-9311ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 7 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Benchmarks to Skills: Low-Rank Factors for LLM EvaluationabstractAbstract Current evaluations of large language models (LLMs) rely heavily on a growing collection of benchmarks and on aggregate benchmark scores, yet it remains unclear what this comparison actually captures, and what these scores reveal about models’ underlying capabilities. Here, we propose a new paradigm for LLM evaluation, by asking whether benchmark performance reflects many independent abilities, or rather, relies on a small number of shared dimensions. To answer this, we apply Factor Analysis (FA) to a massive performance matrix of LLMs versus benchmarks (60 × 44) revealing an intrinsically low-rank structure of that matrix. That is, a small number of latent factors captures most of the structure in the full task space. This low-rank geometry reveals substantial redundancy across existing tasks and explains why many benchmarks appear to be measuring overlapping abilities. We further show that these latent factors correspond to coherent, skill-like, dimensions of LLM behavior. Leveraging this latent skill-space, we deliver three practical tools for LLM evaluation and downstream users: (i) identifying redundant tasks, (ii) profiling new models using a small subset of tasks, and (iii) selecting models aligned with desired skill profiles. Our method provides a solid alternative to the de-facto standard of a single aggregate score, and establishes an interpretable and practical framework for understanding and benchmarking LLM core capabilities. Aviya Maimon, Amir David Nisan Cohen, Gal Vishne, Shauli Ravfogel, Reut Tsarfaty |
Trans. Assoc. Comput. Linguistics | 4 |
| 2026 | RELIC: Evaluating Complex Reasoning via the Recognition of Languages In-ContextabstractAbstract Large language models (LLMs) are increasingly used to solve complex tasks where they must retrieve and compose many pieces of in-context information in long reasoning chains. For many real-world tasks it is hard to accurately gauge how model performance and strategy change as task complexity grows. To evaluate models’ complex reasoning capability in a scalable and verifiable way, we introduce RELIC (Recognition of Languages In-Context), a framework that evaluates an LLM’s ability to decide whether a given string belongs to the context-free language (CFL) generated by a grammar presented in-context. CFL recognition allows us to modulate the intrinsic complexity of the problem by varying grammar size and string length and translate this asymptotic complexity into predictions for ideal LLM performance. We find that even the most advanced reasoning models perform poorly on RELIC, not only failing to appropriately scale their inference compute to keep pace with task difficulty, but even reducing the number of reasoning tokens they use as task complexity increases. We find that these decreases in compute accompany changes in reasoning strategy, as models move from identifying and implementing algorithmic solutions to guessing. For models whose full completions go uninspected, this manifests as “quiet quitting” on hard tasks. Code: https://jpetty.org/relic Jackson Petty, Michael Y. Hu, Shauli Ravfogel, William Merrill, Tal Linzen |
Trans. Assoc. Comput. Linguistics | 4 |
| 2025 | The Medium Is Not the Message: Deconfounding Document Embeddings via Linear Concept ErasureabstractEmbedding-based similarity metrics between text sequences can be influenced not just by the content dimensions we most care about, but can also be biased by spurious attributes like the text's source or language.These document confounders cause problems for many applications, but especially those that need to pool texts from different corpora.This paper shows that a debiasing algorithm that removes information about observed confounders from the encoder representations substantially reduces these biases at a minimal computational cost.Document similarity and clustering metrics improve across every embedding variant and task we evaluate-often dramatically.Interestingly, performance on out-of-distribution benchmarks is not impacted, indicating that the embeddings are not otherwise degraded.1 Yu Fan 0007, Shauli Ravfogel, Mrinmaya Sachan, Elliott Ash, Alexander Miserlis Hoyle |
EMNLP | 3 |
| 2025 | Intrinsic Test of Unlearning Using Parametric Knowledge TracesabstractThe task of "unlearning" certain concepts in large language models (LLMs) has gained attention for its role in mitigating harmful, private, or incorrect outputs.Current evaluations mostly rely on behavioral tests, without monitoring residual knowledge in model parameters, which can be adversarially exploited to recover erased information.We argue that unlearning should also be assessed internally by tracking changes in the parametric traces of unlearned concepts.To this end, we propose a general evaluation methodology that uses vocabulary projections to inspect concepts encoded in model parameters.We apply this approach to localize "concept vectors" -parameter vectors encoding concrete conceptsand construct CONCEPTVECTORS, a benchmark of hundreds of such concepts and their parametric traces in two open-source LLMs.Evaluation on CONCEPTVECTORS shows that existing methods minimally alter concept vectors, mostly suppressing them at inference time, while direct ablation of these vectors removes the associated knowledge and reduces adversarial susceptibility.Our findings reveal limitations of behavior-only evaluations and advocate for parameter-based assessments.We release our code and benchmark at https://github. com/yihuaihong/ConceptVectors.* Bias term is omitted for brevity. Yihuai Hong, Haiqin Yang, Shauli Ravfogel, Mor Geva |
EMNLP | 4 |
| 2025 | Gumbel Counterfactual Generation From Language ModelsabstractUnderstanding and manipulating the causal generation mechanisms in language models is essential for controlling their behavior. Previous work has primarily relied on techniques such as representation surgery---e.g., model ablations or manipulation of linear subspaces tied to specific concepts---to intervene on these models. To understand the impact of interventions precisely, it is useful to examine counterfactuals---e.g., how a given sentence would have appeared had it been generated by the model following a specific intervention. We highlight that counterfactual reasoning is conceptually distinct from interventions, as articulated in Pearl's causal hierarchy. Based on this observation, we propose a framework for generating true string counterfactuals by reformulating language models as a structural equation model using the Gumbel-max trick, which we called Gumbel counterfactual generation.
This reformulation allows us to model the joint distribution over original strings and their counterfactuals resulting from the same instantiation of the sampling noise. We develop an algorithm based on hindsight Gumbel sampling that allows us to infer the latent noise variables and generate counterfactuals of observed strings. Our experiments demonstrate that the approach produces meaningful counterfactuals while at the same time showing that commonly used intervention techniques have considerable undesired side effects. Shauli Ravfogel, Anej Svete, Vésteinn Snæbjarnarson, Ryan Cotterell |
ICLR | 1 |
| 2025 | Preserving Task-Relevant Information Under Linear Concept RemovalabstractModern neural networks often encode unwanted concepts alongside task-relevant information, leading to fairness and interpretability concerns. Existing post-hoc approaches can remove undesired concepts but often degrade useful signals. We introduce SPLINCE—Simultaneous Projection for LINear concept removal and Covariance prEservation—which eliminates sensitive concepts from representations while exactly preserving their covariance with a target label. SPLINCE achieves this via an oblique projection that ``splices out'' the unwanted direction yet protects important label correlations. Theoretically, it is the unique solution that removes linear concept predictability and maintains target covariance with minimal embedding distortion. Empirically, SPLINCE outperforms baselines on benchmarks such as Bias in Bios and Winobias, removing protected attributes while minimally damaging main-task information. Floris Holstege, Shauli Ravfogel, Bram Wouters |
NeurIPS | 2 |
| 2025 | Emergence of Linear Truth Encodings in Language ModelsabstractRecent probing studies reveal that large language models exhibit linear subspaces that separate true from false statements, yet the mechanism behind their emergence is unclear. We introduce a transparent, one-layer transformer toy model that reproduces such truth subspaces end-to-end and exposes one concrete route by which they can arise. We study one simple setting in which truth encoding can emerge: a data distribution where factual statements co-occur with other factual statements (and vice-versa), encouraging the model to learn this distinction in order to lower the LM loss on future tokens. We corroborate this pattern with experiments in pretrained language models. Finally, in the toy setting we observe a two-phase learning dynamic: networks first memorize individual factual associations in a few steps, then---over a longer horizon---learn to linearly separate true from false, which in turn lowers language-modeling loss. Together, these results provide both a mechanistic demonstration and an empirical motivation for how and why linear truth representations can emerge in language models. Shauli Ravfogel, Gilad Yehudai, Tal Linzen, Joan Bruna, Alberto Bietti |
NeurIPS | 1 |
| 2024 | Language Concept Erasure for Language-invariant Dense RetrievalabstractMultilingual models aim for language-invariant representations but still prominently encode language identity.This, along with the scarcity of high-quality parallel retrieval data, limits their performance in retrieval.We introduce LANCER, a multi-task learning framework that improves language-invariant dense retrieval by reducing language-specific signals in the embedding space.Leveraging the notion of linear concept erasure, we design a loss function that penalizes cross-correlation between representations and their language labels.LANCER leverages only English retrieval data and general multilingual corpora, training models to focus on language-invariant retrieval by semantic similarity without necessitating a vast parallel corpus.Experimental results on various datasets show our method consistently improves over baselines, with extensive analyses demonstrating greater language agnosticism. Zhiqi Huang 0002, Puxuan Yu, Shauli Ravfogel, James Allan 0001 |
EMNLP | 3 |
| 2024 | Representation Surgery: Theory and Practice of Affine SteeringabstractLanguage models often exhibit undesirable behavior, e.g., generating toxic or gender-biased text. In the case of neural language models, an encoding of the undesirable behavior is often present in the model’s representations. Thus, one natural (and common) approach to prevent the model from exhibiting undesirable behavior is to steer the model’s representations in a manner that reduces the probability of it generating undesirable text. This paper investigates the formal and empirical properties of steering functions, i.e., transformation of the neural language model’s representations that alter its behavior. First, we derive two optimal, in the least-squares sense, affine steering functions under different constraints. Our theory provides justification for existing approaches and offers a novel, improved steering approach. Second, we offer a series of experiments that demonstrate the empirical effectiveness of the methods in mitigating bias and reducing toxic generation. Shashwat Singh 0001, Shauli Ravfogel, Jonathan Herzig, Roee Aharoni, Ryan Cotterell, Ponnurangam Kumaraguru |
ICML | 2 |
| 2024 | On Affine Homotopy between Language EncodersabstractPre-trained language encoders---functions that represent text as vectors---are an integral component of many NLP tasks.
We tackle a natural question in language encoder analysis: What does it mean for two encoders to be similar?
We contend that a faithful measure of similarity needs to be \emph{intrinsic}, that is, task-independent, yet still be informative of \emph{extrinsic} similarity---the performance on downstream tasks.
It is common to consider two encoders similar if they are \emph{homotopic}, i.e., if they can be aligned through some transformation.
In this spirit, we study the properties of \emph{affine} alignment of language encoders and its implications on extrinsic similarity.
We find that while affine alignment is fundamentally an asymmetric notion of similarity, it is still informative of extrinsic similarity.
We confirm this on datasets of natural language representations.
Beyond providing useful bounds on extrinsic similarity, affine intrinsic similarity also allows us to begin uncovering the structure of the space of pre-trained encoders by defining an order over them. Robin Shing Moon Chan, Reda Boumasmoud, Anej Svete, Qipeng Guo, Zhijing Jin 0001, Shauli Ravfogel, Mrinmaya Sachan, Bernhard Schölkopf, Mennatallah El-Assady, Ryan Cotterell |
NeurIPS | 7 |
| 2023 | Linear Guardedness and its ImplicationsabstractMethods for erasing human-interpretable concepts from neural representations that assume linearity have been found to be tractable and useful.However, the impact of this removal on the behavior of downstream classifiers trained on the modified representations is not fully understood.In this work, we formally define the notion of log-linear guardedness as the inability of an adversary to predict the concept directly from the representation, and study its implications.We show that, in the binary case, under certain assumptions, a downstream log-linear model cannot recover the erased concept.However, we demonstrate that a multiclass log-linear model can be constructed that indirectly recovers the concept in some cases, pointing to the inherent limitations of log-linear guardedness as a downstream bias mitigation technique.These findings shed light on the theoretical limitations of linear erasure methods and highlight the need for further research on the connections between intrinsic and extrinsic bias in neural models. Shauli Ravfogel, Yoav Goldberg, Ryan Cotterell |
ACL (1) | 1 |
| 2023 | The Curious Case of Hallucinatory (Un)answerability: Finding Truths in the Hidden States of Over-Confident Large Language ModelsabstractLarge language models (LLMs) have been shown to possess impressive capabilities, while also raising crucial concerns about the faithfulness of their responses.A primary issue arising in this context is the management of (un)answerable queries by LLMs, which often results in hallucinatory behavior due to overconfidence.In this paper, we explore the behavior of LLMs when presented with (un)answerable queries.We ask: do models represent the fact that the question is (un)answerable when generating a hallucinatory answer?Our results show strong indications that such models encode the answerability of an input query, with the representation of the first decoded token often being a strong indicator.These findings shed new light on the spatial organization within the latent representations of LLMs, unveiling previously unexplored facets of these models.Moreover, they pave the way for the development of improved decoding techniques with better adherence to factual generation, particularly in scenarios where query (un)answerability is a concern. 1 Aviv Slobodkin, Omer Goldman, Avi Caciularu, Ido Dagan, Shauli Ravfogel |
EMNLP | 5 |
| 2023 | LEACE: Perfect linear concept erasure in closed formabstractConcept erasure aims to remove specified features from a representation. It can improve fairness (e.g. preventing a classifier from using gender or race) and interpretability (e.g. removing a concept to observe changes in model behavior). We introduce LEAst-squares Concept Erasure (LEACE), a closed-form method which provably prevents all linear classifiers from detecting a concept while changing the representation as little as possible, as measured by a broad class of norms. We apply LEACE to large language models with a novel procedure called concept scrubbing, which erases target concept information from _every_ layer in the network. We demonstrate our method on two tasks: measuring the reliance of language models on part-of-speech information, and reducing gender bias in BERT embeddings. Our code is available at https://github.com/EleutherAI/concept-erasure. Nora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell, Edward Raff, Stella Biderman |
NeurIPS | 3 |
| 2023 | Linguistic Binding in Diffusion Models: Enhancing Attribute Correspondence through Attention Map AlignmentabstractText-conditioned image generation models often generate incorrect associations between entities and their visual attributes. This reflects an impaired mapping between linguistic binding of entities and modifiers in the prompt and visual binding of the corresponding elements in the generated image. As one example, a query like ``a pink sunflower and a yellow flamingo'' may incorrectly produce an image of a yellow sunflower and a pink flamingo. To remedy this issue, we propose SynGen, an approach which first syntactically analyses the prompt to identify entities and their modifiers, and then uses a novel loss function that encourages the cross-attention maps to agree with the linguistic binding reflected by the syntax. Specifically, we encourage large overlap between attention maps of entities and their modifiers, and small overlap with other entities and modifier words. The loss is optimized during inference, without retraining or fine-tuning the model. Human evaluation on three datasets, including one new and challenging set, demonstrate significant improvements of SynGen compared with current state of the art methods. This work highlights how making use of sentence structure during inference can efficiently and substantially improve the faithfulness of text-to-image generation. Royi Rassin, Eran Hirsch, Daniel Glickman, Shauli Ravfogel, Yoav Goldberg, Gal Chechik |
NeurIPS | 4 |
| 2023 | Visual Comparison of Language Model AdaptationabstractNeural language models are widely used; however, their model parameters often need to be adapted to the specific domains and tasks of an application, which is time- and resource-consuming. Thus, adapters have recently been introduced as a lightweight alternative for model adaptation. They consist of a small set of task-specific parameters with a reduced training time and simple parameter composition. The simplicity of adapter training and composition comes along with new challenges, such as maintaining an overview of adapter properties and effectively comparing their produced embedding spaces. To help developers overcome these challenges, we provide a twofold contribution. First, in close collaboration with NLP researchers, we conducted a requirement analysis for an approach supporting adapter evaluation and detected, among others, the need for both intrinsic (i.e., embedding similarity-based) and extrinsic (i.e., prediction-based) explanation methods. Second, motivated by the gathered requirements, we designed a flexible visual analytics workspace that enables the comparison of adapter properties. In this paper, we discuss several design iterations and alternatives for interactive, comparative visual explanation methods. Our comparative visualizations show the differences in the adapted embedding vectors and prediction outcomes for diverse human-interpretable concepts (e.g., person names, human qualities). We evaluate our workspace through case studies and show that, for instance, an adapter trained on the language debiasing task according to context-0 (decontextualized) embeddings introduces a new type of bias where words (even gender-independent words such as countries) become more similar to female- than male pronouns. We demonstrate that these are artifacts of context-0 embeddings, and the adapter effectively eliminates the gender information from the contextualized word representations. Rita Sevastjanova, Eren Cakmak, Shauli Ravfogel, Ryan Cotterell, Mennatallah El-Assady |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2022 | Adversarial Concept Erasure in Kernel SpaceabstractThe representation space of neural models for textual data emerges in an unsupervised manner during training.Understanding how those representations encode human-interpretable concepts is a fundamental problem.One prominent approach for the identification of concepts in neural representations is searching for a linear subspace whose erasure prevents the prediction of the concept from the representations.However, while many linear erasure algorithms are tractable and interpretable, neural networks do not necessarily represent concepts in a linear manner.To identify non-linearly encoded concepts, we propose a kernelization of a linear minimax game for concept erasure.We demonstrate that it is possible to prevent specific nonlinear adversaries from predicting the concept.However, the protection does not transfer to different nonlinear adversaries.Therefore, exhaustively erasing a non-linearly encoded concept remains an open problem. Shauli Ravfogel, Francisco Vargas 0001, Yoav Goldberg, Ryan Cotterell |
EMNLP | 1 |
| 2022 | Linear Adversarial Concept ErasureabstractModern neural models trained on textual data rely on pre-trained representations that emerge without direct supervision. As these representations are increasingly being used in real-world applications, the inability to control their content becomes an increasingly important problem. In this work, we formulate the problem of identifying a linear subspace that corresponds to a given concept, and removing it from the representation. We formulate this problem as a constrained, linear minimax game, and show that existing solutions are generally not optimal for this task. We derive a closed-form solution for certain objectives, and propose a convex relaxation that works well for others. When evaluated in the context of binary gender removal, the method recovers a low-dimensional subspace whose removal mitigates bias by intrinsic and extrinsic evaluation. Surprisingly, we show that the method—despite being linear—is highly expressive, effectively mitigating bias in the output layers of deep, nonlinear classifiers while maintaining tractability and interpretability. Shauli Ravfogel, Michael Twiton, Yoav Goldberg, Ryan Cotterell |
ICML | 1 |
| 2021 | Counterfactual Interventions Reveal the Causal Effect of Relative Clause Representations on Agreement PredictionabstractWhen language models process syntactically complex sentences, do they use their representations of syntax in a manner that is consistent with the grammar of the language?We propose AlterRep, an intervention-based method to address this question.For any linguistic feature of a given sentence, AlterRep generates counterfactual representations by altering how the feature is encoded, while leaving intact all other aspects of the original representation.By measuring the change in a model's word prediction behavior when these counterfactual representations are substituted for the original ones, we can draw conclusions about the causal effect of the linguistic feature in question on the model's behavior.We apply this method to study how BERT models of different sizes process relative clauses (RCs).We find that BERT variants use RC boundary information during word prediction in a manner that is consistent with the rules of English grammar; this RC boundary information generalizes to a considerable extent across different RC types, suggesting that BERT represents RCs as an abstract linguistic category. Shauli Ravfogel, Grusha Prasad, Tal Linzen, Yoav Goldberg |
CoNLL | 1 |
| 2021 | Contrastive Explanations for Model InterpretabilityabstractContrastive explanations clarify why an event occurred in contrast to another.They are inherently intuitive to humans to both produce and comprehend.We propose a method to produce contrastive explanations in the latent space, via a projection of the input representation, such that only the features that differentiate two potential decisions are captured.Our modification allows model behavior to consider only contrastive reasoning, and uncover which aspects of the input are useful for and against particular decisions.Additionally, for a given input feature, our contrastive explanations can answer for which label, and against which alternative label, is the feature useful.We produce contrastive explanations via both highlevel abstract concept attribution and low-level input token/span attribution for two NLP classification benchmarks.Our findings demonstrate the ability of label-contrastive explanations to provide fine-grained interpretability of model decisions.1 Alon Jacovi, Swabha Swayamdipta, Shauli Ravfogel, Yanai Elazar, Yejin Choi 0001, Yoav Goldberg |
EMNLP (1) | 3 |
| 2021 | Ab Antiquo: Neural Proto-language ReconstructionabstractHistorical linguists have identified regularities in the process of historic sound change.The comparative method utilizes those regularities to reconstruct proto-words based on observed forms in daughter languages.Can this process be efficiently automated?We address the task of proto-word reconstruction, in which the model is exposed to cognates in contemporary daughter languages, and has to predict the proto word in the ancestor language.We provide a novel dataset for this task, encompassing over 8,000 comparative entries, and show that neural sequence models outperform conventional methods applied to this task so far.Error analysis reveals a variability in the ability of neural model to capture different phonological changes, correlating with the complexity of the changes.Analysis of learned embeddings reveals the models learn phonologically meaningful generalizations, corresponding to well-attested phonological shifts documented by historical linguistics. Carlo Meloni, Shauli Ravfogel, Yoav Goldberg |
NAACL-HLT | 2 |
| 2021 | Measuring and Improving Consistency in Pretrained Language ModelsabstractAbstract Consistency of a model—that is, the invariance of its behavior under meaning-preserving alternations in its input—is a highly desirable property in natural language processing. In this paper we study the question: Are Pretrained Language Models (PLMs) consistent with respect to factual knowledge? To this end, we create ParaRel🤘, a high-quality resource of cloze-style query English paraphrases. It contains a total of 328 paraphrases for 38 relations. Using ParaRel🤘, we show that the consistency of all PLMs we experiment with is poor— though with high variance between relations. Our analysis of the representational spaces of PLMs suggests that they have a poor structure and are currently not suitable for representing knowledge robustly. Finally, we propose a method for improving model consistency and experimentally demonstrate its effectiveness.1 Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard H. Hovy, Hinrich Schütze, Yoav Goldberg |
Trans. Assoc. Comput. Linguistics | 3 |
| 2021 | Erratum: Measuring and Improving Consistency in Pretrained Language ModelsabstractAbstract During production of this paper, an error was introduced to the formula on the bottom of the right column of page 1020. In the last two terms of the formula, the n and m subscripts were swapped. The correct formula is:Lc=∑n=1k∑m=n+1kDKL(Qnri∥Qmri)+DKL(Qmri∥Qnri)The paper has been updated. Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard H. Hovy, Hinrich Schütze, Yoav Goldberg |
Trans. Assoc. Comput. Linguistics | 3 |
| 2021 | Amnesic Probing: Behavioral Explanation With Amnesic CounterfactualsabstractAbstract A growing body of work makes use of probing in order to investigate the working of neural models, often considered black boxes. Recently, an ongoing debate emerged surrounding the limitations of the probing paradigm. In this work, we point out the inability to infer behavioral conclusions from probing results, and offer an alternative method that focuses on how the information is being used, rather than on what information is encoded. Our method, Amnesic Probing, follows the intuition that the utility of a property for a given task can be assessed by measuring the influence of a causal intervention that removes it from the representation. Equipped with this new analysis tool, we can ask questions that were not possible before, for example, is part-of-speech information important for word prediction? We perform a series of analyses on BERT to answer these types of questions. Our findings demonstrate that conventional probing performance is not correlated to task importance, and we call for increased scrutiny of claims that draw behavioral or causal conclusions from probing results.1 Yanai Elazar, Shauli Ravfogel, Alon Jacovi, Yoav Goldberg |
Trans. Assoc. Comput. Linguistics | 2 |
| 2020 | Null It Out: Guarding Protected Attributes by Iterative Nullspace ProjectionabstractThe ability to control for the kinds of information encoded in neural representation has a variety of use cases, especially in light of the challenge of interpreting these models.We present Iterative Null-space Projection (INLP), a novel method for removing information from neural representations.Our method is based on repeated training of linear classifiers that predict a certain property we aim to remove, followed by projection of the representations on their null-space.By doing so, the classifiers become oblivious to that target property, making it hard to linearly separate the data according to it.While applicable for multiple uses, we evaluate our method on bias and fairness use-cases, and show that our method is able to mitigate bias in word embeddings, as well as to increase fairness in a setting of multi-class classification. Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, Yoav Goldberg |
ACL | 1 |