Noam Slonim

dblp:62/7001 · DBLP profile ↗
← Back
53ranked-venue papers
10as first author
15since 2021 · last 2025
0000-0001-5171-8264ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 48 · 8 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author
YearPublicationVenuePosition
2025 Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation
abstract
We introduce Debate Speech Evaluation as a novel and challenging benchmark for assessing LLM judges.Evaluating debate speeches requires a deep understanding of the speech at multiple levels, including argument strength and relevance, the coherence and organization of the speech, the appropriateness of its style and tone, and so on.This task involves a unique set of cognitive abilities that previously received limited attention in systematic LLM benchmarking.To explore such skills, we leverage a dataset of over 600 meticulously annotated debate speeches and present the first in-depth analysis of how state-of-the-art LLMs compare to human judges on this task.Our findings reveal a nuanced picture: while larger models can approximate individual human judgments in some respects, they differ substantially in their overall judgment behavior.We also investigate the ability of frontier LLMs to generate persuasive, opinionated speeches, showing that models may perform at a human level on this task.
Noy Sternlicht, Ariel Gera, Roy Bar-Haim, Tom Hope, Noam Slonim
EMNLP5
2024 Efficient Benchmarking (of Language Models)
abstract
Yotam Perlitz, Elron Bandel, Ariel Gera, Ofir Arviv, Liat Ein-Dor, Eyal Shnarch, Noam Slonim, Michal Shmueli-Scheuer, Leshem Choshen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yotam Perlitz, Elron Bandel, Ariel Gera, Ofir Arviv, Liat Ein-Dor, Eyal Shnarch, Noam Slonim, Michal Shmueli-Scheuer, Leshem Choshen
NAACL-HLT7
2023 nBIIG: A Neural BI Insights Generation System for Table Reporting
Yotam Perlitz, Dafna Sheinwald, Noam Slonim, Michal Shmueli-Scheuer
AAAI3
2023 ColD Fusion: Collaborative Descent for Distributed Multitask Finetuning
abstract
We propose a new paradigm to continually evolve pretrained models, denoted ColD Fusion.It provides the benefits of multitask learning but leverages distributed computation with limited communication and eliminates the need for shared data.Consequentially, ColD Fusion can give rise to a synergistic loop, where finetuned models can be recycled to continually improve the pretrained model they are based upon.We show that ColD Fusion yields comparable benefits to multitask training by producing a model that (a) attains strong performance on all of the datasets it was trained on; and (b) is a better starting point for finetuning on unseen datasets.We show that ColD Fusion outperforms RoBERTa and even previous multitask models.Specifically, when training and testing on 35 diverse datasets, ColD Fusion-based model outperforms RoBERTa by 2.33 points on average without any changes to the architecture.1
Shachar Don-Yehiya, Elad Venezian, Colin Raffel, Noam Slonim, Leshem Choshen
ACL (1)4
2023 The Benefits of Bad Advice: Autocontrastive Decoding across Model Layers
abstract
Ariel Gera, Roni Friedman, Ofir Arviv, Chulaka Gunasekara, Benjamin Sznajder, Noam Slonim, Eyal Shnarch. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Ariel Gera, Roni Friedman, Ofir Arviv, R. Chulaka Gunasekara, Benjamin Sznajder, Noam Slonim, Eyal Shnarch
ACL (1)6
2023 Where to start? Analyzing the potential value of intermediate models
abstract
Previous studies observed that finetuned models may be better base models than the vanilla pretrained model.Such a model, finetuned on some source dataset, may provide a better starting point for a new finetuning process on a desired target dataset.Here, we perform a systematic analysis of this intertraining scheme, over a wide range of English classification tasks.Surprisingly, our analysis suggests that the potential intertraining gain can be analyzed independently for the target dataset under consideration, and for a base model being considered as a starting point.Hence, a performant model is generally strong, even if its training data was not aligned with the target dataset.Furthermore, we leverage our analysis to propose a practical and efficient approach to determine if and how to select a base model in real-world settings.Last, we release an updating ranking of best models in the HuggingFace hub per architecture.
Leshem Choshen, Elad Venezian, Shachar Don-Yehiya, Noam Slonim, Yoav Katz
EMNLP4
2023 Active Learning for Natural Language Generation
abstract
The field of Natural Language Generation (NLG) suffers from a severe shortage of labeled data due to the extremely expensive and timeconsuming process involved in manual annotation.A natural approach for coping with this problem is active learning (AL), a well-known machine learning technique for improving annotation efficiency by selectively choosing the most informative examples to label.However, while AL has been well-researched in the context of text classification, its application to NLG remains largely unexplored.In this paper, we present a first systematic study of active learning for NLG, considering a diverse set of tasks and multiple leading selection strategies, and harnessing a strong instruction-tuned model.Our results indicate that the performance of existing AL strategies is inconsistent, surpassing the baseline of random example selection in some cases but not in others.We highlight some notable differences between the classification and generation scenarios, and analyze the selection behaviors of existing AL strategies.Our findings motivate exploring novel approaches for applying AL to generation tasks.
Yotam Perlitz, Ariel Gera, Michal Shmueli-Scheuer, Dafna Sheinwald, Noam Slonim, Liat Ein-Dor
EMNLP5
2023 Efficient Methods for Natural Language Processing: A Survey
abstract
Abstract Recent work in natural language processing (NLP) has yielded appealing results from scaling model parameters and training data; however, using only scale to improve performance means that resource consumption also grows. Such resources include data, time, storage, or energy, all of which are naturally limited and unevenly distributed. This motivates research into efficient methods that require fewer resources to achieve similar results. This survey synthesizes and relates current methods and findings in efficient NLP. We aim to provide both guidance for conducting NLP under limited resources, and point towards promising research directions for developing more efficient methods.
Marcos V. Treviso, Ji-Ung Lee, Tianchu Ji, Betty van Aken, Manuel R. Ciosici, Michael Hassid, Kenneth Heafield, Sara Hooker, Colin Raffel, Pedro Henrique Martins, André F. T. Martins, Jessica Zosa Forde, Peter A. Milder, Edwin Simpson, Noam Slonim, Jesse Dodge, Emma Strubell, Niranjan Balasubramanian, Leon Derczynski, Iryna Gurevych, Roy Schwartz 0001
Trans. Assoc. Comput. Linguistics16
2022 Fortunately, Discourse Markers Can Enhance Language Models for Sentiment Analysis
abstract
In recent years, pretrained language models have revolutionized the NLP world, while achieving state of the art performance in various downstream tasks. However, in many cases, these models do not perform well when labeled data is scarce and the model is expected to perform in the zero or few shot setting. Recently, several works have shown that continual pretraining or performing a second phase of pretraining (inter-training) which is better aligned with the downstream task, can lead to improved results, especially in the scarce data setting. Here, we propose to leverage sentiment-carrying discourse markers to generate large-scale weakly-labeled data, which in turn can be used to adapt language models for sentiment analysis. Extensive experimental results show the value of our approach on various benchmark datasets, including the finance domain. Code, models and data are available at https://github.com/ibm/tslm-discourse-markers.
Liat Ein-Dor, Ilya Shnayderman, Artem Spector, Lena Dankin, Ranit Aharonov, Noam Slonim
AAAI6
2022 Quality Controlled Paraphrase Generation
abstract
Elron Bandel, Ranit Aharonov, Michal Shmueli-Scheuer, Ilya Shnayderman, Noam Slonim, Liat Ein-Dor. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Elron Bandel, Ranit Aharonov, Michal Shmueli-Scheuer, Ilya Shnayderman, Noam Slonim, Liat Ein-Dor
ACL (1)5
2022 Cluster & Tune: Boost Cold Start Performance in Text Classification
abstract
Eyal Shnarch, Ariel Gera, Alon Halfon, Lena Dankin, Leshem Choshen, Ranit Aharonov, Noam Slonim. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Eyal Shnarch, Ariel Gera, Alon Halfon, Lena Dankin, Leshem Choshen, Ranit Aharonov, Noam Slonim
ACL (1)7
2022 Zero-Shot Text Classification with Self-Training
abstract
Recent advances in large pretrained language models have increased attention to zero-shot text classification.In particular, models finetuned on natural language inference datasets have been widely adopted as zero-shot classifiers due to their promising results and offthe-shelf availability.However, the fact that such models are unfamiliar with the target task can lead to instability and performance issues.We propose a plug-and-play method to bridge this gap using a simple self-training approach, requiring only the class names along with an unlabeled dataset, and without the need for domain expertise or trial and error.We show that fine-tuning the zero-shot classifier on its most confident predictions leads to significant performance gains across a wide range of text classification tasks, presumably since self-training adapts the zero-shot model to the task at hand.
Ariel Gera, Alon Halfon, Eyal Shnarch, Yotam Perlitz, Liat Ein-Dor, Noam Slonim
EMNLP6
2022 Multi-Domain Targeted Sentiment Analysis
abstract
Orith Toledo-Ronen, Matan Orbach, Yoav Katz, Noam Slonim. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Orith Toledo-Ronen, Matan Orbach, Yoav Katz, Noam Slonim
NAACL-HLT4
2021 Every Bite Is an Experience: Key Point Analysis of Business Reviews
abstract
Roy Bar-Haim, Lilach Eden, Yoav Kantor, Roni Friedman, Noam Slonim. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Roy Bar-Haim, Lilach Eden, Yoav Kantor, Roni Friedman, Noam Slonim
ACL/IJCNLP (1)5
2021 YASO: A Targeted Sentiment Analysis Evaluation Dataset for Open-Domain Reviews
abstract
Current TSA evaluation in a cross-domain setup is restricted to the small set of review domains available in existing datasets.Such an evaluation is limited, and may not reflect true performance on sites like Amazon or Yelp that host diverse reviews from many domains.To address this gap, we present YASO -a new TSA evaluation dataset of open-domain user reviews.YASO contains 2215 English sentences from dozens of review domains, annotated with target terms and their sentiment.Our analysis verifies the reliability of these annotations, and explores the characteristics of the collected data.Benchmark results using five contemporary TSA systems show there is ample room for improvement on this challenging new dataset.YASO is available at github.com/IBM/yaso-tsa.
Matan Orbach, Orith Toledo-Ronen, Artem Spector, Ranit Aharonov, Yoav Katz, Noam Slonim
EMNLP (1)6
2020 Corpus Wide Argument Mining - A Working Solution
abstract
One of the main tasks in argument mining is the retrieval of argumentative content pertaining to a given topic. Most previous work addressed this task by retrieving a relatively small number of relevant documents as the initial source for such content. This line of research yielded moderate success, which is of limited use in a real-world system. Furthermore, for such a system to yield a comprehensive set of relevant arguments, over a wide range of topics, it requires leveraging a large and diverse corpus in an appropriate manner. Here we present a first end-to-end high-precision, corpus-wide argument mining system. This is made possible by combining sentence-level queries over an appropriate indexing of a very large corpus of newspaper articles, with an iterative annotation scheme. This scheme addresses the inherent label bias in the data and pinpoints the regions of the sample space whose manual labeling is required to obtain high-precision among top-ranked candidates.
Liat Ein-Dor, Eyal Shnarch, Lena Dankin, Alon Halfon, Benjamin Sznajder, Ariel Gera, Carlos Alzate, Martin Gleize, Leshem Choshen, Yufang Hou 0001, Yonatan Bilu, Ranit Aharonov, Noam Slonim
AAAI13
2020 A Large-Scale Dataset for Argument Quality Ranking: Construction and Analysis
abstract
Identifying the quality of free-text arguments has become an important task in the rapidly expanding field of computational argumentation. In this work, we explore the challenging task of argument quality ranking. To this end, we created a corpus of 30,497 arguments carefully annotated for point-wise quality, released as part of this work. To the best of our knowledge, this is the largest dataset annotated for point-wise argument quality, larger by a factor of five than previously released datasets. Moreover, we address the core issue of inducing a labeled score from crowd annotations by performing a comprehensive evaluation of different approaches to this problem. In addition, we analyze the quality dimensions that characterize this dataset. Finally, we present a neural method for argument quality ranking, which outperforms several baselines on our own dataset, as well as previous methods published for another dataset.
Shai Gretz, Roni Friedman, Edo Cohen-Karlik, Assaf Toledo, Dan Lahav, Ranit Aharonov, Noam Slonim
AAAI7
2020 From Arguments to Key Points: Towards Automatic Argument Summarization
abstract
Generating a concise summary from a large collection of arguments on a given topic is an intriguing yet understudied problem.We propose to represent such summaries as a small set of talking points, termed key points, each scored according to its salience.We show, by analyzing a large dataset of crowd-contributed arguments, that a small number of key points per topic is typically sufficient for covering the vast majority of the arguments.Furthermore, we found that a domain expert can often predict these key points in advance.We study the task of argument-to-key point mapping, and introduce a novel large-scale dataset for this task.We report empirical results for an extensive set of experiments with this dataset, showing promising performance.
Roy Bar-Haim, Lilach Eden, Roni Friedman, Yoav Kantor, Dan Lahav, Noam Slonim
ACL6
2020 Out of the Echo Chamber: Detecting Countering Debate Speeches
abstract
An educated and informed consumption of media content has become a challenge in modern times.With the shift from traditional news outlets to social media and similar venues, a major concern is that readers are becoming encapsulated in "echo chambers" and may fall prey to fake news and disinformation, lacking easy access to dissenting views.We suggest a novel task aiming to alleviate some of these concerns -that of detecting articles that most effectively counter the arguments -and not just the stance -made in a given text.We study this problem in the context of debate speeches.Given such a speech, we aim to identify, from among a set of speeches on the same topic and with an opposing stance, the ones that directly counter it.We provide a large dataset of 3685 such speeches (in English), annotated for this relation, which hopefully would be of general interest to the NLP community.We explore several algorithms addressing this task, and while some are successful, all fall short of expert human performance, suggesting room for further research.All data collected during this work is freely available for research 1 .
Matan Orbach, Yonatan Bilu, Assaf Toledo, Dan Lahav, Michal Jacovi, Ranit Aharonov, Noam Slonim
ACL7
2020 Quantitative argument summarization and beyond: Cross-domain key point analysis
abstract
When summarizing a collection of views, arguments or opinions on some topic, it is often desirable not only to extract the most salient points, but also to quantify their prevalence.Work on multi-document summarization has traditionally focused on creating textual summaries, which lack this quantitative aspect.Recent work has proposed to summarize arguments by mapping them to a small set of expert-generated key points, where the salience of each key point corresponds to the number of its matching arguments.The current work advances key point analysis in two important respects: first, we develop a method for automatic extraction of key points, which enables fully automatic analysis, and is shown to achieve performance comparable to a human expert.Second, we demonstrate that the applicability of key point analysis goes well beyond argumentation data.Using models trained on publicly available argumentation datasets, we achieve promising results in two additional domains: municipal surveys and user reviews.An additional contribution is an in-depth evaluation of argument-to-key point matching models, where we substantially outperform previous results.
Roy Bar-Haim, Yoav Kantor, Lilach Eden, Roni Friedman, Dan Lahav, Noam Slonim
EMNLP (1)6
2020 Active Learning for BERT: An Empirical Study
abstract
Liat Ein-Dor, Alon Halfon, Ariel Gera, Eyal Shnarch, Lena Dankin, Leshem Choshen, Marina Danilevsky, Ranit Aharonov, Yoav Katz, Noam Slonim. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Liat Ein-Dor, Alon Halfon, Ariel Gera, Eyal Shnarch, Lena Dankin, Leshem Choshen, Marina Danilevsky, Ranit Aharonov, Yoav Katz, Noam Slonim
EMNLP (1)10
2019 From Surrogacy to Adoption; From Bitcoin to Cryptocurrency: Debate Topic Expansion
abstract
Roy Bar-Haim, Dalia Krieger, Orith Toledo-Ronen, Lilach Edelstein, Yonatan Bilu, Alon Halfon, Yoav Katz, Amir Menczel, Ranit Aharonov, Noam Slonim. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Roy Bar-Haim, Dalia Krieger, Orith Toledo-Ronen, Lilach Edelstein, Yonatan Bilu, Alon Halfon, Yoav Katz, Amir Menczel, Ranit Aharonov, Noam Slonim
ACL (1)10
2019 Argument Invention from First Principles
abstract
Yonatan Bilu, Ariel Gera, Daniel Hershcovich, Benjamin Sznajder, Dan Lahav, Guy Moshkowich, Anael Malet, Assaf Gavron, Noam Slonim. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Yonatan Bilu, Ariel Gera, Daniel Hershcovich, Benjamin Sznajder, Dan Lahav, Guy Moshkowich, Anael Malet, Assaf Gavron, Noam Slonim
ACL (1)9
2019 Are You Convinced? Choosing the More Convincing Evidence with a Siamese Network
abstract
Machines capable of responding and interacting with humans in helpful ways have become ubiquitous.We now expect them to discuss with us the more delicate questions in our world, and they should do so armed with effective arguments.But what makes an argument more persuasive?What will convince you?In this paper, we present a new data set, IBM-EviConv, of pairs of evidence labeled for convincingness, designed to be more challenging than existing alternatives.We also propose a Siamese neural network architecture shown to outperform several baselines on both a prior convincingness data set and our own.Finally, we provide insights into our experimental results and the various kinds of argumentative value our method is capable of detecting.
Martin Gleize, Eyal Shnarch, Leshem Choshen, Lena Dankin, Guy Moshkowich, Ranit Aharonov, Noam Slonim
ACL (1)7
2019 A Dataset of General-Purpose Rebuttal
abstract
Matan Orbach, Yonatan Bilu, Ariel Gera, Yoav Kantor, Lena Dankin, Tamar Lavee, Lili Kotlerman, Shachar Mirkin, Michal Jacovi, Ranit Aharonov, Noam Slonim. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Matan Orbach, Yonatan Bilu, Ariel Gera, Yoav Kantor, Lena Dankin, Tamar Lavee, Lili Kotlerman, Shachar Mirkin, Michal Jacovi, Ranit Aharonov, Noam Slonim
EMNLP/IJCNLP (1)11
2019 Automatic Argument Quality Assessment - New Datasets and Methods
abstract
Assaf Toledo, Shai Gretz, Edo Cohen-Karlik, Roni Friedman, Elad Venezian, Dan Lahav, Michal Jacovi, Ranit Aharonov, Noam Slonim. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Assaf Toledo, Shai Gretz, Edo Cohen-Karlik, Roni Friedman, Elad Venezian, Dan Lahav, Michal Jacovi, Ranit Aharonov, Noam Slonim
EMNLP/IJCNLP (1)9
2018 Towards an argumentative content search engine using weak supervision
abstract
Searching for sentences containing claims in a large text corpus is a key component in developing an argumentative content search engine. Previous works focused on detecting claims in a small set of documents or within documents enriched with argumentative content. However, pinpointing relevant claims in massive unstructured corpora, received little attention. A step in this direction was taken in (Levy et al. 2017), where the authors suggested using a weak signal to develop a relatively strict query for claim–sentence detection. Here, we leverage this work to define weak signals for training DNNs to obtain significantly greater performance. This approach allows to relax the query and increase the potential coverage. Our results clearly indicate that the system is able to successfully generalize from the weak signal, outperforming previously reported results in terms of both precision and coverage. Finally, we adapt our system to solve a recent argument mining task of identifying argumentative sentences in Web texts retrieved from heterogeneous sources, and obtain F1 scores comparable to the supervised baseline.
Ran Levy 0001, Ben Bogin, Shai Gretz, Ranit Aharonov, Noam Slonim
COLING5
2018 Learning Sentiment Composition from Sentiment Lexicons
abstract
Sentiment composition is a fundamental sentiment analysis problem. Previous work relied on manual rules and manually-created lexical resources such as negator lists, or learned a composition function from sentiment-annotated phrases or sentences. We propose a new approach for learning sentiment composition from a large, unlabeled corpus, which only requires a word-level sentiment lexicon for supervision. We automatically generate large sentiment lexicons of bigrams and unigrams, from which we induce a set of lexicons for a variety of sentiment composition processes. The effectiveness of our approach is confirmed through manual annotation, as well as sentiment classification experiments with both phrase-level and sentence-level benchmarks.
Orith Toledo-Ronen, Roy Bar-Haim, Alon Halfon, Charles Jochim, Amir Menczel, Ranit Aharonov, Noam Slonim
COLING7
2018 Project Debater
abstract
Project Debater is the first AI system that was shown to debate humans in a meaningful manner in a full live debate. Developing this system started in 2012, as the next AI Grand Challenge pursued by IBM Research, following the demonstration of Deep Blue in Chess in 1997, and Watson in Jeopardy! In 2011. The Project Debater system was demonstrated for the first time in San Francisco in June 2018, in two full live debates vs. expert human debaters, and correspondingly received massive media attention. This talk will present the challenges in developing this system, its current capabilities and present limitations, as well as how we envision its future.
Noam Slonim
COMMA1
2018 Listening Comprehension over Argumentative Content
abstract
Shachar Mirkin, Guy Moshkowich, Matan Orbach, Lili Kotlerman, Yoav Kantor, Tamar Lavee, Michal Jacovi, Yonatan Bilu, Ranit Aharonov, Noam Slonim. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018.
Shachar Mirkin, Guy Moshkowich, Matan Orbach, Lili Kotlerman, Yoav Kantor, Tamar Lavee, Michal Jacovi, Yonatan Bilu, Ranit Aharonov, Noam Slonim
EMNLP10
2018 Learning Concept Abstractness Using Weak Supervision
abstract
Ella Rabinovich, Benjamin Sznajder, Artem Spector, Ilya Shnayderman, Ranit Aharonov, David Konopnicki, Noam Slonim. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018.
Ella Rabinovich, Benjamin Sznajder, Artem Spector, Ilya Shnayderman, Ranit Aharonov, David Konopnicki, Noam Slonim
EMNLP7
2018 Semantic Relatedness of Wikipedia Concepts - Benchmark Data and a Working Solution
Liat Ein-Dor, Alon Halfon, Yoav Kantor, Ran Levy 0001, Yosi Mass, Ruty Rinott, Eyal Shnarch, Noam Slonim
LREC8
2018 SLIDE - a Sentiment Lexicon of Common Idioms
Charles Jochim, Francesca Bonin, Roy Bar-Haim, Noam Slonim
LREC4
2018 A Recorded Debating Dataset
Shachar Mirkin, Michal Jacovi, Tamar Lavee, Hong-Kwang Jeff Kuo, Samuel Thomas 0001, Leslie Sager, Lili Kotlerman, Elad Venezian, Noam Slonim
LREC9
2017 Stance Classification of Context-Dependent Claims
abstract
Roy Bar-Haim, Indrajit Bhattacharya, Francesco Dinuzzo, Amrita Saha, Noam Slonim. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017.
Roy Bar-Haim, Indrajit Bhattacharya, Francesco Dinuzzo, Amrita Saha, Noam Slonim
EACL (1)5
2017 GRASP: Rich Patterns for Argumentation Mining
abstract
We present the GrASP algorithm for automatically extracting patterns that characterize subtle linguistic phenomena.To that end, GrASP augments each term of input text with multiple layers of linguistic information.These different facets of the text terms are systematically combined to reveal rich patterns.We report highly promising experimental results in several challenging text analysis tasks within the field of Argumentation Mining.We believe that GrASP is general enough to be useful for other domains too.
Eyal Shnarch, Ran Levy 0001, Vikas C. Raykar, Noam Slonim
EMNLP4
2015 Show Me Your Evidence - an Automatic Method for Context Dependent Evidence Detection
abstract
Engaging in a debate with oneself or others to take decisions is an integral part of our day-today life.A debate on a topic (say, use of performance enhancing drugs) typically proceeds by one party making an assertion/claim (say, PEDs are bad for health) and then providing an evidence to support the claim (say, a 2006 study shows that PEDs have psychiatric side effects).In this work, we propose the task of automatically detecting such evidences from unstructured text that support a given claim.This task has many practical applications in decision support and persuasion enhancement in a wide range of domains.We first introduce an extensive benchmark data set tailored for this task, which allows training statistical models and assessing their performance.Then, we suggest a system architecture based on supervised learning to address the evidence detection task.Finally, promising experimental results are reported.
Ruty Rinott, Lena Dankin, Carlos Alzate, Mitesh M. Khapra, Ehud Aharoni, Noam Slonim
EMNLP6
2014 Context Dependent Claim Detection
Ran Levy 0001, Yonatan Bilu, Daniel Hershcovich, Ehud Aharoni, Noam Slonim
COLING5
2013 Hartigan's K-Means Versus Lloyd's K-Means - Is It Time for a Change?
Noam Slonim, Ehud Aharoni, Koby Crammer
IJCAI1
2011 Active Online Classification via Information Maximization
abstract
We propose an online classification approach for co-occurrence data which is based on a simple information theoretic principle. We further show how to properly estimate the uncertainty associated with each prediction of our scheme and demonstrate how to exploit these uncertainty estimates. First, in order to abstain highly uncertain predictions. And second, within an active learning framework, in order to preserve classification accuracy while substantially reducing training set size. Our method is highly efficient in terms of run-time and memory footprint requirements. Experimental results in the domain of text classification demonstrate that the classification accuracy of our method is superior or comparable to other state-of-the-art online classification algorithms.
Noam Slonim, Elad Yom-Tov, Koby Crammer
IJCAI1
2010 Predicting Customer Churn in Mobile Networks through Analysis of Social Groups
abstract
Churn prediction aims to identify subscribers who are about to transfer their business to a competitor. Since the cost associated with customer acquisition is much greater than the cost of customer retention, churn prediction has emerged as a crucial Business Intelligence (BI) application for modern telecommunication operators. The dominant approach to churn prediction is to model individual customers and derive their likelihood of churn using a predictive model. Recent work has shown that analyzing customers' interactions by assessing the social vicinity of recent churners can improve the accuracy of churn prediction. We propose a novel framework, termed Group-First Churn Prediction, which eliminates the a priori requirement of knowing who recently churned. Specifically, our approach exploits the structure of customer interactions to predict which groups of subscribers are most prone to churn, before even a single member in the group has churned. Our method works by identifying closely-knit groups of subscribers using second order social metrics derived from information theoretic principles. The interactions within each group are then analyzed to identify social leaders. Based on Key Performance Indicators that are derived from these groups, a novel statistical model is used to predict the churn of the groups and their members. Our experimental results, which are based on data from a telecommunication operator with approximately 16 million subscribers, demonstrate the unique advantages of the proposed method. We further provide empirical evidence that our method captures social phenomena in a highly significant manner.
Yossi Richter, Elad Yom-Tov, Noam Slonim
SDM3
2009 Parallel Pairwise Clustering
abstract
Given the pairwise affinity relations associated with a set of data items, the goal of a clustering algorithm is to automatically partition the data into a small number of homogeneous clusters. However, since the input size is quadratic in the number of data points, existing algorithms are non feasible for many practical applications. Here, we propose a simple strategy to cluster massive data by randomly splitting the original affinity matrix into small manageable affinity matrices that are clustered independently. Our proposal is most appealing in a parallel computing environment where at each iteration, each worker node clusters a subset of the input data and the results from all workers are then integrated in a master node to create a new clustering partition over the entire data. We demonstrate that this approach yields high quality clustering partitions for various real world problems, even though at each iteration only small fractions of the original data matrix are examined and at no point is the entire affinity matrix stored in memory or even computed. Furthermore, we demonstrate that the proposed algorithm has intriguing stochastic convergence properties that provide further insight into the clustering problem.
Elad Yom-Tov, Noam Slonim
SDM2
2006 Information Bottleneck for Non Co-Occurrence Data
abstract
We present a general model-independent approach to the analysis of data in cases when these data do not appear in the form of co-occurrence of two variables X, Y , but rather as a sample of values of an unknown (stochastic) function Z (X, Y ). For example, in gene expression data, the expression level Z is a function of gene X and condition Y ; or in movie ratings data the rating Z is a function of viewer X and movie Y . The approach represents a consistent extension of the Information Bottleneck method that has previously relied on the availability of co-occurrence statistics. By altering the relevance variable we eliminate the need in the sample of joint distribution of all input variables. This new formulation also enables simple MDL-like model complexity control and prediction of missing values of Z . The approach is analyzed and shown to be on a par with the best known clustering algorithms for a wide range of domains. For the prediction of missing values (collaborative filtering) it improves the currently best known results.
Yevgeny Seldin, Noam Slonim, Naftali Tishby
NIPS2
2006 Multivariate Information Bottleneck
abstract
The information bottleneck (IB) method is an unsupervised model independent data organization technique. Given a joint distribution, p(X, Y), this method constructs a new variable, T, that extracts partitions, or clusters, over the values of X that are informative about Y. Algorithms that are motivated by the IB method have already been applied to text classification, gene expression, neural code, and spectral analysis. Here, we introduce a general principled framework for multivariate extensions of the IB method. This allows us to consider multiple systems of data partitions that are interrelated. Our approach utilizes Bayesian networks for specifying the systems of clusters and which information terms should be maintained. We show that this construction provides insights about bottleneck variations and enables us to characterize the solutions of these variations. We also present four different algorithmic approaches that allow us to construct solutions in practice and apply them to several real-world problems.
Noam Slonim, Nir Friedman, Naftali Tishby
Neural Comput.1
2002 Discriminative Feature Selection via Multiclass Variable Memory Markov Model
Noam Slonim, Gill Bejerano, Shai Fine, Naftali Tishby
ICML1
2002 Maximum Likelihood and the Information Bottleneck
abstract
 that defines partitions over the values of The information bottleneck (IB) method is an information-theoretic formulation , this method constructs for clustering problems. Given a joint distribution a new variable that are informative . Maximum likelihood (ML) of mixture models is a standard statistical about approach to clustering problems. In this paper, we ask: how are the two methods related ? We define a simple mapping between the IB problem and the ML prob- lem for the multinomial mixture model. We show that under this mapping the problems are strongly related. In fact, for uniform input distribution over or for large sample size, the problems are mathematically equivalent. Specifically, in these cases, every fixed point of the IB-functional defines a fixed point of the (log) likelihood and vice versa. Moreover, the values of the functionals at the fixed points are equal under simple transformations. As a result, in these cases, every algorithm that solves one of the problems, induces a solution for the other.
Noam Slonim, Yair Weiss
NIPS1
2002 Unsupervised document classification using sequential information maximization
abstract
We present a novel sequential clustering algorithm which is motivated by the Information Bottleneck (IB) method. In contrast to the agglomerative IB algorithm, the new sequential (sIB) approach is guaranteed to converge to a local maximum of the information with time and space complexity typically linear in the data size. information, as required by the original IB principle. Moreover, the time and space complexity are significantly improved. We apply this algorithm to unsupervised document classification. In our evaluation, on small and medium size corpora, the sIB is found to be consistently superior to all the other clustering methods we examine, typically by a significant margin. Moreover, the sIB results are comparable to those obtained by a supervised Naive Bayes classifier. Finally, we propose a simple procedure for trading cluster's recall to gain higher precision, and show how this approach can extract clusters which match the existing topics of the corpus almost perfectly.
Noam Slonim, Nir Friedman, Naftali Tishby
SIGIR1
2001 Agglomerative Multivariate Information Bottleneck
abstract
The information bottleneck method is an unsupervised model independent data organization technique. Given a joint distribution peA, B), this method con(cid:173) structs a new variable T that extracts partitions, or clusters, over the values of A that are informative about B. In a recent paper, we introduced a general princi(cid:173) pled framework for multivariate extensions of the information bottleneck method that allows us to consider multiple systems of data partitions that are inter-related. In this paper, we present a new family of simple agglomerative algorithms to construct such systems of inter-related clusters. We analyze the behavior of these algorithms and apply them to several real-life datasets.
Noam Slonim, Nir Friedman, Naftali Tishby
NIPS1
2001 Multivariate Information Bottleneck
Nir Friedman, Ori Mosenzon, Noam Slonim, Naftali Tishby
UAI3
2000 Data Clustering by Markovian Relaxation and the Information Bottleneck Method
abstract
We introduce a new, non-parametric and principled, distance based clustering method. This method combines a pairwise based ap(cid:173) proach with a vector-quantization method which provide a mean(cid:173) ingful interpretation to the resulting clusters. The idea is based on turning the distance matrix into a Markov process and then examine the decay of mutual-information during the relaxation of this process. The clusters emerge as quasi-stable structures dur(cid:173) ing this relaxation, and then are extracted using the information bottleneck method. These clusters capture the information about the initial point of the relaxation in the most effective way. The method can cluster data with no geometric or other bias and makes no assumption about the underlying distribution.
Naftali Tishby, Noam Slonim
NIPS2
2000 Document clustering using word clusters via the information bottleneck method
abstract
We present a novel implementation of the recently introduced information bottleneck method for unsupervised document clustering. Given a joint empirical distribution of words and documents, p(x, y), we first cluster the words, Y, so that the obtained word clusters, Ytilde;, maximally preserve the information on the documents. The resulting joint distribution. p(X, Ytilde;), contains most of the original information about the documents, I(X; Ytilde;) ≈ I(X; Y), but it is much less sparse and noisy. Using the same procedure we then cluster the documents, X, so that the information about the word-clusters is preserved. Thus, we first find word-clusters that capture most of the mutual information about to set of documents, and then find document clusters, that preserve the information about the word clusters. We tested this procedure over several document collections based on subsets taken from the standard 20Newsgroups corpus. The results were assessed by calculating the correlation between the document clusters and the correct labels for these documents. Finding from our experiments show that this double clustering procedure, which uses the information bottleneck method, yields significantly superior performance compared to other common document distributional clustering algorithms. Moreover, the double clustering procedure improves all the distributional clustering methods examined here.
Noam Slonim, Naftali Tishby
SIGIR1
1999 Agglomerative Information Bottleneck
Noam Slonim, Naftali Tishby
NIPS1
1998 WebSuite: A Tool Suite for Harnessing Web Data
Catriel Beeri, Gershon Elber, Tova Milo, Yehoshua Sagiv, Oded Shmueli, Naftali Tishby, Yakov A. Kogan, David Konopnicki, Pini Mogilevski, Noam Slonim
WebDB10