EDBT 2026 Demo / reviewers in the wild / expert
Idan Szpektor
dblp:15/6513
· DBLP profile ↗
60ranked-venue papers
8as first author
26since 2021 · last 2026
0009-0002-2746-4027ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 5 first-author · 23 since 2021Databases, data management, data science and information retrieval · 20 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Location Not Found: Exposing Implicit Local and Global Biases in Multilingual LLMsabstractGuy Mor-Lan, Omer Goldman, Matan Eyal, Adi Mayrav Gilady, Sivan Eiger, Idan Szpektor, Avinatan Hassidim, Yossi Matias, Reut Tsarfaty. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Guy Mor, Omer Goldman, Matan Eyal, Adi Mayrav Gilady, Sivan Eiger, Idan Szpektor, Avinatan Hassidim, Yossi Matias, Reut Tsarfaty |
ACL (1) | 6 |
| 2026 | Localizing Factual Inconsistencies in Attributable Text GenerationabstractAbstract There has been an increasing interest in detecting hallucinations in model-generated texts, both manually and automatically, at varying levels of granularity. However, most existing methods fail to precisely pinpoint the errors. In this work, we introduce QASemConsistency, a new formalism for localizing factual inconsistencies in attributable text generation, at a fine-grained level. Drawing inspiration from Neo-Davidsonian formal semantics, we propose decomposing the generated text into minimal predicate-argument level propositions, expressed as simple question-answer (QA) pairs, and assess whether each individual QA pair is supported by a trusted reference text. As each QA pair corresponds to a single semantic relation between a predicate and an argument, QASemConsistency effectively localizes the unsupported information. We first demonstrate the effectiveness of the QASemConsistency methodology for human annotation, by collecting crowdsourced annotations of granular consistency errors, while achieving a substantial inter-annotator agreement. This benchmark includes more than 3K instances spanning various tasks of attributable text generation. We also show that QASemConsistency yields factual consistency scores that correlate well with human judgments. Finally, we implement several methods for automatically detecting localized factual inconsistencies, with both supervised entailment models and LLMs.1 Arie Cattan, Paul Roit, Shiyue Zhang 0001, David Wan, Roee Aharoni, Idan Szpektor, Mohit Bansal, Ido Dagan |
Trans. Assoc. Comput. Linguistics | 6 |
| 2025 | MDCure: A Scalable Pipeline for Multi-Document Instruction-FollowingabstractMulti-document (MD) processing is crucial for LLMs to handle real-world tasks such as summarization and question-answering across large sets of documents.While LLMs have improved at processing long inputs, MD contexts still present unique difficulties, including management of inter-document dependencies, redundancy, and incoherent structures.To address this challenge, we introduce MDCure, a scalable and effective instruction data generation framework to enhance the MD capabilities of LLMs without the computational cost of pretraining or reliance on human-annotated data.MDCure generates high-quality synthetic MD instruction data over sets of articles via targeted prompts.We also introduce MDCureRM, a cost-effective, MD-specific reward model to score and filter generated data based on their training utility for MD settings.MDCure is compatible with open-and closed-source models in addition to policy optimization methods such as PPO, enabling even small opensource models to surpass proprietary LLMs as strong generators of high-quality MD instruction data without further data filtering.With MDCure, we fine-tune a wide variety of LLMs up to 70B parameters in size from the FlanT5, Qwen2, and LLAMA3.1 model families.Extensive evaluations on a wide range of MD and long-context benchmarks spanning various tasks and domains show MDCure consistently improves performance over pre-trained baselines and base models by up to 75.1%. Gabrielle K. Liu, Avi Caciularu, Idan Szpektor, Arman Cohan |
ACL (1) | 4 |
| 2025 | MetaFaith: Faithful Natural Language Uncertainty Expression in LLMsabstractA critical component in the trustworthiness of LLMs is reliable uncertainty communication, yet LLMs often use assertive language when conveying false claims, leading to over-reliance and eroded trust.We present the first systematic study of faithful confidence calibration of LLMs, benchmarking models' ability to use linguistic expressions of uncertainty that faithfully reflect their intrinsic uncertainty, across a comprehensive array of models, datasets, and prompting strategies.Our results demonstrate that LLMs largely fail at this task, and that existing interventions are insufficient: standard prompt approaches provide only marginal gains, and existing, factuality-based calibration techniques can even harm faithful calibration.To address this critical gap, we introduce MetaFaith, a novel prompt-based calibration approach inspired by human metacognition.We show that MetaFaith robustly improves faithful calibration across diverse models and task domains, enabling up to 61% improvement in faithfulness and achieving an 83% win rate over original generations as judged by humans. Gabrielle K. Liu, Gal Yona, Avi Caciularu, Idan Szpektor, Tim G. J. Rudner, Arman Cohan |
EMNLP | 4 |
| 2025 | Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model PerformanceabstractNLP benchmarks rely on standardized datasets for training and evaluating models and are crucial for advancing the field.Traditionally, expert annotations ensure high-quality labels; however, the cost of expert annotation does not scale well with the growing demand for larger datasets required by modern models.While crowd-sourcing provides a more scalable solution, it often comes at the expense of annotation precision and consistency.Recent advancements in large language models (LLMs) offer new opportunities to enhance the annotation process, particularly for detecting label errors in existing datasets.In this work, we consider the recent approach of LLM-as-a-judge, leveraging an ensemble of LLMs to flag potentially mislabeled examples.We conduct a case study on four factual consistency datasets from the TRUE benchmark, spanning diverse NLP tasks, and on SummEval, which uses Likertscale ratings of summary quality across multiple dimensions.We empirically analyze the labeling quality of existing datasets and compare expert, crowd-sourced, and LLM-based annotations in terms of the agreement, label quality, and efficiency, demonstrating the strengths and limitations of each annotation method.Our findings reveal a substantial number of label errors, which, when corrected, induce a significant upward shift in reported model performance.This suggests that many of the LLMs' so-called mistakes are due to label errors rather than genuine model failures.Additionally, we discuss the implications of mislabeled data and propose methods to mitigate them in training to improve performance. Omer Nahum, Nitay Calderon, Orgad Keller, Idan Szpektor, Roi Reichart |
EMNLP | 4 |
| 2025 | LLMs Know More Than They Show: On the Intrinsic Representation of LLM HallucinationsabstractLarge language models (LLMs) often produce errors, including factual inaccuracies, biases, and reasoning failures, collectively referred to as "hallucinations". Recent studies have demonstrated that LLMs' internal states encode information regarding the truthfulness of their outputs, and that this information can be utilized to detect errors. In this work, we show that the internal representations of LLMs encode much more information about truthfulness than previously recognized. We first discover that the truthfulness information is concentrated in specific tokens, and leveraging this property significantly enhances error detection performance. Yet, we show that such error detectors fail to generalize across datasets, implying that---contrary to prior claims---truthfulness encoding is not universal but rather multifaceted. Next, we show that internal representations can also be used for predicting the types of errors the model is likely to make, facilitating the development of tailored mitigation strategies. Lastly, we reveal a discrepancy between LLMs' internal encoding and external behavior: they may encode the correct answer, yet consistently generate an incorrect one. Taken together, these insights deepen our understanding of LLM errors from the model's internal perspective, which can guide future research on enhancing error analysis and mitigation. Hadas Orgad, Michael Toker, Zorik Gekhman, Roi Reichart, Idan Szpektor, Hadas Kotek, Yonatan Belinkov |
ICLR | 5 |
| 2025 | Video-STaR: Self-Training Enables Video Instruction Tuning with Any SupervisionabstractThe performance and reasoning capabilities of Large Multi-modal Models (LMMs) is dependent on the size and quality of their training datasets. However, collecting datasets that support chain-of-thought instruction tuning is highly challenging. Existing video instruction tuning datasets are often derived by prompting large language models with video captions to generate question-answer pairs, which makes them predominantly descriptive rather than reasoning-focused.
Meanwhile, many labeled video datasets with diverse labels and supervision exist -- however, we find that their integration into LMMs is non-trivial.
Herein, we present $\underline{\text{Video}}$ $\underline{\text{S}}\text{elf}$-$\underline{\text{T}}\text{raining}$ $\text{with}$ $\underline{\text{a}}\text{ugmented}$ $\underline{\text{R}}\text{easoning}$ (Video-STaR), the first self-training approach for video instruction tuning.
Video-STaR allows the utilization of *any* labeled video dataset for video instruction tuning.
In Video-STaR, an LMM cycles between instruction generation and finetuning, which we show (I) improves general video understanding and (II) adapts LMMs to novel downstream tasks with existing supervision.
During instruction generation, an LMM is prompted to propose an answer. The answers are then filtered only to those that contain the original video labels, and the LMM is then re-trained on the generated dataset.
By training exclusively on generated answers containing the correct video labels, Video-STaR leverages these existing labels as weak supervision for video instruction tuning.
Our results demonstrate that Video-STaR-augmented LMMs achieve notable improvements in (I) general Video QA, where TempCompass performance improved by 6.1%, *and* (II) downstream tasks, with a 9.9% increase in Kinetics700-QA accuracy and a 4.0% improvement in action quality assessment on FineDiving, while also exhibiting better interpretability. Orr Zohar, Yonatan Bitton, Idan Szpektor, Serena Yeung-Levy |
ICLR | 4 |
| 2025 | Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted CaptionsabstractMoran Yanuka, Assaf Ben-Kish, Yonatan Bitton, Idan Szpektor, Raja Giryes. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Moran Yanuka, Assaf Ben-Kish, Yonatan Bitton, Idan Szpektor, Raja Giryes |
NAACL (Long Papers) | 4 |
| 2025 | 3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language ModelabstractHumans excel at performing complex tasks by leveraging long-term memory across temporal and spatial experiences. In contrast, current Large Language Models (LLMs) struggle to effectively plan and act in dynamic, multi-room 3D environments.
We posit that part of this limitation is due to the lack of proper 3D spatial-temporal memory modeling in LLMs.
To address this, we first introduce 3DMem-Bench, a comprehensive benchmark comprising over 26,000 trajectories and 2,892 embodied tasks, question-answering and captioning, designed to evaluate an agent's ability to reason over long-term memory in 3D environments.
Second, we propose 3DLLM-Mem, a novel dynamic memory management and fusion model for embodied spatial-temporal reasoning and actions in LLMs.
Our model uses working memory tokens, which represents current observations, as queries to selectively attend to and fuse the most useful spatial and temporal features from episodic memory, which stores past observations and interactions. Our approach allows the agent to focus on task-relevant information while maintaining memory efficiency in complex, long-horizon environments.
Experimental results demonstrate that 3DLLM-Mem achieves state-of-the-art performance across various tasks, outperforming the strongest baselines by 16.5\% in success rate on 3DMem-Bench's most challenging in-the-wild embodied tasks. Wenbo Hu 0006, Yining Hong, Leison Gao, Zibu Wei, Xingcheng Yao, Nanyun Peng 0001, Yonatan Bitton, Idan Szpektor, Kai-Wei Chang 0001 |
NeurIPS | 9 |
| 2025 | Contrastive Sequential-Diffusion Learning: Non-Linear and Multi-Scene Instructional Video SynthesisabstractGenerated video scenes for action-centric sequence descriptions, such as recipe instructions and do-it-yourself projects, often include non-linear patterns, where the next video may need to be visually consistent not with the immediately preceding video but with earlier ones. Current multi-scene video synthesis approaches fail to meet these consistency requirements. To address this, we propose a contrastive sequential video diffusion method that selects the most suitable previously generated scene to guide and condition the denoising process of the next scene. The result is a multi-scene video that is grounded in the scene descriptions and coherent w.r.t. the scenes that require visual consistency. Experiments with action-centered data from the real world demonstrate the practicality and improved consistency of our model compared to previous work. Code and examples available at https://github.com/novasearch/CoSeD Vasco Ramos, Yonatan Bitton, Michal Yarom, Idan Szpektor, João Magalhães |
WACV | 4 |
| 2024 | Generating Coherent Sequences of Visual Illustrations for Real-World Manual TasksabstractJoão Bordalo, Vasco Ramos, Rodrigo Valério, Diogo Glória-Silva, Yonatan Bitton, Michal Yarom, Idan Szpektor, Joao Magalhaes. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. João Bordalo, Vasco Ramos, Rodrigo Valerio, Diogo Glória-Silva, Yonatan Bitton, Michal Yarom, Idan Szpektor, João Magalhães |
ACL (1) | 7 |
| 2024 | VideoCon: Robust Video-Language Alignment via Contrast CaptionsabstractDespite being (pre)trained on a massive amount of data, state-of-the-art video-language alignment models are not robust to semantically-plausible contrastive changes in the video captions. Our work addresses this by identifying a broad spectrum of contrast misalignments, such as re-placing entities, actions, and flipping event order, which alignment models should be robust against. To this end, we introduce the VideoCon, a video-language alignment dataset constructed by a large language model that gen-erates plausible contrast video captions and explanations for differences between original and contrast video captions. Then, a generative video-language model is fine-tuned with VideoCon to assess video-language entailment and generate explanations. Our VideoCon-based alignment model significantly outperforms current models. It exhibits a 12-point increase in AUC for the video-language alignment task on human-generated contrast captions. Finally, our model sets new state of the art zero-shot performance in temporally-extensive video-language tasks such as text-to-video retrieval (SSv2-Temporal) and video question answering (ATP-Hard). Moreover, our model shows superior performance on novel videos and human-crafted captions and explanations. Hritik Bansal, Yonatan Bitton, Idan Szpektor, Kai-Wei Chang 0001, Aditya Grover |
CVPR | 3 |
| 2024 | Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment
Brian Gordon, Yonatan Bitton, Yonatan Shafir, Roopal Garg, Xi Chen 0071, Dani Lischinski, Daniel Cohen-Or, Idan Szpektor |
ECCV (57) | 8 |
| 2024 | Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language ModelsabstractImagine observing someone scratching their arm; to understand why, additional context would be necessary. However, spotting a mosquito nearby would immediately offer a likely explanation for the person’s discomfort, thereby alleviating the need for further information. This example illustrates how subtle visual cues can challenge our cognitive skills and demonstrates the complexity of interpreting visual scenarios. To study these skills, we present Visual Riddles, a benchmark aimed to test vision and language models on visual riddles requiring commonsense and world knowledge. The benchmark comprises 400 visual riddles, each featuring a unique image created by a variety of text-to-image models, question, ground-truth answer, textual hint, and attribution. Human evaluation reveals that existing models lag significantly behind human performance, which is at 82% accuracy, with Gemini-Pro-1.5 leading with 40% accuracy. Our benchmark comes with automatic evaluation tasks to make assessment scalable. These findings underscore the potential of Visual Riddles as a valuable resource for enhancing vision and language models’ capabilities in interpreting complex visual scenarios. Data, code, and leaderboard are available at https://visual-riddles.github.io/. Nitzan Guetta, Aviv Slobodkin, Aviya Maimon, Eliya Habba, Royi Rassin, Yonatan Bitton, Idan Szpektor, Amir Globerson, Yuval Elovici |
NeurIPS | 7 |
| 2024 | Multi-turn Reinforcement Learning with Preference Human FeedbackabstractReinforcement Learning from Human Feedback (RLHF) has become the standard approach for aligning Large Language Models (LLMs) with human preferences, allowing LLMs to demonstrate remarkable abilities in various tasks. Existing methods work by emulating the human preference at the single decision (turn) level, limiting their capabilities in settings that require planning or multi-turn interactions to achieve a long-term goal. In this paper, we address this issue by developing novel methods for Reinforcement Learning (RL) from preference feedback between two full multi-turn conversations. In the tabular setting, we present a novel mirror-descent-based policy optimization algorithm for the general multi-turn preference-based RL problem, and prove its convergence to Nash equilibrium. To evaluate performance, we create a new environment, Education Dialogue, where a teacher agent guides a student in learning a random topic, and show that a deep RL variant of our algorithm outperforms RLHF baselines. Finally, we show that in an environment with explicit rewards, our algorithm recovers the same performance as a reward-based RL baseline, despite relying solely on a weaker preference signal. Lior Shani, Aviv Rosenberg 0002, Asaf Cassel, Oran Lang, Daniele Calandriello, Avital Zipori, Hila Noga, Orgad Keller, Bilal Piot, Idan Szpektor, Avinatan Hassidim, Yossi Matias, Rémi Munos |
NeurIPS | 10 |
| 2023 | DisentQA: Disentangling Parametric and Contextual Knowledge with Counterfactual Question AnsweringabstractElla Neeman, Roee Aharoni, Or Honovich, Leshem Choshen, Idan Szpektor, Omri Abend. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Ella Neeman, Roee Aharoni, Or Honovich, Leshem Choshen, Idan Szpektor, Omri Abend |
ACL (1) | 5 |
| 2023 | Factually Consistent Summarization via Reinforcement Learning with Textual Entailment FeedbackabstractPaul Roit, Johan Ferret, Lior Shani, Roee Aharoni, Geoffrey Cideron, Robert Dadashi, Matthieu Geist, Sertan Girgin, Leonard Hussenot, Orgad Keller, Nikola Momchev, Sabela Ramos Garea, Piotr Stanczyk, Nino Vieillard, Olivier Bachem, Gal Elidan, Avinatan Hassidim, Olivier Pietquin, Idan Szpektor. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Paul Roit, Johan Ferret, Lior Shani, Roee Aharoni, Geoffrey Cideron, Robert Dadashi, Matthieu Geist, Sertan Girgin, Léonard Hussenot, Orgad Keller, Nikola Momchev, Sabela Ramos, Piotr Stanczyk, Nino Vieillard, Olivier Bachem, Gal Elidan, Avinatan Hassidim, Olivier Pietquin, Idan Szpektor |
ACL (1) | 19 |
| 2023 | Sentence Retrieval for Open-Ended Dialogue Using Dual Contextual Modeling
Itay Harel, Hagai Taitelbaum, Idan Szpektor, Oren Kurland |
ECIR (1) | 3 |
| 2023 | TrueTeacher: Learning Factual Consistency Evaluation with Large Language ModelsabstractFactual consistency evaluation is often conducted using Natural Language Inference (NLI) models, yet these models exhibit limited success in evaluating summaries.Previous work improved such models with synthetic training data.However, the data is typically based on perturbed human-written summaries, which often differ in their characteristics from real model-generated summaries and have limited coverage of possible factual errors.Alternatively, large language models (LLMs) have recently shown promising results in directly evaluating generative tasks, but are too computationally expensive for practical use.Motivated by these limitations, we introduce TrueTeacher, a method for generating synthetic data by annotating diverse model-generated summaries using a LLM.Unlike prior work, TrueTeacher does not rely on human-written summaries, and is multilingual by nature.Experiments on the TRUE benchmark show that a student model trained using our data, substantially outperforms both the state-of-the-art model with similar capacity, and the LLM teacher.In a systematic study, we compare TrueTeacher to existing synthetic data generation methods and demonstrate its superiority and robustness to domain-shift.We also show that our method generalizes to multilingual scenarios.Lastly, we release our largescale synthetic dataset (1.4M examples), generated using TrueTeacher, and a checkpoint trained on this data.1 Zorik Gekhman, Jonathan Herzig, Roee Aharoni, Chen Elkind, Idan Szpektor |
EMNLP | 5 |
| 2023 | What You See is What You Read? Improving Text-Image Alignment EvaluationabstractAutomatically determining whether a text and a corresponding image are semantically aligned is a significant challenge for vision-language models, with applications in generative text-to-image and image-to-text tasks. In this work, we study methods for automatic text-image alignment evaluation. We first introduce SeeTRUE: a comprehensive evaluation set, spanning multiple datasets from both text-to-image and image-to-text generation tasks, with human judgements for whether a given text-image pair is semantically aligned. We then describe two automatic methods to determine alignment: the first involving a pipeline based on question generation and visual question answering models, and the second employing an end-to-end classification approach by finetuning multimodal pretrained models. Both methods surpass prior approaches in various text-image alignment tasks, with significant improvements in challenging cases that involve complex composition or unnatural images. Finally, we demonstrate how our approaches can localize specific misalignments between an image and a given text, and how they can be used to automatically re-rank candidates in text-to-image generation. Michal Yarom, Yonatan Bitton, Soravit Changpinyo, Roee Aharoni, Jonathan Herzig, Oran Lang, Eran Ofek, Idan Szpektor |
NeurIPS | 8 |
| 2023 | On the Robustness of Dialogue History Representation in Conversational Question Answering: A Comprehensive Study and a New Prompt-based MethodabstractAbstract Most work on modeling the conversation history in Conversational Question Answering (CQA) reports a single main result on a common CQA benchmark. While existing models show impressive results on CQA leaderboards, it remains unclear whether they are robust to shifts in setting (sometimes to more realistic ones), training data size (e.g., from large to small sets) and domain. In this work, we design and conduct the first large-scale robustness study of history modeling approaches for CQA. We find that high benchmark scores do not necessarily translate to strong robustness, and that various methods can perform extremely differently under different settings. Equipped with the insights from our study, we design a novel prompt-based history modeling approach and demonstrate its strong robustness across various settings. Our approach is inspired by existing methods that highlight historic answers in the passage. However, instead of highlighting by modifying the passage token embeddings, we add textual prompts directly in the passage text. Our approach is simple, easy to plug into practically any model, and highly effective, thus we recommend it as a starting point for future model developers. We also hope that our study and insights will raise awareness to the importance of robustness-focused evaluation, in addition to obtaining high leaderboard scores, leading to better CQA systems.1 Zorik Gekhman, Nadav Oved, Orgad Keller, Idan Szpektor, Roi Reichart |
Trans. Assoc. Comput. Linguistics | 4 |
| 2022 | All You May Need for VQA are Image CaptionsabstractSoravit Changpinyo, Doron Kukliansy, Idan Szpektor, Xi Chen, Nan Ding, Radu Soricut. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Soravit Changpinyo, Doron Kukliansky, Idan Szpektor, Xi Chen 0071, Nan Ding 0002, Radu Soricut |
NAACL-HLT | 3 |
| 2022 | TRUE: Re-evaluating Factual Consistency EvaluationabstractOr Honovich, Roee Aharoni, Jonathan Herzig, Hagai Taitelbaum, Doron Kukliansy, Vered Cohen, Thomas Scialom, Idan Szpektor, Avinatan Hassidim, Yossi Matias. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Or Honovich, Roee Aharoni, Jonathan Herzig, Hagai Taitelbaum, Doron Kukliansky, Vered Cohen, Thomas Scialom, Idan Szpektor, Avinatan Hassidim, Yossi Matias |
NAACL-HLT | 8 |
| 2022 | A Dataset for Sentence Retrieval for Open-Ended DialoguesabstractWe address the task of sentence retrieval for open-ended dialogues. The goal is to retrieve sentences from a document corpus that contain information useful for generating the next turn in a given dialogue. Prior work on dialogue-based retrieval focused on specific types of dialogues: either conversational QA or conversational search. To address a broader scope of this task where any type of dialogue can be used, we constructed a dataset that includes open-ended dialogues from Reddit, candidate sentences from Wikipedia for each dialogue and human annotations for the sentences. We report the performance of several retrieval baselines, including neural retrieval models, over the dataset. To adapt neural models to the types of dialogues in the dataset, we explored an approach to induce a large-scale weakly supervised training data from Reddit. Using this training set significantly improved the performance over training on the MS MARCO dataset. Itay Harel, Hagai Taitelbaum, Idan Szpektor, Oren Kurland |
SIGIR | 3 |
| 2021 | What's the Best Place for an AI Conference, Vancouver or _______: Why Completing Comparative Questions is DifficultabstractAlthough large neural language models (LMs) like BERT can be finetuned to yield state-of-the-art results on many NLP tasks, it is often unclear what these models actually learn. Here we study using such LMs to fill in entities in human-authored comparative questions, like ``Which country is older, India or _____?''---i.e., we study the ability of neural LMs to ask (not answer) reasonable questions. We show that accuracy in this fill-in-the-blank task is well-correlated with human judgements of whether a question is reasonable, and that these models can be trained to achieve nearly human-level performance in completing comparative questions in three different subdomains. However, analysis shows that what they learn fails to model any sort of broad notion of which entities are semantically comparable or similar---instead the trained models are very domain-specific, and performance is highly correlated with co-occurrences between specific entities observed in the training set. This is true both for models that are pretrained on general text corpora, as well as models trained on a large corpus of comparison questions. Our study thus reinforces recent results on the difficulty of making claims about a deep model's world knowledge or linguistic competence based on performance on specific benchmark problems. We make our evaluation datasets publicly available to foster future research on complex understanding and reasoning in such models at standards of human interaction. Avishai Zagoury, Einat Minkov, Idan Szpektor, William W. Cohen |
AAAI | 3 |
| 2021 | $Q^2$: Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question AnsweringabstractNeural knowledge-grounded generative models for dialogue often produce content that is factually inconsistent with the knowledge they rely on, making them unreliable and limiting their applicability.Inspired by recent work on evaluating factual consistency in abstractive summarization, we propose an automatic evaluation metric for factual consistency in knowledge-grounded dialogue using automatic question generation and question answering.Our metric, denoted Q 2 , compares answer spans using natural language inference (NLI), instead of token-based matching as done in previous work.To foster proper evaluation, we curate a novel dataset of dialogue system outputs for the Wizard-of-Wikipedia dataset, manually annotated for factual consistency.We perform a thorough meta-evaluation of Q 2 against other metrics using this dataset and two others, where it consistently shows higher correlation with human judgements. Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, Omri Abend |
EMNLP (1) | 5 |
| 2020 | Dynamic Composition for Conversational Domain ExplorationabstractWe study conversational domain exploration (CODEX), where the user’s goal is to enrich her knowledge of a given domain by conversing with an informative bot. Such conversations should be well grounded in high-quality domain knowledge as well as engaging and open-ended. A CODEX bot should be proactive and introduce relevant information even if not directly asked for by the user. The bot should also appropriately pivot the conversation to undiscovered regions of the domain. To address these dialogue characteristics, we introduce a novel approach termed dynamic composition that decouples candidate content generation from the flexible composition of bot responses. This allows the bot to control the source, correctness and quality of the offered content, while achieving flexibility via a dialogue manager that selects the most appropriate contents in a compositional manner. We implemented a CODEX bot based on dynamic composition and integrated it into the Google Assistant . As an example domain, the bot conversed about the NBA basketball league in a seamless experience, such that users were not aware whether they were conversing with the vanilla system or the one augmented with our CODEX bot. Results are positive and offer insights into what makes for a good conversation. To the best of our knowledge, this is the first real user experiment of open-ended dialogues as part of a commercial assistant system. Idan Szpektor, Deborah Cohen, Gal Elidan, Avinatan Hassidim, Orgad Keller, Sayali Kulkarni, Eran Ofek, Sagie Pudinsky, Asaf Revach, Shimi Salant, Yossi Matias |
WWW | 1 |
| 2019 | A Joint Named-Entity Recognizer for Heterogeneous Tag-sets Using a Tag HierarchyabstractWe study a variant of domain adaptation for named-entity recognition where multiple, heterogeneously tagged training sets are available.Furthermore, the test tag-set is not identical to any individual training tag-set.Yet, the relations between all tags are provided in a tag hierarchy, covering the test tags as a combination of training tags.This setting occurs when various datasets are created using different annotation schemes.This is also the case of extending a tag-set with a new tag by annotating only the new tag in a new dataset.We propose to use the given tag hierarchy to jointly learn a neural network that shares its tagging layer among all tag-sets.We compare this model to combining independent models and to a model based on the multitasking approach.Our experiments show the benefit of the tag-hierarchy model, especially when facing non-trivial consolidation of tag-sets. Genady Beryozkin, Yoel Drori, Oren Gilon, Tzvika Hartman, Idan Szpektor |
ACL (1) | 5 |
| 2018 | Care to Share?: Learning to Rank Personal Photos for Public SharingabstractWith mobile devices, users are taking ever-growing numbers of photos every day. These photos are uploaded to social sites such as Facebook and Flickr, often automatically. Yet, the portion of these uploaded photos being publicly shared is low, and on a constant decline. Deciding which photo to share takes considerable time and attention, and many users would rather forfeit the social interaction and engagement than sift through their piles of uploaded photos. In this paper, we introduce a novel task of recommending socially-engaging photos to their creators for public sharing. This will turn a tedious manual chore into a quick, software-assisted process. We provide extensive analysis over a large-scale dataset from the Flickr photo sharing website, which reveals some of the traits of photo sharing in such sites. Additionally, we present a ranking algorithm for the task that comprises three steps:(a) grouping of near-duplicate photos;(b) ranking the photos in each group by their "shareability"; and(c) ranking the groups by their likelihood to contain a shareable photo. A large-scale experiment allows us to evaluate our algorithm and show its benefits compared to competitive baselines and algorithmic alternatives. Ido Guy, Alexander Nus, Dan Pelleg, Idan Szpektor |
WSDM | 4 |
| 2016 | When the Crowd is Not Enough: Improving User Experience with Social Media through Automatic Quality AnalysisabstractSocial media gives voice to the people, but also opens the door to low-quality contributions, which degrade the experience for the majority of users. To address the latter issue, the prevailing solution is to rely on the 'wisdom of the crowds' to promote good content (e.g., via votes or 'like' buttons), or to downgrade bad content. Unfortunately, such crowd feedback may be sparse, subjective, and slow to accumulate. In this pa- per, we investigate the effects, on the users, of automatically filtering question-answering content, using a combination of syntactic, semantic, and social signals. Using this filtering, a large-scale experiment with real users was performed to mea- sure the resulting engagement and satisfaction. To our knowledge, this experiment represents the first reported large-scale user study of automatically curating social media content in real time. Our results show that automated quality filtering indeed improves user engagement, usually aligning with, and often outperforming, crowd-based quality judgments. Dan Pelleg, Oleg Rokhlenko, Idan Szpektor, Eugene Agichtein, Ido Guy |
CSCW | 3 |
| 2016 | Supporting Human Answers for Advice-Seeking Questions in CQA Sites
Liora Braunstain, Oren Kurland, David Carmel, Idan Szpektor, Anna Shtok |
ECIR | 4 |
| 2016 | Syntactic Parsing of Web Queries with Question IntentabstractAccurate automatic processing of Web queries is important for high-quality information retrieval from the Web.While the syntactic structure of a large portion of these queries is trivial, the structure of queries with question intent is much richer.In this paper we therefore address the task of statistical syntactic parsing of such queries.We first show that the standard dependency grammar does not account for the full range of syntactic structures manifested by queries with question intent.To alleviate this issue we extend the dependency grammar to account for segments -independent syntactic units within a potentially larger syntactic structure.We then propose two distant supervision approaches for the task.Both algorithms do not require manually parsed queries for training.Instead, they are trained on millions of (query, page title) pairs from the Community Question Answering (CQA) domain, where the CQA page was clicked by the user who initiated the query in a search engine.Experiments on a new treebank 1 consisting of 5,000 Web queries from the CQA domain, manually parsed using the proposed grammar, show that our algorithms outperform alternative approaches trained on various sources: tens of thousands of manually parsed OntoNotes sentences, millions of unlabeled CQA queries and thousands of manually segmented CQA queries. Yuval Pinter, Roi Reichart, Idan Szpektor |
HLT-NAACL | 3 |
| 2016 | SIGIR 2016 Workshop WebQA II: Web Question Answering Beyond FactoidsabstractWeb search engines have made great progress at answering factoid queries. However, they are not well-tailored for managing more complex questions, especially when they require explanation and/or description. The WebQA workshop series aims at exploring diverse approaches to answering questions on the Web. This year, particular emphasis will be given to Community Question Answering (CQA), where comments by the users engaged in the forum communities can be used to answer new questions. Questions posted on the Web can be short and ambiguous (similarly to Web queries to a search engine). These issues make the WebQA task more challenging than traditional QA, and finding the most effective approaches for it remains an open problem. Alessandro Moschitti, Lluís Màrquez, Preslav Nakov, Eugene Agichtein, Charles L. A. Clarke, Idan Szpektor |
SIGIR | 6 |
| 2016 | Novelty based Ranking of Human Answers for Community QuestionsabstractQuestions and their corresponding answers within a community based question answering (CQA) site are frequently presented as top search results forWeb search queries and viewed by millions of searchers daily. The number of answers for CQA questions ranges from a handful to dozens, and a searcher would be typically interested in the different suggestions presented in various answers for a question. Yet, especially when many answers are provided, the viewer may not want to sift through all answers but to read only the top ones. Prior work on answer ranking in CQA considered the qualitative notion of each answer separately, mainly whether it should be marked as best answer. We propose to promote CQA answers not only by their relevance to the question but also by the diversification and novelty qualities they hold compared to other answers. Specifically, we aim at ranking answers by the amount of new aspects they introduce with respect to higher ranked answers (novelty), on top of their relevance estimation. This approach is common in Web search and information retrieval, yet it was not addressed within the CQA settings before, which is quite different from classic document retrieval. We propose a novel answer ranking algorithm that borrows ideas from aspect ranking and multi-document summarization, but adapts them to our scenario. Answers are ranked in a greedy manner, taking into account their relevance to the question as well as their novelty compared to higher ranked answers and their coverage of important aspects. An experiment over a collection of Health questions, using a manually annotated gold-standard dataset, shows that considering novelty for answer ranking improves the quality of the ranked answer list. Adi Omari, David Carmel, Oleg Rokhlenko, Idan Szpektor |
SIGIR | 4 |
| 2016 | That's Not My Question: Learning to Weight Unmatched Terms in CQA Vertical SearchabstractA fundamental task in Information Retrieval (IR) is term weighting. Early IR theory considered both the presence or absence of all terms in the lexicon for ranking and needed to weight them all. Yet, as the size of lexicons grew and models became too complex, common weighting models preferred to aggregate only the weights of the query terms that are matched in candidate documents. Thus, unmatched term contribution in these models is only considered indirectly, such as in probability smoothing with corpus distribution, or in weight normalization by document length. In this work we propose a novel term weighting model that directly assesses the weights of unmatched terms, and show its benefits. Specifically, we propose a Learning To Rank framework, in which features corresponding to matched terms are also "mirrored" in similar features that account only for unmatched terms. The relative importance of each feature is learned via a click-through query log. As a test case, we consider vertical search in Community-based Question Answering(CQA) sites from Web queries. Queries that result in viewing CQA content often contain fine grained information needs and benefit more from unmatched term weighting. We assess our model both via manual evaluation and via automatic evaluation over a clickthrough log. Our results show consistent improvement in retrieval when unmatched information is taken into account. This holds both when only identical terms are considered matched, and when related terms are matched via distributional similarity. Boaz Petersil, Avihai Mejer, Idan Szpektor, Koby Crammer |
SIGIR | 3 |
| 2016 | Identifying Web Queries with Question IntentabstractVertical selection is the task of predicting relevant verticals for a Web query so as to enrich the Web search results with complementary vertical results. We investigate a novel variant of this task, where the goal is to detect queries with a question intent. Specifically, we address queries for which the user would like an answer with a human touch. We call these CQA-intent queries, since answers to them are typically found in community question answering (CQA) sites. A typical approach in vertical selection is using a vertical's specific language model of relevant queries and computing the query-likelihood for each vertical as a selective criterion. This works quite well for many domains like Shopping, Local and Travel. Yet, we claim that queries with CQA intent are harder to distinguish by modeling content alone, since they cover many different topics. We propose to also take the structure of queries into consideration, reasoning that queries with question intent have quite a different structure than other queries. We present a supervised classification scheme, random forest over word-clusters for variable length texts, which can model the query structure. Our experiments show that it substantially improves classification performance in the CQA-intent selection task compared to content-oriented based classification, especially as query length grows. Gilad Tsur, Yuval Pinter, Idan Szpektor, David Carmel |
WWW | 3 |
| 2015 | Web Question Answering: Beyond Factoids: SIGIR 2015 WorkshopabstractNo abstract available. Eugene Agichtein, David Carmel, Charles L. A. Clarke, Praveen K. Paritosh, Dan Pelleg, Idan Szpektor |
SIGIR | 6 |
| 2015 | Unsupervised acquisition of entailment relations from the WebabstractAbstract Entailment recognition is a primary generic task in natural language inference, whose focus is to detect whether the meaning of one expression can be inferred from the meaning of the other. Accordingly, many NLP applications would benefit from high coverage knowledgebases of paraphrases and entailment rules. To this end, learning such knowledgebases from the Web is especially appealing due to its huge size as well as its highly heterogeneous content, allowing for a more scalable rule extraction of various domains. However, the scalability of state-of-the-art entailment rule acquisition approaches from the Web is still limited. We present a fully unsupervised learning algorithm for Web-based extraction of entailment relations. We focus on increased scalability and generality with respect to prior work, with the potential of a large-scale Web-based knowledgebase. Our algorithm takes as its input a lexical–syntactic template and searches the Web for syntactic templates that participate in an entailment relation with the input template. Experiments show promising results, achieving performance similar to a state-of-the-art unsupervised algorithm, operating over an offline corpus, but with the benefit of learning rules for different domains with no additional effort. Idan Szpektor, Hristo Tanev, Ido Dagan, Bonaventura Coppola, Milen Kouylekov |
Nat. Lang. Eng. | 1 |
| 2014 | Improving Term Weighting for Community Question Answering Search Using Syntactic AnalysisabstractQuery term weighting is a fundamental task in information retrieval and most popular term weighting schemes are primarily based on statistical analysis of term occurrences within the document collection. In this work we study how term weighting may benefit from syntactic analysis of the corpus. Focusing on community question answering (CQA) sites, we take into account the syntactic function of the terms within CQA texts as an important factor affecting their relative importance for retrieval. We analyze a large log of web queries that landed on Yahoo Answers site, showing a strong deviation between the tendencies of different document words to appear in a landing (click-through) query given their syntactic function. To this end, we propose a novel term weighting method that makes use of the syntactic information available for each query term occurrence in the document, on top of term occurrence statistics. The relative importance of each feature is learned via a learning to rank algorithm that utilizes a click-through query log. We examine the new weighting scheme using manual evaluation based on editorial data and using automatic evaluation over the query log. Our experimental results show consistent improvement in retrieval when syntactic information is taken into account. David Carmel, Avihai Mejer, Yuval Pinter, Idan Szpektor |
CIKM | 4 |
| 2014 | Probabilistic Modeling of Joint-context in Distributional SimilarityabstractMost traditional distributional similarity models fail to capture syntagmatic patterns that group together multiple word features within the same joint context.In this work we introduce a novel generic distributional similarity scheme under which the power of probabilistic models can be leveraged to effectively model joint contexts.Based on this scheme, we implement a concrete model which utilizes probabilistic n-gram language models.Our evaluations suggest that this model is particularly wellsuited for measuring similarity for verbs, which are known to exhibit richer syntagmatic patterns, while maintaining comparable or better performance with respect to competitive baselines for nouns.Following this, we propose our scheme as a framework for future semantic similarity models leveraging the substantial body of work that exists in probabilistic language modeling. Oren Melamud, Ido Dagan, Jacob Goldberger, Idan Szpektor, Deniz Yuret |
CoNLL | 4 |
| 2013 | A Two Level Model for Context Sensitive Inference Rules
Oren Melamud, Jonathan Berant, Ido Dagan, Jacob Goldberger, Idan Szpektor |
ACL (1) | 5 |
| 2013 | Generating Synthetic Comparable Questions for News Articles
Oleg Rokhlenko, Idan Szpektor |
ACL (1) | 2 |
| 2013 | Will My Question Be Answered? Predicting "Question Answerability" in Community Question-Answering Sites
Gideon Dror, Yoelle Maarek, Idan Szpektor |
ECML/PKDD (3) | 3 |
| 2013 | From query to question in one click: suggesting synthetic questions to searchersabstractIn Web search, users may remain unsatisfied for several reasons: the search engine may not be effective enough or the query might not reflect their intent. Years of research focused on providing the best user experience for the data available to the search engine. However, little has been done to address the cases in which relevant content for the specific user need has not been posted on the Web yet. One obvious solution is to directly ask other users to generate the missing content using Community Question Answering services such as Yahoo! Answers or Baidu Zhidao. However, formulating a full-fledged question after having issued a query requires some effort. Some previous work proposed to automatically generate natural language questions from a given query, but not for scenarios in which a searcher is presented with a list of questions to choose from. We propose here to generate synthetic questions that can actually be clicked by the searcher so as to be directly posted as questions on a Community Question Answering service. This imposes new constraints, as questions will be actually shown to searchers, who will not appreciate an awkward style or redundancy. To this end, we introduce a learning-based approach that improves not only the relevance of the suggested questions to the original query, but also their grammatical correctness. In addition, since queries are often underspecified and ambiguous, we put a special emphasis on increasing the diversity of suggestions via a novel diversification mechanism. We conducted several experiments to evaluate our approach by comparing it to prior work. The experiments show that our algorithm improves question quality by 14% over prior work and that adding diversification reduced redundancy by 55%. Gideon Dror, Yoelle Maarek, Avihai Mejer, Idan Szpektor |
WWW | 4 |
| 2013 | When relevance is not enough: promoting diversity and freshnessin personalized question recommendationabstractWhat makes a good question recommendation system for community question-answering sites? First, to maintain the health of the ecosystem, it needs to be designed around answerers, rather than exclusively for askers. Next, it needs to scale to many questions and users, and be fast enough to route a newly-posted question to potential answerers within the few minutes before the asker's patience runs out. It also needs to show each answerer questions that are relevant to his or her interests. We have designed and built such a system for Yahoo! Answers, but realized, when testing it with live users, that it was not enough. Idan Szpektor, Yoelle Maarek, Dan Pelleg |
WWW | 1 |
| 2013 | Introduction to special section on paraphrasingabstractNo abstract available. Haifeng Wang 0001, William B. Dolan, Idan Szpektor |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2012 | Learning Verb Inference Rules from Linguistically-Motivated Evidence
Hila Weisman, Jonathan Berant, Idan Szpektor, Ido Dagan |
EMNLP-CoNLL | 3 |
| 2012 | When web search fails, searchers become askers: understanding the transitionabstractWhile Web search has become increasingly effective over the last decade, for many users' needs the required answers may be spread across many documents, or may not exist on the Web at all. Yet, many of these needs could be addressed by asking people via popular Community Question Answering (CQA) services, such as Baidu Knows, Quora, or Yahoo! Answers. In this paper, we perform the first large-scale analysis of how searchers become askers. For this, we study the logs of a major web search engine to trace the transformation of a large number of failed searches into questions posted on a popular CQA site. Specifically, we analyze the characteristics of the queries, and of the patterns of search behavior that precede posting a question; the relationship between the content of the attempted queries and of the posted questions; and the subsequent actions the user performs on the CQA site. Our work develops novel insights into searcher intent and behavior that lead to asking questions to the community, providing a foundation for more effective integration of automated web search and social information seeking. Qiaoling Liu, Eugene Agichtein, Gideon Dror, Yoelle Maarek, Idan Szpektor |
SIGIR | 5 |
| 2012 | Learning from the past: answering new questions with past answersabstractCommunity-based Question Answering sites, such as Yahoo! Answers or Baidu Zhidao, allow users to get answers to complex, detailed and personal questions from other users. However, since answering a question depends on the ability and willingness of users to address the asker's needs, a significant fraction of the questions remain unanswered. We measured that in Yahoo! Answers, this fraction represents 15% of all incoming English questions. At the same time, we discovered that around 25% of questions in certain categories are recurrent, at least at the question-title level, over a period of one year. Anna Shtok, Gideon Dror, Yoelle Maarek, Idan Szpektor |
WWW | 4 |
| 2011 | I want to answer; who has a question?: Yahoo! answers recommender systemabstractYahoo! Answers is currently one of the most popular question answering systems. We claim however that its user experience could be significantly improved if it could route the "right question" to the "right user." Indeed, while some users would rush answering a question such as "what should I wear at the prom?," others would be upset simply being exposed to it. We argue here that Community Question Answering sites in general and Yahoo! Answers in particular, need a mechanism that would expose users to questions they can relate to and possibly answer. Gideon Dror, Yehuda Koren, Yoelle Maarek, Idan Szpektor |
KDD | 4 |
| 2011 | Predicting web searcher satisfaction with existing community-based answersabstractCommunity-based Question Answering (CQA) sites, such as Yahoo! Answers, Baidu Knows, Naver, and Quora, have been rapidly growing in popularity. The resulting archives of posted answers to questions, in Yahoo! Answers alone, already exceed in size 1 billion, and are aggressively indexed by web search engines. In fact, a large number of search engine users benefit from these archives, by finding existing answers that address their own queries. This scenario poses new challenges and opportunities for both search engines and CQA sites. To this end, we formulate a new problem of predicting the satisfaction of web searchers with CQA answers. We analyze a large number of web searches that result in a visit to a popular CQA site, and identify unique characteristics of searcher satisfaction in this setting, namely, the effects of query clarity, query-to-question match, and answer quality. We then propose and evaluate several approaches to predicting searcher satisfaction that exploit these characteristics. To the best of our knowledge, this is the first attempt to predict and validate the usefulness of CQA archives for external searchers, rather than for the original askers. Our results suggest promising directions for improving and exploiting community question answering services in pursuit of satisfying even more Web search queries. Qiaoling Liu, Eugene Agichtein, Gideon Dror, Evgeniy Gabrilovich, Yoelle Maarek, Dan Pelleg, Idan Szpektor |
SIGIR | 7 |
| 2011 | Improving recommendation for long-tail queries via templatesabstractThe ability to aggregate huge volumes of queries over a large population of users allows search engines to build precise models for a variety of query-assistance features such as query recommendation, correction, etc. Yet, no matter how much data is aggregated, the long-tail distribution implies that a large fraction of queries are rare. As a result, most query assistance services perform poorly or are not even triggered on long-tail queries. We propose a method to extend the reach of query assistance techniques (and in particular query recommendation) to long-tail queries by reasoning about rules between query templates rather than individual query transitions, as currently done in query-flow graph models. As a simple example, if we recognize that 'Montezuma' is a city in the rare query "Montezuma surf" and if the rule 'city surf → beach has been observed, we are able to offer "Montezuma beach" as a recommendation, even if the two queries were never observed in a same session. We conducted experiments to validate our hypothesis, first via traditional small-scale editorial assessments but more interestingly via a novel automated large scale evaluation methodology. Our experiments show that general coverage can be relatively increased by 24% using templates without penalizing quality. Furthermore, for 36% of the 95M queries in our query flow graph, which have no out edges and thus could not be served recommendations, we can now offer at least one recommendation in 98% of the cases. Idan Szpektor, Aristides Gionis, Yoelle Maarek |
WWW | 1 |
| 2010 | Directional distributional similarity for lexical inferenceabstractAbstract Distributional word similarity is most commonly perceived as a symmetric relation. Yet, directional relations are abundant in lexical semantics and in many Natural Language Processing (NLP) settings that require lexical inference, making symmetric similarity measures less suitable for their identification. This paper investigates the nature of directional (asymmetric) similarity measures that aim to quantify distributional feature inclusion. We identify desired properties of such measures for lexical inference, specify a particular measure based on Average Precision that addresses these properties, and demonstrate the empirical benefit of directional measures for two different NLP datasets. Lili Kotlerman, Ido Dagan, Idan Szpektor, Maayan Zhitomirsky-Geffet |
Nat. Lang. Eng. | 3 |
| 2009 | Source-Language Entailment Modeling for Translating Unknown Terms
Shachar Mirkin, Lucia Specia, Nicola Cancedda, Ido Dagan, Marc Dymetman, Idan Szpektor |
ACL/IJCNLP | 6 |
| 2008 | Contextual Preferences
Idan Szpektor, Ido Dagan, Roy Bar-Haim, Jacob Goldberger |
ACL | 1 |
| 2008 | Natural Language as the Basis for Meaning Representation and Inference
Ido Dagan, Roy Bar-Haim, Idan Szpektor, Iddo Greental, Eyal Shnarch |
CICLing | 3 |
| 2008 | Learning Entailment Rules for Unary Templates
Idan Szpektor, Ido Dagan |
COLING | 1 |
| 2007 | Instance-based Evaluation of Entailment Rule Acquisition
Idan Szpektor, Eyal Shnarch, Ido Dagan |
ACL | 1 |
| 2006 | Investigating a Generic Paraphrase-Based Approach for Relation Extraction
Lorenza Romano, Milen Kouylekov, Idan Szpektor, Ido Dagan, Alberto Lavelli |
EACL | 3 |
| 2004 | Scaling Web-based Acquisition of Entailment Relations
Idan Szpektor, Hristo Tanev, Ido Dagan, Bonaventura Coppola |
EMNLP | 1 |