Haggai Roitman

dblp:94/4396 · DBLP profile ↗
← Back
48ranked-venue papers
17as first author
7since 2021 · last 2026
0000-0002-5260-2287ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 45 · 17 first-author · 7 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author
YearPublicationVenuePosition
2026 Sequential Recommendation with Generative Intent Prediction Utilizing User Search-Behavior
abstract
Sequential recommendation systems often struggle to accurately predict user preferences when limited to historical browsing data. We present a novel approach that combines recommendation systems with search engine methodologies, introducing a generative intent prediction model that leverages both item view histories and historical search queries. The model is enhanced by incorporating user interaction data from search engine result pages (SERP), leading to more accurate query predictions aligned with actual user behavior. By integrating this intent prediction model into sequential recommendation frameworks through a query expansion-inspired approach, we demonstrate significant performance improvements over traditional methods, particularly in challenging scenarios where conventional approaches fall short.
Guy Elovici, Bracha Shapira, Haggai Roitman, Yotam Eshel
WSDM3
2025 X-Cross: Dynamic Integration of Language Models for Cross-Domain Sequential Recommendation
abstract
As new products are emerging daily, recommendation systems are required to quickly adapt to possible new domains without needing extensive retraining. This work presents ''X-Cross'' -- a novel cross-domain sequential-recommendation model that recommends products in new domains by integrating several domain-specific language models; each model is fine-tuned with low-rank adapters (LoRA). Given a recommendation prompt, operating layer by layer, X-Cross dynamically refines the representation of each source language model by integrating knowledge from all other models. These refined representations are propagated from one layer to the next, leveraging the activations from each domain adapter to ensure domain-specific nuances are preserved while enabling adaptability across domains. Using Amazon datasets for sequential recommendation, X-Cross achieves performance comparable to a model that is fine-tuned with LoRA, while using only 25% of the additional parameters. In cross-domain tasks, such as adapting from Toys domain to Tools, Electronics or Sports, X-Cross demonstrates robust performance, while requiring about 50%-75% less fine-tuning data than LoRA to make fine-tuning effective. Furthermore, X-Cross achieves significant improvement in accuracy over alternative cross-domain baselines. Overall, X-Cross enables scalable and adaptive cross-domain recommendations, reducing computational overhead and providing an efficient solution for data-constrained environments.
Guy Hadad, Haggai Roitman, Yotam Eshel, Bracha Shapira, Lior Rokach
SIGIR2
2025 A Rank-Based Approach to Recommender System's Top-K Queries with Uncertain Scores
abstract
Top- K queries provide a ranked answer using a score that can either be given explicitly or computed from tuple values. Recommender systems use scores, based on user feedback on items with which they interact, to answer top- K queries. Such scores pose the challenge of correctly ranking elements using scores that are more often than not, uncertain. In this work, we address top- K queries based on uncertain scores. We propose to explicitly model the inherent uncertainty in the provided data and to consider a distribution of scores instead of a single score. Rooted in works of database probabilistic ranking, we offer the use of probabilistic ranking as a tool of choice for generating recommendation in the presence of uncertainty. We argue that the ranking approach should be chosen in a manner that maximizes user satisfaction, extending state-of-the-art on quality aspect of top- K answers over uncertain data, their relationship to top- K semantics, and improve ranking with uncertain scores in recommender systems. Towards this end, we introduce RankDist, an algorithm for efficiently computing probability of item position in a ranked recommendation. We show that rank-based (rather than score-based) methods that are computed using RankDist, which were not applied in recommender systems before, offer a guaranteed optimality by expectation and empirical superiority when tested on common benchmarks.
Coral Scharf, Carmel Domshlak, Avigdor Gal, Haggai Roitman
Proc. ACM Manag. Data4
2022 Sequential Modeling with Multiple Attributes for Watchlist Recommendation in E-Commerce
abstract
In e-commerce, the watchlist enables users to track items over time and has emerged as a primary feature, playing an important role in users' shopping journey. Watchlist items typically have multiple attributes whose values may change over time (e.g., price, quantity). Since many users accumulate dozens of items on their watchlist, and since shopping intents change over time, recommending the top watchlist items in a given context can be valuable. In this work, we study the watchlist functionality in e-commerce and introduce a novel watchlist recommendation task. Our goal is to prioritize which watchlist items the user should pay attention to next by predicting the next items the user will click. We cast this task as a specialized sequential recommendation task and discuss its characteristics. Our proposed recommendation model, Trans2D, is built on top of the Transformer architecture, where we further suggest a novel extended attention mechanism (Attention2D) that allows to learn complex item-item, attribute-attribute and item-attribute patterns from sequential-data with multiple item attributes. Using a large-scale watchlist dataset from eBay, we evaluate our proposed model, where we demonstrate its superiority compared to multiple state-of-the-art baselines, many of which are adapted for this task.
Uriel Singer, Haggai Roitman, Yotam Eshel, Alexander Nus, Ido Guy, Or Levi, Idan Hasson, Eliyahu Kiperwasser
WSDM2
2021 PreSizE: Predicting Size in E-Commerce using Transformers
abstract
Recent advances in the e-commerce fashion industry have led to an exploration of novel ways to enhance buyer experience via improved personalization. Predicting a proper size for an item to recommend is an important personalization challenge, and is being studied in this work. Earlier works in this field either focused on modeling explicit buyer fitment feedback or modeling of only a single aspect of the problem (e.g., specific category, brand, etc.). More recent works proposed richer models, either content-based or sequence-based, better accounting for content-based aspects of the problem or better modeling the buyer's online journey. However, both these approaches fail in certain scenarios: either when encountering unseen items (sequence-based models) or when encountering new users (content-based models).
Yotam Eshel, Or Levi, Haggai Roitman, Alexander Nus
SIGIR3
2021 Automatic Form Filling with Form-BERT
abstract
Digital-forms are commonly used for collecting structured information from users. However, filling digital-forms that include a large number of fields is tedious and error-prone. Auto-filling form fields for the user is highly beneficial for improving user experience and potentially collecting more valuable information (in cases where not all fields are mandatory). Online E-commerce marketplaces quite often utilize such forms to collect listing attributes from sellers. In this work, we describe Form-BERT -- a Transformer-based model which is optimized for auto-filling listing attributes given the following inputs: free-text, list of known attribute names, and zero or more attribute values. Form-BERT can be further used iteratively to leverage filled out attributes as the form filling progresses.
Gilad Fuchs, Haggai Roitman, Matan Mandelbrod
SIGIR2
2021 Learning to Rerank Schema Matches
abstract
Schema matching is at the heart of integrating structured and semi-structured data with applications in data warehousing, data analysis recommendations, Web table matching, etc. Schema matching is known as an uncertain process and a common method to overcome this uncertainty introduces a human expert with a ranked list of possible schema matches to choose from, known as top-Kmatching. In this work we propose a learning algorithm that utilizes an innovative set of features to rerank a list of schema matches and improves upon the ranking of the best match. We provide a bound on the size of an initial match list, tying the number of matches with a desired level of confidence in finding the best match. We also propose the use of matching predictors as features in a learning task, and tailored nine new matching predictors for this purpose. The proposed algorithm assists the matching process by introducing a quality set of alternative matches to a human expert. It also serves as a step towards eliminating the involvement of human experts as decision makers in a matching process altogether. A large scale empirical evaluation with real-world benchmark shows the effectiveness of the proposed algorithmic solution.
Avigdor Gal, Haggai Roitman, Roee Shraga
IEEE Trans. Knowl. Data Eng.2
2020 Unsupervised FAQ Retrieval with Question Generation and BERT
abstract
We focus on the task of Frequently Asked Questions (FAQ) retrieval.A given user query can be matched against the questions and/or the answers in the FAQ.We present a fully unsupervised method that exploits the FAQ pairs to train two BERT models.The two models match user queries to FAQ answers and questions, respectively.We alleviate the missing labeled data of the latter by automatically generating high-quality question paraphrases.We show that our model is on par and even outperforms supervised models on existing datasets.
Yosi Mass, Boaz Carmeli, Haggai Roitman, David Konopnicki
ACL3
2020 Conversational Document Prediction to Assist Customer Care Agents
abstract
Jatin Ganhotra, Haggai Roitman, Doron Cohen, Nathaniel Mills, Chulaka Gunasekara, Yosi Mass, Sachindra Joshi, Luis Lastras, David Konopnicki. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Jatin Ganhotra, Haggai Roitman, Doron Cohen 0001, Nathaniel Mills, R. Chulaka Gunasekara, Yosi Mass, Sachindra Joshi, Luis A. Lastras, David Konopnicki
EMNLP (1)2
2020 Ad-hoc Document Retrieval using Weak-Supervision with BERT and GPT2
abstract
We describe a weakly-supervised method for training deep learning models for the task of ad-hoc document retrieval.Our method is based on generative and discriminative models that are trained using weak-supervision based solely on the documents in the corpus.We present an end-to-end retrieval system that starts with traditional information retrieval methods, followed by two deep learning re-rankers.We evaluate our method on three different datasets: a COVID-19 related scientific literature dataset and two news datasets.We show that our method outperforms state-ofthe-art methods; this without the need for the expensive process of manually labeling data.
Yosi Mass, Haggai Roitman
EMNLP (1)2
2020 Web Table Retrieval using Multimodal Deep Learning
abstract
We address the web table retrieval task, aiming to retrieve and rank web tables as whole answers to a given information need. To this end, we formally define web tables as multimodal objects. We then suggest a neural ranking model, termed MTR, which makes a novel use of Gated Multimodal Units (GMUs) to learn a joint-representation of the query and the different table modalities. We further enhance this model with a co-learning approach which utilizes automatically learned query-independent and query-dependent "helper'' labels. We evaluate the proposed solution using both ad hoc queries (WikiTables) and natural language questions (GNQtables). Overall, we demonstrate that our approach surpasses the performance of previously studied state-of-the-art baselines.
Roee Shraga, Haggai Roitman, Guy Feigenblat, Mustafa Canim
SIGIR2
2020 Unsupervised Dual-Cascade Learning with Pseudo-Feedback Distillation for Query-Focused Extractive Summarization
abstract
We propose Dual-CES – a novel unsupervised, query-focused, multi-document extractive summarizer. Dual-CES builds on top of the Cross Entropy Summarizer (CES) and is designed to better handle the tradeoff between saliency and focus in summarization. To this end, Dual-CES employs a two-step dual-cascade optimization approach with saliency-based pseudo-feedback distillation. Overall, Dual-CES significantly outperforms all other state-of-the-art unsupervised alternatives. Dual-CES is even shown to be able to outperform strong supervised summarizers.
Haggai Roitman, Guy Feigenblat, Doron Cohen 0001, Odellia Boni, David Konopnicki
WWW1
2020 Ad Hoc Table Retrieval using Intrinsic and Extrinsic Similarities
abstract
Given a keyword query, the ad hoc table retrieval task aims at retrieving a ranked list of the top-k most relevant tables in a given table corpus. Previous works have primarily focused on designing table-centric lexical and semantic features, which could be utilized for learning-to-rank (LTR) tables. In this work, we make a novel use of intrinsic (passage-based) and extrinsic (manifold-based) table similarities for enhanced retrieval. Using the WikiTables benchmark, we study the merits of utilizing such similarities for this task. To this end, we combine both similarity types via a simple, yet an effective, cascade re-ranking approach. Overall, our proposed approach results in a significantly better table retrieval quality, which even transcends that of strong semantically-rich baselines.
Roee Shraga, Haggai Roitman, Guy Feigenblat, Mustafa Canim
WWW2
2020 ADnEV: Cross-Domain Schema Matching using Deep Similarity Matrix Adjustment and Evaluation
abstract
Schema matching is a process that serves in integrating structured and semi-structured data. Being a handy tool in multiple contemporary business and commerce applications, it has been investigated in the fields of databases, AI, Semantic Web, and data mining for many years. The core challenge still remains the ability to create quality algorithmic matchers, automatic tools for identifying correspondences among data concepts ( e.g. , database attributes). In this work, we offer a novel post processing step to schema matching that improves the final matching outcome without human intervention. We present a new mechanism, similarity matrix adjustment , to calibrate a matching result and propose an algorithm (dubbed ADnEV) that manipulates, using deep neural networks, similarity matrices, created by state-of-the-art algorithmic matchers. ADnEV learns two models that iteratively adjust and evaluate the original similarity matrix. We empirically demonstrate the effectiveness of the proposed algorithmic solution for improving matching results, using real-world benchmark ontology and schema sets. We show that ADnEV can generalize into new domains without the need to learn the domain terminology, thus allowing cross-domain learning. We also show ADnEV to be a powerful tool in handling schemata which matching is particularly challenging. Finally, we show the benefit of using ADnEV in a related integration task of ontology alignment.
Roee Shraga, Avigdor Gal, Haggai Roitman
Proc. VLDB Endow.3
2019 Normalized Query Commitment Revisited
abstract
We revisit the Normalized Query Commitment (NQC) query performance prediction (QPP) method. To this end, we suggest a scaled extension to a discriminative QPP framework and use it to analyze NQC. Using this analysis allows us to redesign NQC and suggest several options for improvement.
Haggai Roitman
SIGIR1
2019 Query Performance Prediction for Pseudo-Feedback-Based Retrieval
abstract
The query performance prediction task (QPP) is estimating retrieval effectiveness in the absence of relevance judgments. Prior work has focused on prediction for retrieval methods based on surface level query-document similarities (e.g., query likelihood). We address the prediction challenge for pseudo-feedback-based retrieval methods which utilize an initial retrieval to induce a new query model; the query model is then used for a second (final) retrieval. Our suggested approach accounts for the presumed effectiveness of the initially retrieved list, its similarity with the final retrieved list and properties of the latter. Empirical evaluation demonstrates the clear merits of our approach.
Haggai Roitman, Oren Kurland
SIGIR1
2018 Heterogeneous Data Integration by Learning to Rerank Schema Matches
abstract
Schema matching is a task at the heart of integrating heterogeneous structured and semi-structured data with applications in data warehousing, process matching, data analysis recommendations, Web table matching, etc. Schema matching is known to be an uncertain process and a common method of overcoming this uncertainty is by introducing a human expert with a ranked list of possible schema matches from which the expert may choose, known astop-Kmatching. In this work we propose a learning algorithm that utilizes an innovative set of features to rerank a list of schema matches and improves upon the ranking of the best match. The proposed algorithm assists the matching process by introducing a quality set of alternative matches to a human expert. It also serves as a step towards eliminating the involvement of human experts as decision makers in a matching process altogether. A large scale empirical evaluation with real-world benchmark shows the effectiveness of the proposed algorithmic solution.
Avigdor Gal, Haggai Roitman, Roee Shraga
ICDM2
2018 Query Performance Prediction using Passage Information
abstract
We focus on the post-retrieval query performance prediction (QPP) task. Specifically, we make a new use of passage information for this task. Using such information we derive a new mean score calibration predictor that provides a more accurate prediction. Using an empirical evaluation over several common TREC benchmarks, we show that, QPP methods that only make use of document-level features are mostly suited for short query prediction tasks; while such methods perform significantly worse in verbose query prediction settings. We further demonstrate that, QPP methods that utilize passage-information are much better suited for verbose settings. Moreover, our proposed predictor, which utilizes both document-level and passage-level features provides a more accurate and consistent prediction for both types of queries. Finally, we show a connection between our predictor and a recently proposed supervised QPP method, which results in an enhanced prediction.
Haggai Roitman
SIGIR1
2017 SummIt: A Tool for Extractive Summarization, Discovery and Analysis
abstract
We propose to demonstrate SummIt -- a tool for extractive summarization, discovery and analysis. The main goal of SummIt is to provide consumable summaries that are driven by users' information intents. To this end, SummIt discovers and analyzes potential intents that can be used for summarization. Given an intent, SummIt generates a summary based on a novel unsupervised, query-focused, extractive, multi-document summarization approach. Using visualization aids, SummIt further allows to analyze a given summary and explore both its narrow and broader context.
Guy Feigenblat, Odellia Boni, Haggai Roitman, David Konopnicki
CIKM3
2017 Unsupervised Query-Focused Multi-Document Summarization using the Cross Entropy Method
abstract
We present a novel unsupervised query-focused multi-document summarization approach. To this end, we generate a summary by extracting a subset of sentences using the Cross-Entropy (CE) Method. The proposed approach is generic and requires no domain knowledge. Using an evaluation over DUC 2005-2007 datasets with several other state-of-the-art baseline methods, we demonstrate that, our approach is both effective and efficient.
Guy Feigenblat, Haggai Roitman, Odellia Boni, David Konopnicki
SIGIR2
2017 An Extended Relevance Model for Session Search
abstract
The session search task aims at best serving the user's information need given her previous search behavior during the session. We propose an extended relevance model that captures the user's dynamic information need in the session. Our relevance modelling approach is directly driven by the user's query reformulation (change) decisions and the estimate of how much the user's search behavior affects such decisions. Overall, we demonstrate that, the proposed approach significantly boosts session search performance.
Nir Levine, Haggai Roitman, Doron Cohen 0001
SIGIR2
2017 An Enhanced Approach to Query Performance Prediction Using Reference Lists
abstract
We address the problem of query performance prediction (QPP) using reference lists. To date, no previous QPP method has been fully successful in generating and utilizing several pseudo-effective and pseudo-ineffective reference lists. In this work, we try to fill the gaps. We first propose a novel unsupervised approach for generating and selecting both types of reference lists using query perturbation and statistical inference. We then propose an enhanced QPP approach that utilizes both types of selected reference lists.
Haggai Roitman
SIGIR1
2016 RecSys'16 Workshop on Deep Learning for Recommender Systems (DLRS)
abstract
We believe that Deep Learning is one of the next big things in Recommendation Systems technology. The past few years have seen the tremendous success of deep neural networks in a number of complex tasks such as computer vision, natural language processing and speech recognition. Despite this, only little work has been published on Deep Learning methods for Recommender Systems. Notable recent application areas are music recommendation, news recommendation, and session-based recommendation. The aim of the workshop is to encourage the application of Deep Learning techniques in Recommender Systems, to promote research in deep learning methods for Recommender Systems, and to bring together researchers from the Recommender Systems and Deep Learning communities.
Alexandros Karatzoglou, Balázs Hidasi, Domonkos Tikk, Oren Sar Shalom, Haggai Roitman, Bracha Shapira, Lior Rokach
RecSys5
2016 When Watson Went to Work: Leveraging Cognitive Computing in the Real World
abstract
No abstract available.
Aya Soffer, David Konopnicki, Haggai Roitman
SIGIR3
2016 From Diversity-based Prediction to Better Ontology & Schema Matching
abstract
Ontology & schema matching predictors assess the quality of matchers in the absence of an exact match. We propose MCD (Match Competitor Deviation), a new diversity-based predictor that compares the strength of a matcher confidence in the correspondence of a concept pair with respect to other correspondences that involve either concept. We also propose to use MCD as a regulator to optimally control a balance between Precision and Recall and use it towards 1:1 matching by combining it with a similarity measure that is based on solving a maximum weight bipartite graph matching (MWBM). Optimizing the combined measure is known to be an NP-Hard problem. Therefore, we propose CEM, an approximation to an optimal match by efficiently scanning multiple possible matches, using rare event estimation. Using a thorough empirical study over several benchmark real-world datasets, we show that MCD outperforms other state-of-the-art predictor and that CEM significantly outperform existing matchers.
Avigdor Gal, Haggai Roitman, Tomer Sagi
WWW2
2014 Using the cross-entropy method to re-rank search results
abstract
We present a novel unsupervised approach to re-ranking an initially retrieved list. The approach is based on the Cross Entropy method applied to permutations of the list, and relies on performance prediction. Using pseudo predictors we establish a lower bound on the prediction quality that is required so as to have our approach significantly outperform the original retrieval. Our experiments serve as a proof of concept demonstrating the considerable potential of the proposed approach. A case in point, only a tiny fraction of the huge space of permutations needs to be explored to attain significant improvements over the original retrieval.
Haggai Roitman, Shay Hummel, Oren Kurland
SIGIR1
2014 A fusion approach to cluster labeling
abstract
We present a novel approach to the cluster labeling task using fusion methods. The core idea of our approach is to weigh labels, suggested by any labeler, according to the estimated labeler's decisiveness with respect to each of its suggested labels. We hypothesize that, a cluster labeler's labeling choice for a given cluster should remain stable even in the presence of a slightly incomplete cluster data. Using state-of-the-art cluster labeling and data fusion methods, evaluated over a large data collection of clusters, we demonstrate that, overall, the cluster labeling fusion methods that further consider the labeler's decisiveness provide the best labeling performance.
Haggai Roitman, Shay Hummel, Michal Shmueli-Scheuer
SIGIR1
2013 Modeling the uniqueness of the user preferences for recommendation systems
abstract
In this paper we propose a novel framework for modeling the uniqueness of the user preferences for recommendation systems. User uniqueness is determined by learning to what extent the user's item preferences deviate from those of an "average user" in the system. Based on this framework, we suggest three different recommendation strategies that trade between uniqueness and conformity. Using two real item datasets, we demonstrate the effectiveness of our uniqueness based recommendation framework.
Haggai Roitman, David Carmel, Yosi Mass, Iris Eiron
SIGIR1
2012 Workshop on multimodal crowd sensing (CrowdSens 2012)
abstract
This paper provides an overview of the 1st International Workshop on Multimodal Crowd Sensing (CrowdSens 2012), held at the 21st ACM International Conference on Information and Knowledge Management (CIKM 2012). This workshop aimed to provide an open forum for researchers from various fields such as fields such as Natural Language Processing, Information Extraction, Data Mining, Information Retrieval, User Modeling and Personalization, Stream Processing, and Sensor Networks, for addressing the challenges of effectively mining, analyzing, fusing, and exploiting information sourced from multimodal physical and social sensor data sources.
Haggai Roitman, Iván Cantador, Miriam Fernández
CIKM1
2012 Bridging the Gaps towards Advanced Data Discovery over Semi-structured Data
Sivan Yogev, Haggai Roitman
ER2
2012 Surfacing time-critical insights from social media
abstract
We propose to demonstrate an end-to-end framework for leveraging time-sensitive and critical social media information for businesses. More specifically, we focus on identifying, structuring, integrating, and exposing timely insights that are essential to marketing services and monitoring reputation over social media. Our system includes components for information extraction from text, entity resolution and integration, analytics, and a user interface.
Alexe Dumitru-Bogdan, Mauricio A. Hernández, Kirsten Hildrum, Rajasekar Krishnamurthy, Georgia Koutrika, Meena Nagarajan, Haggai Roitman, Michal Shmueli-Scheuer, Ioana Stanoi, Chitra Venkatramani, Rohit Wagle
SIGMOD Conference7
2012 On the Relationship between Novelty and Popularity of User-Generated Content
abstract
This work deals with the task of predicting the popularity of user-generated content. We demonstrate how the novelty of newly published content plays an important role in affecting its popularity. More specifically, we study three dimensions of novelty. The first one, termed contemporaneous novelty , models the relative novelty embedded in a new post with respect to contemporary content that was generated by others. The second type of novelty, termed self novelty , models the relative novelty with respect to the user’s own contribution history. The third type of novelty, termed discussion novelty , relates to the novelty of the comments associated by readers with respect to the post content. We demonstrate the contribution of the new novelty measures to estimating blog-post popularity by predicting the number of comments expected for a fresh post. We further demonstrate how novelty based measures can be utilized for predicting the citation volume of academic papers.
David Carmel, Haggai Roitman, Elad Yom-Tov
ACM Trans. Intell. Syst. Technol.2
2012 Folksonomy-Based Term Extraction for Word Cloud Generation
abstract
In this work we study the task of term extraction for word cloud generation in sparsely tagged domains, in which manual tags are scarce. We present a folksonomy-based term extraction method, called tag-boost , which boosts terms that are frequently used by the public to tag content. Our experiments with tag-boost based term extraction over different domains demonstrate tremendous improvement in word cloud quality, as reflected by the agreement between manual tags of the testing items and the cloud’s terms extracted from the items’ content. Moreover, our results demonstrate the high robustness of this approach, as compared to alternative cloud generation methods that exhibit a high sensitivity to data sparseness. Additionally, we show that tag-boost can be effectively applied even in nontagged domains, by using an external rich folksonomy borrowed from a well-tagged domain.
David Carmel, Erel Uziel, Ido Guy, Yosi Mass, Haggai Roitman
ACM Trans. Intell. Syst. Technol.5
2011 Folksonomy-based term extraction for word cloud generation
abstract
In this work we study the task of term extraction for word cloud generation. We present a folksonomy-based term extraction method, called tag-boost, which boosts terms that are frequently used by the public to tag content. Our experiments with tag-boost-based term extraction over different domains demonstrate tremendous improvement in word cloud quality, as reflected by the agreement between extracted terms and manually assigned tags of the testing items. Additionally, we show that tag-boost can be effectively applied even in non-tagged domains, by using an external rich folksonomy borrowed from a well-tagged domain.
David Carmel, Erel Uziel, Ido Guy, Yosi Mass, Haggai Roitman
CIKM5
2011 Search and mining entity-relationship data
abstract
This paper summarizes the details of the first international workshop on search and mining entity-relationship data. This workshop will bridge between IR, DB, and KM researchers to seek novel solutions for search and data mining of rich entity-relationship data and their applications in various domains. We first provide an overview about the workshop. We then briefly discuss the workshop program.
Haggai Roitman, Ralf Schenkel, Marko Grobelnik
CIKM1
2011 Exploratory search over social-medical data
abstract
In this demo we shall present the IBM Patient Empowerment System (PES), and more specifically, its social-medical discovery sub-system. Social and medical data are represented using entities and relationships and are explored using a combination of expressive, yet intuitive, query language, faceted search, and ER graph navigation. While this demonstration focuses on the healthcare domain, the underlining search technology is generic and can be utilized in many other domains. Therefore, this demo has two main contributions. First, we present a novel entity-relationship indexing and retrieval solution, and discuss its implementation challenges. Second, the demonstration depicts a practical entity-relationship discovery technology in a real domain setting within a real IBM system.
Haggai Roitman, Sivan Yogev, Yevgenia Tsimerman, Dae Won Kim, Yossi Mesika
CIKM1
2011 A Dual Framework and Algorithms for Targeted Online Data Delivery
abstract
A variety of emerging online data delivery applications challenge existing techniques for data delivery to human users, applications, or middleware that are accessing data from multiple autonomous servers. In this paper, we develop a framework for formalizing and comparing pull-based solutions and present dual optimization approaches. The first approach, most commonly used nowadays, maximizes user utility under the strict setting of meeting a priori constraints on the usage of system resources. We present an alternative and more flexible approach that maximizes user utility by satisfying all users. It does this while minimizing the usage of system resources. We discuss the benefits of this latter approach and develop an adaptive monitoring solution Satisfy User Profiles (SUPs). Through formal analysis, we identify sufficient optimality conditions for SUP. Using real (RSS feeds) and synthetic traces, we empirically analyze the behavior of SUP under varying conditions. Our experiments show that we can achieve a high degree of satisfaction of user utility when the estimations of SUP closely estimate the real event stream, and has the potential to save a significant amount of system resources. We further show that SUP can exploit feedback to improve user utility with only a moderate increase in resource utilization.
Haggai Roitman, Avigdor Gal, Louiqa Raschid
IEEE Trans. Knowl. Data Eng.1
2010 On the relationship between novelty and popularity of user-generated content
abstract
This work deals with the task of predicting the popularity of user-generated content. We demonstrate how the novelty of newly published content plays an important role in affecting its popularity. We study three dimensions of novelty: contemporaneous novelty, self novelty, and discussion novelty. We demonstrate the contribution of the new novelty measures to estimating blog-post popularity by predicting the number of comments expected for a fresh post. We further demonstrate how novelty based measures can be utilized for predicting the citation volume of academic papers.
David Carmel, Haggai Roitman, Elad Yom-Tov
CIKM2
2010 Social bookmark weighting for search and recommendation
David Carmel, Haggai Roitman, Elad Yom-Tov
VLDB J.2
2009 Who tags the tags?: a framework for bookmark weighting
abstract
In this work we propose a novel framework for bookmark weighting which allows us to estimate the effectiveness of each of the bookmarks individually. We show that by weighting bookmarks according to their estimated quality we can significantly improve search effectiveness. Using empirical evaluation on real data gathered from two large bookmarking systems, we demonstrate the effectiveness of the new framework for search enhancement.
David Carmel, Haggai Roitman, Elad Yom-Tov
CIKM2
2009 Web Monitoring 2.0: Crossing Streams to Satisfy Complex Data Needs
abstract
Web monitoring 2.0 supports the complex information needs of clients who probe multiple information sources and generate mashups by integrating across these volatile streams. A proxy that aims at satisfying multiple customized client profiles will face a scalability challenge in trying to maximize the number of clients served while at the same time fully satisfying complex client needs. In this paper, we introduce an abstraction of complex execution intervals, a combination of time intervals and information streams, to capture complex client needs. Given some budgetary constraints (e.g., bandwidth), we present offline algorithmic solutions for the problem of maximizing completeness of capturing complex profiles.
Haggai Roitman, Avigdor Gal, Louiqa Raschid
ICDE1
2009 Best-Effort Top-k Query Processing Under Budgetary Constraints
abstract
We consider a novel problem of top-k query processing under budget constraints. We provide both a framework and a set of algorithms to address this problem. Existing algorithms for top-k processing are budget-oblivious, i.e., they do not take budget constraints into account when making scheduling decisions, but focus on the performance to compute the final top-k results. Under budget constraints, these algorithms therefore often return results that are a lot worse than the results that can be achieved with a clever, budget-aware scheduling algorithm. This paper introduces novel algorithms for budget-aware top-k processing that produce results that have a significantly higher quality than those of state-of-the-art budget-oblivious solutions.
Michal Shmueli-Scheuer, Chen Li 0001, Yosi Mass, Haggai Roitman, Ralf Schenkel, Gerhard Weikum
ICDE4
2009 Enhancing cluster labeling using wikipedia
abstract
This work investigates cluster labeling enhancement by utilizing Wikipedia, the free on-line encyclopedia. We describe a general framework for cluster labeling that extracts candidate labels from Wikipedia in addition to important terms that are extracted directly from the text. The "labeling quality" of each candidate is then evaluated by several independent judges and the top evaluated candidates are recommended for labeling.
David Carmel, Haggai Roitman, Naama Zwerdling
SIGIR2
2008 Providing Top-K Alternative Schema Matchings with
Haggai Roitman, Avigdor Gal, Carmel Domshlak
ER1
2008 Satisfying Complex Data Needs using Pull-Based Online Monitoring of Volatile Data Sources
abstract
Emerging applications on the Web require better management of volatile data in pull-based environments. In a pull based setting, data may be periodically removed from the server. Data may also become obsolete, no longer serving client needs. In both cases, we consider such data to be volatile. To model such constraints on data usability, and support complex user needs we define profiles to specify which data sources are to be monitored and when. Using a novel abstraction of execution intervals we model complex profiles that access simultaneously several servers to gain from the used data. Given some budgetary constraints (e.g., bandwidth), the paper formalizes the problem of maximizing completeness.
Haggai Roitman, Avigdor Gal, Louiqa Raschid
ICDE1
2008 Capturing Approximated Data Delivery Tradeoffs
abstract
This paper presents a middleware data delivery setting with a proxy that is required to maximize the completeness of captured updates, specified in its clients' profiles, while minimizing at the same time the delay in delivering the updates to clients. The two objectives may conflict when the monitoring budget is limited. Therefore, any solution should consider this tradeoff in satisfying both objectives. We term this problem the "proxy dilemma" and formalize it as a biobjective optimization problem. Such problem occurs in many contemporary applications, such as mobile and sensor networks, and poses scalability challenges in delivering up-to-date data from remote resources to meet client specifications. We present a Pareto set as a formal solution to the proxy dilemma. We discuss the complexity of generating a Pareto set for the proxy dilemma and suggest an approximation scheme to this problem.
Haggai Roitman, Avigdor Gal, Louiqa Raschid
ICDE1
2008 Maintaining dynamic channel profiles on the web
abstract
This work addresses a novel problem of maintaining channel proflies on the Web. Such channel maintenance is essential for next generation of Web 2.0 applications that provide sophisticated search and discovery services over Web information channels. Maintaining a fresh channel profile is extremely difficult due to the the dynamic nature of the channel, especially under the constraint of a limited monitoring budget. We propose a novel monitoring scheme that learns the channels' monitoring rates. The monitoring scheme is further extended to consider the content that is published on the channels. We describe a novelty detection filter that refines the monitoring rate according to the expected rate of novel content published on the channels. We further show how inter-channel profile similarities can be utilized to refine the channel monitoring rates. Using real-world data of Web feeds we study the performance of the monitoring scheme. We experiment with several monitoring policies over a large set of Web feeds and show that a policy based on learning the monitoring rate of the channels, combined with novelty detection, outperforms alternative channel monitoring policies. Our results show that the suggested content-based policy is able to maintain high quality channel profiles under limited monitoring resources.
Haggai Roitman, David Carmel, Elad Yom-Tov
Proc. VLDB Endow.1
2007 Rank Aggregation for Automatic Schema Matching
abstract
Schema matching is a basic operation of data integration, and several tools for automating it have been proposed and evaluated in the database community. Research in this area reveals that there is no single schema matcher that is guaranteed to succeed in finding a good mapping for all possible domains and, thus, an ensemble of schema matchers should be considered. In this paper, we introduce schema metamatching, a general framework for composing an arbitrary ensemble of schema matchers and generating a list of best ranked schema mappings. Informally, schema metamatching stands for computing a "consensus" ranking of alternative mappings between two schemata, given the "individual" graded rankings provided by several schema matchers. We introduce several algorithms for this problem, varying from adaptations of some standard techniques for general quantitative rank aggregation to novel techniques specific to the problem of schema matching, and to combinations of both. We provide a formal analysis of the applicability and relative performance of these algorithms and evaluate them empirically on a set of real-world schemata
Carmel Domshlak, Avigdor Gal, Haggai Roitman
IEEE Trans. Knowl. Data Eng.3