Yosi Mass

dblp:23/1530 · DBLP profile ↗
← Back
24ranked-venue papers
10as first author
2since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 9 first-author · 2 since 2021Databases, data management, data science and information retrieval · 14 · 6 first-authorSecurity and privacy · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorTheory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
9 papers
Information retrieval · 82% Machine learning and data management · 9% Recommender systems · 4%
Artificial intelligence
5 papers
Question answering and dialogue systems · 53% Language models and text generation · 22% Generative modeling · 16%

Topics — the 30 heaviest of 45, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems
long-form question answering
0.812024
Achieving Human Parity in Content-Grounded Datasets Generation · ICLR 2024
Machine learning › Generative modeling
synthetic data generation
0.812024
Achieving Human Parity in Content-Grounded Datasets Generation · ICLR 2024
Natural language and speech › Language models and text generation
text summarization
0.812024
Achieving Human Parity in Content-Grounded Datasets Generation · ICLR 2024
Natural language and speech › Question answering and dialogue systems › intent detection
few-shot intent detection
0.712023
QAID: Question Answering Inspired Few-shot Intent Detection · ICLR 2023
Natural language and speech › Question answering and dialogue systems
intent detection
0.712023
QAID: Question Answering Inspired Few-shot Intent Detection · ICLR 2023
Natural language and speech › Question answering and dialogue systems
conversational search
0.412020
Conversational Document Prediction to Assist Customer Care Agents · EMNLP (1) 2020
Natural language and speech › Information extraction and text analysis › information retrieval
document ranking
0.412020
Conversational Document Prediction to Assist Customer Care Agents · EMNLP (1) 2020
Information retrieval › retrieval models
ad-hoc retrieval
0.412020
Ad-hoc Document Retrieval using Weak-Supervision with BERT and GPT2 · EMNLP (1) 2020
Information retrieval › question answering
FAQ retrieval
0.412020
Unsupervised FAQ Retrieval with Question Generation and BERT · ACL 2020
Information retrieval › reranking
neural re-ranking
0.412020
Ad-hoc Document Retrieval using Weak-Supervision with BERT and GPT2 · EMNLP (1) 2020
Information retrieval
question answering
0.412020
Unsupervised FAQ Retrieval with Question Generation and BERT · ACL 2020
Machine learning and data management
weak supervision
0.412020
Ad-hoc Document Retrieval using Weak-Supervision with BERT and GPT2 · EMNLP (1) 2020
Information retrieval › document processing › document analysis
document representation
0.212016
Semantic Documents Relatedness using Concept Graph Representation · WSDM 2016
Information retrieval › similarity measure
document similarity
0.212016
Semantic Documents Relatedness using Concept Graph Representation · WSDM 2016
Information retrieval › text analysis
semantic relatedness
0.212016
Semantic Documents Relatedness using Concept Graph Representation · WSDM 2016
Recommender systems › user modeling
user preference modeling
0.212013
Modeling the uniqueness of the user preferences for recommendation systems · SIGIR 2013
Information retrieval
retrieval models
0.222012
Language models for keyword search over data graphs · WSDM 2012
Searching XML documents via XML fragments · SIGIR 2003
Information retrieval › reranking
answer re-ranking
0.112012
Language models for keyword search over data graphs · WSDM 2012
Information retrieval
keyword search
0.112012
Language models for keyword search over data graphs · WSDM 2012
Information retrieval
ranking
0.112012
Language models for keyword search over data graphs · WSDM 2012
Natural language and speech › Language models and text generation › pre-trained language model
BERT
0.112020
Unsupervised FAQ Retrieval with Question Generation and BERT · ACL 2020
Natural language and speech › Language models and text generation
pre-trained language model
0.112020
Ad-hoc Document Retrieval using Weak-Supervision with BERT and GPT2 · EMNLP (1) 2020
Information retrieval
distributed information retrieval
0.112011
KMV-peer: a robust and adaptive peer-selection algorithm · WSDM 2011
Information retrieval › distributed information retrieval
peer selection
0.112011
KMV-peer: a robust and adaptive peer-selection algorithm · WSDM 2011
Information retrieval › distributed information retrieval
peer-to-peer search
0.112011
KMV-peer: a robust and adaptive peer-selection algorithm · WSDM 2011
Information retrieval › distributed information retrieval
resource selection
0.112011
KMV-peer: a robust and adaptive peer-selection algorithm · WSDM 2011
Information retrieval
image retrieval
0.112009
A unified inverted index for an efficient image and text retrieval · SIGIR 2009
Information retrieval › indexing
inverted index
0.112009
A unified inverted index for an efficient image and text retrieval · SIGIR 2009
Information retrieval
multimedia analysis and retrieval
0.112009
A unified inverted index for an efficient image and text retrieval · SIGIR 2009
Query processing and optimization
query scheduling
0.112009
Best-Effort Top-k Query Processing Under Budgetary Constraints · ICDE 2009

Methods — techniques the papers use, named apart from their topics

BERT · 1.7weak supervision · 0.9question generation · 0.9generative and discriminative models · 0.9GPT-2 · 0.9large language model · 0.8few-shot learning · 0.7neural retrieval · 0.4neural network embedding · 0.2graph similarity · 0.2closeness centrality · 0.2uniqueness modeling · 0.2language modeling · 0.1trust policy language · 0.0certificate chain collection · 0.0modular architectural representation · 0.0incremental loading · 0.0
YearPublicationVenuePosition
2024 Achieving Human Parity in Content-Grounded Datasets Generation
abstract
The lack of high-quality data for content-grounded generation tasks has been identified as a major obstacle to advancing these tasks. To address this gap, we propose Genie, a novel method for automatically generating high-quality content-grounded data. It consists of three stages: (a) Content Preparation, (b) Generation: creating task-specific examples from the content (e.g., question-answer pairs or summaries). (c) Filtering mechanism aiming to ensure the quality and faithfulness of the generated data. We showcase this methodology by generating three large-scale synthetic data, making wishes, for Long-Form Question-Answering (LFQA), summarization, and information extraction. In a human evaluation, our generated data was found to be natural and of high quality. Furthermore, we compare models trained on our data with models trained on human-written data -- ELI5 and ASQA for LFQA and CNN-DailyMail for Summarization. We show that our models are on par with or outperforming models trained on human-generated data and consistently outperforming them in faithfulness. Finally, we applied our method to create LFQA data within the medical domain and compared a model trained on it with models trained on other domains.
Asaf Yehudai, Boaz Carmeli, Yosi Mass, Ofir Arviv, Nathaniel Mills, Eyal Shnarch, Leshem Choshen
ICLR3
2023 QAID: Question Answering Inspired Few-shot Intent Detection
Asaf Yehudai, Matan Vetzler, Yosi Mass, Koren Lazar, Doron Cohen 0001, Boaz Carmeli
ICLR3
2020 Unsupervised FAQ Retrieval with Question Generation and BERT
abstract
We focus on the task of Frequently Asked Questions (FAQ) retrieval.A given user query can be matched against the questions and/or the answers in the FAQ.We present a fully unsupervised method that exploits the FAQ pairs to train two BERT models.The two models match user queries to FAQ answers and questions, respectively.We alleviate the missing labeled data of the latter by automatically generating high-quality question paraphrases.We show that our model is on par and even outperforms supervised models on existing datasets.
Yosi Mass, Boaz Carmeli, Haggai Roitman, David Konopnicki
ACL1
2020 Conversational Document Prediction to Assist Customer Care Agents
abstract
Jatin Ganhotra, Haggai Roitman, Doron Cohen, Nathaniel Mills, Chulaka Gunasekara, Yosi Mass, Sachindra Joshi, Luis Lastras, David Konopnicki. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Jatin Ganhotra, Haggai Roitman, Doron Cohen 0001, Nathaniel Mills, R. Chulaka Gunasekara, Yosi Mass, Sachindra Joshi, Luis A. Lastras, David Konopnicki
EMNLP (1)6
2020 Ad-hoc Document Retrieval using Weak-Supervision with BERT and GPT2
abstract
We describe a weakly-supervised method for training deep learning models for the task of ad-hoc document retrieval.Our method is based on generative and discriminative models that are trained using weak-supervision based solely on the documents in the corpus.We present an end-to-end retrieval system that starts with traditional information retrieval methods, followed by two deep learning re-rankers.We evaluate our method on three different datasets: a COVID-19 related scientific literature dataset and two news datasets.We show that our method outperforms state-ofthe-art methods; this without the need for the expensive process of manually labeling data.
Yosi Mass, Haggai Roitman
EMNLP (1)1
2018 Word Emphasis Prediction for Expressive Text to Speech
Yosi Mass, Slava Shechtman, Moran Mordechay, Ron Hoory, Oren Sar Shalom, Guy Lev, David Konopnicki
INTERSPEECH1
2018 Semantic Relatedness of Wikipedia Concepts - Benchmark Data and a Working Solution
Liat Ein-Dor, Alon Halfon, Yoav Kantor, Ran Levy 0001, Yosi Mass, Ruty Rinott, Eyal Shnarch, Noam Slonim
LREC5
2016 Semantic Documents Relatedness using Concept Graph Representation
abstract
We deal with the problem of document representation for the task of measuring semantic relatedness between documents. A document is represented as a compact concept graph where nodes represent concepts extracted from the document through references to entities in a knowledge base such as DBpedia. Edges represent the semantic and structural relationships among the concepts. Several methods are presented to measure the strength of those relationships. Concepts are weighted through the concept graph using closeness centrality measure which reflects their relevance to the aspects of the document. A novel similarity measure between two concept graphs is presented. The similarity measure first represents concepts as continuous vectors by means of neural networks. Second, the continuous vectors are used to accumulate pairwise similarity between pairs of concepts while considering their assigned weights. We evaluate our method on a standard benchmark for document similarity. Our method outperforms state-of-the-art methods including ESA (Explicit Semantic Annotation) while our concept graphs are much smaller than the concept vectors generated by ESA. Moreover, we show that by combining our concept graph with ESA, we obtain an even further improvement.
Yuan Ni, Qiongkai Xu, Yosi Mass, Dafna Sheinwald, Huijia Zhu, Shao Sheng Cao
WSDM4
2014 Knowledge Management for Keyword Search over Data Graphs
abstract
This demo presents exploratory keyword search over data graphs by means of semantic facets. The demo starts with a keyword search over data graphs. Answers are first ranked by an existing search engine that considers their textual relevance and semantic structure. The user can then explore the answers through facets of structural patterns (i.e., schemas) as well as through other features. A particular way of presenting answers in a compact form is also supported and is applicable when looking for a single entity that connects the keywords. The demo is based on a working prototype that users can try on their own. It includes five data graphs that are quite diversified. In particular, three of them were generated from relational databases and two - from RDF triples. The demo shows that the system enables users to easily and quickly perform various search tasks by means of exploration, filtering and summarization.
Yosi Mass, Yehoshua Sagiv
CIKM1
2013 Modeling the uniqueness of the user preferences for recommendation systems
abstract
In this paper we propose a novel framework for modeling the uniqueness of the user preferences for recommendation systems. User uniqueness is determined by learning to what extent the user's item preferences deviate from those of an "average user" in the system. Based on this framework, we suggest three different recommendation strategies that trade between uniqueness and conformity. Using two real item datasets, we demonstrate the effectiveness of our uniqueness based recommendation framework.
Haggai Roitman, David Carmel, Yosi Mass, Iris Eiron
SIGIR3
2012 Language models for keyword search over data graphs
abstract
In keyword search over data graphs, an answer is a non-redundant subtree that includes the given keywords. This paper focuses on improving the effectiveness of that type of search. A novel approach that combines language models with structural relevance is described. The proposed approach consists of three steps. First, language models are used to assign dynamic, query-dependent weights to the graph. Those weights complement static weights that are pre-assigned to the graph. Second, an existing algorithm returns candidate answers based on their weights. Third, the candidate answers are re-ranked by creating a language model for each one. The effectiveness of the proposed approach is verified on a benchmark of three datasets: IMDB, Wikipedia and Mondial. The proposed approach outperforms all existing systems on the three datasets, which is a testament to its robustness. It is also shown that the effectiveness can be further improved by augmenting keyword queries with very basic knowledge about the structure.
Yosi Mass, Yehoshua Sagiv
WSDM1
2012 Folksonomy-Based Term Extraction for Word Cloud Generation
abstract
In this work we study the task of term extraction for word cloud generation in sparsely tagged domains, in which manual tags are scarce. We present a folksonomy-based term extraction method, called tag-boost , which boosts terms that are frequently used by the public to tag content. Our experiments with tag-boost based term extraction over different domains demonstrate tremendous improvement in word cloud quality, as reflected by the agreement between manual tags of the testing items and the cloud’s terms extracted from the items’ content. Moreover, our results demonstrate the high robustness of this approach, as compared to alternative cloud generation methods that exhibit a high sensitivity to data sparseness. Additionally, we show that tag-boost can be effectively applied even in nontagged domains, by using an external rich folksonomy borrowed from a well-tagged domain.
David Carmel, Erel Uziel, Ido Guy, Yosi Mass, Haggai Roitman
ACM Trans. Intell. Syst. Technol.4
2011 IQ: The Case for Iterative Querying for Knowledge
Yosi Mass, Maya Ramanath, Yehoshua Sagiv, Gerhard Weikum
CIDR1
2011 Folksonomy-based term extraction for word cloud generation
abstract
In this work we study the task of term extraction for word cloud generation. We present a folksonomy-based term extraction method, called tag-boost, which boosts terms that are frequently used by the public to tag content. Our experiments with tag-boost-based term extraction over different domains demonstrate tremendous improvement in word cloud quality, as reflected by the agreement between extracted terms and manually assigned tags of the testing items. Additionally, we show that tag-boost can be effectively applied even in non-tagged domains, by using an external rich folksonomy borrowed from a well-tagged domain.
David Carmel, Erel Uziel, Ido Guy, Yosi Mass, Haggai Roitman
CIKM4
2011 KMV-peer: a robust and adaptive peer-selection algorithm
abstract
The problem of fully decentralized search over many collections is considered. The objective is to approximate the results of centralized search (namely, using a central index) while controlling the communication cost and involving only a small number of collections. The proposed solution is couched in a peer-to-peer (P2P) network, but can also be applied in other setups. Peers publish per-term summaries of their collections. Specifically, for each term, the range of document scores is divided into intervals; and for each interval, a KMV (K Minimal Values) synopsis of its documents is created. A new peer-selection algorithm uses the KMV synopses and two scoring functions in order to adaptively rank the peers, according to the relevance of their documents to a given query. The proposed method achieves high-quality results while meeting the above criteria of efficiency. In particular, experiments are done on two large, real-world datasets; one is blogs and the other is web data. These experiments show that the algorithm outperforms the state-of-the-art approaches and is robust over different collections, various scoring functions and multi-term queries.
Yosi Mass, Yehoshua Sagiv, Michal Shmueli-Scheuer
WSDM1
2010 A peer-selection algorithm for information retrieval
abstract
A novel method for creating collection summaries is developed, and a fully decentralized peer-selection algorithm is described. This algorithm finds the most promising peers for answering a given query. Specifically, peers publish per-term synopses of their documents. The synopses of a peer for a given term are divided into score intervals and for each interval, a KMV (K Minimal Values) synopsis of its documents is created. The synopses are used to effectively rank peers by their relevance to a multi-term quer. The proposed approach is verified by experiments on a large real-world dataset. In particular, two collections were created from this dataset, each with a different number of peers. Compared to the state-of-the-art approaches, the proposed method is effective and efficient even when documents are randomly distributed among peers
Yosi Mass, Yehoshua Sagiv, Michal Shmueli-Scheuer
CIKM1
2009 A scalable and effective full-text search in P2P networks
abstract
We consider the problem of full-text search involving multi-term queries in a network of self-organizing, autonomous peers. Existing approaches do not scale well with respect to the number of peers, because they either require access to a large number of peers or incur a high communication cost in order to achieve good query results. In this paper, we present a novel algorithmic framework for processing multi-term queries in P2P networks that achieves high recall while using (per-query) a small number of peers and a low communication cost, thereby enabling high query throughput. Our approach is based on per-query peer-selection strategy using two-dimensional histograms of score distributions. A full utilization of the histograms incurs a high communication cost. We show how to drastically reduce this cost by employing a two-phase peer-selection algorithm. We also describe an adaptive approach to peer selection that further increases the recall. Experiments on a large real-world collection show that the recall is indeed high while the number of involved peers and the communication cost are low.
Yosi Mass, Yehoshua Sagiv, Michal Shmueli-Scheuer
CIKM1
2009 Best-Effort Top-k Query Processing Under Budgetary Constraints
abstract
We consider a novel problem of top-k query processing under budget constraints. We provide both a framework and a set of algorithms to address this problem. Existing algorithms for top-k processing are budget-oblivious, i.e., they do not take budget constraints into account when making scheduling decisions, but focus on the performance to compute the final top-k results. Under budget constraints, these algorithms therefore often return results that are a lot worse than the results that can be achieved with a clever, budget-aware scheduling algorithm. This paper introduces novel algorithms for budget-aware top-k processing that produce results that have a significantly higher quality than those of state-of-the-art budget-oblivious solutions.
Michal Shmueli-Scheuer, Chen Li 0001, Yosi Mass, Haggai Roitman, Ralf Schenkel, Gerhard Weikum
ICDE3
2009 A unified inverted index for an efficient image and text retrieval
abstract
We present an efficient method for approximate search in a combination of several metric spaces -- which are a generalization of low level image features -- using an inverted index. Our approximation gives very high recall with subsecond response time on a real data set of one million images extracted from Flickr.
Jonathan Mamou, Yosi Mass, Michal Shmueli-Scheuer, Benjamin Sznajder
SIGIR2
2007 Just in time indexing for up to the second search
abstract
E-commerce and intranet search systems require newly arriving content to be indexed and made available for search within minutes or hours of arrival. Applications such as file system and email search demand even faster turnaround from search systems, requiring new content to become available for search almost instantaneously. However, incrementally updating inverted indices, which are the predominant datastructure used in search engines, is an expensive operation that most systems avoid performing at high rates.
Ronny Lempel, Yosi Mass, Shila Ofek-Koifman, Dafna Sheinwald, Yael Petruschka, Ron Sivan
CIKM2
2003 Searching XML documents via XML fragments
abstract
Most of the work on XML query and search has stemmed from the publishing and database communities, mostly for the needs of business applications. Recently, the Information Retrieval community began investigating the XML search issue to answer information discovery needs. Following this trend, we present here an approach where information needs can be expressed in an approximate manner as pieces of XML documents or "XML fragments" of the same nature as the documents that are being searched. We present an extension of the vector space model for searching XML collections via XML fragments and ranking results by relevance. We describe how we have extended a full-text search engine to comply with this model. The value of the proposed method is demonstrated by the relative high precision of our system, which was among the top performers in the recent INEX workshop. Our results indicate that certain queries are more appropriate than others for the extended vector space model. Specifically, queries with relatively specific contexts but vague information needs are best situated to reap the benefit of this model. Finally our results show that one method may not fit all types of queries and that it could be worthwhile to use different solutions for different applications.
David Carmel, Yoelle Maarek, Matan Mandelbrod, Yosi Mass, Aya Soffer
SIGIR4
2001 Relying Party Credentials Framework
Amir Herzberg, Yosi Mass
CT-RSA2
2000 Access Control Meets Public Key Infrastructure, Or: Assigning Roles to Strangers
abstract
The Internet enables connectivity between many strangers: entities that don't know each other. We present the Trust Policy Language (TPL), used to define the mapping of strangers to predefined business roles, based on certificates issued by third parties. TPL is expressive enough to allow complex policies, e.g. non-monotone (negative) certificates, while being simple enough to allow automated policy checking and processing. Issuers of certificates are either known in advance, or provide sufficient certificates to be considered a trusted authority according to the policy. This allows bottom-up, "grass roots" buildup of trust, as in the real world. We extend, rather than replace, existing role based access control mechanisms. This provides a simple, modular architecture and easy migration from existing systems. Our system automatically collects missing certificates from peer servers. In particular this allows use of standard browsers, which pass only one certificate to the server. We describe our implementation, which can be used as an extension of a Web server or as a separate server with interface to applications.
Amir Herzberg, Yosi Mass, Joris Mihaeli, Dalit Naor, Yiftach Ravid
S&P2
1999 VRCommerce - electronic commerce in virtual reality
abstract
Existing technology and standards allow the creation of threedimensional, virtual-reality browsing experience.Such an interface may be more attractive and natural (at least to some).In particular, with the increase in electronic commerce, the creation of virtual reality stores and shopping malls seems of potential value.However, there are substantial challenges in the use of existing virtual reality tools to create any large space, in particular a store or a shopping mall.The main challenges are the substantial size of the representation of the space (communication and processing overhead), difficulties of navigation using typical UI devices, and support for interconnecting separately designed spaces (stores) into one continuous virtual space (mall).We present an approach to address these challenges, based on limiting the spaces to a modular collection of basic architectural elements such as rooms and hallways.We describe our implementation of virtual-reality, three-dimensional e-commerce, and the VR Commerce toolkit, implemented using standard VRML and Java.VRCommerce is an integrated solution for creation, online operation and navigation in three dimensional malls and stores.It enables continuous navigation between separately designed and managed stores and incremental loading of spaces and objects as needed.Furthermore, VRCommerce offers simplified navigation with less decision making via a two-dimensional `mall directory map` and automated walk modes.
Yosi Mass, Amir Herzberg
EC1