Marie Jacob

dblp:76/747 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 10 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
6 papers
Query processing and optimization · 29% Data integration and cleaning · 28% Data models and query languages · 20%
Computer networks
2 papers
Internet of things and sensor networks · 100%

Topics — the 13 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
preference query
0.212014
A System for Management and Analysis of Preference Data · Proc. VLDB Endow. 2014
Data integration and cleaning › schema mapping
object-relational mapping
0.212013
Incremental mapping compilation in an object-to-relational mapping system · SIGMOD Conference 2013
Data integration and cleaning
schema mapping
0.212013
Incremental mapping compilation in an object-to-relational mapping system · SIGMOD Conference 2013
Data stream processing
continuous query processing
0.112011
Sharing work in keyword search over databases · SIGMOD Conference 2011
Query processing and optimization
multi-query optimization
0.112011
Sharing work in keyword search over databases · SIGMOD Conference 2011
Query processing and optimization
shared computation
0.112011
Sharing work in keyword search over databases · SIGMOD Conference 2011
Internet of things and sensor networks
sensor network query processing
0.112010
Dynamic Join Optimization in Multi-Hop Wireless Sensor Networks · Proc. VLDB Endow. 2010
Information retrieval
keyword query
0.112008
Learning to create data-integrating queries · Proc. VLDB Endow. 2008
Information retrieval
keyword search
0.012011
Sharing work in keyword search over databases · SIGMOD Conference 2011
Query processing and optimization › keyword query processing
keyword search over databases
0.012011
Sharing work in keyword search over databases · SIGMOD Conference 2011
Query processing and optimization › query optimization › join ordering
join optimization
0.012010
Dynamic Join Optimization in Multi-Hop Wireless Sensor Networks · Proc. VLDB Endow. 2010
Distributed and cloud data management
distributed query processing
0.012009
SmartCIS: integrating digital and physical environments · SIGMOD Conference 2009
Internet of things and sensor networks
environmental sensing
0.012009
SmartCIS: integrating digital and physical environments · SIGMOD Conference 2009

Methods — techniques the papers use, named apart from their topics

self-tuning heuristics · 0.2cost modeling · 0.2adaptive learning · 0.2roundtrip validation · 0.2NP-hardness analysis · 0.2query plan scheduling · 0.1pipelined operators · 0.1schema mapping · 0.1provenance annotation · 0.1
YearPublicationVenuePosition
2016 Optimizing Similar Item Recommendations in a Semi-structured Marketplace to Maximize Conversion
abstract
This paper tackles the problem of recommendations in eBay's large semi-structured marketplace. eBay's variable inventory and lack of structured information about listings makes traditional collaborative filtering algorithms difficult to use. We discuss how to overcome these data limitations to produce high quality recommendations in real time with a combination of a customized scalable architecture as well as a widely applicable machine learned ranking model. A pointwise ranking approach is utilized to reduce the ranking problem to a binary classification problem optimized on past user purchase behavior. We present details of a sampling strategy and feature engineering that have been critical to achieve a lift in both purchase through rate (PTR) and revenue.
Yuri M. Brovman, Marie Jacob, Natraj Srinivasan, Stephen Neola, Daniel A. Galron, Ryan Snyder
RecSys2
2015 Analyzing Crowd Rankings
abstract
Ranked data is ubiquitous in real-world applications, arising naturally when users express preferences about products and services, when voters cast ballots in elections, and when funding proposals are evaluated based on their merits or university departments based on their reputation. This paper focuses on crowdsourcing and novel analysis of ranked data. We describe the design of a data collection task in which Amazon MT workers were asked to rank movies. We present results of data analysis, correlating our ranked dataset with IMDb, where movies are rated on a discrete scale rather than ranked. We develop an intuitive measure of worker quality appropriate for this task, where no gold standard answer exists. We propose a model of local structure in ranked datasets, reflecting that subsets of the workers agree in their ranking over subsets of the items, develop a data mining algorithm that identifies such structure, and evaluate in on our dataset. Our dataset is publicly available at https://github.com/stoyanovich/CrowdRank.
Julia Stoyanovich, Marie Jacob, Xuemei Gong
WebDB2
2014 A System for Management and Analysis of Preference Data
abstract
Preference data arises in a wide variety of domains. Over the past decade, we have seen a sharp increase in the volume of preference data, in the diversity of applications that use it, and in the richness of preference data analysis methods. Examples of applications include rank aggregation in genomic data analysis, management of votes in elections, and recommendation systems in e-commerce. However, little attention has been paid to the challenges of building a system for preference-data management, which would help incorporate sophisticated analytics into larger applications, support computational abstractions for usability by data scientists, and enable scaling up to modern volumes. This vision paper proposes a management system for preference data that aims to address these challenges. We adopt the relational database model, and propose extensions that are specialized to handling preference data. Specifically, we introduce a special type of a relation that is designed for preference data, and describe composable operators on preference relations that can be embedded in SQL statements, for convenient reuse across applications.
Marie Jacob, Benny Kimelfeld, Julia Stoyanovich
Proc. VLDB Endow.1
2013 Understanding Local Structure in Ranked Datasets
Julia Stoyanovich, Sihem Amer-Yahia, Susan B. Davidson, Marie Jacob, Tova Milo
CIDR4
2013 Incremental mapping compilation in an object-to-relational mapping system
abstract
In an object-to-relational mapping system (ORM), mapping expressions explain how to expose relational data as objects and how to store objects in tables. If mappings are sufficiently expressive, then it is possible to define lossy mappings. If a user updates an object, stores it in the database based on a lossy mapping, and then retrieves the object from the database, the user might get a different result than the updated state of the object; that is, the mapping might not "roundtrip." To avoid this, the ORM should validate that user-defined mappings roundtrip the data. However, this problem is NP-hard, so mapping validation can be very slow for large or complex mappings.
Philip A. Bernstein, Marie Jacob, Jorge Pérez 0001, Guillem Rull, James F. Terwilliger
SIGMOD Conference2
2011 Sharing work in keyword search over databases
abstract
An important means of allowing non-expert end-users to pose ad hoc queries whether over single databases or data integration systems is through keyword search. Given a set of keywords, the query processor finds matches across different tuples and tables. It computes and executes a set of relational sub-queries whose results are combined to produce the k highest ranking answers. Work on keyword search primarily focuses on single-database, single-query settings: each query is answered in isolation, despite possible overlap between queries posed by different users or at different times; and the number of relevant tables is assumed to be small, meaning that sub-queries can be processed without using cost-based methods to combine work. As we apply keyword search to support ad hoc data integration queries over scientific or other databases on the Web, we must reuse and combine computation. In this paper, we propose an architecture that continuously receives sets of ranked keyword queries, and seeks to reuse work across these queries. We extend multiple query optimization and continuous query techniques, and develop a new query plan scheduling module we call the ATC (based on its analogy to an air traffic controller). The ATC manages the flow of tuples among a multitude of pipelined operators, minimizing the work needed to return the top-k answers for all queries. We also develop techniques to manage the sharing and reuse of state as queries complete and input data streams are exhausted. We show the effectiveness of our techniques in handling queries over real and synthetic data sets.
Marie Jacob, Zachary G. Ives
SIGMOD Conference1
2010 Dynamic Join Optimization in Multi-Hop Wireless Sensor Networks
abstract
To enable smart environments and self-tuning data centers, we are developing the Aspen system for integrating physical sensor data, as well as stream data coming from machine logical state, and database or Web data from the Internet. A key component of this system is a query processor optimized for limited-bandwidth, possibly battery-powered devices with multiple hop wireless radio communications. This query processor is given a portion of a data integration query, possibly including joins among sensors, to execute. Several recent papers have developed techniques for computing joins in sensors, but these techniques are static and are only appropriate for specific join selectivity ratios. We consider the problem of dynamic join optimization for sensor networks, developing solutions that employ cost modeling, as well as adaptive learning and self-tuning heuristics to choose the best algorithm under real and variable selectivity values. We focus on in-network join computation, but our architecture extends to other approaches (and we compare against these). We develop basic techniques assuming selectivities are uniform and known in advance, and optimization can be done on a pairwise basis; we then extend the work to handle joins between multiple pairs, when selectivities are not fully known. We experimentally validate our work at scale using standard datasets.
Svilen R. Mihaylov, Marie Jacob, Zachary G. Ives, Sudipto Guha
Proc. VLDB Endow.2
2009 Interactive Data Integration through Smart Copy & Paste
Zachary G. Ives, Craig A. Knoblock, Steven Minton, Marie Jacob, Partha P. Talukdar, Rattapoom Tuchinda, José Luis Ambite, Maria Muslea, Cenk Gazen
CIDR4
2009 SmartCIS: integrating digital and physical environments
abstract
demonstration SmartCIS: integrating digital and physical environments Share on Authors: Mengmeng Liu University of Pennsylvania, Philadelphia, PA, USA University of Pennsylvania, Philadelphia, PA, USAView Profile , Svilen R. Mihaylov University of Pennsylvania, Philadelphia, PA, USA University of Pennsylvania, Philadelphia, PA, USAView Profile , Zhuowei Bao University of Pennsylvania, Philadelphia, PA, USA University of Pennsylvania, Philadelphia, PA, USAView Profile , Marie Jacob University of Pennsylvania, Philadelphia, PA, USA University of Pennsylvania, Philadelphia, PA, USAView Profile , Zachary G. Ives University of Pennsylvania, Philadelphia, PA, USA University of Pennsylvania, Philadelphia, PA, USAView Profile , Boon Thau Loo University of Pennsylvania, Philadelphia, PA, USA University of Pennsylvania, Philadelphia, PA, USAView Profile , Sudipto Guha University of Pennsylvania, Philadelphia, PA, USA University of Pennsylvania, Philadelphia, PA, USAView Profile Authors Info & Claims SIGMOD '09: Proceedings of the 2009 ACM SIGMOD International Conference on Management of dataJune 2009 Pages 1111–1114https://doi.org/10.1145/1559845.1559996Online:29 June 2009Publication History 6citation263DownloadsMetricsTotal Citations6Total Downloads263Last 12 Months2Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Svilen R. Mihaylov, Zhuowei Bao, Marie Jacob, Zachary G. Ives, Boon Thau Loo, Sudipto Guha
SIGMOD Conference4
2008 Learning to create data-integrating queries
abstract
The number of potentially-related data resources available for querying --- databases, data warehouses, virtual integrated schemas --- continues to grow rapidly. Perhaps no area has seen this problem as acutely as the life sciences, where hundreds of large, complex, interlinked data resources are available on fields like proteomics, genomics, disease studies, and pharmacology. The schemas of individual databases are often large on their own, but users also need to pose queries across multiple sources, exploiting foreign keys and schema mappings. Since the users are not experts, they typically rely on the existence of pre-defined Web forms and associated query templates, developed by programmers to meet the particular scientists' needs. Unfortunately, such forms are scarce commodities, often limited to a single database, and mismatched with biologists' information needs that are often context-sensitive and span multiple databases. We present a system with which a non-expert user can author new query templates and Web forms, to be reused by anyone with related information needs. The user poses keyword queries that are matched against source relations and their attributes; the system uses sequences of associations (e.g., foreign keys, links, schema mappings, synonyms, and taxonomies) to create multiple ranked queries linking the matches to keywords; the set of queries is attached to a Web query form. Now the user and his or her associates may pose specific queries by filling in parameters in the form. Importantly, the answers to this query are ranked and annotated with data provenance, and the user provides feedback on the utility of the answers, from which the system ultimately learns to assign costs to sources and associations according to the user's specific information need, as a result changing the ranking of the queries used to generate results. We evaluate the effectiveness of our method against "gold standard" costs from domain experts and demonstrate the method's scalability.
Partha P. Talukdar, Marie Jacob, Muhammad Salman Mehmood, Koby Crammer, Zachary G. Ives, Fernando Pereira 0003, Sudipto Guha
Proc. VLDB Endow.2