Rimma V. Nehme

dblp:94/81 · DBLP profile ↗
← Back
24ranked-venue papers
12as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 22 · 10 first-authorSecurity and privacy · 1 · 1 first-authorTheory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
16 papers
Query processing and optimization · 42% Distributed and cloud data management · 19% Database system architecture and tuning · 18%
Computer architecture, parallel and distributed computing, and storage systems
6 papers
Performance modeling and evaluation · 48% Distributed systems · 23% Parallel and multicore computing · 16%
Network and information security
3 papers
Authentication and access control · 84% Privacy and data protection · 16%

Topics — the 28 heaviest of 35, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed and cloud data management
data partitioning
0.322014
Resource Bricolage for Parallel Database Systems · Proc. VLDB Endow. 2014
Automated partitioning design in parallel database systems · SIGMOD Conference 2011
Data stream processing
continuous query processing
0.342010
FENCE: Continuous access control enforcement in dynamic data stream environments · ICDE 2010
Tagging Stream Data for Rich Real-Time Services · Proc. VLDB Endow. 2009
A Security Punctuation Framework for Enforcing Access Control on Streaming Data · ICDE 2008
Query processing and optimization › query execution
query progress estimation
0.212016
Operator and Query Progress Estimation in Microsoft SQL Server Live Query Statistics · SIGMOD Conference 2016
Query processing and optimization
query optimization
0.222011
Automated partitioning design in parallel database systems · SIGMOD Conference 2011
Configuration-parametric query optimization for physical design tuning · SIGMOD Conference 2008
Query processing and optimization
parallel query processing
0.212014
Resource Bricolage for Parallel Database Systems · Proc. VLDB Endow. 2014
Distributed systems
fault tolerance
0.212014
Partial results in database systems · SIGMOD Conference 2014
Database system architecture and tuning
parallel database system
0.222012
Automated partitioning design in parallel database systems · SIGMOD Conference 2011
Query optimization in microsoft SQL server PDW · SIGMOD Conference 2012
Distributed and cloud data management
distributed query processing
0.212013
Split query processing in polybase · SIGMOD Conference 2013
Distributed and cloud data management
federated database
0.212013
Split query processing in polybase · SIGMOD Conference 2013
Query processing and optimization › query optimization
distributed query optimization
0.112012
Query optimization in microsoft SQL server PDW · SIGMOD Conference 2012
Query processing and optimization › query execution
query progress indicator
0.112012
GSLPI: A Cost-Based Query Progress Indicator · ICDE 2012
Performance modeling and evaluation › performance prediction
query performance prediction
0.112012
GSLPI: A Cost-Based Query Progress Indicator · ICDE 2012
Debugging and program repair
fault localization
0.112010
Mini-Me: A min-repro system for database software · ICDE 2010
Query processing and optimization
adaptive query processing
0.112009
Query Mesh: Multi-Route Query Processing Technology · Proc. VLDB Endow. 2009
Data stream processing
stream processing systems
0.112009
StreamShield: a stream-centric approach towards security and privacy in data stream environments · SIGMOD Conference 2009
Parallel and multicore computing › parallel computing
parallel database systems
0.112017
Resource bricolage and resource selection for parallel database systems · VLDB J. 2017
Information retrieval › indexing
index compression
0.112008
Transaction time indexing with version compression · Proc. VLDB Endow. 2008
Query processing and optimization › adaptive query processing › adaptive query optimization
parametric query optimization
0.112008
Configuration-parametric query optimization for physical design tuning · SIGMOD Conference 2008
Database system architecture and tuning › database design
physical database design
0.112008
Configuration-parametric query optimization for physical design tuning · SIGMOD Conference 2008
Spatial and temporal data management
temporal databases
0.112008
Transaction time indexing with version compression · Proc. VLDB Endow. 2008
Authentication and access control
access control policy
0.122010
FENCE: Continuous access control enforcement in dynamic data stream environments · ICDE 2010
A Security Punctuation Framework for Enforcing Access Control on Streaming Data · ICDE 2008
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.112014
Resource Bricolage for Parallel Database Systems · Proc. VLDB Endow. 2014
GPUs and heterogeneous computing
heterogeneous resources
0.112014
Resource Bricolage for Parallel Database Systems · Proc. VLDB Endow. 2014
Query processing and optimization › query optimization
cost-based optimization
0.012013
Split query processing in polybase · SIGMOD Conference 2013
Parallel and multicore computing › parallel architecture
massively parallel processing
0.012012
Query optimization in microsoft SQL server PDW · SIGMOD Conference 2012
Database system architecture and tuning
massively parallel processing
0.012011
Automated partitioning design in parallel database systems · SIGMOD Conference 2011
Database system architecture and tuning
database engine testing
0.012010
Mini-Me: A min-repro system for database software · ICDE 2010
Data models and query languages
metadata
0.012009
Tagging Stream Data for Rich Real-Time Services · Proc. VLDB Endow. 2009

Methods — techniques the papers use, named apart from their topics

resource selection · 0.6progress estimation · 0.5security punctuations · 0.4result classification · 0.4linear programming · 0.4analytical framework · 0.4cost model · 0.3data mining · 0.2mapreduce · 0.2cost-based optimization · 0.2simplification transformation · 0.1record and replay · 0.1query algebra · 0.1
YearPublicationVenuePosition
2017 Resource bricolage and resource selection for parallel database systems
Jiexing Li, Jeffrey F. Naughton, Rimma V. Nehme
VLDB J.3
2016 Operator and Query Progress Estimation in Microsoft SQL Server Live Query Statistics
abstract
We describe the design and implementation of the new Live Query Statistics (LQS) feature in Microsoft SQL Server 2016. The functionality includes the display of overall query progress as well as progress of individual operators in the query execution plan. We describe the overall functionality of LQS, give usage examples and detail all areas where we had to extend the current state-of-the-art to build the complete LQS feature. Finally, we evaluate the effect these extensions have on progress estimation accuracy with a series of experiments using a large set of synthetic and real workloads.
Kukjin Lee, Arnd Christian König, Vivek R. Narasayya, Bolin Ding, Surajit Chaudhuri, Brent Ellwein, Alexey Eksarevskiy, Manbeen Kohli, Jacob Wyant, Praneeta Prakash, Rimma V. Nehme, Jiexing Li, Jeffrey F. Naughton
SIGMOD Conference11
2015 Database Optimization in the Cloud: Where Costs, Partial Results, and Consumer Choice Meet
Willis Lang, Rimma V. Nehme, Ian Rae
CIDR2
2014 Partial results in database systems
abstract
As the size and complexity of analytic data processing systems continue to grow, the effort required to mitigate faults and performance skew has also risen. However, in some environments we have encountered, users prefer to continue query execution even in the presence of failures (e.g., the unavailability of certain data sources), and receive a "partial" answer to their query. We explore ways to characterize and classify these partial results, and describe an analytical framework that allows the system to perform coarse to fine-grained analysis to determine the semantics of a partial result. We propose that if the system is equipped with such a framework, in some cases it is better to return and explain partial results than to attempt to avoid them.
Willis Lang, Rimma V. Nehme, Eric Robinson, Jeffrey F. Naughton
SIGMOD Conference2
2014 Resource Bricolage for Parallel Database Systems
abstract
Running parallel database systems in an environment with heterogeneous resources has become increasingly common, due to cluster evolution and increasing interest in moving applications into public clouds. For database systems running in a heterogeneous cluster, the default uniform data partitioning strategy may overload some of the slow machines while at the same time it may under-utilize the more powerful machines. Since the processing time of a parallel query is determined by the slowest machine, such an allocation strategy may result in a significant query performance degradation. We take a first step to address this problem by introducing a technique we call resource bricolage that improves database performance in heterogeneous environments. Our approach quantifies the performance differences among machines with various resources as they process workloads with diverse resource requirements. We formalize the problem of minimizing workload execution time and view it as an optimization problem, and then we employ linear programming to obtain a recommended data partitioning scheme. We verify the effectiveness of our technique with an extensive experimental study on a commercial database system.
Jiexing Li, Jeffrey F. Naughton, Rimma V. Nehme
Proc. VLDB Endow.3
2013 Toward Progress Indicators on Steroids for Big Data Systems
Jiexing Li, Rimma V. Nehme, Jeffrey F. Naughton
CIDR2
2013 Unlocking Cool in Databases (and CS in general)
Rimma V. Nehme
CIDR1
2013 FENCE: continuous access control enforcement in dynamic data stream environments
abstract
In this paper, we address the problem of continuous access control enforcement in dynamic data stream environments, where both data and query security restrictions may potentially change in real-time. We present FENCE framework that ffectively addresses this problem. The distinguishing characteristics of FENCE include: (1) the stream-centric approach to security, (2) the symmetric model for security settings of both continuous queries and streaming data, and (3) two alternative security-aware query processing approaches that can optimize query execution based on regular and security-related selectivities. In FENCE, both data and query security restrictions are modeled symmetrically in the form of security metadata, called "security punctuations" embedded inside data streams. We distinguish between two types of security punctuations, namely, the data security punctuations (or short, dsps) which represent the access control policies of the streaming data, and the query security punctuations (or short, qsps) which describe the access authorizations of the continuous queries. We also present our encoding method to support XACML(eXtensible Access Control Markup Language) standard. We have implemented FENCE in a prototype DSMS and present our performance evaluation. The results of our experimental study show that FENCE's approach has low overhead and can give great performance benefits compared to the alternative security solutions for streaming environments.
Rimma V. Nehme, Hyo-Sang Lim, Elisa Bertino
CODASPY1
2013 Split query processing in polybase
abstract
This paper presents Polybase, a feature of SQL Server PDW V2 that allows users to manage and query data stored in a Hadoop cluster using the standard SQL query language. Unlike other database systems that provide only a relational view over HDFS-resident data through the use of an external table mechanism, Polybase employs a split query processing paradigm in which SQL operators on HDFS-resident data are translated into MapReduce jobs by the PDW query optimizer and then executed on the Hadoop cluster. The paper describes the design and implementation of Polybase along with a thorough performance evaluation that explores the benefits of employing a split query processing paradigm for executing queries that involve both structured data in a relational DBMS and unstructured data in Hadoop. Our results demonstrate that while the use of a split-based query execution paradigm can improve the performance of some queries by as much as 10X, one must employ a cost-based query optimizer that considers a broad set of factors when deciding whether or not it is advantageous to push a SQL operator to Hadoop. These factors include the selectivity factor of the predicate, the relative sizes of the two clusters, and whether or not their nodes are co-located. In addition, differences in the semantics of the Java and SQL languages must be carefully considered in order to avoid altering the expected results of a query.
David J. DeWitt, Alan Halverson, Rimma V. Nehme, Srinath Shankar, Josep Aguilar-Saborit, Artin Avanes, Miro Flasza, Jim Gramling
SIGMOD Conference3
2013 Multi-route query processing and optimization
Rimma V. Nehme, Karen E. Works, Chuan Lei, Elke A. Rundensteiner, Elisa Bertino
J. Comput. Syst. Sci.1
2012 GSLPI: A Cost-Based Query Progress Indicator
abstract
Progress indicators for SQL queries were first published in 2004 with the simultaneous and independent proposals from Chaudhuri et al. and Luo et al. In this paper, we implement both progress indicators in the same commercial RDBMS to investigate their performance. We summarize common cases in which they are both accurate and cases in which they fail to provide reliable estimates. Although there are differences in their performance, much more striking is the similarity in the errors they make due to a common simplifying uniform future speed assumption. While the developers of these progress indicators were aware that this assumption could cause errors, they neither explored how large the errors might be nor did they investigate the feasibility of removing the assumption. To rectify this we propose a new query progress indicator, similar to these early progress indicators but without the uniform speed assumption. Experiments show that on the TPC-H benchmark, on queries for which the original progress indicators have errors up to 30X the query running time, the new progress indicator is accurate to within 10 percent. We also discuss the sources of the errors that still remain and shed some light on what would need to be done to eliminate them.
Jiexing Li, Rimma V. Nehme, Jeffrey F. Naughton
ICDE2
2012 Query optimization in microsoft SQL server PDW
abstract
In recent years, Massively Parallel Processors have increasingly been used to manage and query vast amounts of data. Dramatic performance improvements are achieved through distributed execution of queries across many nodes. Query optimization for such system is a challenging and important problem.
Srinath Shankar, Rimma V. Nehme, Josep Aguilar-Saborit, Mostafa Elhemali, Alan Halverson, Eric Robinson, Mahadevan Sankara Subramanian, David J. DeWitt, César A. Galindo-Legaria
SIGMOD Conference2
2011 Automated partitioning design in parallel database systems
abstract
In recent years, Massively Parallel Processors (MPPs) have gained ground enabling vast amounts of data processing. In such environments, data is partitioned across multiple compute nodes, which results in dramatic performance improvements during parallel query execution. To evaluate certain relational operators in a query correctly, data sometimes needs to be re-partitioned (i.e., moved) across compute nodes. Since data movement operations are much more expensive than relational operations, it is crucial to design a suitable data partitioning strategy that minimizes the cost of such expensive data transfers. A good partitioning strategy strongly depends on how the parallel system would be used. In this paper we present a partitioning advisor that recommends the best partitioning design for an expected workload. Our tool recommends which tables should be replicated (i.e., copied into every compute node) and which ones should be distributed according to specific column(s) so that the cost of evaluating similar workloads is minimized. In contrast to previous work, our techniques are deeply integrated with the underlying parallel query optimizer, which results in more accurate recommendations in a shorter amount of time. Our experimental evaluation using a real MPP system, Microsoft SQL Server 2008 Parallel Data Warehouse, with both real and synthetic workloads shows the effectiveness of the proposed techniques and the importance of deep integration of the partitioning advisor with the underlying query optimizer.
Rimma V. Nehme, Nicolas Bruno
SIGMOD Conference1
2010 Mini-Me: A min-repro system for database software
abstract
Testing and debugging database software is often challenging and time consuming. A very arduous task for DB testers is finding a min-repro - the ¿simplest possible setup¿ that reproduces the original problem. Currently, a great deal of searching for min-repros is carried out manually using non-database-specific tools, which is both slow and error-prone. We propose to demonstrate a system, called Mini-Me1, designed to ease and speed-up the task of finding min-repros in database-related products. Mini-Me employs several effective tools, including: the novel simplification transformations, the high-level language for creating search scripts and automation, the ¿record-and-replay¿ functionality, and the visualization of the search space and results. In addition to the standard application mode, the system can be interacted with in the game mode. The latter can provide an intrinsically motivating environment for developing successful search strategies by DB testers, which can be data-mined and recorded as patterns and used as recommendations for DB testers in the future. Potentially, a system like Mini-Me can save hours of time (for both customers and testers to isolate a problem), which could result in faster fixes and large cost savings to organizations.
Nicolas Bruno, Rimma V. Nehme
ICDE2
2010 FENCE: Continuous access control enforcement in dynamic data stream environments
abstract
In this paper, we present FENCE framework that addresses the problem of continuous access control enforcement in dynamic data stream environments. The distinguishing characteristics of FENCE include: (1) the stream-centric approach to security, (2) the symmetric modeling of security for both continuous queries and streaming data, and (3) security-aware query processing that considers both regular and security-related selectivities. In FENCE, both data and query security restrictions are modeled in the form of streaming security metadata, called ¿security punctuations¿, embedded inside data streams. We have implemented FENCE in a prototype DSMS and briefly summarize our performance observations.
Rimma V. Nehme, Hyo-Sang Lim, Elisa Bertino
ICDE1
2009 Self-tuning query mesh for adaptive multi-route query processing
abstract
In real-life applications, different subsets of data may have distinct statistical properties, e.g., various websites may have diverse visitation rates, different categories of stocks may have dissimilar price fluctuation patterns. For such applications, it can be fruitful to eliminate the commonly made single execution plan assumption and instead execute a query using several plans, each optimally serving a subset of data with particular statistical properties. Furthermore, in dynamic environments, data properties may change continuously, thus calling for adaptivity. The intriguing question is: can we have an execution strategy that (1) is plan-based to leverage on all the benefits of traditional plan-based systems, (2) supports multiple plans each customized for different subset of data, and yet (3) is as adaptive as "plan-less" systems like Eddies? While the recently proposed Query Mesh (QM) approach provides a foundation for such an execution paradigm, it does not address the question of adaptivity required for highly dynamic environments. In this work, we fill this gap by proposing a Self-Tuning Query Mesh (ST-QM) --- an adaptive solution for content-based multi-plan execution engines. ST-QM addresses adaptive query processing by abstracting it as a concept drift problem --- a well-known subject in machine learning. Such abstraction allows to discard adaptivity candidates (i.e., the cases indicating a change in the environment) early in the process if they are insignificant or not "worthwhile" to adapt to, and thus minimize the adaptivity overhead. A unique feature of our aproach is that all logical transformations to the execution strategy get translated into a single inexpensive physical operation --- the classifier change. Our experimental evaluation using a continuous query engine shows the performance benefits of ST-QM approach over the alternatives, namely the non-adaptive and the Eddies-based solutions.
Rimma V. Nehme, Elke A. Rundensteiner, Elisa Bertino
EDBT1
2009 StreamShield: a stream-centric approach towards security and privacy in data stream environments
abstract
We propose to demonstrate the StreamShield, a system designed to address the problem of security and privacy in the context of Data Stream Management Systems (DSMSs). In StreamShield, continuous access control is enforced by taking a novel "stream-centric" approach towards security. Security policies are not persistently stored on the server, but rather are depicted by security metadata, called "security punctuations", and get embedded into streams together with the data. We distinguish between two types of security punctuations: (1) the "data security punctuations" (dsps) describing the data-side security policies, and (2) the "query security punctuations" (qsps) representing the query-side security policies. The advantages of such stream-centric security model include flexibility, dynamicity and speed of enforcement. Furthermore, DSMSs can adapt to not only data-related but also to security-related selectivities, which helps reduce the waste of resources, when few subjects have access to streaming data.
Rimma V. Nehme, Hyo-Sang Lim, Elisa Bertino, Elke A. Rundensteiner
SIGMOD Conference1
2009 Tagging Stream Data for Rich Real-Time Services
abstract
In recent years, data streams have become ubiquitous as technology is improving and the prices of portable devices are falling, e.g., sensor networks, location-based services. Most data streams transmit only data tuples based on which continuous queries are evaluated. In this paper, we propose to enrich data streams with a new type of metadata called streaming tags or short tick-tags . The fundamental premise of tagging is that users can label data using uncontrolled vocabulary, and these tags can be exploited in a wide variety of applications, such as data exploration, data search, and to produce "enriched" with additional semantics, thus more informative query results. In this paper we focus primarily on the problem of continuous query processing with streaming tags and tagged objects, and address the tick-tag semantic issues as well as efficiency concerns. Our main contributions are as follows. First, we specify a general and flexible Stream Tag Framework (or short STF) that supports a stream-centric approach to tagging, and where tick-tags , attached to streaming objects are treated as first-class citizens. Second, under STF, users can query tags explicitly as well as implicitly by outputting the tags of the base data together with query results. Finally, we have implemented STF in a prototype Data Stream Management System, and through a set of performance experiments, we show that the cost of stream tagging is small and the approach is scalable to a large percentage of tagged objects.
Rimma V. Nehme, Elke A. Rundensteiner, Elisa Bertino
Proc. VLDB Endow.1
2009 Query Mesh: Multi-Route Query Processing Technology
abstract
We propose to demonstrate a practical alternative approach to the current state-of-the-art query processing techniques, called the " Query Mesh " (or QM , for short). The main idea of QM is to compute multiple routes (i.e., query plans), each designed for a particular subset of data with distinct statistical properties. Based on the execution routes and the data characteristics, a classifier model is induced and is used to partition new data tuples to assign the best routes for their processing. We propose to demonstrate the QM framework in the streaming context using our demo application, called the " Ubi-City ". We will illustrate the innovative features of QM , including: the QM optimization with the integrated machine learning component, the QM execution using the efficient " Self-Routing Fabric " infrastructure, and finally, the QM adaptive component that performs the online adaptation of QM with near-zero runtime overhead.
Rimma V. Nehme, Karen E. Works, Elke A. Rundensteiner, Elisa Bertino
Proc. VLDB Endow.1
2008 A Security Punctuation Framework for Enforcing Access Control on Streaming Data
abstract
The management of privacy and security in the context of data stream management systems (DSMS) remains largely an unaddressed problem to date. Unlike in traditional DBMSs where access control policies are persistently stored on the server and tend to remain stable, in streaming applications the contexts and with them the access control policies on the real-time data may rapidly change. A person entering a casino may want to immediately block others from knowing his current whereabouts. We thus propose a novel ";stream-centric"; approach, where security restrictions are not persistently stored on the DSMS server, but rather streamed together with the data. Here, the access control policies are expressed via security constraints (called security punctuations, or short, sps) and are embedded into data streams. The advantages of the sp model include flexibility, dynamicity and speed of enforcement. DSMSs can adapt to not only data-related but also security-related selectivities, which helps reduce the waste of resources, when few subjects have access to data. We propose a security-aware query algebra and new equivalence rules together with cost estimations to guide the security-aware query plan optimization. We have implemented the sp framework in a real DSMS. Our experimental results show the validity and the performance advantages of our sp model as compared to alternative access control enforcement solutions for DSMSs.
Rimma V. Nehme, Elke A. Rundensteiner, Elisa Bertino
ICDE1
2008 Configuration-parametric query optimization for physical design tuning
abstract
Automated physical design tuning for database systems has recently become an active area of research and development. Existing tuning tools explore the space of feasible solutions by repeatedly optimizing queries in the input workload for several candidate configurations. This general approach, while scalable, often results in tuning sessions waiting for results from the query optimizer over 90% of the time. In this paper we introduce a novel approach, called Configuration-Parametric Query Optimization, that drastically improves the performance of current tuning tools. By issuing a single optimization call per query, we are able to generate a compact representation of the optimization space that can then produce very efficiently execution plans for the input query under arbitrary configurations. Our experiments show that our technique speeds-up query optimization by 30x to over 450x with virtually no loss in quality, and effectively eliminates the optimization bottleneck in existing tuning tools. Our techniques open the door for new, more sophisticated optimization strategies by eliminating the main bottleneck of current tuning tools.
Nicolas Bruno, Rimma V. Nehme
SIGMOD Conference2
2008 Transaction time indexing with version compression
abstract
Immortal DB is a transaction time database system designed to enable high performance for temporal applications. It is built into a commercial database engine, Microsoft SQL Server. This paper describes how we integrated a temporal indexing technique, the TSB-tree, into Immortal DB to serve as the core access method. The TSB-tree provides high performance access and update for both current and historical data. A main challenge was integrating TSB-tree functionality while preserving original B+tree functionality, including concurrency control and recovery. We discuss the overall architecture, including our unique treatment of index terms, and practical issues such as uncommitted data and log management. Performance is a primary concern. To increase performance, versions are locally delta compressed, exploiting the commonality between adjacent versions of the same record. This technique is also applied to index terms in index pages. There is a tradeoff between query performance and storage space. We discuss optimizing performance regarding this tradeoff throughout the paper. The result of our efforts is a high-performance transaction time database system built into an RDBMS engine, which has not been achieved before. We include a thorough experimental study and analysis that confirms the very good performance that it achieves.
David B. Lomet, Mingsheng Hong, Rimma V. Nehme, Rui Zhang 0003
Proc. VLDB Endow.3
2007 ClusterSheddy : Load Shedding Using Moving Clusters over Spatio-temporal Data Streams
Rimma V. Nehme, Elke A. Rundensteiner
DASFAA1
2006 SCUBA: Scalable Cluster-Based Algorithm for Evaluating Continuous Spatio-temporal Queries on Moving Objects
Rimma V. Nehme, Elke A. Rundensteiner
EDBT1