Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Nikolay Laptev

dblp:96/7449 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
2since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 9 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Systems, architecture and hardware · 1Computer networks · 1Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
10 papers
Query processing and optimization · 39% Data mining · 26% Data stream processing · 11%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Cloud and datacenter computing · 61% Performance modeling and evaluation · 18% Electronic design automation · 15%
Artificial intelligence
1 paper
Probabilistic and Bayesian machine learning · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational finance and economics · 100%

Topics — the 27 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization › query optimization
learned query optimization
0.712023
AutoSteer: Learned Query Optimization for Any SQL Database · Proc. VLDB Endow. 2023
Cloud and datacenter computing
resource allocation
0.412019
Probabilistic Modeling of Computing Demand for Service Level Agreement · IEEE Trans. Serv. Comput. 2019
Query processing and optimization
approximate query processing
0.322013
Very fast estimation for result and accuracy of big data analytics: The EARL system · ICDE 2013
Early Accurate Results for Advanced Analytics on MapReduce · Proc. VLDB Endow. 2012
Web and social media mining › user behavior analysis
user behavior modeling
0.212016
DECT: Distributed Evolving Context Tree for Understanding User Behavior Pattern Evolution · AAAI 2016
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
hidden markov model
0.212015
Inertial Hidden Markov Models: Modeling Change in Multivariate Time Series · AAAI 2015
Data mining
anomaly detection
0.212015
Generic and Scalable Framework for Automated Time-series Anomaly Detection · KDD 2015
Data mining
time series analysis
0.212015
Inertial Hidden Markov Models: Modeling Change in Multivariate Time Series · AAAI 2015
Data mining › anomaly detection
time series anomaly detection
0.212015
Generic and Scalable Framework for Automated Time-series Anomaly Detection · KDD 2015
Data mining › time series analysis
time series segmentation
0.212015
Inertial Hidden Markov Models: Modeling Change in Multivariate Time Series · AAAI 2015
Computational finance and economics › financial market prediction
stock prediction
0.212014
Stock trade volume prediction with Yahoo Finance user browsing behavior · ICDE 2014
Web and social media mining › web usage mining
clickstream analysis
0.212014
Stock trade volume prediction with Yahoo Finance user browsing behavior · ICDE 2014
Data stream processing
complex event processing
0.112012
Optimization of Massive Pattern Queries by Dynamic Configuration Morphing · ICDE 2012
Data stream processing › continuous query processing
early result production
0.112012
Early Accurate Results for Advanced Analytics on MapReduce · Proc. VLDB Endow. 2012
Information retrieval
pattern matching
0.112012
Optimization of Massive Pattern Queries by Dynamic Configuration Morphing · ICDE 2012
Data mining
data stream mining
0.112011
SMM: A data stream management system for knowledge discovery · ICDE 2011
Data stream processing
stream processing systems
0.112011
SMM: A data stream management system for knowledge discovery · ICDE 2011
Performance modeling and evaluation › statistical analysis
extreme value theory
0.112019
Probabilistic Modeling of Computing Demand for Service Level Agreement · IEEE Trans. Serv. Comput. 2019
Performance modeling and evaluation
workload characterization
0.112019
Probabilistic Modeling of Computing Demand for Service Level Agreement · IEEE Trans. Serv. Comput. 2019
Electronic design automation › logic synthesis › logic optimization
common subexpression elimination
0.112009
Xquasher: a tool for efficient computation of multiple linear expressions · DAC 2009
Electronic design automation
high-level synthesis
0.112009
Xquasher: a tool for efficient computation of multiple linear expressions · DAC 2009
Data mining › predictive modeling
forecasting
0.112015
Generic and Scalable Framework for Automated Time-series Anomaly Detection · KDD 2015
Recommender systems › user modeling
user intent modeling
0.112014
Stock trade volume prediction with Yahoo Finance user browsing behavior · ICDE 2014
Data mining
big data analytics
0.012013
Very fast estimation for result and accuracy of big data analytics: The EARL system · ICDE 2013
Query processing and optimization › query optimization › graph query optimization
pattern query optimization
0.012012
Optimization of Massive Pattern Queries by Dynamic Configuration Morphing · ICDE 2012
Query processing and optimization
query optimization
0.012012
Optimization of Massive Pattern Queries by Dynamic Configuration Morphing · ICDE 2012
Parallel and multicore computing › data-parallel programming
mapreduce
0.012012
Early Accurate Results for Advanced Analytics on MapReduce · Proc. VLDB Endow. 2012
Integrated circuit design
digital signal processing circuits
0.012009
Xquasher: a tool for efficient computation of multiple linear expressions · DAC 2009

Methods — techniques the papers use, named apart from their topics

visualization · 0.7reinforcement learning · 0.7bandit optimization · 0.7bootstrapping · 0.5state transition regularization · 0.4expectation-maximization · 0.4probabilistic modeling · 0.4extreme value theory · 0.4web traffic analysis · 0.4correlation analysis · 0.4variable-order markov model · 0.2learning curve prediction · 0.2non-parametric estimation · 0.1power set encoding · 0.1
YearPublicationVenuePosition
2023 QO-Insight: Inspecting Steered Query Optimizers
abstract
Steered query optimizers address the planning mistakes of traditional query optimizers by providing them with hints on a per-query basis, thereby guiding them in the right direction. This paper introduces QO-Insight, a visual tool designed for exploring query execution traces of such steered query optimizers. Although steered query optimizers are typically perceived as black boxes, QO-Insight empowers database administrators and experts to gain qualitative insights and enhance their performance through visual inspection and analysis.
Christoph Anneser, Mario Petruccelli, Nesime Tatbul, David E. Cohen, Zhenggang Xu, Prithviraj Pandian, Nikolay Laptev, Ryan Marcus, Alfons Kemper
Proc. VLDB Endow.7
2023 AutoSteer: Learned Query Optimization for Any SQL Database
abstract
This paper presents AutoSteer, a learning-based solution that automatically drives query optimization in any SQL database that exposes tunable optimizer knobs. AutoSteer builds on the Bandit optimizer (Bao) and extends it with new capabilities (e.g., automated hint-set discovery) to minimize integration effort and facilitate usability in both monolithic and disaggregated SQL systems. We successfully applied AutoSteer on PostgreSQL, PrestoDB, Spark-SQL, MySQL, and DuckDB - five popular open-source database engines with diverse query optimizers. We then conducted a detailed experimental evaluation with public benchmarks (JOB, Stackoverflow, TPC-DS) and a production workload from Meta's PrestoDB deployments. Our evaluation shows that AutoSteer can not only outperform these engines' native query optimizers (e.g., up to 40% improvements for PrestoDB) but can also match the performance of Bao-for-PostgreSQL with reduced human supervision and increased adaptivity, as it replaces Bao's static, expert-picked hint-sets with those that are automatically discovered. We also provide an open-source implementation of AutoSteer together with a visual tool for interactive use by query optimization experts.
Christoph Anneser, Nesime Tatbul, David E. Cohen, Zhenggang Xu, Prithviraj Pandian, Nikolay Laptev, Ryan Marcus
Proc. VLDB Endow.6
2019 Probabilistic Modeling of Computing Demand for Service Level Agreement
abstract
Cloud computing applications must be allocated sufficient resources to comply with Service Level Agreements (SLAs). This paper considers data-driven probabilistic modeling of application resource demand for resource allocation. The modeling method is focused on peak demand and SLA violations and relies on a branch of statistics known as extreme value theory (EVT). Rigorous statistical validation of the proposed model shows that it generalizes better than the alternative. The paper presents a resource allocation algorithm using the model to ensure a given small SLA violation rate. For resource allocation in Yahoo data center, over 50 percent savings are demonstrated using the proposed approach.
Saahil Shenoy, Dimitry M. Gorinevsky, Nikolay Laptev
IEEE Trans. Serv. Comput.3
2016 DECT: Distributed Evolving Context Tree for Understanding User Behavior Pattern Evolution
abstract
Internet user behavior models characterize user browsing dynamics or the transitions among web pages. The models help Internet companies improve their services by accurately targeting customers and providing them the information they want. For instance, specific web pages can be customized and prefetched for individuals based on sequences of web pages they have visited. Existing user behavior models abstracted as time-homogeneous Markov models cannot efficiently model user behavior variation through time. This demo presents DECT, a scalable time-variant variable-order Markov model. DECT digests terabytes of user session data and yields user behavior patterns through time. We realize DECT using Apache Spark and deploy it on top of Yahoo! infrastructure. We demonstrate the benefits of DECT with anomaly detection and ad click rate prediction applications. DECT enables the detection of higher-order path anomalies and provides deep insights into ad click rates with respect to user visiting paths.
Xiaokui Shu, Nikolay Laptev, Danfeng Yao
AAAI2
2016 DECT: Distributed Evolving Context Tree for Mining Web Behavior Evolution
abstract
Internet user behavior models characterize user browsing dynamics or the transitions among web pages. The mod- els help Internet companies improve their services by accu- rately targeting customers and providing them the informa- tion they want. For instance, specic web pages can be cus- tomized and prefetched for individuals based on sequences of web pages they have visited. Existing user behavior mod- els abstracted as time-homogeneous Markov models do not provide efficient support for modeling user behavior varia- tion through time. This paper presents DECT, a scalable time-variant variable-order Markov model. DECT digests terabytes of user session data and yields user behavior pat- terns through time. We realize DECT using Apache Spark. Our implementation is being open-sourced and we deploy DECT on top of Yahoo! infrastructure. We demonstrate the benets of DECT with anomaly detection and ad click rate prediction applications. DECT enables the detection of higher-order path anomalies that are masked out by exist- ing models. DECT also provides insights into ad click rates with respect to user visiting paths.
Xiaokui Shu, Nikolay Laptev, Danfeng Yao
EDBT2
2015 Inertial Hidden Markov Models: Modeling Change in Multivariate Time Series
abstract
Faced with the problem of characterizing systematic changes in multivariate time series in an unsupervised manner, we derive and test two methods of regularizing hidden Markov models for this task. Regularization on state transitions provides smooth transitioning among states, such that the sequences are split into broad, contiguous segments. Our methods are compared with a recent hierarchical Dirichlet process hidden Markov model (HDP-HMM) and a baseline standard hidden Markov model, of which the former suffers from poor performance on moderate-dimensional data and sensitivity to parameter settings, while the latter suffers from rapid state transitioning, over-segmentation and poor performance on a segmentation task involving human activity accelerometer data from the UCI Repository. The regularized methods developed here are able to perfectly characterize change of behavior in the human activity data for roughly half of the real-data test cases, with accuracy of 94% and low variation of information. In contrast to the HDP-HMM, our methods provide simple, drop-in replacements for standard hidden Markov model update rules, allowing standard expectation maximization (EM) algorithms to be used for learning.
George D. Montañez, Saeed Amizadeh, Nikolay Laptev
AAAI3
2015 Generic and Scalable Framework for Automated Time-series Anomaly Detection
abstract
This paper introduces a generic and scalable framework for automated anomaly detection on large scale time-series data. Early detection of anomalies plays a key role in maintaining consistency of person's data and protects corporations against malicious attackers. Current state of the art anomaly detection approaches suffer from scalability, use-case restrictions, difficulty of use and a large number of false positives. Our system at Yahoo, EGADS, uses a collection of anomaly detection and forecasting models with an anomaly filtering layer for accurate and scalable anomaly detection on time-series. We compare our approach against other anomaly detection systems on real and synthetic data with varying time-series characteristics. We found that our framework allows for 50-60% improvement in precision and recall for a variety of use-cases. Both the data and the framework are being open-sourced. The open-sourcing of the data, in particular, represents the first of its kind effort to establish the standard benchmark for anomaly detection.
Nikolay Laptev, Saeed Amizadeh, Ian Flint
KDD1
2014 Stock trade volume prediction with Yahoo Finance user browsing behavior
abstract
Web traffic represents a powerful mirror for various real-world phenomena. For example, it was shown that web search volumes have a positive correlation with stock trading volumes and with the sentiment of investors. Our hypothesis is that user browsing behavior on a domain-specific portal is a better predictor of user intent than web searches.
Ilaria Bordino, Nicolas Kourtellis, Nikolay Laptev, Youssef Billawala
ICDE3
2013 Very fast estimation for result and accuracy of big data analytics: The EARL system
abstract
Approximate results based on samples often provide the only way in which advanced analytical applications on very massive data sets (a.k.a. `big data') can satisfy their time and resource constraints. Unfortunately, methods and tools for the computation of accurate early results are currently not supported in big data systems (e.g., Hadoop). Therefore, we propose a nonparametric accuracy estimation method and system to speedup big data analytics. Our framework is called EARL (Early Accurate Result Library) and it works by predicting the learning curve and choosing the appropriate sample size for achieving the desired error bound specified by the user. The error estimates are based on a technique called bootstrapping that has been widely used and validated by statisticians, and can be applied to arbitrary functions and data distributions. Therefore, this demo will elucidate (a) the functionality of EARL and its intuitive GUI interface whereby first-time users can appreciate the accuracy obtainable from increasing sample sizes by simply viewing the learning curve displayed by EARL, (b) the usability of EARL, whereby conference participants can interact with the system to quickly estimate the sample sizes needed to obtain the desired accuracies or response times, and then compare them against the accuracies and response times obtained in the actual computations.
Nikolay Laptev, Kai Zeng 0002, Carlo Zaniolo
ICDE1
2012 Optimization of Massive Pattern Queries by Dynamic Configuration Morphing
abstract
Complex pattern queries play a critical role in many applications that must efficiently search databases and data streams. Current techniques support the search for multiple patterns using deterministic or non-deterministic automata. In practice however, the static pattern representation does not fully utilize available system resources, subsequently suffering from poor performance. Therefore a low overhead auto-reconfigurable automaton is needed that optimizes pattern matching performance. In this paper, we propose a dynamic system that entails the efficient and reliable evaluation of a very large number of pattern queries on a resource constrained system under changing stress-load. Our system prototype, Morpheus, pre-computes several query pattern representations, named templates, which are then morphed into a required form during run-time. Morpheus uses templates to speed up dynamic automaton reconfiguration. Results from empirical studies confirm the benefits of our approach, with three orders of magnitude improvement achieved in the overall pattern matching performance with the help of dynamic reconfiguration. This is accomplished only with a modest increase in amortized memory usage.
Nikolay Laptev, Carlo Zaniolo
ICDE1
2012 Early Accurate Results for Advanced Analytics on MapReduce
abstract
Approximate results based on samples often provide the only way in which advanced analytical applications on very massive data sets can satisfy their time and resource constraints. Unfortunately, methods and tools for the computation of accurate early results are currently not supported in MapReduce-oriented systems although these are intended for 'big data'. Therefore, we proposed and implemented a non-parametric extension of Hadoop which allows the incremental computation of early results for arbitrary work-flows, along with reliable on-line estimates of the degree of accuracy achieved so far in the computation. These estimates are based on a technique called bootstrapping that has been widely employed in statistics and can be applied to arbitrary functions and data distributions. In this paper, we describe our Early Accurate Result Library (EARL) for Hadoop that was designed to minimize the changes required to the MapReduce framework. Various tests of EARL of Hadoop are presented to characterize the frequent situations where EARL can provide major speed-ups over the current version of Hadoop.
Nikolay Laptev, Kai Zeng 0002, Carlo Zaniolo
Proc. VLDB Endow.1
2011 SMM: A data stream management system for knowledge discovery
abstract
The problem of supporting data mining applications proved to be difficult for database management systems and it is now proving to be very challenging for data stream management systems (DSMSs), where the limitations of SQL are made even more severe by the requirements of continuous queries. The major technical advances that achieved separately on DSMSs and on data stream mining algorithms have failed to converge and produce powerful data stream mining systems. Such systems, however, are essential since the traditional pull-based approach of cache mining is no longer applicable, and the push-based computing mode of data streams and their bursty traffic complicate application development. For instance, to write mining applications with quality of service (QoS) levels approaching those of DSMSs, a mining analyst would have to contend with many arduous tasks, such as support for data buffering, complex storage and retrieval methods, scheduling, fault-tolerance, synopsis-management, load shedding, and query optimization. Our Stream Mill Miner (SMM) system solves these problems by providing a data stream mining workbench that combines the ease of specifying high-level mining tasks, as in Weka, with the performance and QoS guarantees of a DSMS. This is accomplished in three main steps. The first is an open and extensible DSMS architecture where KDD queries can be easily expressed as user-defined aggregates (UDAs) - our system combines that with the efficiency of synoptic data structures and mining-aware load shedding and optimizations. The second key component of SMM is its integrated library of fast mining algorithms that are light enough to be effective on data streams. The third advanced feature of SMM is a Mining Model Definition Language (MMDL) that allows users to define the flow of mining tasks, integrated with a simple box&arrow GUI, to shield the mining analyst from the complexities of lower-level queries. SMM is the first DSMS capable of online mining and this paper describes its architecture, design, and performance on mining queries.
Hetal Thakkar, Nikolay Laptev, Hamid Mousavi 0001, Barzan Mozafari, Vincenzo Russo, Carlo Zaniolo
ICDE2
2009 Xquasher: a tool for efficient computation of multiple linear expressions
abstract
Digital signal processing applications often require the computation of linear systems. These computations can be considerably expensive and require optimizations for lower power consumption, higher throughput, and faster response time. Unfortunately, system designers do not have the necessary tools to take advantage of the wide flexibility in ways to evaluate these expressions. Therefore, we address the problem of efficiently computing a set of linear systems through a tool, Xquasher, that is developed by us to enable elimination of large common subexpression from expressions with an arbitrary number of terms. Xquasher provides a methodology for efficient computation of both single and multiple linear expressions. We also introduce the concept of power set encoding which helps us to provide an effective optimization method and achieves significant improvement over previously published work. Our tool provides optimized designs with 15% less area with the cost of 3% increase in delay by reducing number of additions on average by 45%.
Arash Arfaee, Ali Irturk, Nikolay Laptev, Farzan Fallah, Ryan Kastner
DAC3
2009 Architectural optimization of decomposition algorithms for wireless communication systems
abstract
Matrix decomposition is required in various algorithms used in wireless communication applications. FPGAs strike a balance between ASICs and DSPs, as they have the programmability of software with performance capacity approaching that of a custom hardware implementation. However, FPGA architectures require designers to make a countless number of system, architectural and logic design decisions. By performing design space exploration, a designer can find the optimal device for a specific application, however very few tools exist which can accomplish this task. This paper presents automatic generation and optimization of decomposition methods using a core generator tool, GUSTO, that we developed to enable easy design space exploration with different parameterization options such as resource allocation, bit widths of the data, number of functional units and organization of controllers and interconnects. We present a detailed study of area and throughput tradeoffs of matrix decomposition architectures using different parameterizations.
Ali Irturk, Bridget Benson, Nikolay Laptev, Ryan Kastner
WCNC3