EDBT 2026 Demo / reviewers in the wild / expert
Nikolay Laptev
dblp:96/7449
· DBLP profile ↗
14ranked-venue papers
4as first author
2since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 9 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Systems, architecture and hardware · 1Computer networks · 1Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
10 papers |
Query processing and optimization · 39% Data mining · 26% Data stream processing · 11% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Cloud and datacenter computing · 61% Performance modeling and evaluation · 18% Electronic design automation · 15% | |
| Artificial intelligence
1 paper |
Probabilistic and Bayesian machine learning · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational finance and economics · 100% |
Topics — the 27 heaviest of 31, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization › query optimization
learned query optimization |
0.7 | 1 | 2023 | AutoSteer: Learned Query Optimization for Any SQL Database · Proc. VLDB Endow. 2023 |
Cloud and datacenter computing
resource allocation |
0.4 | 1 | 2019 | Probabilistic Modeling of Computing Demand for Service Level Agreement · IEEE Trans. Serv. Comput. 2019 |
Query processing and optimization
approximate query processing |
0.3 | 2 | 2013 | Very fast estimation for result and accuracy of big data analytics: The EARL system · ICDE 2013 Early Accurate Results for Advanced Analytics on MapReduce · Proc. VLDB Endow. 2012 |
Web and social media mining › user behavior analysis
user behavior modeling |
0.2 | 1 | 2016 | DECT: Distributed Evolving Context Tree for Understanding User Behavior Pattern Evolution · AAAI 2016 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
hidden markov model |
0.2 | 1 | 2015 | Inertial Hidden Markov Models: Modeling Change in Multivariate Time Series · AAAI 2015 |
Data mining
anomaly detection |
0.2 | 1 | 2015 | Generic and Scalable Framework for Automated Time-series Anomaly Detection · KDD 2015 |
Data mining
time series analysis |
0.2 | 1 | 2015 | Inertial Hidden Markov Models: Modeling Change in Multivariate Time Series · AAAI 2015 |
Data mining › anomaly detection
time series anomaly detection |
0.2 | 1 | 2015 | Generic and Scalable Framework for Automated Time-series Anomaly Detection · KDD 2015 |
Data mining › time series analysis
time series segmentation |
0.2 | 1 | 2015 | Inertial Hidden Markov Models: Modeling Change in Multivariate Time Series · AAAI 2015 |
Computational finance and economics › financial market prediction
stock prediction |
0.2 | 1 | 2014 | Stock trade volume prediction with Yahoo Finance user browsing behavior · ICDE 2014 |
Web and social media mining › web usage mining
clickstream analysis |
0.2 | 1 | 2014 | Stock trade volume prediction with Yahoo Finance user browsing behavior · ICDE 2014 |
Data stream processing
complex event processing |
0.1 | 1 | 2012 | Optimization of Massive Pattern Queries by Dynamic Configuration Morphing · ICDE 2012 |
Data stream processing › continuous query processing
early result production |
0.1 | 1 | 2012 | Early Accurate Results for Advanced Analytics on MapReduce · Proc. VLDB Endow. 2012 |
Information retrieval
pattern matching |
0.1 | 1 | 2012 | Optimization of Massive Pattern Queries by Dynamic Configuration Morphing · ICDE 2012 |
Data mining
data stream mining |
0.1 | 1 | 2011 | SMM: A data stream management system for knowledge discovery · ICDE 2011 |
Data stream processing
stream processing systems |
0.1 | 1 | 2011 | SMM: A data stream management system for knowledge discovery · ICDE 2011 |
Performance modeling and evaluation › statistical analysis
extreme value theory |
0.1 | 1 | 2019 | Probabilistic Modeling of Computing Demand for Service Level Agreement · IEEE Trans. Serv. Comput. 2019 |
Performance modeling and evaluation
workload characterization |
0.1 | 1 | 2019 | Probabilistic Modeling of Computing Demand for Service Level Agreement · IEEE Trans. Serv. Comput. 2019 |
Electronic design automation › logic synthesis › logic optimization
common subexpression elimination |
0.1 | 1 | 2009 | Xquasher: a tool for efficient computation of multiple linear expressions · DAC 2009 |
Electronic design automation
high-level synthesis |
0.1 | 1 | 2009 | Xquasher: a tool for efficient computation of multiple linear expressions · DAC 2009 |
Data mining › predictive modeling
forecasting |
0.1 | 1 | 2015 | Generic and Scalable Framework for Automated Time-series Anomaly Detection · KDD 2015 |
Recommender systems › user modeling
user intent modeling |
0.1 | 1 | 2014 | Stock trade volume prediction with Yahoo Finance user browsing behavior · ICDE 2014 |
Data mining
big data analytics |
0.0 | 1 | 2013 | Very fast estimation for result and accuracy of big data analytics: The EARL system · ICDE 2013 |
Query processing and optimization › query optimization › graph query optimization
pattern query optimization |
0.0 | 1 | 2012 | Optimization of Massive Pattern Queries by Dynamic Configuration Morphing · ICDE 2012 |
Query processing and optimization
query optimization |
0.0 | 1 | 2012 | Optimization of Massive Pattern Queries by Dynamic Configuration Morphing · ICDE 2012 |
Parallel and multicore computing › data-parallel programming
mapreduce |
0.0 | 1 | 2012 | Early Accurate Results for Advanced Analytics on MapReduce · Proc. VLDB Endow. 2012 |
Integrated circuit design
digital signal processing circuits |
0.0 | 1 | 2009 | Xquasher: a tool for efficient computation of multiple linear expressions · DAC 2009 |
Methods — techniques the papers use, named apart from their topics
visualization · 0.7reinforcement learning · 0.7bandit optimization · 0.7bootstrapping · 0.5state transition regularization · 0.4expectation-maximization · 0.4probabilistic modeling · 0.4extreme value theory · 0.4web traffic analysis · 0.4correlation analysis · 0.4variable-order markov model · 0.2learning curve prediction · 0.2non-parametric estimation · 0.1power set encoding · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | QO-Insight: Inspecting Steered Query OptimizersabstractSteered query optimizers address the planning mistakes of traditional query optimizers by providing them with hints on a per-query basis, thereby guiding them in the right direction. This paper introduces QO-Insight, a visual tool designed for exploring query execution traces of such steered query optimizers. Although steered query optimizers are typically perceived as black boxes, QO-Insight empowers database administrators and experts to gain qualitative insights and enhance their performance through visual inspection and analysis. Christoph Anneser, Mario Petruccelli, Nesime Tatbul, David E. Cohen, Zhenggang Xu, Prithviraj Pandian, Nikolay Laptev, Ryan Marcus, Alfons Kemper |
Proc. VLDB Endow. | 7 |
| 2023 | AutoSteer: Learned Query Optimization for Any SQL DatabaseabstractThis paper presents AutoSteer, a learning-based solution that automatically drives query optimization in any SQL database that exposes tunable optimizer knobs. AutoSteer builds on the Bandit optimizer (Bao) and extends it with new capabilities (e.g., automated hint-set discovery) to minimize integration effort and facilitate usability in both monolithic and disaggregated SQL systems. We successfully applied AutoSteer on PostgreSQL, PrestoDB, Spark-SQL, MySQL, and DuckDB - five popular open-source database engines with diverse query optimizers. We then conducted a detailed experimental evaluation with public benchmarks (JOB, Stackoverflow, TPC-DS) and a production workload from Meta's PrestoDB deployments. Our evaluation shows that AutoSteer can not only outperform these engines' native query optimizers (e.g., up to 40% improvements for PrestoDB) but can also match the performance of Bao-for-PostgreSQL with reduced human supervision and increased adaptivity, as it replaces Bao's static, expert-picked hint-sets with those that are automatically discovered. We also provide an open-source implementation of AutoSteer together with a visual tool for interactive use by query optimization experts. Christoph Anneser, Nesime Tatbul, David E. Cohen, Zhenggang Xu, Prithviraj Pandian, Nikolay Laptev, Ryan Marcus |
Proc. VLDB Endow. | 6 |
| 2019 | Probabilistic Modeling of Computing Demand for Service Level AgreementabstractCloud computing applications must be allocated sufficient resources to comply with Service Level Agreements (SLAs). This paper considers data-driven probabilistic modeling of application resource demand for resource allocation. The modeling method is focused on peak demand and SLA violations and relies on a branch of statistics known as extreme value theory (EVT). Rigorous statistical validation of the proposed model shows that it generalizes better than the alternative. The paper presents a resource allocation algorithm using the model to ensure a given small SLA violation rate. For resource allocation in Yahoo data center, over 50 percent savings are demonstrated using the proposed approach. Saahil Shenoy, Dimitry M. Gorinevsky, Nikolay Laptev |
IEEE Trans. Serv. Comput. | 3 |
| 2016 | DECT: Distributed Evolving Context Tree for Understanding User Behavior Pattern EvolutionabstractInternet user behavior models characterize user browsing dynamics or the transitions among web pages. The models help Internet companies improve their services by accurately targeting customers and providing them the information they want. For instance, specific web pages can be customized and prefetched for individuals based on sequences of web pages they have visited. Existing user behavior models abstracted as time-homogeneous Markov models cannot efficiently model user behavior variation through time. This demo presents DECT, a scalable time-variant variable-order Markov model. DECT digests terabytes of user session data and yields user behavior patterns through time. We realize DECT using Apache Spark and deploy it on top of Yahoo! infrastructure. We demonstrate the benefits of DECT with anomaly detection and ad click rate prediction applications. DECT enables the detection of higher-order path anomalies and provides deep insights into ad click rates with respect to user visiting paths. Xiaokui Shu, Nikolay Laptev, Danfeng Yao |
AAAI | 2 |
| 2016 | DECT: Distributed Evolving Context Tree for Mining Web Behavior EvolutionabstractInternet user behavior models characterize user browsing dynamics or the transitions among web pages. The mod- els help Internet companies improve their services by accu- rately targeting customers and providing them the informa- tion they want. For instance, specic web pages can be cus- tomized and prefetched for individuals based on sequences of web pages they have visited. Existing user behavior mod- els abstracted as time-homogeneous Markov models do not provide efficient support for modeling user behavior varia- tion through time. This paper presents DECT, a scalable time-variant variable-order Markov model. DECT digests terabytes of user session data and yields user behavior pat- terns through time. We realize DECT using Apache Spark. Our implementation is being open-sourced and we deploy DECT on top of Yahoo! infrastructure. We demonstrate the benets of DECT with anomaly detection and ad click rate prediction applications. DECT enables the detection of higher-order path anomalies that are masked out by exist- ing models. DECT also provides insights into ad click rates with respect to user visiting paths. Xiaokui Shu, Nikolay Laptev, Danfeng Yao |
EDBT | 2 |
| 2015 | Inertial Hidden Markov Models: Modeling Change in Multivariate Time SeriesabstractFaced with the problem of characterizing systematic changes in multivariate time series in an unsupervised manner, we derive and test two methods of regularizing hidden Markov models for this task. Regularization on state transitions provides smooth transitioning among states, such that the sequences are split into broad, contiguous segments. Our methods are compared with a recent hierarchical Dirichlet process hidden Markov model (HDP-HMM) and a baseline standard hidden Markov model, of which the former suffers from poor performance on moderate-dimensional data and sensitivity to parameter settings, while the latter suffers from rapid state transitioning, over-segmentation and poor performance on a segmentation task involving human activity accelerometer data from the UCI Repository. The regularized methods developed here are able to perfectly characterize change of behavior in the human activity data for roughly half of the real-data test cases, with accuracy of 94% and low variation of information. In contrast to the HDP-HMM, our methods provide simple, drop-in replacements for standard hidden Markov model update rules, allowing standard expectation maximization (EM) algorithms to be used for learning. George D. Montañez, Saeed Amizadeh, Nikolay Laptev |
AAAI | 3 |
| 2015 | Generic and Scalable Framework for Automated Time-series Anomaly DetectionabstractThis paper introduces a generic and scalable framework for automated anomaly detection on large scale time-series data. Early detection of anomalies plays a key role in maintaining consistency of person's data and protects corporations against malicious attackers. Current state of the art anomaly detection approaches suffer from scalability, use-case restrictions, difficulty of use and a large number of false positives. Our system at Yahoo, EGADS, uses a collection of anomaly detection and forecasting models with an anomaly filtering layer for accurate and scalable anomaly detection on time-series. We compare our approach against other anomaly detection systems on real and synthetic data with varying time-series characteristics. We found that our framework allows for 50-60% improvement in precision and recall for a variety of use-cases. Both the data and the framework are being open-sourced. The open-sourcing of the data, in particular, represents the first of its kind effort to establish the standard benchmark for anomaly detection. Nikolay Laptev, Saeed Amizadeh, Ian Flint |
KDD | 1 |
| 2014 | Stock trade volume prediction with Yahoo Finance user browsing behaviorabstractWeb traffic represents a powerful mirror for various real-world phenomena. For example, it was shown that web search volumes have a positive correlation with stock trading volumes and with the sentiment of investors. Our hypothesis is that user browsing behavior on a domain-specific portal is a better predictor of user intent than web searches. Ilaria Bordino, Nicolas Kourtellis, Nikolay Laptev, Youssef Billawala |
ICDE | 3 |
| 2013 | Very fast estimation for result and accuracy of big data analytics: The EARL systemabstractApproximate results based on samples often provide the only way in which advanced analytical applications on very massive data sets (a.k.a. `big data') can satisfy their time and resource constraints. Unfortunately, methods and tools for the computation of accurate early results are currently not supported in big data systems (e.g., Hadoop). Therefore, we propose a nonparametric accuracy estimation method and system to speedup big data analytics. Our framework is called EARL (Early Accurate Result Library) and it works by predicting the learning curve and choosing the appropriate sample size for achieving the desired error bound specified by the user. The error estimates are based on a technique called bootstrapping that has been widely used and validated by statisticians, and can be applied to arbitrary functions and data distributions. Therefore, this demo will elucidate (a) the functionality of EARL and its intuitive GUI interface whereby first-time users can appreciate the accuracy obtainable from increasing sample sizes by simply viewing the learning curve displayed by EARL, (b) the usability of EARL, whereby conference participants can interact with the system to quickly estimate the sample sizes needed to obtain the desired accuracies or response times, and then compare them against the accuracies and response times obtained in the actual computations. Nikolay Laptev, Kai Zeng 0002, Carlo Zaniolo |
ICDE | 1 |
| 2012 | Optimization of Massive Pattern Queries by Dynamic Configuration MorphingabstractComplex pattern queries play a critical role in many applications that must efficiently search databases and data streams. Current techniques support the search for multiple patterns using deterministic or non-deterministic automata. In practice however, the static pattern representation does not fully utilize available system resources, subsequently suffering from poor performance. Therefore a low overhead auto-reconfigurable automaton is needed that optimizes pattern matching performance. In this paper, we propose a dynamic system that entails the efficient and reliable evaluation of a very large number of pattern queries on a resource constrained system under changing stress-load. Our system prototype, Morpheus, pre-computes several query pattern representations, named templates, which are then morphed into a required form during run-time. Morpheus uses templates to speed up dynamic automaton reconfiguration. Results from empirical studies confirm the benefits of our approach, with three orders of magnitude improvement achieved in the overall pattern matching performance with the help of dynamic reconfiguration. This is accomplished only with a modest increase in amortized memory usage. Nikolay Laptev, Carlo Zaniolo |
ICDE | 1 |
| 2012 | Early Accurate Results for Advanced Analytics on MapReduceabstractApproximate results based on samples often provide the only way in which advanced analytical applications on very massive data sets can satisfy their time and resource constraints. Unfortunately, methods and tools for the computation of accurate early results are currently not supported in MapReduce-oriented systems although these are intended for 'big data'. Therefore, we proposed and implemented a non-parametric extension of Hadoop which allows the incremental computation of early results for arbitrary work-flows, along with reliable on-line estimates of the degree of accuracy achieved so far in the computation. These estimates are based on a technique called bootstrapping that has been widely employed in statistics and can be applied to arbitrary functions and data distributions. In this paper, we describe our Early Accurate Result Library (EARL) for Hadoop that was designed to minimize the changes required to the MapReduce framework. Various tests of EARL of Hadoop are presented to characterize the frequent situations where EARL can provide major speed-ups over the current version of Hadoop. Nikolay Laptev, Kai Zeng 0002, Carlo Zaniolo |
Proc. VLDB Endow. | 1 |
| 2011 | SMM: A data stream management system for knowledge discoveryabstractThe problem of supporting data mining applications proved to be difficult for database management systems and it is now proving to be very challenging for data stream management systems (DSMSs), where the limitations of SQL are made even more severe by the requirements of continuous queries. The major technical advances that achieved separately on DSMSs and on data stream mining algorithms have failed to converge and produce powerful data stream mining systems. Such systems, however, are essential since the traditional pull-based approach of cache mining is no longer applicable, and the push-based computing mode of data streams and their bursty traffic complicate application development. For instance, to write mining applications with quality of service (QoS) levels approaching those of DSMSs, a mining analyst would have to contend with many arduous tasks, such as support for data buffering, complex storage and retrieval methods, scheduling, fault-tolerance, synopsis-management, load shedding, and query optimization. Our Stream Mill Miner (SMM) system solves these problems by providing a data stream mining workbench that combines the ease of specifying high-level mining tasks, as in Weka, with the performance and QoS guarantees of a DSMS. This is accomplished in three main steps. The first is an open and extensible DSMS architecture where KDD queries can be easily expressed as user-defined aggregates (UDAs) - our system combines that with the efficiency of synoptic data structures and mining-aware load shedding and optimizations. The second key component of SMM is its integrated library of fast mining algorithms that are light enough to be effective on data streams. The third advanced feature of SMM is a Mining Model Definition Language (MMDL) that allows users to define the flow of mining tasks, integrated with a simple box&arrow GUI, to shield the mining analyst from the complexities of lower-level queries. SMM is the first DSMS capable of online mining and this paper describes its architecture, design, and performance on mining queries. Hetal Thakkar, Nikolay Laptev, Hamid Mousavi 0001, Barzan Mozafari, Vincenzo Russo, Carlo Zaniolo |
ICDE | 2 |
| 2009 | Xquasher: a tool for efficient computation of multiple linear expressionsabstractDigital signal processing applications often require the computation of linear systems. These computations can be considerably expensive and require optimizations for lower power consumption, higher throughput, and faster response time. Unfortunately, system designers do not have the necessary tools to take advantage of the wide flexibility in ways to evaluate these expressions. Therefore, we address the problem of efficiently computing a set of linear systems through a tool, Xquasher, that is developed by us to enable elimination of large common subexpression from expressions with an arbitrary number of terms. Xquasher provides a methodology for efficient computation of both single and multiple linear expressions. We also introduce the concept of power set encoding which helps us to provide an effective optimization method and achieves significant improvement over previously published work. Our tool provides optimized designs with 15% less area with the cost of 3% increase in delay by reducing number of additions on average by 45%. Arash Arfaee, Ali Irturk, Nikolay Laptev, Farzan Fallah, Ryan Kastner |
DAC | 3 |
| 2009 | Architectural optimization of decomposition algorithms for wireless communication systemsabstractMatrix decomposition is required in various algorithms used in wireless communication applications. FPGAs strike a balance between ASICs and DSPs, as they have the programmability of software with performance capacity approaching that of a custom hardware implementation. However, FPGA architectures require designers to make a countless number of system, architectural and logic design decisions. By performing design space exploration, a designer can find the optimal device for a specific application, however very few tools exist which can accomplish this task. This paper presents automatic generation and optimization of decomposition methods using a core generator tool, GUSTO, that we developed to enable easy design space exploration with different parameterization options such as resource allocation, bit widths of the data, number of functional units and organization of controllers and interconnects. We present a detailed study of area and throughput tradeoffs of matrix decomposition architectures using different parameterizations. Ali Irturk, Bridget Benson, Nikolay Laptev, Ryan Kastner |
WCNC | 3 |