Valentin Flunkert

dblp:191/6737 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
2since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 since 2021Databases, data management, data science and information retrieval · 3Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Time series and sequential data · 34% Generative modeling · 22% Efficient and distributed learning · 16%
Databases, data mining, and information retrieval
3 papers
Data mining · 86% Machine learning and data management · 14%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 41% Distributed systems · 41% Cloud and datacenter computing · 18%

Topics — the 20 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Time series and sequential data
anomaly detection
0.622022
GluonTS: Probabilistic and Neural Time Series Modeling in Python · J. Mach. Learn. Res. 2020
Neural Contextual Anomaly Detection for Time Series · IJCAI 2022
Machine learning › Generative modeling
generative adversarial network
0.612022
PSA-GAN: Progressive Self Attention GANs for Synthetic Time Series · ICLR 2022
Machine learning › Representation and self-supervised learning › representation learning › sequence representation learning
multivariate time-series representation learning
0.612022
Neural Contextual Anomaly Detection for Time Series · IJCAI 2022
Machine learning › Generative modeling › diffusion model
time series generation
0.612022
PSA-GAN: Progressive Self Attention GANs for Synthetic Time Series · ICLR 2022
Data mining
anomaly detection
0.612022
Neural Contextual Anomaly Detection for Time Series · IJCAI 2022
Data mining › anomaly detection
time series anomaly detection
0.612022
Neural Contextual Anomaly Detection for Time Series · IJCAI 2022
Machine learning › Time series and sequential data › time series modeling
probabilistic forecasting
0.522020
GluonTS: Probabilistic and Neural Time Series Modeling in Python · J. Mach. Learn. Res. 2020
Probabilistic Demand Forecasting at Scale · Proc. VLDB Endow. 2017
Machine learning › Efficient and distributed learning
distributed training
0.412020
Elastic Machine Learning Algorithms in Amazon SageMaker · SIGMOD Conference 2020
Machine learning › Efficient and distributed learning › distributed training › distributed training systems
elastic training
0.412020
Elastic Machine Learning Algorithms in Amazon SageMaker · SIGMOD Conference 2020
Machine learning › Optimization for machine learning
hyperparameter optimization
0.412020
Elastic Machine Learning Algorithms in Amazon SageMaker · SIGMOD Conference 2020
Machine learning › Time series and sequential data
time series modeling
0.412020
GluonTS: Probabilistic and Neural Time Series Modeling in Python · J. Mach. Learn. Res. 2020
Data mining › time series analysis
time series forecasting
0.412019
Forecasting Big Time Series: Theory and Practice · KDD 2019
Data mining › predictive modeling › forecasting
demand prediction
0.312017
Probabilistic Demand Forecasting at Scale · Proc. VLDB Endow. 2017
Machine learning and data management
machine learning pipeline
0.312017
Probabilistic Demand Forecasting at Scale · Proc. VLDB Endow. 2017
High-performance computing
cluster computing
0.312017
Probabilistic Demand Forecasting at Scale · Proc. VLDB Endow. 2017
Distributed systems
distributed machine learning
0.312017
Probabilistic Demand Forecasting at Scale · Proc. VLDB Endow. 2017
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.212016
Bayesian Intermittent Demand Forecasting for Large Inventories · NIPS 2016
Machine learning › Time series and sequential data › time series modeling
demand forecasting
0.212016
Bayesian Intermittent Demand Forecasting for Large Inventories · NIPS 2016
Machine learning › Deep learning architectures and training › attention mechanism
self-attention
0.212022
PSA-GAN: Progressive Self Attention GANs for Synthetic Time Series · ICLR 2022
Computational finance and economics
electronic commerce
0.112016
Bayesian Intermittent Demand Forecasting for Large Inventories · NIPS 2016

Methods — techniques the papers use, named apart from their topics

representation learning · 1.1deep anomaly detection · 1.1resumable training · 0.9incremental training · 0.9feature engineering · 0.9ensembling · 0.9statistical modeling · 0.8progressive self-attention · 0.6GAN · 0.6newton-raphson · 0.5kalman smoothing · 0.5
YearPublicationVenuePosition
2022 PSA-GAN: Progressive Self Attention GANs for Synthetic Time Series
Paul Jeha, Michael Bohlke-Schneider, Shubham Kapoor, Rajbir-Singh Nirwan, Valentin Flunkert, Jan Gasthaus, Tim Januschowski
ICLR6
2022 Neural Contextual Anomaly Detection for Time Series
abstract
We introduce Neural Contextual Anomaly Detection (NCAD), a framework for anomaly detection on time series that scales seamlessly from the unsupervised to supervised setting, and is applicable to both univariate and multivariate time series. This is achieved by combining recent developments in representation learning for multivariate time series, with techniques for deep anomaly detection originally developed for computer vision that we tailor to the time series setting. Our window-based approach facilitates learning the boundary between normal and anomalous classes by injecting generic synthetic anomalies into the available data. NCAD can effectively take advantage of domain knowledge and of any available training labels. We demonstrate empirically on standard benchmark datasets that our approach obtains a state-of-the-art performance in the supervised, semi-supervised, and unsupervised settings.
Christian Carmona, Francois-Xavier Aubet, Valentin Flunkert, Jan Gasthaus
IJCAI3
2020 Elastic Machine Learning Algorithms in Amazon SageMaker
abstract
There is a large body of research on scalable machine learning (ML). Nevertheless, training ML models on large, continuously evolving datasets is still a difficult and costly undertaking for many companies and institutions. We discuss such challenges and derive requirements for an industrial-scale ML platform. Next, we describe the computational model behind Amazon SageMaker, which is designed to meet such challenges. SageMaker is an ML platform provided as part of Amazon Web Services (AWS), and supports incremental training, resumable and elastic learning as well as automatic hyperparameter optimization. We detail how to adapt several popular ML algorithms to its computational model. Finally, we present an experimental evaluation on large datasets, comparing SageMaker to several scalable, JVM-based implementations of ML algorithms, which we significantly outperform with regard to computation time and cost.
Edo Liberty, Zohar S. Karnin, Bing Xiang, Laurence Rouesnel, Baris Coskun, Ramesh Nallapati, Julio Delgado, Amir Sadoughi, Yury Astashonok, Piali Das, Can Balioglu, Saswata Chakravarty, Madhav Jha, Philip Gautier, David Arpin, Tim Januschowski, Valentin Flunkert, Yuyang Wang 0001, Jan Gasthaus, Lorenzo Stella, Syama Sundar Rangapuram, David Salinas, Sebastian Schelter, Alexander J. Smola
SIGMOD Conference17
2020 GluonTS: Probabilistic and Neural Time Series Modeling in Python
abstract
We introduce the Gluon Time Series Toolkit (GluonTS), a Python library for deep learning based time series modeling for ubiquitous tasks, such as forecasting and anomaly detection. GluonTS simplifies the time series modeling pipeline by providing the necessary components and tools for quick model development, efficient experimentation and evaluation. In addition, it contains reference implementations of state-of-the-art time series models that enable simple benchmarking of new algorithms.
Alexander Alexandrov 0001, Konstantinos Benidis, Michael Bohlke-Schneider, Valentin Flunkert, Jan Gasthaus, Tim Januschowski, Danielle C. Maddix, Syama Sundar Rangapuram, David Salinas, Jasper Schulz, Lorenzo Stella, Ali Caner Türkmen, Yuyang Wang 0001
J. Mach. Learn. Res.4
2019 Probabilistic Forecasting with Spline Quantile Function RNNs
abstract
In this paper, we propose a flexible method for probabilistic modeling with conditional quantile functions using monotonic regression splines. The shape of the spline is parameterized by a neural network whose parameters are learned by minimizing the continuous ranked probability score. Within this framework, we propose a method for probabilistic time series forecasting, which combines the modeling capacity of recurrent neural networks with the flexibility of a spline-based representation of the output distribution. Unlike methods based on parametric probability density functions and maximum likelihood estimation, the proposed method can flexibly adapt to different output distributions without manual intervention. We empirically demonstrate the effectiveness of the approach on synthetic and real-world data sets.
Jan Gasthaus, Konstantinos Benidis, Yuyang Wang 0001, Syama Sundar Rangapuram, David Salinas, Valentin Flunkert, Tim Januschowski
AISTATS6
2019 Forecasting Big Time Series: Theory and Practice
abstract
Time series forecasting is a key ingredient in the automation and optimization of business processes: in retail, deciding which products to order and where to store them depends on the forecasts of future demand in different regions; in cloud computing, the estimated future usage of services and infrastructure components guides capacity planning; and workforce scheduling in warehouses and factories requires forecasts of the future workload. Recent years have witnessed a paradigm shift in forecasting techniques and applications, from computer-assisted model- and assumption-based to data-driven and fully-automated. This shift can be attributed to the availability of large, rich, and diverse time series data sources and result in a set of challenges that need to be addressed such as the following. How can we build statistical models to efficiently and effectively learn to forecast from large and diverse data sources? How can we leverage the statistical power of "similar'' time series to improve forecasts in the case of limited observations? What are the implications for building forecasting systems that can handle large data volumes?
Christos Faloutsos, Valentin Flunkert, Jan Gasthaus, Tim Januschowski, Yuyang Wang 0001
KDD2
2017 Probabilistic Demand Forecasting at Scale
abstract
We present a platform built on large-scale, data-centric machine learning (ML) approaches, whose particular focus is demand forecasting in retail. At its core, this platform enables the training and application of probabilistic demand forecasting models, and provides convenient abstractions and support functionality for forecasting problems. The platform comprises of a complex end-to-end machine learning system built on Apache Spark, which includes data preprocessing, feature engineering, distributed learning, as well as evaluation, experimentation and ensembling. Furthermore, it meets the demands of a production system and scales to large catalogues containing millions of items. We describe the challenges of building such a platform and discuss our design decisions. We detail aspects on several levels of the system, such as a set of general distributed learning schemes, our machinery for ensembling predictions, and a high-level dataflow abstraction for modeling complex ML pipelines. To the best of our knowledge, we are not aware of prior work on real-world demand forecasting systems which rivals our approach in terms of scalability.
Joos-Hendrik Böse, Valentin Flunkert, Jan Gasthaus, Tim Januschowski, Dustin Lange, David Salinas, Sebastian Schelter, Matthias W. Seeger, Yuyang Wang 0001
Proc. VLDB Endow.2
2016 Bayesian Intermittent Demand Forecasting for Large Inventories
abstract
We present a scalable and robust Bayesian method for demand forecasting in the context of a large e-commerce platform, paying special attention to intermittent and bursty target statistics. Inference is approximated by the Newton-Raphson algorithm, reduced to linear-time Kalman smoothing, which allows us to operate on several orders of magnitude larger problems than previous related work. In a study on large real-world sales datasets, our method outperforms competing approaches on fast and medium moving items.
Matthias W. Seeger, David Salinas, Valentin Flunkert
NIPS3