VLDB 2026 Research / reviewers in the wild / expert
Erick Draayer
dblp:299/8601
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2026
0009-0006-6026-7585ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deep clustering for large-scale interpretable time series segmentationabstractTime series segmentation (TSS) is often an unsupervised data mining task that partitions a given time series into homogeneous regions. Existing TSS algorithms either scale poorly or perform poorly on complex large-scale time series (TS) commonly observed in real-world applications. This paper introduces Deep Clustering for Time Series Segmentation (DC-TSS). DC-TSS is a domain-agnostic method that uses a three-phase neural-based model to segment a given time series. DC-TSS includes a carefully designed neural architecture and a newly designed data augmentation approach to efficiently learn TS representations, utilizes a neural-based clustering model to refine such representations, designs a novel efficient component to infer segments from clustered TS representation, and provides mechanisms to understand/interpret the segmentation results. We test DC-TSS on 27 multivariate time series datasets, which are much larger and more complex than others typically used in TSS studies. We also test DC-TSS on a more traditional repository of 98 simpler time series datasets. The experiments from both types of dataset provide an in-depth analysis of DC-TSS’s performance and limitations. We compare five variations of our method against seven strong baselines. The results show that DC-TSS significantly outperforms other methods and scales well to larger and more complex datasets and shows some limitation on shorter simple datasets. DC-TSS addresses a growing need for unsupervised TSS algorithms designed to segment large-scale, complex datasets, which are becoming more common as evolving technology allows collecting and storing greater volumes of data. Erick Draayer, Huiping Cao, Qixu Gong |
Data Min. Knowl. Discov. | 1 |
| 2025 | Benchmarking Large Language Models with Integer Sequence Generation TasksabstractWe present a novel benchmark designed to rigorously evaluate the capabilities of large language models (LLMs) in mathematical reasoning and algorithmic code synthesis tasks. The benchmark comprises integer sequence generation tasks sourced from the Online Encyclopedia of Integer Sequences (OEIS), testing LLMs' abilities to accurately and efficiently generate Python code to compute these sequences without using lookup tables. Our comprehensive evaluation includes leading models from OpenAI (including the specialized reasoning-focused o-series), Anthropic, Meta, and Google across a carefully selected set of 1000 OEIS sequences categorized as easy'' orhard.'' Half of these sequences are classical sequences from the early days of OEIS and half were recently added to avoid contamination with the models' training data. To prevent models from exploiting memorized sequence values, we introduce an automated cheating detection mechanism that flags usage of lookup tables, validated by comparison with human expert evaluations. Experimental results demonstrate that reasoning-specialized models (o3, o3-mini, o4-mini from OpenAI, and Gemini 2.5-pro from Google) achieve substantial improvements in accuracy over non-reasoning models, especially on more complex tasks. However, overall model performance on the hard sequences is poor, highlighting persistent challenges in algorithmic reasoning. Our benchmark provides important insights into the strengths and limitations of state-of-the-art LLMs, particularly emphasizing the necessity for further advancements to reliably solve complex mathematical reasoning tasks algorithmically. Daniel O'Malley, Manish Bhattarai, Nishath Rajiv Ranasinghe, Erick Draayer, Javier E. Santos |
NeurIPS | 4 |
| 2024 | Towards Uncertainty Quantification for Time Series Segmentation
Erick Draayer, Huiping Cao |
CIKM | 1 |
| 2021 | Reevaluating the Change Point Detection Problem with Segment-based Bayesian Online DetectionabstractChange point detection is widely used for finding transitions between states of data generation within a time series. Methods for change point detection currently assume this transition is instantaneous and therefore focus on finding a single point of data to classify as a change point. However, this assumption is flawed because many time series actually display short periods of transitions between different states of data generation. Previous work has shown Bayesian Online Change Point Detection (BOCPD) to be the most effective method for change point detection on a wide range of different time series. This paper explores adapting the change point detection algorithms to detect abrupt changes over short periods of time. We design a segment-based mechanism to examine a window of data points within a time series, rather than a single data point, to determine if the window captures abrupt change. We test our segment-based Bayesian change detection algorithm on 36 different time series and compare it to the original BOCPD algorithm. Our results show that, for some of these 36 time series, the segment-based approach for detecting abrupt changes can much more accurately identify change points based on standard metrics. Erick Draayer, Huiping Cao, Yifan Hao 0003 |
CIKM | 1 |