Nestor Cabello

dblp:285/3278 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
4since 2021 · last 2026
0000-0002-6136-599XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Adaptive Rich-kernelized Contrastive Learning for Capacity Enhancement in Collaborative Filtering
abstract
Recent research has shown that single-vector embedding retrieval models face a fundamental bottleneck: the finite dimensionality of single-vector representations limits their capacity to represent arbitrary top-k relevant item combinations, even with perfect training. This inherent bottleneck substantially limits the expressive capacity of such models and reduces their ability to capture complex user-item interaction patterns. To overcome this bottleneck, we propose Adaptive Rich-kernelized Contrastive Learning (ARC), which enhances model expressiveness while maintaining the computational efficiency of single-vector retrieval. ARC replaces the fixed inner product with a learnable spherical kernel family parameterized by a truncated Gegenbauer expansion, thereby increasing the model's effective dimensionality while preserving single-vector efficiency and low-pass inductive bias for generalization. Specifically, we construct a positive-definite kernel on the unit sphere using a positive combination of Gegenbauer polynomials, and adopt a contrastive learning objective to jointly learn the polynomials weights and user–item embeddings. Through data-driven optimization, the adaptive kernel induces a more expressive representation space, enabling the model to better capture complex preferences. From a theoretical perspective, the learned kernel implicitly maps embeddings into a higher-dimensional Reproducing Kernel Hilbert Space, allowing the model to capture a broader range of top-k item combinations before reaching the geometric limit imposed by the embedding dimensionality. Consequently, the proposed ARC framework effectively alleviates the inherent representational bottleneck in traditional single-vector embedding models. Extensive experiments on four real-world datasets demonstrate the effectiveness of our method, showing consistent improvements over strong baselines and state-of-the-art models.
Ling Luo 0002, Nestor Cabello, Lars Kulik
SIGIR3
2025 Ordinal Embedding for Collaborative Filtering: A Unified Regularization for Enhanced Generalization and Interpretability
abstract
Collaborative filtering is a primary paradigm of modern recommender systems. A typical practice is to embed collaborative signals into a latent space and infer recommendation scores based on the similarities between user and item embeddings. Besides inter-type similarities (i.e., user-item relationships), intra-type similarities (i.e., user-user, and item-item) are also essential as they capture the intrinsic structure of users and items. However, many existing recommendation models only learn inter-type similarities using objectives like ranking loss or binary classification loss, while neglecting intra-type similarities. Consequently, the intrinsic structures of users and items are often distorted in the latent space, where users with similar historical interactions diverge more than those dissimilar. In this study, we show the importance of preserving the ordinal relations of intra-type similarities. We provide a theoretical analysis suggesting that preserving intra-type similarity rankings can enhance a model's generalizability and interpretability. In addition, we propose a regularization that enforces a constraint on the rankings of intra-type similarities, ensuring that learning inter-type similarities does not break intrinsic ordinal structures. It can be seamlessly integrated into most latent factor models and can be jointly trained with their original objectives. Extensive experiments on 4 benchmark datasets and 5 representative models show that our ordinal regularization can consistently improve recommendation performance, and enhance the intra-type similarity coherence in the latent space. The results also exhibit enhanced generalizability and interpretability of recommendations.
Ling Luo 0002, Nestor Cabello, Lars Kulik
CIKM3
2025 PULSAR: Advancing Interval-Based Time Series Classification to State-of-the-Art Performance
abstract
State-of-the-art Time Series Classification (TSC) models such as HC2 and MultiRocket-Hydra reach high accuracy but rely on representations that are inherently difficult to interpret. Interval-based classifiers summarize local segments and thus provide more intuitive features. However, they are usually behind the most accurate methods. Our approach PULSAR—Pooled mUlti-scaLe Summaries from rAndomized inteRvals significantly enhances the accuracy of interval-based approaches and is close to MultiRocket-Hydra. PULSAR first converts every series into several representations (raw, derivative, periodogram, and others). It then uses sub-series of varied length and dilation and computes simple local statistics. To build higher-order features, it partitions each sequence of local statistics at multiple depths, applies a randomly chosen set of pooling operators to every partition, and keeps all aggregates from the coarsest level. For finer partitions we apply a supervised feature selection strategy to retain only the most discriminative features. Finally, we concatenate the selected aggregates with the global ones and train an ensemble classifier. We test PULSAR on 142 UCR datasets. It outperforms all interval-based approaches by a statistically significant margin and matches the predictive performance of HC2 and MultiRocket-Hydra. PULSAR sets a new benchmark for interval-based TSC while preserving a feature structure that we can still interpret.
Nestor Cabello, Lars Kulik
ICDM1
2024 Fast, accurate and explainable time series classification through randomization
abstract
Abstract Time series classification(TSC) aims to predict the class label of a given time series, which is critical to a rich set of application areas such as economics and medicine. State-of-the-art TSC methods have mostly focused on classification accuracy, without considering classification speed. However, efficiency is important for big data analysis. Datasets with a large training size or long series challenge the use of the current highly accurate methods, because they are usually computationally expensive. Similarly, classification explainability, which is an important property required by modern big data applications such asappliance modelingand legislation such as theEuropean General Data Protection Regulation, has received little attention. To address these gaps, we propose a novel TSC method – theRandomized-Supervised Time Series Forest(r-STSF). r-STSF is extremely fast and achieves state-of-the-art classification accuracy. It is an efficient interval-based approach that classifies time series according to aggregate values of the discriminatory sub-series (intervals). To achieve state-of-the-art accuracy, r-STSF builds an ensemble of randomized trees using the discriminatory sub-series. It uses four time series representations, nine aggregation functions and a supervised binary-inspired search combined with a feature ranking metric to identify highly discriminatory sub-series. The discriminatory sub-series enable explainable classifications. Experiments on extensive datasets show that r-STSF achieves state-of-the-art accuracy while being orders of magnitude faster than most existing TSC methods and enabling for explanations on the classifier decision.
Nestor Cabello, Elham Naghizade, Jianzhong Qi 0001, Lars Kulik
Data Min. Knowl. Discov.1
2020 Fast and Accurate Time Series Classification Through Supervised Interval Search
abstract
Time series classification (TSC) aims to predict the class label of a given time series. Modern applications such as appliance modelling require to model an abundance of long time series, which makes it difficult to use many state-of-the-art TSC techniques due to their high computational cost and lack of interpretable outputs. To address these challenges, we propose a novel TSC method: the Supervised Time Series Forest (STSF). STSF improves the classification efficiency by examining only a (set of) sub-series of the original time series, and its tree-based structure allows for interpretable outcomes. STSF adapts a top-down approach to search for relevant sub-series in three different time series representations prior to training any tree classifier, where the relevance of a sub-series is measured by feature ranking metrics (i.e., supervision signals). Experiments on extensive real datasets show that STSF achieves comparable accuracy to state-of-the-art TSC methods while being significantly more efficient, enabling TSC for long time series.
Nestor Cabello, Elham Naghizade, Jianzhong Qi 0001, Lars Kulik
ICDM1