Antonios Valkanas

dblp:130/7787 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
7since 2021 · last 2025
0009-0001-1234-0016ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Computer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Deep learning architectures and training · 33% Graph learning · 15% Learning paradigms · 15%
Databases, data mining, and information retrieval
2 papers
Data mining · 72% Machine learning and data management · 28%
Computer graphics and multimedia
1 paper
Computer animation and physical simulation · 100%

Topics — the 16 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training › recurrent neural network
linear recurrent neural network
0.912025
SKOLR: Structured Koopman Operator Linear RNN for Time-Series Forecasting · ICML 2025
Natural language and speech › Language models and text generation › large language model inference
LLM cascades
0.912025
C3PO: Optimized Large Language Model Cascades with Probabilistic Cost Constraints for Reasoning · NeurIPS 2025
Machine learning › Deep learning architectures and training
recurrent neural network
0.912025
SKOLR: Structured Koopman Operator Linear RNN for Time-Series Forecasting · ICML 2025
Data mining
time series analysis
0.912025
SKOLR: Structured Koopman Operator Linear RNN for Time-Series Forecasting · ICML 2025
Data mining › time series analysis
time series forecasting
0.912025
SKOLR: Structured Koopman Operator Linear RNN for Time-Series Forecasting · ICML 2025
Computer animation and physical simulation › motion synthesis
human motion synthesis
0.812024
Motion In-Betweening via Deep $\Delta$Δ-Interpolator · IEEE Trans. Vis. Comput. Graph. 2024
Computer animation and physical simulation › motion synthesis › motion interpolation
motion in-betweening
0.812024
Motion In-Betweening via Deep $\Delta$Δ-Interpolator · IEEE Trans. Vis. Comput. Graph. 2024
Machine learning and data management › continual learning
incremental learning
0.712023
Structure Aware Incremental Learning with Personalized Imitation Weights for Recommender Systems · AAAI 2023
Machine learning › Graph learning › graph neural network › graph neural network architecture
bayesian graph neural networks
0.612022
Bag Graph: Multiple Instance Learning Using Bayesian Graph Neural Networks · AAAI 2022
Machine learning › Graph learning
graph neural network
0.612022
Bag Graph: Multiple Instance Learning Using Bayesian Graph Neural Networks · AAAI 2022
Machine learning › Learning paradigms
multiple instance learning
0.612022
Bag Graph: Multiple Instance Learning Using Bayesian Graph Neural Networks · AAAI 2022
Machine learning › Learning paradigms
weakly supervised learning
0.612022
Bag Graph: Multiple Instance Learning Using Bayesian Graph Neural Networks · AAAI 2022
Machine learning › Trustworthy machine learning › uncertainty estimation
conformal prediction
0.312025
C3PO: Optimized Large Language Model Cascades with Probabilistic Cost Constraints for Reasoning · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning
dynamical system
0.312025
SKOLR: Structured Koopman Operator Linear RNN for Time-Series Forecasting · ICML 2025
Machine learning › Representation and self-supervised learning › dynamical system representation
koopman operator
0.312025
SKOLR: Structured Koopman Operator Linear RNN for Time-Series Forecasting · ICML 2025
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.212023
Structure Aware Incremental Learning with Personalized Imitation Weights for Recommender Systems · AAAI 2023

Methods — techniques the papers use, named apart from their topics

spectral decomposition · 1.7koopman operator theory · 1.7MLP · 1.7spherical linear interpolation · 1.5delta-mode learning · 1.5imitation weights · 1.3graph neural network · 1.3self-supervised learning · 0.9regret minimization · 0.9conformal prediction · 0.9knowledge distillation · 0.7
YearPublicationVenuePosition
2025 MODL: Multilearner Online Deep Learning
abstract
Online deep learning tackles the challenge of learning from data streams by balancing two competing goals: fast learning and deep learning. However, existing research primarily emphasizes deep learning solutions, which are more adept at handling the ”deep” aspect than the ”fast” aspect of online learning. In this work, we introduce an alternative paradigm through a hybrid multilearner approach. We begin by developing a fast online logistic regression learner, which operates without relying on backpropagation. It leverages closed-form recursive updates of model parameters, efficiently addressing the fast learning component of the online learning challenge. This approach is further integrated with a cascaded multilearner design, where shallow and deep learners are co-trained in a cooperative, synergistic manner to solve the online learning problem. We demonstrate that this approach achieves state-of-the-art performance on standard online learning datasets. We make our code available: \url{https://github.com/AntonValk/MODL}
Antonios Valkanas, Boris N. Oreshkin, Mark Coates
AISTATS1
2025 SKOLR: Structured Koopman Operator Linear RNN for Time-Series Forecasting
abstract
Koopman operator theory provides a framework for nonlinear dynamical system analysis and time-series forecasting by mapping dynamics to a space of real-valued measurement functions, enabling a linear operator representation. Despite the advantage of linearity, the operator is generally infinite-dimensional. Therefore, the objective is to learn measurement functions that yield a tractable finite-dimensional Koopman operator approximation. In this work, we establish a connection between Koopman operator approximation and linear Recurrent Neural Networks (RNNs), which have recently demonstrated remarkable success in sequence modeling. We show that by considering an extended state consisting of lagged observations, we can establish an equivalence between a structured Koopman operator and linear RNN updates. Building on this connection, we present SKOLR, which integrates a learnable spectral decomposition of the input signal with a multilayer perceptron (MLP) as the measurement functions and implements a structured Koopman operator via a highly parallel linear RNN stack. Numerical experiments on various forecasting benchmarks and dynamical systems show that this streamlined, Koopman-theory-based design delivers exceptional performance. Our code is available at: https://github.com/networkslab/SKOLR.
Liheng Ma, Antonios Valkanas, Boris N. Oreshkin, Mark Coates
ICML3
2025 C3PO: Optimized Large Language Model Cascades with Probabilistic Cost Constraints for Reasoning
abstract
Large language models (LLMs) have achieved impressive results on complex reasoning tasks, but their high inference cost remains a major barrier to real-world deployment. A promising solution is to use cascaded inference, where small, cheap models handle easy queries, and only the hardest examples are escalated to more powerful models. However, existing cascade methods typically rely on supervised training with labeled data, offer no theoretical generalization guarantees, and provide limited control over test-time computational cost. We introduce **C3PO** (*Cost Controlled Cascaded Prediction Optimization*), a self-supervised framework for optimizing LLM cascades under probabilistic cost constraints. By focusing on minimizing regret with respect to the most powerful model (MPM), C3PO avoids the need for labeled data by constructing a cascade using only unlabeled model outputs. It leverages conformal prediction to bound the probability that inference cost exceeds a user-specified budget. We provide theoretical guarantees on both cost control and generalization error, and show that our optimization procedure is effective even with small calibration sets. Empirically, C3PO achieves state-of-the-art performance across a diverse set of reasoning benchmarks including GSM8K, MATH-500, BigBench-Hard and AIME, outperforming strong LLM cascading baselines in both accuracy and cost-efficiency. Our results demonstrate that principled, label-free cascade optimization can enable scalable LLM deployment.
Antonios Valkanas, Soumyasundar Pal, Pavel Rumiantsev, Yingxue Zhang 0001, Mark Coates
NeurIPS1
2024 Population Monte Carlo With Normalizing Flow
abstract
Adaptive importance sampling (AIS) methods provide a useful alternative to Markov Chain Monte Carlo (MCMC) algorithms for performing inference of intractable distributions. Population Monte Carlo (PMC) algorithms constitute a family of AIS approaches which adapt the proposal distributions iteratively to improve the approximation of the target distribution. Recent work in this area primarily focuses on ameliorating the proposal adaptation procedure for high-dimensional applications. However, most of the AIS algorithms use simple proposal distributions for sampling, which might be inadequate in exploring target distributions with intricate geometries. In this work, we construct expressive proposal distributions in the AIS framework using normalizing flow, an appealing approach for modeling complex distributions. We use an iterative parameter update rule to enhance the approximation of the target distribution. Numerical experiments show that in high-dimensional settings, the proposed algorithm offers significantly improved performance compared to the existing techniques.
Soumyasundar Pal, Antonios Valkanas, Mark Coates
IEEE Signal Process. Lett.2
2024 Motion In-Betweening via Deep $\Delta$Δ-Interpolator
abstract
We show that the task of synthesizing human motion conditioned on a set of key frames can be solved more accurately and effectively if a deep learning based interpolator operates in the delta mode using the spherical linear interpolator as a baseline. We empirically demonstrate the strength of our approach on publicly available datasets achieving state-of-the-art performance. We further generalize these results by showing that the ∆-regime is viable with respect to the reference of the last known frame (also known as the zero-velocity model). This supports the more general conclusion that operating in the reference frame local to input frames is more accurate and robust than in the global (world) reference frame advocated in previous work.
Boris N. Oreshkin, Antonios Valkanas, Félix G. Harvey, Louis-Simon Ménard, Florent Bocquelet, Mark Coates
IEEE Trans. Vis. Comput. Graph.2
2023 Structure Aware Incremental Learning with Personalized Imitation Weights for Recommender Systems
abstract
Recommender systems now consume large-scale data and play a significant role in improving user experience. Graph Neural Networks (GNNs) have emerged as one of the most effective recommender system models because they model the rich relational information. The ever-growing volume of data can make training GNNs prohibitively expensive. To address this, previous attempts propose to train the GNN models incrementally as new data blocks arrive. Feature and structure knowledge distillation techniques have been explored to allow the GNN model to train in a fast incremental fashion while alleviating the catastrophic forgetting problem. However, preserving the same amount of the historical information for all users is sub-optimal since it fails to take into account the dynamics of each user's change of preferences. For the users whose interests shift substantially, retaining too much of the old knowledge can overly constrain the model, preventing it from quickly adapting to the users’ novel interests. In contrast, for users who have static preferences, model performance can benefit greatly from preserving as much of the user's long-term preferences as possible. In this work, we propose a novel training strategy that adaptively learns personalized imitation weights for each user to balance the contribution from the recent data and the amount of knowledge to be distilled from previous time periods. We demonstrate the effectiveness of learning imitation weights via a comparison on five diverse datasets for three state-of-art structure distillation based recommender systems. The performance shows consistent improvement over competitive incremental learning techniques.
Yuening Wang, Yingxue Zhang 0001, Antonios Valkanas, Ruiming Tang, Chen Ma 0001, Jianye Hao, Mark Coates
AAAI3
2022 Bag Graph: Multiple Instance Learning Using Bayesian Graph Neural Networks
abstract
Multiple Instance Learning (MIL) is a weakly supervised learning problem where the aim is to assign labels to sets or bags of instances, as opposed to traditional supervised learning where each instance is assumed to be independent and identically distributed (IID) and is to be labeled individually. Recent work has shown promising results for neural network models in the MIL setting. Instead of focusing on each instance, these models are trained in an end-to-end fashion to learn effective bag-level representations by suitably combining permutation invariant pooling techniques with neural architectures. In this paper, we consider modelling the interactions between bags using a graph and employ Graph Neural Networks (GNNs) to facilitate end-to-end learning. Since a meaningful graph representing dependencies between bags is rarely available, we propose to use a Bayesian GNN framework that can generate a likely graph structure for scenarios where there is uncertainty in the graph or when no graph is available. Empirical results demonstrate the efficacy of the proposed technique for several MIL benchmark tasks and a distribution regression task.
Soumyasundar Pal, Antonios Valkanas, Florence Regol, Mark Coates
AAAI2
2007 Adaptive TDD Synchronisation for WIMAX Access Networks
abstract
In adaptive time division duplex (ATDD) wireless systems, severe co-channel interference conditions can occur if the movable downlink/uplink (UL) TDD boundary is not synchronised among all frames in base stations. To reduce interference outage and to improve a system's spectral efficiency, a new single frequency cell (SFC) network architecture is proposed, which allows for distributed boundary synchronisation (DBS) via inter-sector signalling. SFC-DBS dynamically synchronises TDD boundaries among neighbouring sectors for each frame, thus avoiding sector-to-sector interference, while preserving the ATDD radio resource assignment efficiency. Analysis shows that SFC-DBS achieves an additional 6–11 dB in the average UL signal-to-interference ratio, compared with existing channel assignment schemes, which corresponds to 25–50 % capacity gain subject to traffic asymmetry in different sectors. More importantly, the proposed SFC scheme does not incur any further cost in the frequency planning, whereas the DBS scheme requires only minor system modifications. Compared with interference cancellation via antenna arrays and beamforming, SFC-DBS achieves similar performance, albeit without the cost for complex radio transceivers and multiple antenna elements.
Konstantinos Ntagkounakis, Panagiotis I. Dallas, Bayan S. Sharif, Antonios Valkanas
IET Commun.4