Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Boris N. Oreshkin

dblp:33/1017 · DBLP profile ↗
← Back
22ranked-venue papers
11as first author
8since 2021 · last 2026
0000-0002-1869-2004ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Deep learning architectures and training · 26% Transfer learning and domain adaptation · 18% Segmentation and scene understanding · 11%
Computer graphics and multimedia
2 papers
Computer animation and physical simulation · 100%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 30 heaviest of 35, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Time series and sequential data › time series analysis
time series forecasting
1.122023
NHITS: Neural Hierarchical Interpolation for Time Series Forecasting · AAAI 2023
N-BEATS: Neural basis expansion analysis for interpretable time series forecasting · ICLR 2020
Machine learning › Deep learning architectures and training › recurrent neural network
linear recurrent neural network
0.912025
SKOLR: Structured Koopman Operator Linear RNN for Time-Series Forecasting · ICML 2025
Machine learning › Deep learning architectures and training
recurrent neural network
0.912025
SKOLR: Structured Koopman Operator Linear RNN for Time-Series Forecasting · ICML 2025
Data mining
time series analysis
0.912025
SKOLR: Structured Koopman Operator Linear RNN for Time-Series Forecasting · ICML 2025
Data mining › time series analysis
time series forecasting
0.912025
SKOLR: Structured Koopman Operator Linear RNN for Time-Series Forecasting · ICML 2025
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.832020
Adaptive Cross-Modal Few-shot Learning · NeurIPS 2019
TADAM: Task dependent adaptive metric for improved few-shot learning · NeurIPS 2018
Weakly Supervised Few-shot Object Segmentation using Co-Attention with Visual and Semantic Embeddings · IJCAI 2020
Computer animation and physical simulation › motion synthesis
human motion synthesis
0.812024
Motion In-Betweening via Deep $\Delta$Δ-Interpolator · IEEE Trans. Vis. Comput. Graph. 2024
Computer animation and physical simulation › motion synthesis › motion interpolation
motion in-betweening
0.812024
Motion In-Betweening via Deep $\Delta$Δ-Interpolator · IEEE Trans. Vis. Comput. Graph. 2024
Robotics › Motion planning and robot control › robot control
inverse kinematics
0.612022
ProtoRes: Proto-Residual Network for Pose Authoring via Learned Inverse Kinematics · ICLR 2022
Robotics › Motion planning and robot control › robot control › inverse kinematics
learned inverse kinematics
0.612022
ProtoRes: Proto-Residual Network for Pose Authoring via Learned Inverse Kinematics · ICLR 2022
Computer animation and physical simulation
character animation
0.612022
ProtoRes: Proto-Residual Network for Pose Authoring via Learned Inverse Kinematics · ICLR 2022
Machine learning › Graph learning
graph neural network
0.512021
FC-GAGA: Fully Connected Gated Graph Architecture for Spatio-Temporal Traffic Forecasting · AAAI 2021
Machine learning › Transfer learning and domain adaptation
meta-learning
0.512021
Meta-Learning Framework with Applications to Zero-Shot Time-Series Forecasting · AAAI 2021
Machine learning › Deep learning architectures and training › skip connections
residual connection
0.512021
Meta-Learning Framework with Applications to Zero-Shot Time-Series Forecasting · AAAI 2021
Machine learning › Graph learning › spatio-temporal graph learning
spatio-temporal graph forecasting
0.512021
FC-GAGA: Fully Connected Gated Graph Architecture for Spatio-Temporal Traffic Forecasting · AAAI 2021
Machine learning › Transfer learning and domain adaptation › zero-shot learning
zero-shot forecasting
0.512021
Meta-Learning Framework with Applications to Zero-Shot Time-Series Forecasting · AAAI 2021
Computer vision › Vision and language › cross-modal attention
co-attention
0.412020
Weakly Supervised Few-shot Object Segmentation using Co-Attention with Visual and Semantic Embeddings · IJCAI 2020
Machine learning › Trustworthy machine learning › interpretability › explainable AI
explainable prediction
0.412020
N-BEATS: Neural basis expansion analysis for interpretable time series forecasting · ICLR 2020
Computer vision › Vision and language › cross-modal alignment
visual-semantic embedding
0.412020
Weakly Supervised Few-shot Object Segmentation using Co-Attention with Visual and Semantic Embeddings · IJCAI 2020
Machine learning › Learning paradigms
continual learning
0.412019
AMP: Adaptive Masked Proxies for Few-Shot Segmentation · ICCV 2019
Machine learning › Transfer learning and domain adaptation › few-shot learning
cross-modal few-shot learning
0.412019
Adaptive Cross-Modal Few-shot Learning · NeurIPS 2019
Computer vision › Segmentation and scene understanding › semantic segmentation
few-shot segmentation
0.412019
AMP: Adaptive Masked Proxies for Few-Shot Segmentation · ICCV 2019
Computer vision › Segmentation and scene understanding
instance segmentation
0.412019
AMP: Adaptive Masked Proxies for Few-Shot Segmentation · ICCV 2019
Machine learning › Transfer learning and domain adaptation › meta-learning
metric-based meta-learning
0.412019
Adaptive Cross-Modal Few-shot Learning · NeurIPS 2019
Computer vision › Segmentation and scene understanding
semantic segmentation
0.412019
AMP: Adaptive Masked Proxies for Few-Shot Segmentation · ICCV 2019
Machine learning › Representation and self-supervised learning › representation learning
metric learning
0.312018
TADAM: Task dependent adaptive metric for improved few-shot learning · NeurIPS 2018
Machine learning › Probabilistic and Bayesian machine learning
dynamical system
0.312025
SKOLR: Structured Koopman Operator Linear RNN for Time-Series Forecasting · ICML 2025
Machine learning › Representation and self-supervised learning › dynamical system representation
koopman operator
0.312025
SKOLR: Structured Koopman Operator Linear RNN for Time-Series Forecasting · ICML 2025
Computer vision › Video understanding and tracking
video object segmentation
0.112020
Weakly Supervised Few-shot Object Segmentation using Co-Attention with Visual and Semantic Embeddings · IJCAI 2020
Machine learning › Representation and self-supervised learning › multimodal representation learning
multimodal representation alignment
0.112019
Adaptive Cross-Modal Few-shot Learning · NeurIPS 2019

Methods — techniques the papers use, named apart from their topics

spectral decomposition · 1.7koopman operator theory · 1.7MLP · 1.7spherical linear interpolation · 1.5delta-mode learning · 1.5residual network · 1.1neural network · 1.1neural hierarchical interpolation · 0.7multirate sampling · 0.7fully connected time-series forecasting · 0.5
YearPublicationVenuePosition
2026 TAT: Temporal-Aligned Transformer for Multi-Horizon Peak Demand Forecasting
abstract
Multi-horizon time series forecasting has many practical applications such as demand forecasting. Accurate demand prediction is critical to help make buying and inventory decisions for supply chain management of e-commerce and physical retailers, and such predictions are typically required for future horizons extending tens of weeks. This is especially challenging during high-stake sales events when demand peaks are particularly difficult to predict accurately. However, these events are important not only for managing supply chain operations but also for ensuring a seamless shopping experience for customers. To address this challenge, we propose Temporal-Aligned Transformer (TAT), a multi-horizon forecaster leveraging apriori-known context variables such as holiday and promotion events information for improving predictive performance. Our model consists of an encoder and decoder, both embedded with a novel Temporal Alignment Attention (TAA), designed to learn context-dependent alignment for peak demand forecasting. We conduct extensive empirical analysis on two large-scale proprietary datasets from a large e-commerce retailer. We demonstrate that TAT brings up to 30% accuracy improvement on peak demand forecasting while maintaining competitive overall performance compared to other state-of-the-art methods.
Zhiyuan Zhao 0002, Sitan Yang, Kin G. Olivares, Boris N. Oreshkin, Stan Vitebsky, Michael W. Mahoney, B. Aditya Prakash, Dmitry Efimov
ICDE4
2025 MODL: Multilearner Online Deep Learning
abstract
Online deep learning tackles the challenge of learning from data streams by balancing two competing goals: fast learning and deep learning. However, existing research primarily emphasizes deep learning solutions, which are more adept at handling the ”deep” aspect than the ”fast” aspect of online learning. In this work, we introduce an alternative paradigm through a hybrid multilearner approach. We begin by developing a fast online logistic regression learner, which operates without relying on backpropagation. It leverages closed-form recursive updates of model parameters, efficiently addressing the fast learning component of the online learning challenge. This approach is further integrated with a cascaded multilearner design, where shallow and deep learners are co-trained in a cooperative, synergistic manner to solve the online learning problem. We demonstrate that this approach achieves state-of-the-art performance on standard online learning datasets. We make our code available: \url{https://github.com/AntonValk/MODL}
Antonios Valkanas, Boris N. Oreshkin, Mark Coates
AISTATS2
2025 SKOLR: Structured Koopman Operator Linear RNN for Time-Series Forecasting
abstract
Koopman operator theory provides a framework for nonlinear dynamical system analysis and time-series forecasting by mapping dynamics to a space of real-valued measurement functions, enabling a linear operator representation. Despite the advantage of linearity, the operator is generally infinite-dimensional. Therefore, the objective is to learn measurement functions that yield a tractable finite-dimensional Koopman operator approximation. In this work, we establish a connection between Koopman operator approximation and linear Recurrent Neural Networks (RNNs), which have recently demonstrated remarkable success in sequence modeling. We show that by considering an extended state consisting of lagged observations, we can establish an equivalence between a structured Koopman operator and linear RNN updates. Building on this connection, we present SKOLR, which integrates a learnable spectral decomposition of the input signal with a multilayer perceptron (MLP) as the measurement functions and implements a structured Koopman operator via a highly parallel linear RNN stack. Numerical experiments on various forecasting benchmarks and dynamical systems show that this streamlined, Koopman-theory-based design delivers exceptional performance. Our code is available at: https://github.com/networkslab/SKOLR.
Liheng Ma, Antonios Valkanas, Boris N. Oreshkin, Mark Coates
ICML4
2024 Motion In-Betweening via Deep $\Delta$Δ-Interpolator
abstract
We show that the task of synthesizing human motion conditioned on a set of key frames can be solved more accurately and effectively if a deep learning based interpolator operates in the delta mode using the spherical linear interpolator as a baseline. We empirically demonstrate the strength of our approach on publicly available datasets achieving state-of-the-art performance. We further generalize these results by showing that the ∆-regime is viable with respect to the reference of the last known frame (also known as the zero-velocity model). This supports the more general conclusion that operating in the reference frame local to input frames is more accurate and robust than in the global (world) reference frame advocated in previous work.
Boris N. Oreshkin, Antonios Valkanas, Félix G. Harvey, Louis-Simon Ménard, Florent Bocquelet, Mark Coates
IEEE Trans. Vis. Comput. Graph.1
2023 NHITS: Neural Hierarchical Interpolation for Time Series Forecasting
abstract
Recent progress in neural forecasting accelerated improvements in the performance of large-scale forecasting systems. Yet, long-horizon forecasting remains a very difficult task. Two common challenges afflicting the task are the volatility of the predictions and their computational complexity. We introduce NHITS, a model which addresses both challenges by incorporating novel hierarchical interpolation and multi-rate data sampling techniques. These techniques enable the proposed method to assemble its predictions sequentially, emphasizing components with different frequencies and scales while decomposing the input signal and synthesizing the forecast. We prove that the hierarchical interpolation technique can efficiently approximate arbitrarily long horizons in the presence of smoothness. Additionally, we conduct extensive large-scale dataset experiments from the long-horizon forecasting literature, demonstrating the advantages of our method over the state-of-the-art methods, where NHITS provides an average accuracy improvement of almost 20% over the latest Transformer architectures while reducing the computation time by an order of magnitude (50 times). Our code is available at https://github.com/Nixtla/neuralforecast.
Cristian Challu, Kin G. Olivares, Boris N. Oreshkin, Federico Garza Ramírez, Max Mergenthaler Canseco, Artur Dubrawski
AAAI3
2022 ProtoRes: Proto-Residual Network for Pose Authoring via Learned Inverse Kinematics
Boris N. Oreshkin, Florent Bocquelet, Félix G. Harvey, Bay Raitt, Dominic Laflamme
ICLR1
2021 FC-GAGA: Fully Connected Gated Graph Architecture for Spatio-Temporal Traffic Forecasting
abstract
Forecasting of multivariate time-series is an important problem that has applications in traffic management, cellular network configuration, and quantitative finance. A special case of the problem arises when there is a graph available that captures the relationships between the time-series. In this paper we propose a novel learning architecture that achieves performance competitive with or better than the best existing algorithms, without requiring knowledge of the graph. The key element of our proposed architecture is the learnable fully connected hard graph gating mechanism that enables the use of the state-of-the-art and highly computationally efficient fully connected time-series forecasting architecture in traffic forecasting applications. Experimental results for two public traffic network datasets illustrate the value of our approach, and ablation studies confirm the importance of each element of the architecture. The code is available here: https://github.com/boreshkinai/fc-gaga.
Boris N. Oreshkin, Arezou Amini, Lucy Coyle, Mark Coates
AAAI1
2021 Meta-Learning Framework with Applications to Zero-Shot Time-Series Forecasting
abstract
Can meta-learning discover generic ways of processing time series (TS) from a diverse dataset so as to greatly improve generalization on new TS coming from different datasets? This work provides positive evidence to this using a broad meta-learning framework which we show subsumes many existing meta-learning algorithms. Our theoretical analysis suggests that residual connections act as a meta-learning adaptation mechanism, generating a subset of task-specific parameters based on a given TS input, thus gradually expanding the expressive power of the architecture on-the-fly. The same mechanism is shown via linearization analysis to have the interpretation of a sequential update of the final linear layer. Our empirical results on a wide range of data emphasize the importance of the identified meta-learning mechanisms for successful zero-shot univariate forecasting, suggesting that it is viable to train a neural network on a source TS dataset and deploy it on a different target TS dataset without retraining, resulting in performance that is at least as good as that of state-of-practice univariate forecasting models.
Boris N. Oreshkin, Dmitri Carpov, Nicolas Chapados, Yoshua Bengio
AAAI1
2020 N-BEATS: Neural basis expansion analysis for interpretable time series forecasting
Boris N. Oreshkin, Dmitri Carpov, Nicolas Chapados, Yoshua Bengio
ICLR1
2020 Weakly Supervised Few-shot Object Segmentation using Co-Attention with Visual and Semantic Embeddings
abstract
Significant progress has been made recently in developing few-shot object segmentation methods. Learning is shown to be successful in few-shot segmentation settings, using pixel-level, scribbles and bounding box supervision. This paper takes another approach, i.e., only requiring image-level label for few-shot object segmentation. We propose a novel multi-modal interaction module for few-shot object segmentation that utilizes a co-attention mechanism using both visual and word embedding. Our model using image-level labels achieves 4.8% improvement over previously proposed image-level few-shot object segmentation. It also outperforms state-of-the-art methods that use weak bounding box supervision on PASCAL-5^i. Our results show that few-shot segmentation benefits from utilizing word embeddings, and that we are able to perform few-shot segmentation using stacked joint visual semantic processing with weak image-level labels. We further propose a novel setup, Temporal Object Segmentation for Few-shot Learning (TOSFL) for videos. TOSFL can be used on a variety of public video data such as Youtube-VOS, as demonstrated in both instance-level and category-level TOSFL experiments.
Mennatullah Siam, Naren Doraiswamy, Boris N. Oreshkin, Hengshuai Yao, Martin Jägersand
IJCAI3
2019 AMP: Adaptive Masked Proxies for Few-Shot Segmentation
abstract
Deep learning has thrived by training on large-scale datasets. However, in robotics applications sample efficiency is critical. We propose a novel adaptive masked proxies method that constructs the final segmentation layer weights from few labelled samples. It utilizes multiresolution average pooling on base embeddings masked with the label to act as a positive proxy for the new class, while fusing it with the previously learned class signatures. Our method is evaluated on PASCAL-5idataset and outperforms the state-of-the-art in the few-shot semantic segmentation. Unlike previous methods, our approach does not require a second branch to estimate parameters or prototypes, which enables it to be used with 2-stream motion and appearance based segmentation networks. We further propose a novel setup for evaluating continual learning of object segmentation which we name incremental PASCAL (iPASCAL) where our method outperforms the baseline method. Our code is publicly available at https://github. com/MSiam/AdaptiveMaskedProxies.
Mennatullah Siam, Boris N. Oreshkin, Martin Jägersand
ICCV2
2019 Adaptive Cross-Modal Few-shot Learning
abstract
Metric-based meta-learning techniques have successfully been applied to few-shot classification problems. In this paper, we propose to leverage cross-modal information to enhance metric-based few-shot learning methods. Visual and semantic feature spaces have different structures by definition. For certain concepts, visual features might be richer and more discriminative than text ones. While for others, the inverse might be true. Moreover, when the support from visual information is limited in image classification, semantic representations (learned from unsupervised text corpora) can provide strong prior knowledge and context to help learning. Based on these two intuitions, we propose a mechanism that can adaptively combine information from both modalities according to new image categories to be learned. Through a series of experiments, we show that by this adaptive combination of the two modalities, our model outperforms current uni-modality few-shot learning methods and modality-alignment methods by a large margin on all benchmarks and few-shot scenarios tested. Experiments also show that our model can effectively adjust its focus on the two modalities. The improvement in performance is particularly large when the number of shots is very small.
Chen Xing, Negar Rostamzadeh, Boris N. Oreshkin, Pedro O. Pinheiro
NeurIPS3
2018 TADAM: Task dependent adaptive metric for improved few-shot learning
abstract
Few-shot learning has become essential for producing models that generalize from few examples. In this work, we identify that metric scaling and metric task conditioning are important to improve the performance of few-shot algorithms. Our analysis reveals that simple metric scaling completely changes the nature of few-shot algorithm parameter updates. Metric scaling provides improvements up to 14% in accuracy for certain metrics on the mini-Imagenet 5-way 5-shot classification task. We further propose a simple and effective way of conditioning a learner on the task sample set, resulting in learning a task-dependent metric space. Moreover, we propose and empirically test a practical end-to-end optimization procedure based on auxiliary task co-training to learn a task-dependent metric space. The resulting few-shot learning model based on the task-dependent scaled metric achieves state of the art on mini-Imagenet. We confirm these results on another few-shot dataset that we introduce in this paper based on CIFAR100.
Boris N. Oreshkin, Pau Rodríguez, Alexandre Lacoste
NeurIPS1
2013 Uncertainty Driven Probabilistic Voxel Selection for Image Registration
abstract
This paper presents a novel probabilistic voxel selection strategy for medical image registration in time-sensitive contexts, where the goal is aggressive voxel sampling (e.g., using less than 1% of the total number) while maintaining registration accuracy and low failure rate. We develop a Bayesian framework whereby, first, a voxel sampling probability field (VSPF) is built based on the uncertainty on the transformation parameters. We then describe a practical, multi-scale registration algorithm, where, at each optimization iteration, different voxel subsets are sampled based on the VSPF. The approach maximizes accuracy without committing to a particular fixed subset of voxels. The probabilistic sampling scheme developed is shown to manage the tradeoff between the robustness of traditional random voxel selection (by permitting more exploration) and the accuracy of fixed voxel selection (by permitting a greater proportion of informative voxels).
Boris N. Oreshkin, Tal Arbel
IEEE Trans. Medical Imaging1
2011 Spatial and probabilistic codebook template based head pose estimation from unconstrained environments
abstract
In unconstrained environments, head pose detection can be very challenging due to the joint and arbitrary occurrence of facial expressions, background clutter, partial occlusions and illumination conditions. Despite the wide range of head pose literature, most current methods can address this problem only up to a certain degree, and mostly for restricted scenarios. In this paper, we address the problem of head pose classification from real world images with large appearance variation. We represent each pose with a probabilistic and spatial template learned from facial codewords. The inference of the best template representing a test image is achieved probabilistically and spatially at the codebook. The experimental results are obtained from 5500 video frames collected under different illumination and background conditions. Our probabilistic framework is shown to outperform the current state-of-the-art in head pose classification.
Meltem Demirkus, Boris N. Oreshkin, James J. Clark, Tal Arbel
ICIP2
2010 Efficient delay-tolerant particle filtering through selective processing of out-of-sequence measurements
Boris N. Oreshkin, Mark Coates
FUSION2
2010 Asynchronous distributed particle filter via decentralized evaluation of Gaussian products
Boris N. Oreshkin, Mark Coates
FUSION1
2009 Multi-hop Greedy Gossip with Eavesdropping
Deniz Üstebay, Boris N. Oreshkin, Mark Coates, Michael G. Rabbat
FUSION2
2009 The speed of greed: Characterizing myopic gossip through network voracity
abstract
This paper analyzes the rate of convergence of greedy gossip with eavesdropping (GGE). In previous work, we proposed GGE, a fast gossip algorithm based on exploiting the broadcast nature of wireless communications rather than location information. Assuming all transmissions are wireless broadcasts, nodes can keep track of their neighbors' values by eavesdropping on their communications. Then, when it comes time to gossip, a node greedily and myopically gossips with the neighbor whose value is most different from its own, rather than with a randomly chosen neighbor. Previously, we have proved that GGE converges to the average consensus on connected network topologies and demonstrated that GGE outperforms standard randomized gossip (RG). In this paper we study the rate of convergence of GGE in terms of network voracity which is a topology-dependent constant analogous to the second-largest eigenvalue characterization for RG. Simulations demonstrate that the convergence rate of GGE is superior to existing average consensus algorithms such as geographic gossip.
Deniz Üstebay, Boris N. Oreshkin, Mark Coates, Michael G. Rabbat
ICASSP2
2008 Weak sense Lp error bounds for leader-node distributed particle filters
Boris N. Oreshkin, Mark Coates
FUSION1
2008 Distributed average consensus with increased convergence rate
abstract
The average consensus problem in the distributed signal processing context is addressed by linear iterative algorithms, with asymptotic convergence to the consensus. The convergence of the average consensus for an arbitrary weight matrix satisfying the convergence conditions is unfortunately slow restricting the use of the developed algorithms in applications. In this paper, we propose the use of linear extrapolation methods in order to accelerate distributed linear iterations. We provide analytical and simulation results that demonstrate the validity and effectiveness of the proposed scheme. Finally, we report simulation results showing that the generalized version of our algorithm, when a grid search for the unknown optimum value of mixing parameter is used, significantly outperforms the optimum consensus algorithm based on weight matrix optimization.
Boris N. Oreshkin, Tuncer C. Aysal, Mark Coates
ICASSP1
2008 Optimization of loading factor preventing target cancellation
abstract
Adaptive algorithms based on sample matrix inversion belong to an important class of algorithms used in radar target detection to overcome prior uncertainty of interference covariance. Sample matrix inversion problem is generally ill conditioned. Moreover, the contamination of the empirical covariance matrix by the useful signal leads to significant degradation of performance of this class of adaptive algorithms. Regularization, also known in radar literature as sample covariance loading, can be used to combat both ill conditioning of the original problem and contamination of the empirical covariance by the desired signal. However, the optimum value of loading factor cannot be derived unless strong assumptions are made regarding the structure of covariance matrix and useful signal penetration model. In this paper an iterative algorithm for loading factor optimization based on the maximization of empirical signal to interference plus noise ratio (SINR) is proposed. The proposed solution does not rely on any assumptions regarding the structure of empirical covariance matrix and signal penetration model. The paper also presents simulation examples showing the effectiveness of the proposed solution.
Boris N. Oreshkin, Peter A. Bakulev
ICASSP1