VLDB 2026 Research / reviewers in the wild / expert
Ross Goroshin
dblp:91/8200 · also Rostislav Goroshin
· DBLP profile ↗
12ranked-venue papers
2as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Deep learning architectures and training · 30% Video understanding and tracking · 23% Reinforcement learning · 12% | |
| Software engineering, system software, and programming languages
1 paper |
Empirical software engineering · 100% |
Topics — the 27 heaviest of 28, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking
feature tracking |
0.9 | 1 | 2025 | Tapnext: Tracking Any Point (Tap) as Next Token Prediction · ICCV 2025 |
Machine learning › Generative modeling › autoregressive model
next-token prediction |
0.9 | 1 | 2025 | Tapnext: Tracking Any Point (Tap) as Next Token Prediction · ICCV 2025 |
Computer vision › Video understanding and tracking › feature tracking
tracking any point |
0.9 | 1 | 2025 | Tapnext: Tracking Any Point (Tap) as Next Token Prediction · ICCV 2025 |
Computer vision › Video understanding and tracking › video representation learning
video tokenization |
0.9 | 1 | 2025 | Tapnext: Tracking Any Point (Tap) as Next Token Prediction · ICCV 2025 |
Machine learning › Deep learning architectures and training
autoencoder |
0.8 | 1 | 2024 | Course Correcting Koopman Representations · ICLR 2024 |
Robotics › Motion planning and robot control
dynamic modeling |
0.8 | 1 | 2024 | Course Correcting Koopman Representations · ICLR 2024 |
Machine learning › Deep learning architectures and training › convolutional neural network
convolutional neural network architecture |
0.7 | 2 | 2021 | Impact of Aliasing on Generalization in Deep Convolutional Networks · ICCV 2021 Efficient object localization using Convolutional Networks · CVPR 2015 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.7 | 1 | 2023 | Block-State Transformers · NeurIPS 2023 |
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
auxiliary tasks |
0.7 | 1 | 2023 | Proto-Value Networks: Scaling Representation Learning with Auxiliary Tasks · ICLR 2023 |
Machine learning › Reinforcement learning › function approximation
representation learning for reinforcement learning |
0.7 | 1 | 2023 | Proto-Value Networks: Scaling Representation Learning with Auxiliary Tasks · ICLR 2023 |
Machine learning › Deep learning architectures and training
sequence modeling |
0.7 | 1 | 2023 | Block-State Transformers · NeurIPS 2023 |
Machine learning › Deep learning architectures and training
state space model |
0.7 | 1 | 2023 | Block-State Transformers · NeurIPS 2023 |
Machine learning › Deep learning architectures and training
data augmentation |
0.5 | 1 | 2021 | Impact of Aliasing on Generalization in Deep Convolutional Networks · ICCV 2021 |
Machine learning › Trustworthy machine learning
out-of-distribution generalization |
0.5 | 1 | 2021 | Impact of Aliasing on Generalization in Deep Convolutional Networks · ICCV 2021 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
0.4 | 1 | 2020 | Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few Examples · ICLR 2020 |
Machine learning › Transfer learning and domain adaptation
meta-learning |
0.4 | 1 | 2020 | Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few Examples · ICLR 2020 |
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning |
0.4 | 2 | 2015 | Learning to Linearize Under Uncertainty · NIPS 2015 Unsupervised Learning of Spatiotemporally Coherent Metrics · ICCV 2015 |
Empirical software engineering › benchmarking
benchmark dataset |
0.4 | 1 | 2020 | Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few Examples · ICLR 2020 |
Robotics › Robot navigation and mapping
learning-based navigation |
0.3 | 1 | 2017 | Learning to Navigate in Complex Environments · ICLR (Poster) 2017 |
Machine learning › Reinforcement learning › reinforcement learning for control
navigation policy learning |
0.3 | 1 | 2017 | Learning to Navigate in Complex Environments · ICLR (Poster) 2017 |
Computer vision › Face, body and person analysis
human pose estimation |
0.2 | 1 | 2015 | Efficient object localization using Convolutional Networks · CVPR 2015 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
latent generative model |
0.2 | 1 | 2015 | Learning to Linearize Under Uncertainty · NIPS 2015 |
Machine learning › Representation and self-supervised learning › representation learning
metric learning |
0.2 | 1 | 2015 | Unsupervised Learning of Spatiotemporally Coherent Metrics · ICCV 2015 |
Computer vision › Video understanding and tracking › temporal modeling
temporal consistency |
0.2 | 1 | 2015 | Unsupervised Learning of Spatiotemporally Coherent Metrics · ICCV 2015 |
Computer vision › Video understanding and tracking
video prediction |
0.2 | 1 | 2015 | Learning to Linearize Under Uncertainty · NIPS 2015 |
Natural language and speech › Language models and text generation
language modeling |
0.2 | 1 | 2023 | Block-State Transformers · NeurIPS 2023 |
Computer vision › Image recognition and object detection
image classification |
0.1 | 1 | 2020 | Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few Examples · ICLR 2020 |
Methods — techniques the papers use, named apart from their topics
next-token prediction · 0.9koopman operator theory · 0.8autoencoder · 0.8value function · 0.7transformer · 0.7state space model · 0.7auxiliary tasks · 0.7structural architecture modification · 0.5low-pass filtering · 0.5frequency analysis · 0.5meta-learning · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Tapnext: Tracking Any Point (Tap) as Next Token Prediction
Artem Zholus, Carl Doersch, Yi Yang 0007, Skanda Koppula, Viorica Patraucean, Xu Owen He, Ignacio Rocco, Mehdi S. M. Sajjadi, Sarath Chandar, Ross Goroshin |
ICCV | 10 |
| 2024 | BootsTAP: Bootstrapped Training for Tracking-Any-Point
Carl Doersch, Pauline Luc, Yi Yang 0007, Dilara Gokay, Skanda Koppula, Ankush Gupta 0001, Joseph Heyward, Ignacio Rocco, Ross Goroshin, João Carreira 0001, Andrew Zisserman |
ACCV (2) | 9 |
| 2024 | Course Correcting Koopman RepresentationsabstractKoopman representations aim to learn features of nonlinear dynamical systems (NLDS) which lead to linear dynamics in the latent space. Theoretically, such features can be used to simplify many problems in modeling and control of NLDS. In this work we study autoencoder formulations of this problem, and different ways they can be used to model dynamics, specifically for future state prediction over long horizons. We discover several limitations of predicting future states in the latent space and propose an inference-time mechanism, which we refer to as Periodic Reencoding, for faithfully capturing long term dynamics. We justify this method both analytically and empirically via experiments in low and high dimensional NLDS. Mahan Fathi, Clement Gehring, Jonathan Pilault, David Kanaa, Pierre-Luc Bacon, Ross Goroshin |
ICLR | 6 |
| 2023 | Proto-Value Networks: Scaling Representation Learning with Auxiliary Tasks
Jesse Farebrother, Joshua Greaves, Rishabh Agarwal, Charline Le Lan, Ross Goroshin, Pablo Samuel Castro, Marc G. Bellemare |
ICLR | 5 |
| 2023 | Block-State TransformersabstractState space models (SSMs) have shown impressive results on tasks that require modeling long-range dependencies and efficiently scale to long sequences owing to their subquadratic runtime complexity.
Originally designed for continuous signals, SSMs have shown superior performance on a plethora of tasks, in vision and audio; however, SSMs still lag Transformer performance in Language Modeling tasks.
In this work, we propose a hybrid layer named Block-State Transformer (*BST*), that internally combines an SSM sublayer for long-range contextualization, and a Block Transformer sublayer for short-term representation of sequences.
We study three different, and completely *parallelizable*, variants that integrate SSMs and block-wise attention.
We show that our model outperforms similar Transformer-based architectures on language modeling perplexity and generalizes to longer sequences.
In addition, the Block-State Transformer demonstrates a more than *tenfold* increase in speed at the layer level compared to the Block-Recurrent Transformer when model parallelization is employed. Jonathan Pilault, Mahan Fathi, Orhan Firat, Christopher Joseph Pal, Pierre-Luc Bacon, Ross Goroshin |
NeurIPS | 6 |
| 2021 | Impact of Aliasing on Generalization in Deep Convolutional NetworksabstractWe investigate the impact of aliasing on generalization in Deep Convolutional Networks and show that data augmentation schemes alone are unable to prevent it due to structural limitations in widely used architectures. Drawing insights from frequency analysis theory, we take a closer look at ResNet and EfficientNet architectures and review the trade-off between aliasing and information loss in each of their major components. We show how to mitigate aliasing by inserting non-trainable low-pass filters at key locations, particularly where networks lack the capacity to learn them. These simple architectural changes lead to substantial improvements in generalization on i.i.d. and even more on out-of-distribution conditions, such as image classification under natural corruptions on ImageNet-C [11] and few-shot learning on Meta-Dataset [26]. State-of-the art results are achieved on both datasets without introducing additional trainable parameters and using the default hyper-parameters of open source codebases. Cristina Nader Vasconcelos, Hugo Larochelle, Vincent Dumoulin, Rob Romijnders, Nicolas Le Roux, Ross Goroshin |
ICCV | 6 |
| 2020 | Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few Examples
Eleni Triantafillou, Tyler Zhu, Vincent Dumoulin, Pascal Lamblin, Utku Evci, Kelvin Xu, Ross Goroshin, Carles Gelada, Kevin Swersky, Pierre-Antoine Manzagol, Hugo Larochelle |
ICLR | 7 |
| 2017 | Learning to Navigate in Complex Environments
Piotr Mirowski, Razvan Pascanu, Fabio Viola, Hubert Soyer, Andrew J. Ballard, Andrea Banino, Misha Denil, Ross Goroshin, Laurent Sifre, Koray Kavukcuoglu, Dharshan Kumaran, Raia Hadsell |
ICLR (Poster) | 8 |
| 2015 | Efficient object localization using Convolutional NetworksabstractRecent state-of-the-art performance on human-body pose estimation has been achieved with Deep Convolutional Networks (ConvNets). Traditional ConvNet architectures include pooling and sub-sampling layers which reduce computational requirements, introduce invariance and prevent over-training. These benefits of pooling come at the cost of reduced localization accuracy. We introduce a novel architecture which includes an efficient ‘position refinement’ model that is trained to estimate the joint offset location within a small region of the image. This refinement model is jointly trained in cascade with a state-of-the-art ConvNet model [21] to achieve improved accuracy in human joint location estimation. We show that the variance of our detector approaches the variance of human annotations on the FLIC [20] dataset and outperforms all existing approaches on the MPII-human-pose dataset [1]. Jonathan Tompson, Ross Goroshin, Arjun Jain, Yann LeCun, Christoph Bregler |
CVPR | 2 |
| 2015 | Unsupervised Learning of Spatiotemporally Coherent MetricsabstractCurrent state-of-the-art classification and detection algorithms train deep convolutional networks using labeled data. In this work we study unsupervised feature learning with convolutional networks in the context of temporally coherent unlabeled data. We focus on feature learning from unlabeled video data, using the assumption that adjacent video frames contain semantically similar information. This assumption is exploited to train a convolutional pooling auto-encoder regularized by slowness and sparsity priors. We establish a connection between slow feature learning and metric learning. Using this connection we define "temporal coherence" -- a criterion which can be used to set hyper-parameters in a principled and automated manner. In a transfer learning experiment, we show that the resulting encoder can be used to define a more semantically coherent metric without the use of labels. Ross Goroshin, Joan Bruna, Jonathan Tompson, David Eigen, Yann LeCun |
ICCV | 1 |
| 2015 | Learning to Linearize Under UncertaintyabstractTraining deep feature hierarchies to solve supervised learning tasks has achieving state of the art performance on many problems in computer vision. However, a principled way in which to train such hierarchies in the unsupervised setting has remained elusive. In this work we suggest a new architecture and loss for training deep feature hierarchies that linearize the transformations observed in unlabelednatural video sequences. This is done by training a generative model to predict video frames. We also address the problem of inherent uncertainty in prediction by introducing a latent variables that are non-deterministic functions of the input into the network architecture. Ross Goroshin, Michaël Mathieu, Yann LeCun |
NIPS | 1 |
| 2009 | Automatic Cable Detection in Sonar ImageryabstractThe classical paradigm of line and curve detection in images, as prescribed by the Hough transform, breaks down in cluttered and noisy imagery. In this paper we present an "upgraded" and ultimately more robust approach to line detection in images. The classical approach to line detection in imagery is low-pass filtering, followed by edge detection, followed by the application of the Hough transform. Peaks in the Hough transform correspond to straight line segments in the image. In our approach we replace low pass filtering by anisotropic diffusion; we replace edge detection by phase analysis of frequency components; and finally, lines corresponding to peaks in the Hough transform are statistically analyzed to reveal the most prominent and likely line segments (especially if the line thickness is known a priori) in the context of sampling distributions. The technique is demonstrated on real and synthetic aperture sonar (SAS) imagery. Jason C. Isaacs, Ross Goroshin |
SMC | 2 |