Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ross Goroshin

dblp:91/8200 · also Rostislav Goroshin · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Deep learning architectures and training · 30% Video understanding and tracking · 23% Reinforcement learning · 12%
Software engineering, system software, and programming languages
1 paper
Empirical software engineering · 100%

Topics — the 27 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
feature tracking
0.912025
Tapnext: Tracking Any Point (Tap) as Next Token Prediction · ICCV 2025
Machine learning › Generative modeling › autoregressive model
next-token prediction
0.912025
Tapnext: Tracking Any Point (Tap) as Next Token Prediction · ICCV 2025
Computer vision › Video understanding and tracking › feature tracking
tracking any point
0.912025
Tapnext: Tracking Any Point (Tap) as Next Token Prediction · ICCV 2025
Computer vision › Video understanding and tracking › video representation learning
video tokenization
0.912025
Tapnext: Tracking Any Point (Tap) as Next Token Prediction · ICCV 2025
Machine learning › Deep learning architectures and training
autoencoder
0.812024
Course Correcting Koopman Representations · ICLR 2024
Robotics › Motion planning and robot control
dynamic modeling
0.812024
Course Correcting Koopman Representations · ICLR 2024
Machine learning › Deep learning architectures and training › convolutional neural network
convolutional neural network architecture
0.722021
Impact of Aliasing on Generalization in Deep Convolutional Networks · ICCV 2021
Efficient object localization using Convolutional Networks · CVPR 2015
Machine learning › Deep learning architectures and training
attention mechanism
0.712023
Block-State Transformers · NeurIPS 2023
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
auxiliary tasks
0.712023
Proto-Value Networks: Scaling Representation Learning with Auxiliary Tasks · ICLR 2023
Machine learning › Reinforcement learning › function approximation
representation learning for reinforcement learning
0.712023
Proto-Value Networks: Scaling Representation Learning with Auxiliary Tasks · ICLR 2023
Machine learning › Deep learning architectures and training
sequence modeling
0.712023
Block-State Transformers · NeurIPS 2023
Machine learning › Deep learning architectures and training
state space model
0.712023
Block-State Transformers · NeurIPS 2023
Machine learning › Deep learning architectures and training
data augmentation
0.512021
Impact of Aliasing on Generalization in Deep Convolutional Networks · ICCV 2021
Machine learning › Trustworthy machine learning
out-of-distribution generalization
0.512021
Impact of Aliasing on Generalization in Deep Convolutional Networks · ICCV 2021
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.412020
Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few Examples · ICLR 2020
Machine learning › Transfer learning and domain adaptation
meta-learning
0.412020
Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few Examples · ICLR 2020
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning
0.422015
Learning to Linearize Under Uncertainty · NIPS 2015
Unsupervised Learning of Spatiotemporally Coherent Metrics · ICCV 2015
Empirical software engineering › benchmarking
benchmark dataset
0.412020
Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few Examples · ICLR 2020
Robotics › Robot navigation and mapping
learning-based navigation
0.312017
Learning to Navigate in Complex Environments · ICLR (Poster) 2017
Machine learning › Reinforcement learning › reinforcement learning for control
navigation policy learning
0.312017
Learning to Navigate in Complex Environments · ICLR (Poster) 2017
Computer vision › Face, body and person analysis
human pose estimation
0.212015
Efficient object localization using Convolutional Networks · CVPR 2015
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
latent generative model
0.212015
Learning to Linearize Under Uncertainty · NIPS 2015
Machine learning › Representation and self-supervised learning › representation learning
metric learning
0.212015
Unsupervised Learning of Spatiotemporally Coherent Metrics · ICCV 2015
Computer vision › Video understanding and tracking › temporal modeling
temporal consistency
0.212015
Unsupervised Learning of Spatiotemporally Coherent Metrics · ICCV 2015
Computer vision › Video understanding and tracking
video prediction
0.212015
Learning to Linearize Under Uncertainty · NIPS 2015
Natural language and speech › Language models and text generation
language modeling
0.212023
Block-State Transformers · NeurIPS 2023
Computer vision › Image recognition and object detection
image classification
0.112020
Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few Examples · ICLR 2020

Methods — techniques the papers use, named apart from their topics

next-token prediction · 0.9koopman operator theory · 0.8autoencoder · 0.8value function · 0.7transformer · 0.7state space model · 0.7auxiliary tasks · 0.7structural architecture modification · 0.5low-pass filtering · 0.5frequency analysis · 0.5meta-learning · 0.4
YearPublicationVenuePosition
2025 Tapnext: Tracking Any Point (Tap) as Next Token Prediction
Artem Zholus, Carl Doersch, Yi Yang 0007, Skanda Koppula, Viorica Patraucean, Xu Owen He, Ignacio Rocco, Mehdi S. M. Sajjadi, Sarath Chandar, Ross Goroshin
ICCV10
2024 BootsTAP: Bootstrapped Training for Tracking-Any-Point
Carl Doersch, Pauline Luc, Yi Yang 0007, Dilara Gokay, Skanda Koppula, Ankush Gupta 0001, Joseph Heyward, Ignacio Rocco, Ross Goroshin, João Carreira 0001, Andrew Zisserman
ACCV (2)9
2024 Course Correcting Koopman Representations
abstract
Koopman representations aim to learn features of nonlinear dynamical systems (NLDS) which lead to linear dynamics in the latent space. Theoretically, such features can be used to simplify many problems in modeling and control of NLDS. In this work we study autoencoder formulations of this problem, and different ways they can be used to model dynamics, specifically for future state prediction over long horizons. We discover several limitations of predicting future states in the latent space and propose an inference-time mechanism, which we refer to as Periodic Reencoding, for faithfully capturing long term dynamics. We justify this method both analytically and empirically via experiments in low and high dimensional NLDS.
Mahan Fathi, Clement Gehring, Jonathan Pilault, David Kanaa, Pierre-Luc Bacon, Ross Goroshin
ICLR6
2023 Proto-Value Networks: Scaling Representation Learning with Auxiliary Tasks
Jesse Farebrother, Joshua Greaves, Rishabh Agarwal, Charline Le Lan, Ross Goroshin, Pablo Samuel Castro, Marc G. Bellemare
ICLR5
2023 Block-State Transformers
abstract
State space models (SSMs) have shown impressive results on tasks that require modeling long-range dependencies and efficiently scale to long sequences owing to their subquadratic runtime complexity. Originally designed for continuous signals, SSMs have shown superior performance on a plethora of tasks, in vision and audio; however, SSMs still lag Transformer performance in Language Modeling tasks. In this work, we propose a hybrid layer named Block-State Transformer (*BST*), that internally combines an SSM sublayer for long-range contextualization, and a Block Transformer sublayer for short-term representation of sequences. We study three different, and completely *parallelizable*, variants that integrate SSMs and block-wise attention. We show that our model outperforms similar Transformer-based architectures on language modeling perplexity and generalizes to longer sequences. In addition, the Block-State Transformer demonstrates a more than *tenfold* increase in speed at the layer level compared to the Block-Recurrent Transformer when model parallelization is employed.
Jonathan Pilault, Mahan Fathi, Orhan Firat, Christopher Joseph Pal, Pierre-Luc Bacon, Ross Goroshin
NeurIPS6
2021 Impact of Aliasing on Generalization in Deep Convolutional Networks
abstract
We investigate the impact of aliasing on generalization in Deep Convolutional Networks and show that data augmentation schemes alone are unable to prevent it due to structural limitations in widely used architectures. Drawing insights from frequency analysis theory, we take a closer look at ResNet and EfficientNet architectures and review the trade-off between aliasing and information loss in each of their major components. We show how to mitigate aliasing by inserting non-trainable low-pass filters at key locations, particularly where networks lack the capacity to learn them. These simple architectural changes lead to substantial improvements in generalization on i.i.d. and even more on out-of-distribution conditions, such as image classification under natural corruptions on ImageNet-C [11] and few-shot learning on Meta-Dataset [26]. State-of-the art results are achieved on both datasets without introducing additional trainable parameters and using the default hyper-parameters of open source codebases.
Cristina Nader Vasconcelos, Hugo Larochelle, Vincent Dumoulin, Rob Romijnders, Nicolas Le Roux, Ross Goroshin
ICCV6
2020 Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few Examples
Eleni Triantafillou, Tyler Zhu, Vincent Dumoulin, Pascal Lamblin, Utku Evci, Kelvin Xu, Ross Goroshin, Carles Gelada, Kevin Swersky, Pierre-Antoine Manzagol, Hugo Larochelle
ICLR7
2017 Learning to Navigate in Complex Environments
Piotr Mirowski, Razvan Pascanu, Fabio Viola, Hubert Soyer, Andrew J. Ballard, Andrea Banino, Misha Denil, Ross Goroshin, Laurent Sifre, Koray Kavukcuoglu, Dharshan Kumaran, Raia Hadsell
ICLR (Poster)8
2015 Efficient object localization using Convolutional Networks
abstract
Recent state-of-the-art performance on human-body pose estimation has been achieved with Deep Convolutional Networks (ConvNets). Traditional ConvNet architectures include pooling and sub-sampling layers which reduce computational requirements, introduce invariance and prevent over-training. These benefits of pooling come at the cost of reduced localization accuracy. We introduce a novel architecture which includes an efficient ‘position refinement’ model that is trained to estimate the joint offset location within a small region of the image. This refinement model is jointly trained in cascade with a state-of-the-art ConvNet model [21] to achieve improved accuracy in human joint location estimation. We show that the variance of our detector approaches the variance of human annotations on the FLIC [20] dataset and outperforms all existing approaches on the MPII-human-pose dataset [1].
Jonathan Tompson, Ross Goroshin, Arjun Jain, Yann LeCun, Christoph Bregler
CVPR2
2015 Unsupervised Learning of Spatiotemporally Coherent Metrics
abstract
Current state-of-the-art classification and detection algorithms train deep convolutional networks using labeled data. In this work we study unsupervised feature learning with convolutional networks in the context of temporally coherent unlabeled data. We focus on feature learning from unlabeled video data, using the assumption that adjacent video frames contain semantically similar information. This assumption is exploited to train a convolutional pooling auto-encoder regularized by slowness and sparsity priors. We establish a connection between slow feature learning and metric learning. Using this connection we define "temporal coherence" -- a criterion which can be used to set hyper-parameters in a principled and automated manner. In a transfer learning experiment, we show that the resulting encoder can be used to define a more semantically coherent metric without the use of labels.
Ross Goroshin, Joan Bruna, Jonathan Tompson, David Eigen, Yann LeCun
ICCV1
2015 Learning to Linearize Under Uncertainty
abstract
Training deep feature hierarchies to solve supervised learning tasks has achieving state of the art performance on many problems in computer vision. However, a principled way in which to train such hierarchies in the unsupervised setting has remained elusive. In this work we suggest a new architecture and loss for training deep feature hierarchies that linearize the transformations observed in unlabelednatural video sequences. This is done by training a generative model to predict video frames. We also address the problem of inherent uncertainty in prediction by introducing a latent variables that are non-deterministic functions of the input into the network architecture.
Ross Goroshin, Michaël Mathieu, Yann LeCun
NIPS1
2009 Automatic Cable Detection in Sonar Imagery
abstract
The classical paradigm of line and curve detection in images, as prescribed by the Hough transform, breaks down in cluttered and noisy imagery. In this paper we present an "upgraded" and ultimately more robust approach to line detection in images. The classical approach to line detection in imagery is low-pass filtering, followed by edge detection, followed by the application of the Hough transform. Peaks in the Hough transform correspond to straight line segments in the image. In our approach we replace low pass filtering by anisotropic diffusion; we replace edge detection by phase analysis of frequency components; and finally, lines corresponding to peaks in the Hough transform are statistically analyzed to reveal the most prominent and likely line segments (especially if the line thickness is known a priori) in the context of sampling distributions. The technique is demonstrated on real and synthetic aperture sonar (SAS) imagery.
Jason C. Isaacs, Ross Goroshin
SMC2