Drew Linsley

dblp:194/2308 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0002-9722-7839ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 7 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Deep learning architectures and training · 26% Trustworthy machine learning · 23% Image recognition and object detection · 13%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 18 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
recurrent neural network
1.742021
Tracking Without Re-recognition in Humans and Machines · NeurIPS 2021
Stable and expressive recurrent vision models · NeurIPS 2020
Recurrent neural circuits for contour detection · ICLR 2020
Computer vision › Video understanding and tracking
object tracking
1.422025
Tracking objects that change in appearance with phase synchrony · ICLR 2025
Tracking Without Re-recognition in Humans and Machines · NeurIPS 2021
Machine learning › Trustworthy machine learning
interpretability
1.222023
Unlocking Feature Visualization for Deep Network with MAgnitude Constrained Optimization · NeurIPS 2023
Harmonizing the object recognition strategies of deep neural networks with humans · NeurIPS 2022
Machine learning › Trustworthy machine learning › interpretability
attribution methods
0.712023
Unlocking Feature Visualization for Deep Network with MAgnitude Constrained Optimization · NeurIPS 2023
Computer vision › 3D vision
biological vision modeling
0.712023
Performance-optimized deep neural networks are evolving into worse models of inferotemporal visual cortex · NeurIPS 2023
Machine learning › Trustworthy machine learning › interpretability › visual explanation
feature visualization
0.712023
Unlocking Feature Visualization for Deep Network with MAgnitude Constrained Optimization · NeurIPS 2023
Machine learning › Representation and self-supervised learning › representation matching › feature alignment
neural representation alignment
0.712023
Performance-optimized deep neural networks are evolving into worse models of inferotemporal visual cortex · NeurIPS 2023
Computer vision › Image recognition and object detection › object detection
contour detection
0.412020
Recurrent neural circuits for contour detection · ICLR 2020
Machine learning › Efficient and distributed learning
memory-efficient training
0.412020
Stable and expressive recurrent vision models · NeurIPS 2020
Computer vision › Segmentation and scene understanding
perceptual grouping
0.412020
Disentangling neural mechanisms for perceptual grouping · ICLR 2020
Machine learning › Deep learning architectures and training
attention mechanism
0.412019
Learning what and where to attend · ICLR (Poster) 2019
Machine learning › Deep learning architectures and training › recurrent neural network › gated recurrent network
gated recurrent unit
0.312018
Learning long-range spatial dependencies with horizontal gated recurrent units · NeurIPS 2018
Machine learning › Deep learning architectures and training › recurrent neural network
complex-valued recurrent neural network
0.312025
Tracking objects that change in appearance with phase synchrony · ICLR 2025
Machine learning › Trustworthy machine learning
robustness
0.312025
The 3D-PC: a benchmark for visual perspective taking in humans and machines · ICLR 2025
Computer vision › Image recognition and object detection
object recognition
0.212023
Performance-optimized deep neural networks are evolving into worse models of inferotemporal visual cortex · NeurIPS 2023
Computer vision › Video understanding and tracking › object tracking
transformer-based tracking
0.112021
Tracking Without Re-recognition in Humans and Machines · NeurIPS 2021
Computer vision › Segmentation and scene understanding
panoptic segmentation
0.112020
Stable and expressive recurrent vision models · NeurIPS 2020
Image and video processing › image segmentation
contour detection
0.112020
Recurrent neural circuits for contour detection · ICLR 2020

Methods — techniques the papers use, named apart from their topics

text prompting · 0.9phase synchrony · 0.9linear probing · 0.9fine-tuning · 0.9deep neural network · 0.9complex-valued recurrent neural network · 0.9phase spectrum optimization · 0.7magnitude-constrained optimization · 0.7representational alignment analysis · 0.6neural harmonizer · 0.6recurrent neural network · 0.4
YearPublicationVenuePosition
2025 The 3D-PC: a benchmark for visual perspective taking in humans and machines
abstract
Visual perspective taking (VPT) is the ability to perceive and reason about the perspectives of others. It is an essential feature of human intelligence, which develops over the first decade of life and requires an ability to process the 3D structure of visual scenes. A growing number of reports have indicated that deep neural networks (DNNs) become capable of analyzing 3D scenes after training on large image datasets. We investigated if this emergent ability for 3D analysis in DNNs is sufficient for VPT with the 3D perception challenge (3D-PC): a novel benchmark for 3D perception in humans and DNNs. The 3D-PC is comprised of three 3D-analysis tasks posed within natural scene images: (i.) a simple test of object depth order, (ii.) a basic VPT task (VPT-basic), and (iii.) a more challenging version of VPT (VPT-perturb) designed to limit the effectiveness of "shortcut" visual strategies. We tested human participants (N=33) and linearly probed or text-prompted over 300 DNNs on the challenge and found that nearly all of the DNNs approached or exceeded human accuracy in analyzing object depth order. Surprisingly, DNN accuracy on this task correlated with their object recognition performance. In contrast, there was an extraordinary gap between DNNs and humans on VPT-basic. Humans were nearly perfect, whereas most DNNs were near chance. Fine-tuning DNNs on VPT-basic brought them close to human performance, but they, unlike humans, dropped back to chance when tested on VPT-perturb. Our challenge demonstrates that the training routines and architectures of today's DNNs are well-suited for learning basic 3D properties of scenes and objects but are ill-suited for reasoning about these properties like humans do. We release our 3D-PC datasets and code to help bridge this gap in 3D perception between humans and machines.
Drew Linsley, Peisen Zhou, Alekh Karkada Ashok, Akash Nagaraj, Gaurav Gaonkar, Francis E. Lewis, Zygmunt Pizlo, Thomas Serre
ICLR1
2025 Tracking objects that change in appearance with phase synchrony
abstract
Objects we encounter often change appearance as we interact with them. Changes in illumination (shadows), object pose, or the movement of non-rigid objects can drastically alter available image features. How do biological visual systems track objects as they change? One plausible mechanism involves attentional mechanisms for reasoning about the locations of objects independently of their appearances --- a capability that prominent neuroscience theories have associated with computing through neural synchrony. Here, we describe a novel deep learning circuit that can learn to precisely control attention to features separately from their location in the world through neural synchrony: the complex-valued recurrent neural network (CV-RNN). Next, we compare object tracking in humans, the CV-RNN, and other deep neural networks (DNNs), using FeatureTracker: a large-scale challenge that asks observers to track objects as their locations and appearances change in precisely controlled ways. While humans effortlessly solved FeatureTracker, state-of-the-art DNNs did not. In contrast, our CV-RNN behaved similarly to humans on the challenge, providing a computational proof-of-concept for the role of phase synchronization as a neural substrate for tracking appearance-morphing objects as they move about.
Sabine Muzellec, Drew Linsley, Alekh Karkada Ashok, Ennio Mingolla, Girik Malik, Rufin VanRullen, Thomas Serre
ICLR2
2023 Unlocking Feature Visualization for Deep Network with MAgnitude Constrained Optimization
abstract
Feature visualization has gained significant popularity as an explainability method, particularly after the influential work by Olah et al. in 2017. Despite its success, its widespread adoption has been limited due to issues in scaling to deeper neural networks and the reliance on tricks to generate interpretable images. Here, we describe MACO, a simple approach to address these shortcomings. It consists in optimizing solely an image's phase spectrum while keeping its magnitude constant to ensure that the generated explanations lie in the space of natural images. Our approach yields significantly better results -- both qualitatively and quantitatively -- unlocking efficient and interpretable feature visualizations for state-of-the-art neural networks. We also show that our approach exhibits an attribution mechanism allowing to augment feature visualizations with spatial importance. Furthermore, we enable quantitative evaluation of feature visualizations by introducing 3 metrics: transferability, plausibility, and alignment with natural images. We validate our method on various applications and we introduce a website featuring MACO visualizations for all classes of the ImageNet dataset, which will be made available upon acceptance. Overall, our study unlocks feature visualizations for the largest, state-of-the-art classification networks without resorting to any parametric prior image model, effectively advancing a field that has been stagnating since 2017 (Olah et al, 2017).
Thomas Fel, Thibaut Boissin, Victor Boutin, Agustin M. Picard, Paul Novello, Julien Colin, Drew Linsley, Tom Rousseau, Rémi Cadène, Lore Goetschalckx, Laurent Gardes, Thomas Serre
NeurIPS7
2023 Performance-optimized deep neural networks are evolving into worse models of inferotemporal visual cortex
abstract
One of the most impactful findings in computational neuroscience over the past decade is that the object recognition accuracy of deep neural networks (DNNs) correlates with their ability to predict neural responses to natural images in the inferotemporal (IT) cortex. This discovery supported the long-held theory that object recognition is a core objective of the visual cortex, and suggested that more accurate DNNs would serve as better models of IT neuron responses to images. Since then, deep learning has undergone a revolution of scale: billion parameter-scale DNNs trained on billions of images are rivaling or outperforming humans at visual tasks including object recognition. Have today's DNNs become more accurate at predicting IT neuron responses to images as they have grown more accurate at object recognition? Surprisingly, across three independent experiments, we find that this is not the case. DNNs have become progressively worse models of IT as their accuracy has increased on ImageNet. To understand why DNNs experience this trade-off and evaluate if they are still an appropriate paradigm for modeling the visual system, we turn to recordings of IT that capture spatially resolved maps of neuronal activity elicited by natural images. These neuronal activity maps reveal that DNNs trained on ImageNet learn to rely on different visual features than those encoded by IT and that this problem worsens as their accuracy increases. We successfully resolved this issue with the neural harmonizer, a plug-and-play training routine for DNNs that aligns their learned representations with humans. Our results suggest that harmonized DNNs break the trade-off between ImageNet accuracy and neural prediction accuracy that assails current DNNs and offer a path to more accurate models of biological vision. Our work indicates that the standard approach for modeling IT with task-optimized DNNs needs revision, and other biological constraints, including human psychophysics data, are needed to accurately reverse-engineer the visual cortex.
Drew Linsley, Ivan F. Rodriguez Rodriguez, Thomas Fel, Michael Arcaro, Saloni Sharma, Margaret S. Livingstone, Thomas Serre
NeurIPS1
2022 Harmonizing the object recognition strategies of deep neural networks with humans
abstract
The many successes of deep neural networks (DNNs) over the past decade have largely been driven by computational scale rather than insights from biological intelligence. Here, we explore if these trends have also carried concomitant improvements in explaining the visual strategies humans rely on for object recognition. We do this by comparing two related but distinct properties of visual strategies in humans and DNNs: where they believe important visual features are in images and how they use those features to categorize objects. Across 84 different DNNs trained on ImageNet and three independent datasets measuring the where and the how of human visual strategies for object recognition on those images, we find a systematic trade-off between DNN categorization accuracy and alignment with human visual strategies for object recognition. \textit{State-of-the-art DNNs are progressively becoming less aligned with humans as their accuracy improves}. We rectify this growing issue with our neural harmonizer: a general-purpose training routine that both aligns DNN and human visual strategies and improves categorization accuracy. Our work represents the first demonstration that the scaling laws that are guiding the design of DNNs today have also produced worse models of human vision. We release our code and data at https://serre-lab.github.io/Harmonization to help the field build more human-like DNNs.
Thomas Fel, Ivan F. Rodriguez Rodriguez, Drew Linsley, Thomas Serre
NeurIPS3
2022 How and What to Learn: Taxonomizing Self-Supervised Learning for 3D Action Recognition
abstract
There are two competing standards for self-supervised learning in action recognition from 3D skeletons. Su et al., 2020 [31] used an auto-encoder architecture and an image reconstruction objective function to achieve state-of-the-art performance on the NTU60 C-View benchmark. Rao et al., 2020 [23] used Contrastive learning in the latent space to achieve state-of-the-art performance on the NTU60 C-Sub benchmark. Here, we reconcile these disparate approaches by developing a taxonomy of self-supervised learning for action recognition. We observe that leading approaches generally use one of two types of objective functions: those that seek to reconstruct the input from a latent representation ("Attractive" learning) versus those that also try to maximize the representations distinctiveness ("Contrastive" learning). Independently, leading approaches also differ in how they implement these objective functions: there are those that optimize representations in the decoder output space and those which optimize representations in the network’s latent space (encoder output). We find that combining these approaches leads to larger gains in performance and tolerance to transformation than is achievable by any individual method, leading to state-of-the-art performance on three standard action recognition datasets. We include links to our code and data.
Amor Ben Tanfous, Aimen Zerroug, Drew Linsley, Thomas Serre
WACV3
2022 Understanding the Computational Demands Underlying Visual Reasoning
abstract
Visual understanding requires comprehending complex visual relations between objects within a scene. Here, we seek to characterize the computational demands for abstract visual reasoning. We do this by systematically assessing the ability of modern deep convolutional neural networks (CNNs) to learn to solve the synthetic visual reasoning test (SVRT) challenge, a collection of 23 visual reasoning problems. Our analysis reveals a novel taxonomy of visual reasoning tasks, which can be primarily explained by both the type of relations (same-different versus spatial-relation judgments) and the number of relations used to compose the underlying rules. Prior cognitive neuroscience work suggests that attention plays a key role in humans' visual reasoning ability. To test this hypothesis, we extended the CNNs with spatial and feature-based attention mechanisms. In a second series of experiments, we evaluated the ability of these attention networks to learn to solve the SVRT challenge and found the resulting architectures to be much more efficient at solving the hardest of these visual reasoning tasks. Most important, the corresponding improvements on individual tasks partially explained our novel taxonomy. Overall, this work provides a granular computational account of visual reasoning and yields testable neuroscience predictions regarding the differential need for feature-based versus spatial attention depending on the type of visual reasoning problem.
Mohit Vaishnav, Rémi Cadène, Andrea Alamia, Drew Linsley, Rufin VanRullen, Thomas Serre
Neural Comput.4
2021 Tracking Without Re-recognition in Humans and Machines
abstract
Imagine trying to track one particular fruitfly in a swarm of hundreds. Higher biological visual systems have evolved to track moving objects by relying on both their appearance and their motion trajectories. We investigate if state-of-the-art spatiotemporal deep neural networks are capable of the same. For this, we introduce PathTracker, a synthetic visual challenge that asks human observers and machines to track a target object in the midst of identical-looking "distractor" objects. While humans effortlessly learn PathTracker and generalize to systematic variations in task design, deep networks struggle. To address this limitation, we identify and model circuit mechanisms in biological brains that are implicated in tracking objects based on motion cues. When instantiated as a recurrent network, our circuit model learns to solve PathTracker with a robust visual strategy that rivals human performance and explains a significant proportion of their decision-making on the challenge. We also show that the success of this circuit model extends to object tracking in natural videos. Adding it to a transformer-based architecture for object tracking builds tolerance to visual nuisances that affect object appearance, establishing the new state of the art on the large-scale TrackingNet challenge. Our work highlights the importance of understanding human vision to improve computer vision.
Drew Linsley, Girik Malik, Junkyung Kim, Lakshmi Narasimhan Govindarajan, Ennio Mingolla, Thomas Serre
NeurIPS1
2020 Disentangling neural mechanisms for perceptual grouping
Junkyung Kim, Drew Linsley, Kalpit Thakkar, Thomas Serre
ICLR2
2020 Recurrent neural circuits for contour detection
Drew Linsley, Junkyung Kim, Alekh Karkada Ashok, Thomas Serre
ICLR1
2020 Stable and expressive recurrent vision models
abstract
Primate vision depends on recurrent processing for reliable perception. A growing body of literature also suggests that recurrent connections improve the learning efficiency and generalization of vision models on classic computer vision challenges. Why then, are current large-scale challenges dominated by feedforward networks? We posit that the effectiveness of recurrent vision models is bottlenecked by the standard algorithm used for training them, "back-propagation through time" (BPTT), which has O(N) memory-complexity for training an N step model. Thus, recurrent vision model design is bounded by memory constraints, forcing a choice between rivaling the enormous capacity of leading feedforward models or trying to compensate for this deficit through granular and complex dynamics. Here, we develop a new learning algorithm, "contractor recurrent back-propagation" (C-RBP), which alleviates these issues by achieving constant O(1) memory-complexity with steps of recurrent processing. We demonstrate that recurrent vision models trained with C-RBP can detect long-range spatial dependencies in a synthetic contour tracing task that BPTT-trained models cannot. We further show that recurrent vision models trained with C-RBP to solve the large-scale Panoptic Segmentation MS-COCO challenge outperform the leading feedforward approach, with fewer free parameters. C-RBP is a general-purpose learning algorithm for any application that can benefit from expansive recurrent dynamics. Code and data are available at https://github.com/c-rbp.
Drew Linsley, Alekh Karkada Ashok, Lakshmi Narasimhan Govindarajan, Rex G. Liu, Thomas Serre
NeurIPS1
2019 Learning what and where to attend
Drew Linsley, Dan Shiebler, Sven Eberhardt, Thomas Serre
ICLR (Poster)1
2018 Learning long-range spatial dependencies with horizontal gated recurrent units
abstract
Progress in deep learning has spawned great successes in many engineering applications. As a prime example, convolutional neural networks, a type of feedforward neural networks, are now approaching -- and sometimes even surpassing -- human accuracy on a variety of visual recognition tasks. Here, however, we show that these neural networks and their recent extensions struggle in recognition tasks where co-dependent visual features must be detected over long spatial ranges. We introduce a visual challenge, Pathfinder, and describe a novel recurrent neural network architecture called the horizontal gated recurrent unit (hGRU) to learn intrinsic horizontal connections -- both within and across feature columns. We demonstrate that a single hGRU layer matches or outperforms all tested feedforward hierarchical baselines including state-of-the-art architectures with orders of magnitude more parameters.
Drew Linsley, Junkyung Kim, Vijay Veerabadran, Charles Windolf, Thomas Serre
NeurIPS1