EDBT 2026 Demo / reviewers in the wild / expert
Drew Linsley
dblp:194/2308
· DBLP profile ↗
13ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0002-9722-7839ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 7 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Deep learning architectures and training · 26% Trustworthy machine learning · 23% Image recognition and object detection · 13% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 18 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
recurrent neural network |
1.7 | 4 | 2021 | Tracking Without Re-recognition in Humans and Machines · NeurIPS 2021 Stable and expressive recurrent vision models · NeurIPS 2020 Recurrent neural circuits for contour detection · ICLR 2020 |
Computer vision › Video understanding and tracking
object tracking |
1.4 | 2 | 2025 | Tracking objects that change in appearance with phase synchrony · ICLR 2025 Tracking Without Re-recognition in Humans and Machines · NeurIPS 2021 |
Machine learning › Trustworthy machine learning
interpretability |
1.2 | 2 | 2023 | Unlocking Feature Visualization for Deep Network with MAgnitude Constrained Optimization · NeurIPS 2023 Harmonizing the object recognition strategies of deep neural networks with humans · NeurIPS 2022 |
Machine learning › Trustworthy machine learning › interpretability
attribution methods |
0.7 | 1 | 2023 | Unlocking Feature Visualization for Deep Network with MAgnitude Constrained Optimization · NeurIPS 2023 |
Computer vision › 3D vision
biological vision modeling |
0.7 | 1 | 2023 | Performance-optimized deep neural networks are evolving into worse models of inferotemporal visual cortex · NeurIPS 2023 |
Machine learning › Trustworthy machine learning › interpretability › visual explanation
feature visualization |
0.7 | 1 | 2023 | Unlocking Feature Visualization for Deep Network with MAgnitude Constrained Optimization · NeurIPS 2023 |
Machine learning › Representation and self-supervised learning › representation matching › feature alignment
neural representation alignment |
0.7 | 1 | 2023 | Performance-optimized deep neural networks are evolving into worse models of inferotemporal visual cortex · NeurIPS 2023 |
Computer vision › Image recognition and object detection › object detection
contour detection |
0.4 | 1 | 2020 | Recurrent neural circuits for contour detection · ICLR 2020 |
Machine learning › Efficient and distributed learning
memory-efficient training |
0.4 | 1 | 2020 | Stable and expressive recurrent vision models · NeurIPS 2020 |
Computer vision › Segmentation and scene understanding
perceptual grouping |
0.4 | 1 | 2020 | Disentangling neural mechanisms for perceptual grouping · ICLR 2020 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.4 | 1 | 2019 | Learning what and where to attend · ICLR (Poster) 2019 |
Machine learning › Deep learning architectures and training › recurrent neural network › gated recurrent network
gated recurrent unit |
0.3 | 1 | 2018 | Learning long-range spatial dependencies with horizontal gated recurrent units · NeurIPS 2018 |
Machine learning › Deep learning architectures and training › recurrent neural network
complex-valued recurrent neural network |
0.3 | 1 | 2025 | Tracking objects that change in appearance with phase synchrony · ICLR 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.3 | 1 | 2025 | The 3D-PC: a benchmark for visual perspective taking in humans and machines · ICLR 2025 |
Computer vision › Image recognition and object detection
object recognition |
0.2 | 1 | 2023 | Performance-optimized deep neural networks are evolving into worse models of inferotemporal visual cortex · NeurIPS 2023 |
Computer vision › Video understanding and tracking › object tracking
transformer-based tracking |
0.1 | 1 | 2021 | Tracking Without Re-recognition in Humans and Machines · NeurIPS 2021 |
Computer vision › Segmentation and scene understanding
panoptic segmentation |
0.1 | 1 | 2020 | Stable and expressive recurrent vision models · NeurIPS 2020 |
Image and video processing › image segmentation
contour detection |
0.1 | 1 | 2020 | Recurrent neural circuits for contour detection · ICLR 2020 |
Methods — techniques the papers use, named apart from their topics
text prompting · 0.9phase synchrony · 0.9linear probing · 0.9fine-tuning · 0.9deep neural network · 0.9complex-valued recurrent neural network · 0.9phase spectrum optimization · 0.7magnitude-constrained optimization · 0.7representational alignment analysis · 0.6neural harmonizer · 0.6recurrent neural network · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The 3D-PC: a benchmark for visual perspective taking in humans and machinesabstractVisual perspective taking (VPT) is the ability to perceive and reason about the perspectives of others. It is an essential feature of human intelligence, which develops over the first decade of life and requires an ability to process the 3D structure of visual scenes. A growing number of reports have indicated that deep neural networks (DNNs) become capable of analyzing 3D scenes after training on large image datasets. We investigated if this emergent ability for 3D analysis in DNNs is sufficient for VPT with the 3D perception challenge (3D-PC): a novel benchmark for 3D perception in humans and DNNs. The 3D-PC is comprised of three 3D-analysis tasks posed within natural scene images: (i.) a simple test of object depth order, (ii.) a basic VPT task (VPT-basic), and (iii.) a more challenging version of VPT (VPT-perturb) designed to limit the effectiveness of "shortcut" visual strategies. We tested human participants (N=33) and linearly probed or text-prompted over 300 DNNs on the challenge and found that nearly all of the DNNs approached or exceeded human accuracy in analyzing object depth order. Surprisingly, DNN accuracy on this task correlated with their object recognition performance. In contrast, there was an extraordinary gap between DNNs and humans on VPT-basic. Humans were nearly perfect, whereas most DNNs were near chance. Fine-tuning DNNs on VPT-basic brought them close to human performance, but they, unlike humans, dropped back to chance when tested on VPT-perturb. Our challenge demonstrates that the training routines and architectures of today's DNNs are well-suited for learning basic 3D properties of scenes and objects but are ill-suited for reasoning about these properties like humans do. We release our 3D-PC datasets and code to help bridge this gap in 3D perception between humans and machines. Drew Linsley, Peisen Zhou, Alekh Karkada Ashok, Akash Nagaraj, Gaurav Gaonkar, Francis E. Lewis, Zygmunt Pizlo, Thomas Serre |
ICLR | 1 |
| 2025 | Tracking objects that change in appearance with phase synchronyabstractObjects we encounter often change appearance as we interact with them. Changes in illumination (shadows), object pose, or the movement of non-rigid objects can drastically alter available image features. How do biological visual systems track objects as they change? One plausible mechanism involves attentional mechanisms for reasoning about the locations of objects independently of their appearances --- a capability that prominent neuroscience theories have associated with computing through neural synchrony. Here, we describe a novel deep learning circuit that can learn to precisely control attention to features separately from their location in the world through neural synchrony: the complex-valued recurrent neural network (CV-RNN). Next, we compare object tracking in humans, the CV-RNN, and other deep neural networks (DNNs), using FeatureTracker: a large-scale challenge that asks observers to track objects as their locations and appearances change in precisely controlled ways. While humans effortlessly solved FeatureTracker, state-of-the-art DNNs did not. In contrast, our CV-RNN behaved similarly to humans on the challenge, providing a computational proof-of-concept for the role of phase synchronization as a neural substrate for tracking appearance-morphing objects as they move about. Sabine Muzellec, Drew Linsley, Alekh Karkada Ashok, Ennio Mingolla, Girik Malik, Rufin VanRullen, Thomas Serre |
ICLR | 2 |
| 2023 | Unlocking Feature Visualization for Deep Network with MAgnitude Constrained OptimizationabstractFeature visualization has gained significant popularity as an explainability method, particularly after the influential work by Olah et al. in 2017. Despite its success, its widespread adoption has been limited due to issues in scaling to deeper neural networks and the reliance on tricks to generate interpretable images. Here, we describe MACO, a simple approach to address these shortcomings. It consists in optimizing solely an image's phase spectrum while keeping its magnitude constant to ensure that the generated explanations lie in the space of natural images. Our approach yields significantly better results -- both qualitatively and quantitatively -- unlocking efficient and interpretable feature visualizations for state-of-the-art neural networks. We also show that our approach exhibits an attribution mechanism allowing to augment feature visualizations with spatial importance. Furthermore, we enable quantitative evaluation of feature visualizations by introducing 3 metrics: transferability, plausibility, and alignment with natural images. We validate our method on various applications and we introduce a website featuring MACO visualizations for all classes of the ImageNet dataset, which will be made available upon acceptance.
Overall, our study unlocks feature visualizations for the largest, state-of-the-art classification networks without resorting to any parametric prior image model, effectively advancing a field that has been stagnating since 2017 (Olah et al, 2017). Thomas Fel, Thibaut Boissin, Victor Boutin, Agustin M. Picard, Paul Novello, Julien Colin, Drew Linsley, Tom Rousseau, Rémi Cadène, Lore Goetschalckx, Laurent Gardes, Thomas Serre |
NeurIPS | 7 |
| 2023 | Performance-optimized deep neural networks are evolving into worse models of inferotemporal visual cortexabstractOne of the most impactful findings in computational neuroscience over the past decade is that the object recognition accuracy of deep neural networks (DNNs) correlates with their ability to predict neural responses to natural images in the inferotemporal (IT) cortex. This discovery supported the long-held theory that object recognition is a core objective of the visual cortex, and suggested that more accurate DNNs would serve as better models of IT neuron responses to images. Since then, deep learning has undergone a revolution of scale: billion parameter-scale DNNs trained on billions of images are rivaling or outperforming humans at visual tasks including object recognition. Have today's DNNs become more accurate at predicting IT neuron responses to images as they have grown more accurate at object recognition?
Surprisingly, across three independent experiments, we find that this is not the case. DNNs have become progressively worse models of IT as their accuracy has increased on ImageNet. To understand why DNNs experience this trade-off and evaluate if they are still an appropriate paradigm for modeling the visual system, we turn to recordings of IT that capture spatially resolved maps of neuronal activity elicited by natural images. These neuronal activity maps reveal that DNNs trained on ImageNet learn to rely on different visual features than those encoded by IT and that this problem worsens as their accuracy increases. We successfully resolved this issue with the neural harmonizer, a plug-and-play training routine for DNNs that aligns their learned representations with humans. Our results suggest that harmonized DNNs break the trade-off between ImageNet accuracy and neural prediction accuracy that assails current DNNs and offer a path to more accurate models of biological vision. Our work indicates that the standard approach for modeling IT with task-optimized DNNs needs revision, and other biological constraints, including human psychophysics data, are needed to accurately reverse-engineer the visual cortex. Drew Linsley, Ivan F. Rodriguez Rodriguez, Thomas Fel, Michael Arcaro, Saloni Sharma, Margaret S. Livingstone, Thomas Serre |
NeurIPS | 1 |
| 2022 | Harmonizing the object recognition strategies of deep neural networks with humansabstractThe many successes of deep neural networks (DNNs) over the past decade have largely been driven by computational scale rather than insights from biological intelligence. Here, we explore if these trends have also carried concomitant improvements in explaining the visual strategies humans rely on for object recognition. We do this by comparing two related but distinct properties of visual strategies in humans and DNNs: where they believe important visual features are in images and how they use those features to categorize objects. Across 84 different DNNs trained on ImageNet and three independent datasets measuring the where and the how of human visual strategies for object recognition on those images, we find a systematic trade-off between DNN categorization accuracy and alignment with human visual strategies for object recognition. \textit{State-of-the-art DNNs are progressively becoming less aligned with humans as their accuracy improves}. We rectify this growing issue with our neural harmonizer: a general-purpose training routine that both aligns DNN and human visual strategies and improves categorization accuracy. Our work represents the first demonstration that the scaling laws that are guiding the design of DNNs today have also produced worse models of human vision. We release our code and data at https://serre-lab.github.io/Harmonization to help the field build more human-like DNNs. Thomas Fel, Ivan F. Rodriguez Rodriguez, Drew Linsley, Thomas Serre |
NeurIPS | 3 |
| 2022 | How and What to Learn: Taxonomizing Self-Supervised Learning for 3D Action RecognitionabstractThere are two competing standards for self-supervised learning in action recognition from 3D skeletons. Su et al., 2020 [31] used an auto-encoder architecture and an image reconstruction objective function to achieve state-of-the-art performance on the NTU60 C-View benchmark. Rao et al., 2020 [23] used Contrastive learning in the latent space to achieve state-of-the-art performance on the NTU60 C-Sub benchmark. Here, we reconcile these disparate approaches by developing a taxonomy of self-supervised learning for action recognition. We observe that leading approaches generally use one of two types of objective functions: those that seek to reconstruct the input from a latent representation ("Attractive" learning) versus those that also try to maximize the representations distinctiveness ("Contrastive" learning). Independently, leading approaches also differ in how they implement these objective functions: there are those that optimize representations in the decoder output space and those which optimize representations in the network’s latent space (encoder output). We find that combining these approaches leads to larger gains in performance and tolerance to transformation than is achievable by any individual method, leading to state-of-the-art performance on three standard action recognition datasets. We include links to our code and data. Amor Ben Tanfous, Aimen Zerroug, Drew Linsley, Thomas Serre |
WACV | 3 |
| 2022 | Understanding the Computational Demands Underlying Visual ReasoningabstractVisual understanding requires comprehending complex visual relations between objects within a scene. Here, we seek to characterize the computational demands for abstract visual reasoning. We do this by systematically assessing the ability of modern deep convolutional neural networks (CNNs) to learn to solve the synthetic visual reasoning test (SVRT) challenge, a collection of 23 visual reasoning problems. Our analysis reveals a novel taxonomy of visual reasoning tasks, which can be primarily explained by both the type of relations (same-different versus spatial-relation judgments) and the number of relations used to compose the underlying rules. Prior cognitive neuroscience work suggests that attention plays a key role in humans' visual reasoning ability. To test this hypothesis, we extended the CNNs with spatial and feature-based attention mechanisms. In a second series of experiments, we evaluated the ability of these attention networks to learn to solve the SVRT challenge and found the resulting architectures to be much more efficient at solving the hardest of these visual reasoning tasks. Most important, the corresponding improvements on individual tasks partially explained our novel taxonomy. Overall, this work provides a granular computational account of visual reasoning and yields testable neuroscience predictions regarding the differential need for feature-based versus spatial attention depending on the type of visual reasoning problem. Mohit Vaishnav, Rémi Cadène, Andrea Alamia, Drew Linsley, Rufin VanRullen, Thomas Serre |
Neural Comput. | 4 |
| 2021 | Tracking Without Re-recognition in Humans and MachinesabstractImagine trying to track one particular fruitfly in a swarm of hundreds. Higher biological visual systems have evolved to track moving objects by relying on both their appearance and their motion trajectories. We investigate if state-of-the-art spatiotemporal deep neural networks are capable of the same. For this, we introduce PathTracker, a synthetic visual challenge that asks human observers and machines to track a target object in the midst of identical-looking "distractor" objects. While humans effortlessly learn PathTracker and generalize to systematic variations in task design, deep networks struggle. To address this limitation, we identify and model circuit mechanisms in biological brains that are implicated in tracking objects based on motion cues. When instantiated as a recurrent network, our circuit model learns to solve PathTracker with a robust visual strategy that rivals human performance and explains a significant proportion of their decision-making on the challenge. We also show that the success of this circuit model extends to object tracking in natural videos. Adding it to a transformer-based architecture for object tracking builds tolerance to visual nuisances that affect object appearance, establishing the new state of the art on the large-scale TrackingNet challenge. Our work highlights the importance of understanding human vision to improve computer vision. Drew Linsley, Girik Malik, Junkyung Kim, Lakshmi Narasimhan Govindarajan, Ennio Mingolla, Thomas Serre |
NeurIPS | 1 |
| 2020 | Disentangling neural mechanisms for perceptual grouping
Junkyung Kim, Drew Linsley, Kalpit Thakkar, Thomas Serre |
ICLR | 2 |
| 2020 | Recurrent neural circuits for contour detection
Drew Linsley, Junkyung Kim, Alekh Karkada Ashok, Thomas Serre |
ICLR | 1 |
| 2020 | Stable and expressive recurrent vision modelsabstractPrimate vision depends on recurrent processing for reliable perception. A growing body of literature also suggests that recurrent connections improve the learning efficiency and generalization of vision models on classic computer vision challenges. Why then, are current large-scale challenges dominated by feedforward networks? We posit that the effectiveness of recurrent vision models is bottlenecked by the standard algorithm used for training them, "back-propagation through time" (BPTT), which has O(N) memory-complexity for training an N step model. Thus, recurrent vision model design is bounded by memory constraints, forcing a choice between rivaling the enormous capacity of leading feedforward models or trying to compensate for this deficit through granular and complex dynamics. Here, we develop a new learning algorithm, "contractor recurrent back-propagation" (C-RBP), which alleviates these issues by achieving constant O(1) memory-complexity with steps of recurrent processing. We demonstrate that recurrent vision models trained with C-RBP can detect long-range spatial dependencies in a synthetic contour tracing task that BPTT-trained models cannot. We further show that recurrent vision models trained with C-RBP to solve the large-scale Panoptic Segmentation MS-COCO challenge outperform the leading feedforward approach, with fewer free parameters. C-RBP is a general-purpose learning algorithm for any application that can benefit from expansive recurrent dynamics. Code and data are available at https://github.com/c-rbp. Drew Linsley, Alekh Karkada Ashok, Lakshmi Narasimhan Govindarajan, Rex G. Liu, Thomas Serre |
NeurIPS | 1 |
| 2019 | Learning what and where to attend
Drew Linsley, Dan Shiebler, Sven Eberhardt, Thomas Serre |
ICLR (Poster) | 1 |
| 2018 | Learning long-range spatial dependencies with horizontal gated recurrent unitsabstractProgress in deep learning has spawned great successes in many engineering applications. As a prime example, convolutional neural networks, a type of feedforward neural networks, are now approaching -- and sometimes even surpassing -- human accuracy on a variety of visual recognition tasks. Here, however, we show that these neural networks and their recent extensions struggle in recognition tasks where co-dependent visual features must be detected over long spatial ranges. We introduce a visual challenge, Pathfinder, and describe a novel recurrent neural network architecture called the horizontal gated recurrent unit (hGRU) to learn intrinsic horizontal connections -- both within and across feature columns. We demonstrate that a single hGRU layer matches or outperforms all tested feedforward hierarchical baselines including state-of-the-art architectures with orders of magnitude more parameters. Drew Linsley, Junkyung Kim, Vijay Veerabadran, Charles Windolf, Thomas Serre |
NeurIPS | 1 |