EDBT 2026 Demo / reviewers in the wild / expert
Alekh Karkada Ashok
dblp:230/2212 · also Alekh Ashok
· DBLP profile ↗
5ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Deep learning architectures and training · 20% Trustworthy machine learning · 16% 3D vision · 15% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 10 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking
object tracking |
0.9 | 1 | 2025 | Tracking objects that change in appearance with phase synchrony · ICLR 2025 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.9 | 2 | 2020 | Stable and expressive recurrent vision models · NeurIPS 2020 Recurrent neural circuits for contour detection · ICLR 2020 |
Natural language and speech › Language models and text generation › alignment
human alignment |
0.7 | 1 | 2023 | Computing a human-like reaction time metric from stable recurrent vision models · NeurIPS 2023 |
Machine learning › Trustworthy machine learning
interpretability |
0.7 | 1 | 2023 | Computing a human-like reaction time metric from stable recurrent vision models · NeurIPS 2023 |
Computer vision › Image recognition and object detection › object detection
contour detection |
0.4 | 1 | 2020 | Recurrent neural circuits for contour detection · ICLR 2020 |
Machine learning › Efficient and distributed learning
memory-efficient training |
0.4 | 1 | 2020 | Stable and expressive recurrent vision models · NeurIPS 2020 |
Machine learning › Deep learning architectures and training › recurrent neural network
complex-valued recurrent neural network |
0.3 | 1 | 2025 | Tracking objects that change in appearance with phase synchrony · ICLR 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.3 | 1 | 2025 | The 3D-PC: a benchmark for visual perspective taking in humans and machines · ICLR 2025 |
Computer vision › Segmentation and scene understanding
panoptic segmentation |
0.1 | 1 | 2020 | Stable and expressive recurrent vision models · NeurIPS 2020 |
Image and video processing › image segmentation
contour detection |
0.1 | 1 | 2020 | Recurrent neural circuits for contour detection · ICLR 2020 |
Methods — techniques the papers use, named apart from their topics
text prompting · 0.9phase synchrony · 0.9linear probing · 0.9fine-tuning · 0.9deep neural network · 0.9complex-valued recurrent neural network · 0.9recurrent neural network · 0.9subjective logic theory · 0.7evidence accumulation · 0.7backpropagation through time · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The 3D-PC: a benchmark for visual perspective taking in humans and machinesabstractVisual perspective taking (VPT) is the ability to perceive and reason about the perspectives of others. It is an essential feature of human intelligence, which develops over the first decade of life and requires an ability to process the 3D structure of visual scenes. A growing number of reports have indicated that deep neural networks (DNNs) become capable of analyzing 3D scenes after training on large image datasets. We investigated if this emergent ability for 3D analysis in DNNs is sufficient for VPT with the 3D perception challenge (3D-PC): a novel benchmark for 3D perception in humans and DNNs. The 3D-PC is comprised of three 3D-analysis tasks posed within natural scene images: (i.) a simple test of object depth order, (ii.) a basic VPT task (VPT-basic), and (iii.) a more challenging version of VPT (VPT-perturb) designed to limit the effectiveness of "shortcut" visual strategies. We tested human participants (N=33) and linearly probed or text-prompted over 300 DNNs on the challenge and found that nearly all of the DNNs approached or exceeded human accuracy in analyzing object depth order. Surprisingly, DNN accuracy on this task correlated with their object recognition performance. In contrast, there was an extraordinary gap between DNNs and humans on VPT-basic. Humans were nearly perfect, whereas most DNNs were near chance. Fine-tuning DNNs on VPT-basic brought them close to human performance, but they, unlike humans, dropped back to chance when tested on VPT-perturb. Our challenge demonstrates that the training routines and architectures of today's DNNs are well-suited for learning basic 3D properties of scenes and objects but are ill-suited for reasoning about these properties like humans do. We release our 3D-PC datasets and code to help bridge this gap in 3D perception between humans and machines. Drew Linsley, Peisen Zhou, Alekh Karkada Ashok, Akash Nagaraj, Gaurav Gaonkar, Francis E. Lewis, Zygmunt Pizlo, Thomas Serre |
ICLR | 3 |
| 2025 | Tracking objects that change in appearance with phase synchronyabstractObjects we encounter often change appearance as we interact with them. Changes in illumination (shadows), object pose, or the movement of non-rigid objects can drastically alter available image features. How do biological visual systems track objects as they change? One plausible mechanism involves attentional mechanisms for reasoning about the locations of objects independently of their appearances --- a capability that prominent neuroscience theories have associated with computing through neural synchrony. Here, we describe a novel deep learning circuit that can learn to precisely control attention to features separately from their location in the world through neural synchrony: the complex-valued recurrent neural network (CV-RNN). Next, we compare object tracking in humans, the CV-RNN, and other deep neural networks (DNNs), using FeatureTracker: a large-scale challenge that asks observers to track objects as their locations and appearances change in precisely controlled ways. While humans effortlessly solved FeatureTracker, state-of-the-art DNNs did not. In contrast, our CV-RNN behaved similarly to humans on the challenge, providing a computational proof-of-concept for the role of phase synchronization as a neural substrate for tracking appearance-morphing objects as they move about. Sabine Muzellec, Drew Linsley, Alekh Karkada Ashok, Ennio Mingolla, Girik Malik, Rufin VanRullen, Thomas Serre |
ICLR | 3 |
| 2023 | Computing a human-like reaction time metric from stable recurrent vision modelsabstractThe meteoric rise in the adoption of deep neural networks as computational models of vision has inspired efforts to ``align” these models with humans. One dimension of interest for alignment includes behavioral choices, but moving beyond characterizing choice patterns to capturing temporal aspects of visual decision-making has been challenging. Here, we sketch a general-purpose methodology to construct computational accounts of reaction times from a stimulus-computable, task-optimized model. Specifically, we introduce a novel metric leveraging insights from subjective logic theory summarizing evidence accumulation in recurrent vision models. We demonstrate that our metric aligns with patterns of human reaction times for stimulus manipulations across four disparate visual decision-making tasks spanning perceptual grouping, mental simulation, and scene categorization. This work paves the way for exploring the temporal alignment of model and human visual strategies in the context of various other cognitive tasks toward generating testable hypotheses for neuroscience. Links to the code and data can be found on the project page: https://serre-lab.github.io/rnn_rts_site/. Lore Goetschalckx, Lakshmi Narasimhan Govindarajan, Alekh Karkada Ashok, Aarit Ahuja, David L. Sheinberg, Thomas Serre |
NeurIPS | 3 |
| 2020 | Recurrent neural circuits for contour detection
Drew Linsley, Junkyung Kim, Alekh Karkada Ashok, Thomas Serre |
ICLR | 3 |
| 2020 | Stable and expressive recurrent vision modelsabstractPrimate vision depends on recurrent processing for reliable perception. A growing body of literature also suggests that recurrent connections improve the learning efficiency and generalization of vision models on classic computer vision challenges. Why then, are current large-scale challenges dominated by feedforward networks? We posit that the effectiveness of recurrent vision models is bottlenecked by the standard algorithm used for training them, "back-propagation through time" (BPTT), which has O(N) memory-complexity for training an N step model. Thus, recurrent vision model design is bounded by memory constraints, forcing a choice between rivaling the enormous capacity of leading feedforward models or trying to compensate for this deficit through granular and complex dynamics. Here, we develop a new learning algorithm, "contractor recurrent back-propagation" (C-RBP), which alleviates these issues by achieving constant O(1) memory-complexity with steps of recurrent processing. We demonstrate that recurrent vision models trained with C-RBP can detect long-range spatial dependencies in a synthetic contour tracing task that BPTT-trained models cannot. We further show that recurrent vision models trained with C-RBP to solve the large-scale Panoptic Segmentation MS-COCO challenge outperform the leading feedforward approach, with fewer free parameters. C-RBP is a general-purpose learning algorithm for any application that can benefit from expansive recurrent dynamics. Code and data are available at https://github.com/c-rbp. Drew Linsley, Alekh Karkada Ashok, Lakshmi Narasimhan Govindarajan, Rex G. Liu, Thomas Serre |
NeurIPS | 2 |