EDBT 2026 Demo / reviewers in the wild / expert
Colin Conwell
dblp:309/6127
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Representation and self-supervised learning · 26% Video understanding and tracking · 22% 3D vision · 13% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% |
Topics — the 14 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking
action recognition |
0.9 | 1 | 2025 | Modeling dynamic social vision highlights gaps between deep learning and humans · ICLR 2025 |
Computer vision › 3D vision › 3d scene understanding
dynamic scene understanding |
0.9 | 1 | 2025 | Modeling dynamic social vision highlights gaps between deep learning and humans · ICLR 2025 |
Machine learning › Representation and self-supervised learning › representation matching
feature alignment |
0.9 | 1 | 2025 | Training the Untrainable: Introducing Inductive Bias via Representational Alignment · NeurIPS 2025 |
Machine learning › Learning theory
inductive bias |
0.9 | 1 | 2025 | Training the Untrainable: Introducing Inductive Bias via Representational Alignment · NeurIPS 2025 |
Machine learning › Transfer learning and domain adaptation
knowledge transfer |
0.9 | 1 | 2025 | Training the Untrainable: Introducing Inductive Bias via Representational Alignment · NeurIPS 2025 |
Knowledge, reasoning and agents › Multi-agent systems › agent interaction
multi-agent interaction |
0.9 | 1 | 2025 | Modeling dynamic social vision highlights gaps between deep learning and humans · ICLR 2025 |
Computer vision › Video understanding and tracking › activity recognition
social interaction recognition |
0.9 | 1 | 2025 | Modeling dynamic social vision highlights gaps between deep learning and humans · ICLR 2025 |
Machine learning › Representation and self-supervised learning › computational neuroscience › neural coding
brain encoding models |
0.8 | 1 | 2024 | Revealing Vision-Language Integration in the Brain with Multimodal Networks · ICML 2024 |
Bioinformatics and computational biology › neuroscience
neuroinformatics |
0.8 | 1 | 2024 | Revealing Vision-Language Integration in the Brain with Multimodal Networks · ICML 2024 |
Machine learning › Representation and self-supervised learning › representation analysis
representational similarity analysis |
0.5 | 1 | 2021 | Neural Regression, Representational Similarity, Model Zoology & Neural Taskonomy at Scale in Rodent Visual Cortex · NeurIPS 2021 |
Bioinformatics and computational biology
computational neuroscience |
0.5 | 1 | 2021 | Neural Regression, Representational Similarity, Model Zoology & Neural Taskonomy at Scale in Rodent Visual Cortex · NeurIPS 2021 |
Natural language and speech › Language models and text generation › large language model evaluation
human-model comparison |
0.3 | 1 | 2025 | Modeling dynamic social vision highlights gaps between deep learning and humans · ICLR 2025 |
Computer vision › Image recognition and object detection
object recognition |
0.3 | 1 | 2025 | Training the Untrainable: Introducing Inductive Bias via Representational Alignment · NeurIPS 2025 |
Computer vision › 3D vision › biological vision modeling
visual cortex modeling |
0.1 | 1 | 2021 | Neural Regression, Representational Similarity, Model Zoology & Neural Taskonomy at Scale in Rodent Visual Cortex · NeurIPS 2021 |
Methods — techniques the papers use, named apart from their topics
stereoencephalography · 1.5cross-attention · 1.5contrastive learning · 1.5neural regression · 1.0benchmarking · 1.0video captioning · 0.9neural response prediction · 0.9neural distance function · 0.9layerwise representational similarity · 0.9knowledge distillation · 0.9representational similarity analysis · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Modeling dynamic social vision highlights gaps between deep learning and humansabstractDeep learning models trained on computer vision tasks are widely considered the most successful models of human vision to date. The majority of work that supports this idea evaluates how accurately these models predict behavior and brain responses to static images of objects and scenes. Real-world vision, however, is highly dynamic, and far less work has evaluated deep learning models on human responses to moving stimuli, especially those that involve more complicated, higher-order phenomena like social interactions. Here, we extend a dataset of natural videos depicting complex multi-agent interactions by collecting human-annotated sentence captions for each video, and we benchmark 350+ image, video, and language models on behavior and neural responses to the videos. As in prior work, we find that many vision models reach the noise ceiling in predicting visual scene features and responses along the ventral visual stream (often considered the primary neural substrate of object and scene recognition). In contrast, vision models poorly predict human action and social interaction ratings and neural responses in the lateral stream (a neural pathway theorized to specialize in dynamic, social vision), though video models show a striking advantage in predicting mid-level lateral stream regions. Language models (given human sentence captions of the videos) predict action and social ratings better than image and video models, but perform poorly at predicting neural responses in the lateral stream. Together, these results identify a major gap in AI's ability to match human social vision and provide insights to guide future model development for dynamic, natural contexts. Kathy Garcia, Emalie McMahon, Colin Conwell, Michael F. Bonner, Leyla Isik |
ICLR | 3 |
| 2025 | Training the Untrainable: Introducing Inductive Bias via Representational AlignmentabstractWe demonstrate that architectures which traditionally are considered to be ill-suited for a task can be trained using inductive biases from another architecture. We call a network untrainable when it overfits, underfits, or converges to poor results even when tuning their hyperparameters. For example, fully connected networks overfit on object recognition while deep convolutional networks without residual connections underfit. The traditional answer is to change the architecture to impose some inductive bias, although the nature of that bias is unknown. We introduce guidance, where a guide network steers a target network using a neural distance function. The target minimizes its task loss plus a layerwise representational similarity against the frozen guide. If the guide is trained, this transfers over the architectural prior and knowledge of the guide to the target. If the guide is untrained, this transfers over only part of the architectural prior of the guide. We show that guidance prevents FCN overfitting on ImageNet, narrows the vanilla RNN–Transformer gap, boosts plain CNNs toward ResNet accuracy, and aids Transformers on RNN-favored tasks. We further identify that guidance-driven initialization alone can mitigate FCN overfitting. Our method provides a mathematical tool to investigate priors and architectures, and in the long term, could automate architecture design. Vighnesh Subramaniam, David Mayo, Colin Conwell, Tomaso A. Poggio, Boris Katz, Brian Cheung, Andrei Barbu |
NeurIPS | 3 |
| 2024 | Revealing Vision-Language Integration in the Brain with Multimodal NetworksabstractWe use (multi)modal deep neural networks (DNNs) to probe for sites of multimodal integration in the human brain by predicting stereoencephalography (SEEG) recordings taken while human subjects watched movies. We operationalize sites of multimodal integration as regions where a multimodal vision-language model predicts recordings better than unimodal language, unimodal vision, or linearly-integrated language-vision models. Our target DNN models span different architectures (e.g., convolutional networks and transformers) and multimodal training techniques (e.g., cross-attention and contrastive learning). As a key enabling step, we first demonstrate that trained vision and language models systematically outperform their randomly initialized counterparts in their ability to predict SEEG signals. We then compare unimodal and multimodal models against one another. Because our target DNN models often have different architectures, number of parameters, and training sets (possibly obscuring those differences attributable to integration), we carry out a controlled comparison of two models (SLIP and SimCLR), which keep all of these attributes the same aside from input modality. Using this approach, we identify a sizable number of neural sites (on average 141 out of 1090 total sites or 12.94%) and brain regions where multimodal integration seems to occur. Additionally, we find that among the variants of multimodal training techniques we assess, CLIP-style training is the best suited for downstream prediction of the neural activity in these sites. Vighnesh Subramaniam, Colin Conwell, Christopher Wang, Gabriel Kreiman, Boris Katz, Ignacio Cases, Andrei Barbu |
ICML | 2 |
| 2021 | Neural Regression, Representational Similarity, Model Zoology & Neural Taskonomy at Scale in Rodent Visual CortexabstractHow well do deep neural networks fare as models of mouse visual cortex? A majority of research to date suggests results far more mixed than those produced in the modeling of primate visual cortex. Here, we perform a large-scale benchmarking of dozens of deep neural network models in mouse visual cortex with both representational similarity analysis and neural regression. Using the Allen Brain Observatory's 2-photon calcium-imaging dataset of activity in over 6,000 reliable rodent visual cortical neurons recorded in response to natural scenes, we replicate previous findings and resolve previous discrepancies, ultimately demonstrating that modern neural networks can in fact be used to explain activity in the mouse visual cortex to a more reasonable degree than previously suggested. Using our benchmark as an atlas, we offer preliminary answers to overarching questions about levels of analysis (e.g. do models that better predict the representations of individual neurons also predict representational similarity across neural populations?); questions about the properties of models that best predict the visual system overall (e.g. is convolution or category-supervision necessary to better predict neural activity?); and questions about the mapping between biological and artificial representations (e.g. does the information processing hierarchy in deep nets match the anatomical hierarchy of mouse visual cortex?). Along the way, we catalogue a number of models (including vision transformers, MLP-Mixers, normalization free networks, Taskonomy encoders and self-supervised models) outside the traditional circuit of convolutional object recognition. Taken together, our results provide a reference point for future ventures in the deep neural network modeling of mouse visual cortex, hinting at novel combinations of mapping method, architecture, and task to more fully characterize the computational motifs of visual representation in a species so central to neuroscience, but with a perceptual physiology and ecology markedly different from the ones we study in primates. Colin Conwell, David Mayo, Andrei Barbu, Michael A. Buice, George Alvarez, Boris Katz |
NeurIPS | 1 |