George A. Alvarez

dblp:79/8258 · also George Angelo Alvarez · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
4since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Deep learning architectures and training · 37% Image recognition and object detection · 26% Trustworthy machine learning · 19%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 71% Computational science and engineering · 29%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › self-supervised visual representation learning
self-supervised vision model
0.912025
Visual Anagrams Reveal Hidden Differences in Holistic Shape Processing Across Vision Models · NeurIPS 2025
Computer vision › Image recognition and object detection
shape recognition
0.912025
Visual Anagrams Reveal Hidden Differences in Holistic Shape Processing Across Vision Models · NeurIPS 2025
Machine learning › Deep learning architectures and training › regularization
dropout
0.812024
Manipulating dropout reveals an optimal balance of efficiency and robustness in biological and machine visual systems · ICLR 2024
Machine learning › Deep learning architectures and training
regularization
0.812024
Manipulating dropout reveals an optimal balance of efficiency and robustness in biological and machine visual systems · ICLR 2024
Machine learning › Deep learning architectures and training
feedback loop
0.712023
Cognitive Steering in Deep Neural Networks via Long-Range Modulatory Feedback Connections · NeurIPS 2023
Computer vision › Image recognition and object detection
visual recognition
0.712023
Cognitive Steering in Deep Neural Networks via Long-Range Modulatory Feedback Connections · NeurIPS 2023
Machine learning › Trustworthy machine learning
robustness
0.422024
Manipulating dropout reveals an optimal balance of efficiency and robustness in biological and machine visual systems · ICLR 2024
Cognitive Steering in Deep Neural Networks via Long-Range Modulatory Feedback Connections · NeurIPS 2023
Machine learning › Trustworthy machine learning
interpretability
0.312025
Visual Anagrams Reveal Hidden Differences in Holistic Shape Processing Across Vision Models · NeurIPS 2025
Machine learning › Trustworthy machine learning › robustness › spurious correlation
texture bias
0.312025
Visual Anagrams Reveal Hidden Differences in Holistic Shape Processing Across Vision Models · NeurIPS 2025
Bioinformatics and computational biology
computational neuroscience
0.212024
Manipulating dropout reveals an optimal balance of efficiency and robustness in biological and machine visual systems · ICLR 2024
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.212023
Cognitive Steering in Deep Neural Networks via Long-Range Modulatory Feedback Connections · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
approximate inference
0.112009
Explaining human multiple object tracking as resource-constrained approximate inference in a dynamic probabilistic model · NIPS 2009
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › sequential monte carlo
particle filtering
0.112009
Explaining human multiple object tracking as resource-constrained approximate inference in a dynamic probabilistic model · NIPS 2009
Computational science and engineering
computational cognitive science
0.112009
Explaining human multiple object tracking as resource-constrained approximate inference in a dynamic probabilistic model · NIPS 2009
Machine learning › Efficient and distributed learning › inference efficiency
resource-constrained inference
0.012009
Explaining human multiple object tracking as resource-constrained approximate inference in a dynamic probabilistic model · NIPS 2009

Methods — techniques the papers use, named apart from their topics

fMRI representational comparison · 1.5eigenspectrum analysis · 1.5representational similarity analysis · 0.9attention masking · 0.9modulatory feedback · 0.7cognitive steering · 0.7rao-blackwellized particle filter · 0.2ideal-observer model · 0.1ideal observer model · 0.1
YearPublicationVenuePosition
2025 Visual Anagrams Reveal Hidden Differences in Holistic Shape Processing Across Vision Models
abstract
Humans are able to recognize objects based on both local texture cues and the configuration of object parts, yet contemporary vision models primarily harvest local texture cues, yielding brittle, non-compositional features. Work on shape-vs-texture bias has pitted shape and texture representations in opposition, measuring shape relative to texture, ignoring the possibility that models (and humans) can simultaneously rely on both types of cues, and obscuring the absolute quality of both types of representation. We therefore recast shape evaluation as a matter of absolute configural competence, operationalized by the Configural Shape Score (CSS), which (i) measures the ability to recognize both images in Object-Anagram pairs that preserve local texture while permuting global part arrangement to depict different object categories. Across 86 convolutional, transformer, and hybrid models, CSS (ii) uncovers a broad spectrum of configural sensitivity with fully self-supervised and language-aligned transformers -- exemplified by DINOv2, SigLIP2 and EVA-CLIP -- occupying the top end of the CSS spectrum. Mechanistic probes reveal that (iii) high-CSS networks depend on long-range interactions: radius-controlled attention masks abolish performance showing a distinctive U-shaped integration profile, and representational-similarity analyses expose a mid-depth transition from local to global coding. A BagNet control, whose receptive fields straddle patch seams, remains at chance (iv), ruling out any "border-hacking" strategies. Finally, (v) we show that configural shape score also predicts other shape-dependent evals (e.g., foreground bias, spectral and noise robustness). Overall, we propose that the path toward truly robust, generalizable, and human-like vision systems may not lie in forcing an artificial choice between shape and texture, but rather in architectural and learning frameworks that seamlessly integrate both local-texture and global configural shape.
Fenil R. Doshi, Thomas Fel, Talia Konkle, George A. Alvarez
NeurIPS4
2025 A feedforward mechanism for human-like contour integration
abstract
Deep neural network models provide a powerful experimental platform for exploring core mechanisms underlying human visual perception, such as perceptual grouping and contour integration-the process of linking local edge elements to arrive at a unified perceptual representation of a complete contour. Here, we demonstrate that feedforward convolutional neural networks (CNNs) fine-tuned on contour detection show this human-like capacity, but without relying on mechanisms proposed in prior work, such as lateral connections, recurrence, or top-down feedback. We identified two key properties needed for ImageNet pre-trained, feed-forward models to yield human-like contour integration: first, progressively increasing receptive field structure served as a critical architectural motif to support this capacity; and second, biased fine-tuning for contour-detection specifically for gradual curves (~20 degrees) resulted in human-like sensitivity to curvature. We further demonstrate that fine-tuning ImageNet pretrained models uncovers other hidden human-like capacities in feed-forward networks, including uncrowding (reduced interference from distractors as the number of distractors increases), which is considered a signature of human perceptual grouping. Thus, taken together these results provide a computational existence proof that purely feedforward hierarchical computations are capable of implementing gestalt "good continuation" and perceptual organization needed for human-like contour-integration and uncrowding. More broadly, these results raise the possibility that in human vision, later stages of processing play a more prominent role in perceptual-organization than implied by theories focused on recurrence and early lateral connections.
Fenil R. Doshi, Talia Konkle, George A. Alvarez
PLoS Comput. Biol.3
2024 Manipulating dropout reveals an optimal balance of efficiency and robustness in biological and machine visual systems
abstract
According to the efficient coding hypothesis, neural populations encode information optimally when representations are high-dimensional and uncorrelated. However, such codes may carry a cost in terms of generalization and robustness. Past empirical studies of early visual cortex (V1) in rodents have suggested that this tradeoff indeed constrains sensory representations. However, it remains unclear whether these insights generalize across the hierarchy of the human visual system, and particularly to object representations in high-level occipitotemporal cortex (OTC). To gain new empirical clarity, here we develop a family of object recognition models with parametrically varying dropout proportion $p$, which induces systematically varying dimensionality of internal responses (while controlling all other inductive biases). We find that increasing dropout produces an increasingly smooth, low-dimensional representational space. Optimal robustness to lesioning is observed at around 70% dropout, after which both accuracy and robustness decline. Representational comparison to large-scale 7T fMRI data from occipitotemporal cortex in the Natural Scenes Dataset reveals that this optimal degree of dropout is also associated with maximal emergent neural predictivity. Finally, using new techniques for achieving denoised estimates of the eigenspectrum of human fMRI responses, we compare the rate of eigenspectrum decay between model and brain feature spaces. We observe that the match between model and brain representations is associated with a common balance between efficiency and robustness in the representational space. These results suggest that varying dropout may reveal an optimal point of balance between the efficiency of high-dimensional codes and the robustness of low dimensional codes in hierarchical vision systems.
Jacob S. Prince, Gabriel Fajardo, George A. Alvarez, Talia Konkle
ICLR3
2023 Cognitive Steering in Deep Neural Networks via Long-Range Modulatory Feedback Connections
abstract
Given the rich visual information available in each glance, humans can internally direct their visual attention to enhance goal-relevant information---a capacity often absent in standard vision models. Here we introduce cognitively and biologically-inspired long-range modulatory pathways to enable `cognitive steering’ in vision models. First, we show that models equipped with these feedback pathways naturally show improved image recognition, adversarial robustness, and increased brain alignment, relative to baseline models. Further, these feedback projections from the final layer of the vision backbone provide a meaningful steering interface, where goals can be specified as vectors in the output space. We show that there are effective ways to steer the model that dramatically improve recognition of categories in composite images of multiple categories, succeeding where baseline feed-forward models without flexible steering fail. And, our multiplicative modulatory motif prevents rampant hallucination of the top-down goal category, dissociating what the model is looking for, from what it is looking at. Thus, these long-range modulatory pathways enable new behavioral capacities for goal-directed visual encoding, offering a flexible communication interface between cognitive and visual systems.
Talia Konkle, George A. Alvarez
NeurIPS2
2014 Multi-modal Symbolic Representations of Number: Everything You Ever Wanted to Know About Mental Abacus, but Were Afraid to Ask
David Barner, George A. Alvarez, Mahesh Srinivasan, Neon Brooks, Susan Goldin-Meadow, Jessica Sullivan, Katie Wagner, Michael C. Frank
CogSci2
2009 Explaining human multiple object tracking as resource-constrained approximate inference in a dynamic probabilistic model
abstract
Multiple object tracking is a task commonly used to investigate the architecture of human visual attention. Human participants show a distinctive pattern of successes and failures in tracking experiments that is often attributed to limits on an object system, a tracking module, or other specialized cognitive structures. Here we use a computational analysis of the task of object tracking to ask which human failures arise from cognitive limitations and which are consequences of inevitable perceptual uncertainty in the tracking task. We find that many human performance phenomena, measured through novel behavioral experiments, are naturally produced by the operation of our ideal observer model (a Rao-Blackwelized particle filter). The tradeoff between the speed and number of objects being tracked, however, can only arise from the allocation of a flexible cognitive resource, which can be formalized as either memory or attention.
Ed Vul, Michael C. Frank, George A. Alvarez, Josh Tenenbaum
NIPS3