VLDB 2026 Research / reviewers in the wild / expert
Stephanie Fu
dblp:270/1541
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2024
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
3D vision · 24% Trustworthy machine learning · 17% Language models and text generation · 14% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 100% | |
| Computer graphics and multimedia
1 paper |
Image and video coding · 100% |
Topics — the 17 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
image retrieval |
0.8 | 2 | 2024 | Axiomatic Explanations for Visual Search, Retrieval, and Similarity Learning · ICLR 2022 OpenStreetView-5M: The Many Roads to Global Visual Geolocation · CVPR 2024 |
Machine learning › Transfer learning and domain adaptation
cross-task transfer |
0.8 | 1 | 2024 | When does perceptual alignment benefit vision representations? · NeurIPS 2024 |
Machine learning › Deep learning architectures and training › neural network layer design
feature upsampling |
0.8 | 1 | 2024 | FeatUp: A Model-Agnostic Framework for Features at Any Resolution · ICLR 2024 |
Natural language and speech › Language models and text generation › alignment
human-model alignment |
0.8 | 1 | 2024 | Evaluating Multiview Object Consistency in Humans and Image Models · NeurIPS 2024 |
Computer vision › 3D vision › visual localization › geo-localization
image geo-localization |
0.8 | 1 | 2024 | OpenStreetView-5M: The Many Roads to Global Visual Geolocation · CVPR 2024 |
Computer vision › 3D vision › implicit neural representation
neural field |
0.8 | 1 | 2024 | FeatUp: A Model-Agnostic Framework for Features at Any Resolution · ICLR 2024 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
self-supervised visual representation learning |
0.8 | 1 | 2024 | A Vision Check-up for Language Models · CVPR 2024 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.8 | 1 | 2024 | FeatUp: A Model-Agnostic Framework for Features at Any Resolution · ICLR 2024 |
Machine learning › Representation and self-supervised learning › representation learning
visual representation learning |
0.8 | 1 | 2024 | When does perceptual alignment benefit vision representations? · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › interpretability
visual explanation |
0.7 | 1 | 2023 | DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data · NeurIPS 2023 |
Image and video coding › image quality assessment
perceptual similarity |
0.7 | 1 | 2023 | DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data · NeurIPS 2023 |
Machine learning › Trustworthy machine learning › interpretability › attribution methods
axiomatic attribution |
0.6 | 1 | 2022 | Axiomatic Explanations for Visual Search, Retrieval, and Similarity Learning · ICLR 2022 |
Machine learning › Trustworthy machine learning
interpretability |
0.6 | 1 | 2022 | Axiomatic Explanations for Visual Search, Retrieval, and Similarity Learning · ICLR 2022 |
Computer vision › 3D vision
depth estimation |
0.2 | 1 | 2024 | FeatUp: A Model-Agnostic Framework for Features at Any Resolution · ICLR 2024 |
Information retrieval › image retrieval
large-scale image retrieval |
0.2 | 1 | 2024 | OpenStreetView-5M: The Many Roads to Global Visual Geolocation · CVPR 2024 |
Usability and user experience research › experimental design
behavioral experiments |
0.2 | 1 | 2024 | Evaluating Multiview Object Consistency in Humans and Image Models · NeurIPS 2024 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.2 | 1 | 2023 | DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
spatial representation learning · 1.5multi-scale evaluation · 1.5image encoder benchmarking · 1.5triplet learning · 0.8multi-view consistency loss · 0.8large language model · 0.8implicit neural representation · 0.8human similarity judgment finetuning · 0.8code-based image representation · 0.8behavioral experiments · 0.8behavioral experiment · 0.8synthetic data generation · 0.7human similarity judgments · 0.7axiomatic attribution · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | OpenStreetView-5M: The Many Roads to Global Visual GeolocationabstractDetermining the location of an image anywhere on Earth is a complex visual task, which makes it particularly relevant for evaluating computer vision algorithms. Yet, the absence of standard, large-scale, open-access datasets with reliably localizable images has limited its potential. To address this issue, we introduce OpenStreetView-5M, a large-scale, open-access dataset comprising over 5.1 million georeferenced street view images, covering 225 countries and territories. In contrast to existing benchmarks, we enforce a strict train/test separation, allowing us to evaluate the relevance of learned geographical features beyond mere memorization. To demonstrate the utility of our dataset, we conduct an extensive benchmark of various state-of-the-art image encoders, spatial representations, and training strategies. All associated codes and models can be found at github.com/gastruc/osv5m. Guillaume Astruc, Nicolas Dufour, Ioannis Siglidis, Constantin Aronssohn, Nacim Bouia, Stephanie Fu, Romain Loiseau, Van Nguyen Nguyen, Charles Raude, Elliot Vincent, Lintao Xu, Loïc Landrieu |
CVPR | 6 |
| 2024 | A Vision Check-up for Language ModelsabstractWhat does learning to model relationships between strings teach Large Language Models (LLMs) about the visual world? We systematically evaluate LLMs' abilities to generate and recognize an assortment of visual concepts of increasing complexity and then demonstrate how a preliminary visual representation learning system can be trained using models of text. As language models lack the ability to consume or output visual information as pixels, we use code to represent images in our study. Although LLM-generated images do not look like natural images, results on image generation and the ability of models to correct these generated images indicate that precise modeling of strings can teach language models about numerous aspects of the visual world. Furthermore, experiments on self-supervised visual representation learning, utilizing images generated with text models, highlight the potential to train vision models capable of making semantic assessments of natural images using just LLMs. Pratyusha Sharma, Tamar Rott Shaham, Manel Baradad Jurjo, Adrián Rodríuez-Muñoz, Shivam Duggal, Phillip Isola, Antonio Torralba 0001, Stephanie Fu |
CVPR | 8 |
| 2024 | FeatUp: A Model-Agnostic Framework for Features at Any ResolutionabstractDeep features are a cornerstone of computer vision research, capturing image semantics and enabling the community to solve downstream tasks even in the zero- or few-shot regime. However, these features often lack the spatial resolution to directly perform dense prediction tasks like segmentation and depth prediction because models aggressively pool information over large areas. In this work, we introduce FeatUp, a task- and model-agnostic framework to restore lost spatial information in deep features. We introduce two variants of FeatUp: one that guides features with high-resolution signal in a single forward pass, and one that fits an implicit model to a single image to reconstruct features at any resolution. Both approaches use a multi-view consistency loss with deep analogies to NeRFs. Our features retain their original semantics and can be swapped into existing applications to yield resolution and performance gains even without re-training. We show that FeatUp significantly outperforms other feature upsampling and image super-resolution approaches in class activation map generation, transfer learning for segmentation and depth prediction, and end-to-end training for semantic segmentation. Stephanie Fu, Mark Hamilton, Laura E. Brandt, Axel Feldmann, Zhoutong Zhang, William T. Freeman |
ICLR | 1 |
| 2024 | Evaluating Multiview Object Consistency in Humans and Image ModelsabstractWe introduce a benchmark to directly evaluate the alignment between human observers and vision models on a 3D shape inference task. We leverage an experimental design from the cognitive sciences: given a set of images, participants identify which contain the same/different objects, despite considerable viewpoint variation. We draw from a diverse range of images that include common objects (e.g., chairs) as well as abstract shapes (i.e., procedurally generated 'nonsense' objects). After constructing over 2000 unique image sets, we administer these tasks to human participants, collecting 35K trials of behavioral data from over 500 participants. This includes explicit choice behaviors as well as intermediate measures, such as reaction time and gaze data. We then evaluate the performance of common vision models (e.g., DINOv2, MAE, CLIP). We find that humans outperform all models by a wide margin. Using a multi-scale evaluation approach, we identify underlying similarities and differences between models and humans: while human-model performance is correlated, humans allocate more time/processing on challenging trials. All images, data, and code can be accessed via our project page. Tyler Bonnen, Stephanie Fu, Yutong Bai, Thomas P. O'Connell, Yoni Friedman, Nancy Kanwisher, Josh Tenenbaum, Alexei A. Efros |
NeurIPS | 2 |
| 2024 | When does perceptual alignment benefit vision representations?abstractHumans judge perceptual similarity according to diverse visual attributes, including scene layout, subject location, and camera pose. Existing vision models understand a wide range of semantic abstractions but improperly weigh these attributes and thus make inferences misaligned with human perception.
While vision representations have previously benefited from human preference alignment in contexts like image generation, the utility of perceptually aligned representations in more general-purpose settings remains unclear. Here, we investigate how aligning vision model representations to human perceptual judgments impacts their usability in standard computer vision tasks. We finetune state-of-the-art models on a dataset of human similarity judgments for synthetic image triplets and evaluate them across diverse computer vision tasks. We find that aligning models to perceptual judgments yields representations that improve upon the original backbones across many downstream tasks, including counting, semantic segmentation, depth estimation, instance retrieval, and retrieval-augmented generation. In addition, we find that performance is widely preserved on other tasks, including specialized out-of-distribution domains such as in medical imaging and 3D environment frames. Our results suggest that injecting an inductive bias about human perceptual knowledge into vision models can make them better representation learners. Shobhita Sundaram, Stephanie Fu, Lukas Muttenthaler, Netanel Tamir, Lucy Chai, Simon Kornblith, Trevor Darrell, Phillip Isola |
NeurIPS | 2 |
| 2023 | DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic DataabstractCurrent perceptual similarity metrics operate at the level of pixels and patches. These metrics compare images in terms of their low-level colors and textures, but fail to capture mid-level similarities and differences in image layout, object pose, and semantic content. In this paper, we develop a perceptual metric that assesses images holistically. Our first step is to collect a new dataset of human similarity judgments over image pairs that are alike in diverse ways. Critical to this dataset is that judgments are nearly automatic and shared by all observers. To achieve this we use recent text-to-image models to create synthetic pairs that are perturbed along various dimensions. We observe that popular perceptual metrics fall short of explaining our new data, and we introduce a new metric, DreamSim, tuned to better align with human perception. We analyze how our metric is affected by different visual attributes, and find that it focuses heavily on foreground objects and semantic content while also being sensitive to color and layout. Notably, despite being trained on synthetic data, our metric generalizes to real images, giving strong results on retrieval and reconstruction tasks. Furthermore, our metric outperforms both prior learned metrics and recent large vision models on these tasks. Our project page: https://dreamsim-nights.github.io/ Stephanie Fu, Netanel Tamir, Shobhita Sundaram, Lucy Chai, Richard Zhang 0001, Tali Dekel, Phillip Isola |
NeurIPS | 1 |
| 2022 | Axiomatic Explanations for Visual Search, Retrieval, and Similarity Learning
Mark Hamilton, Scott M. Lundberg, Stephanie Fu, Lei Zhang 0001, William T. Freeman |
ICLR | 3 |