Jeffrey Gu

dblp:280/0994 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
3D vision · 26% Representation and self-supervised learning · 24% Transfer learning and domain adaptation · 8%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 18 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning › geometric representation learning
hyperbolic representation learning
1.322024
Hyperbolic Deep Learning in Computer Vision: A Survey · Int. J. Comput. Vis. 2024
Capturing implicit hierarchical structure in 3D biomedical images with self-supervised hyperbolic representations · NeurIPS 2021
Machine learning › Deep learning architectures and training
hypernetwork
0.912025
Foundation Models Secretly Understand Neural Network Weights: Enhancing Hypernetwork Architectures with Foundation Models · ICLR 2025
Computer vision › 3D vision
implicit neural representation
0.912025
Foundation Models Secretly Understand Neural Network Weights: Enhancing Hypernetwork Architectures with Foundation Models · ICLR 2025
Machine learning › Representation and self-supervised learning
pre-training
0.912025
BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature · CVPR 2025
Computer vision › Vision and language
vision-language model
0.912025
BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature · CVPR 2025
Computer vision › Image recognition and object detection
deep learning for vision
0.812024
Hyperbolic Deep Learning in Computer Vision: A Survey · Int. J. Comput. Vis. 2024
Computer vision › 3D vision
human mesh recovery
0.712023
NeMo: 3D Neural Motion Fields from Multiple Video Instances of the Same Action · CVPR 2023
Computer vision › 3D vision › 3d motion analysis
human motion reconstruction
0.712023
NeMo: 3D Neural Motion Fields from Multiple Video Instances of the Same Action · CVPR 2023
Computer vision › Face, body and person analysis
human pose estimation
0.712023
NeMo: 3D Neural Motion Fields from Multiple Video Instances of the Same Action · CVPR 2023
Machine learning › Transfer learning and domain adaptation
meta-learning
0.712023
Generalizable Neural Fields as Partially Observed Neural Processes · ICCV 2023
Computer vision › 3D vision › implicit neural representation
neural field
0.712023
Generalizable Neural Fields as Partially Observed Neural Processes · ICCV 2023
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
neural processes
0.712023
Generalizable Neural Fields as Partially Observed Neural Processes · ICCV 2023
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
0.512021
Capturing implicit hierarchical structure in 3D biomedical images with self-supervised hyperbolic representations · NeurIPS 2021
Machine learning › Generative modeling
variational autoencoder
0.512021
Capturing implicit hierarchical structure in 3D biomedical images with self-supervised hyperbolic representations · NeurIPS 2021
Machine learning › Transfer learning and domain adaptation
foundation model adaptation
0.312025
Foundation Models Secretly Understand Neural Network Weights: Enhancing Hypernetwork Architectures with Foundation Models · ICLR 2025
Medical and health informatics › medical imaging
medical image analysis
0.312025
BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature · CVPR 2025
Computer vision › Segmentation and scene understanding
biomedical image segmentation
0.112021
Capturing implicit hierarchical structure in 3D biomedical images with self-supervised hyperbolic representations · NeurIPS 2021
Computer vision › Segmentation and scene understanding › image segmentation
unsupervised segmentation
0.112021
Capturing implicit hierarchical structure in 3D biomedical images with self-supervised hyperbolic representations · NeurIPS 2021

Methods — techniques the papers use, named apart from their topics

streaming pretraining · 1.7contrastive learning · 1.7transformer · 0.9foundation model · 0.9hyperbolic geometry · 0.8neural radiance field · 0.7monocular human mesh recovery · 0.7hypernetwork · 0.7gradient-based meta-learning · 0.7gyroplane convolutional layer · 0.5
YearPublicationVenuePosition
2025 BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature
abstract
The development of vision-language models (VLMs) is driven by large-scale and diverse multi-modal datasets. However, progress toward generalist biomedical VLMs is limited by the lack of annotated, publicly accessible datasets across biology and medicine. Existing efforts are limited to narrow domains, missing the full diversity of biomedical knowledge encoded in scientific literature. To address this gap, we introduce BIOMEDICA: a scalable, open-source framework to extract, annotate, and serialize the entirety of the PubMed Central Open Access subset into an easy-to-use, publicly accessible dataset. Our framework produces a comprehensive archive with over 24 million unique image-text pairs from over 6 million articles. Metadata and expert-guided annotations are additionally provided.We demonstrate the utility and accessibility of our resource by releasing BMC-CLIP, a suite of CLIP-style models continuously pre-trained on BIOMEDICA dataset via streaming (eliminating the need to download 27 TB of data locally). On average, our models achieve state-of-the-art performance across 40 tasks — spanning pathology, radiology, ophthalmology, dermatology, surgery, molecular biology, parasitology, and cell biology — excelling in zero-shot classification with 6.56% average improvement (as high as 29.8% and 17.5% in dermatology and ophthalmology, respectively) and stronger image-text retrieval while using 10x less compute. To foster reproducibility and collaboration, we release our codebase1,2and dataset3to the broader research community.
Alejandro Lozano, Min Woo Sun, James Burgess, Liangyu Chen 0005, Jeffrey J. Nirschl, Jeffrey Gu, Iván López 0001, Josiah Aklilu, Anita Rau, Austin Wolfgang Katzer, Collin Chiu, Alfred Seunghoon Song, Robert Tibshirani, Serena Yeung-Levy
CVPR6
2025 Foundation Models Secretly Understand Neural Network Weights: Enhancing Hypernetwork Architectures with Foundation Models
abstract
Large pre-trained models, or foundation models, have shown impressive performance when adapted to a variety of downstream tasks, often out-performing specialized models. Hypernetworks, neural networks that generate some or all of the parameters of another neural network, have become an increasingly important technique for conditioning and generalizing implicit neural representations (INRs), which represent signals or objects such as audio or 3D shapes using a neural network. However, despite the potential benefits of incorporating foundation models in hypernetwork methods, this research direction has not been investigated, likely due to the dissimilarity of the weight generation task with other visual tasks. To address this gap, we (1) show how foundation models can improve hypernetworks with Transformer-based architectures, (2) provide an empirical analysis of the benefits of foundation models for hypernetworks through the lens of the generalizable INR task, showing that leveraging foundation models improves performance, generalizability, and data efficiency across a variety of algorithms and modalities. We also provide further analysis in examining the design space of foundation model-based hypernetworks, including examining the choice of foundation models, algorithms, and the effect of scaling foundation models.
Jeffrey Gu, Serena Yeung-Levy
ICLR1
2024 Hyperbolic Deep Learning in Computer Vision: A Survey
abstract
Abstract Deep representation learning is a ubiquitous part of modern computer vision. While Euclidean space has been the de facto standard manifold for learning visual representations, hyperbolic space has recently gained rapid traction for learning in computer vision. Specifically, hyperbolic learning has shown a strong potential to embed hierarchical structures, learn from limited samples, quantify uncertainty, add robustness, limit error severity, and more. In this paper, we provide a categorization and in-depth overview of current literature on hyperbolic learning for computer vision. We research both supervised and unsupervised literature and identify three main research themes in each direction. We outline how hyperbolic learning is performed in all themes and discuss the main research problems that benefit from current advances in hyperbolic learning for computer vision. Moreover, we provide a high-level intuition behind hyperbolic geometry and outline open research questions to further advance research in this direction.
Pascal Mettes, Mina Ghadimi Atigh, Martin Keller-Ressel, Jeffrey Gu, Serena Yeung-Levy
Int. J. Comput. Vis.4
2023 NeMo: 3D Neural Motion Fields from Multiple Video Instances of the Same Action
abstract
The task of reconstructing 3D human motion has wide-ranging applications. The gold standard Motion capture (MoCap) systems are accurate but inaccessible to the general public due to their cost, hardware, and space constraints. In contrast, monocular human mesh recovery (HMR) methods are much more accessible than MoCap as they take single-view videos as inputs. Replacing the multi-view MoCap systems with a monocular HMR method would break the current barriers to collecting accurate 3D motion thus making exciting applications like motion analysis and motion-driven animation accessible to the general public. However, the performance of existing HMR methods degrades when the video contains challenging and dynamic motion that is not in existing MoCap datasets used for training. This reduces its appeal as dynamic motion is frequently the target in 3D motion recovery in the aforementioned applications. Our study aims to bridge the gap between monocular HMR and multi-view MoCap systems by leveraging information shared across multiple video instances of the same action. We introduce the Neural Motion (NeMo) field. It is optimized to represent the underlying 3D motions across a set of videos of the same action. Empirically, we show that NeMo can recover 3D motion in sports using videos from the Penn Action dataset, where NeMo outperforms existing HMR methods in terms of 2D keypoint detection. To further validate NeMo using 3D metrics, we collected a small MoCap dataset mimicking actions in Penn Action, and show that NeMo achieves better 3D reconstruction compared to various baselines.
Kuan-Chieh Wang, Zhenzhen Weng, Maria Xenochristou, João Pedro Araújo 0001, Jeffrey Gu, C. Karen Liu, Serena Yeung-Levy
CVPR5
2023 Generalizable Neural Fields as Partially Observed Neural Processes
abstract
Neural fields, which represent signals as a function parameterized by a neural network, are a promising alternative to traditional discrete vector or grid-based representations. Compared to discrete representations, neural representations both scale well with increasing resolution, are continuous, and can be many-times differentiable. However, given a dataset of signals that we would like to represent, having to optimize a separate neural field for each signal is inefficient, and cannot capitalize on shared information or structures among signals. Existing generalization methods view this as a meta-learning problem and employ gradient-based meta-learning to learn an initialization which is then fine-tuned with test-time optimization, or learn hypernetworks to produce the weights of a neural field. We instead propose a new paradigm that views the large-scale training of neural representations as a part of a partially-observed neural process framework, and leverage neural process algorithms to solve this task. We demonstrate that this approach outperforms both state-of-the-art gradient-based meta-learning approaches and hypernetwork approaches.
Jeffrey Gu, Kuan-Chieh Wang, Serena Yeung-Levy
ICCV1
2021 Capturing implicit hierarchical structure in 3D biomedical images with self-supervised hyperbolic representations
abstract
We consider the task of representation learning for unsupervised segmentation of 3D voxel-grid biomedical images. We show that models that capture implicit hierarchical relationships between subvolumes are better suited for this task. To that end, we consider encoder-decoder architectures with a hyperbolic latent space, to explicitly capture hierarchical relationships present in subvolumes of the data. We propose utilizing a 3D hyperbolic variational autoencoder with a novel gyroplane convolutional layer to map from the embedding space back to 3D images. To capture these relationships, we introduce an essential self-supervised loss---in addition to the standard VAE loss---which infers approximate hierarchies and encourages implicitly related subvolumes to be mapped closer in the embedding space. We present experiments on synthetic datasets along with a dataset from the medical domain to validate our hypothesis.
Joy Hsu, Jeffrey Gu, Gong-Her Wu, Wah Chiu, Serena Yeung-Levy
NeurIPS2
2021 Staying in shape: learning invariant shape representations using contrastive learning
abstract
Creating representations of shapes that are invariant to isometric or almost-isometric transformations has long been an area of interest in shape analysis, since enforcing invariance allows the learning of more effective and robust shape representations. Most existing invariant shape representations are handcrafted, and previous work on learning shape representations do not focus on producing invariant representations. To solve the problem of learning unsupervised invariant shape representations, we use contrastive learning, which produces discriminative representations through learning invariance to user-specified data augmentations. To produce representations that are specifically isometry and almost-isometry invariant, we propose new data augmentations that randomly sample these transformations. We show experimentally that our method outperforms previous unsupervised learning approaches in both effectiveness and robustness.
Jeffrey Gu, Serena Yeung-Levy
UAI1