VLDB 2026 Research / reviewers in the wild / expert
Aaron Adcock
dblp:133/2099 · also Aaron B. Adcock
· DBLP profile ↗
8ranked-venue papers
2as first author
5since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Image recognition and object detection · 30% Representation and self-supervised learning · 29% Trustworthy machine learning · 23% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 82% Web and social media mining · 18% |
Topics — the 14 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
fairness |
1.3 | 2 | 2023 | GeoDE: a Geographically Diverse Evaluation Dataset for Object Recognition · NeurIPS 2023 FACET: Fairness in Computer Vision Evaluation Benchmark · ICCV 2023 |
Computer vision › Image recognition and object detection
visual recognition |
1.2 | 2 | 2023 | The effectiveness of MAE pre-pretraining for billion-scale pretraining · ICCV 2023 Revisiting Weakly Supervised Pre-Training of Visual Perception Models · CVPR 2022 |
Machine learning › Representation and self-supervised learning › pre-training
foundation model pretraining |
0.7 | 1 | 2023 | The effectiveness of MAE pre-pretraining for billion-scale pretraining · ICCV 2023 |
Machine learning › Trustworthy machine learning › dataset bias
geographic bias |
0.7 | 1 | 2023 | GeoDE: a Geographically Diverse Evaluation Dataset for Object Recognition · NeurIPS 2023 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked autoencoder |
0.7 | 1 | 2023 | The effectiveness of MAE pre-pretraining for billion-scale pretraining · ICCV 2023 |
Computer vision › Image recognition and object detection
object recognition |
0.7 | 1 | 2023 | GeoDE: a Geographically Diverse Evaluation Dataset for Object Recognition · NeurIPS 2023 |
Machine learning › Representation and self-supervised learning
pre-training |
0.6 | 1 | 2022 | Revisiting Weakly Supervised Pre-Training of Visual Perception Models · CVPR 2022 |
Machine learning › Representation and self-supervised learning › pre-training
weakly supervised pre-training |
0.6 | 1 | 2022 | Revisiting Weakly Supervised Pre-Training of Visual Perception Models · CVPR 2022 |
Machine learning › Transfer learning and domain adaptation
zero-shot transfer |
0.6 | 1 | 2022 | Revisiting Weakly Supervised Pre-Training of Visual Perception Models · CVPR 2022 |
Machine learning › Efficient and distributed learning
hardware acceleration |
0.5 | 1 | 2021 | PyTorchVideo: A Deep Learning Library for Video Understanding · ACM Multimedia 2021 |
Data mining › structured data mining › graph mining › network structure analysis
core-periphery structure |
0.2 | 1 | 2013 | Tree-Like Structure in Large Social and Information Networks · ICDM 2013 |
Data mining › structured data mining
graph mining |
0.2 | 1 | 2013 | Tree-Like Structure in Large Social and Information Networks · ICDM 2013 |
Web and social media mining
social network analysis |
0.1 | 1 | 2016 | Social Hash: An Assignment Framework for Optimizing Distributed Systems Operations on Social Networks · NSDI 2016 |
Graph algorithms and graph theory › graph decomposition
tree decomposition |
0.0 | 1 | 2013 | Tree-Like Structure in Large Social and Information Networks · ICDM 2013 |
Methods — techniques the papers use, named apart from their topics
self-supervised pretraining · 0.7masked autoencoding · 0.7intersectional analysis · 0.7dataset collection · 0.7self-supervised learning · 0.6residual network · 0.6hashtag supervision · 0.6multimodal data loading · 0.5deep learning library · 0.5k-core decomposition · 0.3hyperbolicity · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | FACET: Fairness in Computer Vision Evaluation BenchmarkabstractComputer vision models have known performance disparities across attributes such as gender and skin tone. This means during tasks such as classification and detection, model performance differs for certain classes based on the demographics of the people in the image. These disparities have been shown to exist, but until now there has not been a unified approach to measure these differences for common use-cases of computer vision models. We present a new benchmark named FACET (FAirness in Computer Vision EvaluaTion), a large, publicly available evaluation set of 32k images for some of the most common vision tasks - image classification, object detection and segmentation. For every image in FACET, we hired expert reviewers to manually annotate person-related attributes such as perceived skin tone and hair type, manually draw bounding boxes and label fine-grained person-related classes such as disk jockey or guitarist. In addition, we use FACET to benchmark state-of-the-art vision models and present a deeper understanding of potential performance disparities and challenges across sensitive demographic attributes. With the exhaustive annotations collected, we probe models using single demographics attributes as well as multiple attributes using an intersectional approach (e.g. hair color and perceived skin tone). Our results show that classification, detection, segmentation, and visual grounding models exhibit performance disparities across demographic attributes and intersections of attributes. These harms suggest that not all people represented in datasets receive fair and equitable treatment in these vision tasks. We hope current and future results using our benchmark will contribute to fairer, more robust vision models. FACET is available publicly at https://facet.metademolab.com. Laura Gustafson, Chloé Rolland, Nikhila Ravi, Quentin Duval, Aaron Adcock, Cheng-Yang Fu, Melissa Hall, Candace Ross |
ICCV | 5 |
| 2023 | The effectiveness of MAE pre-pretraining for billion-scale pretrainingabstractThis paper revisits the standard pretrain-then-finetune paradigm used in computer vision for visual recognition tasks. Typically, state-of-the-art foundation models are pretrained using large scale (weakly) supervised datasets with billions of images. We introduce an additional pre-pretraining stage that is simple and uses the self-supervised MAE technique to initialize the model. While MAE has only been shown to scale with the size of models, we find that it scales with the size of the training dataset as well. Thus, our MAE-based pre-pretraining scales with both model and data size making it applicable for training foundation models. Pre-pretraining consistently improves both the model convergence and the downstream transfer performance across a range of model scales (millions to billions of parameters), and dataset sizes (millions to billions of images). We measure the effectiveness of pre-pretraining on 10 different visual recognition tasks spanning image classification, video recognition, object detection, low-shot classification and zero-shot recognition. Our largest model achieves new state-of-the-art results on iNaturalist-18 (91.3%), 1-shot ImageNet-1k (62.1%), and zero-shot transfer on Food-101 (96.2%). Our study reveals that model initialization plays a significant role, even for web-scale pretraining with billions of images. Mannat Singh, Quentin Duval, Kalyan Vasudev Alwala, Haoqi Fan 0001, Vaibhav Aggarwal, Aaron Adcock, Armand Joulin, Piotr Dollár, Christoph Feichtenhofer, Ross B. Girshick, Rohit Girdhar, Ishan Misra |
ICCV | 6 |
| 2023 | GeoDE: a Geographically Diverse Evaluation Dataset for Object RecognitionabstractCurrent dataset collection methods typically scrape large amounts of data from the web. While this technique is extremely scalable, data collected in this way tends to reinforce stereotypical biases, can contain personally identifiable information, and typically originates from Europe and North America. In this work, we rethink the dataset collection paradigm and introduce GeoDE, a geographically diverse dataset with 61,940 images from 40 classes and 6 world regions, and no personally identifiable information, collected by soliciting images from people across the world. We analyse GeoDE to understand differences in images collected in this manner compared to web-scraping. Despite the smaller size of this dataset, we demonstrate its use as both an evaluation and training dataset, allowing us to highlight shortcomings in current models, as well as demonstrate improved performance even when training on this small dataset. We release the full dataset and code at https://geodiverse-data-collection.cs.princeton.edu/ Vikram V. Ramaswamy, Sing Yu Lin, Dora Zhao, Aaron Adcock, Laurens van der Maaten, Deepti Ghadiyaram, Olga Russakovsky |
NeurIPS | 4 |
| 2022 | Revisiting Weakly Supervised Pre-Training of Visual Perception ModelsabstractModel pre-training is a cornerstone of modern visual recognition systems. Although fully supervised pre-training on datasets like ImageNet is still the de-facto standard, recent studies suggest that large-scale weakly supervised pretraining can outperform fully supervised approaches. This paper revisits weakly-supervised pre-training of models using hashtag supervision with modern versions of residual networks and the largest-ever dataset of images and corresponding hashtags. We study the performance of the resulting models in various transfer-learning settings including zero-shot transfer. We also compare our models with those obtained via large-scale self-supervised learning. We find our weakly-supervised models to be very competitive across all settings, and find they substantially outperform their self-supervised counterparts. We also include an investigation into whether our models learned potentially troubling associations or stereotypes. Overall, our results provide a compelling argument for the use of weakly supervised learning in the development of visual recognition systems. Our models, Supervised Weakly through hashtAGs (SWAG), are available publicly. Mannat Singh, Laura Gustafson, Aaron Adcock, Vinicius de Freitas Reis, Bugra Gedik, Raj Prateek Kosaraju, Dhruv Mahajan 0001, Ross B. Girshick, Piotr Dollár, Laurens van der Maaten |
CVPR | 3 |
| 2021 | PyTorchVideo: A Deep Learning Library for Video UnderstandingabstractWe introduce PyTorchVideo, an open-source deep-learning library that provides a rich set of modular, efficient, and reproducible components for a variety of video understanding tasks, including classification, detection, self-supervised learning, and low-level processing. The library covers a full stack of video understanding tools including multimodal data loading, transformations, and models that reproduce state-of-the-art performance. PyTorchVideo further supports hardware acceleration that enables real-time inference on mobile devices. The library is based on PyTorch and can be used by any training framework; for example, PyTorchLightning, PySlowFast, or Classy Vision. PyTorchVideo is available at https://pytorchvideo.org/. Haoqi Fan 0001, Tullie Murrell, Kalyan Vasudev Alwala, Yanghao Li, Yilei Li, Nikhila Ravi, Meng Li 0004, Haichuan Yang, Jitendra Malik, Ross B. Girshick, Matt Feiszli, Aaron Adcock, Wan-Yen Lo, Christoph Feichtenhofer |
ACM Multimedia | 14 |
| 2016 | Social Hash: An Assignment Framework for Optimizing Distributed Systems Operations on Social Networks
Alon Shalita, Brian Karrer, Igor Kabiljo, Alessandro Presta, Aaron Adcock, Herald Kllapi, Michael Stumm |
NSDI | 6 |
| 2014 | Classification of hepatic lesions using the matching metric
Aaron Adcock, Daniel L. Rubin, Gunnar E. Carlsson |
Comput. Vis. Image Underst. | 1 |
| 2013 | Tree-Like Structure in Large Social and Information NetworksabstractAlthough large social and information networks are often thought of as having hierarchical or tree-like structure, this assumption is rarely tested. We have performed a detailed empirical analysis of the tree-like properties of realistic informatics graphs using two very different notions of tree-likeness: Gromov's d-hyperbolicity, which is a notion from geometric group theory that measures how tree-like a graph is in terms of its metric structure, and tree decompositions, tools from structural graph theory which measure how tree-like a graph is in terms of its cut structure. Although realistic informatics graphs often do not have meaningful tree-like structure when viewed with respect to the simplest and most popular metrics, e.g., the value of d or the tree width, we conclude that many such graphs do have meaningful tree-like structure when viewed with respect to more refined metrics, e.g., a size-resolved notion of d or a closer analysis of the tree decompositions. We also show that, although these two rigorous notions of tree-likeness capture very different tree-like structures in worst-case, for realistic informatics graphs they empirically identify surprisingly similar structure. We interpret this tree-like structure in terms of the recently-characterized "nested core-periphery" property of large informatics graphs, and we show that the fast and scalable k-core heuristic can be used to identify this tree-like structure. Aaron Adcock, Blair D. Sullivan, Michael W. Mahoney |
ICDM | 1 |