EDBT 2026 Demo / reviewers in the wild / expert
David Fan 0001
dblp:03/3063-1
· DBLP profile ↗
10ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-9217-5451ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scaling Language-Free Visual Representation LearningabstractVisual Self-Supervised Learning (SSL) currently underperforms Contrastive Language-Image Pretraining (CLIP) in multimodal settings such as Visual Question Answering (VQA). This multimodal gap is often attributed to the semantics introduced by language supervision, even though visual SSL and CLIP models are often trained on different data. In this work, we ask the question: "Do visual self-supervised approaches lag behind CLIP due to the lack of language supervision, or differences in the training data?" We study this question by training both visual SSL and CLIP models on the same MetaCLIP data, and leveraging VQA as a diverse testbed for vision encoders. In this controlled setup, visual SSL models scale better than CLIP models in terms of data and model capacity, and visual SSL performance does not saturate even after scaling up to 7B parameters. Consequently, we observe visual SSL methods achieve CLIP-level performance on a wide range of VQA and classic vision benchmarks. These findings demonstrate that pure visual SSL can match language-supervised visual pretraining at scale, opening new opportunities for vision-centric representation learning. David Fan 0001, Shengbang Tong, Jiachen Zhu 0002, Koustuv Sinha, Zhuang Liu 0003, Xinlei Chen, Michael G. Rabbat, Nicolas Ballas, Yann LeCun, Amir Bar, Saining Xie |
ICCV | 1 |
| 2025 | MetaMorph: Multimodal Understanding and Generation via Instruction TuningabstractIn this work, we propose Visual-Predictive Instruction Tuning (VPiT) - a simple and effective extension to visual instruction tuning that enables a pretrained LLM to quickly morph into an unified autoregressive model capable of generating both text and visual tokens. VPiT teaches an LLM to predict discrete text tokens and continuous visual tokens from any input sequence of image and text data curated in an instruction-following format. Our empirical investigation reveals several intriguing properties of VPiT: (1) visual generation ability emerges as a natural byproduct of improved visual understanding, and can be unlocked efficiently with a small amount of generation data; (2) while we find understanding and generation to be mutually beneficial, understanding data contributes to both capabilities more effectively than generation data. Building upon these findings, we train our MetaMorph model and achieve competitive performance on both visual understanding and generation. In visual generation, MetaMorph can leverage the world knowledge and reasoning abilities gained from LLM pretraining, and overcome common failure modes exhibited by other generation models. Our results suggest that LLMs may have strong "prior" vision capabilities that can be efficiently adapted to both visual understanding and generation with a relatively simple instruction tuning process. Shengbang Tong, David Fan 0001, Jiachen Zhu 0002, Yunyang Xiong, Xinlei Chen, Koustuv Sinha, Michael G. Rabbat, Yann LeCun, Saining Xie, Zhuang Liu 0003 |
ICCV | 2 |
| 2025 | Now you see Me: Context-Aware Automatic Audio DescriptionabstractAudio Description (AD) plays a pivotal role as an application system aimed at guaranteeing accessibility in multimedia content, which provides additional narrations at suitable intervals to describe visual elements, catering specifically to the needs of visually impaired audiences. In this paper, we introduce CA3D, the pioneering unified ContextAware Automatic Audio Description system that provides AD event scripts with precise locations in the long cinematic content. Specifically, CA3D system consists of: 1) a Temporal Feature Enhancement Module to efficiently capture longer term dependencies, 2) an anchor-based AD event detector with feature suppression module that localizes the AD events and extracts discriminative feature for AD generation, and 3) a self-refinement module that leverages the generated output to tweak AD event boundaries from coarse to fine. Unlike conventional methods which rely on metadata and ground truth AD timestamp for AD detection and generation tasks, the proposed CA3D is the first end-to-end trainable system that only uses visual cue. Extensive experiments demonstrate that the proposed CA3D improves existing architectures for both AD event detection and script generation metrics, establishing the new state-of-the-art performances in the AD automation. Seon-Ho Lee, Jue Wang 0010, David Fan 0001, Linda Liu, Vimal Bhat, Xinyu Li 0003 |
WACV | 3 |
| 2025 | GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-Grained Video-Language LearningabstractIn various video-language learning tasks, the challenge of achieving cross-modality alignment with multi-grained data persists. We propose a method to tackle this challenge from two crucial perspectives: data and modeling. Given the absence of a multi-grained video-text pretraining dataset, we introduce a Granularity EXpansion (GEX) method with Integration and Compression operations to expand the granularity of a single-grained dataset. To better model multi-grained data, we introduce an Iterative Approximation Module (IAM), which embeds multi-grained videos and texts into a unified, low-dimensional semantic space while preserving essential information for cross-modal alignment. Furthermore, GEXIA is highly scalable with no restrictions on the number of video-text granularities for alignment. We evaluate our work on three categories of video tasks across seven benchmark datasets, showcasing state-of-the-art or comparable performance. Remarkably, our model excels in tasks involving long-form video understanding, even though the pretraining dataset only contains short video clips. Jue Wang 0010, David Fan 0001, Zhenlin Xu, Linda Liu, Vimal Bhat, Xinyu Li 0003 |
WACV | 4 |
| 2024 | Text-Guided Video Masked Autoencoder
David Fan 0001, Jue Wang 0010, Shuai Liao, Vimal Bhat, Xinyu Li 0003 |
ECCV (5) | 1 |
| 2024 | Video Token Merging for Long Video UnderstandingabstractAs the scale of data and models for video understanding rapidly expand, handling long-form video input in transformer-based models presents a practical challenge. Rather than resorting to input sampling or token dropping, which may result in information loss, token merging shows promising results when used in collaboration with transformers. However, the application of token merging for long-form video processing is not trivial. We begin with the premise that token merging should not rely solely on the similarity of video tokens; the saliency of tokens should also be considered. To address this, we explore various video token merging strategies for long-form video classification, starting with a simple extension of image token merging, moving to region-concentrated merging, and finally proposing a learnable video token merging (VTM) algorithm that dynamically merges tokens based on their saliency. Extensive experimental results show that we achieve better or comparable performances on the LVU, COIN, and Breakfast datasets. Moreover, our approach significantly reduces memory costs by 84% and boosts throughput by approximately 6.89 times compared to baseline algorithms. Seon-Ho Lee, Jue Wang 0010, David Fan 0001, Xinyu Li 0003 |
NeurIPS | 4 |
| 2023 | Motion-Guided Masking for Spatiotemporal Representation LearningabstractSeveral recent works have directly extended the image masked autoencoder (MAE) with random masking into video domain, achieving promising results. However, unlike images, both spatial and temporal information are important for video understanding. This suggests that the random masking strategy that is inherited from the image MAE is less effective for video MAE. This motivates the design of a novel masking algorithm that can more efficiently make use of video saliency. Specifically, we propose a motion-guided masking algorithm (MGM) which leverages motion vectors to guide the position of each mask over time. Crucially, these motion-based correspondences can be directly obtained from information stored in the compressed format of the video, which makes our method efficient and scalable. On two challenging large-scale video benchmarks (Kinetics-400 and Something-Something V2), we equip video MAE with our MGM and achieve up to +1.3% improvement compared to previous state-of-the-art methods. Additionally, our MGM achieves equivalent performance to previous video MAE using up to 66% fewer training epochs. Lastly, we show that MGM generalizes better to downstream transfer learning and domain adaptation tasks on the UCF101, HMDB51, and Diving48 datasets, achieving up to +4.9% improvement compared to baseline methods. David Fan 0001, Jue Wang 0010, Shuai Liao, Yi Zhu 0001, Vimal Bhat, Hector J. Santos-Villalobos, Rohith MV, Xinyu Li 0003 |
ICCV | 1 |
| 2023 | MEGA: Multimodal Alignment Aggregation and Distillation For Cinematic Video SegmentationabstractPrevious research has studied the task of segmenting cinematic videos into scenes and into narrative acts. However, these studies have overlooked the essential task of multimodal alignment and fusion for effectively and efficiently processing long-form videos (> 60min). In this paper, we introduce Multimodal alignmEnt aGgregation and distillAtion (MEGA) for cinematic long-video segmentation. MEGA tackles the challenge by leveraging multiple media modalities. The method coarsely aligns inputs of variable lengths and different modalities with alignment positional encoding. To maintain temporal synchronization while reducing computation, we further introduce an enhanced bottleneck fusion layer which uses temporal alignment. Additionally, MEGA employs a novel contrastive loss to synchronize and transfer labels across modalities, enabling act segmentation from labeled synopsis sentences on video shots. Our experimental results show that MEGA outperforms state-of-the-art methods on MovieNet dataset for scene segmentation (with an Average Precision improvement of +1.19%) and on TRIPOD dataset for act segmentation (with a Total Agreement improvement of +5.51%). Najmeh Sadoughi, Xinyu Li 0003, Avijit Vajpayee, David Fan 0001, Bing Shuai, Hector J. Santos-Villalobos, Vimal Bhat, Rohith MV |
ICCV | 4 |
| 2021 | Shot Contrastive Self-Supervised Learning for Scene Boundary DetectionabstractScenes play a crucial role in breaking the storyline of movies and TV episodes into semantically cohesive parts. However, given their complex temporal structure, finding scene boundaries can be a challenging task requiring large amounts of labeled training data. To address this challenge, we present a self-supervised shot contrastive learning approach (ShotCoL) to learn a shot representation that maximizes the similarity between nearby shots compared to randomly selected shots. We show how to apply our learned shot representation for the task of scene boundary detection to offer state-of-the-art performance on the MovieNet [33] dataset while requiring only ~25% of the training labels, using 9× fewer model parameters and offering 7× faster runtime. To assess the effectiveness of ShotCoL on novel applications of scene boundary detection, we take on the problem of finding timestamps in movies and TV episodes where video-ads can be inserted while offering a minimally disruptive viewing experience. To this end, we collected a new dataset called AdCuepoints with 3, 975 movies and TV episodes, 2.2 million shots and 19, 119 minimally disruptive ad cue-point labels. We present a thorough empirical analysis on this dataset demonstrating the effectiveness of ShotCoL for ad cue-points detection. Shixing Chen, Xiaohan Nie, David Fan 0001, Dongqing Zhang, Vimal Bhat, Raffay Hamid |
CVPR | 3 |
| 2020 | OASIS: A Large-Scale Dataset for Single Image 3D in the WildabstractSingle-view 3D is the task of recovering 3D properties such as depth and surface normals from a single image. We hypothesize that a major obstacle to single-image 3D is data. We address this issue by presenting Open Annotations of Single Image Surfaces (OASIS), a dataset for single-image 3D in the wild consisting of annotations of detailed 3D geometry for 140,000 images. We train and evaluate leading models on a variety of single-image 3D tasks. We expect OASIS to be a useful resource for 3D vision research. Project site: https://pvl.cs.princeton.edu/OASIS. Shengyi Qian 0001, David Fan 0001, Noriyuki Kojima, Max Hamilton, Jia Deng 0001 |
CVPR | 3 |