VLDB 2026 Research / reviewers in the wild / expert
Gene Chou
dblp:322/1261
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Generating 3D-Consistent Videos from Unposed Internet PhotosabstractWe address the problem of generating videos from unposed internet photos. A handful of input images serve as keyframes, and our model interpolates between them to simulate a path moving between the cameras. Given random images, a model’s ability to capture underlying geometry, recognize scene identity, and relate frames in terms of camera position and orientation reflects a fundamental understanding of 3D structure and scene layout. However, existing video models such as Luma Dream Machine fail at this task. We design a self-supervised method that takes advantage of the consistency of videos and variability of multiview internet photos to train a scalable, 3D-aware video model without any 3D annotations such as camera parameters. We validate that our method outperforms all baselines in terms of geometric and appearance consistency. We also show our model benefits applications that enable camera control, such as 3D Gaussian Splatting. Our results suggest that we can scale up scene-level 3D learning using only 2D data such as videos and multiview internet photos. Gene Chou, Kai Zhang 0045, Sai Bi, Hao Tan 0002, Zexiang Xu, Fujun Luan, Bharath Hariharan, Noah Snavely |
CVPR | 1 |
| 2025 | FlashDepth: Real-Time Streaming Video Depth Estimation at 2K ResolutionabstractA versatile video depth estimation model should (1) be accurate and consistent across frames, (2) produce high-resolution depth maps, and (3) support real-time streaming. We propose FlashDepth, a method that satisfies all three requirements, performing depth estimation on a 2044x1148 streaming video at 24 FPS. We show that, with careful modifications to pretrained single-image depth models, these capabilities are enabled with relatively little data and training. We evaluate our approach across multiple unseen datasets against state-of-the-art depth models, and find that ours outperforms them in terms of boundary sharpness and speed by a significant margin, while maintaining competitive accuracy. We hope our model will enable various applications that require high-resolution depth, such as video editing, and online decision-making, such as robotics. We release all code and model weights at https://github.com/Eyeline-Research/FlashDepth Gene Chou, Wenqi Xian, Guandao Yang, Mohamed Abdelfattah, Bharath Hariharan, Noah Snavely, Ning Yu 0006, Paul E. Debevec |
ICCV | 1 |
| 2025 | Generalist YOLO: Towards Real-Time End-to-End Multi-Task Visual Language ModelsabstractGeneralist models, capable of handling multiple modalities and tasks simultaneously, are currently one of the hottest research topics. However, due to interference between different tasks during the training process, existing generalist models require a very large decoder to achieve good results in various tasks, which makes real-time prediction difficult for current generalist models. This paper introduces Generalist YOLO, which takes a significant step towards real-time prediction systems for visual language generalist models. The proposed Generalist YOLO uses a unified encoder to reduce conflicts between different tasks, thereby decreasing the complexity required by the decoder. It also introduces a primary-secondary co-attention mechanism that allows different tasks to learn together more effectively, achieving high efficiency and high accuracy. We propose a semantically consistent asymmetric training strategy, allowing various tasks to benefit from performance improvements brought by the latest research results in various fields. The proposed Generalist YOLO achieves excellent results on various vision and language tasks based on MS COCO. While maintaining high accuracy across all tasks, it is 135 times faster than existing generalist models. The source code is released on GitHub at https://github.com/WongKinYiu/GeneralistYOLO. Hung-Shuo Chang, Chien-Yao Wang, Richard Robert Wang, Gene Chou, Hong-Yuan Mark Liao |
WACV | 4 |
| 2024 | MegaScenes: Scene-Level View Synthesis at Scale
Joseph Tung, Gene Chou, Ruojin Cai, Guandao Yang, Kai Zhang 0045, Gordon Wetzstein, Bharath Hariharan, Noah Snavely |
ECCV (29) | 2 |
| 2023 | Diffusion-SDF: Conditional Generative Modeling of Signed Distance FunctionsabstractProbabilistic diffusion models have achieved state-of-the-art results for image synthesis, inpainting, and text-to-image tasks. However, they are still in the early stages of generating complex 3D shapes. This work proposes Diffusion-SDF, a generative model for shape completion, single-view reconstruction, and reconstruction of real-scanned point clouds. We use neural signed distance functions (SDFs) as our 3D representation to parameterize the geometry of various signals (e.g., point clouds, 2D images) through neural networks. Neural SDFs are implicit functions and diffusing them amounts to learning the reversal of their neural network weights, which we solve using a custom modulation module. Extensive experiments show that our method is capable of both realistic unconditional generation and conditional generation from partial inputs. This work expands the domain of diffusion models from learning 2D, explicit representations, to 3D, implicit representations. Code is released at https://github.com/princeton-computational-imaging/Diffusion-SDF. Gene Chou, Yuval Bahat, Felix Heide |
ICCV | 1 |
| 2023 | Thin On-Sensor Nanophotonic Array CamerasabstractToday's commodity camera systems rely on compound optics to map light originating from the scene to positions on the sensor where it gets recorded as an image. To record images without optical aberrations, i.e., deviations from Gauss' linear model of optics, typical lens systems introduce increasingly complex stacks of optical elements which are responsible for the height of existing commodity cameras. In this work, we investigate flat nanophotonic computational cameras as an alternative that employs an array of skewed lenslets and a learned reconstruction approach. The optical array is embedded on a metasurface that, at 700 nm height, is flat and sits on the sensor cover glass at 2.5 mm focal distance from the sensor. To tackle the highly chromatic response of a metasurface and design the array over the entire sensor, we propose a differentiable optimization method that continuously samples over the visible spectrum and factorizes the optical modulation for different incident fields into individual lenses. We reconstruct a megapixel image from our flat imager with a learned probabilistic reconstruction method that employs a generative diffusion model to sample an implicit prior. To tackle scene-dependent aberrations in broadband , we propose a method for acquiring paired captured training data in varying illumination conditions. We assess the proposed flat camera design in simulation and with an experimental prototype, validating that the method is capable of recovering images from diverse scenes in broadband with a single nanophotonic layer. Praneeth Chakravarthula, Jipeng Sun, Chenyang Lei, Gene Chou, Mario Bijelic, Johannes Froesch, Arka Majumdar, Felix Heide |
ACM Trans. Graph. | 5 |
| 2022 | GenSDF: Two-Stage Learning of Generalizable Signed Distance FunctionsabstractWe investigate the generalization capabilities of neural signed distance functions (SDFs) for learning 3D object representations for unseen and unlabeled point clouds. Existing methods can fit SDFs to a handful of object classes and boast fine detail or fast inference speeds, but do not generalize well to unseen shapes. We introduce a two-stage semi-supervised meta-learning approach that transfers shape priors from labeled to unlabeled data to reconstruct unseen object categories. The first stage uses an episodic training scheme to simulate training on unlabeled data and meta-learns initial shape priors. The second stage then introduces unlabeled data with disjoint classes in a semi-supervised scheme to diversify these priors and achieve generalization. We assess our method on both synthetic data and real collected point clouds. Experimental results and analysis validate that our approach outperforms existing neural SDF methods and is capable of robust zero-shot inference on 100+ unseen classes. Code can be found at https://github.com/princeton-computational-imaging/gensdf Gene Chou, Ilya Chugunov, Felix Heide |
NeurIPS | 1 |