Brandon Yushan Feng

dblp:284/2193 · also Brandon Y. Feng · DBLP profile ↗
← Back
21ranked-venue papers
6as first author
20since 2021 · last 2025
0000-0001-7003-9128ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 17 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Repurposing Pre-trained Video Diffusion Models for Event-based Video Interpolation
abstract
Video Frame Interpolation aims to recover realistic missing frames between observed frames, generating a high-frame-rate video from a low-frame-rate video. However, without additional guidance, the large motion between frames makes this problem ill-posed. Event-based Video Frame Interpolation (EVFI) addresses this challenge by using sparse, high-temporal-resolution event measurements as motion guidance. This guidance allows EVFI methods to significantly outperform frame-only methods. However, to date, EVFI methods have relied on a limited set of paired event-frame training data, severely limiting their performance and generalization capabilities. In this work, we overcome the limited data challenge by adapting pre-trained video diffusion models trained on internet-scale datasets to EVFI. We experimentally validate our approach on real-world EVFI datasets, including a new one that we introduce. Our method outperforms existing methods and generalizes across cameras far better than existing approaches.
Jingxi Chen, Brandon Yushan Feng, Haoming Cai, Tianfu Wang 0007, Levi Burner, Dehao Yuan, Cornelia Fermüller, Christopher A. Metzler, Yiannis Aloimonos
CVPR2
2025 Parametric Shadow Control for Portrait Generation in Text-to-Image Diffusion Models
abstract
Text-to-image diffusion models excel at generating diverse portraits, but lack intuitive shadow control. Existing editing approaches, as post-processing, struggle to offer effective manipulation across diverse styles. Additionally, these methods either rely on expensive real-world light-stage data collection or require extensive computational resources for training. To address these limitations, we introduce Shadow Director, a method that extracts and manipulates hidden shadow attributes within well-trained diffusion models. Our approach uses a small estimation network that requires only a few thousand synthetic images and hours of training-no costly real-world light-stage data needed. Shadow Director enables parametric and intuitive control over shadow shape, placement, and intensity during portrait generation while preserving artistic integrity and identity across diverse styles. Despite training only on synthetic data built on real-world identities, it generalizes effectively to generated portraits with diverse styles, making it a more accessible and resource-friendly solution.
Haoming Cai, Tsung-Wei Huang, Shiv Gehlot, Brandon Yushan Feng, Sachin Shah, Guan-Ming Su, Christopher A. Metzler
ICCV4
2025 IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering
abstract
Vision-language models (VLMs) excel at descriptive tasks, but whether they truly understand scenes from visual observations remains uncertain. We introduce IR3D-Bench, a benchmark challenging VLMs to demonstrate understanding through active creation rather than passive recognition. Grounded in the analysis-by-synthesis paradigm, IR3D-Bench tasks Vision-Language Agents (VLAs) with actively using programming and rendering tools to recreate the underlying 3D structure of an input image, achieving agentic inverse rendering through tool use. This ''understanding-by-creating'' approach probes the tool-using generative capacity of VLAs, moving beyond the descriptive or conversational capacity measured by traditional scene understanding benchmarks. We provide a comprehensive suite of metrics to evaluate geometric accuracy, spatial relations, appearance attributes, and overall plausibility. Initial experiments on agentic inverse rendering powered by various state-of-the-art VLMs highlight current limitations, particularly in visual precision rather than basic tool usage. IR3D-Bench, including data and evaluation protocols, is released to facilitate systematic study and development of tool-using VLAs towards genuine scene understanding by creating.
Hengyu Liu 0007, Chenxin Li, Yipeng Wu, Wuyang Li, Zhiqin Yang, Zhenyuan Zhang 0001, Yunlong Lin, Sirui Han, Brandon Yushan Feng
NeurIPS10
2024 Learning to Estimate 6DoF Pose from Limited Data: A Few-Shot, Generalizable Approach using RGB Images
abstract
The accurate estimation of six degrees-of-freedom (6DoF) object poses is essential for many applications in robotics and augmented reality. However, existing methods for 6DoF pose estimation often depend on CAD templates or dense support views, restricting their usefulness in real-world situations. In this study, we present a new cascade framework named Cas6D for few-shot 6DoF pose estimation that is generalizable and uses only RGB images. To address the false positives of target object detection in the extreme few-shot setting, our framework utilizes a self-supervised pre-trained ViT to learn robust feature representations. Then, we initialize the nearest top-K pose candidates based on similarity score and refine the initial poses using feature pyramids to formulate and update the cascade warped feature volume, which encodes context at increasingly finer scales. By discretizing the pose search range using multiple pose bins and progressively narrowing the pose search range in each stage using predictions from the previous stage, Cas6D can overcome the large gap between pose candidates and ground truth poses, which is a common failure mode in sparse-view scenarios. Experimental results on the LINEMOD and GenMOP datasets demonstrate that Cas6D outperforms state-of-the-art methods by 9.2% and 3.8& accuracy (Proj-5) under the 32-shot setting compared to OnePose++ and Gen6D. Our framework also performs best under the full-shot setting with all support views. Code is avilable at https://github.com/github.com/paulpanwang/Cas6D.
Panwang Pan, Zhiwen Fan, Brandon Yushan Feng, Peihao Wang, Chenxin Li, Zhangyang Wang
3DV3
2024 Seeing the World through Your Eyes
abstract
The reflective nature of the human eye is an under-appreciated source of information about what the world around us looks like. By imaging the eyes of a moving person, we capture multiple views of a scene outside the camera's direct line of sight through the reflections in the eyes. In this paper, we reconstruct a radiance field beyond the camera's line of sight using portrait images containing eye reflections. This task is challenging due to 1) the difficulty of accurately estimating eye poses and 2) the entangled appearance of the iris textures and the scene reflections. To address these, our method jointly optimizes the cornea poses, the radiance field depicting the scene, and the observer's eye iris texture. We further present a regularization prior on the iris texture to improve scene reconstruction quality. Through various experiments on synthetic and real-world captures featuring people with varied eye colors, and lighting conditions, we demonstrate the feasibility of our approach to recover the radiance field using cornea reflections.
Hadi Alzayer, Kevin Zhang 0003, Brandon Yushan Feng, Christopher A. Metzler, Jia-Bin Huang 0001
CVPR3
2024 WaveMo: Learning Wavefront Modulations to See Through Scattering
abstract
Imaging through scattering media is a fundamental and pervasive challenge infields ranging from medical diagnos-tics to astronomy. A promising strategy to overcome this challenge is wavefront modulation, which induces measure-ment diversity during image acquisition. Despite its importance, designing optimal wavefront modulations to image through scattering remains under-explored. This paper in-troduces a novel learning-based framework to address the gap. Our approach jointly optimizes wavefront modulations and a computationally lightweight feedforward “proxy” re-construction network. This network is trained to recover scenes obscured by scattering, using measurements that are modified by these modulations. The learned modulations produced by our framework generalize effectively to un-seen scattering scenarios and exhibit remarkable versatility. During deployment, the learned modulations can be decou-pled from the proxy network to augment other more computationally expensive restoration algorithms. Through ex-tensive experiments, we demonstrate our approach signifi-cantly advances the state of the art in imaging through scat-tering media. Our project webpage is at https://wavemo-2024.github.io/.
Mingyang Xie, Haiyun Guo, Brandon Yushan Feng, Lingbo Jin, Ashok Veeraraghavan, Christopher A. Metzler
CVPR3
2024 Flash-Splat: 3D Reflection Removal with Flash Cues and Gaussian Splats
Mingyang Xie, Haoming Cai, Sachin Shah, Brandon Yushan Feng, Jia-Bin Huang 0001, Christopher A. Metzler
ECCV (82)5
2024 PhysDreamer: Physics-Based Interaction with 3D Objects via Video Generation
Hong-Xing Yu, Rundi Wu, Brandon Yushan Feng, Changxi Zheng, Noah Snavely, Jiajun Wu 0001, William T. Freeman
ECCV (2)4
2024 EndoSparse: Real-Time Sparse View Synthesis of Endoscopic Scenes using Gaussian Splatting
Chenxin Li, Brandon Yushan Feng, Yifan Liu 0010, Hengyu Liu 0007, Cheng Wang 0043, Weihao Yu 0005, Yixuan Yuan
MICCAI (6)2
2024 👦 Endora: Video Generation Models as Endoscopy Simulators
Chenxin Li, Hengyu Liu 0007, Yifan Liu 0010, Brandon Yushan Feng, Wuyang Li, Xinyu Liu 0001, Zhen Chen 0013, Yixuan Yuan
MICCAI (6)4
2024 Temporally Consistent Atmospheric Turbulence Mitigation with Neural Representations
abstract
Atmospheric turbulence, caused by random fluctuations in the atmosphere's refractive index, introduces complex spatio-temporal distortions in imagery captured at long range. Video Atmospheric Turbulence Mitigation (ATM) aims to restore videos affected by these distortions. However, existing video ATM methods, both supervised and self-supervised, struggle to maintain temporally consistent mitigation across frames, leading to visually incoherent results. This limitation arises from the stochastic nature of atmospheric turbulence, which varies across space and time. Inspired by the observation that atmospheric turbulence induces high-frequency temporal variations, we propose ConVRT, a novel framework for consistent video restoration through turbulence. ConVRT introduces a neural video representation that explicitly decouples spatial and temporal information into a spatial content field and a temporal deformation field, enabling targeted regularization of the network's temporal representation capability. By leveraging the low-pass filtering properties of the regularized temporal representations, ConVRT effectively mitigates turbulence-induced temporal frequency variations and promotes temporal consistency. Furthermore, our training framework seamlessly integrates supervised pre-training on synthetic turbulence data with self-supervised learning on real-world videos, significantly improving the temporally consistent mitigation of ATM methods on diverse real-world data. More information can be found on our project page: https://convrt-2024.github.io/
Haoming Cai, Jingxi Chen, Brandon Yushan Feng, Weiyun Jiang, Mingyang Xie, Kevin Zhang 0003, Cornelia Fermüller, Yiannis Aloimonos, Ashok Veeraraghavan, Christopher A. Metzler
NeurIPS3
2024 Neural Subspaces for Light Fields
abstract
We introduce a framework for compactly representing light field content with the novel concept of neural subspaces. While the recently proposed neural light field representation achieves great compression results by encoding a light field into a single neural network, the unified design is not optimized for the composite structures exhibited in light fields. Moreover, encoding every part of the light field into one network is not ideal for applications that require rapid transmission and decoding. We recognize this problem's connection to subspace learning. We present a method that uses several small neural networks, specializing in learning the neural subspace for a particular light field segment. Moreover, we propose an adaptive weight sharing strategy among those small networks, improving parameter efficiency. In effect, this strategy enables a concerted way to track the similarity among nearby neural subspaces by leveraging the layered structure of neural networks. Furthermore, we develop a soft-classification technique to enhance the color prediction accuracy of neural representations. Our experimental results show that our method better reconstructs the light field than previous methods on various light field scenes. We further demonstrate its successful deployment on encoding light fields with irregular viewpoint layout and dynamic scene content.
Brandon Yushan Feng, Amitabh Varshney
IEEE Trans. Vis. Comput. Graph.1
2024 HoloCamera: Advanced Volumetric Capture for Cinematic-Quality VR Applications
abstract
High-precision virtual environments are increasingly important for various education, simulation, training, performance, and entertainment applications. We present HoloCamera, an innovative volumetric capture instrument to rapidly acquire, process, and create cinematic-quality virtual avatars and scenarios. The HoloCamera consists of a custom-designed free-standing structure with 300 high-resolution RGB cameras mounted with uniform spacing spanning the four sides and the ceiling of a room-sized studio. The light field acquired from these cameras is streamed through a distributed array of GPUs that interleave the processing and transmission of 4K resolution images. The distributed compute infrastructure that powers these RGB cameras consists of 50 Jetson AGX Xavier boards, with each processing unit dedicated to driving and processing imagery from six cameras. A high-speed Gigabit Ethernet network fabric seamlessly interconnects all computing boards. In this systems paper, we provide an in-depth description of the steps involved and lessons learned in constructing such a cutting-edge volumetric capture facility that can be generalized to other such facilities. We delve into the techniques employed to achieve precise frame synchronization and spatial calibration of cameras, careful determination of angled camera mounts, image processing from the camera sensors, and the need for a resilient and robust network infrastructure. To advance the field of volumetric capture, we are releasing a high-fidelity static light-field dataset, which will serve as a benchmark for further research and applications of cinematic-quality volumetric light fields.
Jonathan Heagerty, Shuvra S. Bhattacharyya, Sujal Bista, Barbara Brawn, Brandon Yushan Feng, Susmija Jabbireddy, Joseph F. JáJá, Hernisa Kacorri, David Li 0001, Derek Yarnell, Matthias Zwicker, Amitabh Varshney
IEEE Trans. Vis. Comput. Graph.7
2023 Continuous Levels of Detail for Light Field Networks
David Li 0001, Brandon Yushan Feng, Amitabh Varshney
BMVC2
2023 3D Motion Magnification: Visualizing Subtle Motions with Time-Varying Radiance Fields
abstract
Motion magnification helps us visualize subtle, imperceptible motion. However, prior methods only work for 2D videos captured with a fixed camera. We present a 3D motion magnification method that can magnify subtle motions from scenes captured by a moving camera, while supporting novel view rendering. We represent the scene with time-varying radiance fields and leverage the Eulerian principle for motion magnification to extract and amplify the variation of the embedding of a fixed point over time. We study and validate our proposed principle for 3D motion magnification using both implicit and tri-plane-based radiance fields as our underlying 3D scene representation. We evaluate the effectiveness of our method on both synthetic and real-world scenes captured under various camera setups.
Brandon Yushan Feng, Hadi Alzayer, Michael Rubinstein, William T. Freeman, Jia-Bin Huang 0001
ICCV1
2023 StegaNeRF: Embedding Invisible Information within Neural Radiance Fields
abstract
Recent advancements in neural rendering have paved the way for a future marked by the widespread distribution of visual data through the sharing of Neural Radiance Field (NeRF) model weights. However, while established techniques exist for embedding ownership or copyright information within conventional visual data such as images and videos, the challenges posed by the emerging NeRF format have remained unaddressed. In this paper, we introduce StegaNeRF, an innovative approach for steganographic information embedding within NeRF renderings. We have meticulously developed an optimization framework that enables precise retrieval of hidden information from images generated by NeRF, while ensuring the original visual quality of the rendered images to remain intact. Through rigorous experimentation, we assess the efficacy of our methodology across various potential deployment scenarios. Furthermore, we delve into the insights gleaned from our analysis. StegaNeRF represents an initial foray into the intriguing realm of infusing NeRF renderings with customizable, imperceptible, and recoverable information, all while minimizing any discernible impact on the rendered images. For more details, please visit our project page: https://xggnet.github.io/StegaNeRF/
Chenxin Li, Brandon Yushan Feng, Zhiwen Fan, Panwang Pan, Zhangyang Wang
ICCV2
2022 PRIF: Primary Ray-Based Implicit Function
Brandon Yushan Feng, Yinda Zhang 0001, Danhang Tang, Ruofei Du, Amitabh Varshney
ECCV (3)1
2022 VIINTER: View Interpolation with Implicit Neural Representations of Images
abstract
We present VIINTER, a method for view interpolation by interpolating the implicit neural representation (INR) of the captured images. We leverage the learned code vector associated with each image and interpolate between these codes to achieve viewpoint transitions. We propose several techniques that significantly enhance the interpolation quality. VIINTER signifies a new way to achieve view interpolation without constructing 3D structure, estimating camera poses, or computing pixel correspondence. We validate the effectiveness of VIINTER on several multi-view scenes with different types of camera layout and scene composition. As the development of INR of images (as opposed to surface or volume) has centered around tasks like image fitting and super-resolution, with VIINTER, we show its capability for view interpolation and offer a promising outlook on using INR for image manipulation tasks.
Brandon Yushan Feng, Susmija Jabbireddy, Amitabh Varshney
SIGGRAPH Asia1
2021 SIGNET: Efficient Neural Representation for Light Fields
abstract
We present a novel neural representation for light field content that enables compact storage and easy local reconstruction with high fidelity. We use a fully-connected neural network to learn the mapping function between each light field pixel’s coordinates and its corresponding color values. Since neural networks that simply take in raw coordinates are unable to accurately learn data containing fine details, we present an input transformation strategy based on the Gegenbauer polynomials, which previously showed theoretical advantages over the Fourier basis. We conduct experiments that show our Gegenbauer-based design combined with sinusoidal activation functions leads to a better light field reconstruction quality than a variety of network designs, including those with Fourier-inspired techniques introduced by prior works. Moreover, our SInusoidal Gegenbauer NETwork, or SIGNET, can represent light field scenes more compactly than the state-of-the-art compression methods while maintaining a comparable reconstruction quality. SIGNET also innately allows random access to encoded light field pixels due to its functional design. We further demonstrate that SIGNET’s super-resolution capability without any additional training.
Brandon Yushan Feng, Amitabh Varshney
ICCV1
2021 GazeChat: Enhancing Virtual Conferences with Gaze-aware 3D Photos
abstract
Communication software such as Clubhouse and Zoom has evolved to be an integral part of many people’s daily lives. However, due to network bandwidth constraints and concerns about privacy, cameras in video conferencing are often turned off by participants. This leads to a situation in which people can only see each others’ profile images, which is essentially an audio-only experience. Even when switched on, video feeds do not provide accurate cues as to who is talking to whom. This paper introduces GazeChat, a remote communication system that visually represents users as gaze-aware 3D profile photos. This satisfies users’ privacy needs while keeping online conversations engaging and efficient. GazeChat uses a single webcam to track whom any participant is looking at, then uses neural rendering to animate all participants’ profile images so that participants appear to be looking at each other. We have conducted a remote user study (N=16) to evaluate GazeChat in three conditions: audio conferencing with profile photos, GazeChat, and video conferencing. Based on the results of our user study, we conclude that GazeChat maintains the feeling of presence while preserving more privacy and requiring lower bandwidth than video conferencing, provides a greater level of engagement than to audio conferencing, and helps people to better understand the structure of their conversation.
Zhenyi He, Keru Wang, Brandon Yushan Feng, Ruofei Du, Ken Perlin
UIST3
2020 Deep Depth Estimation on 360° Images with a Double Quaternion Loss
abstract
While 360° images are becoming ubiquitous due to popularity of panoramic content, they cannot directly work with most of the existing depth estimation techniques developed for perspective images. In this paper, we present a deep-learning-based framework of estimating depth from 360° images. We present an adaptive depth refinement procedure that refines depth estimates using normal estimates and pixel-wise uncertainty scores. We introduce double quaternion approximation to combine the loss of the joint estimation of depth and surface normal. Furthermore, we use the double quaternion formulation to also measure stereo consistency between the horizontally displaced depth maps, leading to a new loss function for training a depth estimation CNN. Results show that the new double-quaternion-based loss and the adaptive depth refinement procedure lead to better network performance. Our proposed method can be used with monocular as well as stereo images. When evaluated on several datasets, our method surpasses state-of-the-art methods on most metrics.
Brandon Yushan Feng, Wangjue Yao, Zheyuan Liu 0003, Amitabh Varshney
3DV1