VLDB 2026 Research / reviewers in the wild / expert
Yifan Peng 0001
dblp:123/7941 · also Evan Yifan Peng, Yifan (Evan) Peng
· DBLP profile ↗
41ranked-venue papers
5as first author
29since 2021 · last 2026
0000-0003-0667-2599ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 40 · 5 first-author · 28 since 2021Artificial intelligence and machine learning · 8 · 5 since 2021Human-computer interaction and ubiquitous computing · 7 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Color Compressive Hologram Synthesis with Learned Wave PropagationabstractHolographic displays are a promising technology for delivering immersive, true 3D visualization in virtual and augmented reality applications. However, generating high-fidelity phase-only holograms remains challenging, especially with the demand for efficient compression to handle the substantial data inherent in high-resolution holographic streaming. Existing techniques often struggle to balance the trade-off between optical display quality and compression efficiency, and jointly optimizing these aspects is still in its infancy. This work presents a learning-empowered multi-color hologram compression scheme that utilizes a pre-trained, camera-calibrated wave propagation model, especially for unfiltered holographic display configurations with compact form factors. In particular, the inter-color processing leverages the inherent redundancy across color channels, allowing for efficient compression. By incorporating the learned camera-calibrated wave propagation model into our training process, we can achieve superior optical display quality and compression rates. Experiments demonstrate that our method realizes a reduction in bits per pixel (bpp) of 44% to 74% over representative baselines at the same quality level. We envision the proposed compressive hologram synthesis scheme establishing a new benchmark for high-fidelity holographic reconstruction at lower bitrates, marking a significant advance towards the deployment of holography-empowered visual media systems. Hyunmin Ban, Wenbin Zhou 0001, Yifan Peng 0001 |
Comput. Vis. Media | 4 |
| 2026 | HoloPathTracer: Fast and Accurate Wave Path Tracing for HolographyabstractHolography offers unique advantages for delivering perceptual realism while preserving compact form factors in VR/AR. Its perceptual quality, however, hinges on encoding rich wavefronts of photorealistic scenes into interference patterns and then incoherently multiplexing the resulting wave fields for perception. Existing CGH paradigms decouple radiance estimation from wave propagation by pre-rendering radiance on discretized scene sectors. This separation between radiometric and wave-optical computation inherently limits the range of focus cues and visual effects that can be faithfully reproduced, including depth- and view-continuity, and physically based material behaviors such as glossy or mirror-like reflection and refraction. We present a physically accurate yet computationally efficient wave optics rendering framework leveraging path tracing to encode full 3D visual cues into phase holograms. Specifically, we employ a Monte Carlo method to solve both the rendering equation and the Rayleigh-Sommerfeld integral simultaneously. Our algorithm is fully compatible with modern graphics techniques and can generate multiple time-multiplexed random holograms with minimal additional time cost via Path Reuse. By employing a fast approximation with an ambient radiance cache, we realize an order of magnitude convergence speed improvement. The resulting coherent wave fields that inherently encode comprehensive visual effects are converted into phase-only holograms under complex-amplitude supervision. Through extensive simulations and experimental validations on a spatial light modulator-based display prototype, we demonstrate faithful holographic reconstructions of natural 3D cues and complex materials, including realistic defocus blur, view-dependent effects, as well as appearance highlights and reflections. Wenbin Zhou 0001, Jiankai Xing, Suyeon Choi, Yifan Peng 0001 |
ACM Trans. Graph. | 6 |
| 2026 | Towards Edge Holography via Implicit Neural Representation and CompressionabstractHolographic displays offer the promise of realistic 3D visualization for virtual and augmented wearable solutions. Nevertheless, existing computer-generated holography (CGH) methods often struggle with either a high computational burden or limited display realism. While the emerging cloud-edge computing mechanism can enable the real-time streaming of holograms, classic image compression techniques struggle to efficiently encode and decode the substantial high-frequency information inherent in hologram data. In light of these challenges, we present a display-aware and lightweight CGH framework, leveraging implicit neural representations (INRs) and camera-calibrated wave propagation, to generate and compress high-fidelity phase-only holograms. Specifically, our approach interprets hologram generation as a continuous function approximation problem, enabling the network, with reduced parameters, to effectively learn the inherent periodicity and high-frequency components of 2D and 3D hologram data. To enable efficient deployment, we further incorporate quantization-aware training, followed by entropy coding. Experimental results evaluated on an unfiltered holographic display prototype demonstrate that the proposed INR-CGH retains image quality comparable to that of existing optimization-based methods in both 2D and 3D scenarios. In addition, our compact INR representation achieves up to 11× compression rate with minimal quality degradation and can be further reduced via quantization-aware training. The resulting model enables ≥250 fps in decoding speed, paving the way towards edge holography. Hyunmin Ban, Wenbin Zhou 0001, Yifan Peng 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2026 | EventTracer: Fast Path Tracing-Based Event Stream RenderingabstractSimulating event streams from 3D scenes has become a common practice in event-based vision research, as it meets the demand for large-scale, high temporal frequency data without setting up expensive hardware devices or undertaking extensive data collections. Yet existing methods in this direction typically work with noiseless RGB frames that are costly to render, and therefore their simulations are often unrealistically low in temporal resolution. In this work, we propose EventTracer, a path tracing-based rendering pipeline that simulates high-fidelity event sequences from complex 3D scenes in an efficient and physics-aware manner. Specifically, we speed up the rendering process via low sample-per-pixel (SPP) path tracing, and train a lightweight event spiking network to denoise the resulting RGB videos into realistic event sequences. Our EventTracerpipeline runs at a speed of $\sim$∼1 minutes per second of 360p video, and it inherits the merit of accurate spatiotemporal modeling from its path tracing backbone. We show through the Real2Sim and Sim2Real tests that EventTracercaptures higher-fidelity scene details and demonstrates a greater similarity to real-world event data than alternative event simulators, which establishes it as a potential tool for creating large-scale event-RGB datasets, narrowing the sim-to-real gap in event-based vision, and boosting various downstream applications. Xiaoyang Bai, Jinfan Lu, Pengfei Shen, Edmund Y. Lam, Yifan Peng 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | Learned Binocular-Encoding Optics for RGBD Imaging Using Joint Stereo and Focus CuesabstractExtracting high-fidelity RGBD information from two-dimensional (2D) images is essential for various visual computing applications. Stereo imaging, as a reliable passive imaging technique for obtaining three-dimensional (3D) scene information, has benefited greatly from deep learning advancements. However, existing stereo depth estimation algorithms struggle to perceive high-frequency information and resolve high-resolution depth maps in realistic camera settings with large depth variations. These algorithms commonly neglect the hardware parameter configuration, limiting the potential for achieving optimal solutions solely through software-based design strategies.This work presents a hardware-software co-designed RGBD imaging framework that leverages both stereo and focus cues to reconstruct texture-rich color images along with detailed depth maps over a wide depth range. A pair of rank-2 parameterized diffractive optical elements (DOEs) is employed to encode perpendicular complementary information optically during stereo acquisitions. Additionally, we employ an IGEV-UNet-fused neural network tailored to the proposed rank-2 encoding for stereo matching and image reconstruction. Through prototyping a stereo camera with customized DOEs, our deep stereo imaging paradigm has demonstrated superior performance over existing monocular and stereo imaging systems in both image PSNR by 2.96 dB gain and depth accuracy in high-frequency details across distances from 0.67 to 8 meters. Yuhui Liu, Liangxun Ou, Qiang Fu 0002, Hadi Amata, Wolfgang Heidrich, Yifan Peng 0001 |
CVPR | 6 |
| 2025 | Glossy Object Reconstruction with Cost-effective Polarized AcquisitionabstractThe challenge of image-based 3D reconstruction for glossy objects lies in separating diffuse and specular components on glossy surfaces from captured images, a task complicated by the ambiguity in discerning lighting conditions and material properties using RGB data alone. While state-of-the-art methods rely on tailored and/or high-end equipment for data acquisition, which can be cumbersome and time-consuming, this work introduces a scalable polarization-aided approach that employs cost-effective acquisition tools. By attaching a linear polarizer to readily available RGB cameras, multi-view polarization images can be captured without the need for advance calibration or precise measurements of the polarizer angle, substantially reducing system construction costs. The proposed approach represents polarimetric BRDF, Stokes vectors, and polarization states of object surfaces as neural implicit fields. These fields, combined with the polarizer angle, are retrieved by optimizing the rendering loss of input polarized images. By leveraging fundamental physical principles for the implicit representation of polarization rendering, our method demonstrates superiority over existing techniques through experiments in public datasets and real captured images on both reconstruction and novel view synthesis. Bojian Wu, Yifan Peng 0001, Ruizhen Hu, Xiaowei Zhou 0001 |
CVPR | 2 |
| 2025 | Enhanced Velocity Field Modeling for Gaussian Video ReconstructionabstractHigh-fidelity 3D video reconstruction is essential for enabling real-time rendering of dynamic scenes with realistic motion in VR/AR. The deformation field paradigm of 3D Gaussian splatting has achieved near-photorealistic results in video reconstruction due to the great representation capability of deep deformation networks. However, in videos with complex motion and significant scale variations, deformation networks often overfit to irregular Gaussian trajectories, leading to suboptimal visual quality. Moreover, the gradient-based densification strategy designed for static scene reconstruction proves inadequate to address the absence of dynamic content. In light of these challenges, we propose a flow-empowered velocity field modeling scheme tailored for Gaussian video reconstruction, dubbed FlowGaussian-VR. It consists of two core components: a velocity field rendering (VFR) pipeline which enables optical flow-based optimization, and a flow-assisted adaptive densification (FAD) strategy that adjusts the number and size of Gaussians in dynamic regions. We also explore a temporal velocity refinement (TVR) post-processing algorithm to further estimate and correct noise in Gaussian trajectories via extended Kalman filtering. We validate our model's effectiveness on multi-view dynamic reconstruction and novel view synthesis with real-world datasets containing challenging motion scenarios, demonstrating not only notable visual improvements (over 2.5 dB gain in PSNR) and less blurry artifacts in dynamic textures, but also regularized and trackable per-Gaussian trajectories. Xiaoyang Bai, Tongchen Zhang, Pengfei Shen, Weiwei Xu 0003, Yifan Peng 0001 |
ISMAR | 6 |
| 2025 | S3 Imagery: Specular Shading from Scratch-AnisotropyabstractScratch-represented 3D visual arts can create compelling visual effects by manipulating light reflections across surfaces. Established works, such as those involving scratch holograms, have realized impressive multi-view imagery effects of reflection arts. However, creating a continuous view of 3D virtual objects with shading effects, especially view-dependent shading remains a challenge. Yet, most reported works are demonstrated on planar surfaces, leaving exploring the potential benefits of leveraging curved surfaces for diverse imagery scenarios an interesting research avenue. This work explores the continuous view-dependent imagery with rich shading effects via scratch-based reflection, whose design space has the potential to be extended to arbitrary curved surfaces. This is achieved by solving the ordinary differential equations under constraints calculated from established bidirectional reflectance distribution function models to optimize scratch distribution on substrate surfaces. Importantly, we create real-world examples by manufacturing optimized reflectors using off-the-shelf carving machines, delivering state-of-the-art specular view-dependent imagery that features continuous and realistic shading effects on both planar and developable curved surfaces. Pengfei Shen, Feifan Qu, Ruizhen Hu, Yifan Peng 0001 |
SIGGRAPH Asia | 5 |
| 2025 | NeLiF: Neural Lighting Function Generation for Real-Time Indoor RenderingabstractRecent advances in neural rendering have mainly focused on modeling radiance fields with neural representations, often overlooking the underlying mechanisms for producing various lighting effects, and consequently leading to the limited adaptability to dynamic scenes. Lighting effects, such as highlights, shadows, and indirect illuminations, are typically computed using physically-based rendering methods like path tracing, which can be computationally intensive for complex indoor luminaires. Although several recent studies have aimed to model global illumination effects with neural representations, they commonly suffer from long training times or poor generalizability to new scenes. Addressing these challenges, this work presents a novel neural lighting function generation model capable of synthesizing diverse lighting effects in real time for unseen dynamic scenes and complex indoor luminaires, achieving results comparable to state-of-the-art rendering pipelines. Our model operates in two stages. First, multi-view observation images of the luminaire are captured to encode a compact, scene-independent 3D neural lighting field. Subsequently, light information is sampled from this neural lighting field and integrated with G-buffers and shadow clues to produce the shading results. In parallel, we employ a state-of-the-art generative model together with our training-free Inverse HDR Splatting module to generate HDR 3D Gaussians representing the luminaire. This strategy capitalizes on the powerful generalization capabilities of advanced generative models, enabling efficient and accurate appearance reconstruction for a diverse range of complex luminaires. In our experiments, the model trained on a dataset of 10,000 modern indoor scenes and thousands of illuminations demonstrates strong generalizability, high efficiency, and visually convincing results across a wide range of test scenes, highlighting its potential as a practical and flexible solution for high-fidelity, real-time neural indoor rendering. Hongtao Sheng, Yuchi Huo, Chuankun Zheng, Guangzhi Han, Yifan Peng 0001, Bin Zang, Hao Zhu 0004, Rui Tang 0015, Rui Wang 0004, Hujun Bao |
SIGGRAPH Asia | 5 |
| 2025 | 4D Gaussian Videos with Motion LayeringabstractOnline free-view navigation in volumetric videos requires high-quality rendering and real-time streaming in order to provide immersive user experiences. However, existing methods ( e.g. , dynamic NeRF and 3DGS) may not handle dynamic scenes with complex motions, and their models may not be streamable due to storage and bandwidth constraints. In this paper, we propose a novel 4D Gaussian Video (4DGV) approach that enables the creation and streaming of photorealistic, volumetric videos for dynamic scenes over the Internet. The core of our 4DGV is a novel streamable group of Gaussians (GOG) representation based on motion layering. Each GOG consists of static and dynamic points obtained via lifting 2D segmentation into 3D in motion layering, where the deformation of each dynamic point is represented as the temporal offset of its attributes. We also adaptively convert static points back to dynamic points to handle the appearance change, (e.g. , moving shadows and reflections), of static objects through optimization. To support real-time streaming of 4DGVs, we show that by applying quantization on Gaussian attributes and H.265 encoding on deformation offsets, our GOG representation can be significantly compressed (to around 6% of the original model size) without sacrificing the accuracy (PSNR loss less than 0.01dB). Extensive experiments on standard benchmarks demonstrate that our method outperforms state-of-the-art volumetric video approaches, with superior rendering quality and minimum storage overheads. Pinxuan Dai, Peiquan Zhang, Ke Xu 0010, Yifan Peng 0001, Dandan Ding, Yujun Shen, Yin Yang 0002, Xinguo Liu, Rynson W. H. Lau, Weiwei Xu 0003 |
ACM Trans. Graph. | 5 |
| 2025 | A Fully-statistical Wave Scattering Model for Heterogeneous SurfacesabstractHeterogeneous surfaces exhibit spatially varying geometry and material, and therefore admit diverse appearances. Existing computer graphics works can only model heterogeneity using explicit structures or statistical parameters that describe a coarser level of detail. We extend the boundary by introducing a new model that describes the heterogeneous surfaces fully statistically at the microscopic level, with rich geometry and material details that are comparable to the wavelengths of light. We treat the heterogeneous surfaces as a mixture of stochastic vector processes. We adapt the well-known generalized Harvey-Shack theory to quantify the mean scattered intensity, i.e., the BRDF of these surfaces. We further explore the covariance statistic of the scattered field and derive its rank-1 decomposition. This leads to a practical algorithm that samples the speckles (fluctuating intensities) from the statistics, enriching the appearance without explicit definition of heterogeneous surfaces. The formulations are analytic, and we validate the quantities by comprehensive numerical simulations. Our heterogeneous surface model demonstrates various applications including corrosion (natural), particle deposition (man-made), and height-correlated mixture (artistic). Code for this paper is available at https://github.com/Rendering-at-ZJU/HeteroSurface. Zhengze Liu, Yuchi Huo, Yifan Peng 0001, Rui Wang 0004 |
ACM Trans. Graph. | 3 |
| 2025 | MoFlow: Motion-Guided Flows for Recurrent Rendered Frame PredictionabstractRendering realistic images in real-time on high-frame-rate display devices poses considerable challenges, even with advanced graphics cards. This stimulates a demand for frame prediction technologies to boost frame rates. The key to these algorithms is to exploit spatiotemporal coherence by warping rendered pixels with motion representations. However, existing motion estimation methods can suffer from low precision, high overhead, and incomplete support for visual effects. In this article, we present a rendered frame prediction framework with a novel motion representation, dubbed motion-guided flow (MoFlow) , aiming at overcoming the intrinsic limitations of optical flow and motion vectors and precisely capture the dynamics of intricate geometries, lighting, and translucent objects. Notably, we construct MoFlows using a recurrent feature streaming network, which specializes in learning latent motion features from multiple frames. The results of extensive experiments demonstrate that, compared to state-of-the-art methods, our method achieves superior visual quality and temporal stability with lower latency. The recurrent mechanism allows our method to predict single or multiple consecutive frames, increasing the frame rate by over 2×. The proposed approach represents a flexible pipeline to meet the demands of various graphics applications, devices, and scenarios. Zhizhen Wu, Zhilong Yuan, Chenyu Zuo, Yazhen Yuan, Yifan Peng 0001, Guiyang Pu, Rui Wang 0004, Yuchi Huo |
ACM Trans. Graph. | 5 |
| 2025 | GO-NeRF: Generating Objects in Neural Radiance Fields for Virtual Reality Content CreationabstractVirtual environments (VEs) are pivotal for virtual, augmented, and mixed reality systems. Despite advances in 3D generation and reconstruction, the direct creation of 3D objects within an established 3D scene (represented as NeRF) for novel VE creation remains a relatively unexplored domain. This process is complex, requiring not only the generation of high-quality 3D objects but also their seamless integration into the existing scene. To this end, we propose a novel pipeline featuring an intuitive interface, dubbed GO-NeRF. Our approach takes text prompts and user-specified regions as inputs and leverages the scene context to generate 3D objects within the scene. We employ a compositional rendering formulation that effectively integrates the generated 3D objects into the scene, utilizing optimized 3D-aware opacity maps to avoid unintended modifications to the original scene. Furthermore, we develop tailored optimization objectives and training strategies to enhance the model's ability to capture scene context and mitigate artifacts, such as floaters, that may occur while optimizing 3D objects within the scene. Extensive experiments conducted on both forward-facing and 360°scenes demonstrate the superior performance of our proposed method in generating objects that harmonize with surrounding scenes and synthesizing high-quality novel view images. The code will be at https://daipengwa.github.io/G0-NeRF/. Peng Dai 0003, Feitong Tan, Xin Yu 0004, Yifan Peng 0001, Yinda Zhang 0001, Xiaojuan Qi 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | Multi-illumination-interfered Neural Holography with Expanded EyeboxabstractHolography has immense potential for near-eye displays in virtual and augmented reality (VR/AR), providing natural 3D depth cues through wavefront reconstruction. However, balancing the field of view (FOV) with the eyebox remains challenging, constrained by the étendue limitation. Additionally, holographic image quality is often compromised due to differences between actual wave propagation and simulation models. This study addresses these by expanding the eyebox via multi-angle illumination, and enhancing image quality with end-to-end pupil-aware hologram optimization. Further, energy efficiency is improved by incorporating higher-order diffractions and pupil constraints. We explore a Pupil-HOGD algorithm for multi-angle illumination and validate it with a dual-angle holographic display prototype. Integrated with camera calibration and tracked eye position, the developed Pupil-HOGD algorithm improves image quality and expands the eyebox by 50% horizontally. We envision this approach extends the space-bandwidth product (SBP) of holographic displays, enabling broader applications in immersive, high-quality visual computing. Xinxing Xia, Pengfei Mi, Yiqing Tao, Wenbin Zhou 0001, Yingjie Yu, Yifan Peng 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2024 | Learned Scanpaths Aid Blind Panoramic Video Quality AssessmentabstractPanoramic videos have the advantage of providing an immersive and interactive viewing experience. Nevertheless, their spherical nature gives rise to various and uncertain user viewing behaviors, which poses significant challenges for panoramic video quality assessment (PVQA). In this work, we propose an end-to-end optimized, blind PVQA method with explicit modeling of user viewing patterns through visual scanpaths. Our method consists of two modules: a scanpath generator and a quality assessor. The scanpath generator is initially trained to predict future scanpaths by minimizing their expected code length and then jointly optimized with the quality assessor for quality prediction. Our blind PVQA method enables direct quality assessment of panoramic images by treating them as videos composed of identical frames. Experiments on three public panoramic image and video quality datasets, encompassing both synthetic and authentic distortions, validate the superiority of our blind PVQA model over existing methods. Kanglong Fan, Wen Wen 0007, Mu Li 0005, Yifan Peng 0001, Kede Ma |
CVPR | 4 |
| 2024 | Local Gaussian Density Mixtures for Unstructured Lumigraph RenderingabstractPSNR 29.47 PSNR 28.80 PSNR 27.28 PSNR 30.84 PSNR 26.69 PSNR 26. Xiuchao Wu, Jiamin Xu, Chi Wang 0004, Yifan Peng 0001, Qixing Huang, James Tompkin 0001, Weiwei Xu 0003 |
SIGGRAPH Asia | 4 |
| 2024 | Neural Bokeh: Learning Lens Blur for Computational Videography and Out-of-Focus Mixed RealityabstractWe present Neural Bokeh, a deep learning approach for synthesizing convincing out-of-focus effects with applications in Mixed Reality (MR) image and video compositing. Unlike existing approaches that solely learn the amount of blur for out-of-focus areas, our approach captures the overall characteristic of the bokeh to enable the seamless integration of rendered scene content into real images, ensuring a consistent lens blur over the resulting MR composition. Our method learns spatially varying blur shapes, i.e., bokeh, from a dataset of real images acquired using the physical camera that is used to capture the photograph or video of the MR composition. Accordingly, those learned blur shapes mimic the characteristics of the physical lens. As the run-time and the resulting quality of Neural Bokeh increase with the resolution of input images, we employ low-resolution images for the MR view finding at runtime and high-resolution renderings for compositing with high-resolution photographs or videos in an offline process. We envision a variety of applications, including visual enhancement of image and video compositing containing creative utilization of out-of-focus effects. David Mandl, Shohei Mori, Peter Mohr, Yifan Peng 0001, Tobias Langlotz, Dieter Schmalstieg, Denis Kalkofen |
VR | 4 |
| 2024 | LightFormer: Light-Oriented Global Neural Rendering in Dynamic SceneabstractThe generation of global illumination in real time has been a long-standing challenge in the graphics community, particularly in dynamic scenes with complex illumination. Recent neural rendering techniques have shown great promise by utilizing neural networks to represent the illumination of scenes and then decoding the final radiance. However, incorporating object parameters into the representation may limit their effectiveness in handling fully dynamic scenes. This work presents a neural rendering approach, dubbed LightFormer , that can generate realistic global illumination for fully dynamic scenes, including dynamic lighting, materials, cameras, and animated objects, in real time. Inspired by classic many-lights methods, the proposed approach focuses on the neural representation of light sources in the scene rather than the entire scene, leading to the overall better generalizability. The neural prediction is achieved by leveraging the virtual point lights and shading clues for each light. Specifically, two stages are explored. In the light encoding stage, each light generates a set of virtual point lights in the scene, which are then encoded into an implicit neural light representation, along with screen-space shading clues like visibility. In the light gathering stage, a pixel-light attention mechanism composites all light representations for each shading point. Given the geometry and material representation, in tandem with the composed light representations of all lights, a lightweight neural network predicts the final radiance. Experimental results demonstrate that the proposed LightFormer can yield reasonable and realistic global illumination in fully dynamic scenes with real-time performance. Haocheng Ren, Yuchi Huo, Yifan Peng 0001, Hongtao Sheng, Weidong Xue, Hongxiang Huang, Jingzhen Lan, Rui Wang 0004, Hujun Bao |
ACM Trans. Graph. | 3 |
| 2024 | Learned Multi-aperture Color-coded Optics for Snapshot Hyperspectral ImagingabstractLearned optics, which incorporate lightweight diffractive optics, coded-aperture modulation, and specialized image-processing neural networks, have recently garnered attention in the field of snapshot hyperspectral imaging (HSI). While conventional methods typically rely on a single lens element paired with an off-the-shelf color sensor, these setups, despite their widespread availability, present inherent limitations. First, the Bayer sensor's spectral response curves are not optimized for HSI applications, limiting spectral fidelity of the reconstruction. Second, single lens designs rely on a single diffractive optical element (DOE) to simultaneously encode spectral information and maintain spatial resolution across all wavelengths, which constrains spectral encoding capabilities. This work investigates a multi-channel lens array combined with aperture-wise color filters, all co-optimized alongside an image reconstruction network. This configuration enables independent spatial encoding and spectral response for each channel, improving optical encoding across both spatial and spectral dimensions. Specifically, we validate that the method achieves over a 5dB improvement in PSNR for spectral reconstruction compared to existing single-diffractive lens and coded-aperture techniques. Experimental validation further confirmed that the method is capable of recovering up to 31 spectral bands within the 429--700 nm range in diverse indoor and outdoor environments. Zheng Shi 0003, Xiong Dun, Haoyu Wei, Siyu Dong, Zhanshan Wang 0002, Xinbin Cheng, Felix Heide, Yifan Peng 0001 |
ACM Trans. Graph. | 8 |
| 2023 | Learning A Room with the Occ-SDF Hybrid: Signed Distance Function Mingled with Occupancy Aids Scene RepresentationabstractImplicit neural rendering, using signed distance function (SDF) representation with geometric priors like depth or surface normal, has made impressive strides in the surface reconstruction of large-scale scenes. However, applying this method to reconstruct a room-level scene from images may miss structures in low-intensity areas and/or small, thin objects. We have conducted experiments on three datasets to identify limitations of the original color rendering loss and priors-embedded SDF scene representation.Our findings show that the color rendering loss creates an optimization bias against low-intensity areas, resulting in gradient vanishing and leaving these areas unoptimized. To address this issue, we propose a feature-based color rendering loss that utilizes non-zero feature values to bring back optimization signals. Additionally, the SDF representation can be influenced by objects along a ray path, disrupting the monotonic change of SDF values when a single object is present. Accordingly, we explore using the occupancy representation, which encodes each point separately and is unaffected by objects along a querying ray. Our experimental results demonstrate that the joint forces of the feature-based rendering loss and Occ-SDF hybrid representation scheme can provide high-quality reconstruction results, especially in challenging room-level scenarios. The code is available at https://github.com/shawLyu/Occ-SDF-Hybrid Xiaoyang Lyu, Peng Dai 0003, Zizhang Li, Dongyu Yan, Yifan Peng 0001, Xiaojuan Qi 0001 |
ICCV | 6 |
| 2023 | Message from the ISMAR 2023 Science and Technology Conference Program Chairs
Gerd Bruder, Anne-Hélène Olivier, Andrew Cunningham, Yifan Peng 0001, Jens Grubert, Ian Williams 0001 |
ISMAR | 4 |
| 2023 | Adaptive Recurrent Frame Prediction with Learnable Motion VectorsabstractThe utilization of dedicated ray tracing graphics cards has revolutionized the production of stunning visual effects in real-time rendering. However, the demand for high frame rates and high resolutions remains a challenge. The pixel warping approach is a crucial technique for increasing frame rate and resolution by exploiting the spatio-temporal coherence. To this end, existing super-resolution and frame prediction methods rely heavily on motion vectors from rendering engine pipelines to track object movements. This work builds upon state-of-the-art heuristic approaches by exploring a novel adaptive recurrent frame prediction framework that integrates learnable motion vectors. Our framework supports the prediction of transparency, particles, and texture animations, with improved motion vectors that capture shading, reflections, and occlusions, in addition to geometry movements. In addition, we introduce a feature streaming neural network, dubbed FSNet, that allows for the adaptive prediction of one or multiple sequential frames. Extensive experiments against state-of-the-art methods demonstrate that FSNet can operate at lower latency with significant visual enhancements and can upscale frame rates by at least two times. This approach offers a flexible pipeline to improve the rendering frame rates of various graphics applications and devices. Zhizhen Wu, Chenyu Zuo, Yuchi Huo, Yazhen Yuan, Yifan Peng 0001, Guiyang Pu, Rui Wang 0004, Hujun Bao |
SIGGRAPH Asia | 5 |
| 2023 | Off-Axis Layered Displays: Hybrid Direct-View/Near-Eye Mixed Reality with Focus CuesabstractThis work introduces off-axis layered displays, the first approach to stereoscopic direct-view displays with support for focus cues. Off-axis layered displays combine a head-mounted display with a traditional direct-view display for encoding a focal stack and thus, for providing focus cues. To explore the novel display architecture, we present a complete processing pipeline for the real-time computation and post-render warping of off-axis display patterns. In addition, we build two prototypes using a head-mounted display in combination with a stereoscopic direct-view display, and a more widely available monoscopic direct-view display. In addition we show how extending off-axis layered displays with an attenuation layer and with eye-tracking can improve image quality. We thoroughly analyze each component in a technical evaluation and present examples captured through our prototypes. Christoph Ebner, Peter Mohr, Tobias Langlotz, Yifan Peng 0001, Dieter Schmalstieg, Gordon Wetzstein, Denis Kalkofen |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | Learning scale awareness in keypoint extraction and description
Xuelun Shen, Cheng Wang 0003, Xin Li 0003, Yifan Peng 0001, Chenglu Wen, Ming Cheng 0002 |
Pattern Recognit. | 4 |
| 2022 | Video See-Through Mixed Reality with Focus CuesabstractThis work introduces the first approach to video see-through mixed reality with full support for focus cues. By combining the flexibility to adjust the focus distance found in varifocal designs with the robustness to eye-tracking error found in multifocal designs, our novel display architecture reliably delivers focus cues over a large workspace. In particular, we introduce gaze-contingent layered displays and mixed reality focal stacks, an efficient representation of mixed reality content that lends itself to fast processing for driving layered displays in real time. We thoroughly evaluate this approach by building a complete end-to-end pipeline for capture, render, and display of focus cues in video see-through displays that uses only off-the-shelf hardware and compute components. Christoph Ebner, Shohei Mori, Peter Mohr, Yifan Peng 0001, Dieter Schmalstieg, Gordon Wetzstein, Denis Kalkofen |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | Color Contrast Enhanced Rendering for Optical See-Through Head-Mounted DisplaysabstractMost commercially available optical see-through head-mounted displays (OST-HMDs) utilize optical combiners to simultaneously visualize the physical background and virtual objects. The displayed images perceived by users are a blend of rendered pixels and background colors. Enabling high fidelity color perception in mixed reality (MR) scenarios using OST-HMDs is an important but challenging task. We propose a real-time rendering scheme to enhance the color contrast between virtual objects and the surrounding background for OST-HMDs. Inspired by the discovery of color perception in psychophysics, we first formulate the color contrast enhancement as a constrained optimization problem. We then design an end-to-end algorithm to search the optimal complementary shift in both chromaticity and luminance of the displayed color. This aims at enhancing the contrast between virtual objects and the real background as well as keeping the consistency with the original displayed color. We assess the performance of our approach using a simulated OST-HMD environment and an off-the-shelf OST-HMD. Experimental results from objective evaluations and subjective user studies demonstrate that the proposed approach makes rendered virtual objects more distinguishable from the surrounding background, thereby bringing a better visual experience. Yunjin Zhang, Rui Wang 0004, Yifan Peng 0001, Wei Hua 0002, Hujun Bao |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2021 | Depth from Defocus with Learned Optics for Imaging and Occlusion-aware Depth EstimationabstractMonocular depth estimation remains a challenging problem, despite significant advances in neural network architectures that leverage pictorial depth cues alone. Inspired by depth from defocus and emerging point spread function engineering approaches that optimize programmable optics end-to-end with depth estimation networks, we propose a new and improved framework for depth estimation from a single RGB image using a learned phase-coded aperture. Our optimized aperture design uses rotational symmetry constraints for computational efficiency, and we jointly train the optics and the network using an occlusion-aware image formation model that provides more accurate defocus blur at depth discontinuities than previous techniques do. Using this framework and a custom prototype camera, we demonstrate state-of-the art image and depth estimation quality among end-to-end optimized computational cameras in simulation and experiment. Hayato Ikoma, Cindy M. Nguyen, Christopher A. Metzler, Yifan Peng 0001, Gordon Wetzstein |
ICCP | 4 |
| 2021 | Neural 3D holography: learning accurate wave propagation models for 3D holographic virtual and augmented reality displaysabstractHolographic near-eye displays promise unprecedented capabilities for virtual and augmented reality (VR/AR) systems. The image quality achieved by current holographic displays, however, is limited by the wave propagation models used to simulate the physical optics. We propose a neural network-parameterized plane-to-multiplane wave propagation model that closes the gap between physics and simulation. Our model is automatically trained using camera feedback and it outperforms related techniques in 2D plane-to-plane settings by a large margin. Moreover, it is the first network-parameterized model to naturally extend to 3D settings, enabling high-quality 3D computer-generated holography using a novel phase regularization strategy of the complex-valued wave field. The efficacy of our approach is demonstrated through extensive experimental evaluation with both VR and optical see-through AR display prototypes. Suyeon Choi, Manu Gopakumar, Yifan Peng 0001, Jonghyun Kim 0006, Gordon Wetzstein |
ACM Trans. Graph. | 3 |
| 2021 | Differentiable Compound Optics and Processing Pipeline Optimization for End-to-end Camera DesignabstractMost modern commodity imaging systems we use directly for photography—or indirectly rely on for downstream applications—employ optical systems of multiple lenses that must balance deviations from perfect optics, manufacturing constraints, tolerances, cost, and footprint. Although optical designs often have complex interactions with downstream image processing or analysis tasks, today’s compound optics are designed in isolation from these interactions. Existing optical design tools aim to minimize optical aberrations, such as deviations from Gauss’ linear model of optics, instead of application-specific losses, precluding joint optimization with hardware image signal processing (ISP) and highly parameterized neural network processing. In this article, we propose an optimization method for compound optics that lifts these limitations. We optimize entire lens systems jointly with hardware and software image processing pipelines, downstream neural network processing, and application-specific end-to-end losses. To this end, we propose a learned, differentiable forward model for compound optics and an alternating proximal optimization method that handles function compositions with highly varying parameter dimensions for optics, hardware ISP, and neural nets. Our method integrates seamlessly atop existing optical design tools, such as Zemax . We can thus assess our method across many camera system designs and end-to-end applications. We validate our approach in an automotive camera optics setting—together with hardware ISP post processing and detection—outperforming classical optics designs for automotive object detection and traffic light state detection. For human viewing tasks, we optimize optics and processing pipelines for dynamic outdoor scenarios and dynamic low-light imaging. We outperform existing compartmentalized design or fine-tuning methods qualitatively and quantitatively, across all domain-specific applications tested. Ethan Tseng, Ali Mosleh 0002, Fahim Mannan, Karl St. Arnaud, Yifan Peng 0001, Alexander Braun 0001, Derek Nowrouzezahrai, Jean-François Lalonde, Felix Heide |
ACM Trans. Graph. | 6 |
| 2020 | Deep Optics for Single-Shot High-Dynamic-Range ImagingabstractHigh-dynamic-range (HDR) imaging is crucial for many applications. Yet, acquiring HDR images with a single shot remains a challenging problem. Whereas modern deep learning approaches are successful at hallucinating plausible HDR content from a single low-dynamic-range (LDR) image, saturated scene details often cannot be faithfully recovered. Inspired by recent deep optical imaging approaches, we interpret this problem as jointly training an optical encoder and electronic decoder where the encoder is parameterized by the point spread function (PSF) of the lens, the bottleneck is the sensor with a limited dynamic range, and the decoder is a convolutional neural network (CNN). The lens surface is then jointly optimized with the CNN in a training phase; we fabricate this optimized optical element and attach it as a hardware add-on to a conventional camera during inference. In extensive simulations and with a physical prototype, we demonstrate that this end-to-end deep optical imaging approach to single-shot HDR imaging outperforms both purely CNN-based approaches and other PSF engineering approaches. Christopher A. Metzler, Hayato Ikoma, Yifan Peng 0001, Gordon Wetzstein |
CVPR | 3 |
| 2020 | Neural holography with camera-in-the-loop trainingabstractHolographic displays promise unprecedented capabilities for direct-view displays as well as virtual and augmented reality applications. However, one of the biggest challenges for computer-generated holography (CGH) is the fundamental tradeoff between algorithm runtime and achieved image quality, which has prevented high-quality holographic image synthesis at fast speeds. Moreover, the image quality achieved by most holographic displays is low, due to the mismatch between the optical wave propagation of the display and its simulated model. Here, we develop an algorithmic CGH framework that achieves unprecedented image fidelity and real-time framerates. Our framework comprises several parts, including a novel camera-in-the-loop optimization strategy that allows us to either optimize a hologram directly or train an interpretable model of the optical wave propagation and a neural network architecture that represents the first CGH algorithm capable of generating full-color high-quality holographic images at 1080p resolution in real time. Yifan Peng 0001, Suyeon Choi, Nitish Padmanaban, Gordon Wetzstein |
ACM Trans. Graph. | 1 |
| 2020 | End-to-end Learned, Optically Coded Super-resolution SPAD CameraabstractSingle Photon Avalanche Photodiodes (SPADs) have recently received a lot of attention in imaging and vision applications due to their excellent performance in low-light conditions, as well as their ultra-high temporal resolution. Unfortunately, like many evolving sensor technologies, image sensors built around SPAD technology currently suffer from a low pixel count. In this work, we investigate a simple, low-cost, and compact optical coding camera design that supports high-resolution image reconstructions from raw measurements with low pixel counts. We demonstrate this approach for regular intensity imaging, depth imaging, as well transient imaging. Our method uses an end-to-end framework to simultaneously optimize the optical design and a reconstruction network for obtaining super-resolved images from raw measurements. The optical design space is that of an engineered point spread function (implemented with diffractive optics), which can be considered an optimized anti-aliasing filter to preserve as much high-resolution information as possible despite imaging with a low pixel count, low fill-factor SPAD array. We further investigate a deep network for reconstruction. The effectiveness of this joint design and reconstruction approach is demonstrated for a range of different applications, including high-speed imaging, and time of flight depth imaging, as well as transient imaging. While our work specifically focuses on low-resolution SPAD sensors, similar approaches should prove effective for other emerging image sensor technologies with low pixel counts and low fill-factors. Qilin Sun 0001, Jian Zhang 0018, Xiong Dun, Bernard Ghanem, Yifan Peng 0001, Wolfgang Heidrich |
ACM Trans. Graph. | 5 |
| 2019 | Wirtinger holography for near-eye displaysabstractNear-eye displays using holographic projection are emerging as an exciting display approach for virtual and augmented reality at high-resolution without complex optical setups --- shifting optical complexity to computation. While precise phase modulation hardware is becoming available, phase retrieval algorithms are still in their infancy, and holographic display approaches resort to heuristic encoding methods or iterative methods relying on various relaxations. In this work, we depart from such existing approximations and solve the phase retrieval problem for a hologram of a scene at a single depth at a given time by revisiting complex Wirtinger derivatives, also extending our framework to render 3D volumetric scenes. Using Wirtinger derivatives allows us to pose the phase retrieval problem as a quadratic problem which can be minimized with first-order optimization methods. The proposed Wirtinger Holography is flexible and facilitates the use of different loss functions, including learned perceptual losses parametrized by deep neural networks, as well as stochastic optimization methods. We validate this framework by demonstrating holographic reconstructions with an order of magnitude lower error, both in simulation and on an experimental hardware prototype. Praneeth Chakravarthula, Yifan Peng 0001, Joel S. Kollin, Henry Fuchs, Felix Heide |
ACM Trans. Graph. | 2 |
| 2019 | Holographic near-eye displays based on overlap-add stereogramsabstractHolographic near-eye displays are a key enabling technology for virtual and augmented reality (VR/AR) applications. Holographic stereograms (HS) are a method of encoding a light field into a hologram, which enables them to natively support view-dependent lighting effects. However, existing HS algorithms require the choice of a hogel size, forcing a tradeoff between spatial and angular resolution. Based on the fact that the short-time Fourier transform (STFT) connects a hologram to its observable light field, we develop the overlap-add stereogram (OLAS) as the correct method of "inverting" the light field into a hologram via the STFT. The OLAS makes more efficient use of the information contained within the light field than previous HS algorithms, exhibiting better image quality at a range of distances and hogel sizes. Most remarkably, the OLAS does not degrade spatial resolution with increasing hogel size, overcoming the spatio-angular resolution tradeoff that previous HS algorithms face. Importantly, the optimal hogel size of previous methods typically varies with the depth of every object in a scene, making the OLAS not only a hogel size-invariant method, but also nearly scene independent. We demonstrate the performance of the OLAS both in simulation and on a prototype near-eye display system, showing focusing capabilities and view-dependent effects. Nitish Padmanaban, Yifan Peng 0001, Gordon Wetzstein |
ACM Trans. Graph. | 2 |
| 2019 | Learned large field-of-view imaging with thin-plate opticsabstractTypical camera optics consist of a system of individual elements that are designed to compensate for the aberrations of a single lens. Recent computational cameras shift some of this correction task from the optics to post-capture processing, reducing the imaging optics to only a few optical elements. However, these systems only achieve reasonable image quality by limiting the field of view (FOV) to a few degrees - effectively ignoring severe off-axis aberrations with blur sizes of multiple hundred pixels. In this paper, we propose a lens design and learned reconstruction architecture that lift this limitation and provide an order of magnitude increase in field of view using only a single thin-plate lens element. Specifically, we design a lens to produce spatially shift-invariant point spread functions, over the full FOV, that are tailored to the proposed reconstruction architecture. We achieve this with a mixture PSF, consisting of a peak and and a low-pass component, which provides residual contrast instead of a small spot size as in traditional lens designs. To perform the reconstruction, we train a deep network on captured data from a display lab setup, eliminating the need for manual acquisition of training data in the field. We assess the proposed method in simulation and experimentally with a prototype camera system. We compare our system against existing single-element designs, including an aspherical lens and a pinhole, and we compare against a complex multielement lens, validating high-quality large field-of-view (i.e. 53°) imaging performance using only a single thin-plate element. Yifan Peng 0001, Qilin Sun 0001, Xiong Dun, Gordon Wetzstein, Wolfgang Heidrich, Felix Heide |
ACM Trans. Graph. | 1 |
| 2018 | Depth and Transient Imaging With Compressive SPAD Array CamerasabstractTime-of-flight depth imaging and transient imaging are two imaging modalities that have recently received a lot of interest. Despite much research, existing hardware systems are limited either in terms of temporal resolution or are prohibitively expensive. Arrays of Single Photon Avalanche Diodes (SPADs) promise to fill this gap by providing higher temporal resolution at an affordable cost. Unfortunately SPAD arrays are to date only available in relatively small resolutions. In this work we aim to overcome the spatial resolution limit of SPAD arrays by employing a compressive sensing camera design. Using a DMD and custom optics, we achieve an image resolution of up to 800×400 on SPAD Arrays of resolution 64×32. Using our new data fitting model for the time histograms, we suppress the noise while abstracting the phase and amplitude information, so as to realize a temporal resolution of a few tens of picoseconds. Qilin Sun 0001, Xiong Dun, Yifan Peng 0001, Wolfgang Heidrich |
CVPR | 3 |
| 2018 | Focal sweep imaging with multi-focal diffractive opticsabstractDepth-dependent defocus results in a limited depth-of-field in consumer-level cameras. Computational imaging provides alternative solutions to resolve all-in-focus images with the assistance of designed optics and algorithms. In this work, we extend the concept of focal sweep from refractive optics to diffractive optics, where we fuse multiple focal powers onto one single element. In contrast to state-of-the-art sweep models, ours can generate better-conditioned point spread function (PSF) distributions along the expected depth range with drastically shortened (40%) sweep distance. Further by encoding axially asymmetric PSFs subject to color channels, and then sharing sharp information across channels, we preserve details as well as color fidelity. We prototype two diffractive imaging systems that work in the monochromatic and RGB color domain. Experimental results indicate that the depth-of-field can be significantly extended with fewer artifacts remaining after the deconvolution. Yifan Peng 0001, Xiong Dun, Qilin Sun 0001, Felix Heide, Wolfgang Heidrich |
ICCP | 1 |
| 2018 | End-to-end optimization of optics and image processing for achromatic extended depth of field and super-resolution imagingabstractIn typical cameras the optical system is designed first; once it is fixed, the parameters in the image processing algorithm are tuned to get good image reproduction. In contrast to this sequential design approach, we consider joint optimization of an optical system (for example, the physical shape of the lens) together with the parameters of the reconstruction algorithm. We build a fully-differentiable simulation model that maps the true source image to the reconstructed one. The model includes diffractive light propagation, depth and wavelength-dependent effects, noise and nonlinearities, and the image post-processing. We jointly optimize the optical parameters and the image processing algorithm parameters so as to minimize the deviation between the true and reconstructed image, over a large set of images. We implement our joint optimization method using autodifferentiation to efficiently compute parameter gradients in a stochastic optimization algorithm. We demonstrate the efficacy of this approach by applying it to achromatic extended depth of field and snapshot super-resolution imaging. Vincent Sitzmann, Steven Diamond, Yifan Peng 0001, Xiong Dun, Stephen P. Boyd, Wolfgang Heidrich, Felix Heide, Gordon Wetzstein |
ACM Trans. Graph. | 3 |
| 2017 | Revisiting Cross-Channel Information Transfer for Chromatic Aberration CorrectionabstractImage aberrations can cause severe degradation in image quality for consumer-level cameras, especially under the current tendency to reduce the complexity of lens designs in order to shrink the overall size of modules. In simplified optical designs, chromatic aberration can be one of the most significant causes for degraded image quality, and it can be quite difficult to remove in post-processing, since it results in strong blurs in at least some of the color channels. In this work, we revisit the pixel-wise similarity between different color channels of the image and accordingly propose a novel algorithm for correcting chromatic aberration based on this cross-channel correlation. In contrast to recent weak prior-based models, ours uses strong pixel-wise fitting and transfer, which lead to significant quality improvements for large chromatic aberrations. Experimental results on both synthetic and real world images captured by different optical systems demonstrate that the chromatic aberration can be significantly reduced using our approach. Tiancheng Sun, Yifan Peng 0001, Wolfgang Heidrich |
ICCV | 2 |
| 2017 | Mix-and-match holographyabstractComputational caustics and light steering displays offer a wide range of interesting applications, ranging from art works and architectural installations to energy efficient HDR projection. In this work we expand on this concept by encoding several target images into pairs of front and rear phase-distorting surfaces. Different target holograms can be decoded by mixing and matching different front and rear surfaces under specific geometric alignments. Our approach, which we call mix-and-match holography, is made possible by moving from a refractive caustic image formation process to a diffractive, holographic one. This provides the extra bandwidth that is required to multiplex several images into pairing surfaces. We derive a detailed image formation model for the setting of holographic projection displays, as well as a multiplexing method based on a combination of phase retrieval methods and complex matrix factorization. We demonstrate several application scenarios in both simulation and physical prototypes. Yifan Peng 0001, Xiong Dun, Qilin Sun 0001, Wolfgang Heidrich |
ACM Trans. Graph. | 1 |
| 2016 | The diffractive achromat full spectrum computational imaging with diffractive opticsabstractDiffractive optical elements (DOEs) have recently drawn great attention in computational imaging because they can drastically reduce the size and weight of imaging devices compared to their refractive counterparts. However, the inherent strong dispersion is a tremendous obstacle that limits the use of DOEs in full spectrum imaging, causing unacceptable loss of color fidelity in the images. In particular, metamerism introduces a data dependency in the image blur, which has been neglected in computational imaging methods so far. We introduce both a diffractive achromat based on computational optimization, as well as a corresponding algorithm for correction of residual aberrations. Using this approach, we demonstrate high fidelity color diffractive-only imaging over the full visible spectrum. In the optical design, the height profile of a diffractive lens is optimized to balance the focusing contributions of different wavelengths for a specific focal length. The spectral point spread functions (PSFs) become nearly identical to each other, creating approximately spectrally invariant blur kernels. This property guarantees good color preservation in the captured image and facilitates the correction of residual aberrations in our fast two-step deconvolution without additional color priors. We demonstrate our design of diffractive achromat on a 0.5mm ultrathin substrate by photolithography techniques. Experimental results show that our achromatic diffractive lens produces high color fidelity and better image quality in the full visible spectrum. Yifan Peng 0001, Qiang Fu 0002, Felix Heide, Wolfgang Heidrich |
ACM Trans. Graph. | 1 |