Sizhuo Ma

dblp:227/4634 · DBLP profile ↗
← Back
24ranked-venue papers
8as first author
20since 2021 · last 2026
0000-0003-0092-9744ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 6 first-author · 17 since 2021Artificial intelligence and machine learning · 12 · 3 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Snapmoji: Instant Generation of Animatable Dual-Stylized Avatars
abstract
Despite the increasing popularity of avatar systems such as Snapchat Bitmojis, existing production avatar platforms face several limitations, such as a limited number of predefined assets, tedious customization processes, and inefficient rendering requirements. Addressing these shortcomings, we introduce Snapmoji, an avatar generation system that instantly creates 3D avatars, and enables customization in a process we call dual-stylization. Snapmoji first maps a selfie of a user to a primary avatar (e.g., Bitmoji style) using a new technique we name Gaussian Domain Adaptation (GDA), then applies a secondary style (e.g., skeleton, yarn, toy) to the primary avatar, all while preserving the user’s identity. The generated 3D avatars can then be rendered an animated on mobile devices at 30–40 FPS.
Eric Ming Chen, Di Liu 0003, Sizhuo Ma, Michael Vasilkovsky, Bing Zhou 0001, Wenzhou Wang, Jiahao Luo, Dimitris N. Metaxas, Vincent Sitzmann, Jian Wang 0100
WACV3
2026 Velocity Disambiguation for Video Frame Interpolation
abstract
Existing video frame interpolation (VFI) methods blindly predict where each object is at a specific timestep $t$t ("time indexing"), which struggles to predict precise object movements. Given two images of a baseball, there are infinitely many possible trajectories: accelerating or decelerating, straight or curved. This often results in blurry frames as the method averages out these possibilities. Instead of forcing the network to learn this complicated time-to-location mapping implicitly together with predicting the frames, we provide the network with an explicit hint on how far the object has traveled between start and end frames, a novel approach termed "distance indexing". This method offers a clearer learning goal for models, reducing the uncertainty tied to object speeds. We further observed that, even with this extra guidance, objects can still be blurry especially when they are equally far from both input frames (i.e., halfway in-between), due to the directional ambiguity in long-range motion. To solve this, we propose an iterative reference-based estimation strategy that breaks down a long-range prediction into several short-range steps. When integrating our plug-and-play strategies into state-of-the-art learning-based models, they exhibit markedly sharper outputs and superior perceptual quality in arbitrary time interpolations, using a uniform distance indexing map in the same format as time indexing without requiring extra computation. Furthermore, we demonstrate that if additional latency is acceptable, a continuous map estimator can be employed to compute a pixel-wise dense distance indexing using multiple nearby frames. Combined with efficient multi-frame refinement, this extension can further disambiguate complex motion, thus enhancing performance both qualitatively and quantitatively. Additionally, the ability to manually specify distance indexing allows for independent temporal manipulation of each object, providing a novel tool for video editing tasks such as re-timing.
Zhihang Zhong, Wei Wang 0333, Xiao Sun 0001, Yu Qiao 0001, Gurunandan Krishnan, Sizhuo Ma, Jian Wang 0100
IEEE Trans. Pattern Anal. Mach. Intell.7
2025 Exposure-Limited Image Enhancement with Generative Diffusion Prior
abstract
Many consumer cameras are equipped with 8-bit image sensors, which often struggle to capture scenes with a High Dynamic Range (HDR). This limitation can result in overexposed or underexposed regions, a loss of fine details due to low bit-depth compression, skewed color distributions, and noticeable noise in dark areas. Traditional Standard Dynamic Range (SDR) image enhancement methods typically focus on color mapping by expanding the color range and adjusting brightness. However, they often fail to restore details in dynamic range extremes, i.e. regions where pixel values approach the minimum or maximum limits. We define “exposure-limited image enhancement” as the process of enhancing images with large missing areas due to exposure issues within the SDR space, which differs from existing “mis-exposed image enhancement” methods primarily aimed at correcting color distributions. To enhance these exposure-limited images and overcome the limitations of current models, we propose a novel two-stage approach. In the first stage, we remap color and brightness to a suitable range while preserving existing details. In the second stage, we use a diffusion prior to generate content in severely overexposed or underexposed regions, which are otherwise lost during capture. Notably, this generative refinement module can also serve as a plug-and-play component alongside existing enhancement methods. Extensive experiments demonstrate that our method significantly improves image quality and detail, outperforming state-of-the-art techniques in dynamic range extremes. The project page is at https://Sagiri0208.github.io.
Baiang Li, Sizhuo Ma, Yanhong Zeng, Xiaogang Xu 0002, Youqing Fang, Zhao Zhang 0001, Jian Wang 0100, Kai Chen 0026
ICCP2
2025 DiffBody: Human Body Image Restoration with Generative Diffusion Prior
abstract
Human body image restoration is crucial for various applications but remains challenging due to the limitations of generative models: General image restoration methods built on generative models may generate unnatural textures, noticeable structural misalignments, and significant loss of fine details. To address these shortcomings, we present DiffBody, a novel human body-aware diffusion model that incorporates domain-specific knowledge to significantly enhance restoration quality. Our approach adopts a two-stage framework: (1) a multi-branch joint diffusion model generates preliminary priors, including normal and depth maps supported by a robust reconstruction pre-processing step; (2) a restoration stage refines the output using a body-prior ControlNet and a color adapter, ensuring structural accuracy and color consistency. Extensive quantitative evaluations, qualitative evaluations, and user studies validate the superior performance of DiffBody in producing perceptually high-quality human body restoration results. Code is available at https://github.com/yimingz1218/DiffBody.
Lionel Z. Wang, Sizhuo Ma, Xinjie Li 0002, Zhihang Zhong, Jian Wang 0100
ICCP3
2025 Privacy-Enabled Parallax Display
abstract
Privacy filters for displays are designed to obfuscate or hide visual content from unintended observers, while making the displayed information visible only to selected viewers. Existing privacy filters suffer from either wide viewing field (low selectivity) or limited user positioning. To solve this dilemma, we propose a display technology that allows a narrow but adaptive viewing field that can be directed to arbitrary user location. While conventional parallax barriers provide such capability of modulating the light field according to the user location, it suffers from repeated views. Our key observation is that this view repetition originates from the periodicity of barrier patterns, and we propose a privacy-enabled parallax display based on randomized barrier design. In addition to randomizing the locations of 1D slits, we also propose breaking down the slits into pinholes and randomizing their 2D locations, which results in privacy-preservation along the vertical direction as well. We build a hardware prototype using two off-the-shelf liquid-crystal displays. Experiments show that the proposed randomized parallax barrier can direct to the user a narrow viewing field of about ±6°, providing a significantly improved privacy protection as compared to traditional privacy screens.
Sizhuo Ma, Karl Bayer, Gurunandan Krishnan, Mohit Gupta 0001, Shree K. Nayar
VR1
2025 Personalized Restoration via Dual-Pivot Tuning
abstract
Generative diffusion models can serve as priors, ensuring that image restoration solutions adhere to natural image manifolds. For facial images, however, personalized priors are essential to accurately reconstruct individual-specific facial features. We propose Dual-Pivot Tuning - a simple yet effective two-stage approach to personalize blind restoration systems while preserving general prior integrity. Our key observation is that for efficient personalization, the diffusion model should be tuned around a fixed textual pivot in the first step, while in the second step a guiding network should be tuned in a generic (non-personalized) manner, using the personalized diffusion model as a fixed "pivot". This approach ensures that personalization does not interfere with the restoration process, producing results with a natural appearance that show high fidelity to both identity and degraded image attributes. We conducted extensive experiments with images of widely recognized individuals, evaluating our approach both qualitatively and quantitatively against relevant baselines. Notably, our personalized prior not only achieves superior identity fidelity, but also outperforms state-of-the-art generic priors in terms of overall image quality. Project webpage is https://personalized-restoration.github.io/ and code is available at https://github.com/personalized-restoration/personalized-restoration.
Pradyumna Chari, Sizhuo Ma, Daniil Ostashev, Achuta Kadambi, Gurunandan Krishnan, Jian Wang 0100, Kfir Aberman
IEEE Trans. Image Process.2
2025 Privacy-Preserving Visual Localization With Event Cameras
abstract
We consider the problem of client-server localization, where edge device users communicate visual data with the service provider for locating oneself against a pre-built 3D map. This localization paradigm is a crucial component for location-based services in AR/VR or mobile applications, as it is not trivial to store large-scale 3D maps and process fast localization on resource-limited edge devices. Nevertheless, conventional client-server localization systems possess numerous challenges in computational efficiency, robustness, and privacy-preservation during data transmission. Our work aims to jointly solve these challenges with a localization pipeline based on event cameras. By using event cameras, our system consumes low energy and maintains small memory bandwidth. Then during localization, we propose applying event-to-image conversion and leverage mature image-based localization, which achieves robustness even in low-light or fast-moving scenes. To further enhance privacy protection, we introduce privacy protection techniques at two levels. Network level protection aims to hide the entire user's view in private scenes using a novel split inference approach, while sensor level protection aims to hide sensitive user details such as faces with light-weight filtering. Both methods involve small client-side computation and localization performance loss, while significantly mitigating the feeling of insecurity as revealed in our user study. We thus project our method to serve as a building block for practical location-based services using event cameras.
Young Min Kim 0001, Ramzi Zahreddine, Weston A. Welge, Gurunandan Krishnan, Sizhuo Ma, Jian Wang 0100
IEEE Trans. Image Process.6
2024 DSL-FIQA: Assessing Facial Image Quality via Dual-Set Degradation Learning and Landmark-Guided Transformer
abstract
Generic Face Image Quality Assessment (GFIQA) evalu-ates the perceptual quality of facial images, which is crucial in improving image restoration algorithms and selecting high-quality face images for downstream tasks. We present a novel transformer-based method for GFIQA, which is aided by two unique mechanisms. First, a “Dual-Set Degradation Representation Learning” (DSL) mechanism uses facial images with both synthetic and real degradations to decouple degradation from content, ensuring gen-eralizability to real-world scenarios. This self-supervised method learns degradation features on a global scale, pro-viding a robust alternative to conventional methods that use local patch information in degradation learning. Second, our transformer leverages facial landmarks to emphasize visually salient parts of a face image in evaluating its per-ceptual quality. We also introduce a balanced and diverse Comprehensive Generic Face IQA (CGFIQA-40k) dataset of 40K images carefully designed to overcome the biases, in particular the imbalances in skin tone and gender represen-tation, in existing datasets. Extensive analysis and evaluation demonstrate the robustness of our method, marking a significant improvement over prior methods.
Gurunandan Krishnan, Sy-Yen Kuo, Sizhuo Ma, Jian Wang 0100
CVPR5
2024 RobustSAM: Segment Anything Robustly on Degraded Images
abstract
Segment Anything Model (SAM) has emerged as a transformative approach in image segmentation, acclaimed for its robust zero-shot segmentation capabilities and flexible prompting system. Nonetheless, its performance is challenged by images with degraded quality. Addressing this limitation, we propose the Robust Segment Anything Model (RobustSAM), which enhances SAM's performance on low-quality images while preserving its promptability and zero-shot generalization. Our method leverages the pre-trained SAM model with only marginal parameter increments and computational requirements. The additional parameters of RobustSAM can be optimized within 30 hours on eight GPUs, demonstrating its feasibility and practicality for typical research laboratories. We also introduce the Robust-Seg dataset, a collection of 688K image-mask pairs with different degradations designed to train and evaluate our model optimally. Extensive experiments across various segmentation tasks and datasets confirm RobustSAM's superior performance, especially under zero-shot conditions, underscoring its potential for extensive real-world application. Additionally, our method has been shown to effectively improve the performance of SAM-based downstream tasks such as single image dehazing and deblurring.
Yu-Jiet Vong, Sy-Yen Kuo, Sizhuo Ma, Jian Wang 0100
CVPR4
2024 Holodepth: Programmable Depth-Varying Projection via Computer-Generated Holography
Dorian Chan, Matthew O'Toole, Sizhuo Ma, Jian Wang 0100
ECCV (61)3
2024 Delving Deep into Engagement Prediction of Short Videos
Dasong Li, Baili Lu, Hongsheng Li 0001, Sizhuo Ma, Gurunandan Krishnan, Jian Wang 0100
ECCV (54)5
2024 Clearer Frames, Anytime: Resolving Velocity Ambiguity in Video Frame Interpolation
Zhihang Zhong, Gurunandan Krishnan, Xiao Sun 0001, Yu Qiao 0001, Sizhuo Ma, Jian Wang 0100
ECCV (33)5
2024 DisCO: Portrait Distortion Correction with Perspective-Aware 3D GANs
Zhixiang Wang 0001, Yu-Lun Liu 0001, Jia-Bin Huang 0001, Shin'ichi Satoh 0001, Sizhuo Ma, Gurunandan Krishnan, Jian Wang 0100
Int. J. Comput. Vis.5
2024 Perspective-Aligned AR Mirror with Under-Display Camera
abstract
Augmented reality (AR) mirrors are novel displays that have great potential for commercial applications such as virtual apparel try-on. Typically the camera is placed beside the display, leading to distorted perspectives during user interaction. In this paper, we present a novel approach to address this problem by placing the camera behind a transparent display, thereby providing users with a perspective-aligned experience. Simply placing the camera behind the display can compromise image quality due to optical effects. We meticulously analyze the image formation process, and present an image restoration algorithm that benefits from physics-based data synthesis and network design. Our method significantly improves image quality and outperforms existing methods especially on the underexplored wire and backscatter artifacts. We then carefully design a full AR mirror system including display and camera selection, real-time processing pipeline, and mechanical design. Our user study demonstrates that the system is exceptionally well-received by users, highlighting its advantages over existing camera configurations not only as an AR mirror, but also for video conferencing. Our work represents a step forward in the development of AR mirrors, with potential applications in retail, cosmetics, fashion, etc. The image restoration dataset and code are available at https://perspective-armirror.github.io/.
Jian Wang 0100, Sizhuo Ma, Karl Bayer, Yi Zhang 0108, Peihao Wang, Bing Zhou 0001, Shree K. Nayar, Gurunandan Krishnan
ACM Trans. Graph.2
2023 Energy-Efficient Adaptive 3D Sensing
abstract
Active depth sensing achieves robust depth estimation but is usually limited by the sensing range. Naively increasing the optical power can improve sensing range but induces eye-safety concerns for many applications, including autonomous robots and augmented reality. In this paper, we propose an adaptive active depth sensor that jointly optimizes range, power consumption, and eye-safety. The main observation is that we need not project light patterns to the entire scene but only to small regions of interest where depth is necessary for the application and passive stereo depth estimation fails. We theoretically compare this adaptive sensing scheme with other sensing strategies, such as full-frame projection, line scanning, and point scanning. We show that, to achieve the same maximum sensing distance, the proposed method consumes the least power while having the shortest (best) eye-safety distance. We implement this adaptive sensing scheme with two hardware prototypes, one with a phase-only spatial light modulator (SLM) and the other with a micro-electro-mechanical (MEMS) mirror and diffractive optical elements (DOE). Experimental results validate the advantage of our method and demonstrate its capability of acquiring higher quality geometry adaptively. Please see our project website for video results and code: https://btilmon.github.io/e3d.html.
Brevin Tilmon, Zhanghao Sun, Sanjeev J. Koppal, Georgios Evangelidis 0002, Ramzi Zahreddine, Gurunandan Krishnan, Sizhuo Ma, Jian Wang 0100
CVPR8
2023 Be Real in Scale: Swing for True Scale in Dual Camera Mode
abstract
Many mobile AR apps that use the front-facing camera can benefit significantly from knowing the metric scale of the user’s face. However, the true scale of the face is hard to measure because monocular vision suffers from a fundamental ambiguity in scale. The methods based on prior knowledge about the scene either have a large error or are not easily accessible. In this paper, we propose a new method to measure the face scale by a simple user interaction: the user only needs to swing the phone to capture two selfies while using the recently popular Dual Camera mode. This mode allows simultaneous streaming of the front camera and the rear cameras and has become a key feature in many social apps. A computer vision method is applied to first estimate the absolute motion of the phone from the images captured by two rear cameras, and then calculate the point cloud of the face by triangulation. We develop a prototype mobile app to validate the proposed method. Our user study shows that the proposed method is favored compared to existing methods because of its high accuracy and ease of use. Our method can be built into Dual Camera mode and can enable a wide range of applications (e.g., virtual try-on for online shopping, true-scale 3D face modeling, gaze tracking, and face anti-spoofing) by introducing true scale to smartphone-based XR. The code is available at https://github.com/ruiyu0/Swing-for-True-Scale.
Rui Yu 0002, Jian Wang 0100, Sizhuo Ma, Sharon X. Huang, Gurunandan Krishnan
ISMAR3
2023 QfaR: Location-Guided Scanning of Visual Codes from Long Distances
abstract
Visual codes such as QR codes provide a low-cost and convenient communication channel between physical objects and mobile devices, but typically operate when the code and the device are in close physical proximity. We propose a system, called QfaR, which enables mobile devices to scan visual codes across long distances even where the image resolution of the visual codes is extremely low. QfaR is based on location-guided code scanning, where we utilize a crowd-sourced database of physical locations of codes. Our key observation is that if the approximate location of the codes and the user is known, the space of possible codes can be dramatically pruned down. Then, even if every "single bit" from the low-resolution code cannot be recovered, QfaR can still identify the visual code from the pruned list with high probability. By applying computer vision techniques, QfaR is also robust against challenging imaging conditions, such as tilt, motion blur, etc. Experimental results with common iOS and Android devices show that QfaR can significantly enhance distances at which codes can be scanned, e.g., 3.6cm-sized codes can be scanned at a distance of 7.5 meters, and 0.5m-sized codes at about 100 meters. QfaR has many potential applications, and beyond our diverse experiments, we also conduct a simple case study on its use for efficiently scanning QR code-based badges to estimate event attendance.
Sizhuo Ma, Jian Wang 0100, Wenzheng Chen, Suman Banerjee 0001, Mohit Gupta 0001, Shree K. Nayar
MobiCom1
2023 Burst Vision Using Single-Photon Cameras
abstract
Single-photon avalanche diodes (SPADs) are novel image sensors that record the arrival of individual photons at extremely high temporal resolution. In the past, they were only available as single pixels or small-format arrays, for various active imaging applications such as LiDAR and microscopy. Recently, high-resolution SPAD arrays up to 3.2 megapixel have been realized, which for the first time may be able to capture sufficient spatial details for general computer vision tasks, purely as a passive sensor. However, existing vision algorithms are not directly applicable on the binary data captured by SPADs. In this paper, we propose developing quanta vision algorithms based on burst processing for extracting scene information from SPAD photon streams. With extensive real-world data, we demonstrate that current SPAD arrays, along with burst processing as an example plug-and-play algorithm, are capable of a wide range of downstream vision tasks in extremely challenging imaging conditions including fast motion, low light (< 5 lux) and high dynamic range. To our knowledge, this is the first attempt to demonstrate the capabilities of SPAD sensors for a wide gamut of real-world computer vision tasks including object detection, pose estimation, SLAM, and text recognition. We hope this work will inspire future research into developing computer vision algorithms in extreme scenarios using single-photon cameras.
Sizhuo Ma, Paul Mos, Edoardo Charbon, Mohit Gupta 0001
WACV1
2023 Seeing Photons in Color
abstract
Megapixel single-photon avalanche diode (SPAD) arrays have been developed recently, opening up the possibility of deploying SPADs as generalpurpose passive cameras for photography and computer vision. However, most previous work on SPADs has been limited to monochrome imaging. We propose a computational photography technique that reconstructs high-quality color images from mosaicked binary frames captured by a SPAD array, even for high-dyanamic-range (HDR) scenes with complex and rapid motion. Inspired by conventional burst photography approaches, we design algorithms that jointly denoise and demosaick single-photon image sequences. Based on the observation that motion effectively increases the color sample rate, we design a blue-noise pseudorandom RGBW color filter array for SPADs, which is tailored for imaging dark, dynamic scenes. Results on simulated data, as well as real data captured with a fabricated color SPAD hardware prototype shows that the proposed method can reconstruct high-quality images with minimal color artifacts even for challenging low-light, HDR and fast-moving scenes. We hope that this paper, by adding color to computational single-photon imaging, spurs rapid adoption of SPADs for real-world passive imaging applications.
Sizhuo Ma, Varun Sundar, Paul Mos, Claudio Bruschini, Edoardo Charbon, Mohit Gupta 0001
ACM Trans. Graph.1
2022 Single-Photon Structured Light
abstract
We present a novel structured light technique that uses Single Photon Avalanche Diode (SPAD) arrays to enable 3D scanning at high-frame rates and low-light levels. This technique, called “Single-Photon Structured Light”, works by sensing binary images that indicates the presence or absence of photon arrivals during each exposure; the SPAD array is used in conjunction with a high-speed binary projector, with both devices operated at speeds as high as 20 kHz. The binary images that we acquire are heavily influenced by photon noise and are easily corrupted by ambient sources of light. To address this, we develop novel temporal sequences using error correction codes that are designed to be robust to short-range effects like projector and camera defocus as well as resolution mismatch between the two devices. Our lab prototype is capable of 3D imaging in challenging scenarios involving objects with extremely low albedo or undergoing fast motion, as well as scenes under strong ambient illumination.
Varun Sundar, Sizhuo Ma, Aswin C. Sankaranarayanan, Mohit Gupta 0001
CVPR2
2020 Inertial Safety from Structured Light
Sizhuo Ma, Mohit Gupta 0001
ECCV (23)1
2020 Differential Scene Flow from Light Field Gradients
Sizhuo Ma, Brandon M. Smith 0001, Mohit Gupta 0001
Int. J. Comput. Vis.1
2020 Quanta burst photography
abstract
Single-photon avalanche diodes (SPADs) are an emerging sensor technology capable of detecting individual incident photons, and capturing their time-of-arrival with high timing precision. While these sensors were limited to singlepixel or low-resolution devices in the past, recently, large (up to 1 MPixel) SPAD arrays have been developed. These single-photon cameras (SPCs) are capable of capturing high-speed sequences of binary single-photon images with no read noise. We present quanta burst photography, a computational photography technique that leverages SPCs as passive imaging devices for photography in challenging conditions, including ultra low-light and fast motion. Inspired by recent success of conventional burst photography, we design algorithms that align and merge binary sequences captured by SPCs into intensity images with minimal motion blur and artifacts, high signal-to-noise ratio (SNR), and high dynamic range. We theoretically analyze the SNR and dynamic range of quanta burst photography, and identify the imaging regimes where it provides significant benefits. We demonstrate, via a recently developed SPAD array, that the proposed method is able to generate high-quality images for scenes with challenging lighting, complex geometries, high dynamic range and moving objects. With the ongoing development of SPAD arrays, we envision quanta burst photography finding applications in both consumer and scientific photography.
Sizhuo Ma, Arin C. Ulku, Claudio Bruschini, Edoardo Charbon, Mohit Gupta 0001
ACM Trans. Graph.1
2018 3D Scene Flow from 4D Light Field Gradients
Sizhuo Ma, Brandon M. Smith 0001, Mohit Gupta 0001
ECCV (8)1