Huiqiang Sun

dblp:83/10152 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0002-3653-3613ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021
YearPublicationVenuePosition
2026 BokehFlow: Depth-Free Controllable Bokeh Rendering via Flow Matching
abstract
Bokeh rendering simulates the shallow depth-of-field effect in photography, enhancing visual aesthetics and guiding viewer attention to regions of interest. Although recent approaches perform well, rendering controllable bokeh without additional depth inputs remains a significant challenge. Existing classical and neural controllable methods rely on accurate depth maps, while generative approaches often struggle with limited controllability and efficiency. In this paper, we propose BokehFlow, a depth-free framework for controllable bokeh rendering based on flow matching. BokehFlow directly synthesizes photorealistic bokeh effects from all-in-focus images, eliminating the need for depth inputs. It employs a cross-attention mechanism to enable semantic control over both focus regions and blur intensity via text prompts. To support training and evaluation, we collect and synthesize four datasets. Extensive experiments demonstrate that BokehFlow achieves visually compelling bokeh effects and offers precise control, outperforming existing depth-dependent and generative methods in both rendering quality and efficiency.
Yachuan Huang, Xianrui Luo, Liao Shen, Jiaqi Li 0007, Huiqiang Sun, Zihao Huang 0001, Zhiguo Cao 0001
AAAI6
2026 BokehCrafter: Taming Video Diffusion Models for Controllable Bokeh Rendering
abstract
Bokeh is used in photography to emphasize the selected subject by smoothly blurring the out-of-focus region with appealing highlights. While recent advances have achieved impressive results in rendering realistic blur, existing frameworks typically rely on disparity maps and bokeh-relevant inputs (e.g., focal distance and blur size), and face significant challenges in video bokeh rendering due to limited temporal consistency. In this paper, we propose BokehCrafter, the first video diffusion framework that generates temporally coherent and visually pleasing bokeh effects from all-in-focus video inputs under user-friendly input conditions. Specifically, we leverage a dual-stream attention mechanism, integrating a reference image branch and a rendering instruction branch. We propose a Bokeh Image Extraction (BIE) module and a CLIP-based text encoder to extract image and text features, respectively, whose outputs are fused via a Text-Image Fusion (TIF) module to enable fine-grained and controllable bokeh rendering. To support the novel capabilities of our model, we construct Video Bokeh Scenes (VBS), a large-scale dataset containing a wide variety of bokeh videos with corresponding rendering instructions, across various scenes and rendering settings. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art methods in both bokeh rendering quality and temporal consistency.
Liao Shen, Jiaqi Li 0007, Tianqi Liu 0003, Huiqiang Sun, Zihao Huang 0001, Yachuan Huang, Xianrui Luo, Zhiguo Cao 0001
AAAI5
2025 DoF-Gaussian: Controllable Depth-of-Field for 3D Gaussian Splatting
abstract
Recent advances in 3D Gaussian Splatting (3D-GS) have shown remarkable success in representing 3D scenes and generating high-quality, novel views in real-time. However, 3D-GS and its variants assume that input images are captured based on pinhole imaging and are fully in focus. This assumption limits their applicability, as real-world images often feature shallow depth-of-field (DoF). In this paper, we introduce DoF-Gaussian, a controllable depth-of-field method for 3D-GS. We develop a lens-based imaging model based on geometric optics principles to control DoF effects. To ensure accurate scene geometry, we incorporate depth priors adjusted per scene, and we apply defocus-to-focus adaptation to minimize the gap in the circle of confusion. We also introduce a synthetic dataset to assess refocusing capabilities and the model’s ability to learn precise lens parameters. Our framework is customizable and supports various interactive applications. Extensive experiments confirm the effectiveness of our method. Our project is available at https://dof-gaussian.github.io/.
Liao Shen, Tianqi Liu 0003, Huiqiang Sun, Jiaqi Li 0007, Zhiguo Cao 0001, Wei Li 0319, Chen Change Loy
CVPR3
2025 MuGS: Multi-Baseline Generalizable Gaussian Splatting Reconstruction
Yaopeng Lou, Li Shen 0008, Tianqi Liu 0003, Jiaqi Li 0007, Zihao Huang 0001, Huiqiang Sun, Zhiguo Cao 0001
ICCV6
2025 Dual-Camera All-in-Focus Neural Radiance Fields
abstract
We present the first framework capable of synthesizing the all-in-focus neural radiance field (NeRF) from inputs without manual refocusing. Without refocusing, the camera will automatically focus on the fixed object for all views, and current NeRF methods typically using one camera fail due to the consistent defocus blur and a lack of sharp reference. To restore the all-in-focus NeRF, we introduce the dual-camera from smartphones, where the ultra-wide camera has a wider depth-of-field (DoF) and the main camera possesses a higher resolution. The dual camera pair saves the high-fidelity details from the main camera and uses the ultra-wide camera's deep DoF as reference for all-in-focus restoration. To this end, we first implement spatial warping and color matching to align the dual camera, followed by a defocus-aware fusion module with learnable defocus parameters to predict a defocus map and fuse the aligned camera pair. We also build a multi-view dataset that includes image pairs of the main and ultra-wide cameras in a smartphone. Extensive experiments on this dataset verify that our solution, termed DC-NeRF, can produce high-quality all-in-focus novel views and compares favorably against strong baselines quantitatively and qualitatively. We further show DoF applications of DC-NeRF with adjustable blur intensity and focal plane, including refocusing and split diopter.
Xianrui Luo, Zijin Wu, Juewen Peng, Huiqiang Sun, Zhiguo Cao 0001, Guosheng Lin
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Dynamic View Synthesis From Small Camera Motion Videos
abstract
Novel view synthesis for dynamic 3D scenes poses a significant challenge. Many notable efforts use NeRF-based approaches to address this task and yield impressive results. However, these methods rely heavily on sufficient motion parallax in the input images or videos. When the camera motion range becomes limited or even stationary (i.e., small camera motion), existing methods encounter two primary challenges: incorrect representation of scene geometry and inaccurate estimation of camera parameters. These challenges make prior methods struggle to produce satisfactory results or even become ineffective. To address the first challenge, we propose a novel Distribution-based Depth Regularization (DDR) that ensures the rendering weight distribution to align with the true distribution. Specifically, unlike previous methods that use depth loss to calculate the error of the expectation, we calculate the expectation of the error by using Gumbel-softmax to differentiably sample points from discrete rendering weight distribution. Additionally, we introduce constraints that enforce the volume density of spatial points before the object boundary along the ray to be near zero, ensuring that our model learns the correct geometry of the scene. To demystify the DDR, we further propose a visualization tool that enables observing the scene geometry representation at the rendering weight level. For the second challenge, we incorporate camera parameter learning during training to enhance the robustness of our model to camera parameters. We conduct extensive experiments to demonstrate the effectiveness of our approach in representing scenes with small camera motion input, and our results compare favorably to state-of-the-art methods.
Huiqiang Sun, Xingyi Li 0005, Juewen Peng, Liao Shen, Zhiguo Cao 0001, Ke Xian, Guosheng Lin
IEEE Trans. Vis. Comput. Graph.1
2024 3D Multi-frame Fusion for Video Stabilization
abstract
In this paper, we present RStab, a novel framework for video stabilization that integrates 3D multi-frame fusion through volume rendering. Departing from conventional methods, we introduce a 3D multi-frame perspective to generate stabilized images, addressing the challenge of full-frame generation while preserving structure. The core of our RStab framework lies in Stabilized Rendering (SR), a volume rendering module, fusing multi-frame information in 3D space. Specifically, SR involves warping features and colors from multiple frames by projection, fusing them into descriptors to render the stabilized image. However, the precision of warped information depends on the projection accuracy, a factor significantly influenced by dynamic regions. In response, we introduce the Adaptive Ray Range (ARR) module to integrate depth priors, adaptively defining the sampling range for the projection process. Additionally, we propose Color Correction (CC) assisting geometric constraints with optical flow for accurate color aggregation. Thanks to the three modules, our RStab demonstrates superior performance compared with previous stabilizers in the field of view (FOV), image quality, and video stability across various datasets.
Weiyue Zhao, Tianqi Liu 0003, Huiqiang Sun, Baopu Li, Zhiguo Cao 0001
CVPR5
2024 DyBluRF: Dynamic Neural Radiance Fields from Blurry Monocular Video
abstract
Recent advancements in dynamic neural radiance field methods have yielded remarkable outcomes. However, these approaches rely on the assumption of sharp input images. When faced with motion blur, existing dynamic NeRF methods often struggle to generate high-quality novel views. In this paper, we propose DyBluRF, a dynamic radiance field approach that synthesizes sharp novel views from a monocular video affected by motion blur. To account for motion blur in input images, we simultaneously capture the camera trajectory and object Discrete Cosine Transform (DCT) trajectories within the scene. Additionally, we employ a global cross-time rendering approach to ensure consistent temporal coherence across the entire scene. We curate a dataset comprising diverse dynamic scenes that are specifically tailored for our task. Experimental results on our dataset demonstrate that our method outperforms existing approaches in generating sharp novel views from motion-blurred inputs while maintaining spatial-temporal consistency of the scene.
Huiqiang Sun, Xingyi Li 0005, Liao Shen, Ke Xian, Zhiguo Cao 0001
CVPR1
2024 Dynamic Neural Radiance Field from Defocused Monocular Video
Xianrui Luo, Huiqiang Sun, Juewen Peng, Zhiguo Cao 0001
ECCV (5)2
2024 DreamMover: Leveraging the Prior of Diffusion Models for Image Interpolation with Large Motion
Liao Shen, Tianqi Liu 0003, Huiqiang Sun, Baopu Li, Jianming Zhang 0001, Zhiguo Cao 0001
ECCV (15)3
2023 3D Cinemagraphy from a Single Image
abstract
We present 3D Cinemagraphy, a new technique that mar-ries 2D image animation with 3D photography. Given a single still image as input, our goal is to generate a video that contains both visual content animation and camera motion. We empirically find that naively combining existing 2D image animation and 3D photography methods leads to obvious artifacts or inconsistent animation. Our key insight is that representing and animating the scene in 3D space offers a natural solution to this task. To this end, we first convert the input image into feature-based layered depth images using predicted depth values, followed by unprojecting them to a feature point cloud. To animate the scene, we perform motion estimation and lift the 2D motion into the 3D scene flow. Finally, to resolve the problem of hole emer-gence as points move forward, we propose to bidirectionally displace the point cloud as per the scene flow and synthe-size novel views by separately projecting them into target image planes and blending the results. Extensive experiments demonstrate the effectiveness of our method. A user study is also conducted to validate the compelling rendering results of our method.
Xingyi Li 0005, Zhiguo Cao 0001, Huiqiang Sun, Jianming Zhang 0001, Ke Xian, Guosheng Lin
CVPR3
2023 Make-It-4D: Synthesizing a Consistent Long-Term Dynamic Scene Video from a Single Image
abstract
We study the problem of synthesizing a long-term dynamic video from only a single image. This is challenging since it requires consistent visual content movements given large camera motions. Existing methods either hallucinate inconsistent perpetual views or struggle with long camera trajectories. To address these issues, it is essential to estimate the underlying 4D (including 3D geometry and scene motion) and fill in the occluded regions. To this end, we present Make-It-4D, a novel method that can generate a consistent long-term dynamic video from a single image. On the one hand, we utilize layered depth images (LDIs) to represent a scene, and they are then unprojected to form a feature point cloud. To animate the visual content, the feature point cloud is displaced based on the scene flow derived from motion estimation and the corresponding camera pose. Such 4D representation enables our method to maintain the global consistency of the generated dynamic video. On the other hand, we fill in the occluded regions by using a pre-trained diffusion model to inpaint and outpaint the input image. This enables our method to work under large camera motions. Benefiting from our design, our method can be training-free which saves a significant amount of training time. Experimental results demonstrate the effectiveness of our approach, which showcases compelling rendering results.
Liao Shen, Xingyi Li 0005, Huiqiang Sun, Juewen Peng, Ke Xian, Zhiguo Cao 0001, Guosheng Lin
ACM Multimedia3
2020 Joint low-rank project embedding and optimal mean principal component analysis
abstract
Principal component analysis (PCA) is the most widely used unsupervised dimensionality reduction approach. A number of variants of PCA have been proposed to improve the robustness of the algorithm. However, the existing methods either cannot select the useful features consistently or is still sensitive to outliers. In order to reveal the intrinsic manifold structure and preserve the global structure of data, it is needed to learn more efficient optimal projection matrix for sample sets with outliers. To this end, the authors propose a novel PCA, named low‐rank project embedding and optimal mean principal component analysis (abbreviated as LRPE‐OMPCA), which can learn the optimal mean and the optimal projection matrix and preserve the global geometric information and discriminative structure captured by the self‐representation coefficient weight matrix into the low‐dimensional embedding subspace. Thus, not only can the proposed method further reduce the influence of outliers but also can discard the useless features, which effectively improve the robustness of the method. An effective iterative algorithm to solve the LRPE‐OMPCA is designed. Experimental results on several image databases illustrate the robustness and effectiveness of the proposed method.
Xiuhong Chen, Huiqiang Sun
IET Image Process.2
2019 L 2, 1-norm-based sparse principle component analysis with trace norm regularised term
abstract
Principal component analysis (PCA) is the most widely used unsupervised dimensionality reduction approach. However, most PCA based on the squared reconstruction errors assume that all training samples have been centred, which make them not robust to outliers or noises in the samples and will depress their performance of classification accuracy. On the other hand, when there are various correlations in the training samples, the l 1 ‐norm regularisation encounters instability problems. To address the above problems, the authors propose a novel L 2,1 ‐norm‐based sparse PCA with the trace norm regularised term (abbreviated to OMSPCA‐L21‐TN) to learn the optimal projection matrix and optimal mean simultaneously, where the objective function in model consists of the L 2,1 ‐norm‐based reconstruction error and the trace‐norm‐based regularised term of the projection vectors involved the sample matrix. Thus, not only can the authors’ method obtain the sparse features and reduce the effect of noise and outliers but also be adaptive to the correlation of the training samples. An effective optimisation solution is also given. The experimental results on some publicly available datasets demonstrate that the proposed approach is feasible and effective.
Xiuhong Chen, Huiqiang Sun
IET Image Process.2