Boning Liu 0001

dblp:243/6845-1 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0002-3014-6049ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
YearPublicationVenuePosition
2026 Splat-SAP: Feed-Forward Gaussian Splatting for Human-Centered Scene with Scale-Aware Point Map Reconstruction
abstract
We present Splat-SAP, a feed-forward approach to render novel views of human-centered scenes from binocular cameras with large sparsity. Gaussian Splatting has shown its promising potential in rendering tasks, but it typically necessitates per-scene optimization with dense input views. Although some recent approaches achieve feed-forward Gaussian Splatting rendering through geometry priors obtained by multi-view stereo, such approaches still require largely overlapped input views to establish the geometry prior. To bridge this gap, we leverage pixel-wise point map reconstruction to represent geometry which is robust to large sparsity for its independent view modeling. In general, we propose a two-stage learning strategy. In stage 1, we transform the point map into real space via an iterative affinity learning process, which facilitates camera control in the following. In stage 2, we project point maps of two input views onto the target view plane and refine such geometry via stereo matching. Furthermore, we anchor Gaussian primitives on this refined plane in order to render high-quality images. As a metric representation, the scale-aware point map in stage 1 is trained in a self-supervised manner without 3D supervision and stage 2 is supervised with photo-metric loss. We collect multi-view human-centered data and demonstrate that our method improves both the stability of point map reconstruction and the visual quality of free-viewpoint rendering.
Boyao Zhou, Shunyuan Zheng, Zhanfeng Liao, Zihan Ma 0011, Hanzhang Tu, Boning Liu 0001, Yebin Liu
AAAI6
2024 GaussianAvatar: Towards Realistic Human Avatar Modeling from a Single Video via Animatable 3D Gaussians
abstract
We present GaussianAvatar, an efficient approach to cre-ating realistic human avatars with dynamic 3D appear-ances from a single video. We start by introducing animat-able 3D Gaussians to explicitly represent humans in var-ious poses and clothing styles. Such an explicit and ani-matable representation can fuse 3D appearances more effi-ciently and consistently from 2D observations. Our repre-sentation is further augmented with dynamic properties to support pose-dependent appearance modeling, where a dy-namic appearance network along with an optimizable feature tensor is designed to learn the motion-to-appearance mapping. Moreover, by leveraging the differentiable motion condition, our method enables a joint optimization of motions and appearances during avatar modeling, which helps to tackle the long-standing issue of inaccurate motion esti-mation in monocular settings. The efficacy of GaussianA-vatar is validated on both the public dataset and our col-lected dataset, demonstrating its superior performances in terms of appearance quality and rendering efficiency. The code and dataset are available at https://github.com/aipixel/GaussianAvatar.
Liangxiao Hu, Hongwen Zhang 0001, Yuxiang Zhang 0006, Boyao Zhou, Boning Liu 0001, Shengping Zhang, Liqiang Nie
CVPR5
2024 GPS-Gaussian: Generalizable Pixel-Wise 3D Gaussian Splatting for Real-Time Human Novel View Synthesis
abstract
We present a new approach, termed GPS-Gaussian, for synthesizing novel views of a character in a real-time manner. The proposed method enables 2K-resolution rendering under a sparse-view camera setting. Unlike the original Gaussian Splatting or neural implicit rendering methods that necessitate per-subject optimizations, we introduce Gaussian parameter maps defined on the source views and regress directly Gaussian Splatting properties for instant novel view synthesis without any fine-tuning or optimization. To this end, we train our Gaussian parameter regression module on a large amount of human scan data, jointly with a depth estimation module to lift 2D parameter maps to 3D space. The proposed framework is fully differentiable and experiments on several datasets demonstrate that our method outperforms state-of-the-art methods while achieving an exceeding rendering speed. The code is available at https://github.com/aipixel/GPS-Gaussian.
Shunyuan Zheng, Boyao Zhou, Ruizhi Shao, Boning Liu 0001, Shengping Zhang, Liqiang Nie, Yebin Liu
CVPR4
2024 Implicit Surface Representation Using Epanechnikov Mixture Regression
abstract
We propose a regression-based implicit surface representation using mixture-of-experts based on the Epanechnikov kernel (EK), a mathematical framework that does not depend on neural networks. The modeling method is implemented using signed distance fields (SDF), modeled using the expectation-maximization algorithm to iterate an optimal set of parameters of Epanechnikov mixture regression. The proposed pipeline achieves better reconstruction than the SDF itself and can be upsampled through mixture-of-experts-based interpolation without extra parameters and processing. Furthermore, the proposed method can efficiently realize data compression compared to meshes and SDF. As for the kernel theory, EK demonstrates a more accurate surface recovery than the Gaussian ones, which expands the applications for Epanechnikov-related theories and also shows potential for theoretical substitution for Gaussian-based modeling and representation.
Boning Liu 0001, Zerong Zheng, Yebin Liu
IEEE Signal Process. Lett.1
2024 5-D Epanechnikov Mixture-of-Experts in Light Field Image Compression
abstract
In this study, we propose a modeling-based compression approach for dense/lenslet light field images captured by Plenoptic 2.0 with square microlenses. This method employs the 5-D Epanechnikov Kernel (5-D EK) and its associated theories. Owing to the limitations of modeling larger image block using the Epanechnikov Mixture Regression (EMR), a 5-D Epanechnikov Mixture-of-Experts using Gaussian Initialization (5-D EMoE-GI) is proposed. This approach outperforms 5-D Gaussian Mixture Regression (5-D GMR). The modeling aspect of our coding framework utilizes the entire EI and the 5D Adaptive Model Selection (5-D AMLS) algorithm. The experimental results demonstrate that the decoded rendered images produced by our method are perceptually superior, outperforming High Efficiency Video Coding (HEVC) and JPEG 2000 at a bit depth below 0.06bpp.
Boning Liu 0001, Yan Zhao 0012, Xiaomeng Jiang, Xingguang Ji, Shigang Wang 0003, Yebin Liu
IEEE Trans. Image Process.1
2023 Tensor4D: Efficient Neural 4D Decomposition for High-Fidelity Dynamic Reconstruction and Rendering
abstract
We present Tensor4D, an efficient yet effective approach to dynamic scene modeling. The key of our solution is an efficient 4D tensor decomposition method so that the dynamic scene can be directly represented as a 4D spatio-temporal tensor. To tackle the accompanying memory issue, we decompose the 4D tensor hierarchically by projecting it first into three time-aware volumes and then nine compact feature planes. In this way, spatial information over time can be simultaneously captured in a compact and memory-efficient manner. When applying Tensor4D for dynamic scene reconstruction and rendering, we further factorize the 4D fields to different scales in the sense that structural motions and dynamic detailed changes can be learned from coarse to fine. The effectiveness of our method is validated on both synthetic and real-world scenes. Extensive experiments show that our method is able to achieve high-quality dynamic reconstruction and rendering from sparse-view camera rigs or even a monocular camera. The code and dataset will be released at https://github.com/DSaurus/Tensor4D.
Ruizhi Shao, Zerong Zheng, Hanzhang Tu, Boning Liu 0001, Hongwen Zhang 0001, Yebin Liu
CVPR4
2023 AvatarReX: Real-time Expressive Full-body Avatars
abstract
We present AvatarReX, a new method for learning NeRF-based full-body avatars from video data. The learnt avatar not only provides expressive control of the body, hands and the face together, but also supports real-time animation and rendering. To this end, we propose a compositional avatar representation, where the body, hands and the face are separately modeled in a way that the structural prior from parametric mesh templates is properly utilized without compromising representation flexibility. Furthermore, we disentangle the geometry and appearance for each part. With these technical designs, we propose a dedicated deferred rendering pipeline, which can be executed at a real-time framerate to synthesize high-quality free-view images. The disentanglement of geometry and appearance also allows us to design a two-pass training strategy that combines volume rendering and surface rendering for network training. In this way, patch-level supervision can be applied to force the network to learn sharp appearance details on the basis of geometry estimation. Overall, our method enables automatic construction of expressive full-body avatars with real-time rendering capability, and can generate photo-realistic images with dynamic details for novel body motions and facial expressions.
Zerong Zheng, Xiaochen Zhao, Hongwen Zhang 0001, Boning Liu 0001, Yebin Liu
ACM Trans. Graph.4
2022 4D Epanechnikov Mixture Regression in LF Image Compression
abstract
With the emergence of light field imaging in recent years, the compression of its elementary image array (EIA) has become a significant problem. Our coding framework includes modeling and reconstruction. For the modeling, the covariance-matrix form of the 4D Epanechnikov kernel (4D EK) and its correlated statistics were deduced to obtain the 4-D Epanechnikov mixture models (4-D EMMs). A 4D Epanechnikov mixture regression (4D EMR) was proposed based on this 4D EK, and a 4D adaptive model selection (4D AMLS) algorithm was designed to realize the optimal modeling for a pseudo video sequence (PVS) of the extracted key-EIA. A linear function based reconstruction (LFBR) was proposed based on the correlation between adjacent elementary images (EIs). The decoded images realized a clear outline reconstruction and superior coding efficiency compared to high-efficiency video coding (HEVC) and JPEG 2000 below approximately 0.05 bpp. This work realized an unprecedented theoretical application by (1) proposing the 4D Epanechnikov kernel theory, (2) exploiting the 4D Epanechnikov mixture regression and its application in the modeling of the pseudo video sequence of light field images, (3) using 4D adaptive model selection for the optimal number of models, and (4) employing a linear function-based reconstruction according to the content similarity.
Boning Liu 0001, Yan Zhao 0012, Xiaomeng Jiang, Shigang Wang 0003
IEEE Trans. Circuits Syst. Video Technol.1
2021 3-D Epanechnikov Mixture Regression in integral imaging compression
Boning Liu 0001, Yan Zhao 0012, Xiaomeng Jiang, Shigang Wang 0003
J. Vis. Commun. Image Represent.1
2021 Three-dimensional Epanechnikov mixture regression in image coding
abstract
Kernel methods have been studied extensively in recent years. We propose a three-dimensional (3-D) Epanechnikov Mixture Regression (EMR) based on our Epanechnikov Kernel (EK) and realize a complete framework for image coding. In our research, we deduce the covariance-matrix form of 3-D Epanechnikov kernels and their correlated statistics to obtain the Epanechnikov mixture models. To apply our theories to image coding, we propose the 3-D EMR which can better model an image in smaller blocks compared with the conventional Gaussian Mixture Regression (GMR). The regressions are all based on our improved Expectation-Maximization (EM) algorithm with mean square error optimization. Finally, we design an Adaptive Mode Selection (AMS) algorithm to realize the best model pattern combination for coding. Our recovered image has clear outlines and superior coding efficiency compared to JPEG below 0.25bpp. Our work realizes an unprecedented theory application by: (1) enriching the theory of Epanechnikov kernel, (2) improving the EM algorithm using MSE optimization, (3) exploiting the EMR and its application in image coding, and (4) AMS optimal modeling combined with Gaussian and Epanechnikov kernel.
Boning Liu 0001, Yan Zhao 0012, Xiaomeng Jiang, Shigang Wang 0003
Signal Process.1
2019 An Image Coding Approach Based on Mixture-of-experts Regression Using Epanechnikov Kernel
abstract
In this paper, we propose an optimal modeling framework for image compression using EMM (Epanechnikov Mixture Model). Epanechnikov Kernel and its correlated statistics are basement of our Epanechnikov Mixture Regression (EMR). In our scheme, the stochastic processes of the pixel values are modelled as an EMM with K experts in three-dimensional space and then we use EMR to search for the optimal solution, whose parameters are determined through EM (Expectation-Maximization) algorithm. In the process of regression, the conditional density is the regression kernel function. Experimental results show that the proposed scheme is effective especially for the image with complex texture without consuming extra bits compared to Gaussian Mixture Regression (GMR).
Boning Liu 0001, Yan Zhao 0012, Xiaomeng Jiang, Shigang Wang 0003
ICASSP1