Bingzheng Liu

dblp:282/0394 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-6949-4147ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Lightweight stereo image super-resolution via adaptive pruning and bridge distillation
Zhe Zhang 0041, Bingzheng Liu, Lei Chen 0091, Pengzhi Li, Yidan Zhang 0002, Jianjun Lei 0001
Knowl. Based Syst.2
2025 Advancing Generalizable Occlusion Modeling for Neural Human Radiance Field
abstract
Generalizable human neural rendering aims to render the target views of the human body by leveraging source views and the skinned multi-person linear (SMPL) model. Despite exhibiting promising performance, the target views rendered by previous methods usually contain corrupted parts of the human body. Two primary challenges hinder high-quality human neural rendering. These challenges involve non-correspondences between 2D pixels and 3D SMPL vertices induced by self-occlusion of the human body and erroneous appearance predictions caused by occlusion between the source and target views. To solve these two challenges, we propose an advancing generalizable occlusion modeling method for the neural human radiance field, in which the hurdles from the self-occlusion of the human body and the occlusion between source and target views are explored and solved. Specifically, to alleviate the non-correspondence problem induced by self-occlusion, a geometry perception module is designed to obtain 3D geometric representations of SMPL vertices, enabling the prediction of accurate density values. Furthermore, a visibility aggregation module is designed to estimate the visibility maps with respect to different source views by utilizing the predicted density. Then, the complementary information among multiple source views is integrated with the support of the visibility maps in the visibility aggregation module, thus effectively addressing the occlusion between views. Experiments on the ZJU-MoCap and THUman datasets show that the proposed method achieves promising performance compared with the existing state-of-the-art methods.
Bingzheng Liu, Jianjun Lei 0001, Bo Peng 0007, Zhe Zhang 0041, Qingming Huang
IEEE Trans. Multim.1
2024 Unsupervised Single-View Synthesis Network via Style Guidance and Prior Distillation
abstract
View synthesis aims to learn a view transformation and synthesize the target views from a single or multiple source views. Although previous view synthesis methods have obtained promising performance, they heavily rely on the supervision of the target view. In this paper, we propose an unsupervised single-view synthesis network (USVS-Net) to learn the view transformation without the supervision of the target view. Specifically, with the usage of only a single source view, a style-guidance view synthesis model is proposed to learn an intrinsic representation, which intends to describe the object from a reference pose. With the intrinsic representation, the view transformation is learned to boost the learning of the unsupervised single-view synthesis. Then, taking the style-guidance view synthesis model as the teacher, a prior-distillation view synthesis model is further presented as the student to learn a more direct view transformation. By utilizing the proposed method, high-quality target views are synthesized in a time-efficient manner. Experiments on both synthetic and real-scene datasets show that despite the lack of supervision of the target view, the proposed method achieves promising results compared with the existing view synthesis methods.
Bingzheng Liu, Bo Peng 0007, Zhe Zhang 0041, Qingming Huang, Nam Ling, Jianjun Lei 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 Self-Constructing Stereo Correspondences for Unsupervised Multi-View Stereo
abstract
Existing unsupervised Multi-View Stereo (MVS) methods generally construct supervision on the basis of the photometric consistency loss, which suffers from unreliable supervision and limited scalability. In this paper, a novel unsupervised MVS framework with Self-constructed Stereo Correspondences, termed SSC-MVS, is proposed to provide reliable supervision for the network and improve scalability of unsupervised MVS. Specifically, a pseudo depth-based learning strategy is first presented to supervise the MVS network with a pseudo depth, which is used to characterize the accurate stereo correspondences. Additionally, a consistency-based training mechanism is designed, where the depth consistency between two differently-augmented inputs is constrained to further improve the robustness of the network in real MVS scenes. Experimental results on widely-used MVS datasets demonstrate that the proposed SSC-MVS obtains the state-of-the-art performance among the unsupervised methods and has the potential to outperform the fully-supervised methods. The code is available athttps://github.com/jzhu98/ssc-mvs.
Bo Peng 0007, Bingzheng Liu, Qingming Huang, Jianjun Lei 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 Self-Supervised Monocular Depth Estimation via Binocular Geometric Correlation Learning
abstract
Monocular depth estimation aims to infer a depth map from a single image. Although supervised learning-based methods have achieved remarkable performance, they generally rely on a large amount of labor-intensively annotated data. Self-supervised methods, on the other hand, do not require any annotation of ground-truth depth and have recently attracted increasing attention. In this work, we propose a self-supervised monocular depth estimation network via binocular geometric correlation learning. Specifically, considering the inter-view geometric correlation, a binocular cue prediction module is presented to generate the auxiliary vision cue for the self-supervised learning of monocular depth estimation. Then, to deal with the occlusion in depth estimation, an occlusion interference attenuated constraint is developed to guide the supervision of the network by inferring the occlusion region and producing paired occlusion masks. Experimental results on two popular benchmark datasets have demonstrated that the proposed network obtains competitive results compared to state-of-the-art self-supervised methods and achieves comparable results to some popular supervised methods.
Bo Peng 0007, Jianjun Lei 0001, Bingzheng Liu, Haifeng Shen, Wanqing Li 0001, Qingming Huang
ACM Trans. Multim. Comput. Commun. Appl.4
2023 Novel View Synthesis from a Single Unposed Image via Unsupervised Learning
abstract
Novel view synthesis aims to generate novel views from one or more given source views. Although existing methods have achieved promising performance, they usually require paired views with different poses to learn a pixel transformation. This article proposes an unsupervised network to learn such a pixel transformation from a single source image. In particular, the network consists of a token transformation module that facilities the transformation of the features extracted from a source image into an intrinsic representation with respect to a pre-defined reference pose and a view generation module that synthesizes an arbitrary view from the representation. The learned transformation allows us to synthesize a novel view from any single source image of an unknown pose. Experiments on the widely used view synthesis datasets have demonstrated that the proposed network is able to produce comparable results to the state-of-the-art methods despite the fact that learning is unsupervised and only a single source image is required for generating a novel view. The code will be available upon the acceptance of the article.
Bingzheng Liu, Jianjun Lei 0001, Bo Peng 0007, Chuanbo Yu, Wanqing Li 0001, Nam Ling
ACM Trans. Multim. Comput. Commun. Appl.1
2022 Transferring knowledge from monocular completion for self-supervised monocular depth estimation
Bingzheng Liu, Liying Xu, Zhe Zhang 0041
Multim. Tools Appl.3
2022 A 32 × 32-Pixel Flash LiDAR Sensor With Noise Filtering for High-Background Noise Applications
abstract
This article introduces a pulsed laser direct time-of-flight (dTOF) flash light detection and ranging (LiDAR) sensor fabricated in 0.18-$\mu \text{m}$HV CMOS technology. The chip includes$32\times 32$macro pixels and 1024 time-to-digital converters (TDCs). A noise filtering circuit with different threshold ($\text{N}_{\mathrm {th}}$) configuration is adopted in each macro pixel [formed by four single-photon avalanche diodes (SPADs)], which can suppress strong background light (BG) induced pile-up. To verify the imaging function and effectiveness of the noise filtering circuit, two systems are implemented (System1 for indoor imaging measurement and System2 for outdoor distance measurement). With the help of the noise filtering circuit and a reasonable signal-to-background noise ratio (SBR), the maximum detection range outdoors with reasonable accuracy can be greatly extended (from 12m @ Nth= 1 of System2 to more than 20m @ Nth= 2 of System2 under 70klux of background noise). The counter in the noise filtering circuit can be reused to get intensity information. A robust 13-bit TDC with a reliable reset is introduced. Thanks to a dedicated START/STOP logic and a Schmitt trigger, large TDC quantization errors can be avoided. It achieves a 200ps resolution (LSB) and exhibits an INLp-pof 3.55LSB and a DNLp-pof 0.53LSB. The maximum inter-frame rate can reach 270kfps with 16 IOs operating at speed of 500MHz. Combining 9k inter-frames to get one frame, a frame rate of 30 is achieved for an indoor 3-D imaging. For outdoor measurement, more laser pulses should be accumulated.
Jin Hu 0006, Bingzheng Liu, Rui Ma 0007, Maliang Liu, Zhangming Zhu
IEEE Trans. Circuits Syst. I Regul. Pap.2
2022 A 16-Channel Analog CMOS SiPM With On-Chip Front-End for D-ToF LiDAR
abstract
This article presents a 16-channel analog silicon photomultiplier (SiPM) with on-chip front-end for direct time-of-flight (D-ToF) LiDAR applications. The proposed receiver is mainly composed of 16-channel SiPM, variable gain amplifier (VGA) and time-to-digital converter (TDC). Each SiPM channel consists of 256 microcells, and their outputs are connected to a common output terminal in parallel. A novel active quenching circuit is adopted to reduce the long exponential tail in conventional SiPM and enable the capability of multi-echo detection. Current steering circuits are adopted within microcells to make the output current of SiPM immune to SPAD gain variations. The receiver was fabricated in 180-nm HV CMOS technology and integrated into the 16-line LiDAR prototype with optical components. Measurement results show that the sensor is capable of 20 m range imaging with 3 cm accuracy under 40 klux background light conditions. With the mechanical scanning system, a high-resolution image ($240\times16$) can be obtained.
Maliang Liu, Bingzheng Liu, Jin Hu 0006, Dong Li 0046, Jiaji Ma 0001, Zekun Chu, Rui Ma 0007, Zhangming Zhu
IEEE Trans. Circuits Syst. I Regul. Pap.2
2022 Multi-Modality MR Image Synthesis via Confidence-Guided Aggregation and Cross-Modality Refinement
abstract
Magnetic resonance imaging (MRI) can provide multi-modality MR images by setting task-specific scan parameters, and has been widely used in various disease diagnosis and planned treatments. However, in practical clinical applications, it is often difficult to obtain multi-modality MR images simultaneously due to patient discomfort, and scanning costs, etc. Therefore, how to effectively utilize the existing modality images to synthesize missing modality image has become a hot research topic. In this paper, we propose a novel confidence-guided aggregation and cross-modality refinement network (CACR-Net) for multi-modality MR image synthesis, which effectively utilizes complementary and correlative information of multiple modalities to synthesize high-quality target-modality images. Specifically, to effectively utilize the complementary modality-specific characteristics, a confidence-guided aggregation module is proposed to adaptively aggregate the multiple target-modality images generated from multiple source-modality images by using the corresponding confidence maps. Based on the aggregated target-modality image, a cross-modality refinement module is presented to further refine the target-modality image by mining correlative information among the multiple source-modality images and aggregated target-modality image. By training the proposed CACR-Net in an end-to-end manner, high-quality and sharp target-modality MR images are effectively synthesized. Experimental results on the widely used benchmark demonstrate that the proposed method outperforms state-of-the-art methods.
Bo Peng 0007, Bingzheng Liu, Yi Bin, Lili Shen, Jianjun Lei 0001
IEEE J. Biomed. Health Informatics2