Zhengda Lu

dblp:193/9259 · DBLP profile ↗
← Back
21ranked-venue papers
2as first author
20since 2021 · last 2026
0000-0002-9581-2268ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-author · 16 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DehazeGS: Seeing Through Fog with 3D Gaussian Splatting
abstract
Current novel view synthesis methods are typically designed for high-quality and clean input images. However, in foggy scenes, scattering and attenuation can significantly degrade the quality of rendering. Although NeRF-based dehazing approaches have been developed, their reliance on deep fully connected neural networks and per-ray sampling strategies leads to high computational costs. Furthermore, NeRF's implicit representation limits its ability to recover fine-grained details from hazy scenes. To overcome these limitations, we propose DehazeGS, the first physics-driven 3D Gaussian Splatting (3DGS) framework for dehazing. We adopt an explicit Gaussian representation to model fog formation via a physically consistent forward rendering process, enabling reconstruction and rendering of fog-free scenes using only multi-view foggy images as input. Specifically, based on the atmospheric scattering model, we simulate the formation of fog by establishing the transmission function directly on Gaussian primitives via depth-to-transmission mapping. During training, we jointly learn the atmospheric light and scattering coefficients while optimizing the Gaussian representation of foggy scenes. At inference time, we remove the effects of scattering and attenuation in Gaussian distributions and directly render the scene to obtain dehazed views. Experiments on both real-world and synthetic foggy datasets demonstrate that DehazeGS achieves state-of-the-art performance.
Yiqun Wang 0001, Aiheng Jiang, Zhengda Lu, Jianwei Guo 0003, Yong Li 0023, Hongxing Qin, Xiaopeng Zhang 0001
AAAI4
2025 An Industrial Multi-machining Feature Dataset and Contrastive Learning-Based Network for Feature Recognition
Haochen He, Zhengda Lu, Haiyong Jiang, Yiqun Wang 0001, Jun Xiao 0005
CGI (2)3
2025 Empowering Vector Graphics with Consistently Arbitrary Viewing and View-dependent Visibility
abstract
This work presents a novel text-to-vector graphics generation approach, Dream3DVG, allowing for arbitrary viewpoint viewing, progressive detail optimization, and view-dependent occlusion awareness. Our approach is a dual-branch optimization framework, consisting of an auxiliary 3D Gaussian Splatting optimization branch and a 3D vector graphics optimization branch. The introduced 3DGS branch can bridge the domain gaps between text prompts and vector graphics with more consistent guidance. Moreover, 3DGS allows for progressive detail control by scheduling classifier-free guidance, facilitating guiding vector graphics with coarse shapes at the initial stages and finer details at later stages. We also improve the view-dependent occlusions by devising a visibility-awareness rendering module. Extensive results on 3D sketches and 3D iconographies, demonstrate the superiority of the method on different abstraction levels of details, cross-view consistency, and occlusion-aware stroke culling. Code is available at https://github.com/chenxinl/Dream3DVG.git.
Jun Xiao 0005, Zhengda Lu, Yiqun Wang 0001, Haiyong Jiang
CVPR3
2025 LSF-Animation: Label-Free Speech-Driven Facial Animation via Implicit Feature Representation
abstract
Speech-driven 3D facial animation has attracted increasing interest since its potential to generate expressive and temporally synchronized digital humans. While recent works have begun to explore emotion-aware animation, they still depend on explicit one-hot encodings to represent identity and emotion with given emotion and identity labels, which limits their ability to generalize to unseen speakers. Moreover, the emotional cues inherently present in speech are often neglected, limiting the naturalness and adaptability of generated animations. In this work, we propose LSF-Animation, a novel framework that eliminates the reliance on explicit emotion and identity feature representations. Specifically, LSF-Animation implicitly extracts emotion information from speech and captures the identity features from a neutral facial mesh, enabling improved generalization to unseen speakers and emotional states without requiring manual labels. Furthermore, we introduce a Hierarchical Interaction Fusion Block (HIFB), which employs a fusion token to integrate dual transformer features and more effectively integrate emotional, motion-related and identity-related cues. Extensive experiments conducted on the 3DMEAD dataset demonstrate that our method surpasses recent state-of-the-art approaches in terms of emotional expressiveness, identity generalization, and animation realism. The source code will be released at: https://github.com/Dogter521/LSF-Animation.
Chuanqing Zhuang, Chenxi Jin, Zhengda Lu, Yiqun Wang 0001, Wu Liu 0005, Jun Xiao 0005
SIGGRAPH Asia4
2025 L2-GNN: Graph neural networks with fast spectral filters using twice linear parameterization
abstract
To improve learning on irregular 3D shapes, such as meshes with varying discretizations and point clouds with different samplings, we propose L 2 -GNN, a new graph neural network that approximates the spectral filters using twice linear parameterization. First, we parameterize the spectral filters using wavelet filter basis functions. The parameterization allows for an enlarged receptive field of graph convolutions, which can simultaneously capture low-frequency and high-frequency information. Second, we parameterize the wavelet filter basis functions using Chebyshev polynomial basis functions. This parameterization reduces the computational complexity of graph convolutions while maintaining robustness to the change of mesh discretization and point cloud sampling. Our L 2 -GNN based on the fast spectral filter can be used for shape correspondence, classification, and segmentation tasks on non-regular mesh or point cloud data. Experimental results show that our method outperforms the current state of the art in terms of both quality and efficiency.
Siying Huang, Zhengda Lu, Hongxing Qin, Huaiwen Zhang, Yiqun Wang 0001
Graph. Model.3
2025 BGPSeg: Boundary-Guided Primitive Instance Segmentation of Point Clouds
abstract
Point cloud primitive instance segmentation is critical for understanding the geometric shapes of man-made objects. Existing learning-based methods mainly focus on learning high-dimensional feature representations of points and further perform clustering or region growing to obtain corresponding primitive instances. However, these features generally cannot accurately represent the discriminability between instances, especially near the boundaries or in regions with small differences in geometric properties. This limitation often leads to over- or under-segmentation of geometric primitives. On the other hand, the boundaries of different primitives are the direct features that distinguish them and thus utilizing boundary information to guide feature learning and clustering is crucial for this task. In this paper, we propose a novel framework BGPSeg for point cloud primitive instance segmentation that utilizes boundary-guided feature extraction and clustering. Specifically, we first introduce a boundary-guided feature extractor with the additional input of a boundary probability map, which utilizes boundary-guided sampling and a boundary transformer to enhance feature discrimination among points crossing geometric boundaries. Furthermore, we propose a boundary-guided primitive clustering module, which combines boundary clues and geometric feature discrimination for clustering to further improve the segmentation performance. Finally, we demonstrate the effectiveness of our BGPSeg with a series of comparison and ablation experiments while achieving the state-of-the-art primitive instance segmentation. Our code is available at https://github.com/fz-20/BGPSeg.
Chuanqing Zhuang, Zhengda Lu, Yiqun Wang 0001, Lupeng Liu, Jun Xiao 0005
IEEE Trans. Image Process.3
2025 Exploring Structural Lines for Interior Floorplan Segmentation
Bingchen Yang, Haiyong Jiang, Zhengda Lu, Jun Xiao 0005
Vis. Comput.3
2024 SpikeGS: Learning 3D Gaussian Fields from Continuous Spike Stream
Xin Peng 0005, Zhengda Lu, Laurent Kneip, Yiqun Wang 0001
ACCV (10)3
2024 FC-4DFS: Frequency-controlled Flexible 4D Facial Expression Synthesizing
abstract
4D facial expression synthesizing is a critical problem in the fields of computer vision and graphics. Current methods lack flexibility and smoothness when simulating the inter-frame motion of expression sequences. In this paper, we propose a frequency-controlled 4D facial expression synthesizing method, FC-4DFS. Specifically, we introduce a frequency-controlled LSTM network to generate 4D facial expression sequences frame by frame from a given neutral landmark with a given length. Meanwhile, we propose a temporal coherence loss to enhance the perception of temporal sequence motion and improve the accuracy of relative displacements. Furthermore, we designed a Multi-level Identity-Aware Displacement Network based on a cross-attention mechanism to reconstruct the 4D facial expression sequences from landmark sequences. Finally, our FC-4DFS achieves flexible and SOTA generation results of 4D facial expression sequences with different lengths on CoMA and Florence4D datasets. The code will be available on GitHub.
Chuanqing Zhuang, Zhengda Lu, Yiqun Wang 0001, Jun Xiao 0005
ACM Multimedia3
2024 SDFReg: Learning Signed Distance Functions for Point Cloud Registration
Leida Zhang, Zhengda Lu, Yiqun Wang 0001
PRCV (6)2
2024 DepthGAN: GAN-based depth generation from semantic layouts
abstract
Existing GAN-based generative methods are typically used for semantic image synthesis. We pose the question of whether GAN-based architectures can generate plausible depth maps and find that existing methods have difficulty in generating depth maps which reasonably represent 3D scene structure due to the lack of global geometric correlations. Thus, we propose DepthGAN, a novel method of generating a depth map using a semantic layout as input to aid construction, and manipulation of well-structured 3D scene point clouds. Specifically, we first build a feature generation model with a cascade of semantically-aware transformer blocks to obtain depth features with global structural information. For our semantically aware transformer block, we propose a mixed attention module and a semantically aware layer normalization module to better exploit semantic consistency for depth features generation. Moreover, we present a novel semantically weighted depth synthesis module, which generates adaptive depth intervals for the current scene. We generate the final depth map by using a weighted combination of semantically aware depth weights for different depth ranges. In this manner, we obtain a more accurate depth map. Extensive experiments on indoor and outdoor datasets demonstrate that DepthGAN achieves superior results both quantitatively and visually for the depth generation task.
Jun Xiao 0005, Yiqun Wang 0001, Zhengda Lu
Comput. Vis. Media4
2024 Efficient Single Correspondence Voting for Point Cloud Registration
abstract
3D point cloud registration is a crucial task in a variety of fields, including remote sensing mapping, computer vision, virtual reality, and autonomous driving. However, this task is still challenging due to the challenges of noise, non-uniformity, partial overlap, and repeated local features in large scene point clouds. In this paper, we propose an efficient single correspondence voting method for large scene point cloud registration. Specifically, we first propose an efficient hypothetical transformation prediction method called SCVC, which determines the 5 degrees of freedom of the transformation through one correspondence, and then uses Hough voting to determine the last degree of freedom. This algorithm can significantly improve the accuracy of registration in both indoor and outdoor scenes. On the other hand, we propose a more robust transformation verification function called VDIR, which can obtain the optimal registration result of two raw point clouds. Finally, we conduct a series of experiments that demonstrate that our method achieves state-of-the-art performance on four real-world datasets: 3DMatch, 3DLoMatch, KITTI, and WHU-TLS. Our code is available at https://github.com/xingxuejun1989/SCVC.
Xuejun Xing, Zhengda Lu, Yiqun Wang 0001, Jun Xiao 0005
IEEE Trans. Image Process.2
2023 RC-Net: Row and Column Network with Text Feature for Parsing Floor Plan Images
Weiliang Meng, Zhengda Lu, Jianwei Guo 0003, Jun Xiao 0005, Xiaopeng Zhang 0001
J. Comput. Sci. Technol.3
2023 SPDET: Edge-Aware Self-Supervised Panoramic Depth Estimation Transformer With Spherical Geometry
abstract
Panoramic depth estimation has become a hot topic in 3D reconstruction techniques with its omnidirectional spatial field of view. However, panoramic RGB-D datasets are difficult to obtain due to the lack of panoramic RGB-D cameras, thus limiting the practicality of supervised panoramic depth estimation. Self-supervised learning based on RGB stereo image pairs has the potential to overcome this limitation due to its low dependence on datasets. In this work, we propose the SPDET, an edge-aware self-supervised panoramic depth estimation network that combines the transformer with a spherical geometry feature. Specifically, we first introduce the panoramic geometry feature to construct our panoramic transformer and reconstruct high-quality depth maps. Furthermore, we introduce the pre-filtered depth-image-based rendering method to synthesize the novel view image for self-supervision. Meanwhile, we design an edge-aware loss function to improve the self-supervised depth estimation for panorama images. Finally, we demonstrate the effectiveness of our SPDET with a series of comparison and ablation experiments while achieving the state-of-the-art self-supervised monocular panoramic depth estimation. Our code and models are available at https://github.com/zcq15/SPDET.
Chuanqing Zhuang, Zhengda Lu, Yiqun Wang 0001, Jun Xiao 0005, Ying Wang 0030
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 NR-MVSNet: Learning Multi-View Stereo Based on Normal Consistency and Depth Refinement
abstract
Multi-view Stereo (MVS) aims to reconstruct a 3D point cloud model from multiple views. In recent years, learning-based MVS methods have received a lot of attention and achieved excellent performance compared with traditional methods. However, these methods still have apparent shortcomings, such as the accumulative error in the coarse-to-fine strategy and the inaccurate depth hypotheses based on the uniform sampling strategy. In this paper, we propose the NR-MVSNet, a coarse-to-fine structure with the depth hypotheses based on the normal consistency (DHNC) module, and the depth refinement with reliable attention (DRRA) module. Specifically, we design the DHNC module to generate more effective depth hypotheses, which collects the depth hypotheses from neighboring pixels with the same normals. As a result, the predicted depth can be smoother and more accurate, especially in texture-less and repetitive-texture regions. On the other hand, we update the initial depth map in the coarse stage by the DRRA module, which can combine attentional reference features and cost volume features to improve the depth estimation accuracy in the coarse stage and address the accumulative error problem. Finally, we conduct a series of experiments on the DTU, BlendedMVS, Tanks & Temples, and ETH3D datasets. The experimental results demonstrate the efficiency and robustness of our NR-MVSNet compared with the state-of-the-art methods. Our implementation is available at https://github.com/wdkyh/NR-MVSNet.
Jingliang Li, Zhengda Lu, Yiqun Wang 0001, Jun Xiao 0005, Ying Wang 0030
IEEE Trans. Image Process.2
2022 ACDNet: Adaptively Combined Dilated Convolution for Monocular Panorama Depth Estimation
abstract
Depth estimation is a crucial step for 3D reconstruction with panorama images in recent years. Panorama images maintain the complete spatial information but introduce distortion with equirectangular projection. In this paper, we propose an ACDNet based on the adaptively combined dilated convolution to predict the dense depth map for a monocular panoramic image. Specifically, we combine the convolution kernels with different dilations to extend the receptive field in the equirectangular projection. Meanwhile, we introduce an adaptive channel-wise fusion module to summarize the feature maps and get diverse attention areas in the receptive field along the channels. Due to the utilization of channel-wise attention in constructing the adaptive channel-wise fusion module, the network can capture and leverage the cross-channel contextual information efficiently. Finally, we conduct depth estimation experiments on three datasets (both virtual and real-world) and the experimental results demonstrate that our proposed ACDNet substantially outperforms the current state-of-the-art (SOTA) methods. Our codes and model parameters are accessed in https://github.com/zcq15/ACDNet.
Chuanqing Zhuang, Zhengda Lu, Yiqun Wang 0001, Jun Xiao 0005, Ying Wang 0030
AAAI2
2022 DS-MVSNet: Unsupervised Multi-view Stereo via Depth Synthesis
abstract
In recent years, supervised or unsupervised learning-based MVS methods achieved excellent performance compared with traditional methods. However, these methods only use the probability volume computed by cost volume regularization to predict reference depths and this manner cannot mine enough information from the probability volume. Furthermore, the unsupervised methods usually try to use two-step or additional inputs for training which make the procedure more complicated. In this paper, we propose the DS-MVSNet, an end-to-end unsupervised MVS structure with the source depths synthesis. To mine the information in probability volume, we creatively synthesize the source depths by splattering the probability volume and depth hypotheses to source views. Meanwhile, we propose the adaptive Gaussian sampling and improved adaptive bins sampling approach that improve the depths hypotheses accuracy. On the other hand, we utilize the source depths to render the reference images and propose depth consistency loss and depth smoothness loss. These can provide additional guidance according to photometric and geometric consistency in different views without additional inputs. Finally, we conduct a series of experiments on the DTU dataset and Tanks $&$ Temples dataset that demonstrate the efficiency and robustness of our DS-MVSNet compared with the state-of-the-art methods.
Jingliang Li, Zhengda Lu, Yiqun Wang 0001, Ying Wang 0030, Jun Xiao 0005
ACM Multimedia2
2022 A Novel Rock-Mass Point Cloud Registration Method Based on Feature Line Extraction and Feature Point Matching
abstract
Registration will directly affect the quality of overall rock-mass point cloud, which is the basis of 3-D reconstruction for rock mass. Advanced methods establish correspondence by extracting various features that remain unchanged. Although these methods have made great progress, they analyze the local characteristics of each sample point, which leads to be inefficient. In this article, we select registration interesting points from feature lines that were extracted based on supervoxel and innovatively introduce the “clustering, primary matching, and coarse registration” strategy, which effectively reduces the complexity of calculating the corresponding relationship during point cloud registration. Finally, the iterative closest point (ICP) algorithm is used to optimize the result of coarse registration. By selecting registration interesting points from the extracted feature lines, the proposed method inherits the robustness of feature lines to noise, initial position, and so on. The experimental results prove that the coarse registration and refined registration results of the proposed method both have high accuracy and efficiency.
Lupeng Liu, Jun Xiao 0005, Yunbiao Wang, Zhengda Lu, Ying Wang 0030
IEEE Trans. Geosci. Remote. Sens.4
2021 Extracting Cycle-aware Feature Curve Networks from 3D Models
Zhengda Lu, Jianwei Guo 0003, Jun Xiao 0005, Ying Wang 0030, Xiaopeng Zhang 0001, Dong-Ming Yan 0001
Comput. Aided Des.1
2021 Data-driven floor plan understanding in rural residential buildings via deep recognition
Zhengda Lu, Jianwei Guo 0003, Weiliang Meng, Jun Xiao 0005, Xiaopeng Zhang 0001
Inf. Sci.1
2020 Quasi Fourier-Mellin Transform for Affine Invariant Features
abstract
Fourier-Mellin transform (FMT) has been widely used for the extraction of rotation- and scale-invariant features. However, affine transform is a more reasonable approximation model for real viewpoint change. Due to shearing, the integral along the angular direction in the calculation of FMT cannot be used to extract the inherent features of an image undergoing affine transform. To eliminate the effect of shearing, whitening transform should be conducted on the integral along the radial direction. FMT can hardly be modified by conventional whitening-based methods with low computational cost due to additional processes. In this paper, two factors are constructed and embedded into FMT. Quasi Fourier-Mellin transform (QFMT) is proposed. The embedding of these factors is equivalent to whitening transform and can eliminate the effect of shearing in the affine transform. In particular, QFMT can also be calculated by integrating along the radial direction followed by integrating along the angular direction, as in FMT. Based on QFMT, the quasi Fourier-Mellin descriptor (QFMD) is constructed for the extraction of affine invariant features. Some experiments have also been conducted to test the performance of the proposed method.
Zhengda Lu, Yuan Yan Tang, Zhou Yuan
IEEE Trans. Image Process.2