Xianwei Zheng

dblp:146/8854 · DBLP profile ↗
← Back
45ranked-venue papers
6as first author
36since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 2 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 2 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2027 Topological signal processing over product cell complexes
Xiaona Zeng, Xianwei Zheng, Jiantao Zhou 0001, Xutao Li 0004
Signal Process.2
2026 Deferred Poisoning: Making the Model More Vulnerable via Hessian Singularization
abstract
Recent studies have shown that deep learning models are very vulnerable to poisoning attacks. Many defense methods have been proposed to address this issue. However, traditional poisoning attacks are not as threatening as commonly believed. This is because they often cause differences in how the model performs on the training set compared to the validation set. Such inconsistency can alert defenders that their data has been poisoned, allowing them to take the necessary defensive actions. In this paper, we introduce a more threatening type of poisoning attack called the Deferred Poisoning Attack. This new attack allows the model to function normally during the training and validation phases but makes it very sensitive to evasion attacks or even natural noise. We achieve this by ensuring the poisoned model's loss function has a similar value as a normally trained model at each input sample but with a large local curvature. A similar model loss ensures that there is no obvious inconsistency between the training and validation accuracy, demonstrating high stealthiness. On the other hand, the large curvature implies that a small perturbation may cause a significant increase in model loss, leading to substantial performance degradation, which reflects a worse robustness. We fulfill this purpose by making the model have singular Hessian information at the optimal point via our proposed Singularization Regularization term. We have conducted both theoretical and empirical analyses of the proposed method and validated its effectiveness through experiments on image classification tasks. Furthermore, we have confirmed the hazards of this form of poisoning attack under more general scenarios using natural noise, offering a new perspective for research in the field of security.
Yuhao He 0001, Jinyu Tian 0001, Xianwei Zheng, Li Dong 0006, Yuanman Li, Jiantao Zhou 0001
AAAI3
2026 Real-time 3D Object Detection with Inference-Aligned Learning
abstract
Real-time 3D object detection from point clouds is essential for dynamic scene understanding in applications such as augmented reality, robotics, and navigation. We introduce a novel Spatial-prioritized and Rank-aware 3D object detection (SR3D) framework for indoor point clouds, to bridge the gap between how detectors are trained and how they are evaluated. This gap stems from the lack of spatial reliability and ranking awareness during training, which conflicts with the ranking-based prediction selection used at inference. Such a training-inference gap hampers the model’s ability to learn representations aligned with inference-time behavior. To address the limitation, SR3D consists of two components tailored to the spatial nature of point clouds during training: a novel spatial-prioritized optimal transport assignment that dynamically emphasizes well-located and spatially reliable samples, and a rank-aware adaptive self-distillation scheme that adaptively injects ranking perception via a self-distillation paradigm. Extensive experiments on ScanNet V2 and SUN RGB-D show that SR3D effectively bridges the training-inference gap and significantly outperforms prior methods in accuracy while maintaining real-time speed.
Xianwei Zheng, Zimin Xia, Linwei Yue, Nan Xue 0001
AAAI2
2026 Seeing Through Satellite Images at Street Views
abstract
This paper studies the task of SatStreet-view synthesis, which aims to render photorealistic street-view panorama images and videos given a satellite image and specified camera positions or trajectories. Our approach involves learning a satellite image conditioned neural radiance field from paired images captured from both satellite and street viewpoints, which comes to be a challenging learning problem due to the sparse-view nature and the extremely large viewpoint changes between satellite and street-view images. We tackle the challenges based on a task-specific observation that street-view specific elements, including the sky and illumination effects, are only visible in street-view panoramas, and present a novel approach, Sat2Density++, to accomplish the goal of photo-realistic street-view panorama rendering by modeling these street-view specific elements in neural networks. In the experiments, our method is evaluated on both urban and suburban scene datasets, demonstrating that Sat2Density++ is capable of rendering photorealistic street-view panoramas that are consistent across multiple views and faithful to the satellite image.
Ming Qian, Bin Tan 0002, Qiuyu Wang, Xianwei Zheng, Hanjiang Xiong, Gui-Song Xia, Yujun Shen, Nan Xue 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Topological Feature Fusion for Multi-Channel No-Reference Point Cloud Quality Assessment
Dianjun Xu, Aizhong Peng, Qiman Zhong, Xianwei Zheng
IEEE Signal Process. Lett.5
2025 ScaleLSD: Scalable Deep Line Segment Detection Streamlined
abstract
This paper studies the problem of Line Segment Detection (LSD) for the characterization of line geometry in images, with the aim of learning a domain-agnostic robust LSD model that works well for any natural images. With the focus of scalable self-supervised learning of LSD, we revisit and streamline the fundamental designs of (deep and non-deep) LSD approaches to have a high-performing and efficient LSD learner, dubbed as ScaleLSD, for the curation of line geometry at scale from over 10M unlabeled real-world images. Our ScaleLSD works very well to detect much more number of line segments from any natural images even than the pioneered non-deep LSD approach, having a more complete and accurate geometric characterization of images using line segments. Experimentally, our proposed ScaleLSD is comprehensively testified under zero-shot protocols in detection performance, single-view 3D geometry estimation, two-view line segment matching, and multiview 3D line mapping, all with excellent performance obtained. Based on the thorough evaluation, our ScaleLSD is observed to be the first deep approach that outperforms the pioneered non-deep LSD in all aspects we have tested, significantly expanding and reinforcing the versatility of the line geometry of images.
Zeran Ke, Bin Tan 0002, Xianwei Zheng, Yujun Shen, Tianfu Wu 0001, Nan Xue 0001
CVPR3
2025 BWFormer: Building Wireframe Reconstruction from Airborne LiDAR Point Cloud with Transformer
abstract
In this paper, we present BWFormer, a novel Transformerbased model for building wireframe reconstruction from airborne LiDAR point cloud. The problem is solved in a ground-up manner here by detecting the building corners in 2D, lifting and connecting them in 3D space afterwards with additional data augmentation. Due to the 2.5D characteristic of the airborne LiDAR point cloud, we simplify the problem by projecting the points on the ground plane to produce a 2D height map. With the height map, a heat map is first generated with pixel-wise corner likelihood to predict the possible 2D corners. Then, 3D corners are predicted by a Transformer-based network with extra height embedding initialization. This 2D-to-3D corner detection strategy reduces the search space significantly. To recover the topological connections among the corners, edges are finally predicted from the height map with the proposed edge attention mechanism, which extracts holistic features and preserves local details simultaneously. In addition, due to the limited datasets in the field and the irregularity of the point clouds, a conditional latent diffusion model for LiDAR scanning simulation is utilized for data augmentation. BW-Former surpasses other state-of-the-art methods, especially in reconstruction completeness. Our code is available at: https : //github.com/3dv-casia/BWformer/.
Lingjie Zhu, Hanqiao Ye, Shangfeng Huang, Xiang Gao 0009, Xianwei Zheng, Shuhan Shen
CVPR6
2025 CoMatcher: Multi-View Collaborative Feature Matching
abstract
This paper proposes a multi-view collaborative matching strategy for reliable track construction in complex scenarios. We observe that the pairwise matching paradigms applied to image set matching often result in ambiguous estimation when the selected independent pairs exhibit significant occlusions or extreme viewpoint changes. This challenge primarily stems from the inherent uncertainty in interpreting intricate 3D structures based on limited two-view observations, as the 3D-to-2D projection leads to significant information loss. To address this, we introduce CoMatcher, a deep multi-view matcher to (i) leverage complementary context cues from different views to form a holistic 3D scene understanding and (ii) utilize cross-view projection consistency to infer a reliable global solution. Building on CoMatcher, we develop a groupwise framework that fully exploits cross-view relationships for large-scale matching tasks. Extensive experiments on various complex scenarios demonstrate the superiority of our method over the mainstream two-view matching paradigm.
Zimin Xia, Mingyue Dong, Shuhan Shen, Linwei Yue, Xianwei Zheng
CVPR6
2025 PointSC: A Novel Simplicial Complex-Based Neural Network for Point Cloud Classification
Xuran Yao, Xianwei Zheng
ICONIP (2)3
2025 MuSPaCSA: Multi-Scale Parallel-Channel Self-Attention Network for Point Cloud Classification and Segmentation
abstract
Point cloud classification and segmentation are fundamental tasks in 3D computer vision. Recently, deep learning-based methods, particularly 3D Transformers, have demonstrated their effectiveness across a variety of point cloud tasks. However, transformer-based methods embed position information into feature vectors, which can introduce a significant computational cost. Additionally, these approaches often struggle to adaptively extract different features across varying receptive fields, which limits their performance in various tasks. To address these challenges, we propose a novel Multi-Scale Parallel-Channel Self-Attention (MuSPaCSA) network, designed with a multi-scale feature extraction architecture by stacking Parallel-Channel Self-Attention (PaCSA) layers for classification and segmentation tasks. Specifically, our MuSPaCSA employs the PaCSA module to extract essential semantic and spatial features. The core components of the PaCSA module include the Semantic-Spatial Integration (SSI) and Adaptive Self-Attention (ASA) modules. The SSI module employs a parallel-channel approach to integrate semantic and spatial information, enabling the representation of high-dimensional structural features in point clouds. The ASA module calculates adaptive weights to aggregate rich, high-dimensional structural features from neighboring nodes in a lightweight manner. Through the multi-scale feature fusion architecture of MuSPaCSA, local and global features, as well as semantic and spatial features, are effectively integrated, significantly enhancing the model’s representational capacity. Extensive experiments demonstrate that our model achieves superior performance and results with lower computational cost compared to competing methods.
Xuran Yao, Xianwei Zheng
IROS2
2025 TRHCN: A Topology-Based Region Hierarchical Convolutional Network for EEG Emotion Recognition
Xuran Yao, Xianwei Zheng, Xutao Li 0004
PRCV (2)3
2025 Shape Activated CAM Learning for Weakly Supervised Remote Sensing Semantic Segmentation
abstract
Class activation map (CAM) based weakly-supervised semantic segmentation (WSSS) of remote sensing (RS) images has attracted extensive research interests for its potential in reducing annotation cost. However, challenged by unconstrained activation issue, existing methods struggle to delineate object boundaries clearly, making them particularly difficult to separate multiple densely packed objects, which are common in RS images. By conducting an in-depth analysis of RS image characteristics, we observed a strong correlation between object shapes and their semantics. Inspired by this finding, we propose an Intrinsic Shape Activation Network (ISANet) to learn the category-relevant shape priors as geometry constraints for target-focused region activation in WSSS of RS images. The key idea is to distill the intrinsic shape priors from the hybrid features that are deterministic in classification. Specifically, we adopt a dual-branch architecture to decouple the learning of shape and texture features and leverage a shape awareness alignment module to generate boundary-clear CAMs for computing pseudo labels. In this way, CAMs are generated with perception of target shapes, which increases the completeness of activation regions and alleviates the ultrarange responses. Extensive experiments demonstrates the superiority of our method in delineating densely-packed objects with clear contours, which is especially beneficial for separating multiple targets in RS images. Our method improves the mIoU of the state-of-the-art method by 7.9% and 3.3% on the NWPU VHR-10 and iSAID dataset respectively.
He Chen 0004, Mingyue Dong, Linwei Yue, Xianwei Zheng, Jun Li 0009, Jianya Gong
IEEE Trans. Geosci. Remote. Sens.4
2025 Holistic Response Lifting for Weakly Supervised Land-Cover Classification
Qiyuan Ma, Xianwei Zheng, Linxi Huan, Linwei Yue, Gui-Song Xia, Jianya Gong
IEEE Trans. Geosci. Remote. Sens.2
2025 DuPMAM: An Efficient Dual Perception Framework Equipped With a Sharp Testing Strategy for Point Cloud Analysis
abstract
The challenges in point cloud analysis are primarily attributed to the irregular and unordered nature of the data. Numerous existing approaches, inspired by the Transformer, introduce attention mechanisms to extract the 3D geometric features. However, these intricate geometric extractors incur high computational overhead and unfavorable inference latency. To tackle this predicament, in this paper, we propose a lightweight and faster attention-based network, named Dual Perception MAM (DuPMAM), for point cloud analysis. Specifically, we present a novel simple Point Multiplicative Attention Mechanism (PMAM). It is implemented solely through single feed-forward fully connected layers, hence leading to lower model complexity and superior inference speed. Based on that, we further devise a dual perception strategy by constructing both a local attention block and a global attention block to learn fine-grained geometric and overall representational features, respectively. Consequently, compared to the existing approaches, our method has excellent perception of local details and global contours of the point cloud objects. In addition, we ingeniously design a Graph-Multiscale Perceptual Field (GMPF) testing strategy for model performance enhancement. It has significant advantage over the traditional voting strategy and is generally applicable to point cloud tasks, encompassing classification, part segmentation and indoor scene segmentation. Empowered by the GMPF testing strategy, DuPMAM delivers the new State-of-the-Art on the real-world dataset ScanObjectNN, the synthetic dataset ModelNet40 and the part segmentation dataset ShapeNet, and compared to the recent GB-Net, our DuPMAM trains 6 times faster and tests 2 times faster.
Xianwei Zheng, Zhulun Yang, Xutao Li 0004, Jiantao Zhou 0001, Yuanman Li
IEEE Trans. Multim.2
2025 Tiny Data Is Sufficient: A Generalizable CNN Architecture for Temporal Domain Long Sequence Identification
abstract
Deep learning (DL) models have made remarkable progress in various sequence processing tasks. It is widely acknowledged that these models heavily rely on numerous training data and finely tuned parameters. Recent studies highlighted that conventional convolutions in deep networks may hamper feature processing efficacy, particularly in long temporal sequence analysis, due to their limited feature representation capabilities. To tackle these challenges, this article introduces a novel generalizable convolutional neural network (GeCNN) architecture tailored for temporal domain long sequence identification. Our framework incorporates three key components: the generic convolutional neural network (CNN), selective CNN, and multiple pooling layers. The generic CNN implements customizable hyper-convolutional operations through non-linear convolvers, thereby enhancing feature representation effectiveness and significantly improving accuracy. Subsequently, the selective CNN is designed to abate the demand for large training data by focusing on various subsequences. We propose the homogeneous striding principle and the partial homogeneous striding theorem to theoretically support the method. The multiple pooling combines eight distinct pooling operations to mitigate the statistical information loss problem typically associated with single pooling actions. Experimental results demonstrate that our GeCNN architecture achieves superior performance with shallow networks and tiny data compared to existing deep networks. The accuracy of the best-trained model surpasses the ResNet and self-attention-based models by 9.51% and 16.79% utilizing only 0.18% of data for training in the GTZAN dataset. Meanwhile, the accuracy of the optimal model overtakes the other two models by 5.35% and 10.16% while using merely 1.56% of data for training in the PLAID dataset.
Chen Li 0058, Xianwei Zheng, Chuangquan Chen, Zicong Deng, Yiqing Shu
IEEE Trans. Neural Networks Learn. Syst.2
2024 Learning Deformable Hypothesis Sampling for Accurate PatchMatch Multi-View Stereo
abstract
This paper introduces a learnable Deformable Hypothesis Sampler (DeformSampler) to address the challenging issue of noisy depth estimation in faithful PatchMatch multi-view stereo (MVS). We observe that the heuristic depth hypothesis sampling modes employed by PatchMatch MVS solvers are insensitive to (i) the piece-wise smooth distribution of depths across the object surface and (ii) the implicit multi-modal distribution of depth prediction probabilities along the ray direction on the surface points. Accordingly, we develop DeformSampler to learn distribution-sensitive sample spaces to (i) propagate depths consistent with the scene's geometry across the object surface and (ii) fit a Laplace Mixture model that approaches the point-wise probabilities distribution of the actual depths along the ray direction. We integrate DeformSampler into a learnable PatchMatch MVS system to enhance depth estimation in challenging areas, such as piece-wise discontinuous surface boundaries and weakly-textured regions. Experimental results on DTU and Tanks & Temples datasets demonstrate its superior performance and generalization capabilities compared to state-of-the-art competitors. Code is available at https://github.com/Geo-Tell/DS-PMNet.
Xianwei Zheng, Hanjiang Xiong
AAAI3
2024 Stratified Avatar Generation from Sparse Observations
abstract
Estimating 3D full-body avatars from AR/VR devices is essential for creating immersive experiences in AR/VR Applications. This task is challenging due to the limited in-put from Head Mounted Devices, which capture only sparse observations from the head and hands. Predicting the full-body avatars, particularly the lower body, from these sparse observations presents significant difficulties. In this paper, we are inspired by the inherent property of the kinematic tree defined in the Skinned Multi-Person Linear (SMPL) model, where the upper body and lower body share only one common ancestor node, bringing the potential of de-coupled reconstruction. We propose a stratified approach to decouple the conventional full-body avatar reconstruction pipeline into two stages, with the reconstruction of the up-per body first and a subsequent reconstruction of the lower body conditioned on the previous stage. To implement this straightforward idea, we leverage the latent diffusion model as a powerful probabilistic generator, and train it to fol-low the latent distribution of decoupled motions explored by a VQ-VAE encoder-decoder model. Extensive experiments on AMASS mocap dataset demonstrate our state-of-the-art performance in the reconstruction of full-body motions.
Quankai Gao, Xianwei Zheng, Nan Xue 0001
CVPR4
2024 PanoPose: Self-supervised Relative Pose Estimation for Panoramic Images
abstract
Scaled relative pose estimation, i.e., estimating relative rotation and scaled relative translation between two images, has always been a major challenge in global Structure-from-Motion (SfM). This difficulty arises because the two-view relative translation computed by traditional geometric vision methods, e.g. the five-point algorithm, is scaleless. Many researchers have proposed diverse translation averaging methods to solve this problem. Instead of solving the problem in the motion averaging phase, we focus on estimating scaled relative pose with the help of panoramic cameras and deep neural networks. In this paper, a novel network, namely PanoPose, is proposed to estimate the relative motion in a fully self-supervised manner and a global SfM pipeline is built for panorama images. The proposed PanoPose comprises a depth-net and a pose-net, with self-supervision achieved by reconstructing the reference image from its neighboring images based on the estimated depth and relative pose. To maintain precise pose estimation under large viewing angle differences, we randomly rotate the panoramic images and pre-train the posenet with images before and after the rotation. To enhance scale accuracy, a fusion block is introduced to incorporate depth information into pose estimation. Extensive experiments on panoramic SfM datasets demonstrate the effectiveness of PanoPose compared with state-of-the-arts.
Diantao Tu, Hainan Cui, Xianwei Zheng, Shuhan Shen
CVPR3
2024 Easing 3D Pattern Reasoning with Side-View Features for Semantic Scene Completion
Linxi Huan, Mingyue Dong, Linwei Yue, Shuhan Shen, Xianwei Zheng
ECCV (69)5
2024 PolyRoom: Room-Aware Transformer for Floorplan Reconstruction
Lingjie Zhu, Hanqiao Ye, Xiang Gao 0009, Xianwei Zheng, Shuhan Shen
ECCV (50)6
2024 Alleviating Semantic Uncertainty in Cams for Weakly-Supervised Land-Cover Classification
abstract
Image-level weakly-supervised land-cover classification (WSLC) relies on image labels to free land-cover classification networks from expensive pixel-level annotations. Existing weakly-supervised methods commonly generate class activation maps (CAMs) from a trained classification network to produce pixel-level pseudo labels for training land-cover classification networks. However, CAMs struggle to completely and accurately capture land-covers in remote sensing images due to the lack of precise localization information in weak supervision. To alleviate semantic uncertainty in CAMs, this paper proposes an Expanding-Separating (ES) scheme to generate high-quality CAMs of land-covers. The ES scheme expands activation areas in CAMs by lifting activation values in inactivated regions and suppresses CAM noises by decreasing values in confusing regions. Experimental results on the DeepGlobe dataset demonstrate that the proposed ES scheme significantly improves the quality of CAMs and benefits the training of land-cover classification, achieving state-of-the-art performance.
Qiyuan Ma, Linxi Huan, Mingyue Dong, Xianwei Zheng, Linwei Yue
IGARSS4
2024 Generative DEM Void Filling With Terrain Feature-Guided Transfer Learning Assisted by Remote Sensing Images
abstract
The quality of digital elevation models (DEMs) is easily affected by data voids in regions with complex terrain conditions. Numerous methods have been proposed to fill DEM voids by effectively exploiting the topographic information from neighboring areas or auxiliary DEMs. However, few studies have considered the integration of multi-modal data, which can provide valuable supplementary information in the areas with no high-quality reference DEM data. In this letter, we propose a generative DEM void filling method by exploring the integration of optical remote sensing images. The core idea is to utilize the image textures to infer the elevation values in the void regions with terrain texture-guided transfer learning. Specifically, the image context attention module (ICAM) is used to preliminarily estimate the missing topographic features by searching the similar patches with the guidance of image context. The terrain feature-guided residual pixel attention block (TFG-RPAB) is then employed to refine the void-filled features by transferring the image textures to topographic features. Finally, the void-filled DEM can be obtained by decoding the reconstructed topographic features. The results shows that the RMSE of RSAGAN is improved by 14.5% to 71.5% when DEM void filling. Both quantitative and qualitative evaluations demonstrate the superiority of the proposed method over the competitive methods in terms of DEM void filling. The source code is available at https://github.com/gaobingcug/RSAGAN.
Linwei Yue, Xianwei Zheng
IEEE Geosci. Remote. Sens. Lett.3
2024 Spatiotemporal Gated Graph Transformer for EEG-Based Emotion Recognition
abstract
The availability of accurate and reliable electroencephalography (EEG) signal data makes emotion recognition feasible. In recent years, an increasing number of deep learning methods have been applied to emotion recognition tasks based on EEG signals, but they all have their own drawbacks, including the inability to simultaneously capture spatial and temporal information, the loss of spatial information, and even the improper selection of data scales. To address the above problems, this paper proposes a novel method for EEG-based emotion recognition named Spatiotemporal Gated Graph Transformer (SGGT). The method takes advantage of graph structure to enrich spatial information and uses a two-tower transformer to encode spatial information and temporal information, respectively. For the spatial feature extraction, we adopt four encoding methods to compensate for the shortcomings of traditional transformers in graph learning tasks. Different from the previous processing methods, our method uses graph pooling to compress the graph, forcing the model to learn the relationship between different time stamps. In addition, a gating mechanism is used to merge the two-tower transformer to realize the fusion of temporal and spatial features. Our method is validated on the SEED, SEED-IV and SEED-V datasets, and it outperforms the baseline methods.
Yadong Chang, Xianwei Zheng, Xutao Li 0004
IEEE Signal Process. Lett.2
2023 Shape Anchor Guided Holistic Indoor Scene Understanding
abstract
This paper proposes a shape anchor guided learning strategy (AncLearn) for robust holistic indoor scene under-standing. We observe that the search space constructed by current methods for proposal feature grouping and instance point sampling often introduces massive noise to instance detection and mesh reconstruction. Accordingly, we develop AncLearn to generate anchors that dynamically fit instance surfaces to (i) unmix noise and target-related features for offering reliable proposals at the detection stage, and (ii) reduce outliers in object point sampling for directly providing well-structured geometry priors without segmentation during reconstruction. We embed AncLearn into a reconstruction-from-detection learning system (AncRec) to generate high-quality semantic scene models in a purely instance-oriented manner. Experiments conducted on the challenging ScanNetv2 dataset demonstrate that our shape anchor-based method consistently achieves state-of-the-art performance in terms of 3D object detection, layout estimation, and shape reconstruction. The code will be available at https://github.com/Geo-Tell/AncRec.
Mingyue Dong, Linxi Huan, Hanjiang Xiong, Shuhan Shen, Xianwei Zheng
ICCV5
2023 Holistic Geometric Feature Learning for Structured Reconstruction
abstract
The inference of topological principles is a key problem in structured reconstruction. We observe that wrongly predicted topological relationships are often incurred by the lack of holistic geometry clues in low-level features. Inspired by the fact that massive signals can be compactly described with frequency analysis, we experimentally explore the efficiency and tendency of learning structure geometry in the frequency domain. Accordingly, we propose a frequency-domain feature learning strategy (F-Learn) to fuse scattered geometric fragments holistically for topology-intact structure reasoning. Benefiting from the parsimonious design, the F-Learn strategy can be easily deployed into a deep reconstructor with a lightweight model modification. Experiments demonstrate that the F-Learn strategy can effectively introduce structure awareness into geometric primitive detection and topology inference, bringing significant performance improvement to final structured reconstruction. Code and pre-trained models are available at https://github.com/Geo-Tell/F-Learn.
Ziqiong Lu, Linxi Huan, Qiyuan Ma, Xianwei Zheng
ICCV4
2022 HoW-3D: Holistic 3D Wireframe Perception from a Single Image
abstract
This paper studies the problem of holistic 3D wireframe perception (HoW-3D), a new task of perceiving both the visible 3D wireframes and the invisible ones from single-view 2D images. As the non-front surfaces of an object cannot be directly observed in a single view, estimating the nonline-of-sight (NLOS) geometries in HoW-3D is a fundamentally challenging problem and remains open in computer vision. We study the problem of HoW-3D by proposing an ABC-HoW benchmark, which is created on top of CAD models sourced from the ABC-dataset with 12k single-view images and the corresponding holistic 3D wireframe models. With our large-scale ABC-HoW benchmark available, we present a novel Deep Spatial Gestalt (DSG) model to learn the visible junctions and line segments as the basis and then infer the NLOS 3D structures from the visible cues by following the Gestalt principles of human vision systems. In our experiments, we demonstrate that our DSG model performs very well in inferring the holistic 3D wireframes from single-view images. Compared with the strong baseline methods, our DSG model outperforms the previous wire-frame detectors in detecting the invisible line geometry in single-view images and is even very competitive with prior arts that take high-fidelity PointCloud as inputs on reconstructing 3D wireframes.
Bin Tan 0002, Nan Xue 0001, Tianfu Wu 0001, Xianwei Zheng, Gui-Song Xia
3DV5
2022 PointFormer: A Dual Perception Attention-Based Network for Point Cloud Classification
Zhulun Yang, Xianwei Zheng, Yadong Chang
ACCV (1)3
2022 Conditional GAN for Point Cloud Generation
Zhulun Yang, Xianwei Zheng, Yadong Chang
ACCV (7)3
2022 Unmixing Convolutional Features for Crisp Edge Detection
abstract
This article presents a context-aware tracing strategy (CATS) for crisp edge detection with deep edge detectors, based on an observation that the localization ambiguity of deep edge detectors is mainly caused by the mixing phenomenon of convolutional neural networks: Feature mixing in edge classification and side mixing during fusing side predictions. The CATS consists of two modules: A novel tracing loss that performs feature unmixing by tracing boundaries for better side edge learning, and a context-aware fusion block that tackles the side mixing by aggregating the complementary merits of learned side edges. Experiments demonstrate that the proposed CATS can be integrated into modern deep edge detectors to improve localization accuracy. With the vanilla VGG16 backbone, in terms of BSDS500 dataset, our CATS improves the F-measure (ODS) of the RCF and BDCN deep edge detectors by 12 and 6 percent, respectively when evaluating without using the morphological non-maximal suppression scheme for edge detection.
Linxi Huan, Nan Xue 0001, Xianwei Zheng, Wei He 0003, Jianya Gong, Gui-Song Xia
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Graph filter design by ring-decomposition for 2-connected graphs
Zhulun Yang, Xianwei Zheng, Zehua Yu
Signal Process.2
2022 Locally Nonlinear Affine Verification for Multisensor Image Matching
abstract
Matching local features between two overlapped images is a fundamental task in photogrammetry and remote sensing. However, images acquired by multiple sensors often differ substantially in properties, thus posing a great challenge to the robustness and flexibility of feature matching methods. In this article, we propose a locally non-linear affine verification (LAV) method for robust multisensor image matching. The main idea of the LAV is the development of a nonlinear regression formulation that practically models the nonlinear deviation of a real surface around a point from its tangent plane during affine verification. Specifically, we start by selecting a restricted set of reliable and well-distributed putative matches as the matching seeds and assign them with neighbors to construct search spaces. In each search space, the regression seeks the smoothest affine model consistent with the latent correct matches, thereby deriving a set of affine parameters to verify correspondence hypotheses for true matches. The verification can be extended to all nearest neighbor matches to discover additional inlier matches. Evaluation on multisensor image datasets with different extents of variations in viewpoint, scale, illumination, and appearance shows that the proposed LAV consistently outperforms existing methods. LAV can achieve a considerable number of high-quality matches, in cases where existing methods provide few or no correct matches.
Xianwei Zheng, Mingyue Dong, Gui-Song Xia, Hanjiang Xiong
IEEE Trans. Geosci. Remote. Sens.2
2022 A Gather-to-Guide Network for Remote Sensing Semantic Segmentation of RGB and Auxiliary Image
abstract
Convolutional neural network (CNN)-based feature fusion of RGB and auxiliary remote sensing data is known to enable improved semantic segmentation. However, such fusion is challengeable because of the substantial variance in data characteristics and quality (e.g., data uncertainties and misalignment) between two modality data. In this article, we propose a unified gather-to-guide network (G2GNet) for remote sensing semantic segmentation of RGB and auxiliary data. The key aspect of the proposed architecture is a novel gather-to-guide module (G2GM) that consists of a feature gatherer and a feature guider. The feature gatherer generates a set of cross-modal descriptors by absorbing the complementary merits of RGB and auxiliary modality data. The feature guider calibrates the RGB feature response by using the channel-wise guide weights extracted from the cross-modal descriptors. In this way, the G2GM can perform RGB feature calibration with different modality data in a gather-to-guide fashion, thus preserving the informative features while suppressing redundant and noisy information. Extensive experiments conducted on two benchmark datasets show that the proposed G2GNet is robust to data uncertainties while also improving the semantic segmentation performance of RGB and auxiliary remote sensing data.
Xianwei Zheng, Xiujie Wu, Linxi Huan, Wei He 0003, Hongyan Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Weighted Error Entropy-Based Information Theoretic Learning for Robust Subspace Representation
abstract
In most of the existing representation learning frameworks, the noise contaminating the data points is often assumed to be independent and identically distributed (i.i.d.), where the Gaussian distribution is often imposed. This assumption, though greatly simplifies the resulting representation problems, may not hold in many practical scenarios. For example, the noise in face representation is usually attributable to local variation, random occlusion, and unconstrained illumination, which is essentially structural, and hence, does not satisfy the i.i.d. property or the Gaussianity. In this article, we devise a generic noise model, referred to as independent and piecewise identically distributed (i.p.i.d.) model for robust presentation learning, where the statistical behavior of the underlying noise is characterized using a union of distributions. We demonstrate that our proposed i.p.i.d. model can better describe the complex noise encountered in practical scenarios and accommodate the traditional i.i.d. one as a special case. Assisted by the proposed noise model, we then develop a new information-theoretic learning framework for robust subspace representation through a novel minimum weighted error entropy criterion. Thanks to the superior modeling capability of the i.p.i.d. model, our proposed learning method achieves superior robustness against various types of noise. When applying our scheme to the subspace clustering and image recognition problems, we observe significant performance gains over the existing approaches.
Yuanman Li, Jiantao Zhou 0001, Jinyu Tian 0001, Xianwei Zheng, Yuan Yan Tang
IEEE Trans. Neural Networks Learn. Syst.4
2021 Robust Feature Matching Using Motion Consistency and Geometrical Constraint for UAV Images
abstract
Establishing feature correspondence between unmanned aerial vehicle (UAV) images is a fundamental task in photogrammetry and remote sensing. However, existing methods still suffer from noisy matches due to the cluttered ground objects presented in UAV images. In this paper, we proposed a novel feature matching method combining motion smoothness and geometrical constraint for UAV images. Given a pair of UAV images (a left one and a right one), once the putative matches were generated, we first divide the left image into a certain number of non-overlapping regions. Then, for local features in each region of the left image, we find their corresponding matching points in the right image, and perform a DBSCAN [1] to cluster the found points into groups. Finally, we determine the group (of points) in the right image which preserves the motion consistency with the region (of points) in the left image, and satisfy a local polar coordinate system-based geometric constraint, as the matching group of that region. Extensive experiments conducted on a set of different conditioned UAV image pairs, show that the proposed method can achieve good performance in terms of precision and recall, outperforming those comparison methods.
Hanjiang Xiong, Xianwei Zheng
IGARSS3
2021 Structured Building Extraction from High-Resolution Satellite Images with a Hybrid Convolutional Neural Network
abstract
Detecting buildings with structure information (e.g., rooflines) from satellite images is a significant yet challenging task. Existing methods usually suffer from the spatial and spectral diversity and complexity of the architectures. In this paper, a deep learning-based approach is proposed to extract structured building rooflines. We use convolutional neural networks to detect corner and line segment primitives. Meanwhile, a collaborative branch of semantic annotation information is combined to obtain the building segmentation map, which ensures the spatial and topological relations of the extracted primitives. Experiments on the SpaceNet dataset show that our proposed approach improves the accuracy of building extraction. Furthermore, the planar graph representation promotes three-dimensional (3D) reconstruction and other subsequent applications.
Hanjiang Xiong, Jianya Gong, Xianwei Zheng
IGARSS4
2021 Multi-windowed vertex-frequency analysis for signals on undirected graphs
Xianwei Zheng, Cuiming Zou, Li Dong 0006, Jiantao Zhou 0001
Comput. Commun.1
2019 Robust Subspace Clustering With Independent and Piecewise Identically Distributed Noise Modeling
abstract
Most of the existing subspace clustering (SC) frameworks assume that the noise contaminating the data is generated by an independent and identically distributed (i.i.d.) source, where the Gaussianity is often imposed. Though these assumptions greatly simplify the underlying problems, they do not hold in many real-world applications. For instance, in face clustering, the noise is usually caused by random occlusions, local variations and unconstrained illuminations, which is essentially structural and hence satisfies neither the i.i.d. property nor the Gaussianity. In this work, we propose an independent and piecewise identically distributed (i.p.i.d.) noise model, where the i.i.d. property only holds locally. We demonstrate that the i.p.i.d. model better characterizes the noise encountered in practical scenarios, and accommodates the traditional i.i.d. model as a special case. Assisted by this generalized noise model, we design an information theoretic learning (ITL) framework for robust SC through a novel minimum weighted error entropy (MWEE) criterion. Extensive experimental results show that our proposed SC scheme significantly outperforms the state-of-the-art competing algorithms.
Yuanman Li, Jiantao Zhou 0001, Xianwei Zheng, Jinyu Tian 0001, Yuan Yan Tang
CVPR3
2019 Multi-Level Fusion of the Multi-Receptive Fields Contextual Networks and Disparity Network for Pairwise Semantic Stereo
abstract
In this paper, we propose a multi-level fusion framework to address the pairwise semantic stereo issue. For disparity estimation, we adopt the pyramid stereo matching network. For semantic segmentation, the single segmentation network is proposed with respect to the left image, along with the disparity fusion segmentation network for the combination of semantic features and disparity features. Specifically, the multi-receptive fusion block is designed and employed to fully extract and fuse the contextual information. Finally, the refined segmentation result is obtained via yet another fusion of the multi-model results. The proposed method achieved a mean intersection over union (mIoU) of 79.05%, an average endpoint error (EPE) of 1.3966, and an mIoU-3 of 77.75%, ranking first in the Pairwise Semantic Stereo Challenge of the 2019 IEEE GRSS Data Fusion Contest [1],[2].
Hongyu Chen 0003, Manhui Lin, Hongyan Zhang 0001, Gui-Song Xia, Xianwei Zheng, Liangpei Zhang 0001
IGARSS6
2019 Multi-Level Downsampling of Graph Signals via Improved Maximum Spanning Trees
abstract
Graph signal processing (GSP) is an emerging field in the signal processing community. Novel GSP-based transforms, such as graph Fourier transform and graph wavelet filter banks, have been successfully utilized in image processing and pattern recognition. As a rapidly developing research area, graph signal processing aims to extend classical signal processing techniques to signals with irregular underlying structures. One of the hot topics in GSP is to develop multi-scale transforms such that novel GSP-based techniques can be applied in image processing or other related areas. For designing graph signal multi-scale frameworks, downsampling operations that ensuring multi-level downsampling should be specifically constructed. Among the existing downsampling methods in graph signal processing, the state-of-the-art method was constructed based on the maximum spanning tree (MST). However, when using this method for multi-level downsampling of graph signals defined on unweighted densely connected graphs, such as social network data, the sampling rates are not close to [Formula: see text]. This phenomenon is summarized as a new problem and called downsampling unbalance problem in this paper. Due to the unbalance, MST-based downsampling method cannot be applied to construct graph signal multi-scale transforms. In this paper, we propose a novel and efficient method to detect and reduce the downsampling unbalance generated by the MST-based method. For any given graph signal, we apply the graph density to construct a measurement of the downsampling unbalance generated by the MST-based method. If a graph signal has large unbalance possibility, the multi-level downsampling is conducted after the MST is improved. The experimental results on synthetic and real-world social network data show that downsampling unbalance can be efficiently detected and then reduced by our method.
Xianwei Zheng, Yuan Yan Tang, Jiantao Zhou 0001, Jianjia Pan, Shouzhi Yang, Youfa Li, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.1
2019 Block sparse representation for pattern classification: Theory, extensions and applications
Yulong Wang 0002, Yuan Yan Tang, Luoqing Li, Xianwei Zheng
Pattern Recognit.4
2017 Indoor pedestrian trajectory tracking based on activity recognition
abstract
Indoor location services (LBSs) have attracted a great deal of attention in recent years, the user's indoor trajectories play an important role in LBSs. In this paper, we combine pedestrian dead reckoning (PDR), human activity recognition (HAR) and landmarks to achieve a good accuracy for indoor localization. The core idea of our research is using PDR to estimate the user's location, and the cumulative error of PDR is reduced by landmarks, which sensed by HAR. In addition, we use step-based classification method to improve the activity recognition accuracy, and we use a magnetometer to aid in identifying door opening activities. The experimental results show that our method can achieve a high degree of accuracy for indoor pedestrian trajectory tracking.
Hanjiang Xiong, Xianwei Zheng, Yan Zhou 0006
IGARSS3
2016 Maximal level estimation and unbalance reduction for graph signal downsampling
abstract
The emerging field of graph signal processing requires a solid design of downsampling operation for graph signals to extend pattern recognition, machine learning and signal processing techniques into the graph setting. The state-of-the-art downsampling method is constructed upon the maximum spanning trees of the graphs. However, under the framework of this method, unbalanced downsampling often occurs for signals defined on densely connected unweighted graphs, such as social network data. The unbalance also significantly reduces the maximal downsampling level, making it smaller than the level we expect. In applications, the maximal level must be estimated to ensure that it is larger than the expected level; meanwhile, the unbalance has to be reduced, if it occurs. In this paper, we propose a novel method to jointly estimate the maximal level and reduce the downsampling unbalance. This method also offers an estimation of the possibility of unbalanced downsampling. If a graph signal is classified to be with high unbalance possibility, the maximum spanning tree will be updated to generate a balanced downsampling. The simulation results on synthesis and real world data support the theoretical analysis.
Xianwei Zheng, Yuan Yan Tang, Jiantao Zhou 0001, Patrick Shen-Pei Wang
ICPR1
2016 A hybrid swarm optimization for neural network training with application in stock price forecasting
abstract
A improved swarm optimization method based on particle swarm optimization (PSO) and simplified swarm optimization (SSO) is proposed to adjust the weight in artificial neural network. This method is a modification of traditional PSO and SSO, and combines them to a new optimization method (PSOSSO for short). The proposed method overcomes some of the drawbacks of SSO and improves its ability to train the weight of ANN. In the experiments, the PSOSSO is employed to train fuzzy wavelet neural network (FWNN) forecasting model to predict the prices of Hong Kong Hang Seng Index. The experimental results present that the PSOSSO is more efficient than traditional PSO and SSO methods.
Jianjia Pan, Yuan Yan Tang, Yulong Wang 0002, Xianwei Zheng, Huiwu Luo, Patrick Shen-Pei Wang
SMC4
2016 Improving unbalanced downsampling via maximum spanning trees for graph signals
abstract
The state-of-the-art downsampling method for graph signals has been constructed by using maximum spanning trees (MSTs) of the graphs. For the graph signals defined on unweighted densely connected graphs, such as social network data, the sampling rates via MST-based downsampling are not close to 1/2, leading to a unbalanced downsampling phenomenon on multi-level downsampling. The unbalance hinders the applications of MST-based downsampling on constructing graph signal multiscale transforms, such as graph wavelet decomposition and multiscale pyramid transform. In this paper, we propose a simple but efficient method to improve the performance of the MST-based method on downsampling balance. For every graph signal, we first propose an unbalance possibility to measure the unbalance of the MST-based downsampling. If the unbalance possibility is high, the downsampling will be conducted on an improved MST, which is constructed by rearranging the structure of the MST to reduce the downsampling unbalance. The experiment results on synthesis graph signal show that the proposed improved MST leads to balanced downsampling. That is, the sampling rates produced by the improved MST are closer to 1/2 in multi-level downsampling than the original MST-based method.
Xianwei Zheng, Yuan Yan Tang, Jiantao Zhou 0001, Patrick Shen-Pei Wang
SMC1
2016 Adaptive Multiscale Decomposition of Graph Signals
abstract
This paper proposes an adaptive multiscale decomposition algorithm for graph signals. We develop two types of graph signal cost functions: α-sparsity functional and graph signal entropies, to capture the energy compaction of the signal components. The adaptive decomposition can then be constructed by applying a minimum cost constraint during the full subband decomposition. The proposed adaptive decomposition is shown to outperform graph wavelet decomposition in compressing nonpiecewise constant graph signals.
Xianwei Zheng, Yuan Yan Tang, Jianjia Pan, Jiantao Zhou 0001
IEEE Signal Process. Lett.1