Shibiao Xu

dblp:76/8401 · DBLP profile ↗
← Back
113ranked-venue papers
7as first author
92since 2021 · last 2026
0000-0003-4037-9900ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 61 · 4 first-author · 48 since 2021Artificial intelligence and machine learning · 39 · 1 first-author · 32 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 2 first-author · 12 since 2021Computer networks · 11 · 11 since 2021Systems, architecture and hardware · 6 · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
YearPublicationVenuePosition
2026 RealNet: Efficient and Unsupervised Detection of AI-Generated Images via Real-Only Representation Learning
abstract
Detecting AI-generated images remains a persistent challenge, as existing detectors often struggle to generalize to forgeries produced by previously unseen generative models. This generalization gap mainly stems from entanglement with semantic content and overfitting to model-specific artifacts. Moreover, many state-of-the-art methods rely on large pre-trained backbones or computationally intensive pipelines, which limit their applicability in real-world, resource-constrained environments. We propose RealNet, a lightweight and unsupervised framework that constructs a disentangled, forgery-aware representation space using only real images. RealNet first extracts semantic-agnostic representations through a dual adversarial denoising mechanism, producing compact features with low intra-class variance. These representations are then perturbed in feature space to generate pseudo-negative samples, which are combined with the original real features to train a lightweight discriminator, enabling robust detection without any dependence on synthetic images during training. Comprehensive evaluations across GAN, diffusion, and emerging VAR-based paradigms demonstrate that RealNet achieves superior cross-model generalization and robustness. RealNet surpasses previous state-of-the-art approaches by 4.51% in accuracy and 3.93% in average precision, while maintaining significantly lower computational cost. Furthermore, we introduce a medically relevant synthetic image dataset and show RealNet remains effective under severe distribution shifts, highlighting its potential for deployment in high-stakes real-world scenarios. Together, these advantages position RealNet as a practical, scalable and socially impactful solution for robust AI-generated image detection.
Shuaibo Li, Laixin Zhang, Wei Ma 0008, Jianwei Guo 0003, Shibiao Xu, Zhijie Qiu, Hongbin Zha
AAAI5
2026 Spatial-Frequency Domain Complementary Learning for Robust Cross-Modal Hashing
abstract
Deep cross-modal hashing has achieved remarkable success in cross-modal retrieval due to its fast retrieval speed and low storage cost, but it highly vulnerable to adversarial attacks. Mainstream defense methods rely on adversarial training, which often induces a robustness–standard performance trade-off, leading to the learning of only limited robust features and degraded standard performance. To address this issue, we propose a Spatial-Frequency Domain Complementary Learning (SFCL) framework to overcome these two challenges by: 1) exploiting the complementarity of spatial and frequency features to learn more comprehensive and adversarially robust features, addressing the limited robustness of existing defenses; 2) by supplementing frequency-domain information, it avoids the performance degradation commonly caused by adversarial training. Specifically, SFCL consists of two modules: a Spatial-Frequency Robust Gating (SFRG) module, which selects robust features and strengthens complementarity via a conditional mutual information-based loss; and a Robustness-Aware Feature Fusion (RAFF) module, which performs bidirectional feature interaction and fusion. Extensive experiments demonstrate significant robustness gains over existing state-of-the-art methods, along with improved standard performance.
Gang Zhou 0001, Shibiao Xu, Xiaolong Zheng 0001
ICMR2
2026 Spectral-Adaptive Adversarial Hashing for Robust Image Retrieval
abstract
Deep hashing is widely used in large-scale image retrieval systems due to its efficient retrieval performance. However, its susceptibility to adversarial attacks limits its security in practical applications. Adversarial training is the most effective method for improving robustness, but it often leads to a significant trade-off between robustness and retrieval accuracy. In this paper, we conduct spectral analysis and find that generating high-quality hash codes requires wide-frequency response models, whereas adversarial training forces the model into spectral collapse, degrading it to a low-frequency response model and weakening its discriminability. To address this issue, we propose a Spectral-Adaptive Adversarial Hashing (SAAH) framework, which selectively preserves discriminative and task-relevant frequency components while suppressing adversarially unstable ones, enabling robust hashing without sacrificing retrieval performance. Extensive experiments on benchmark datasets demonstrate that SAAH consistently achieves a superior balance between retrieval accuracy and adversarial robustness, achieving the best performance in both retrieval accuracy and robustness compared with existing robust hashing methods.
Gang Zhou 0001, Shibiao Xu, Xiaolong Zheng 0001, Daniel Dajun Zeng
SIGIR2
2026 FurniScene: A Large-scale 3D Room Dataset with Intricate Furnishing Scenes
Yuxi Wang 0001, Junran Peng, Genghao Zhang, Chuanchen Luo, Shibiao Xu, Man Zhang 0005, Zhaoxiang Zhang 0001
Int. J. Comput. Vis.5
2026 Deep Learning for 3-D Lane Detection in Autonomous Driving: A Survey
abstract
3D lane detection has become a critical component in the perception task of autonomous vehicles. Unlike 2D lane detection, which operates in the image plane, 3D lane detection estimates the spatial layout of lanes in real-world coordinates, enabling fine-grained localization, map construction, and planning. However, the task remains challenging due to depth ambiguity, sensor limitations, and diverse road conditions. Existing surveys mostly focus on 2D or organize 3D lane detection by sensor modality, lacking a systematic treatment of algorithmic designs. In this paper, we present a comprehensive survey of deep learning-based 3D lane detection methods. We introduce a dual-axis taxonomy that jointly considers modeling paradigms and representation spaces. Based on this framework, we categorize existing methods into four primary paradigms: geometry-based, end-to-end, query-based, and implicit field-based. We analyze how each paradigm interacts with spatial representations such as image, BEV, 3D, and topological spaces. For each category, we review representative frameworks, architectural principles, and performance trade-offs. We also provide an extensive summary of public datasets, evaluation metrics, and state-of-the-art results across multiple benchmarks. Finally, we identify current limitations and outline future research directions toward robust, scalable, and interpretable 3D lane detection.
Xiaoqiang Teng, Zuo Chen, Shunpeng Chen, Shibiao Xu, Zhihao Hao, Deke Guo, Hai-Sheng Li 0002
IEEE Internet Things J.5
2026 Online knowledge distillation optimization based on Multi-Student model Multi-Task collaborative learning
Shibiao Xu, Shanshan Mo, Changwei Wang 0001, Hetong Wang, Rongtao Xu, Li Guo 0004
Knowl. Based Syst.1
2026 Rating-aware argument generation for movie reviews with multimodal large language models and a new dataset
Wenjie Hua, Quan Fang, Muyi Sun, Shibiao Xu, Man Zhang 0005
Multim. Syst.4
2026 Pinching-Antenna Systems (PASS)-Enabled Secure Wireless Communications
abstract
A novel pinching-antenna systems (PASS)-enabled secure wireless communication framework is proposed. By dynamically adjusting the positions of dielectric particles, namely pinching antennas (PAs), along the waveguides, PASS introduces a novel concept of pinching beamforming to enhance the performance of physical layer security. A fundamental PASS-enabled secure communication system is considered with one legitimate user and one eavesdropper. Both single-waveguide and multiple-waveguide scenarios are studied. 1) For the single-waveguide scenario, the secrecy rate (SR) maximization is formulated to optimize the pinching beamforming. A PA-wise successive tuning (PAST) algorithm is proposed, which ensures constructive signal superposition at the legitimate user while inducing a destructive legitimate signal at the eavesdropper. 2) For the multiple-waveguide scenario, artificial noise (AN) is employed to further improve secrecy performance. A pair of practical transmission architectures are developed:waveguide division (WD)andwaveguide multiplexing (WM). The key difference lies in whether each waveguide carries a single type of signal or a mixture of signals with baseband beamforming. For the SR maximization problem under the WD case, a two-stage algorithm is developed, where the pinching beamforming is designed with the PAST algorithm and the baseband power allocation among AN and legitimate signals is solved using successive convex approximation (SCA). For the WM case, an alternating optimization algorithm is developed, where the baseband beamforming is optimized with SCA and the pinching beamforming is designed employing particle swarm optimization. Numerical results demonstrate that i) PASS can significantly improve the secrecy performance over conventional antenna systems in both scenarios; ii) the proposed PAST algorithm for the single-waveguide scenario is efficient, especially when the number of PAs is even or large; iii) WM provides higher and more stable performance at the cost of increased complexity, while WD serves as a simple yet scalable alternative, which is effective when a large number of PAs are deployed.
Guangyu Zhu 0007, Xidong Mu, Li Guo 0004, Shibiao Xu, Yuanwei Liu, Naofal Al-Dhahir
IEEE Trans. Commun.4
2026 TeDri:Teacher-Driven Region Knowledge Distillation
abstract
Knowledge distillation as a practical tool to enhance the performance of small-capacity student networks on downstream tasks comes at the cost of a lengthy distillation process due to the online inference of teacher networks, especially when there is a large capacity gap between them. Therefore, in this paper, we propose a fast distillation framework called TeDri based on region images by offline saving relevant regional information and its teacher guidance. Specifically, first, to alleviate the lack of diversity caused by the fixed augmentation path in region images, we propose Teacher-driven MixUp strategies with mild intensity and advocate binding the mixing factor$\lambda$with teacher guidance confidence, where more confident category representations dominate the MixUp process. Furthermore, recognizing the need to evaluate these randomly cropped regions, and we propose region contrastive learning, encourage the student network to mimic the region partitioning behavior of the teacher, promoting a comprehensive understanding of global semantic content from multiple local perspectives. Finally, we introduce region mutual learning, employing spatial constraints among regions to require the student network towards consistent content interpretation across localized regions. Experiments on CIFAR-100 and ImageNet-1 K validate the effectiveness of the proposed TeDri, achieving competitive performance while significantly reducing training time.
Changwei Wang 0001, Rongtao Xu, Xingtian Pei, Shibiao Xu, Wenbo Xu 0003, Li Guo 0004
IEEE Trans. Knowl. Data Eng.6
2026 Adaptive in Adapter: Boosting Open-Vocabulary Semantic Segmentation With Adaptive Dropout Adapter
abstract
Open-vocabulary semantic segmentation is a challenging multimedia task that requires segmentation and recognition of unseen word classes during the testing phase. Recent works bridge the gap between closed and open-vocabulary recognition by introducing large-scale visual language models such as CLIP with cross-modal alignment capabilities. To preserve multimodal alignment capabilities, it is common to freeze the parameters of the CLIP and then add additional learnable components such as adapters to expand to downstream tasks. However, for the open-vocabulary semantic segmentation task, the plain adapter suffers from overfitting the closed-vocabulary classes and impairs performance on the open-vocabulary unseen classes. In addition, since CLIP is trained to perform image-level alignment can cause the network to over-focus on partially discriminative regions, resulting in incomplete segmentation masks. To alleviate the above problems, we introduce adaptive dropout adapters to release theAdaptiveInAdapter (i.e.AIA) from the following two aspects:i)A Generalization Feature Selection Adapter (GFSA) is proposed to improve the generalization of network over unseen classes.ii)A Discriminative Region Mask Adapter (DRMA) is proposed for retrofitting CLIP backbone, has provided region free biased features for segmentation mask generation. Meanwhile, our proposed AIA achieves the current state-of-the-art performance on several open-vocabulary semantic segmentation benchmarks. Code is available athttps://github.com/clearxu/AIA.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Jiguang Zhang, Xiaoqiang Teng, Weiliang Meng, Xiaopeng Zhang 0001
IEEE Trans. Multim.5
2026 STAR-RIS Assisted SWIPT Systems: Active or Passive?
abstract
A simultaneously transmitting and reflecting reconfigurable intelligent surface (STAR-RIS) assisted simultaneous wireless information and power transfer (SWIPT) system is investigated. Both active and passive STAR-RISs are considered. Passive STAR-RISs can be cost-efficiently fabricated to large aperture sizes with significant near-field regions, but the design flexibility is limited by the coupled phase-shifts. Active STAR-RISs can further amplify signals and have independent phase-shifts, but their aperture sizes are relatively small due to the high cost. To characterize and compare their performance, a power consumption minimization problem is formulated by jointly designing the beamforming at the access point (AP) and the STAR-RIS, subject to both the power and information quality-of-service requirements. To solve the resulting highly-coupled non-convex problem, the original problem is first decomposed into simpler subproblems and then an alternating optimization framework is proposed. For the passive STAR-RIS, the coupled phase-shift constraint is tackled by employing a vector-driven weighted penalty method. While for the active STAR-RIS, the independent phase-shift is optimized with AP beamforming via matrix-driven semidefinite programming, and the amplitude matrix is updated using convex optimization techniques in each iteration. Numerical results show that: 1) given the same aperture sizes, the active STAR-RIS exhibits superior performance over the passive one when the aperture size is small, but the performance gap decreases with the increase in aperture size; and 2) given identical power budgets, the passive STAR-RIS is generally preferred, whereas the active STAR-RIS typically suffers performance loss for balancing between the hardware power and the amplification power.
Guangyu Zhu 0007, Xidong Mu, Li Guo 0004, Ao Huang, Shibiao Xu
IEEE Trans. Wirel. Commun.5
2025 Focus on Local: Finding Reliable Discriminative Regions for Visual Place Recognition
abstract
Visual Place Recognition (VPR) is aimed at predicting the location of a query image by referencing a database of geotagged images. For VPR task, often fewer discriminative local regions in an image produce important effects while mundane background regions do not contribute or even cause perceptual aliasing because of easy overlap. However, existing methods lack precisely modeling and full exploitation of these discriminative regions. In addition, the lack of pixel-level correspondence supervision in the VPR dataset hinders further improvement of the local feature matching capability in the re-ranking stage. In this paper, we propose the Focus on Local (FoL) approach to stimulate the performance of image retrieval and re-ranking in VPR simultaneously by mining and exploiting reliable discriminative local regions in images and introducing pseudo-correlation supervision. First, we design two losses, Extraction-Aggregation Spatial Alignment Loss (SAL) and Foreground-Background Contrast Enhancement Loss (CEL), to explicitly model reliable discriminative local regions and use them to guide the generation of global representations and efficient re-ranking. Second, we introduce a weakly-supervised local feature training strategy based on pseudo-correspondences obtained from aggregating global features to alleviate the lack of local correspondences ground truth for the VPR task. Third, we suggest an efficient re-ranking pipeline that is efficiently and precisely based on discriminative region guidance. Finally, experimental results show that our FoL achieves the state-of-the-art on multiple VPR benchmarks in both image retrieval and re-ranking stages and also significantly outperforms existing two-stage VPR methods in terms of computational efficiency.
Changwei Wang 0001, Shunpeng Chen, Rongtao Xu, Jiguang Zhang, Haoran Yang 0003, Yu Zhang 0133, Kexue Fu 0001, Shide Du, Zhiwei Xu 0005, Longxiang Gao, Li Guo 0004, Shibiao Xu
AAAI14
2025 Novel View Synthesis Under Large-Deviation Viewpoint for Autonomous Driving
abstract
Novel view synthesis is a critical task in autonomous driving. Although 3D Gaussian Splatting (3D-GS) has shown success in generating novel views, it faces challenges in maintaining high-quality rendering when viewpoints deviate significantly from the training set. This difficulty primarily stems from complex lighting conditions and geometric inconsistencies in texture-less regions. To address these issues, we propose an attention-based illumination model that leverages light fields from neighboring views, enhancing the realism of synthesized images. Additionally, we propose a geometry optimization method using planar homography to improve geometric consistency in texture-less regions. Our experiments demonstrate substantial improvements in synthesis quality for large-deviation viewpoints, validating the effectiveness of our approach.
Jiguang Zhang, Shibiao Xu, Chengwei Pan
AAAI4
2025 CAE-DFKD: Bridging the Transferability Gap in Data-Free Knowledge Distillation
abstract
Data-Free Knowledge Distillation (DFKD) enables the knowledge transfer from the given pre-trained teacher network to the target student model without access to the real training data. Existing DFKD methods focus primarily on improving image recognition performance on associated datasets, often neglecting the crucial aspect of the transferability of learned representations. In this paper, we propose Category-Aware Embedding Data-Free Knowledge Distillation (CAE-DFKD), which addresses at the embedding level the limitations of previous rely on image-level methods to improve model generalization but fail when directly applied to DFKD. The superiority and flexibility of CAE-DFKD are extensively evaluated, including: i.) Significant efficiency advantages resulting from altering the generator training paradigm; ii.) Competitive performance with existing DFKD state-of-the-art methods on image recognition tasks; iii.) Remarkable transferability of data-free learned representations demonstrated in downstream tasks.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Yu Zhang 0133, Jie Zhou 0001, Li Guo 0004
DAC5
2025 SCS: Spatially Consistent Self-Supervised approach for One-Shot Anatomical Landmark Detection
abstract
Landmark detection is essential in medical image analysis, serving as the foundation for many downstream tasks. In recent years, supervised anatomical landmark detection models have achieved remarkable success, but typically require large amounts of labeled data for training, which is challenging to obtain due to the expertise and time needed for accurate annotation. Rather than relying on costly expert annotations, this paper focuses on leveraging a one-shot method for automated annotation. To this end, we propose a Spatially Consistent Self-supervised approach (SCS) within a two-stage framework for one-shot anatomical landmark detection. In the first stage, we design a multi-scale contrastive self-supervised method that leverages the inherent spatial consistency of medical images, characterized by clear structures and similar patterns, to extract global and local features. During inference, pseudo-labels are generated based on the one-shot template. In the second stage, we train a supervised model using the pseudo-labels and mitigate label noise through a mask and multi-task approach. Our method is evaluated on three widely-used public X-ray datasets, achieving state-of-the-art performance across almost all metrics with the Mean Radial Error (MRE) reduced to 1.97mm on the Cephalometric dataset, 1.43mm on the Hand dataset, and 6.58mm on the Chest dataset, thereby demonstrating the effectiveness of our Spatially Consistent Self-supervised approach.
Li Guo 0004, Shibiao Xu
ICASSP5
2025 DCSA-UNet: Lightweight UNet with Dual Cross-Shaped Attention For Skin Lesion Segmentation
abstract
Skin cancer is a prevalent and life-threatening disease where early detection significantly improves survival rates. However, existing segmentation models are often too large and computationally intensive, limiting their applicability in resource-constrained medical scenarios. To address this, we propose DCSA-UNet, a lightweight U-Net-based segmentation model designed for efficient and effective skin lesion segmentation. The model integrates a novel Dual Cross-Shaped Attention (DCSA) mechanism, which ensures computational complexity grows linearly with both the number of tokens and the receptive field width. Additionally, the Mask-Guided Progressive Fusion (MGPF) module addresses the feature mismatch between encoder and decoder in lightweight models through progressive multi-scale integration. By combining these innovations, DCSA-UNet achieves state-of-the-art performance with significantly reduced computational and storage requirements. Compared to existing methods, it reduces parameter counts by 59% and GFLOPs by 2%, while improving mIoU by 0.78% and DSC by 0.48%. These results highlight its potential as a robust solution for real-time skin lesion segmentation. https://github.com/Litpill/DCSA-UNet
Li Guo 0004, Shibiao Xu
ICME5
2025 MVPS: Multi-View Adaptive Prompt Synergy for Zero-shot Anomaly Detection
abstract
Zero-shot anomaly detection (ZSAD) in industrial domains faces significant challenges due to the diverse manifestations of anomalies across scales and semantic levels. Existing methods, relying on single prompt spaces, struggle to generalize across these variations. So they exhibit limited generalization across scales and semantic levels. We propose Multi-View Adaptive Prompting Synergy (MVPS), a novel framework that establishes multiple scale-aware prompt spaces to enhance anomaly detection generalization. MVPS comprises three synergistic components: Hierarchical Multi-modal Prompt Tuning (HMPT) for generating scale-aware prompts, Dual-stream Prompt Tuning Orchestration (DPTO) for achieving robust cross-modal feature alignment, and Multi-View Prompt Composition Learning (MVPCL) for effective scale feature perception. This approach enables comprehensive capture anomaly feature across multiple semantic levels and scales, overcoming limitations of single-space representations and scale-insensitive feature alignment. Extensive experiments on seven benchmark datasets demonstrate that MVPS achieves state-of-the-art performance, exhibiting superior generalization capability across diverse anomaly categories and industrial domain. The code is available at https://github.com/MLY-0546/mvps.
Longzhao Huang, Changwei Wang 0001, Rongtao Xu, Shibiao Xu
ICME6
2025 Complementary Information Guided Occupancy Prediction via Multi-Level Representation Fusion
abstract
Camera-based occupancy prediction is a main-stream approach for 3D perception in autonomous driving, aiming to infer complete 3D scene geometry and semantics from 2D images. Almost existing methods focus on improving performance through structural modifications, such as lightweight backbones and complex cascaded frameworks, with good yet limited performance. Few studies explore from the perspective of representation fusion, leaving the rich diversity of features in 2D images underutilized. Motivated by this, we propose CIGOcc, a two-stage occupancy prediction framework based on multi-level representation fusion. CIGOcc extracts segmentation, graphics, and depth features from an input image and introduces a deformable multi-level fusion mechanism to fuse these three multi-level features. Additionally, CIGOcc incorporates knowledge distilled from SAM to further enhance prediction accuracy. Without increasing training costs, CIGOcc achieves state-of-the-art performance on the SemanticKITTI benchmark. The code is provided in the supplementary material and will be released project page.
Rongtao Xu, Jinzhou Lin 0001, Jialei Zhou, Jiahua Dong 0001, Changwei Wang 0001, Ruisheng Wang 0001, Li Guo 0004, Shibiao Xu, Xiaodan Liang
ICRA8
2025 DiffusionIMU: Diffusion-Based Inertial Navigation with Iterative Motion Refinement
abstract
Inertial navigation enables self-contained localization using only Inertial Measurement Units (IMUs), making it widely applicable in various domains such as navigation, augmented reality, and robotics. However, existing methods suffer from drift accumulation due to the sensor noise and difficulty capturing long-range temporal dependencies, limiting their robustness and accuracy. To address these challenges, we propose DiffusionIMU, a novel diffusion-based framework for inertial navigation. DiffusionIMU enhances direct velocity regression from IMU data through an iterative generative denoising process, progressively refining motion state estimation. It integrates the noise-adaptive feature modulation for sensor variability handling, the feature alignment mechanism for representation consistency, and the diffusion-based temporal modeling to decrease accumulated drift. Experiments show that DiffusionIMU consistently outperforms existing methods, demonstrating superior generalization to unseen users while alleviating the impact of the sensor noise.
Xiaoqiang Teng, Shibiao Xu, Zhihao Hao, Deke Guo, Hai-Sheng Li 0002, Weiliang Meng, Xiaopeng Zhang 0001
IJCAI3
2025 3D-MoRe: Unified Modal-Contextual Reasoning for Embodied Question Answering
abstract
With the growing need for diverse and scalable data in indoor scene tasks, such as question answering and dense captioning, we propose 3D-MoRe, a novel paradigm designed to generate large-scale 3D-language datasets by lever-aging the strengths of foundational models. The framework integrates key components, including multi-modal embedding, cross-modal interaction, and a language model decoder, to process natural language instructions and 3D scene data. This approach facilitates enhanced reasoning and response generation in complex 3D environments. Using the ScanNet 3D scene dataset, along with text annotations from ScanQA and ScanRefer, 3D-MoRe generates 62,000 question-answer (QA) pairs and 73,000 object descriptions across 1,513 scenes. We also employ various data augmentation techniques and implement semantic filtering to ensure high-quality data. Experiments on ScanQA demonstrate that 3D-MoRe significantly outperforms state-of-the-art baselines, with the CIDEr score improving by 2.15%. Similarly, on ScanRefer, our approach achieves a notable increase in [email protected] by 1.84%, highlighting its effectiveness in both tasks. Our code and generated datasets will be publicly released to benefit the community, and both can be accessed on the https://3D-MoRe.github.io.
Rongtao Xu, Mingming Yu, Dong An 0002, Shunpeng Chen, Changwei Wang 0001, Li Guo 0004, Xiaodan Liang, Shibiao Xu
IROS9
2025 ARPDR++: Exploiting local-global temporal modeling for smartphone-based indoor pedestrian localization
Xiaoqiang Teng, Shibiao Xu, Deke Guo, Yulan Guo, Pengfei Xu 0013, Runbo Hu
Comput. Networks2
2025 FDBPL: Faster distillation-based prompt learning for region-aware vision-language models adaptation
Changwei Wang 0001, Rongtao Xu, Longzhao Huang, Wenbo Xu 0003, Li Guo 0004, Shibiao Xu
Expert Syst. Appl.9
2025 C2Fi-NeRF: Coarse to fine inversion NeRF for 6D pose estimation
Jiguang Zhang, Zhaohui Zhang 0002, Xuxiang Feng, Shibiao Xu, Rongtao Xu, Changwei Wang 0001, Kexue Fu 0001, Jiaxi Sun, Weilong Ding 0001
Expert Syst. Appl.4
2025 VANE-IN: Velocity Auto-Encoder for Inertial Navigation
abstract
Data-driven inertial navigation is crucial for mobile computing applications, such as navigation, augmented reality, and robotics. It typically depends on a trained velocity regression network (VRN) to estimate velocities from inertial measurement unit (IMU) data, enabling position determination through integration. However, using prior velocity information for feature representation in inertial navigation remains underexplored. This work introduces a framework called velocity auto-encoder for inertial navigation (VANE-IN), which employs a Teacher-Student scheme to enhance VRN performance by encoding the velocity. Specifically, a velocity auto-encoder (VANE) is proposed as a student model to distill prior velocity insights from the training dataset, which is then guided by the VRN acting as the teacher model. Additionally, an attention mechanism is introduced to fuse these insights into the features of the VRN. To this end, the VANE-IN achieves a state-of-the-art position accuracy on the RoNIN benchmarks. Our experimental results demonstrate that the VANE-IN achieves approximately 5% performance improvements over existing methods regarding position accuracy.
Xiaoqiang Teng, Shibiao Xu, Deke Guo, Hai-Sheng Li 0002
IEEE Internet Things J.3
2025 Enhancing User Fairness in Wireless Powered Communication Networks With STAR-RIS
abstract
A simultaneously transmitting and reflecting reconfigurable-intelligent-surface (STAR-RIS)-assisted wireless powered communication network (WPCN) is proposed, where two energy-limited devices first harvest energy from a hybrid access point (HAP) and then use that energy to transmit information back. To fully eliminate thedoubly-near-far-effect in WPCNs, two STAR-RIS operating protocol-driven transmission strategies, namely energy splitting nonorthogonal multiple access (ES-NOMA) and time switching time division multiple access (TS-TDMA) are proposed. For each strategy, the corresponding optimization problem is formulated to maximize the minimum throughput by jointly optimizing time allocation, user transmit power, active HAP beamforming, and passive STAR-RIS beamforming. For ES-NOMA, the resulting intractable problem is solved via a two-layer algorithm, which exploits the 1-D search and block coordinate descent methods in an iterative manner. For TS-TDMA, the optimal active beamforming and passive beamforming are first determined according to the maximum-ratio transmission beamformer. Then, the optimal solution of the time allocation variables is obtained by solving a standard convex problem. Numerical results show that: 1) the STAR-RIS can achieve considerable performance improvements for both strategies compared to the conventional RIS; 2) TS-TDMA is preferred for single-antenna scenarios, whereas ES-NOMA is better suited for multiantenna scenarios; and 3) the superiority of ES-NOMA over TS-TDMA is enhanced as the number of STAR-RIS elements increases.
Guangyu Zhu 0007, Xidong Mu, Li Guo 0004, Ao Huang, Shibiao Xu
IEEE Internet Things J.5
2025 Movable-Element STARS-Assisted Near-Field Wideband Communications
abstract
A novel movable-element simultaneously transmitting and reflecting surface (ME-STARS)-assisted near-field wideband communication framework is proposed. In particular, the position of each STARS element can be adjusted to combat the significant wideband beam squint issue in the near field instead of using costly true-time delay components. Four practical ME-STARS element movement modes are proposed, namely region-based (RB), horizontal-based (HB), vertical-based (VB), and diagonal-based (DB) modes. Based on this, a near-field wideband multi-user downlink communication scenario is considered, where a sum rate maximization problem is formulated by jointly optimizing the base station (BS) precoding, ME-STARS beamforming, and element positions. To solve this intractable problem, a two-layer algorithm is developed. For the inner layer, the block coordinate descent optimization framework is utilized to solve the BS precoding and ME-STARS beamforming in an iterative manner. For the outer layer, the particle swarm optimization-based heuristic search method is employed to determine the desired element positions. Numerical results show that: 1) the ME-STARSs can effectively address the beam squint for near-field wideband communications compared to conventional STARSs with fixed element positions; 2) the RB mode achieves the most efficient beam squint effect mitigation, while the DB mode achieves the best trade-off between performance gain and hardware overhead; and 3) an increase in the number of ME-STARS elements or BS subcarriers substantially improves the system performance.
Guangyu Zhu 0007, Xidong Mu, Li Guo 0004, Ao Huang, Shibiao Xu
IEEE Internet Things J.5
2025 Generalization Boosted Adapter for Open-Vocabulary Segmentation
abstract
Vision-language models (VLMs) have demonstrated remarkable open-vocabulary object recognition capabilities, motivating their adaptation for dense prediction tasks like segmentation. However, directly applying VLMs to such tasks remains challenging due to their lack of pixel-level granularity and the limited data available for fine-tuning, leading to overfitting and poor generalization. To address these limitations, we propose Generalization Boosted Adapter (GBA), a novel adapter strategy that enhances the generalization and robustness of VLMs for open-vocabulary segmentation. GBA comprises two core components: (1) a Style Diversification Adapter (SDA) that decouples features into amplitude and phase components, operating solely on the amplitude to enrich the feature space representation while preserving semantic consistency; and (2) a Correlation Constraint Adapter (CCA) that employs cross-attention to establish tighter semantic associations between text categories and target regions, suppressing irrelevant low-frequency “noise” information and avoiding erroneous associations. Through the synergistic effect of the shallow SDA and the deep CCA, GBA effectively alleviates overfitting issues and enhances the semantic relevance of feature representations. As a simple, efficient, and plug-and-play component, GBA can be flexibly integrated into various CLIP-based methods, demonstrating broad applicability and achieving state-of-the-art performance on multiple open-vocabulary segmentation benchmarks. Code are available athttps://github.com/clearxu/BGA.
Changwei Wang 0001, Xuxiang Feng, Rongtao Xu, Longzhao Huang, Li Guo 0004, Shibiao Xu
IEEE Trans. Circuits Syst. Video Technol.8
2025 DFMC: Feature-Driven Data-Free Knowledge Distillation
abstract
Data-Free Knowledge Distillation (DFKD) enables knowledge transfer from teacher networks without access to the real dataset. However, generator-based DFKD methods often suffer from insufficient diversity or low-confidence in synthetic images, negatively impacting student network performance. This paper introduces DFMC, a generative feature-driven framework to mitigate the inherent limitations of DFKD. We propose exploiting semantic description between generative feature domains to guide augmentation strategies, avoiding random abstract inputs caused by inconsistent semantic quality. Then, by applying noise to the generative features, we produce contrastive learning pairs indirectly, limiting the sampling range of the feature domain to encourage the student network to learn domain-invariant features. Finally, we guide the student network to deeply mimic the teacher’s layer-wise implicit classification behavior for the augmented synthetic images. Extensive experiments across various datasets and downstream tasks demonstrate the effectiveness of DFMC, achieving significant improvements while preventing student networks from overfitting to semantic ambiguous images.
Rongtao Xu, Changwei Wang 0001, Shunpeng Chen, Shibiao Xu, Guangyuan Xu, Li Guo 0004
IEEE Trans. Circuits Syst. Video Technol.6
2025 Segment Anything Model Is a Good Teacher for Local Feature Learning
abstract
Local feature detection and description play an important role in many computer vision tasks, which are designed to detect and describe keypoints in any scene and any downstream task. Data-driven local feature learning methods need to rely on pixel-level correspondence for training. However, a vast number of existing approaches ignored the semantic information on which humans rely to describe image pixels. In addition, it is not feasible to enhance generic scene keypoints detection and description simply by using traditional common semantic segmentation models because they can only recognize a limited number of coarse-grained object classes. In this paper, we propose SAMFeat to introduce SAM (segment anything model), a foundation model trained on 11 million images, as a teacher to guide local feature learning. SAMFeat learns additional semantic information brought by SAM and thus is inspired by higher performance even with limited training samples. To do so, first, we construct an auxiliary task of Attention-weighted Semantic Relation Distillation (ASRD), which adaptively distillates feature relations with category-agnostic semantic information learned by the SAM encoder into a local feature learning network, to improve local feature description using semantic discrimination. Second, we develop a technique called Weakly Supervised Contrastive Learning Based on Semantic Grouping (WSC), which utilizes semantic groupings derived from SAM as weakly supervised signals, to optimize the metric space of local descriptors. Third, we design an Edge Attention Guidance (EAG) to further improve the accuracy of local feature detection and description by prompting the network to pay more attention to the edge region guided by SAM. SAMFeat's performance on various tasks, such as image matching on HPatches, and long-term visual localization on Aachen Day-Night showcases its superiority over previous local features. The release code is available at https://github.com/vignywang/SAMFeat.
Jingqian Wu, Rongtao Xu, Zach Wood-Doughty, Changwei Wang 0001, Shibiao Xu, Edmund Y. Lam
IEEE Trans. Image Process.5
2025 Token Masking Transformer for Weakly Supervised Object Localization
abstract
Weakly supervised object localization (WSOL) is both a promising and challenging task that aims to achieve object localization exclusively through image category labels for supervision. Visual transformers have recently been applied to WSOL, demonstrating significant success through the exploitation of long-range feature dependencies in self-attention mechanisms. However, the transformer-based approach suffers from the same partial activation problem as the CNN-based approach due to the use of the classification task to train self-attention map, i.e., only a few discriminative regions are assigned high attention response and thus the localization map does not cover the whole object. To alleviate this problem, we propose a plug-and-play Token Masking Transformer (TMT) method to help transformer-based WSOL methods to obtain a more complete localization map by dynamic discriminative token masking. Specifically, a batch-wise discriminative token selection strategy is first introduced to flexibly determine the tokens to be masked in each image. Then, we design a token masking transformer block to perform token masking and inspire the network to mine more object-related tokens. Besides, we also design an intermediate token activation loss to further improve the performance of TMT by imposing constraints on intermediate tokens. Extensive experiments demonstrate that our TMT can substantially improve the performance of existing transformer-based methods without increasing the computational cost, and achieves state-of-the-art performance on two mainstream benchmarks.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Man Zhang 0005, Xiaopeng Zhang 0001
IEEE Trans. Multim.4
2025 SRIF: Data-Free Knowledge Distillation via Stable Regulation and Input Filtering
abstract
Data-free knowledge distillation (DFKD) enables knowledge transfer from a pre-trained teacher to a student network without accessing the real dataset. However, generator-based DFKD methods struggle to ensure that the synthetic images accurately reflect the real dataset distribution. The update of the generator network relies heavily on teacher category guidance, but varying teacher prediction accuracy across categories leads to inconsistent synthetic image quality. Such variations introduce a distribution shift between synthetic and real datasets, negatively impacting student network performance during knowledge distillation. To address this challenge, we propose the SRIF, comprising two components: Student-Driven Flexible Filtering (SDFF) and Re-weighting for Independent Regularization (RIR). SDFF filters out synthetic images affected by the category distribution shift during data generation, producing a more reliable dataset. RIR, applied during distillation, encourages the student to learn stable causal relationships through sample reweighting. Both components flexibly integrate into existing DFKD frameworks, improving performance while reducing training costs.
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Jie Zhou 0001, Longxiang Gao, Wenbo Xu 0003, Li Guo 0004
IEEE Trans. Multim.4
2024 QAGait: Revisit Gait Recognition from a Quality Perspective
abstract
Gait recognition is a promising biometric method that aims to identify pedestrians from their unique walking patterns. Silhouette modality, renowned for its easy acquisition, simple structure, sparse representation, and convenient modeling, has been widely employed in controlled in-the-lab research. However, as gait recognition rapidly advances from in-the-lab to in-the-wild scenarios, various conditions raise significant challenges for silhouette modality, including 1) unidentifiable low-quality silhouettes (abnormal segmentation, severe occlusion, or even non-human shape), and 2) identifiable but challenging silhouettes (background noise, non-standard posture, slight occlusion). To address these challenges, we revisit gait recognition pipeline and approach gait recognition from a quality perspective, namely QAGait. Specifically, we propose a series of cost-effective quality assessment strategies, including Maxmial Connect Area and Template Match to eliminate background noises and unidentifiable silhouettes, Alignment strategy to handle non-standard postures. We also propose two quality-aware loss functions to integrate silhouette quality into optimization within the embedding space. Extensive experiments demonstrate our QAGait can guarantee both gait reliability and performance enhancement. Furthermore, our quality assessment strategies can seamlessly integrate with existing gait datasets, showcasing our superiority. Code is available at https://github.com/wzb-bupt/QAGait.
Zengbin Wang, Saihui Hou, Man Zhang 0005, Xu Liu 0008, Chunshui Cao, Yongzhen Huang, Shibiao Xu
AAAI8
2024 Spectral Prompt Tuning: Unveiling Unseen Classes for Zero-Shot Semantic Segmentation
abstract
Recently, CLIP has found practical utility in the domain of pixel-level zero-shot segmentation tasks. The present landscape features two-stage methodologies beset by issues such as intricate pipelines and elevated computational costs. While current one-stage approaches alleviate these concerns and incorporate Visual Prompt Training (VPT) to uphold CLIP's generalization capacity, they still fall short in fully harnessing CLIP's potential for pixel-level unseen class demarcation and precise pixel predictions. To further stimulate CLIP's zero-shot dense prediction capability, we propose SPT-SEG, a one-stage approach that improves CLIP's adaptability from image to pixel. Specifically, we initially introduce Spectral Prompt Tuning (SPT), incorporating spectral prompts into the CLIP visual encoder's shallow layers to capture structural intricacies of images, thereby enhancing comprehension of unseen classes. Subsequently, we introduce the Spectral Guided Decoder (SGD), utilizing both high and low-frequency information to steer the network's spatial focus towards more prominent classification features, enabling precise pixel-level prediction outcomes. Through extensive experiments on two public datasets, we demonstrate the superiority of our method over state-of-the-art approaches, performing well across all classes and particularly excelling in handling unseen classes.
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Li Guo 0004, Man Zhang 0005, Xiaopeng Zhang 0001
AAAI4
2024 UnionFormer: Unified-Learning Transformer with Multi-View Representation for Image Manipulation Detection and Localization
abstract
We present UnionFormer, a novel framework that inte-grates tampering clues across three views by unified learning for image manipulation detection and localization. Specifically, we construct a BSFI-Net to extract tampering features from RGB and noise views, achieving enhanced responsive-ness to boundary artifacts while modulating spatial consis-tency at different scales. Additionally, to explore the incon-sistency between objects as a new view of clues, we combine object consistency modeling with tampering detection and localization into a three-task unified learning process, allowing them to promote and improve mutually. Therefore, we acquire a unified manipulation discriminative representation under multi-scale supervision that consolidates information from three views. This integration facilitates highly effective concurrent detection and localization of tampering. We perform extensive experiments on diverse datasets, and the results show that the proposed approach outperforms state-of-the-art methods in tampering detection and localization.
Shuaibo Li, Wei Ma 0008, Jianwei Guo 0003, Shibiao Xu, Benchong Li, Xiaopeng Zhang 0001
CVPR4
2024 MIM-HD: Making Smaller Masked Autoencoder Better with Efficient Distillation
abstract
Self-supervised learning and knowledge distillation intersect to achieve exceptional performance on downstream tasks across diverse network capacities. This paper introduces MIM-HD, which implements enhancements for masked image modeling (MIM) distillation, in two key aspects. First, a vision transformer head-level relation adaptive distillation approach is proposed, allowing the student to dynamically draw multi-source knowledge from the teacher based on its evolving state, compatible with scenarios where teacher-student transformer block head count differs. Second, to address the overemphasis on the encoder and neglect of the decoder role in maintaining representation consistency in previous MIM distillations, a dual-view decoding strategy for latent visual representations is introduced, reusing the teacher’s decoder to alleviate MIM burdens on smaller networks. MIM-HD effectiveness is demonstrated through evaluations on ADE20K (mIoU) and ImageNet-1K (Acc), achieving +1.4% and +0.5% improved performance, respectively, compared to state-of-the-art methods, with substantial advantages on smaller pre-training datasets. Moreover, MIM-HD achieves superior efficiency, reducing pre-training epochs from 300 to 100.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Li Guo 0004, Jiguang Zhang, Xiaoqiang Teng, Wenbo Xu 0003
ECAI5
2024 Large-Scale STAR-RIS Assisted Mixed Near- and Far-Field SWIPT
abstract
A large-scale simultaneously transmitting and reflecting reconfigurable intelligent surface (STAR-RIS) assisted simultaneous wireless information and power transfer (SWIPT) system is investigated. In contrast to conventional single-field (near or far) models, a mixed near- and far-field SWIPT is considered, where a STAR-RIS with coupled phase-shift is utilized to support near-field energy devices and far-field information users. Under this setup, a transmit power minimization problem is formulated by jointly designing the access point beamforming and the STARRIS beamforming, subject to the quality of service requirements for power and information. To solve this intractable problem, a weight penalty based alternating optimization algorithm is proposed. Finally, numerical results validate the effectiveness of the proposed scheme. Moreover, the coupled phase-shift associated with the STAR-RIS has a more pronounced effect on far-field information users than on near-field energy devices.
Guangyu Zhu 0007, Xidong Mu, Li Guo 0004, Ao Huang, Shibiao Xu
GLOBECOM5
2024 HCF-Net: Hierarchical Context Fusion Network for Infrared Small Object Detection
abstract
Infrared small object detection is an important computer vision task involving the recognition and localization of tiny objects in infrared images, which usually contain only a few pixels. However, it encounters difficulties due to the diminutive size of the objects and the generally complex backgrounds in infrared images. In this paper, we propose a deep learning method, HCF-Net, that significantly improves infrared small object detection performance through multiple practical modules. Specifically, it includes the parallelized patch-aware attention (PPA) module, dimension-aware selective integration (DASI) module, and multi-dilated channel refiner (MDCR) module. The PPA module uses a multi-branch feature extraction strategy to capture feature information at different scales and levels. The DASI module enables adaptive channel selection and fusion. The MDCR module captures spatial features of different receptive field ranges through multiple depth-separable convolutional layers. Extensive experimental results on the SIRST infrared single-frame image dataset show that the proposed HCF-Net performs well, surpassing other traditional and deep learning models. Code is available at https://github.com/zhengshuchen/HCFNet.
Shibiao Xu, ShuChen Zheng, Rongtao Xu, Changwei Wang 0001, Jiguang Zhang, Xiaoqiang Teng, Ao Li 0002, Li Guo 0004
ICME1
2024 DefFusion: Deformable Multimodal Representation Fusion for 3D Semantic Segmentation
abstract
The complementarity between camera and LiDAR data makes fusion methods a promising approach to improve 3D semantic segmentation performance. Recent transformer-based methods have also demonstrated superiority in segmentation. However, multimodal solutions incorporating transformers are underexplored and face two key inherent difficulties: over-attention and noise from different modal data. To overcome these challenges, we propose a Deformable Multimodal Representation Fusion (DefFusion) framework consisting mainly of a Deformable Representation Fusion Transformer and Dynamic Representation Augmentation Modules. The Deformable Representation Fusion Transformer introduces the deformable mechanism in multimodal fusion, avoiding over-attention and improving efficiency by adaptively modeling a 2D key/value set for a given 3D query, thus enabling multimodal fusion with higher flexibility. To enhance the 2D representation and 3D representation, the Dynamic Representation Enhancement Module is proposed to dynamically remove noise in the input representation via Dynamic Grouped Representation Generation and Dynamic Mask Generation. Extensive experiments validate that our model achieves the best 3D semantic segmentation performance on SemanticKITTI and NuScenes benchmarks.
Rongtao Xu, Changwei Wang 0001, Duzhen Zhang, Man Zhang 0005, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
ICRA5
2024 LRGAN: Learnable Weighted Recurrent Generative Adversarial Network for End-to-End Shadow Generation
abstract
In augmented reality(AR) applications, it is a challenging task to generate virtual object shadows while maintaining the precision and consistency of virtual and real areas. To achieve the above target, we propose a learnable weighted recurrent generative adversarial network(LRGAN) for end-to-end shadow generation. Without any additional computational overhead, LRGAN only needs to analyze the background context to create a bridge between the target shadows and the background. Our model incorporates multiple progressive steps to recurrently compute the precise reference masks, based on which a fine-grained shadow generation module generates the shadows. A learnable weighted fusion module, which can normalize pixel values to deal with pixel overflow, fuses the generated shadows with the original image. In addition, we adopt the combined method of module training and the whole model training. Experimental results show that our proposed LRGAN not only improves the plausibility of shadow location and shape but also achieves color harmony in the shadow areas. In the absence of other prior knowledge or post-processing, it outperforms the State-of-the-Art end-to-end methods.
Junsheng Xue, Hai Huang 0001, Zhong Zhou, Shibiao Xu, Aoran Chen
IJCNN4
2024 A Comprehensive Investigation on Speaker Augmentation for Speaker Recognition
abstract
Data augmentation (DA) has played a pivotal role in the success of deep speaker recognition.Current DA techniques primarily focus on speaker-preserving augmentation, which does not change the speaker trait of the speech and does not create new speakers.Recent research has shed light on the potential of speaker augmentation, which generates new speakers to enrich the training dataset.In this study, we delve into two speaker augmentation approaches: speed perturbation (SP) and vocal tract length perturbation (VTLP).Despite the empirical utilization of both methods, a comprehensive investigation into their efficacy is lacking.Our study, conducted using two public datasets, VoxCeleb and CN-Celeb, revealed that both SP and VTLP are proficient at generating new speakers, leading to significant performance improvements in speaker recognition.Furthermore, they exhibit distinct properties in sensitivity to perturbation factors and data complexity, hinting at the potential benefits of their fusion.Our research underscores the substantial potential of speaker augmentation, highlighting the importance of in-depth exploration and analysis.
Shibiao Xu, Lantian Li, Dong Wang 0013
INTERSPEECH2
2024 StableMoFusion: Towards Robust and Efficient Diffusion-based Motion Generation Framework
abstract
Thanks to the powerful generative capacity of diffusion models, recent years have witnessed rapid progress in human motion generation. Existing diffusion-based methods employ disparate network architectures and training strategies. The effect of the design of each component is still unclear. In addition, the iterative denoising process consumes considerable computational overhead, which is prohibitive for real-time scenarios such as virtual characters and humanoid robots. For this reason, we first conduct a comprehensive investigation into network architectures, training strategies, and inference process. Based on the profound analysis, we tailor each component for efficient high-quality human motion generation. Despite the promising performance, the tailored model still suffers from foot skating which is an ubiquitous issue in diffusion-based solutions. To eliminate footskate, we identify foot-ground contact and correct foot motions along the denoising process. By organically combining these well-designed components together, we present StableMoFusion, a robust and efficient framework for human motion generation. Extensive experimental results show that our StableMoFusion performs favorably against current state-of-the-art methods.
Chuanchen Luo, Yuxi Wang 0001, Shibiao Xu, Zhaoxiang Zhang 0001, Man Zhang 0005, Junran Peng
ACM Multimedia5
2024 Visual Harmony: LLM's Power in Crafting Coherent Indoor Scenes from Images
Genghao Zhang, Yuxi Wang 0001, Chuanchen Luo, Shibiao Xu, Junran Peng, Man Zhang 0005
PRCV (6)4
2024 Incomplete multi-view clustering via local and global bagging of anchor graphs
Ao Li 0002, Haoyue Xu, Hailu Yang 0001, Shibiao Xu
Expert Syst. Appl.5
2024 Key-point-guided adaptive convolution and instance normalization for continuous transitive face reenactment of any person
abstract
Abstract Face reenactment technology is widely applied in various applications. However, the reconstruction effects of existing methods are often not quite realistic enough. Thus, this paper proposes a progressive face reenactment method. First, to make full use of the key information, we propose adaptive convolution and instance normalization to encode the key information into all learnable parameters in the network, including the weights of the convolution kernels and the means and variances in the normalization layer. Second, we present continuous transitive facial expression generation according to all the weights of the network generated by the key points, resulting in the continuous change of the image generated by the network. Third, in contrast to classical convolution, we apply the combination of depth‐ and point‐wise convolutions, which can greatly reduce the number of weights and improve the efficiency of training. Finally, we extend the proposed face reenactment method to the face editing application. Comprehensive experiments demonstrate the effectiveness of the proposed method, which can generate a clearer and more realistic face from any person and is more generic and applicable than other methods.
Shibiao Xu, Miao Hua, Jiguang Zhang, Zhaohui Zhang 0002, Xiaopeng Zhang 0001
Comput. Animat. Virtual Worlds1
2024 Graph t-SNE multi-view autoencoder for joint clustering and completion of incomplete multi-view data
Ao Li 0002, Shibiao Xu, Yuan Cheng 0004
Knowl. Based Syst.3
2024 DomainFeat: Learning Local Features With Domain Adaptation
abstract
Accurate and efficient keypoint detection and description is a fundamental step in various computer vision tasks. In this paper, we extract robust descriptors and detect accurate keypoints by learning local Features with Domain adaptation (DomainFeat). Specifically, our Domainfeat includes image-level domain invariance supervision, pixel-level domain consistency supervision, Pixel-Adaptive keypoint Detection(PA-Det), and cross-domain dataset with domain stable point supervision. First, we introduce the image-level domain invariance supervision to make the high-level feature distributions from different domains close by fusing domain-invariant representations in the decoder. Furthermore, to compensate for the inconsistency between descriptors corresponding to the keypoints at the pixel level, we propose the pixel-level domain consistency supervision. Then we present the Pixel-Adaptive keypoint Detection to efficiently detect accurate keypoints, which can improve accuracy by enhancing the local consistency of heatmaps. Finally, we propose an efficient approach to construct data and supervision labels in diverse domains, which can tackle complex application scenarios. With these novel modules and supervision methods, our DomainFeat can make feature detectors more accurate and descriptors more robust. Extensive experiments confirm that Domainfeat achieves state-of-the-art performance on benchmarks such as Aachen-Day-Night localization, HPatches image matching, and the challenging DNIM dataset.
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Weiliang Meng, Bin Fan 0001, Xiaopeng Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 Exploring Intrinsic Discrimination and Consistency for Weakly Supervised Object Localization
abstract
Weakly supervised object localization (WSOL) is a challenging and promising task that aims to localize objects solely based on the supervision of image category labels. In the absence of annotated bounding boxes, WSOL methods must employ the intrinsic properties of the image classification task pipeline to generate object localizations. In this work, we propose a WSOL method for exploring the Intrinsic Discrimination and Consistency in the image classification task pipeline, and call it as IDC. First, we develop a Triplet Metrics Based Foreground Modeling (TMFM) framework to directly predict object foreground regions using intrinsic discrimination. Unlike Class Activation Map (CAM) based methods that also rely on intrinsic discrimination, our TMFM framework alleviates the problem of only focusing on the most discriminative parts by optimizing foreground and background regions synergistically. Second, we design a Dual Geometric Transformation Consistency Constraints (DGTC2) training strategy to introduce additional supervision and regularization constraints for WSOL by leveraging intrinsic geometric transformation consistency. The proposed pixel-wise and object-wise consistency constraint losses cost-effectively provide spontaneous supervision for WSOL. Extensive experiments show that our IDC method achieves significant and consistent performance gains compared to existing state-of-the-art WSOL approaches. Code is available at: https://github.com/vignywang/IDC.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Ruisheng Wang 0001, Xiaopeng Zhang 0001
IEEE Trans. Image Process.3
2024 SkinFormer: Learning Statistical Texture Representation With Transformer for Skin Lesion Segmentation
abstract
Accurate skin lesion segmentation from dermoscopic images is of great importance for skin cancer diagnosis. However, automatic segmentation of melanoma remains a challenging task because it is difficult to incorporate useful texture representations into the learning process. Texture representations are not only related to the local structural information learned by CNN, but also include the global statistical texture information of the input image. In this paper, we propose a transFormer network (SkinFormer) that efficiently extracts and fuses statistical texture representation for Skin lesion segmentation. Specifically, to quantify the statistical texture of input features, a Kurtosis-guided Statistical Counting Operator is designed. We propose Statistical Texture Fusion Transformer and Statistical Texture Enhance Transformer with the help of Kurtosis-guided Statistical Counting Operator by utilizing the transformer's global attention mechanism. The former fuses structural texture information and statistical texture information, and the latter enhances the statistical texture of multi-scale features. Extensive experiments on three publicly available skin lesion datasets validate that our SkinFormer outperforms other SOAT methods, and our method achieves 93.2% Dice score on ISIC 2018. It can be easy to extend SkinFormer to segment 3D images in the future.
Rongtao Xu, Changwei Wang 0001, Jiguang Zhang, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
IEEE J. Biomed. Health Informatics4
2024 PSTNet: Enhanced Polyp Segmentation With Multi-Scale Alignment and Frequency Domain Integration
abstract
Accurate segmentation of colorectal polyps in colonoscopy images is crucial for effective diagnosis and management of colorectal cancer (CRC). However, current deep learning-based methods primarily rely on fusing RGB information across multiple scales, leading to limitations in accurately identifying polyps due to restricted RGB domain information and challenges in feature misalignment during multi-scale aggregation. To address these limitations, we propose the Polyp Segmentation Network with Shunted Transformer (PSTNet), a novel approach that integrates both RGB and frequency domain cues present in the images. PSTNet comprises three key modules: the Frequency Characterization Attention Module (FCAM) for extracting frequency cues and capturing polyp characteristics, the Feature Supplementary Alignment Module (FSAM) for aligning semantic information and reducing misalignment noise, and the Cross Perception localization Module (CPM) for synergizing frequency cues with high-level semantics to achieve efficient polyp segmentation. Extensive experiments on challenging datasets demonstrate PSTNet's significant improvement in polyp segmentation accuracy across various metrics, consistently outperforming state-of-the-art methods. The integration of frequency domain cues and the novel architectural design of PSTNet contribute to advancing computer-assisted polyp segmentation, facilitating more accurate diagnosis and management of CRC.
Rongtao Xu, Changwei Wang 0001, Xiuli Li, Shibiao Xu, Li Guo 0004
IEEE J. Biomed. Health Informatics5
2024 DTTCNet: Time-to-Collision Estimation With Autonomous Emergency Braking Using Multi-Scale Transformer Network
abstract
The rapid advancement of autonomous driving technologies has brought the significance of Autonomous Emergency Braking (AEB) systems, which are paramount in mitigating collision risk and elevating road safety by preemptively applying brakes when a potential collision is detected. Within the core mechanisms of AEB systems, the Time-to-Collision (TTC) estimation plays a pivotal role, in quantitatively determining the criticality and timing for initiating braking interventions. However, existing TTC estimation approaches exhibit sensitivity to diverse driving scenarios, compromising the performance of AEB systems, especially in instantaneous situations. To address these issues, this paper presents DTTCNet, a novel supervised deep learning model for TTC estimation that leverages multi-scale transformer architectures and multi-task losses, thereby enhancing precision and boosting system performance. The DTTCNet first extracts spatiotemporal features from raw sensor data and utilizes a supervised training strategy. The multi-scale transformer architecture effectively captures variations across different scales, while the multi-task loss function optimizes the network training performance. Our experimental results on a challenging dataset demonstrate that DTTCNet achieves approximately 20% performance improvements over existing methods in terms of accuracy. This signifies a promising approach to augmenting the safety of autonomous driving systems with the integration of aftermarket mobile devices (e.g., Mobileye and Bosch products).
Xiaoqiang Teng, Shibiao Xu, Deke Guo, Yulan Guo, Weiliang Meng, Xiaopeng Zhang 0001
IEEE Trans. Mob. Comput.2
2024 Wave-Like Class Activation Map With Representation Fusion for Weakly-Supervised Semantic Segmentation
abstract
The Class Activation Map (CAM) is widely used to generate pseudo-labels for Weakly Supervised Semantic Segmentation (WSSS), while it does not adequately consider the modeling of foreground-independent information, resulting in prone to false positive pixels. In this paper, we propose a Wave-like Class Activation Map (WaveCAM) from the perspective of representation fusion and dynamic aggregation representation to alleviate the above problem. Specifically, our WaveCAM includes the foreground-aware representation modeling that enhances perception of foreground information, and the foreground-independent representation modeling that enhances perception of foreground-independent information, and a representation-adaptive fusion module that fuses the two representations. Both representations are expressed as wave functions with amplitude and phase to dynamically aggregate representations and extract semantic information after initialization, and they are fused through the adaptive fusion module to obtain an output containing rich semantic information. Extensive experiments on PASCAL VOC 2012 dataset and MS COCO 2014 dataset validate that our WaveCAM can easily embed multi-stage WSSS and end-to-end WSSS, achieving the state-of-the-art performance.
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
IEEE Trans. Multim.3
2024 Accurate Lung Nodule Segmentation With Detailed Representation Transfer and Soft Mask Supervision
abstract
Accurate lung lesion segmentation from computed tomography (CT) images is crucial to the analysis and diagnosis of lung diseases, such as COVID-19 and lung cancer. However, the smallness and variety of lung nodules and the lack of high-quality labeling make the accurate lung nodule segmentation difficult. To address these issues, we first introduce a novel segmentation mask named " soft mask," which has richer and more accurate edge details description and better visualization, and develop a universal automatic soft mask annotation pipeline to deal with different datasets correspondingly. Then, a novel network with detailed representation transfer and soft mask supervision (DSNet) is proposed to process the input low-resolution images of lung nodules into high-quality segmentation results. Our DSNet contains a special detailed representation transfer module (DRTM) for reconstructing the detailed representation to alleviate the small size of lung nodules images and an adversarial training framework with soft mask for further improving the accuracy of segmentation. Extensive experiments validate that our DSNet outperforms other state-of-the-art methods for accurate lung nodule segmentation, and has strong generalization ability in other accurate medical segmentation tasks with competitive results. Besides, we provide a new challenging lung nodules segmentation dataset for further studies (https://drive.google.com/file/d/15NNkvDTb_0Ku0IoPsNMHezJRTH1Oi1wm/view?usp=sharing).
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Jun Xiao 0005, Xiaopeng Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 Robust Resource Allocation for STAR-RIS Assisted SWIPT Systems
abstract
A simultaneously transmitting and reflecting reconfigurable intelligent surface (STAR-RIS) assisted simultaneous wireless information and power transfer (SWIPT) system is proposed. More particularly, an STAR-RIS is deployed to assist in the information/power transfer from a multi-antenna access point (AP) to multiple single-antenna information users (IUs) and energy users (EUs), where two practical STAR-RIS operating protocols, namely energy splitting (ES) and time switching (TS), are employed. Under the imperfect channel state information (CSI) condition, a multi-objective optimization problem (MOOP) framework, that simultaneously maximizes the minimum data rate and minimum harvested power, is employed to investigate the fundamental rate-energy trade-off between IUs and EUs. To obtain the optimal robust resource allocation strategy, the MOOP is first transformed into a single-objective optimization problem (SOOP) via the ϵ-constraint method, which is then reformulated by approximating semi-infinite inequality constraints with the S-procedure. For ES, an alternating optimization (AO)-based algorithm is proposed to jointly design AP active beamforming and STAR-RIS passive beamforming, where a penalty method is leveraged in STAR-RIS beamforming design. Furthermore, the developed algorithm is extended to optimize the time allocation policy and beamforming vectors in a two-layer iterative manner for TS. Numerical results reveal that: 1) deploying STAR-RISs achieves a significant performance gain over conventional RISs, especially in terms of harvested power for EUs; 2) the ES protocol obtains a better user fairness performance when focusing only on IUs or EUs, while the TS protocol yields a better balance between IUs and EUs; 3) the imperfect CSI affects IUs more significantly than EUs, whereas TS can confer a more robust design to attenuate these effects.
Guangyu Zhu 0007, Xidong Mu, Li Guo 0004, Ao Huang, Shibiao Xu
IEEE Trans. Wirel. Commun.5
2023 Self Correspondence Distillation for End-to-End Weakly-Supervised Semantic Segmentation
abstract
Efficiently training accurate deep models for weakly supervised semantic segmentation (WSSS) with image-level labels is challenging and important. Recently, end-to-end WSSS methods have become the focus of research due to their high training efficiency. However, current methods suffer from insufficient extraction of comprehensive semantic information, resulting in low-quality pseudo-labels and sub-optimal solutions for end-to-end WSSS. To this end, we propose a simple and novel Self Correspondence Distillation (SCD) method to refine pseudo-labels without introducing external supervision. Our SCD enables the network to utilize feature correspondence derived from itself as a distillation target, which can enhance the network's feature learning process by complementing semantic information. In addition, to further improve the segmentation accuracy, we design a Variation-aware Refine Module to enhance the local consistency of pseudo-labels by computing pixel-level variation. Finally, we present an efficient end-to-end Transformer-based framework (TSCD) via SCD and Variation-aware Refine Module for the accurate WSSS task. Extensive experiments on the PASCAL VOC 2012 and MS COCO 2014 datasets demonstrate that our method significantly outperforms other state-of-the-art methods. Our code is available at https://github.com/Rongtao-Xu/RepresentationLearning/tree/main/SCD-AAAI2023.
Rongtao Xu, Changwei Wang 0001, Jiaxi Sun, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
AAAI4
2023 Audio-Driven Lips and Expression on 3D Human Face
Weiliang Meng, Shibiao Xu, Xiaopeng Zhang 0001
CGI4
2023 Robust Beamforming Design for STAR-RIS Assisted SWIPT Systems
abstract
A simultaneously transmitting and reflecting reconfigurable intelligent surface (STAR-RIS) assisted simultaneous wireless information and power transfer (SWIPT) framework is proposed. More particularly, an STAR-RIS is deployed to assist in SWIPT from a multi-antenna access point (AP) to multiple single-antenna information users (IUs) and energy users (EUs). Due to the near-passive operation of the STAR-RIS, a more practical setup under the assumption of imperfect channel state information is investigated. The max-min fairness optimization problem is formulated to maximize the minimum power harvested by EUs, subject to the signal-to-interference-plus-noise ratio (SINR) constraints for IUs. To tackle this non-convex problem, an alternating optimization (AO) based algorithm is proposed for robust beamforming design. We first approximate the semiinfinite inequality constraints with S-procedure, then the AP active beamforming and the STAR-RIS passive beamforming are alternatively designed, where a penalty based approach is leveraged for STAR-RIS reconfiguration. Numerical results demonstrate that: i) the significant performance gains can be achieved by the proposed scheme over the baseline schemes; and ii) more STAR-RIS elements and higher SINR requirements weaken the robustness of the EU performance in terms of energy harvesting.
Guangyu Zhu 0007, Xidong Mu, Li Guo 0004, Ao Huang, Shibiao Xu
ICC5
2023 Treating Pseudo-labels Generation as Image Matting for Weakly Supervised Semantic Segmentation
abstract
Generating accurate pseudo-labels under the supervision of image categories is a crucial step in Weakly Supervised Semantic Segmentation (WSSS). In this work, we propose a Mat-Label pipeline that provides a fresh way to treat WSSS pseudo-labels generation as an image matting task. By taking a trimap as input which specifies the foreground, background and unknown regions, the image matting task outputs an object mask with fine edges. The intuition behind our Mat-Label is that generating trimap is much easier than generating pseudo-labels directly under weakly supervised setting. Although current CAM-based methods are off-the-shelf solutions for generating a trimap, they suffer from cross-category and foreground-background pixel prediction confusion. To solve this problem, we develop a Double Decoupled Class Activation Map (D2CAM) for Mat-Label to generate a high-quality trimap. By drawing on the idea of metric learning, we explicitly model class activation map with category decoupling and foreground-background decoupling. We also design two simple yet effective refinement constraints for D2CAM to stabilize optimization and eliminate non-exclusive activation. Extensive experiments validate that our Mat-Label achieves substantial and consistent performance gains compared to current state-of-the-art WSSS approaches.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
ICCV3
2023 FeaCo: Reaching Robust Feature-Level Consensus in Noisy Pose Conditions
abstract
Collaborative perception offers a promising solution to overcome challenges such as occlusion and long-range data processing. However, limited sensor accuracy leads to noisy poses that misalign observations among vehicles. To address this problem, we propose the FeaCo, which achieves robust Feature-level Consensus among collaborating agents in noisy pose conditions without additional training. We design an efficient Pose-error Rectification Module (PRM) to align derived feature maps from different vehicles, reducing the adverse effect of noisy pose and bandwidth requirements. We also provide an effective multi-scale Cross-level Attention Module (CAM) to enhance information aggregation and interaction between various scales. Our FeaCo outperforms all other localization rectification methods, as validated on both the collaborative perception simulation dataset OPV2V and real-world dataset V2V4Real, reducing heading error and enhancing localization accuracy across various error levels. Our code is available at: https://github.com/jmgu0212/FeaCo.git.
Jiaming Gu, Muyang Zhang, Weiliang Meng, Shibiao Xu, Jiguang Zhang, Xiaopeng Zhang 0001
ACM Multimedia5
2023 LandmarkGait: Intrinsic Human Parsing for Gait Recognition
abstract
Gait recognition is an emerging biometric technology for identifying pedestrians based on their unique walking patterns. In past gait recognition, global-based methods are inadequate to meet the growing demand for accuracy, while commonly used part-based methods provided coarse and inaccurate feature representation for specific body parts. Human parsing appears to be a better option for accurately representing specific and complete body parts in gait recognition. However, its practical application in gait recognition is often hindered by missing RGB modality, lack of annotated body parts, and difficulty in balancing parsing quantity and quality. To address this issue, we propose LandmarkGait, an accessible and alternative parsing-based solution for gait recognition. LandmarkGait introduces an unsupervised landmark discovery network to transform the dense silhouette into a finite set of landmarks with remarkable consistency across various conditions. By grouping landmarks subsets corresponding to distinct body part regions, following a reconstruction task and further refinement from high-quality input silhouettes, we can directly obtain fine-grained parsing results from original binary silhouettes in an unsupervised manner. Moreover, we also develop a multi-scale feature extractor that simultaneously captures global and parsing feature representations based on the integrity and flexibility of specific body parts. Extensive experiments demonstrate that our LandmarkGait can extract more stable features and exhibit significant performance improvement under all conditions, especially in various dressing conditions. Code is available at https://github.com/wzb-bupt/LandmarkGait.
Zengbin Wang, Saihui Hou, Man Zhang 0005, Xu Liu 0008, Chunshui Cao, Yongzhen Huang, Shibiao Xu
ACM Multimedia7
2023 Deep Deformation Detail Synthesis for Thin Shell Models
abstract
Abstract In physics‐based cloth animation, rich folds and detailed wrinkles are achieved at the cost of expensive computational resources and huge labor tuning. Data‐driven techniques make efforts to reduce the computation significantly by utilizing a preprocessed database. One type of methods relies on human poses to synthesize fitted garments, but these methods cannot be applied to general cloth animations. Another type of methods adds details to the coarse meshes obtained through simulation, which does not have such restrictions. However, existing works usually utilize coordinate‐based representations which cannot cope with large‐scale deformation, and requires dense vertex correspondences between coarse and fine meshes. Moreover, as such methods only add details, they require coarse meshes to be sufficiently close to fine meshes, which can be either impossible, or require unrealistic constraints to be applied when generating fine meshes. To address these challenges, we develop a temporally and spatially as‐consistent‐as‐possible deformation representation (named TS‐ACAP) and design a DeformTransformer network to learn the mapping from low‐resolution meshes to ones with fine details. This TS‐ACAP representation is designed to ensure both spatial and temporal consistency for sequential large‐scale deformations from cloth animations. With this TS‐ACAP representation, our DeformTransformer network first utilizes two mesh‐based encoders to extract the coarse and fine features using shared convolutional kernels, respectively. To transduct the coarse features to the fine ones, we leverage the spatial and temporal Transformer network that consists of vertex‐level and frame‐level attention mechanisms to ensure detail enhancement and temporal coherence of the prediction. Experimental results show that our method is able to produce reliable and realistic animations in various datasets at high frame rates with superior detail synthesis abilities compared to existing methods.
Lin Gao 0004, Jie Yang 0038, Shibiao Xu, Juntao Ye, Xiaopeng Zhang 0001, Yukun Lai
Comput. Graph. Forum4
2023 Automatic polyp segmentation via image-level and surrounding-level context fusion deep neural network
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
Eng. Appl. Artif. Intell.3
2023 Dual-stream Representation Fusion Learning for accurate medical image segmentation
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
Eng. Appl. Artif. Intell.3
2023 Attention Weighted Local Descriptors
abstract
Local features detection and description are widely used in many vision applications with high industrial and commercial demands. With large-scale applications, these tasks raise high expectations for both the accuracy and speed of local features. Most existing studies on local features learning focus on the local descriptions of individual keypoints, which neglect their relationships established from global spatial awareness. In this paper, we present AWDesc with a consistent attention mechanism (CoAM) that opens up the possibility for local descriptors to embrace image-level spatial awareness in both the training and matching stages. For local features detection, we adopt local features detection with feature pyramid to obtain more stable and accurate keypoints localization. For local features description, we provide two versions of AWDesc to cope with different accuracy and speed requirements. On the one hand, we introduce Context Augmentation to address the inherent locality of convolutional neural networks by injecting non-local context information, so that local descriptors can "look wider to describe better". Specifically, well-designed Adaptive Global Context Augmented Module (AGCA) and Diverse Surrounding Context Augmented Module (DSCA) are proposed to construct robust local descriptors with context information from global to surrounding. On the other hand, we design an extremely lightweight backbone network coupled with the proposed special knowledge distillation strategy to achieve the best trade-off in accuracy and speed. What is more, we perform thorough experiments on image matching, homography estimation, visual localization, and 3D reconstruction tasks, and the results demonstrate that our method surpasses the current state-of-the-art local descriptors. Code is available at: https://github.com/vignywang/AWDesc.
Changwei Wang 0001, Rongtao Xu, Ke Lu 0002, Shibiao Xu, Weiliang Meng, Bin Fan 0001, Xiaopeng Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Toward Accurate and Efficient Road Extraction by Leveraging the Characteristics of Road Shapes
abstract
Automatically extracting roads from very high resolution (VHR) remote sensing images is of great importance in a wide range of remote sensing applications. However, complex shapes of roads (i.e., long, geometrically deformed, and thin) always affected the extraction accuracy, which is one of the challenges of road extraction. Based on the insight into road shape characteristics, we propose a novel road shape aware network (RSANet) to achieve efficient and accurate road extraction. First, we introduce the Efficient Strip Transformer Module (ESTM) to efficiently capture the global context to model the long-distance dependence required by the long roads. Second, we design a Geometric Deformation Estimation Module (GDEM) to adaptively extract the context from the shape deformation caused by shooting roads from different perspectives. Third, we provide a simple but effective Road Edge Focal Loss (REF loss) to make the network focus on optimizing the pixels around the road to alleviate the unbalanced distribution of foreground and background pixels caused by the roads being too thin. Finally, we conduct extensive evaluations on public datasets to verify the effectiveness of RSANet and each of the proposed components. Experiments validate that our RSANet outperforms state-of-the-art methods for road extraction in remote sensing images.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Ruisheng Wang 0001, Jiguang Zhang, Xiaopeng Zhang 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 RSSFormer: Foreground Saliency Enhancement for Remote Sensing Land-Cover Segmentation
abstract
High spatial resolution (HSR) remote sensing images contain complex foreground-background relationships, which makes the remote sensing land cover segmentation a special semantic segmentation task. The main challenges come from the large-scale variation, complex background samples and imbalanced foreground-background distribution. These issues make recent context modeling methods sub-optimal due to the lack of foreground saliency modeling. To handle these problems, we propose a Remote Sensing Segmentation framework (RSSFormer), including Adaptive TransFormer Fusion Module, Detail-aware Attention Layer and Foreground Saliency Guided Loss. Specifically, from the perspective of relation-based foreground saliency modeling, our Adaptive Transformer Fusion Module can adaptively suppress background noise and enhance object saliency when fusing multi-scale features. Then our Detail-aware Attention Layer extracts the detail and foreground-related information via the interplay of spatial attention and channel attention, which further enhances the foreground saliency. From the perspective of optimization-based foreground saliency modeling, our Foreground Saliency Guided Loss can guide the network to focus on hard samples with low foreground saliency responses to achieve balanced optimization. Experimental results on LoveDA datasets, Vaihingen datasets, Potsdam datasets and iSAID datasets validate that our method outperforms existing general semantic segmentation methods and remote sensing segmentation methods, and achieves a good compromise between computational overhead and accuracy. Our code is available at https://github.com/Rongtao-Xu/RepresentationLearning/tree/main/RSSFormer-TIP2023.
Rongtao Xu, Changwei Wang 0001, Jiguang Zhang, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
IEEE Trans. Image Process.4
2023 CNDesc: Cross Normalization for Local Descriptors Learning
abstract
For a long time, the local descriptors learning benefited from the use of L2 normalization, which projects the descriptor space onto the hypersphere. However, there is no free lunch in the world. Although hypersphere description space stabilizes the optimization and improves the repeatability of the descriptors, it causes the descriptors to have a denser distribution, which reduces the discrimination between descriptors and leads to some incorrect matches. To alleviate this problem, we propose the learnablecross normalizationtechnology as an alternative to L2 normalization, which can achieve a consistent improvement in several of the current popular local descriptors. In addition, we propose an ER-Backbone that can efficiently reuse features in descriptors extraction and an IDC Loss that can provide an image-level description space distribution consistency constraint to further stimulate the performance of the local descriptors. Based on the above innovations, we provide a novel local descriptors extraction method named CNDesc. We perform experiments on image matching, homography estimation, 3D reconstruction, and visual localization tasks, and the results demonstrate that our CNDesc surpasses the current state-of-the-art local descriptors. Our code is available athttps://github.com/vignywang/CNDesc.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
IEEE Trans. Multim.3
2023 Image Manipulation Localization Using Attentional Cross-Domain CNN Features
abstract
Along with the advancement of manipulation technologies, image modification is becoming increasingly convenient and imperceptible. To tackle the challenging image tampering detection problem, this article presents an attentional cross-domain deep architecture, which can be trained end-to-end. This architecture is composed of three convolutional neural network (CNN) streams to extract three types of features, including visual perception, resampling, and local inconsistency features, from spatial and frequency domains. The multitype and cross-domain features are then combined to formulate hybrid features to distinguish manipulated regions from nonmanipulated parts. Compared with other deep architectures, the proposed one spans a more complementary and discriminative feature space by integrating richer types of features from different domains in a unified end-to-end trainable framework and thus can better capture artifacts caused by different types of manipulations. In addition, we design and train a module called tampering discriminative attention network (TDA-Net) to highlight suspicious parts. These part-level representations are then integrated with the global ones to further enhance the discriminating capability of the hybrid features. To adequately train the proposed architecture, we synthesize a large dataset containing various types of manipulations based on DRESDEN and COCO. Experiments on four public datasets demonstrate that the proposed model can localize various manipulations and achieve the state-of-the-art performance. We also conduct ablation studies to verify the effectiveness of each stream and the TDA-Net module.
Shuaibo Li, Shibiao Xu, Wei Ma 0008, Qiu Zong
IEEE Trans. Neural Networks Learn. Syst.2
2023 HTCViT: an effective network for image classification and segmentation based on natural disaster datasets
Wei Li 0237, Muyang Zhang, Weiliang Meng, Shibiao Xu, Xiaopeng Zhang 0001
Vis. Comput.5
2022 MTLDesc: Looking Wider to Describe Better
abstract
Limited by the locality of convolutional neural networks, most existing local features description methods only learn local descriptors with local information and lack awareness of global and surrounding spatial context. In this work, we focus on making local descriptors ``look wider to describe better'' by learning local Descriptors with More Than Local information (MTLDesc). Specifically, we resort to context augmentation and spatial attention mechanism to make the descriptors obtain non-local awareness. First, Adaptive Global Context Augmented Module and Diverse Local Context Augmented Module are proposed to construct robust local descriptors with context information from global to local. Second, we propose the Consistent Attention Weighted Triplet Loss to leverage spatial attention awareness in both optimization and matching of local descriptors. Third, Local Features Detection with Feature Pyramid is proposed to obtain more stable and accurate keypoints localization. With the above innovations, the performance of the proposed MTLDesc significantly surpasses the current state-of-the-art local descriptors on HPatches, Aachen Day-Night localization and InLoc indoor localization benchmarks. Our code is available at https://github.com/vignywang/MTLDesc.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Bin Fan 0001, Xiaopeng Zhang 0001
AAAI4
2022 DOMAINDESC: Learning Local Descriptors With Domain Adaptation
abstract
Robust and efficient local descriptor is crucial in a wide range of applications. In this paper, we propose a novel descriptor DomainDesc which is invariant as much as possible by learning local Descriptor with Domain adaptation. We design the feature-level domain adaptation loss to improve robustness of our DomainDesc by punishing inconsistent high-level feature distributions of different images, while we present the pixel-level cross-domain consistency loss to compensate for the inconsistency between the descriptors corresponding to the keypoints at the pixel level. Besides, we adopt a new architecture to make the descriptor contain as much information as possible, and combine triplet loss and cross-domain consistency loss for descriptor supervision to ensure the distinguished ability of our descriptor. Finally, we give a cross-domain dataset generation strategy to quickly construct our training dataset for diverse domains to adapt to complex application scenarios. Experiments validate that our DomainDesc achieves state-of-the-art performances on HPatches image matching benchmark and Aachen-Day-Night localization benchmark.
Rongtao Xu, Changwei Wang 0001, Bin Fan 0001, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
ICASSP5
2022 Softgan: Towards Accurate Lung Nodule Segmentation via Soft Mask Supervision
abstract
Accurate lung nodule segmentation from Computed Tomog-raphy (CT) images is crucial to the analysis and diagnosis of lung diseases such as COVID-19 and lung cancer. How-ever, due to the variety of lung nodules and the lack of high-quality labeling, accurate lung nodule segmentation is still a challenging problem. In this paper, we propose a novel paradigm including an automatic accurate annotation pipeline and a segmentation network for this task. First, we introduce a new segmentation mask representation named Soft Mask which has richer and more accurate edge details description and better visualization, and we design a universal automatic Soft Mask annotation pipeline to deal with different datasets. Besides, we provide a new challenging lung nodules segmen-tation dataset with traditional binarized masks and our soft masks for further studies. Second, we propose an effective network called SoftGAN that includes an improved back-bone and an adversarial training framework with Soft Mask, in order to improve the performance of accurate lung nodules segmentation. Extensive experiments validate that our Soft-GAN outperforms the state-of-the-art methods for accurate lung nodule segmentation. [Datasetrelease]
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Jun Xiao 0005, Qimin Peng, Xiaopeng Zhang 0001
ICME3
2022 GeoROS: Georeferenced Real-time Orthophoto Stitching with Unmanned Aerial Vehicle
abstract
Simultaneous orthophoto stitching during the flight of Unmanned Aerial Vehicles (UAV) can greatly promote the practicability and instantaneity of diverse applications such as emergency disaster rescue, digital agriculture, and cadastral survey, which is of remarkable interest in aerial photogrammetry. However, the inaccurately estimated camera poses and the intuitive fusion strategy of existing methods lead to misalignment and distortion artifacts in orthophoto mosaics. To address these issues, we propose a Georeferenced Real-time Orthophoto Stitching method (GeoROS), which can achieve efficient and accurate camera pose estimation through exploiting geolocation information in monocular visual simultaneous localization and mapping (SLAM) and fuse transformed images via orthogonality-preserving criterion. Specifically, in the SLAM process, georeferenced tracking is employed to acquire high-quality initial camera poses with a geolocation based motion model and facilitate non-linear pose optimization. Meanwhile, we design a georeferenced mapping scheme by introducing robust geolocation constraints in joint optimization of camera poses and the position of landmarks. Finally, aerial images warped with localized cameras are fused by considering both the orthogonality of camera orientation relative to the ground plane and the pixel centrality to fulfill global orthorectification. Besides, we construct two datasets with global navigation satellite system (GNSS) information of different scenarios and validate the superiority of our GeoROS method compared with state-of-the-art methods in accuracy and efficiency.
Guangze Gao, Mengke Yuan, Jiaming Gu, Weiliang Meng, Shibiao Xu, Xiaopeng Zhang 0001
IROS6
2022 DA-Net: Dual Branch Transformer and Adaptive Strip Upsampling for Retinal Vessels Segmentation
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
MICCAI (2)3
2022 Caption-Aware Medical VQA via Semantic Focusing and Progressive Cross-Modality Comprehension
abstract
Medical Visual Question Answering as a specific-domain task requires substantive prior knowledge of medicine. However, deep learning techniques encounter severe problems of limited supervision due to the scarcity of well-annotated large-scale medical VQA datasets. As an alternative to facing the data limitation problem, image captioning can be introduced to learn summary information about the picture, which is beneficial to question answering. To this end, we propose a caption-aware VQA method that can read the summary information of image content and clinic diagnoses from plenty of medical images and answer the medical question with richer multimodality features. The proposed method consists of two novel components emphasizing semantic locations and semantic content respectively. Firstly, to extract and leverage the semantic locations implied in image captioning, similarity analysis is designed to summarize the attention maps generated from image captioning by their relevance and guide the visual model to focus on the semantic-rich regions. Besides, to combine the semantic content in the generated captions, we propose a Progressive Compact Bilinear Interactions structure to achieve cross-modality comprehension over the image, question and caption features by performing bilinear attention in a gradual manner. Qualitative and quantitative experiments on various medical datasets exhibit the superiority of the proposed approach compared to the state-of-the-art methods.
Fu'ze Cong, Shibiao Xu, Li Guo 0004, Yinbing Tian
ACM Multimedia2
2022 One-stage object detection knowledge distillation via adversarial learning
Yongqiang Zhang 0007, Mingli Ding, Shibiao Xu, Yancheng Bai
Appl. Intell.4
2022 Instance segmentation of biological images using graph convolutional network
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
Eng. Appl. Artif. Intell.4
2022 Multiview Feature Aggregation for Facade Parsing
abstract
Facade image parsing is essential to the semantic understanding and 3-D reconstruction of urban scenes. Considering the occlusion and appearance ambiguity in single-view images and the easy acquisition of multiple views, in this letter, we propose a multiview enhanced deep architecture for facade parsing. The highlight of this architecture is a cross-view feature aggregation module that can learn to choose and fuse useful convolutional neural network (CNN) features from nearby views to enhance the representation of a target view. Benefitting from the multiview enhanced representation, the proposed architecture can better deal with the ambiguity and occlusion issues. Moreover, our cross-view feature aggregation module can be straightforwardly integrated into existing single-image parsing frameworks. Extensive comparison experiments and ablation studies are conducted to demonstrate the good performance of the proposed method and the validity and transportability of the cross-view feature aggregation module.
Wenguang Ma, Shibiao Xu, Wei Ma 0008, Hongbin Zha
IEEE Geosci. Remote. Sens. Lett.2
2022 Correction to: A hybrid convolutional architecture for accurate image manipulation localization at the pixel-level
Jiguang Zhang, Shibiao Xu
Multim. Tools Appl.3
2022 RTSfM: Real-Time Structure From Motion for Mosaicing and DSM Mapping of Sequential Aerial Images With Low Overlap
abstract
Inspired by simultaneous localization and mapping (SLAM) style workflow, this article presented an online sequential structure from motion (SfM) solution for high-frequency video and large baseline high-resolution aerial images with high efficiency and novel precision. First, as traditional SLAM systems are not good in processing low overlap images, based on our novel hierarchical feature matching paradigm with multihomography and BoW, we proposed a robust tracking method where the relative pose and its scale are estimated separately followed by a joint optimization by considering both perspective-n-point (PnP) and epipolar constraints. Second, to further optimize the camera poses for the sparse map and dense pointcloud reconstruction, we provided a graph-based optimization with reprojection and GPS constraints, which make the camera trajectory and map georeferenced. We also incrementally generated the dense point cloud in real time from keyframes after local mapping optimization. Finally, we use a publicly available aerial image dataset with sequences of different environments, to evaluate the effectiveness of the proposed method, meanwhile, the robust performance of our solution is demonstrated with applications of high-quality aerial images mosaic and digital surface model (DSM) reconstruction in real time. Compared with the state-of-the-art SLAM and traditional SfM methods, the presented system can output large-scale high-quality ortho-mosaic and DSM in real time with the low computational cost.
Lin Chen 0042, Xishan Zhang, Shibiao Xu, Shuhui Bu, Hongkai Jiang, Pengcheng Han, Ke Li 0005
IEEE Trans. Geosci. Remote. Sens.4
2022 Progressive Feature Learning for Facade Parsing With Occlusions
abstract
Existing deep models for facade parsing often fail in classifying pixels in heavily occluded regions of facade images due to the difficulty in feature representation of these pixels. In this paper, we solve facade parsing with occlusions by progressive feature learning. To this end, we locate the regions contaminated by occlusions via Bayesian uncertainty evaluation on categorizing each pixel in these regions. Then, guided by the uncertainty, we propose an occlusion-immune facade parsing architecture in which we progressively re-express the features of pixels in each contaminated region from easy to hard. Specifically, the outside pixels, which have reliable context from visible areas, are re-expressed at early stages; the inner pixels are processed at late stages when their surroundings have been decontaminated at the earlier stages. In addition, at each stage, instead of using regular square convolution kernels, we design a context enhancement module (CEM) with directional strip kernels, which can aggregate structural context to re-express facade pixels. Extensive experiments on popular facade datasets demonstrate that the proposed method achieves state-of-the-art performance.
Wenguang Ma, Shibiao Xu, Wei Ma 0008, Xiaopeng Zhang 0001, Hongbin Zha
IEEE Trans. Image Process.2
2022 Anomaly Matters: An Anomaly-Oriented Model for Medical Visual Question Answering
abstract
Medical images contain various abnormal regions, most of which are closely related to the lesions or diseases. The abnormality or lesion is one of the major concerns during clinical practice and therefore becomes the key in answering questions about medical images. However, the recent efforts still focus on constructing a generic Visual Question Answering framework for medical-domain tasks, which is not adequate for practical medical requirements and applications. In this paper, we present two novel medical-specific modules named multiplication anomaly sensitive module and residual anomaly sensitive module to utilize weakly supervised anomaly localization information in medical Visual Question Answering. Firstly, the proposed multiplication anomaly sensitive module designed for anomaly-related questions can mask the feature of the whole image according to the anomaly location map. Secondly, the residual anomaly sensitive module could learn a flexible anomaly feature while preserving the information of the original questioned image, which is more helpful in answering anomaly-unrelated questions. Thirdly, the transformer decoder and multi-task learning strategy are combined to further enhance the question-reasoning ability and the model generalization performance. Finally, qualitative and quantitative experiments on a variety of medical datasets exhibit the superiority of the proposed approaches compared to the state-of-the-art methods.
Fu'ze Cong, Shibiao Xu, Li Guo 0004, Yinbing Tian
IEEE Trans. Medical Imaging2
2022 Triple-strip attention mechanism-based natural disaster images classification and segmentation
Mengke Yuan, Jiaming Gu, Weiliang Meng, Shibiao Xu, Xiaopeng Zhang 0001
Vis. Comput.5
2021 A Periodic Frame Learning Approach for Accurate Landmark Localization in M-Mode Echocardiography
abstract
Anatomical landmark localization has been a key challenge for medical image analysis. Existing researches mostly adopt CNN as the main architecture for landmark localization while they are not applicable to process image modalities with periodic structure. In this paper, we propose a novel two-stage frame-level detection and heatmap regression model for accurate landmark localization in m-mode echocardiography, which promotes better integration between global context information and local appearance. Specifically, a periodic frame detection module with LSTM is designed to model periodic context and detect frames of systole and diastole from original echocardiography. Next, a CNN based heatmap regression model is introduced to predict landmark localization in each systolic or diastolic local region. Experiment results show that the proposed model achieves average distance error of 9.31, which is at a reduction by 24% comparing to baseline models.
Yinbing Tian, Shibiao Xu, Li Guo 0004, Fu'ze Cong
ICASSP2
2021 Towards Effective Adversarial Attack Against 3D Point Cloud Classification
abstract
In the domain of 3D point cloud classification, deep learning based classifiers have made significant progress, while they have been also proven to be vulnerable on the adversarial at-tack at the same time. Some recent works employ the attack methods that devised for image classification such as projected gradient descent (PGD) to attack the 3D classifiers, but their performances seem quite limited when faced with statistical operations including point cloud denoising and point cloud upsampling. In this paper, we propose ‘SmoothAttack’, a new attack that can craft adversarial point clouds robust to statistical operations. SmoothAttack can be easily applied in both global constraint and pointwise constraint. Besides, we analyze the directions of perturbations onto the point cloud during the iteration process, where SmoothAttack can some-how stabilize the direction and make full use of the adversarial budgets. Experiments validate that our ‘SmoothAttack’ can raise the attack success rates against statistical defenses up to 98% for untargeted attack and 91% for targeted attack on ModelNet40 database when fooling the classifiers Point-Net and DGCNN.
Chengcheng Ma, Weiliang Meng, Baoyuan Wu, Shibiao Xu, Xiaopeng Zhang 0001
ICME4
2021 DC-Net: Dual Context Network for 2D Medical Image Segmentation
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
MICCAI (1)3
2021 HMMN: Online metric learning for human re-identification via hard sample mining memory network
Pengcheng Han, Qing Li 0018, Cunbao Ma, Shibiao Xu, Shuhui Bu, Ke Li 0005
Eng. Appl. Artif. Intell.4
2021 Pyramid ALKNet for Semantic Parsing of Building Facade Image
abstract
The semantic parsing of building facade images is a fundamental yet challenging task in urban scene understanding. Existing works sought to tackle this task by using facade grammars or convolutional neural networks (CNNs). The former can hardly generate parsing results coherent with real images while the latter often fails to capture relationships among facade elements. In this letter, we propose a pyramid atrous large kernel (ALK) network (ALKNet) for the semantic segmentation of facade images. The pyramid ALKNet captures long-range dependencies among building elements by using ALK modules in multiscale feature maps. It makes full use of the regular structures of facades to aggregate useful nonlocal context information and thereby is capable of dealing with challenging image regions caused by occlusions, ambiguities, and so on. Experiments on both rectified and unrectified facade data sets show that ALKNet has better performances than those of state-of-the-art methods.
Wenguang Ma, Wei Ma 0008, Shibiao Xu, Hongbin Zha
IEEE Geosci. Remote. Sens. Lett.3
2021 A hybrid convolutional architecture for accurate image manipulation localization at the pixel-level
Jiguang Zhang, Shibiao Xu
Multim. Tools Appl.3
2021 Accurate Rock-Mass Extraction From Terrestrial Laser Point Clouds via Multiscale and Multiview Convolutional Feature Representation
abstract
Existing 3-D object extraction methods on terrestrial laser point clouds are further developed through filtering and labeling. However, such predefined features are heuristically designed to process generic object point clouds. Thus, existing abilities are insufficient to handle specific rock-mass point clouds. Given the complexity and diversity of terrestrial environments, the effective removal of vegetation points from rock-mass point clouds is particularly challenging. To address such problems, this study presents a novel approach for 3-D rock-mass point clouds labeling by using convolutional feature learning based on distribution priors with multiple scales and views. First, to extract discriminative features of each point for classification, we propose novel multiview supporting planes to analyze the spatial distribution and structure of its neighboring points for each category. Second, we define the multiscale spatial distribution matrix on a grid representation (e.g., the number of points projected into each cell). Last, the statistical information of points is nonlinearly combined and hierarchically compressed to generate a compact and effective convolutional feature representation for classification. The effectiveness of the proposed method is evaluated via experiments on rock-mass point clouds from different scenes. Compared with existing extraction approaches, experimental results indicate the superiority of the proposed method in terms of the precision and recall.
Yunbiao Wang, Shibiao Xu, Jun Xiao 0005, Ying Wang 0030, Lupeng Liu
IEEE Trans. Geosci. Remote. Sens.2
2021 Fast Georeferenced Aerial Image Stitching With Absolute Rotation Averaging and Planar- Restricted Pose Graph
abstract
Accurate digital orthophoto map generation from high-resolution aerial images is important in various applications. Compared with the existing commercial software and the current state-of-the-art mosaicing systems, a novel fast georeferenced orthophoto mosaicing framework is proposed in this study. The framework can adapt to the challenging requirements of high-accuracy orthoimage generations with relatively fast speed, even if the overlap rate is low. We provide appearance and spatial correlation-constrained fast low-overlap neighbor candidate query and matching. On the basis of GPS information, we introduce an absolute position and rotation-averaging strategy for global pose initialization, which is essential for the high convergence and efficiency of nonconvex pose optimization of every image. We also propose a planar-restricted global pose graph optimization method. The optimization is extremely efficient and robust considering that point clouds are parameterized to planes. Finally, we apply a matching graph-based exposure compensation and region reduction algorithm for large-scale and high-resolution image fusion with high efficiency and novel precision. Experimental results demonstrate that our method can achieve the state-of-the-art performance while maintaining high precision and robustness.
Guochen Liu, Shibiao Xu, Shuhui Bu, Hongkai Jiang
IEEE Trans. Geosci. Remote. Sens.3
2021 KGSNet: Key-Point-Guided Super-Resolution Network for Pedestrian Detection in the Wild
abstract
In real-world scenarios (i.e., in the wild), pedestrians are often far from the camera (i.e., small scale), and they often gather together and occlude with each other (i.e., heavily occluded). However, detecting these small-scale and heavily occluded pedestrians remains a challenging problem for the existing pedestrian detection methods. We argue that these problems arise because of two factors: 1) insufficient resolution of feature maps for handling small-scale pedestrians and 2) lack of an effective strategy for extracting body part information that can directly deal with occlusion. To solve the above-mentioned problems, in this article, we propose a key-point-guided super-resolution network (coined KGSNet) for detecting these small-scale and heavily occluded pedestrians in the wild. Specifically, to address factor 1), a super-resolution network is first trained to generate a clear super-resolution pedestrian image from a small-scale one. In the super-resolution network, we exploit key points of the human body to guide the super-resolution network to recover fine details of the human body region for easier pedestrian detection. To address factor 2), a part estimation module is proposed to encode the semantic information of different human body parts where four semantic body parts (i.e., head and upper/middle/bottom body) are extracted based on the key points. Finally, based on the generated clear super-resolved pedestrian patches padded with the extracted semantic body part images at the image level, a classification network is trained to further distinguish pedestrians/backgrounds from the inputted proposal regions. Both proposed networks (i.e., super-resolution network and classification network) are optimized in an alternating manner and trained in an end-to-end fashion. Extensive experiments on the challenging CityPersons data set demonstrate the effectiveness of the proposed method, which achieves superior performance over previous state-of-the-art methods, especially for those small-scale and heavily occluded instances. Beyond this, we also achieve state-of-the-art performance (i.e., 3.89% MR-2on the reasonable subset) on the Caltech data set.
Yongqiang Zhang 0007, Yancheng Bai, Mingli Ding, Shibiao Xu, Bernard Ghanem
IEEE Trans. Neural Networks Learn. Syst.4
2021 Tensor-Based Reliable Multiview Similarity Learning for Robust Spectral Clustering on Uncertain Data
abstract
Similarity graph learning is the most key technique for multiview spectral clustering. However, existing methods fail when applied to uncertain data contaminated with various types of noise in an open environment. Due to the damaged structure by noise, unreliable similar relationships are learned, which extends similarity inconsistency among views. Moreover, the high-order correlation hidden in graphs are ignored generally. To address these problems, we propose a reliable similarity learning scheme for multiview clustering on uncertain data. This method can significantly improve spectral clustering performance in a noisy environment, and the contributions of our scheme include the following three aspects: 1) Uncertain data subspace reconstruction and adaptive graph learning are combined to construct a view-specific graph from high-quality recovered data, thus improving robustness. 2) A low-rank tensor constraint is utilized to facilitate multiview fusion, where the latent high-order correlation among view graphs will be fully explored when learning the consensus graph structure. 3) Data recovery, view-specific graphs, and latent consensus tensor structure are assembled into a unified framework, to be optimized jointly for mutual benefit. Our study also develops an efficient algorithm for obtaining overall solutions. The experimental results on several datasets demonstrate that our proposed approach shows significant improvements in robustness and evaluation metrics over the comparison methods.
Ao Li 0002, Jiajia Chen 0004, Mengke Yuan, Shibiao Xu, Guanglu Sun
IEEE Trans. Reliab.6
2020 MLIFeat: Multi-level Information Fusion Based Deep Local Features
Jinge Wang 0006, Shibiao Xu, Xiaopeng Zhang 0001
ACCV (3)3
2020 DenseFusion: Large-Scale Online Dense Pointcloud and DSM Mapping for UAVs
abstract
With the rapidly developing unmanned aerial vehicles, the requirements of generating maps efficiently and quickly are increasing. To realize online mapping, we develop a real-time dense mapping framework named DenseFusion which can incrementally generates dense geo-referenced 3D point cloud, digital orthophoto map (DOM) and digital surface model (DSM) from sequential aerial images with optional GPS information. The proposed method works in real-time on standard CPUs even for processing high resolution images. Based on the advanced monocular SLAM, our system first estimates appropriate camera poses and extracts effective keyframes, and next constructs virtual stereo-pair from consecutive frame to generate pruned dense 3D point clouds; then a novel realtime DSM fusion method is proposed which can incrementally process dense point cloud. Finally, a high efficiency visualization system is developed to adopt dynamic levels of detail (LoD) method, which makes it render dense point cloud and DSM smoothly. The performance of the proposed method is evaluated through qualitative and quantitative experiments. The results indicate that compared to traditional structure from motion based approaches, the presented framework is able to output both large-scale high-quality DOM and DSM in real-time with low computational cost.
Lin Chen 0042, Shibiao Xu, Shuhui Bu, Pengcheng Han
IROS3
2020 Efficient Joint Gradient Based Attack Against SOR Defense for 3D Point Cloud Classification
abstract
Deep learning based classifiers on 3D point cloud data have been shown vulnerable to adversarial examples, while a defense strategy named Statistical Outlier Removal (SOR) is widely adopted to defend adversarial examples successfully, by discarding outlier points in the point cloud.
Chengcheng Ma, Weiliang Meng, Baoyuan Wu, Shibiao Xu, Xiaopeng Zhang 0001
ACM Multimedia4
2020 Learning across views for stereo image completion
abstract
Stereo image completion (SIC) is to fill holes existing in a pair of stereo images. SIC is more complicated than single image repairing, which needs to complete the pair of images while keeping their stereoscopic consistency. In recent years, deep learning has been introduced into single image repairing but seldom used for SIC. The authors present a novel deep learning‐based approach for SIC. In their method, an X‐shaped fully convolutional network (called SICNet) is proposed and designed to complete stereo images, which is composed of two branches of convolutional neural network layers to encode the context of the left and right images separately, a fusion module for stereo‐interactive completion, and two branches of decoders to produce completed left and right images, respectively. In consideration of both inter‐view and intra‐view cues, they introduce auxiliary networks and define comprehensive losses to train SICNet to perform single‐view coherent and cross‐view consistent completion simultaneously. Extensive experiments are conducted to show the state‐of‐the‐art performances of the proposed approach and its key components.
Wei Ma 0008, Mana Zheng, Wenguang Ma, Shibiao Xu, Xiaopeng Zhang 0001
IET Comput. Vis.4
2020 High accuracy correspondence field estimation via MST based patch matching
Feihu Zhang, Shibiao Xu, Xiaopeng Zhang 0001
Multim. Tools Appl.2
2020 Unsupervised Multi-View Constrained Convolutional Network for Accurate Depth Estimation
abstract
Accurate depth estimation from images is a fundamental problem in computer vision. In this paper, we propose an unsupervised learning based method to predict high-quality depth map from multiple images. A novel multi-view constrained DenseDepthNet is designed for this task. Our DenseDepthNet can effectively leverage both the low-level and high-level features of input images and generate appealing results, especially with sharp details. We employ the public datasets KITTI and Cityscapes for training in an end-to-end unsupervised fashion. A novel depth consistency loss based on multi-view geometry constraint is also applied to the corresponding points across pairwise images, which helps to improve the quality of predicted depth maps significantly. We conduct comprehensive evaluations on our DenseDepthNet and our depth consistency loss function. Experiments validate that our method outperforms the state-of-the-art unsupervised methods and produce comparable results with supervised methods.
Shibiao Xu, Baoyuan Wu, Weiliang Meng, Xiaopeng Zhang 0001
IEEE Trans. Image Process.2
2020 Realistic Procedural Plant Modeling from Multiple View Images
abstract
In this paper, we describe a novel procedural modeling technique for generating realistic plant models from multi-view photographs. The realism is enhanced via visual and spatial information acquired from images. In contrast to previous approaches that heavily rely on user interaction to segment plants or recover branches in images, our method automatically estimates an accurate depth map of each image and extracts a 3D dense point cloud by exploiting an efficient stereophotogrammetry approach. Taking this point cloud as a soft constraint, we fit a parametric plant representation to simulate the plant growth progress. In this way, we are able to synthesize parametric plant models from real data provided by photos and 3D point clouds. We demonstrate the robustness of the proposed approach by modeling various plants with complex branching structures and significant self-occlusions. We also demonstrate that the proposed framework can be used to reconstruct ground-covering plants, such as bushes and shrubs which have been given little attention in the literature. The effectiveness of our approach is validated by visually and quantitatively comparing with the state-of-the-art approaches.
Jianwei Guo 0003, Shibiao Xu, Dong-Ming Yan 0001, Zhanglin Cheng, Marc Jaeger 0002, Xiaopeng Zhang 0001
IEEE Trans. Vis. Comput. Graph.2
2019 GSLAM: A General SLAM Framework and Benchmark
abstract
SLAM technology has recently seen many successes and attracted the attention of high-technological companies. However, how to unify the interface of existing or emerging algorithms, and effectively perform benchmark about the speed, robustness and portability are still problems. In this paper, we propose a novel SLAM platform named GSLAM, which not only provides evaluation functionality, but also supplies useful toolkit for researchers to quickly develop their SLAM systems. Our core contribution is an universal, cross-platform and full open-source SLAM interface for both research and commercial usage, which is aimed to handle interactions with input dataset, SLAM implementation, visualization and applications in an unified framework. Through this platform, users can implement their own functions for better performance with plugin form and further boost the application to practical usage of the SLAM.
Shibiao Xu, Shuhui Bu, Hongkai Jiang, Pengcheng Han
ICCV2
2019 Nonlinear Asymmetric Multi-Valued Hashing
abstract
Most existing hashing methods resort to binary codes for large scale similarity search, owing to the high efficiency of computation and storage. However, binary codes lack enough capability in similarity preservation, resulting in less desirable performance. To address this issue, we propose Nonlinear Asymmetric Multi-Valued Hashing (NAMVH) supported by two distinct non-binary embeddings. Specifically, a real-valued embedding is used for representing the newly-coming query by an ideally nonlinear transformation. Besides, a multi-integer-embedding is employed for compressing the whole database, which is modeled by Binary Sparse Representation (BSR) with fixed sparsity. With these two non-binary embeddings, NAMVH preserves more precise similarities between data points and enables access to the incremental extension with database samples evolving dynamically. To perform meaningful asymmetric similarity computation for efficient semantic search, these embeddings are jointly learnt by preserving the pairwise label-based similarity. Technically, this results in a mixed integer programming problem, which is efficiently solved by a well-designed alternative optimization method. Extensive experiments on seven large scale datasets demonstrate that our approach not only outperforms the existing binary hashing methods in search accuracy, but also retains their query and storage efficiency.
Cheng Da, Gaofeng Meng, Shiming Xiang, Kun Ding 0001, Shibiao Xu, Qing Yang 0002, Chunhong Pan
IEEE Trans. Pattern Anal. Mach. Intell.5
2019 Learning graph structure via graph convolutional networks
Jianlong Chang, Gaofeng Meng, Shibiao Xu, Shiming Xiang, Chunhong Pan
Pattern Recognit.4
2018 Interactive stereo image segmentation via adaptive prior selection
Wei Ma 0008, Shibiao Xu, Xiaopeng Zhang 0001
Multim. Tools Appl.3
2018 Accurate blind deblurring using salientpatch-based prior for large-size images
Chengcheng Ma, Jiguang Zhang, Shibiao Xu, Weiliang Meng, Runping Xi, G. Hemanth Kumar, Xiaopeng Zhang 0001
Multim. Tools Appl.3
2018 Real-time pedestrian detection via hierarchical convolutional feature
Dongming Yang, Jiguang Zhang, Shibiao Xu, Shuiying Ge, G. Hemanth Kumar, Xiaopeng Zhang 0001
Multim. Tools Appl.3
2018 Automatic Building Rooftop Extraction From Aerial Images via Hierarchical RGB-D Priors
abstract
Accurate building rooftop extraction from high-resolution aerial images is of crucial importance in a wide range of applications. Owing to the varying appearance and large-scale range of scene objects, especially for building rooftops in different scales and heights, single-scale or individual prior-based extraction technique is insufficient in pursuing efficient, generic, and accurate extraction results. The trend toward integrating multiscale or several cue techniques appears to be the best way; thus, such integration is the focus of this paper. We first propose a novel salient rooftop detector integrating four correlative RGB-D priors (depth cue, uniqueness prior, shape prior, and transition surface prior) for improved rooftop extraction to address the preceding complex issues mentioned. Then, these correlative cues are computed from image layers created by our multilevel segmentation and further fused into the state-of-the-art high-order conditional random field (CRF) framework to locate the rooftop. Finally, an iterative optimization strategy is applied for high-quality solving, which can robustly handle varying appearance of building rooftops. Performance evaluations in the SZTAKI-INRIA benchmark data sets show that our method outperforms the traditional color-based algorithm and the original high-order CRF algorithm and its variants. The proposed algorithm is also evaluated and found to produce consistently satisfactory results for various large-scale, real-world data sets.
Shibiao Xu, Xingjia Pan, Er Li, Baoyuan Wu, Shuhui Bu, Weiming Dong, Shiming Xiang, Xiaopeng Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2017 AMVH: Asymmetric Multi-Valued hashing
abstract
Most existing hashing methods resort to binary codes for similarity search, owing to the high efficiency of computation and storage. However, binary codes lack enough capability in similarity preservation, resulting in less desirable performance. To address this issue, we propose an asymmetric multi-valued hashing method supported by two different non-binary embeddings. (1) A real-valued embedding is used for representing the newly-coming query. (2) A multi-integer-embedding is employed for compressing the whole database, which is modeled by binary sparse representation with fixed sparsity. With these two non-binary embeddings, the similarities between data points can be preserved precisely. To perform meaningful asymmetric similarity computation for efficient semantic search, these embeddings are jointly learnt by preserving the label-based similarity. Technically, this results in a mixed integer programming problem, which is efficiently solved by alternative optimization. Extensive experiments on three multilabel datasets demonstrate that our approach not only outperforms the existing binary hashing methods in search accuracy, but also retains their query and storage efficiency.
Cheng Da, Shibiao Xu, Kun Ding 0001, Gaofeng Meng, Shiming Xiang, Chunhong Pan
CVPR2
2017 Automatic Road Detection and Centerline Extraction via Cascaded End-to-End Convolutional Neural Network
abstract
Accurate road detection and centerline extraction from very high resolution (VHR) remote sensing imagery are of central importance in a wide range of applications. Due to the complex backgrounds and occlusions of trees and cars, most road detection methods bring in the heterogeneous segments; besides for the centerline extraction task, most current approaches fail to extract a wonderful centerline network that appears smooth, complete, as well as single-pixel width. To address the above-mentioned complex issues, we propose a novel deep model, i.e., a cascaded end-to-end convolutional neural network (CasNet), to simultaneously cope with the road detection and centerline extraction tasks. Specifically, CasNet consists of two networks. One aims at the road detection task, whose strong representation ability is well able to tackle the complex backgrounds and occlusions of trees and cars. The other is cascaded to the former one, making full use of the feature maps produced formerly, to obtain the good centerline extraction. Finally, a thinning algorithm is proposed to obtain smooth, complete, and single-pixel width road centerline network. Extensive experiments demonstrate that CasNet outperforms the state-of-the-art methods greatly in learning quality and learning speed. That is, CasNet exceeds the comparing methods by a large margin in quantitative performance, and it is nearly 25 times faster than the comparing methods. Moreover, as another contribution, a large and challenging road centerline data set for the VHR remote sensing image will be publicly available for further studies.
Ying Wang 0008, Shibiao Xu, Hongzhen Wang, Shiming Xiang, Chunhong Pan
IEEE Trans. Geosci. Remote. Sens.3
2016 Interactive Stereo Image Segmentation With RGB-D Hybrid Constraints
abstract
This letter presents an approach to extracting a target object interactively from a given pair of stereo images. First, a user marks a few parts of the object and background in either of the two views with strokes. The marked pixels are used to generate the prior models of the foreground and background. Second, a graph is constructed with constraints formulated by the priors of foreground/background, similarities between intraview neighbor pixels and correspondences between interview pixels. Third, two segments of the foreground are extracted from the two views by optimization of the graph via graph cut. Traditional methods generally define the priors and neighbor similarities in RGB space. Differently, the proposed method integrates disparity distributions of foreground/background to enrich the priors and defines the similarity metric between neighbor pixels in RGB-D space. The proposed method that utilizes RGB-D hybrid constraints generates stereo segments with accuracies higher than those obtained by state-of-the-art methods.
Wei Ma 0008, Luwei Yang, Shibiao Xu, Xiaopeng Zhang 0001
IEEE Signal Process. Lett.4
2015 Robust Rooftop Extraction From Visible Band Images Using Higher Order CRF
abstract
In this paper, we propose a robust framework for building extraction in visible band images. We first get an initial classification of the pixels based on an unsupervised presegmentation. Then, we develop a novel conditional random field (CRF) formulation to achieve accurate rooftops extraction, which incorporates pixel-level information and segment-level information for the identification of rooftops. Comparing with the commonly used CRF model, a higher order potential defined on segment is added in our model, by exploiting region consistency and shape feature at segment level. Our experiments show that the proposed higher order CRF model outperforms the state-of-the-art methods both at pixel and object levels on rooftops with complex structures and sizes in challenging environments.
Er Li, John Femiani 0001, Shibiao Xu, Xiaopeng Zhang 0001, Peter Wonka
IEEE Trans. Geosci. Remote. Sens.3
2015 PM-PM: PatchMatch With Potts Model for Object Segmentation and Stereo Matching
abstract
This paper presents a unified variational formulation for joint object segmentation and stereo matching, which takes both accuracy and efficiency into account. In our approach, depth-map consists of compact objects, each object is represented through three different aspects: 1) the perimeter in image space; 2) the slanted object depth plane; and 3) the planar bias, which is to add an additional level of detail on top of each object plane in order to model depth variations within an object. Compared with traditional high quality solving methods in low level, we use a convex formulation of the multilabel Potts Model with PatchMatch stereo techniques to generate depth-map at each image in object level and show that accurate multiple view reconstruction can be achieved with our formulation by means of induced homography without discretization or staircasing artifacts. Our model is formulated as an energy minimization that is optimized via a fast primal-dual algorithm, which can handle several hundred object depth segments efficiently. Performance evaluations in the Middlebury benchmark data sets show that our method outperforms the traditional integer-valued disparity strategy as well as the original PatchMatch algorithm and its variants in subpixel accurate disparity estimation. The proposed algorithm is also evaluated and shown to produce consistently good results for various real-world data sets (KITTI benchmark data sets and multiview benchmark data sets).
Shibiao Xu, Feihu Zhang, Xiaofei He 0001, Xukun Shen, Xiaopeng Zhang 0001
IEEE Trans. Image Process.1
2013 Statistical learning based facial animation
abstract
To synthesize real-time and realistic facial animation, we present an effective algorithm which combines image- and geometry-based methods for facial animation simulation. Considering the numerous motion units in the expression coding system, we present a novel simplified motion unit based on the basic facial expression, and construct the corresponding basic action for a head model. As image features are difficult to obtain using the performance driven method, we develop an automatic image feature recognition method based on statistical learning, and an expression image semi-automatic labeling method with rotation invariant face detection, which can improve the accuracy and efficiency of expression feature identification and training. After facial animation redirection, each basic action weight needs to be computed and mapped automatically. We apply the blend shape method to construct and train the corresponding expression database according to each basic action, and adopt the least squares method to compute the corresponding control parameters for facial animation. Moreover, there is a pre-integration of diffuse light distribution and specular light distribution based on the physical method, to improve the plausibility and efficiency of facial rendering. Our work provides a simplification of the facial motion unit, an optimization of the statistical training process and recognition process for facial animation, solves the expression parameters, and simulates the subsurface scattering effect in real time. Experimental results indicate that our method is effective and efficient, and suitable for computer animation and interactive applications.
Shibiao Xu, Guanghui Ma, Weiliang Meng, Xiaopeng Zhang 0001
J. Zhejiang Univ. Sci. C1
2012 Real-time ink simulation using a grid-particle method
Shibiao Xu, Xing Mei, Weiming Dong, Xiaopeng Zhang 0001
Comput. Graph.1