Rongtao Xu

dblp:93/4025 · DBLP profile ↗
← Back
83ranked-venue papers
20as first author
66since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 32 · 7 first-author · 32 since 2021Artificial intelligence and machine learning · 30 · 7 first-author · 30 since 2021Computer networks · 11 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal Interaction
abstract
Dense video captioning jointly localizes and captions salient events in untrimmed videos. Recent methods primarily focus on leveraging additional prior knowledge and advanced multi-task architectures to achieve competitive performance. However, these pipelines rely on implicit modeling that uses frame-level or fragmented video features, failing to capture the temporal coherence across event sequences and comprehensive semantics within visual contexts. To address this, we propose an explicit temporal-semantic modeling framework called Context-Aware Cross-Modal Interaction (CACMI), which leverages both latent temporal characteristics within videos and linguistic semantics from text corpus. Specifically, our model consists of two core components: Cross-modal Frame Aggregation aggregates relevant frames to extract temporally coherent, event-aligned textual features through cross-modal retrieval; and Context-aware Feature Enhancement utilizes query-guided attention to integrate visual dynamics with pseudo-event semantics. Extensive experiments on the ActivityNet Captions and YouCook2 datasets demonstrate that CACMI achieves the state-of-the-art performance on dense video captioning task.
Mingda Jia, Weiliang Meng, Zenghuang Fu, Ju Xin, Rongtao Xu, Jiguang Zhang, Xiaopeng Zhang 0001
AAAI8
2026 MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation
abstract
Multi-subject video generation aims to synthesize videos from textual prompts and multiple reference images, ensuring that each subject preserves natural scale and visual fidelity. However, current methods face two challenges: scale inconsistency, where variations in subject size lead to unnatural generation, and permutation sensitivity, where the order of reference inputs causes subject distortion. In this paper, we propose MoFu, a unified framework that tackles both challenges. For scale inconsistency, we introduce Scale-Aware Modulation (SMO), an LLM-guided module that extracts implicit scale cues from the prompt and modulates features to ensure consistent subject sizes. To address permutation sensitivity, we present a simple yet effective Fourier Fusion strategy that processes the frequency information of reference features via the Fast Fourier Transform to produce a unified representation. Besides, we design a Scale-Permutation Stability Loss to jointly encourage scale-consistent and permutation-invariant generation. To further evaluate these challenges, we establish a dedicated benchmark with controlled variations in subject scale and reference permutation. Extensive experiments demonstrate that MoFu significantly outperforms existing methods in preserving natural scale, subject fidelity, and overall visual quality.
Run Ling, Ke Cao 0001, Ao Ma 0005, Runze He, Changwei Wang 0001, Rongtao Xu, Yihua Shao, Zhanjie Zhang, Guibing Guo, Jingjing Lv, Junjie Shen 0008, Ching Law, Xingwei Wang 0001
AAAI8
2026 Online knowledge distillation optimization based on Multi-Student model Multi-Task collaborative learning
Shibiao Xu, Shanshan Mo, Changwei Wang 0001, Hetong Wang, Rongtao Xu, Li Guo 0004
Knowl. Based Syst.7
2026 EvolveNav: Empowering LLM-Based Vision-Language Navigation via Self-Improving Embodied Reasoning
abstract
Recent studies have revealed the potential of training open-source Large Language Models (LLMs) to unleash LLMs' reasoning ability for enhancing vision-language navigation (VLN) performance, and simultaneously mitigate the domain gap between LLMs' training corpus and the VLN task. However, these approaches predominantly adopt straightforward input-output mapping paradigms, causing the mapping learning difficult and the navigational decisions unexplainable. Chain-of-Thought (CoT) training is a promising way to improve both navigational decision accuracy and interpretability, while the complexity of the navigation task makes the perfect CoT labels unavailable and may lead to overfitting through pure CoT supervised fine-tuning. To address these issues, we propose EvolveNav, a novel sElf-improving embodied reasoning paradigm that realizes adaptable and generalizable navigational reasoning for boosting LLM-based vision-language Navigation. Specifically, EvolveNav involves a two-stage training process: (1) Formalized CoT Supervised Fine-Tuning, where we train the model with curated formalized CoT labels to first activate the model's navigational reasoning capabilities, and simultaneously increase the reasoning speed; (2) Self-Reflective Post-Training, where the model is iteratively trained with its own reasoning outputs as self-enriched CoT labels to enhance the supervision diversity. A self-reflective auxiliary task is also designed to encourage the model to learn correct reasoning patterns by contrasting with wrong ones. Experimental results under both task-specific and cross-task training paradigms demonstrate the consistent superiority of EvolveNav over previous LLM-based VLN approaches on various popular benchmarks, including R2R, REVERIE, CVDN, and SOON. EvolveNav open avenues for exploring effective self-improving reasoning paradigms, enabling building agents capable of self-evolving for promoting LLM-based embodied AI research.
Bingqian Lin, Yunshuang Nie, Khun Loun Zai, Ziming Wei 0001, Mingfei Han 0002, Rongtao Xu, Minzhe Niu, Jianhua Han, Hanwang Zhang, Liang Lin 0004, Bokui Chen, Cewu Lu, Xiaodan Liang
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 SdaPS*: A Novel Source-Free Domain Adaption Method for Point Cloud Primitive Segmentation
Shaohu Wang, Yuchuang Tong, Rongtao Xu, Zhengtao Zhang
IEEE Trans. Ind. Informatics3
2026 TeDri:Teacher-Driven Region Knowledge Distillation
abstract
Knowledge distillation as a practical tool to enhance the performance of small-capacity student networks on downstream tasks comes at the cost of a lengthy distillation process due to the online inference of teacher networks, especially when there is a large capacity gap between them. Therefore, in this paper, we propose a fast distillation framework called TeDri based on region images by offline saving relevant regional information and its teacher guidance. Specifically, first, to alleviate the lack of diversity caused by the fixed augmentation path in region images, we propose Teacher-driven MixUp strategies with mild intensity and advocate binding the mixing factor$\lambda$with teacher guidance confidence, where more confident category representations dominate the MixUp process. Furthermore, recognizing the need to evaluate these randomly cropped regions, and we propose region contrastive learning, encourage the student network to mimic the region partitioning behavior of the teacher, promoting a comprehensive understanding of global semantic content from multiple local perspectives. Finally, we introduce region mutual learning, employing spatial constraints among regions to require the student network towards consistent content interpretation across localized regions. Experiments on CIFAR-100 and ImageNet-1 K validate the effectiveness of the proposed TeDri, achieving competitive performance while significantly reducing training time.
Changwei Wang 0001, Rongtao Xu, Xingtian Pei, Shibiao Xu, Wenbo Xu 0003, Li Guo 0004
IEEE Trans. Knowl. Data Eng.3
2026 Low Harmonic MSK/GMSK Backscatter Based on Active Transistor Load
Yibing Yang, Ming Liu 0010, Gongpu Wang, Rongtao Xu, Wei Gong 0001, Bo Ai 0001
IEEE Trans. Mob. Comput.4
2026 Adaptive in Adapter: Boosting Open-Vocabulary Semantic Segmentation With Adaptive Dropout Adapter
abstract
Open-vocabulary semantic segmentation is a challenging multimedia task that requires segmentation and recognition of unseen word classes during the testing phase. Recent works bridge the gap between closed and open-vocabulary recognition by introducing large-scale visual language models such as CLIP with cross-modal alignment capabilities. To preserve multimodal alignment capabilities, it is common to freeze the parameters of the CLIP and then add additional learnable components such as adapters to expand to downstream tasks. However, for the open-vocabulary semantic segmentation task, the plain adapter suffers from overfitting the closed-vocabulary classes and impairs performance on the open-vocabulary unseen classes. In addition, since CLIP is trained to perform image-level alignment can cause the network to over-focus on partially discriminative regions, resulting in incomplete segmentation masks. To alleviate the above problems, we introduce adaptive dropout adapters to release theAdaptiveInAdapter (i.e.AIA) from the following two aspects:i)A Generalization Feature Selection Adapter (GFSA) is proposed to improve the generalization of network over unseen classes.ii)A Discriminative Region Mask Adapter (DRMA) is proposed for retrofitting CLIP backbone, has provided region free biased features for segmentation mask generation. Meanwhile, our proposed AIA achieves the current state-of-the-art performance on several open-vocabulary semantic segmentation benchmarks. Code is available athttps://github.com/clearxu/AIA.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Jiguang Zhang, Xiaoqiang Teng, Weiliang Meng, Xiaopeng Zhang 0001
IEEE Trans. Multim.3
2025 Focus on Local: Finding Reliable Discriminative Regions for Visual Place Recognition
abstract
Visual Place Recognition (VPR) is aimed at predicting the location of a query image by referencing a database of geotagged images. For VPR task, often fewer discriminative local regions in an image produce important effects while mundane background regions do not contribute or even cause perceptual aliasing because of easy overlap. However, existing methods lack precisely modeling and full exploitation of these discriminative regions. In addition, the lack of pixel-level correspondence supervision in the VPR dataset hinders further improvement of the local feature matching capability in the re-ranking stage. In this paper, we propose the Focus on Local (FoL) approach to stimulate the performance of image retrieval and re-ranking in VPR simultaneously by mining and exploiting reliable discriminative local regions in images and introducing pseudo-correlation supervision. First, we design two losses, Extraction-Aggregation Spatial Alignment Loss (SAL) and Foreground-Background Contrast Enhancement Loss (CEL), to explicitly model reliable discriminative local regions and use them to guide the generation of global representations and efficient re-ranking. Second, we introduce a weakly-supervised local feature training strategy based on pseudo-correspondences obtained from aggregating global features to alleviate the lack of local correspondences ground truth for the VPR task. Third, we suggest an efficient re-ranking pipeline that is efficiently and precisely based on discriminative region guidance. Finally, experimental results show that our FoL achieves the state-of-the-art on multiple VPR benchmarks in both image retrieval and re-ranking stages and also significantly outperforms existing two-stage VPR methods in terms of computational efficiency.
Changwei Wang 0001, Shunpeng Chen, Rongtao Xu, Jiguang Zhang, Haoran Yang 0003, Yu Zhang 0133, Kexue Fu 0001, Shide Du, Zhiwei Xu 0005, Longxiang Gao, Li Guo 0004, Shibiao Xu
AAAI4
2025 MindTuner: Cross-Subject Visual Decoding with Visual Fingerprint and Semantic Correction
abstract
Decoding natural visual scenes from brain activity has flourished, with extensive research in single-subject tasks and, however, less in cross-subject tasks. Reconstructing high-quality images in cross-subject tasks is a challenging problem due to profound individual differences between subjects and the scarcity of data annotation. In this work, we proposed MindTuner for cross-subject visual decoding, which achieves high-quality and rich semantic reconstructions using only 1 hour of fMRI training data benefiting from the phenomena of visual fingerprint in the human visual system and a novel fMRI-to-text alignment paradigm. Firstly, we pre-train a multi-subject model among 7 subjects and fine-tune it with scarce data on new subjects, where LoRAs with Skip-LoRAs are utilized to learn the visual fingerprint. Then, we take the image modality as the intermediate pivot modality to achieve fMRI-to-text alignment, which achieves impressive fMRI-to-text retrieval performance and corrects fMRI-to-image reconstruction with fine-tuned semantics. The results of both qualitative and quantitative analyses demonstrate that MindTuner surpasses state-of-the-art cross-subject visual decoding models on the Natural Scenes Dataset (NSD), whether using training data of 1 hour or 40 hours.
Zixuan Gong, Qi Zhang 0020, Guangyin Bao, Rongtao Xu, Liang Hu 0004, Duoqian Miao 0001
AAAI5
2025 PanoDiT: Panoramic Videos Generation with Diffusion Transformer
abstract
As immersive experiences become increasingly popular, panoramic video has garnered significant attention in both research and applications. The high cost associated with capturing panoramic video underscores the need for efficient prompt-based generation methods. Although recent text-to-video (T2V) diffusion techniques have shown potential in standard video generation, they face challenges when applied to panoramic videos due to substantial differences in content and motion patterns. In this paper, we propose PanoDiT, a framework that utilizes the Diffusion Transformer (DiT) architecture to generate panoramic videos from text descriptions. Unlike traditional methods that rely on UNet-based denoising, our method leverages a transformer architecture for denoising, incorporating both temporal and global attention mechanisms. This ensures coherent frame generation and smooth motion transitions, offering distinct advantages in long-horizon generation tasks. To further enhance motion and consistency in the generated videos, we introduce DTM-LoRA and two panoramic-specific losses. Compared to previous methods, our PanoDiT achieves state-of-the-art performance across various evaluation metrics and user study, with code is available in the supplementary material.
Muyang Zhang, Yuzhi Chen, Rongtao Xu, Changwei Wang 0001, Weiliang Meng, Jianwei Guo 0003, Xiaopeng Zhang 0001
AAAI3
2025 AASD: Accelerate Inference by Aligning Speculative Decoding in Multimodal Large Language Models
abstract
Multimodal Large Language Models (MLLMs) have achieved notable success in visual instruction tuning, yet their inference is time-consuming due to the auto-regressive decoding of Large Language Model (LLM) backbone. Traditional methods for accelerating inference, including model compression and migration from language model acceleration, often compromise output quality or face challenges in effectively integrating multimodal features. To address these issues, we propose AASD, a novel framework for Accelerating inference with refined KV Cache and Aligning speculative decoding in MLLMs. Our approach leverages the target model’s cached KeyValue (KV) pairs to extract vital information for generating draft tokens, enabling efficient speculative decoding. To reduce the computational burden associated with long multimodal token sequences, we introduce a KV Projector to compress the KV Cache while maintaining representational fidelity. Additionally, we design a Target-Draft Attention mechanism that optimizes the alignment between the draft model and the target model, achieving the benefits of real inference scenarios with minimal computational overhead. Extensive experiments on mainstream MLLMs demonstrate that our method achieves up to a $2 \times$ inference speedup without sacrificing accuracy. This study not only provides an effective and lightweight solution for accelerating MLLM inference but also introduces a novel alignment strategy for speculative decoding in multimodal contexts, laying a strong foundation for future research in efficient MLLMs. Code is availiable at https://github.com/transcend-0/ASD
Muyang Zhang, Weiguang Pang, Yuzhi Chen, Rongtao Xu, Kexue Fu 0001, Changwei Wang 0001, Longxiang Gao
DAC6
2025 CAE-DFKD: Bridging the Transferability Gap in Data-Free Knowledge Distillation
abstract
Data-Free Knowledge Distillation (DFKD) enables the knowledge transfer from the given pre-trained teacher network to the target student model without access to the real training data. Existing DFKD methods focus primarily on improving image recognition performance on associated datasets, often neglecting the crucial aspect of the transferability of learned representations. In this paper, we propose Category-Aware Embedding Data-Free Knowledge Distillation (CAE-DFKD), which addresses at the embedding level the limitations of previous rely on image-level methods to improve model generalization but fail when directly applied to DFKD. The superiority and flexibility of CAE-DFKD are extensively evaluated, including: i.) Significant efficiency advantages resulting from altering the generator training paradigm; ii.) Competitive performance with existing DFKD state-of-the-art methods on image recognition tasks; iii.) Remarkable transferability of data-free learned representations demonstrated in downstream tasks.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Yu Zhang 0133, Jie Zhou 0001, Li Guo 0004
DAC3
2025 Structured Preference Optimization for Vision-Language Long-Horizon Task Planning
abstract
Xiwen Liang, Min Lin, Weiqi Ruan, Rongtao Xu, Yuecheng Liu, Jiaqi Chen, Bingqian Lin, Yuzheng Zhuang, Xiaodan Liang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Xiwen Liang, Weiqi Ruan, Rongtao Xu, Yuecheng Liu, Bingqian Lin, Yuzheng Zhuang, Xiaodan Liang
EMNLP4
2025 Channel Estimation and Data Detection in Backscatter Communications with Phase Noise
Ziqi Cui, Gongpu Wang, Rongtao Xu, Ming Zeng 0002, Chintha Tellambura
GLOBECOM3
2025 $A_{0}$: An Affordance-Aware Hierarchical Model for General Robotic Manipulation
Rongtao Xu, Youpeng Wen, Haoting Yang, Jianzheng Huang, Zhe Li 0008, Kaidong Zhang, Liqiong Wang, Yuxuan Kuang, Meng Cao 0002, Feng Zheng 0001, Xiaodan Liang
ICCV1
2025 RoBridge: A Hierarchical Architecture Bridging Cognition and Execution for General Robotic Manipulation
abstract
Operating robots in open-ended scenarios with diverse tasks is a crucial research and application direction in robotics. While recent progress in natural language processing and large multimodal models has enhanced robots' ability to understand complex instructions, robot manipulation still faces the procedural skill dilemma and the declarative skill dilemma in open environments. Existing methods often compromise cognitive and executive capabilities. To address these challenges, in this paper, we propose RoBridge, a hierarchical intelligent architecture for general robotic manipulation. It consists of a high-level cognitive planner (HCP) based on a large-scale pre-trained vision-language model (VLM), an invariant operable representation (IOR) serving as a symbolic bridge, and a generalist embodied agent (GEA). RoBridge maintains the declarative skill of VLM and unleashes the procedural skill of reinforcement learning, effectively bridging the gap between cognition and execution. RoBridge demonstrates significant performance improvements over existing baselines, achieving a 75% success rate on new tasks and an 83% average success rate in sim-to-real generalization using only five real-world data samples per task. This work represents a significant step towards integrating cognitive reasoning with physical execution in robotic systems, offering a new paradigm for general robotic manipulation.
Kaidong Zhang, Rongtao Xu, Pengzhen Ren, Junfan Lin, Hefeng Wu, Xiaodan Liang
ICCV2
2025 MVPS: Multi-View Adaptive Prompt Synergy for Zero-shot Anomaly Detection
abstract
Zero-shot anomaly detection (ZSAD) in industrial domains faces significant challenges due to the diverse manifestations of anomalies across scales and semantic levels. Existing methods, relying on single prompt spaces, struggle to generalize across these variations. So they exhibit limited generalization across scales and semantic levels. We propose Multi-View Adaptive Prompting Synergy (MVPS), a novel framework that establishes multiple scale-aware prompt spaces to enhance anomaly detection generalization. MVPS comprises three synergistic components: Hierarchical Multi-modal Prompt Tuning (HMPT) for generating scale-aware prompts, Dual-stream Prompt Tuning Orchestration (DPTO) for achieving robust cross-modal feature alignment, and Multi-View Prompt Composition Learning (MVPCL) for effective scale feature perception. This approach enables comprehensive capture anomaly feature across multiple semantic levels and scales, overcoming limitations of single-space representations and scale-insensitive feature alignment. Extensive experiments on seven benchmark datasets demonstrate that MVPS achieves state-of-the-art performance, exhibiting superior generalization capability across diverse anomaly categories and industrial domain. The code is available at https://github.com/MLY-0546/mvps.
Longzhao Huang, Changwei Wang 0001, Rongtao Xu, Shibiao Xu
ICME4
2025 Complementary Information Guided Occupancy Prediction via Multi-Level Representation Fusion
abstract
Camera-based occupancy prediction is a main-stream approach for 3D perception in autonomous driving, aiming to infer complete 3D scene geometry and semantics from 2D images. Almost existing methods focus on improving performance through structural modifications, such as lightweight backbones and complex cascaded frameworks, with good yet limited performance. Few studies explore from the perspective of representation fusion, leaving the rich diversity of features in 2D images underutilized. Motivated by this, we propose CIGOcc, a two-stage occupancy prediction framework based on multi-level representation fusion. CIGOcc extracts segmentation, graphics, and depth features from an input image and introduces a deformable multi-level fusion mechanism to fuse these three multi-level features. Additionally, CIGOcc incorporates knowledge distilled from SAM to further enhance prediction accuracy. Without increasing training costs, CIGOcc achieves state-of-the-art performance on the SemanticKITTI benchmark. The code is provided in the supplementary material and will be released project page.
Rongtao Xu, Jinzhou Lin 0001, Jialei Zhou, Jiahua Dong 0001, Changwei Wang 0001, Ruisheng Wang 0001, Li Guo 0004, Shibiao Xu, Xiaodan Liang
ICRA1
2025 3D-MoRe: Unified Modal-Contextual Reasoning for Embodied Question Answering
abstract
With the growing need for diverse and scalable data in indoor scene tasks, such as question answering and dense captioning, we propose 3D-MoRe, a novel paradigm designed to generate large-scale 3D-language datasets by lever-aging the strengths of foundational models. The framework integrates key components, including multi-modal embedding, cross-modal interaction, and a language model decoder, to process natural language instructions and 3D scene data. This approach facilitates enhanced reasoning and response generation in complex 3D environments. Using the ScanNet 3D scene dataset, along with text annotations from ScanQA and ScanRefer, 3D-MoRe generates 62,000 question-answer (QA) pairs and 73,000 object descriptions across 1,513 scenes. We also employ various data augmentation techniques and implement semantic filtering to ensure high-quality data. Experiments on ScanQA demonstrate that 3D-MoRe significantly outperforms state-of-the-art baselines, with the CIDEr score improving by 2.15%. Similarly, on ScanRefer, our approach achieves a notable increase in [email protected] by 1.84%, highlighting its effectiveness in both tasks. Our code and generated datasets will be publicly released to benefit the community, and both can be accessed on the https://3D-MoRe.github.io.
Rongtao Xu, Mingming Yu, Dong An 0002, Shunpeng Chen, Changwei Wang 0001, Li Guo 0004, Xiaodan Liang, Shibiao Xu
IROS1
2025 PhyBlock: A Progressive Benchmark for Physical Understanding and Planning via 3D Block Assembly
abstract
While vision-language models (VLMs) have demonstrated promising capabilities in reasoning and planning for embodied agents, their ability to comprehend physical phenomena, particularly within structured 3D environments, remains severely limited. To close this gap, we introduce PhyBlock, a progressive benchmark designed to assess VLMs on physical understanding and planning through robotic 3D block assembly tasks. PhyBlock integrates a novel four-level cognitive hierarchy assembly task alongside targeted Visual Question Answering (VQA) samples, collectively aimed at evaluating progressive spatial reasoning and fundamental physical comprehension, including object properties, spatial relationships, and holistic scene understanding. PhyBlock includes 2600 block tasks (400 assembly tasks, 2200 VQA tasks) and evaluates models across three key dimensions: partial completion, failure diagnosis, and planning robustness. We benchmark 23 state-of-the-art VLMs, highlighting their strengths and limitations in physically grounded, multi-step planning. Our empirical findings indicate that the performance of VLMs exhibits pronounced limitations in high-level planning and reasoning capabilities, leading to a notable decline in performance for the growing complexity of the tasks.Error analysis reveals persistent difficulties in spatial orientation and dependency reasoning.We position PhyBlock as a unified testbed to advance embodied reasoning, bridging vision-language understanding and real-world physical problem-solving.
Jiajun Wen 0003, Rongtao Xu, Xiwen Liang, Bingqian Lin, Ziming Wei 0001, Haokun Lin, Mingfei Han 0002, Meng Cao 0002, Bokui Chen, Ivan Laptev, Xiaodan Liang
NeurIPS4
2025 Enhanced Fingerprint Localization with GAN-Augmented Datasets for Deep Learning
abstract
This paper proposes a two-stage fingerprint localization framework to address the challenges of excessive reliance on massive real reference points and high computational complexity. The proposed framework employs distinct neural networks on its two stages. First, a generative adversarial network (GAN)-based dataset augmentation is implemented to synthesize virtual fingerprints, expanding radio map coverage and reducing reliance on massive reference points. Second, a hierarchical localization model based on deep learning is introduced, using a convolutional neural network (CNN) for rapid and coarse-grained localization with reduced complexity, followed by a probabilistic deep neural network (DNN) to achieve refined and accurate positioning. Simulation results show that the proposed framework achieves a Top-1 accuracy of 81.2% and a Top-2 accuracy of 98%, outperforming benchmark algorithms. Furthermore, training with a dataset of 57% GAN-generated fingerprints reduces localization error by 30.9%, demonstrating the effectiveness of dataset augmentation strategy.
Junliang Lin, Rongtao Xu
VTC2025-Fall4
2025 Dual prototypes contrastive learning based semi-supervised segmentation method for intelligent medical applications
Tianai Yue, Rongtao Xu, Jingqian Wu, Wenjie Yang 0005, Shide Du, Changwei Wang 0001
Eng. Appl. Artif. Intell.2
2025 FDBPL: Faster distillation-based prompt learning for region-aware vision-language models adaptation
Changwei Wang 0001, Rongtao Xu, Longzhao Huang, Wenbo Xu 0003, Li Guo 0004, Shibiao Xu
Expert Syst. Appl.4
2025 C2Fi-NeRF: Coarse to fine inversion NeRF for 6D pose estimation
Jiguang Zhang, Zhaohui Zhang 0002, Xuxiang Feng, Shibiao Xu, Rongtao Xu, Changwei Wang 0001, Kexue Fu 0001, Jiaxi Sun, Weilong Ding 0001
Expert Syst. Appl.5
2025 Data-Efficient Learning-Based Iterative Optimization Method With Time-Varying Prediction Horizon for Multiagent Collaboration
abstract
Learning-based strategy can be well integrated with model-based optimal control to facilitate cooperative multiagent control through the Internet of Things (IoT). In this work, we propose a data-efficient learning-based iterative optimization method with time-varying prediction horizon (TV-LIO) for multiagent collaboration. Our method builds a multiagent optimization problem by introducing a time-domain guided terminal set and an approximated general cost. We collect the historical agent states at previous iterations as a dataset to reconstruct the general cost and the terminal set iteratively, forming closed-loop data-efficient learning. We consider the influence of the predictive time domain on the optimality and feasibility of the optimization problem and design a time-domain recursive updating mechanism to determine the optimal predictive horizon for each agent at the epoch. The continuous feasibility, stability, and recursive convergence of the proposed method are analyzed theoretically. Unlike the traditional optimization approaches that rely on a preplaned reference path, the proposed method integrates the trajectory planning and tracking control for multiple agents. After several iterations, the general cost of the optimization problem monotonically decreases and the optimal states are finally obtained. The proposed approach is validated and the results demonstrate that our approach can obtain the optimal-cost strategy and trajectories with optimizing time domains for the multiagent system.
Bowen Wang 0005, Xinle Gong, Yafei Wang 0001, Rongtao Xu, Hongcheng Huang
IEEE Internet Things J.4
2025 Constraint-Aware Zero-Shot Vision-Language Navigation in Continuous Environments
abstract
We address the task of Vision-Language Navigation in Continuous Environments (VLN-CE) under the zero-shot setting. Zero-shot VLN-CE is particularly challenging due to the absence of expert demonstrations for training and minimal environment structural prior to guide navigation. To confront these challenges, we propose a Constraint-Aware Navigator (CA-Nav), which reframes zero-shot VLN-CE as a sequential, constraint-aware sub-instruction completion process. CA-Nav continuously translates sub-instructions into navigation plans using two core modules: the Constraint-Aware Sub-instruction Manager (CSM) and the Constraint-Aware Value Mapper (CVM). CSM defines the completion criteria for decomposed sub-instructions as constraints and tracks navigation progress by switching sub-instructions in a constraint-aware manner. CVM, guided by CSM's constraints, generates a value map on the fly and refines it using superpixel clustering to improve navigation stability. CA-Nav achieves the state-of-the-art performance on two VLN-CE benchmarks, surpassing the previous best method by 12% and 13% in Success Rate on the validation unseen splits of R2R-CE and RxR-CE, respectively. Moreover, CA-Nav demonstrates its effectiveness in real-world robot deployments across various indoor scenes and instructions.
Dong An 0002, Yan Huang 0008, Rongtao Xu, Yifei Su, Yonggen Ling, Ian D. Reid 0001, Liang Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 A Model-Prediction-Based Hierarchical Personalized Federated Learning Framework With Distributed Resource Optimization
abstract
With the number of users growing rapidly, the federated learning (FL) algorithm gradually moves towards a multi-layer framework to ensure learning performance. In this article, considering the non-independent and identically distributed scenario, a new hierarchical federated meta-learning (HFML) framework is studied. The Hessian-free Model-Agnostic Meta-Learning is introduced into our model to personalize the local models of edge users (EUs), which is more computationally efficient than the traditional meta-learning. To alleviate the learning performance reduction due to the scarce available bandwidth resources, a multilayer perceptron model prediction scheme based on the attention mechanism is deployed at the side of edge nodes (ENs). To achieve the tradeoff between learning time and model accuracy, the semi-synchronous cloud aggregation mechanism based on the learning states and parameter freshness is proposed. The convergence analysis of the proposed HFML algorithm is also provided to prove that the upper bound of the loss decay exists. To solve the complex nonconvex optimization problem whose target is to maximize the learning efficiency of HFML, considering device selection and communication resource allocation, a decentralized algorithm based on Jacobi-Proximal ADMM (JP-ADMM) is proposed. Extensive simulations are performed to demonstrate the effectiveness of the proposed method. Particularly, compared with the traditional hierarchical federated learning algorithm, the proposed HFML achieves better learning performance while reducing the latency.
Rongtao Xu, Bo Ai 0001
IEEE Trans. Commun.3
2025 PolarBEVU: Multi-Camera 3D Object Detection in Polar Bird's-Eye View via Unprojection
abstract
3D object detection from a Bird’s Eye View (BEV) has emerged as a novel perception paradigm for autonomous driving scenarios. While most current 3D object detection methods still rely on the conventional Cartesian coordinates, they fail to align with the non-aligned coordinate system inherent in image geometry. The Polar coordinates, on the other hand, better fit with the geometric shape corresponding to the perception of cameras. However, transforming between coordinate systems introduces distortions in the perception information, resulting in issues such as “Weak Adaptability to Heatmap Distribution” and “Offset in the Center Point of the Bounding Box.” To address these challenges, this paper proposes a cutting-edge 3D object detection model named PolarBEVU, which leverages the bird’s-eye view under the Polar coordinates along with multi-camera unprojection. The model introduces an innovative “Deformable Uniform Heatmap Distribution” method that adjusts heatmap computations based on box shapes, generating high-quality heatmaps and effectively resolving the issue of “Weak Adaptability to Heatmap Distribution.” Moreover, the model incorporates the concept of “Dynamic High-risk Regression Region” to enhance the accuracy and robustness of the center point regression at the bounding box, thus mitigating the issue of “Offset in the Center Point of the Bounding Box.” In extensive experiments on the nuScenes dataset, PolarBEVU achieves impressive results with 49.9% mAP and 57.4% NDS on the test set, surpassing other comparative approaches and reaching the state-of-the-art (SOTA) performance among methods utilizing Polar coordinates. This clearly demonstrates the efficacy and superiority of PolarBEVU. In addition, the model is successfully deployed on Nvidia Jetson AGX Orin, showcasing real-time inference speeds of 31.42ms. These findings affirm PolarBEVU’s potential for practical applications. Code is available athttps://github.com/JLUrob/PolarBEVU.
Minghui Hou, Chuanhao Lyu, Gang Wang 0013, Baorui Ma, Rongtao Xu, Jue Hu, Xiaopeng Fan 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 Generalization Boosted Adapter for Open-Vocabulary Segmentation
abstract
Vision-language models (VLMs) have demonstrated remarkable open-vocabulary object recognition capabilities, motivating their adaptation for dense prediction tasks like segmentation. However, directly applying VLMs to such tasks remains challenging due to their lack of pixel-level granularity and the limited data available for fine-tuning, leading to overfitting and poor generalization. To address these limitations, we propose Generalization Boosted Adapter (GBA), a novel adapter strategy that enhances the generalization and robustness of VLMs for open-vocabulary segmentation. GBA comprises two core components: (1) a Style Diversification Adapter (SDA) that decouples features into amplitude and phase components, operating solely on the amplitude to enrich the feature space representation while preserving semantic consistency; and (2) a Correlation Constraint Adapter (CCA) that employs cross-attention to establish tighter semantic associations between text categories and target regions, suppressing irrelevant low-frequency “noise” information and avoiding erroneous associations. Through the synergistic effect of the shallow SDA and the deep CCA, GBA effectively alleviates overfitting issues and enhances the semantic relevance of feature representations. As a simple, efficient, and plug-and-play component, GBA can be flexibly integrated into various CLIP-based methods, demonstrating broad applicability and achieving state-of-the-art performance on multiple open-vocabulary segmentation benchmarks. Code are available athttps://github.com/clearxu/BGA.
Changwei Wang 0001, Xuxiang Feng, Rongtao Xu, Longzhao Huang, Li Guo 0004, Shibiao Xu
IEEE Trans. Circuits Syst. Video Technol.4
2025 DFMC: Feature-Driven Data-Free Knowledge Distillation
abstract
Data-Free Knowledge Distillation (DFKD) enables knowledge transfer from teacher networks without access to the real dataset. However, generator-based DFKD methods often suffer from insufficient diversity or low-confidence in synthetic images, negatively impacting student network performance. This paper introduces DFMC, a generative feature-driven framework to mitigate the inherent limitations of DFKD. We propose exploiting semantic description between generative feature domains to guide augmentation strategies, avoiding random abstract inputs caused by inconsistent semantic quality. Then, by applying noise to the generative features, we produce contrastive learning pairs indirectly, limiting the sampling range of the feature domain to encourage the student network to learn domain-invariant features. Finally, we guide the student network to deeply mimic the teacher’s layer-wise implicit classification behavior for the augmented synthetic images. Extensive experiments across various datasets and downstream tasks demonstrate the effectiveness of DFMC, achieving significant improvements while preventing student networks from overfitting to semantic ambiguous images.
Rongtao Xu, Changwei Wang 0001, Shunpeng Chen, Shibiao Xu, Guangyuan Xu, Li Guo 0004
IEEE Trans. Circuits Syst. Video Technol.2
2025 Segment Anything Model Is a Good Teacher for Local Feature Learning
abstract
Local feature detection and description play an important role in many computer vision tasks, which are designed to detect and describe keypoints in any scene and any downstream task. Data-driven local feature learning methods need to rely on pixel-level correspondence for training. However, a vast number of existing approaches ignored the semantic information on which humans rely to describe image pixels. In addition, it is not feasible to enhance generic scene keypoints detection and description simply by using traditional common semantic segmentation models because they can only recognize a limited number of coarse-grained object classes. In this paper, we propose SAMFeat to introduce SAM (segment anything model), a foundation model trained on 11 million images, as a teacher to guide local feature learning. SAMFeat learns additional semantic information brought by SAM and thus is inspired by higher performance even with limited training samples. To do so, first, we construct an auxiliary task of Attention-weighted Semantic Relation Distillation (ASRD), which adaptively distillates feature relations with category-agnostic semantic information learned by the SAM encoder into a local feature learning network, to improve local feature description using semantic discrimination. Second, we develop a technique called Weakly Supervised Contrastive Learning Based on Semantic Grouping (WSC), which utilizes semantic groupings derived from SAM as weakly supervised signals, to optimize the metric space of local descriptors. Third, we design an Edge Attention Guidance (EAG) to further improve the accuracy of local feature detection and description by prompting the network to pay more attention to the edge region guided by SAM. SAMFeat's performance on various tasks, such as image matching on HPatches, and long-term visual localization on Aachen Day-Night showcases its superiority over previous local features. The release code is available at https://github.com/vignywang/SAMFeat.
Jingqian Wu, Rongtao Xu, Zach Wood-Doughty, Changwei Wang 0001, Shibiao Xu, Edmund Y. Lam
IEEE Trans. Image Process.2
2025 Token Masking Transformer for Weakly Supervised Object Localization
abstract
Weakly supervised object localization (WSOL) is both a promising and challenging task that aims to achieve object localization exclusively through image category labels for supervision. Visual transformers have recently been applied to WSOL, demonstrating significant success through the exploitation of long-range feature dependencies in self-attention mechanisms. However, the transformer-based approach suffers from the same partial activation problem as the CNN-based approach due to the use of the classification task to train self-attention map, i.e., only a few discriminative regions are assigned high attention response and thus the localization map does not cover the whole object. To alleviate this problem, we propose a plug-and-play Token Masking Transformer (TMT) method to help transformer-based WSOL methods to obtain a more complete localization map by dynamic discriminative token masking. Specifically, a batch-wise discriminative token selection strategy is first introduced to flexibly determine the tokens to be masked in each image. Then, we design a token masking transformer block to perform token masking and inspire the network to mine more object-related tokens. Besides, we also design an intermediate token activation loss to further improve the performance of TMT by imposing constraints on intermediate tokens. Extensive experiments demonstrate that our TMT can substantially improve the performance of existing transformer-based methods without increasing the computational cost, and achieves state-of-the-art performance on two mainstream benchmarks.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Man Zhang 0005, Xiaopeng Zhang 0001
IEEE Trans. Multim.3
2025 OV-BIS: Open-Vocabulary Boundary Guide Zero-Shot 3D Instance Segmentation
abstract
Open vocabulary 3D instance segmentation aims to align 3D instance segmentation results with natural language text, thereby achieving semantic prediction without relying on predefined class labels for specific scenes, which has been widely used in the field of multimedia. Current open vocabulary 3D instance segmentation methods mainly rely on 2D masks provided by various 2D segmentation foundation models. However, in complex scenes, the calculation of 2D masks often struggles to balance over-segmentation of large objects and under-segmentation of small objects. In this paper, we introduce OV-BIS, a novel zero-shot open vocabulary 3D instance segmentation method that leverages instance boundary information to improve 3D semantic segmentation performance. The key insight of our method is that the edge map as 3D boundary projection is suitable for multi-scale tasks and capable of compensating for the weakness of 2D masks in multi-scale adaptability for complex scenes. Our method aggregates multiview edge maps and 2D masks, iteratively guiding the merging of over-segmented point clouds with regions growing to cluster 3D primitives into distinct 3D instances. By projecting 3D instances onto images and using CLIP to calculate semantic features from multiple perspectives with an outliers filter, 3D semantic instance segmentation has been achieved. Experiments on multiple datasets demonstrate the superiority of our method.
Tinghao Yi, Shaohu Wang, Zhengtao Zhang, Changwei Wang 0001, Dong-Ming Yan 0001, Rongtao Xu, Enhong Chen
IEEE Trans. Multim.6
2025 SRIF: Data-Free Knowledge Distillation via Stable Regulation and Input Filtering
abstract
Data-free knowledge distillation (DFKD) enables knowledge transfer from a pre-trained teacher to a student network without accessing the real dataset. However, generator-based DFKD methods struggle to ensure that the synthetic images accurately reflect the real dataset distribution. The update of the generator network relies heavily on teacher category guidance, but varying teacher prediction accuracy across categories leads to inconsistent synthetic image quality. Such variations introduce a distribution shift between synthetic and real datasets, negatively impacting student network performance during knowledge distillation. To address this challenge, we propose the SRIF, comprising two components: Student-Driven Flexible Filtering (SDFF) and Re-weighting for Independent Regularization (RIR). SDFF filters out synthetic images affected by the category distribution shift during data generation, producing a more reliable dataset. RIR, applied during distillation, encourages the student to learn stable causal relationships through sample reweighting. Both components flexibly integrate into existing DFKD frameworks, improving performance while reducing training costs.
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Jie Zhou 0001, Longxiang Gao, Wenbo Xu 0003, Li Guo 0004
IEEE Trans. Multim.2
2025 PDFT: parameter-diminish fine-tuning for transformer-based models
Muyang Zhang, Weiliang Meng, Mingda Jia, Jiaming Gu, Yihua Shao, Changwei Wang 0001, Rongtao Xu, Xiaopeng Zhang 0001
Vis. Comput.7
2024 Spectral Prompt Tuning: Unveiling Unseen Classes for Zero-Shot Semantic Segmentation
abstract
Recently, CLIP has found practical utility in the domain of pixel-level zero-shot segmentation tasks. The present landscape features two-stage methodologies beset by issues such as intricate pipelines and elevated computational costs. While current one-stage approaches alleviate these concerns and incorporate Visual Prompt Training (VPT) to uphold CLIP's generalization capacity, they still fall short in fully harnessing CLIP's potential for pixel-level unseen class demarcation and precise pixel predictions. To further stimulate CLIP's zero-shot dense prediction capability, we propose SPT-SEG, a one-stage approach that improves CLIP's adaptability from image to pixel. Specifically, we initially introduce Spectral Prompt Tuning (SPT), incorporating spectral prompts into the CLIP visual encoder's shallow layers to capture structural intricacies of images, thereby enhancing comprehension of unseen classes. Subsequently, we introduce the Spectral Guided Decoder (SGD), utilizing both high and low-frequency information to steer the network's spatial focus towards more prominent classification features, enabling precise pixel-level prediction outcomes. Through extensive experiments on two public datasets, we demonstrate the superiority of our method over state-of-the-art approaches, performing well across all classes and particularly excelling in handling unseen classes.
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Li Guo 0004, Man Zhang 0005, Xiaopeng Zhang 0001
AAAI2
2024 AC-CAM: Affinity-Aware Contrast CAM for Weakly-Supervised Semantic Segmentation on MRI Brain Tumor
abstract
Applying the latest visual transformer (ViT) to Weakly-Supervised Semantic Segmentation (WSSS) can compensate for the local perception limitations of CNN, but it also brings about the over-smoothing problem, that is, the final patch labels tend to be uniform. To overcome this challenge, we present an Affinity-Aware Contrast Class Activation Maps (AC-CAM) framework aimed at enhancing WSSS for MRI Brain Tumor analysis by exploiting only image-level labels. We propose two main components: the Affinity-Aware Token Contrast Module (ATCM) and the Affinity-Aware Refine Module (ARM). ATCM utilizes semantic affinities from attention maps to improve the contrast between patch tokens, effectively reducing the over-smoothing tendency of Vision Transformers (ViT). ARM refines the pseudo labels further, incorporating RGB and affinity information to capture the intricate details of the target objects. Our approach capitalizes on the global feature capturing capabilities of ViT, producing more accurate pseudo-labels for WSSS. The framework is optimized through a composite loss function that ensures the consistency of representations for positive token pairs and discriminability for negative ones. Experiments show that our method achieves state-of-the-art performance on the BraTS 2021 dataset.
Jingqian Wu, Changwei Wang 0001, Duzhen Zhang, Rongtao Xu
BIBM6
2024 MIM-HD: Making Smaller Masked Autoencoder Better with Efficient Distillation
abstract
Self-supervised learning and knowledge distillation intersect to achieve exceptional performance on downstream tasks across diverse network capacities. This paper introduces MIM-HD, which implements enhancements for masked image modeling (MIM) distillation, in two key aspects. First, a vision transformer head-level relation adaptive distillation approach is proposed, allowing the student to dynamically draw multi-source knowledge from the teacher based on its evolving state, compatible with scenarios where teacher-student transformer block head count differs. Second, to address the overemphasis on the encoder and neglect of the decoder role in maintaining representation consistency in previous MIM distillations, a dual-view decoding strategy for latent visual representations is introduced, reusing the teacher’s decoder to alleviate MIM burdens on smaller networks. MIM-HD effectiveness is demonstrated through evaluations on ADE20K (mIoU) and ImageNet-1K (Acc), achieving +1.4% and +0.5% improved performance, respectively, compared to state-of-the-art methods, with substantial advantages on smaller pre-training datasets. Moreover, MIM-HD achieves superior efficiency, reducing pre-training epochs from 300 to 100.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Li Guo 0004, Jiguang Zhang, Xiaoqiang Teng, Wenbo Xu 0003
ECAI3
2024 HCF-Net: Hierarchical Context Fusion Network for Infrared Small Object Detection
abstract
Infrared small object detection is an important computer vision task involving the recognition and localization of tiny objects in infrared images, which usually contain only a few pixels. However, it encounters difficulties due to the diminutive size of the objects and the generally complex backgrounds in infrared images. In this paper, we propose a deep learning method, HCF-Net, that significantly improves infrared small object detection performance through multiple practical modules. Specifically, it includes the parallelized patch-aware attention (PPA) module, dimension-aware selective integration (DASI) module, and multi-dilated channel refiner (MDCR) module. The PPA module uses a multi-branch feature extraction strategy to capture feature information at different scales and levels. The DASI module enables adaptive channel selection and fusion. The MDCR module captures spatial features of different receptive field ranges through multiple depth-separable convolutional layers. Extensive experimental results on the SIRST infrared single-frame image dataset show that the proposed HCF-Net performs well, surpassing other traditional and deep learning models. Code is available at https://github.com/zhengshuchen/HCFNet.
Shibiao Xu, ShuChen Zheng, Rongtao Xu, Changwei Wang 0001, Jiguang Zhang, Xiaoqiang Teng, Ao Li 0002, Li Guo 0004
ICME4
2024 DefFusion: Deformable Multimodal Representation Fusion for 3D Semantic Segmentation
abstract
The complementarity between camera and LiDAR data makes fusion methods a promising approach to improve 3D semantic segmentation performance. Recent transformer-based methods have also demonstrated superiority in segmentation. However, multimodal solutions incorporating transformers are underexplored and face two key inherent difficulties: over-attention and noise from different modal data. To overcome these challenges, we propose a Deformable Multimodal Representation Fusion (DefFusion) framework consisting mainly of a Deformable Representation Fusion Transformer and Dynamic Representation Augmentation Modules. The Deformable Representation Fusion Transformer introduces the deformable mechanism in multimodal fusion, avoiding over-attention and improving efficiency by adaptively modeling a 2D key/value set for a given 3D query, thus enabling multimodal fusion with higher flexibility. To enhance the 2D representation and 3D representation, the Dynamic Representation Enhancement Module is proposed to dynamically remove noise in the input representation via Dynamic Grouped Representation Generation and Dynamic Mask Generation. Extensive experiments validate that our model achieves the best 3D semantic segmentation performance on SemanticKITTI and NuScenes benchmarks.
Rongtao Xu, Changwei Wang 0001, Duzhen Zhang, Man Zhang 0005, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
ICRA1
2024 NeuroClips: Towards High-fidelity and Smooth fMRI-to-Video Reconstruction
abstract
Reconstruction of static visual stimuli from non-invasion brain activity fMRI achieves great success, owning to advanced deep learning models such as CLIP and Stable Diffusion. However, the research on fMRI-to-video reconstruction remains limited since decoding the spatiotemporal perception of continuous visual experiences is formidably challenging. We contend that the key to addressing these challenges lies in accurately decoding both high-level semantics and low-level perception flows, as perceived by the brain in response to video stimuli. To the end, we propose NeuroClips, an innovative framework to decode high-fidelity and smooth video from fMRI. NeuroClips utilizes a semantics reconstructor to reconstruct video keyframes, guiding semantic accuracy and consistency, and employs a perception reconstructor to capture low-level perceptual details, ensuring video smoothness. During inference, it adopts a pre-trained T2V diffusion model injected with both keyframes and low-level perception flows for video reconstruction. Evaluated on a publicly available fMRI-video dataset, NeuroClips achieves smooth high-fidelity video reconstruction of up to 6s at 8FPS, gaining significant improvements over state-of-the-art models in various metrics, e.g., a 128% improvement in SSIM and an 81% improvement in spatiotemporal metrics. Our project is available at https://github.com/gongzix/NeuroClips.
Zixuan Gong, Guangyin Bao, Qi Zhang 0020, Zhongwei Wan, Duoqian Miao 0001, Shoujin Wang, Lei Zhu 0003, Changwei Wang 0001, Rongtao Xu, Liang Hu 0004, Yu Zhang 0133
NeurIPS9
2024 Harmonic Long-Range Backscatter with Frequency-Shifted Lightweight Tag
abstract
In this paper, we introduce a simplified long-range (LoRa) backscatter system that enables a lightweight tag to communicate with a remote transceiver using chirp carrier and harmonic backscatter. The key idea is twofold: first, we delegate the chirp waveform generation from the tag to the transceiver, thereby maintaining lightweight tag design; second, we introduce a frequency shift on the tag during its carrier modulation to create harmonic backscatter, thereby mitigating self-interference. To refine the framework, we first detail the system model, including the design for harmonic backscatter. We then present an efficient method to address synchronization and detection issues at the transceiver. Finally, we prototype the proposed system and evaluate its performance. Experimental results demonstrate that in a corridor environment with a carrier power of 0 dBm, the tag can backscatter a signal from the transceiver at a maximum bit rate of 4 kb/s over a distance of 50 meters.
Junliang Lin, Xiannan Zhang, Rongtao Xu, Gongpu Wang, Tony Q. S. Quek
VTC Fall3
2024 Versatile-Modulation and Megabit-Rate Backscatter System: Design, Implementation, and Experimental Results
abstract
Due to its almost zero power consumption and low-hardware costs, backscatter technology has recently gained considerable attention. However, traditional backscatter systems employ basic binary modulations to transmit data at a relatively low rate. In this article, we introduce VITAS, a new backscatter system that enables the tag to communicate with the transceiver using versatile modulations at a megabit rate. First, we present the design of the tag, which supports both binary and higher order modulations. We also develop a series of assembly instructions for the tag controller to generate modulated symbols with minimal clock cycles, thereby enhancing the data rate for a given clock frequency. Then, we design the transceiver to coordinate the backscatter communications with the tag. The transceiver incorporates a specific symbol synchronization algorithm to address the accumulated timing errors resulting from the tag unstable clock during high-rate transmission. Furthermore, we fabricate the tag on a four-layer acrlong PCB and implement the transceiver on a acrlong USRP X310. Finally, we provide experimental results to show that the VITAS system is capable of transmitting acrshort QPSK symbols at a maximum rate of 3 Mbit/s while consuming only$9.63~\mu \text{W}$(3.21 pJ/bit).
Junliang Lin, Gongpu Wang, Rongtao Xu, Yongjun Xu 0002, Xusheng Wei, Yunyong Zhang
IEEE Internet Things J.3
2024 DomainFeat: Learning Local Features With Domain Adaptation
abstract
Accurate and efficient keypoint detection and description is a fundamental step in various computer vision tasks. In this paper, we extract robust descriptors and detect accurate keypoints by learning local Features with Domain adaptation (DomainFeat). Specifically, our Domainfeat includes image-level domain invariance supervision, pixel-level domain consistency supervision, Pixel-Adaptive keypoint Detection(PA-Det), and cross-domain dataset with domain stable point supervision. First, we introduce the image-level domain invariance supervision to make the high-level feature distributions from different domains close by fusing domain-invariant representations in the decoder. Furthermore, to compensate for the inconsistency between descriptors corresponding to the keypoints at the pixel level, we propose the pixel-level domain consistency supervision. Then we present the Pixel-Adaptive keypoint Detection to efficiently detect accurate keypoints, which can improve accuracy by enhancing the local consistency of heatmaps. Finally, we propose an efficient approach to construct data and supervision labels in diverse domains, which can tackle complex application scenarios. With these novel modules and supervision methods, our DomainFeat can make feature detectors more accurate and descriptors more robust. Extensive experiments confirm that Domainfeat achieves state-of-the-art performance on benchmarks such as Aachen-Day-Night localization, HPatches image matching, and the challenging DNIM dataset.
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Weiliang Meng, Bin Fan 0001, Xiaopeng Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 Exploring Intrinsic Discrimination and Consistency for Weakly Supervised Object Localization
abstract
Weakly supervised object localization (WSOL) is a challenging and promising task that aims to localize objects solely based on the supervision of image category labels. In the absence of annotated bounding boxes, WSOL methods must employ the intrinsic properties of the image classification task pipeline to generate object localizations. In this work, we propose a WSOL method for exploring the Intrinsic Discrimination and Consistency in the image classification task pipeline, and call it as IDC. First, we develop a Triplet Metrics Based Foreground Modeling (TMFM) framework to directly predict object foreground regions using intrinsic discrimination. Unlike Class Activation Map (CAM) based methods that also rely on intrinsic discrimination, our TMFM framework alleviates the problem of only focusing on the most discriminative parts by optimizing foreground and background regions synergistically. Second, we design a Dual Geometric Transformation Consistency Constraints (DGTC2) training strategy to introduce additional supervision and regularization constraints for WSOL by leveraging intrinsic geometric transformation consistency. The proposed pixel-wise and object-wise consistency constraint losses cost-effectively provide spontaneous supervision for WSOL. Extensive experiments show that our IDC method achieves significant and consistent performance gains compared to existing state-of-the-art WSOL approaches. Code is available at: https://github.com/vignywang/IDC.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Ruisheng Wang 0001, Xiaopeng Zhang 0001
IEEE Trans. Image Process.2
2024 SkinFormer: Learning Statistical Texture Representation With Transformer for Skin Lesion Segmentation
abstract
Accurate skin lesion segmentation from dermoscopic images is of great importance for skin cancer diagnosis. However, automatic segmentation of melanoma remains a challenging task because it is difficult to incorporate useful texture representations into the learning process. Texture representations are not only related to the local structural information learned by CNN, but also include the global statistical texture information of the input image. In this paper, we propose a transFormer network (SkinFormer) that efficiently extracts and fuses statistical texture representation for Skin lesion segmentation. Specifically, to quantify the statistical texture of input features, a Kurtosis-guided Statistical Counting Operator is designed. We propose Statistical Texture Fusion Transformer and Statistical Texture Enhance Transformer with the help of Kurtosis-guided Statistical Counting Operator by utilizing the transformer's global attention mechanism. The former fuses structural texture information and statistical texture information, and the latter enhances the statistical texture of multi-scale features. Extensive experiments on three publicly available skin lesion datasets validate that our SkinFormer outperforms other SOAT methods, and our method achieves 93.2% Dice score on ISIC 2018. It can be easy to extend SkinFormer to segment 3D images in the future.
Rongtao Xu, Changwei Wang 0001, Jiguang Zhang, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
IEEE J. Biomed. Health Informatics1
2024 PSTNet: Enhanced Polyp Segmentation With Multi-Scale Alignment and Frequency Domain Integration
abstract
Accurate segmentation of colorectal polyps in colonoscopy images is crucial for effective diagnosis and management of colorectal cancer (CRC). However, current deep learning-based methods primarily rely on fusing RGB information across multiple scales, leading to limitations in accurately identifying polyps due to restricted RGB domain information and challenges in feature misalignment during multi-scale aggregation. To address these limitations, we propose the Polyp Segmentation Network with Shunted Transformer (PSTNet), a novel approach that integrates both RGB and frequency domain cues present in the images. PSTNet comprises three key modules: the Frequency Characterization Attention Module (FCAM) for extracting frequency cues and capturing polyp characteristics, the Feature Supplementary Alignment Module (FSAM) for aligning semantic information and reducing misalignment noise, and the Cross Perception localization Module (CPM) for synergizing frequency cues with high-level semantics to achieve efficient polyp segmentation. Extensive experiments on challenging datasets demonstrate PSTNet's significant improvement in polyp segmentation accuracy across various metrics, consistently outperforming state-of-the-art methods. The integration of frequency domain cues and the novel architectural design of PSTNet contribute to advancing computer-assisted polyp segmentation, facilitating more accurate diagnosis and management of CRC.
Rongtao Xu, Changwei Wang 0001, Xiuli Li, Shibiao Xu, Li Guo 0004
IEEE J. Biomed. Health Informatics2
2024 Wave-Like Class Activation Map With Representation Fusion for Weakly-Supervised Semantic Segmentation
abstract
The Class Activation Map (CAM) is widely used to generate pseudo-labels for Weakly Supervised Semantic Segmentation (WSSS), while it does not adequately consider the modeling of foreground-independent information, resulting in prone to false positive pixels. In this paper, we propose a Wave-like Class Activation Map (WaveCAM) from the perspective of representation fusion and dynamic aggregation representation to alleviate the above problem. Specifically, our WaveCAM includes the foreground-aware representation modeling that enhances perception of foreground information, and the foreground-independent representation modeling that enhances perception of foreground-independent information, and a representation-adaptive fusion module that fuses the two representations. Both representations are expressed as wave functions with amplitude and phase to dynamically aggregate representations and extract semantic information after initialization, and they are fused through the adaptive fusion module to obtain an output containing rich semantic information. Extensive experiments on PASCAL VOC 2012 dataset and MS COCO 2014 dataset validate that our WaveCAM can easily embed multi-stage WSSS and end-to-end WSSS, achieving the state-of-the-art performance.
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
IEEE Trans. Multim.1
2024 Accurate Lung Nodule Segmentation With Detailed Representation Transfer and Soft Mask Supervision
abstract
Accurate lung lesion segmentation from computed tomography (CT) images is crucial to the analysis and diagnosis of lung diseases, such as COVID-19 and lung cancer. However, the smallness and variety of lung nodules and the lack of high-quality labeling make the accurate lung nodule segmentation difficult. To address these issues, we first introduce a novel segmentation mask named " soft mask," which has richer and more accurate edge details description and better visualization, and develop a universal automatic soft mask annotation pipeline to deal with different datasets correspondingly. Then, a novel network with detailed representation transfer and soft mask supervision (DSNet) is proposed to process the input low-resolution images of lung nodules into high-quality segmentation results. Our DSNet contains a special detailed representation transfer module (DRTM) for reconstructing the detailed representation to alleviate the small size of lung nodules images and an adversarial training framework with soft mask for further improving the accuracy of segmentation. Extensive experiments validate that our DSNet outperforms other state-of-the-art methods for accurate lung nodule segmentation, and has strong generalization ability in other accurate medical segmentation tasks with competitive results. Besides, we provide a new challenging lung nodules segmentation dataset for further studies (https://drive.google.com/file/d/15NNkvDTb_0Ku0IoPsNMHezJRTH1Oi1wm/view?usp=sharing).
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Jun Xiao 0005, Xiaopeng Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2023 Self Correspondence Distillation for End-to-End Weakly-Supervised Semantic Segmentation
abstract
Efficiently training accurate deep models for weakly supervised semantic segmentation (WSSS) with image-level labels is challenging and important. Recently, end-to-end WSSS methods have become the focus of research due to their high training efficiency. However, current methods suffer from insufficient extraction of comprehensive semantic information, resulting in low-quality pseudo-labels and sub-optimal solutions for end-to-end WSSS. To this end, we propose a simple and novel Self Correspondence Distillation (SCD) method to refine pseudo-labels without introducing external supervision. Our SCD enables the network to utilize feature correspondence derived from itself as a distillation target, which can enhance the network's feature learning process by complementing semantic information. In addition, to further improve the segmentation accuracy, we design a Variation-aware Refine Module to enhance the local consistency of pseudo-labels by computing pixel-level variation. Finally, we present an efficient end-to-end Transformer-based framework (TSCD) via SCD and Variation-aware Refine Module for the accurate WSSS task. Extensive experiments on the PASCAL VOC 2012 and MS COCO 2014 datasets demonstrate that our method significantly outperforms other state-of-the-art methods. Our code is available at https://github.com/Rongtao-Xu/RepresentationLearning/tree/main/SCD-AAAI2023.
Rongtao Xu, Changwei Wang 0001, Jiaxi Sun, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
AAAI1
2023 Task Relation Distillation and Prototypical Pseudo Label for Incremental Named Entity Recognition
abstract
Incremental Named Entity Recognition (INER) involves the sequential learning of new entity types without accessing the training data of previously learned types. However, INER faces the challenge of catastrophic forgetting specific for incremental learning, further aggravated by background shift (i.e., old and future entity types are labeled as the non-entity type in the current task). To address these challenges, we propose a method called task Relation Distillation and Prototypical pseudo label (RDP) for INER. Specifically, to tackle catastrophic forgetting, we introduce a task relation distillation scheme that serves two purposes: 1) ensuring inter-task semantic consistency across different incremental learning tasks by minimizing inter-task relation distillation loss, and 2) enhancing the model's prediction confidence by minimizing intra-task self-entropy loss. Simultaneously, to mitigate background shift, we develop a prototypical pseudo label strategy that distinguishes old entity types from the current non-entity type using the old model. This strategy generates high-quality pseudo labels by measuring the distances between token embeddings and type-wise prototypes. We conducted extensive experiments on ten INER settings of three benchmark datasets (i.e., CoNLL2003, I2B2, and OntoNotes5). The results demonstrate that our method achieves significant improvements over the previous state-of-the-art methods, with an average increase of 6.08% in Micro F1 score and 7.71% in Macro F1 score.
Duzhen Zhang, Hongliu Li, Wei Cong, Rongtao Xu, Jiahua Dong 0001, Xiuyi Chen
CIKM4
2023 Treating Pseudo-labels Generation as Image Matting for Weakly Supervised Semantic Segmentation
abstract
Generating accurate pseudo-labels under the supervision of image categories is a crucial step in Weakly Supervised Semantic Segmentation (WSSS). In this work, we propose a Mat-Label pipeline that provides a fresh way to treat WSSS pseudo-labels generation as an image matting task. By taking a trimap as input which specifies the foreground, background and unknown regions, the image matting task outputs an object mask with fine edges. The intuition behind our Mat-Label is that generating trimap is much easier than generating pseudo-labels directly under weakly supervised setting. Although current CAM-based methods are off-the-shelf solutions for generating a trimap, they suffer from cross-category and foreground-background pixel prediction confusion. To solve this problem, we develop a Double Decoupled Class Activation Map (D2CAM) for Mat-Label to generate a high-quality trimap. By drawing on the idea of metric learning, we explicitly model class activation map with category decoupling and foreground-background decoupling. We also design two simple yet effective refinement constraints for D2CAM to stabilize optimization and eliminate non-exclusive activation. Extensive experiments validate that our Mat-Label achieves substantial and consistent performance gains compared to current state-of-the-art WSSS approaches.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
ICCV2
2023 Automatic polyp segmentation via image-level and surrounding-level context fusion deep neural network
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
Eng. Appl. Artif. Intell.2
2023 Dual-stream Representation Fusion Learning for accurate medical image segmentation
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
Eng. Appl. Artif. Intell.1
2023 Attention Weighted Local Descriptors
abstract
Local features detection and description are widely used in many vision applications with high industrial and commercial demands. With large-scale applications, these tasks raise high expectations for both the accuracy and speed of local features. Most existing studies on local features learning focus on the local descriptions of individual keypoints, which neglect their relationships established from global spatial awareness. In this paper, we present AWDesc with a consistent attention mechanism (CoAM) that opens up the possibility for local descriptors to embrace image-level spatial awareness in both the training and matching stages. For local features detection, we adopt local features detection with feature pyramid to obtain more stable and accurate keypoints localization. For local features description, we provide two versions of AWDesc to cope with different accuracy and speed requirements. On the one hand, we introduce Context Augmentation to address the inherent locality of convolutional neural networks by injecting non-local context information, so that local descriptors can "look wider to describe better". Specifically, well-designed Adaptive Global Context Augmented Module (AGCA) and Diverse Surrounding Context Augmented Module (DSCA) are proposed to construct robust local descriptors with context information from global to surrounding. On the other hand, we design an extremely lightweight backbone network coupled with the proposed special knowledge distillation strategy to achieve the best trade-off in accuracy and speed. What is more, we perform thorough experiments on image matching, homography estimation, visual localization, and 3D reconstruction tasks, and the results demonstrate that our method surpasses the current state-of-the-art local descriptors. Code is available at: https://github.com/vignywang/AWDesc.
Changwei Wang 0001, Rongtao Xu, Ke Lu 0002, Shibiao Xu, Weiliang Meng, Bin Fan 0001, Xiaopeng Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Toward Accurate and Efficient Road Extraction by Leveraging the Characteristics of Road Shapes
abstract
Automatically extracting roads from very high resolution (VHR) remote sensing images is of great importance in a wide range of remote sensing applications. However, complex shapes of roads (i.e., long, geometrically deformed, and thin) always affected the extraction accuracy, which is one of the challenges of road extraction. Based on the insight into road shape characteristics, we propose a novel road shape aware network (RSANet) to achieve efficient and accurate road extraction. First, we introduce the Efficient Strip Transformer Module (ESTM) to efficiently capture the global context to model the long-distance dependence required by the long roads. Second, we design a Geometric Deformation Estimation Module (GDEM) to adaptively extract the context from the shape deformation caused by shooting roads from different perspectives. Third, we provide a simple but effective Road Edge Focal Loss (REF loss) to make the network focus on optimizing the pixels around the road to alleviate the unbalanced distribution of foreground and background pixels caused by the roads being too thin. Finally, we conduct extensive evaluations on public datasets to verify the effectiveness of RSANet and each of the proposed components. Experiments validate that our RSANet outperforms state-of-the-art methods for road extraction in remote sensing images.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Ruisheng Wang 0001, Jiguang Zhang, Xiaopeng Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 RSSFormer: Foreground Saliency Enhancement for Remote Sensing Land-Cover Segmentation
abstract
High spatial resolution (HSR) remote sensing images contain complex foreground-background relationships, which makes the remote sensing land cover segmentation a special semantic segmentation task. The main challenges come from the large-scale variation, complex background samples and imbalanced foreground-background distribution. These issues make recent context modeling methods sub-optimal due to the lack of foreground saliency modeling. To handle these problems, we propose a Remote Sensing Segmentation framework (RSSFormer), including Adaptive TransFormer Fusion Module, Detail-aware Attention Layer and Foreground Saliency Guided Loss. Specifically, from the perspective of relation-based foreground saliency modeling, our Adaptive Transformer Fusion Module can adaptively suppress background noise and enhance object saliency when fusing multi-scale features. Then our Detail-aware Attention Layer extracts the detail and foreground-related information via the interplay of spatial attention and channel attention, which further enhances the foreground saliency. From the perspective of optimization-based foreground saliency modeling, our Foreground Saliency Guided Loss can guide the network to focus on hard samples with low foreground saliency responses to achieve balanced optimization. Experimental results on LoveDA datasets, Vaihingen datasets, Potsdam datasets and iSAID datasets validate that our method outperforms existing general semantic segmentation methods and remote sensing segmentation methods, and achieves a good compromise between computational overhead and accuracy. Our code is available at https://github.com/Rongtao-Xu/RepresentationLearning/tree/main/RSSFormer-TIP2023.
Rongtao Xu, Changwei Wang 0001, Jiguang Zhang, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
IEEE Trans. Image Process.1
2023 CNDesc: Cross Normalization for Local Descriptors Learning
abstract
For a long time, the local descriptors learning benefited from the use of L2 normalization, which projects the descriptor space onto the hypersphere. However, there is no free lunch in the world. Although hypersphere description space stabilizes the optimization and improves the repeatability of the descriptors, it causes the descriptors to have a denser distribution, which reduces the discrimination between descriptors and leads to some incorrect matches. To alleviate this problem, we propose the learnablecross normalizationtechnology as an alternative to L2 normalization, which can achieve a consistent improvement in several of the current popular local descriptors. In addition, we propose an ER-Backbone that can efficiently reuse features in descriptors extraction and an IDC Loss that can provide an image-level description space distribution consistency constraint to further stimulate the performance of the local descriptors. Based on the above innovations, we provide a novel local descriptors extraction method named CNDesc. We perform experiments on image matching, homography estimation, 3D reconstruction, and visual localization tasks, and the results demonstrate that our CNDesc surpasses the current state-of-the-art local descriptors. Our code is available athttps://github.com/vignywang/CNDesc.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
IEEE Trans. Multim.2
2022 MTLDesc: Looking Wider to Describe Better
abstract
Limited by the locality of convolutional neural networks, most existing local features description methods only learn local descriptors with local information and lack awareness of global and surrounding spatial context. In this work, we focus on making local descriptors ``look wider to describe better'' by learning local Descriptors with More Than Local information (MTLDesc). Specifically, we resort to context augmentation and spatial attention mechanism to make the descriptors obtain non-local awareness. First, Adaptive Global Context Augmented Module and Diverse Local Context Augmented Module are proposed to construct robust local descriptors with context information from global to local. Second, we propose the Consistent Attention Weighted Triplet Loss to leverage spatial attention awareness in both optimization and matching of local descriptors. Third, Local Features Detection with Feature Pyramid is proposed to obtain more stable and accurate keypoints localization. With the above innovations, the performance of the proposed MTLDesc significantly surpasses the current state-of-the-art local descriptors on HPatches, Aachen Day-Night localization and InLoc indoor localization benchmarks. Our code is available at https://github.com/vignywang/MTLDesc.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Bin Fan 0001, Xiaopeng Zhang 0001
AAAI2
2022 DOMAINDESC: Learning Local Descriptors With Domain Adaptation
abstract
Robust and efficient local descriptor is crucial in a wide range of applications. In this paper, we propose a novel descriptor DomainDesc which is invariant as much as possible by learning local Descriptor with Domain adaptation. We design the feature-level domain adaptation loss to improve robustness of our DomainDesc by punishing inconsistent high-level feature distributions of different images, while we present the pixel-level cross-domain consistency loss to compensate for the inconsistency between the descriptors corresponding to the keypoints at the pixel level. Besides, we adopt a new architecture to make the descriptor contain as much information as possible, and combine triplet loss and cross-domain consistency loss for descriptor supervision to ensure the distinguished ability of our descriptor. Finally, we give a cross-domain dataset generation strategy to quickly construct our training dataset for diverse domains to adapt to complex application scenarios. Experiments validate that our DomainDesc achieves state-of-the-art performances on HPatches image matching benchmark and Aachen-Day-Night localization benchmark.
Rongtao Xu, Changwei Wang 0001, Bin Fan 0001, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
ICASSP1
2022 Softgan: Towards Accurate Lung Nodule Segmentation via Soft Mask Supervision
abstract
Accurate lung nodule segmentation from Computed Tomog-raphy (CT) images is crucial to the analysis and diagnosis of lung diseases such as COVID-19 and lung cancer. How-ever, due to the variety of lung nodules and the lack of high-quality labeling, accurate lung nodule segmentation is still a challenging problem. In this paper, we propose a novel paradigm including an automatic accurate annotation pipeline and a segmentation network for this task. First, we introduce a new segmentation mask representation named Soft Mask which has richer and more accurate edge details description and better visualization, and we design a universal automatic Soft Mask annotation pipeline to deal with different datasets. Besides, we provide a new challenging lung nodules segmen-tation dataset with traditional binarized masks and our soft masks for further studies. Second, we propose an effective network called SoftGAN that includes an improved back-bone and an adversarial training framework with Soft Mask, in order to improve the performance of accurate lung nodules segmentation. Extensive experiments validate that our Soft-GAN outperforms the state-of-the-art methods for accurate lung nodule segmentation. [Datasetrelease]
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Jun Xiao 0005, Qimin Peng, Xiaopeng Zhang 0001
ICME2
2022 DA-Net: Dual Branch Transformer and Adaptive Strip Upsampling for Retinal Vessels Segmentation
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
MICCAI (2)2
2022 Channel Estimation and Optimal Training Design for Ambient Backscatter Communication Systems under Sensitivity Constraint
abstract
Ambient backscatter communication (AmBC) is a thriving paradigm of wireless communication. It enables wireless sensors such as tags to harvest energy from the ubiquitous radio frequency (RF) energy and to communicate without batteries, thus enlarging the service life and cutting down the cost of wireless sensors. It stimulates the revolution of the IoT and has received much attention from academia and industry. However, the sensitivity, a practical constraint below which the backscatter device can not be activated to transmit data, is often overlooked in basic and applied research. In this paper, we study the channel estimation problem of the AmBC system with sensitivity constraint. We first formulate the system model of the AmBC system with sensitivity constraint, then give the structure of the two-part training sequence. After that, we introduce the maximum-likelihood (ML) channel estimator and propose an optimal training design to minimize the channel estimation mean-square error. Finally, simulation results are provided to validate our analysis.
Ziqi Cui, Gongpu Wang, Xusheng Wei, Rongtao Xu, Xia Chen 0006
VTC Fall4
2022 Instance segmentation of biological images using graph convolutional network
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
Eng. Appl. Artif. Intell.1
2021 DC-Net: Dual Context Network for 2D Medical Image Segmentation
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
MICCAI (1)1
2020 Design and Implementation on a LoRa System with Edge Computing
abstract
The Long Range (LoRa) systems usually process all the computing tasks on the LoRa central server remotely, which brings large latency to Internet of Things (IoT) applications. In this paper, we propose a new design of a LoRa system with edge computing at the LoRa gateway. Our design enables that some of the time computing tasks for latency-sensitive applications can be dealt with timely. The implementation details of the LoRa gateway are presented along with functionality of each component. Finally, comprehensive experiments are conducted to evaluate the performance of the proposed system. The results show that the proposed system can decrease the latency of IoT applications and balance the workloads between the LoRa central server and the LoRa gateway.
Zhiming Liu 0014, Lu Hou 0001, Rongtao Xu, Kan Zheng
WCNC4
2019 An Adaptive Clustering Scheme Based on Modified Density-Based Spatial Clustering of Applications with Noise Algorithm in Ultra-Dense Networks
abstract
As a key technology to increase the system capacity in 5th generation (5G) mobile communications systems, ultra-dense networks (UDN) is proposed by deploying high-density wireless access points in hot spots. To mitigate the serious inter-cell interference(ICI) arising in UDN, we propose an adaptive clustering scheme as the basis for coordinated multipoint transmission and reception (CoMP), which has been proven to effectively eliminate interference. Two machine learning algorithms, Density-Based Spatial Clustering of Applications with Noise (DBSCAN) and Particle Swarm Optimization (PSO), are introduced into the design of clustering scheme. Simulation results have shown that the proposed scheme can achieve a higher system throughput compared with modified K-means scheme. Furthermore, to be consistent with the concept of green communication, geographically isolated points could be identified and processed to save communicate resources with the proposed scheme.
Yuting Ren, Rongtao Xu
VTC Fall2
2019 A Novel Rate and Channel Control Scheme Based on Data Extraction Rate for LoRa Networks
abstract
Long Range (LoRa) has become one of the most popular Low Power Wide Area (LPWA) technologies, which provides a desirable trade-off among communication range, battery life, and deployment cost. In LoRa networks, several transmission parameters can be allocated to ensure efficient and reliable communication. For example, the configuration of the spreading factor allows tuning the data rate and the transmission distance. However, how to dynamically adjust the setting that minimizes the collision probability while meeting the required communication performance is an open challenge. This paper proposes a novel Data Rate and Channel Control (DRCC) scheme for LoRa networks so as to improve wireless resource utilization and support a massive number of LoRa nodes. The scheme estimates channel conditions based on the short-term Data Extraction Rate (DER), and opportunistically adjusts the spreading factor to adapt the variation of channel conditions. Furthermore, the channel control is carried out to balance the link load of all available channels with the global information of the channel usage, which is able to lower the access collisions under dense deployments. Our experiments demonstrate that the proposed DRCC performs well on improving the reliability and capacity compared with other spreading factor allocation schemes in dense deployment scenarios.
Jinyu Xing, Lu Hou 0001, Rongtao Xu, Kan Zheng
WCNC4
2017 A Field Trial on Opportunities for Improving the Unlicensed Spectrum Utilization of LTE
abstract
The cellular data traffic has been dramatically increased, a typical approach is considering the operation of the Long Term Evolution (LTE) system in the unlicensed spectrum. Although there are a large number of theoretical researches on LTE-U, field trial results are still lacking. In this paper, we build a real time field trial in 5.8GHz unlicensed bands. The typical indoor scenario is deployed and we make field trials to evaluate the performance of LTE-U and Wi-Fi including coverage and capacity. Specifically, a methodology to determine the proper Clear Channel Assessment energy detection (CCA-ED) threshold for LTE-U is proposed to implement the friendly coexistence between LTE-U and Wi-Fi systems. Furthermore, experiments are also provided to validate the feasibility of the suggested method in various scenarios. Test results show that LTE-U can greatly improve the spectrum efficiency and optimize wireless resources. At the same time, a proper CCA- ED threshold is necessary for different systems coexisting friendly and fairly.
Rongtao Xu, Hantao Li, Wenfang Tang
VTC Spring3
2017 A Portable SDR Non-Orthogonal Multiple Access Testbed for 5G Networks
abstract
Non-orthogonal multiple access (NOMA) is envisioned to be one of the promising radio access techniques for the fifth generation (5G) mobile networks. In this paper, a portable NOMA testbed based on software defined radio (SDR) is developed by transporting our NOMA system to mini personal computers (PCs). Moreover, the NOMA testbed has been enhanced from 5 MHz bandwidth to 10 MHz bandwidth. As the computation complexity grows higher with the increase of system bandwidth, the portable NOMA testbed is not competent to fully demonstrate the performance of our NOMA system due to the limitation of mini PC. For a better understanding of the NOMA testbed, the system architecture and scenario are introduced in brief. Then, the implementation of NOMA transceiver and the protocol stack are described separately. Finally, a series of experiments are carried out to evaluate its performance loss compared with the original NOMA system based on desktops. The experimental results indicate that the performance of processors may be a short slab for the development of portable SDR-based testbeds towards 5G networks.
Xingguang Wei, Zhiming Geng, Haitao Liu 0015, Kan Zheng, Rongtao Xu
VTC Spring5
2017 Design and prototyping of low-power wide area networks for critical infrastructure monitoring
abstract
Low‐energy critical infrastructure monitoring (LECIM) networks is essential for the monitoring of infrastructure facilities in smart cities. One critical requirement of an LECIM network is its wide coverage of up to several kilometres by using a star topology instead of the tree or mesh networks. In meeting this requirement, this study develops a system with a transceiver of extremely high receiver sensitivity based on the IEEE 802.15.4k physical layer specifications. To reduce the energy consumption, the modulation schemes suitable for low complexity detection are chosen for the data transmission in the design. Also, an efficient parallel preamble and payload data detection are adopted at the access point of the proposed LECIM to acquire concurrent packets from respective nodes. Meanwhile, a data‐aided dynamic timing adjustment scheme is proposed for data field detection to rapidly and adaptively synchronise to the long duration of data packet. Furthermore, a testbed is implemented using a software‐defined radio to demonstrate the effectiveness of the proposed system design.
Rongtao Xu, Kan Zheng, Xianbin Wang 0001
IET Commun.1
2016 A Software Defined Radio Based IEEE 802.15.4k Testbed for M2M Applications
abstract
The IEEE 802.15.4k standard has defined the phys- ical and multiple media access (MAC) layer for low-energy critical infrastructure monitoring (LECIM) networks, which can be used to monitor infrastructure facilities including industrial metering. The main features of LECIM networks are minimal infrastructure with star topology, long range communication with high receiver sensitivity, very limited energy supplied devices. Based on IEEE 802.15.4k specifications, we have designed and developed the prototypes of end device (ED) and access point (AP) using software defined radio technology. The end device is implemented with an ARM-based MCU and a RF module, while the access point is realized by GNURadio and universal software radio peripheral (USRP). A novel parallel preamble and payload detection is applied at AP to acquire multiple packets from respective ED instead of collision avoidance. Furthermore, the field trails are conducted in urban area to demonstrate and evaluate the effectiveness of testbed design.
Rongtao Xu, Lei Lei 0004, Kan Zheng, Hengyang Shen
VTC Fall1
2014 Time-Selective and Frequency-Selective Relay-Based Channel Capacity for Wireless Communication Systems in High-Speed Railway Environment
abstract
One open problem for wireless communication in the high-speed railway environment is channel capacity. In this paper, we study the problem of ergodic capacity for the time-selective and frequency-selective wireless channels in the high-speed railway environment. Firstly, we build up the channel model for the wireless communication between mobile users on the high-speed train and the base station along the railway via the train antennas. Next both the upper bound and the lower bound for the ergodic channel capacity are derived on the basis of this model. Furthermore, the closed-form expression for the ergodic channel capacity is obtained at high SNR. Finally, simulation results are provided to corroborate our proposed studies.
Yang Liu 0048, Zhangdui Zhong, Gongpu Wang, Rongtao Xu
VTC Spring4
2012 Location assistant beamforming for high speed railway
abstract
In this paper, we introduce an algorithm of location assistant direction of departure (DOD) based beamforming scheme which is suitable for wireless communication in highspeed railway (HSR). The beamforming weights are calculated according to the locations of mobile station (MS) and base station (BS) together with the topology of multiple transmit antennas. The purpose of the algorithm is to make phase adjustment at the transmitter such that the output signal-to-noise ratio (SNR) at the receiver can be improved. The performance of beamforming with ideal DOD and estimated DOD which has deviation caused by location errors are both evaluated. Results show that beamforming with ideal DOD greatly outperforms all other transmission modes, i.e., transmit diversity and single-input single-output (SISO), while the inevitable DOD estimation error degrades the overall performance.
Rongtao Xu, Zhangdui Zhong
IWCMC2
2012 Transmission schemes for high-speed railway: Direct or relay?
abstract
Two transmission schemes exist in wireless communication system for high-speed railway: the direct transmission and relay-based transmission. The main purpose of this paper is to evaluate the performance of these two transmission schemes. Specifically, the closed-form expressions of ergodic capacity for each transmission scheme are derived. By comparing these obtained capacity expressions, we find that the relay-based transmission has better performance in the low and intermediate SNR, while the direct transmission becomes a better choice at high SNR. It is also shown that the signal amplitude loss resulted from signal traversing through the carriage is a key parameter in determining the ergodic capacity for direct transmission scheme. Finally, the simulation results demonstrate an exact match with our theoretical findings.
Zhangdui Zhong, Rongtao Xu, Gongpu Wang
IWCMC3
2011 Low complexity Kolmogorov-Smirnov modulation classification
abstract
Kolmogorov-Smirnov (K-S) test-a non-parametric method to measure the goodness of fit, is applied for automatic modulation classification (AMC) in this paper. The basic procedure involves computing the empirical cumulative distribution function (ECDF) of some decision statistic derived from the received signal, and comparing it with the CDFs of the signal under each candidate modulation format. The K-S-based modulation classifier is first developed for AWGN channel, then it is applied to OFDM-SDMA systems to cancel multiuser interference. Regarding the complexity issue of K-S modulation classification, we propose a low-complexity method based on the robustness of the K-S classifier. Extensive simulation results demonstrate that compared with the traditional cumulant-based classifiers, the proposed K-S classifier offers superior classification performance and requires less number of signal samples (thus is fast).
Fanggang Wang 0001, Rongtao Xu, Zhangdui Zhong
WCNC2
2010 A Novel Algorithm to Control Contents Selectively for Vehicular Communication Networks
abstract
With the development of recent vehicular communication technologies, distributing multimedia contents in the vehicular communication networks (VCNs) has become more and more popular, to provide conveniences and entertainment services during the time of driving. However, as multimedia contents are changed and updated dynamically, how to keep the consistency between the original and these replicas in VCNs is very important. Therefore, this paper designs a novel algorithm to control the consistency for the VCNs. In our proposal, after the analyses of the status of road-side units, on-board units and local geographical information, we divide all replicas into two groups, where one is necessary for update and the other are not. Then, we compare the cost to update replicas by using wireless and wired connection, and propose a method to make selection between them. The performance of our proposal is tested by simulation experiments. And the results show that our method can reduce the delay successfully.
Zhou Su 0001, Pinyi Ren, Rongtao Xu, Jiro Katto, Yasuhiko Yasuda
VTC Fall3
2010 Analytical Results for the Performance of MIMO Systems in Frequency Selective Fading Channels
abstract
The performance of spatial multiplexing-based multiple-input multiple output (MIMO) systems using zero forcing (ZF) detector in frequency selective fading channels is analyzed in this paper. The approximation of a linear combination of Wishart matrices is used to derive the probability density function (p.d.f.) of output signal-to-noise ratio (SNR) expression. Analytical error rate expressions for the system are obtained with the assumption that inter-path interference is omitted. Simulations are carried out to evaluate the analytical results. We relate the diversity order of frequency selective fading channels with the multipath power delay profile.
Rongtao Xu, Jiann-Mou Chen, Zhou Su 0001
VTC Fall1
2009 Performance Evaluation of MIMO Systems with Transmit Antenna Selection over Correlated Rayleigh Fading Channels
abstract
In this paper, we investigate MIMO systems with transmit antenna selection over an intra-class correlated Rayleigh fading channel. In our study, one single transmit antenna with an aim to maximizing the total received signal-to-noise ratio (SNR) is selected for transmission, while it is assumed that the receive side use all the available antennas. Using binary-phase-shift-keying (BPSK) modulation as an illustration, we derive the exact bit error rates (BERs) of two diversity schemes, namely transmit antenna selection and full complexity schemes. Moreover, we compare the asymptotic performance of the two diversity schemes analytically. The effect of correlation among the sub-channels on the SNR degradation is also determined in terms of the correlation coefficient and the number of transmit/receive antennas. Finally, results show that the transmit antenna selection scheme can achieve better performance than the full complexity scheme at the expense of a small amount of feedback information from the receiver.
Rongtao Xu, Zhangdui Zhong, Jiann-Mou Chen
VTC Spring1
2009 Approximation to the Capacity of Rician Fading MIMO Channels
abstract
In this paper, the capacity of multiple-input multiple-output (MIMO) systems over a Rician fading channel is evaluated. We show that the capacity of MIMO Rician channels can be well approximated by that of correlated MIMO Rayleigh channels. We take the Rician channel with rank-1 mean matrices as an example, the asymptotic capacities at low and high signal- to-noise ratio (SNR) regions are analyzed. Finally, the asymptotic capacity loss of Rician channels relative to Rayleigh channels is derived.
Rongtao Xu, Zhangdui Zhong, Jiann-Mou Chen
VTC Spring1
2007 A Novel Approach to Analyzing V-BLAST MIMO Systems with Two Transmit Antennas
abstract
A novel approach to analyzing the performance of the V-BLAST multi-input multi-output systems with two transmit antennas is presented in this letter. Based on the properties of Wishart matrices, we derive the exact SNR distributions in the first and second detection steps when optimal detection ordering is used. Closed-form analytical expressions for the bit error rates are then given. The effect of optimal ordering on the diversity order and SNR is evaluated. The results are found to be consistent with those previously published by other researchers
Rongtao Xu, Francis C. M. Lau 0002
IEEE Trans. Wirel. Commun.1
2005 Analytical approach of V-BLAST performance with two transmit antennas
abstract
An analytical approach to the performance analysis of the Vertical Bell Laboratories Space-Time (V-BLAST) multi-input multi-output (MIMO) systems with two transmit antennas is presented in this paper. We derive the exact SNR distribution in the first and second detection steps when optimal detection ordering is used. A closed-form analytical expression for the bit error rate (BER) is then given. The effect of the optimal ordering with two transmit antennas on the diversity order and SNR is also evaluated. Finally, we study the system performance under unbalanced transmit power.
Rongtao Xu, Francis C. M. Lau 0002
WCNC1