EDBT 2026 Demo / reviewers in the wild / expert
Yang Luo 0001
dblp:68/3710-1
· DBLP profile ↗
30ranked-venue papers
0as first author
30since 2021 · last 2026
0000-0002-4576-5934ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 12 · 12 since 2021Artificial intelligence and machine learning · 9 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tensor-Based Joint Subarray Channel Estimation for Terahertz Ultra-Massive MIMO Systems
Jingyue Cao, Chunbo Luo, Yang Luo 0001 |
ICC | 3 |
| 2026 | Joint Optimization of Multi-UAV Trajectory and User Association for Integrated Sensing and Communication
Hongda Wu, Chunbo Luo, Siyan Gu, Yang Luo 0001 |
ICC | 4 |
| 2026 | Toward Reliable Multimodal Beam Prediction in mmWave Communications via Probabilistic Embedding and Uncertainty-AwareabstractAccurate beam prediction in millimeter-wave bands for 6G networks is challenged by complex channel dynamics, limiting the reliability of RF-only solutions. Recent advances in multimodal fusion leveraging non-RF cues have shown potential, yet existing approaches predominantly focus on fusion architectures while overlooking modality robustness under diverse conditions, thereby constraining their effectiveness. To overcome these limitations, we propose the Advanced Multimodal Representation Network (AMR-Net), a unified framework for robust cross-domain beam prediction. AMR-Net introduces probabilistic latent representations, modeling each sample as a learnable distribution to enhance generalization, and incorporates supervised contrastive learning on reparameterized embeddings to improve discriminability. Moreover, an uncertainty-aware composite fusion module is devised, integrating entropy, divergence, and confidence-based metrics to generate dynamic, sample-specific modality weights, thereby enabling reliable inference. Comprehensive evaluations on public vehicle-to-infrastructure (V2I) benchmarks, training under sunny conditions and testing across both matched and unseen scenarios such as varying weather conditions, different signal-to-noise ratio (SNR) levels, and partial modality loss, validate the effectiveness of AMR-Net. The results show that AMR-Net consistently surpasses state-of-the-art multimodal baselines in both accuracy and robustness, highlighting its superior generalization ability and real-world applicability. Zhiwen Deng 0001, Mengfan Xue, Shuai Xiong, Chunbo Luo, Yang Luo 0001 |
IEEE Internet Things J. | 5 |
| 2026 | Tensor-Based Hybrid-Field Channel Estimation for Extremely Large-Scale Massive MIMO SystemsabstractChannel estimation in extremely large-scale multiple-input multiple-output (XL-MIMO) systems presents substantial challenges, especially under low pilot overhead. This paper proposes a novel tensor-based hybrid-field (TBHF) channel estimation scheme for XL-MIMO. We first develop a unified hybrid-field tensor decomposition framework that models received signals as a third-order low-rank tensor, from which channel parameters, including angle of arrival/departure (AoA/AoD), distance, time delay, and path gain, are jointly estimated via factor matrices. To distinguish far-field and near-field components, we propose a path classification mechanism that does not require prior knowledge of path component proportions. For near-field estimation, we introduce an angle-adaptive distance grid design and a two-step estimation strategy, enhancing accuracy and reducing complexity. We analyze the uniqueness of the Canonical Polyadic Decomposition, demonstrating the potential of TBHF to substantially reduce training overhead. The Cramér-Rao Bound for hybrid-field parameter estimation is derived, and simulations verify that TBHF outperforms existing methods in normalized mean square error. Jingyue Cao, Chunbo Luo, Yang Luo 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2026 | Tensor-Based Channel Estimation for Terahertz Ultra-Massive MIMO SystemsabstractChannel estimation in ultra-massive MIMO (UM-MIMO) systems with an array-of-subarrays (AoSA) architecture is challenged by high-dimensional channels and spherical-wave propagation. To address these issues, we propose a novel tensorbased joint SA (TB-JSA) channel estimation framework for UM-MIMO systems by representing the received signal of each subarray (SA) as a third-order low-rank tensor. The proposed TB-JSA framework enables joint channel estimation across multiple SAs while restricting the computationally intensive tensor decomposition to a carefully selected subset of SAs. Specifically, we first develop an SA selection strategy based on the energy-diversity criterion, which significantly reduces the overall computational complexity without sacrificing estimation performance. Subsequently, we propose a joint SA processing scheme that exploits the inherent spatial correlations among the selected SAs. The proposed method applies Canonical Polyadic Decomposition on the received SA signal tensors to extract path parameters, resolves permutation ambiguities through path matching, and jointly estimates the global parameters using a robust weighted least squares. This joint estimation enables accurate reconstruction of the full channel. Simulation results demonstrate that TB-JSA consistently outperforms all baseline methods in terms of accuracy, with estimation performance approaching the Cram´er–Rao Bound (CRB). Moreover, adjusting the number of selected SAs enables a controllable trade-off between estimation accuracy and computational efficiency. Jingyue Cao, Chunbo Luo, Yang Luo 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2025 | Improving Multimodal Learning via Imbalanced Learning
Shicai Wei, Chunbo Luo, Yang Luo 0001 |
ICCV | 3 |
| 2025 | Boosting Multimodal Learning via Disentangled Gradient LearningabstractMultimodal learning often encounters the under-optimized problem and may have worse performance than unimodal learning. Existing methods attribute this problem to the imbalanced learning between modalities and rebalance them through gradient modulation. However, they fail to explain why the dominant modality in multimodal models also underperforms that in unimodal learning. In this work, we reveal the optimization conflict between the modality encoder and modality fusion module in multimodal models. Specifically, we prove that the cross-modal fusion in multimodal models decreases the gradient passed back to each modality encoder compared with unimodal models. Consequently, the performance of each modality in the multimodal model is inferior to that in the unimodal model. To this end, we propose a disentangled gradient learning (DGL) framework to decouple the optimization of the modality encoder and modality fusion module in the multimodal model. DGL truncates the gradient back-propagated from the multimodal loss to the modality encoder and replaces it with the gradient from unimodal loss. Besides, DGL removes the gradient back-propagated from the unimodal loss to the modality fusion module. This helps eliminate the gradient interference between the modality encoder and modality fusion module while ensuring their respective optimization processes. Finally, extensive experiments on multiple types of modalities, tasks, and frameworks with dense cross-modal interaction demonstrate the effectiveness and versatility of the proposed DGL. Code is available at \href{https://github.com/shicaiwei123/ICCV2025-GDL}{https://github.com/shicaiwei123/ICCV2025-GDL} Shicai Wei, Chunbo Luo, Yang Luo 0001 |
ICCV | 3 |
| 2025 | A Task-Centered Algorithm for AAV-Assisted Communications Based on Deep Reinforcement LearningabstractAutonomous aerial vehicles (AAVs) can be harnessed to provide temporary communication resources and act as a relay to assist users in transmitting their tasks to an external base station or cloud. How to allocate AAV resources to assist users has become an important topic in AAV-assisted communications. Although existing studies have considered terrain characteristics, communication throughput and user equality, task properties have often been overlooked, which would impact the overall quality of service. Due to the multiple complex attributes of tasks and dynamic complexity of the environment, optimization often involves high dimensionality and dynamism, making it difficult for traditional algorithms to distinguish primary tasks and effectively transmit important information. We thus propose a multiconstrained nonconvex joint optimization problem modeled as a partially observed Markov decision process, which brings the resource optimization challenges of AAV communication resources. This article further proposes a deep reinforcement learning (DRL)-based algorithm named weighted-K-means-DDPG (WK-Means-DDPG). The innovation of this algorithm lies in the combination of the traditional Weighted K-Means algorithm and DRL algorithm deep deterministic policy gradient (DDPG) to solve the deployment and resource allocation problems of AAVs and improve the performance of the reinforcement learning algorithm. Simulation results show that our algorithm outperforms state of the art and baseline algorithms in transmitting important information and can be generalized to different size of areas, number of users, and number of AAVs. Longcheng Yan, Chunbo Luo, Xiangyuan Jiang, Yang Luo 0001 |
IEEE Internet Things J. | 4 |
| 2024 | Scale Decoupled DistillationabstractLogit knowledge distillation attracts increasing attention due to its practicality in recent studies. However, it of-ten suffers inferior performance compared to the feature knowledge distillation. In this paper, we argue that existing log it-based methods may be sub-optimal since they only leverage the global logit output that couples multiple se-mantic knowledge. This may transfer ambiguous knowl-edge to the student and mislead its learning. To this end, we propose a simple but effective method, i.e., Scale De-coupled Distillation (SDD), for logit knowledge distillation. SDD decouples the global logit output into multi-ple local logit outputs and establishes distillation pipelines for them. This helps the student to mine and inherit fine-grained and unambiguous logit knowledge. Moreover, the decoupled knowledge can be further divided into consis-tent and complementary logit knowledge that transfers the semantic information and sample ambiguity, respectively. By increasing the weight of complementary parts, SDD can guide the student to focus more on ambiguous samples, im-proving its discrimination ability. Extensive experiments on several benchmark datasets demonstrate the effective-ness of SDD for wide teacher-student pairs, especially in the fine-grained classification task. Code is available at: https://github.comishicaiwei123/SDD-CVPR2024 Shicai Wei, Chunbo Luo, Yang Luo 0001 |
CVPR | 3 |
| 2024 | Robust Multimodal Learning via Representation Decoupling
Shicai Wei, Yang Luo 0001, Yuji Wang, Chunbo Luo |
ECCV (42) | 2 |
| 2024 | AE-DENet: Enhancement for Deep Learning-based Channel Estimation in OFDM SystemsabstractDeep learning (DL)-based methods have demonstrated remarkable achievements in addressing orthogonal frequency division multiplexing (OFDM) channel estimation challenges. However, existing DL-based methods mainly rely on separate real and imaginary inputs while ignoring the inherent correlation between the two streams, such as amplitude and phase information that are fundamental in communication signal processing. This paper proposes AE-DENet, a novel autoencoder(AE)-based data enhancement network to improve the performance of existing DL-based channel estimation methods. AE-DENet focuses on enriching the classic least square (LS) estimation input commonly used in DL-based methods by employing a learning-based data enhancement method, which extracts interaction features from the real and imaginary components and fuses them with the original real/imaginary streams to generate an enhanced input for better channel inference. Experimental findings in terms of the mean square error (MSE) results demonstrate that the proposed method enhances the performance of all state-of-the-art DL-based channel estimators with negligible added complexity. Furthermore, the proposed approach is shown to be robust to channel variations and high user mobility. Ephrem Fola, Yang Luo 0001, Chunbo Luo |
GLOBECOM | 2 |
| 2024 | Convolution Meets Transformer: Efficient Hybrid Transformer for Semantic Segmentation with Very High Resolution ImageryabstractIn this paper, we introduce an efficient and lightweight hybrid Transformer architecture, ingeniously integrating convolutions within Transformer blocks for semantic segmentation of remote sensing Very High Resolution (VHR) imagery. To simultaneously avoid the high computational complexity in the shallow layers and capture the local representations of the VHR images, we propose the Group-Team Convolution Modulation (GTCM) module that uses convolutions to approximate the effect of attention mechanisms and modulates features in channel dimension. Additionally, to enlarge the effective receptive field (ERF) in the decoder, based on the grouping philosophy, we adopt dilated convolutions with multiple dilated rates to further enhance the performance. The superiority and efficiency of our proposed hybrid structure are demonstrated by outperforming state-of-the-art methods on the Vaihingen and Potsdam datasets with relatively lower complexity and fewer parameters. Yuji Wang, Ruojun Zhao, Shicai Wei, Jingchen Ni, Yang Luo 0001, Chunbo Luo |
IGARSS | 6 |
| 2024 | UAV-Based Emergency Communications: An Iterative Two-Stage Multiagent Soft Actor-Critic Approach for Optimal Association and Dynamic DeploymentabstractThis paper investigates future emergency wireless communication systems based on multiple unmanned vehicles cooperative deployment. A terrestrial carrier vehicle with wireless communication and management capabilities are deployed to release multiple unmanned aerial vehicles (UAVs) which will serve as aerial mobile stations (UAV-BSs) to cover a disaster affected area, forming an emergency Internet of Things (IoT) network. Under the proposed system architecture, we formulate a joint optimization challenge considering the UAV-BSs’ dynamic deployment positions and the association policy between user equipments (UEs) and BSs to maximize the throughput and coverage in dynamic scenarios as a time-varying mixed-integer non-convex sequential programming (MINSP) problem. To solve this problem, we first investigate the impact of decision delay caused by physical networking and computing environment on system performance to illustrate the urgent need for efficient algorithms. Then, a two-stage iterative training algorithm called centralized training multi-agent soft actor-critic with branch-and-cut (CT-MASAC-BAC) is proposed for computing globally optimal solutions. Numerical results show that CT-MASAC-BAC outperforms the heuristic algorithms and other benchmark deep reinforcement learning algorithms in terms of system utility. Furthermore, the experimental results show that the proposed algorithm is scalable with an increasing number of deployed UAV-BSs, contributing to potentially increased performance with more serving UAV-BSs. Yingjie Cao, Yang Luo 0001, Haifen Yang, Chunbo Luo |
IEEE Internet Things J. | 2 |
| 2024 | Jointly Optimize Throughput and Localization Accuracy: UAV Trajectory Design for Multiuser Integrated Communication and SensingabstractUnmanned aerial vehicle (UAV) is becoming a crucial aerial platform to provide emergency or enhanced communication and sensing services benefited from its unique features, including agile mobility and high probability of Line-of-Sight coverage. In this article, we investigate the novel UAV trajectory design problem where the UAV communicates with multiple users and simultaneously senses the positions of multiple targets, via the integrated communication and sensing design. To evaluate the overall system utility, we first derive the communication throughput and localization error Cramér-Rao bound (CRB) model and highlight the coupled challenge and tradeoffs on UAV’s trajectory. We thus model the three typical integrated sensing and communication scenarios: 1) communication centric; 2) sensing centric; and 3) tradeoff scenarios. We further introduce path discretization to support in-flight communication and hovering sensing, which make it possible to jointly optimize the UAV trajectory, communication throughput and target localization estimation error CRB with limited complexity. Because this optimization problem involves integer programming caused by multiuser association, we propose a novel trajectory initialization framework based on traveling salesman problem to determine the service order for sensing targets and the initial UAV trajectory. To address the nonconvexity of this optimization problem, we propose an efficient iterative algorithm using the successive convexity approximation technique to obtain the approximate optimal solution, which is dynamically reconfigured and optimized with updated information in flight. Extensive numerical results demonstrate that the proposed algorithm achieves superior performance in the tradeoff between average achievable rate and CRB in all three scenarios. Siyan Gu, Chunbo Luo, Yang Luo 0001, Xiaoguang Ma |
IEEE Internet Things J. | 3 |
| 2024 | Chromosomal Mutation-Inspired Radio Augmentation for Enhanced Automatic Modulation ClassificationabstractAs communication environments grow increasingly complex, efficient modulation scheme identification is critical. Automatic modulation classification (AMC) with deep learning has proven effective under diverse and noisy conditions, but its success hinges on the quality and diversity of training data. This article tackles the challenge of acquiring diverse training data sets through an innovative data augmentation method inspired by chromosomal mutations from genetic algorithms. Designed for I/Q modulation signals, the method introduces six radio augmentations: 1) interstitial deletion; 2) terminal deletion; 3) inversion; 4) breakage; 5) ring; and 6) translocation. These augmentations enrich and diversify training data, enhancing the adaptability of AMC models. Experiments show a ninefold expansion of the training sample space, significantly boosting the benchmark AMC models’ performance. Notably, our method achieves state-of-the-art results with CLDNN on RML2016.10A and RML2016.10B, with mean accuracies of 67.11% and 69.02%, respectively. It also improves the Transformer-based TRN model’s mean accuracy on RML2016.10B by 9.55%. Our approach effectively addresses data scarcity in deep learning-based AMC and offers promising avenues for future communication systems. Xitong Pu, Chunbo Luo, Yihao Yin, Zijian Liu 0004, Yang Luo 0001 |
IEEE Internet Things J. | 5 |
| 2024 | Gradient Decoupled Learning With Unimodal Regularization for Multimodal Remote Sensing ClassificationabstractThe joint use of multisource remote-sensing data for Earth observation has drawn much attention due to its robust performance. Although many methods have been proposed to fuse multimodal data, they tend to improve the interaction of different modality data while ignoring the optimization of each modality. Existing studies show that high-performance modalities will suppress the learning of weak ones, leading to under-optimized multimodal learning. To this end, we propose a general framework called gradient decoupled network (GDNet) to assist the multimodal remote sensing (RS) classification. GDNet guides each modality encoder in the multimodal model to learn probabilistic representations instead of deterministic ones. This helps decouple their gradient, reducing their influence on each other and encouraging them to learn the modality-specific information. Then, we further introduce the unimodal regularization for each modality encoder to align their logit output with the multimodal one and label distribution simultaneously. This helps introduce independent gradient paths for each morality encoder to accelerate their optimization when preserving the modality-share information. Finally, extensive experiments conducted on three benchmark datasets demonstrate that the proposed GDNet can effectively address the under-optimized problem in multimodal RS image classification. Code is available athttps://github.com/shicaiwei123/TGRS-GDNet. Shicai Wei, Chunbo Luo, Xiaoguang Ma, Yang Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Privileged Modality Learning via Multimodal HallucinationabstractLearning based on multimodal data has attracted increasing interest recently. While a variety of sensory modalities can be collected for training, not all of them are always available in practical scenarios, which raises the challenge to infer with incomplete modality. This article presents a general framework termed multimodal hallucination (MMH) to bridge the gap between ideal training scenarios and real-world deployment scenarios with incomplete modality data by transferring the complete multimodal knowledge to the hallucination network with incomplete modality input. Compared with the modality hallucination methods that restore privileged modalities information for late fusion, the proposed framework not only helps to preserve the crucial cross-modal cues but relates the study in complete modalities and in incomplete modalities. Then, we introduce two strategies called region-aware distillation and discrepancy-aware distillation to transfer the response-based and joint-representation-based knowledge of pre-trained multimodal networks, respectively. Region-aware distillation establishes and weights knowledge transferring pipelines between the response of multimodal and hallucination networks at multiple regions, which guides the hallucination network to focus on discriminative regions and avoid wasted gradients. Discrepancy-aware distillation guides the hallucination network to mimic the local inter-sample distance of multimodal representations, which enables the hallucination network to acquire the inter-class discrimination refined by multimodal cues. Extensive experiments on multimodal action recognition and face anti-spoofing demonstrate the proposed multimodal hallucination framework can overcome the problem of incomplete modality input in various scenes and achieve state-of-the-art performance.https://github.com/shicaiwei123/TMM-MMH Shicai Wei, Chunbo Luo, Yang Luo 0001, Jialang Xu |
IEEE Trans. Multim. | 3 |
| 2024 | Throughput Maximization of Dynamic TDD Networks With a Full-Duplex UAV-BSabstractDynamic time division duplex (D-TDD) is a promising technology for the future networks, enabling communication nodes to change uplink and downlink traffic opportunistically in the time domain to improve communication resource management. In this article, we propose an innovative system architecture where a full-duplex unmanned aerial vehicle (UAV) serves as a base station (BS) to enhance the performance of D-TDD networks composed of multiple full-duplex ground users and provides essential mobility and duplexing flexibility. We theoretically and numerically analyze two typical communication scenarios under this architecture, namely single transceiver pair (single-pair case) and multiple transceiver pairs (multi-pair case), at any time slot. In the single-pair case, the full-duplex UAV-BS communicates with at most one uplink user and one downlink user at the same time. In the multi-pair case, the UAV-BS supports multiple uplink and downlink users simultaneously. We formulate the throughput optimization problems for each case by jointly considering the uplink/downlink scheduling, power allocation, energy consumption and UAV-BS trajectory. Efficient algorithms are developed to deal with the non-convex problems and proved to be convergent. In addition, extensive simulation results show that our algorithms can cater to asymmetric uplink and downlink traffic and outperform other competitive benchmarks. Chunbo Luo, Haifen Yang, Yang Luo 0001 |
IEEE Trans. Wirel. Commun. | 4 |
| 2023 | MMANet: Margin-Aware Distillation and Modality-Aware Regularization for Incomplete Multimodal LearningabstractMultimodal learning has shown great potentials in numerous scenes and attracts increasing interest recently. However, it often encounters the problem of missing modality data and thus suffers severe performance degradation in practice. To this end, we propose a general framework called MMANet to assist incomplete multimodal learning. It consists of three components: the deployment network used for inference, the teacher network transferring comprehensive multimodal information to the deployment network, and the regularization network guiding the deployment network to balance weak modality combinations. Specifically, we propose a novel margin-aware distillation (MAD) to assist the information transfer by weighing the sample contribution with the classification uncertainty. This encourages the deployment network to focus on the samples near decision boundaries and acquire the refined inter-class margin. Besides, we design a modality-aware regularization (MAR) algorithm to mine the weak modality combinations and guide the regularization network to calculate prediction loss for them. This forces the deployment network to improve its representation ability for the weak modality combinations adaptively. Finally, extensive experiments on multimodal classification and segmentation tasks demonstrate that our MMANet outperforms the state-of-the-art significantly. Code is available at: https://github.com/shicaiwei123/MMANet Shicai Wei, Chunbo Luo, Yang Luo 0001 |
CVPR | 3 |
| 2023 | Throughput Maximization of Flexible Duplex Networks with an In-Band Full-Duplex UAV-BSabstractFlexible duplex is regarded as a promising technology for the fifth generation (5G) and beyond networks. In flexible duplex networks, communication nodes are capable of changing uplink and downlink traffic adaptively to improve communication resource management. In this article, we employ an in-band full-duplex (IBFD) unmanned aerial vehicle (UAV) serving as a base station (BS) to improve the performance of flexible duplex networks. In order to maximize the system throughput, we formulate an optimization problem by jointly designing the uplink/downlink scheduling, power allocation and UAV trajectory. Sequential parametric convex approximation (SPCA) and successive convex approximation (SCA) methods are utilized to provide an efficient solution for this highly non-convex problem. Furthermore, numerical simulation results show that our algorithm improves the system throughput compared to other designs. Yang Luo 0001, Chunbo Luo |
ICC | 2 |
| 2023 | ExpertNet: Defeat noisy labels by deep expert consultation paradigm for pneumoconiosis staging on chest radiographs
Wenjian Sun, Dongsheng Wu, Yang Luo 0001, Lu Liu 0001, Hongjing Zhang, Houjun Zheng, Jiang Shen, Chunbo Luo |
Expert Syst. Appl. | 3 |
| 2023 | A Multi-Modal Hypergraph Neural Network via Parametric Filtering and Feature SamplingabstractIn the real world, relationships between objects are often complex, involving multiple variables and modes. Hypergraph neural networks possess the capability to capture and represent such intricate relationships by deriving and inheriting their graph-based counterparts. Nevertheless, both graph and hypergraph neural networks suffer from the problem of over-smoothing when multiple graph convolution layers are stacked. To address this issue, this article introduces the Multi-modal Hypergraph Neural Network with Parametric Filtering and Feature Sampling (MHNet) to encode complex hypergraph features and mitigate over-smoothing. The proposed approach uses hypergraph structures to model high-order and multi-modal data correlations, a polynomial hypergraph filter to dynamically extract multi-scale node features through parametric polynomial fitting, and a feature sampling strategy to learn from sparse and labeled samples while avoiding overfitting. Experimental results on four hypergraph datasets and two multi-modal visual datasets demonstrate that the proposed MHNet outperforms state-of-the-art algorithms. Zijian Liu 0004, Yang Luo 0001, Xitong Pu, Geyong Min, Chunbo Luo |
IEEE Trans. Big Data | 2 |
| 2023 | Diversity-Guided Distillation With Modality-Center Regularization for Robust Multimodal Remote Sensing Image ClassificationabstractMultimodal learning has shown great potential in remote sensing image classification and attracted increasing interest in the community. Although it is preferable to collect multiple modalities for training, not all of them are available in practical scenarios. To this end, we propose a general diversity-guided distillation network (DGDNet) with modality-center regularization to facilitate accurate model inference when modalities are missing. Compared with existing modality reconstruction methods, DGDNet does not need prior knowledge of the missing modality and can handle various missing modalities via only one model. Specifically, DGDNet consists of two components: the deployment network extracting the modality-invariant representation for robust inference and the teacher network transferring comprehensive multimodal information to the deployment network. This enables the deployment network to learn the modality invariant and specific information simultaneously while maintaining robustness for incomplete modality input. In particular, we design a novel diversity-guided distillation method that transfers knowledge by matching the feature diversity. This helps overcome the representation heterogeneity when encouraging the deployment network to learn modality-specific information. Besides, a modality-center regularization strategy is proposed to address the unbalanced training of teacher and deployment networks by constraining the intra-class inter-modality variations. This helps alleviate the underfitting for weak modality, improving the model performance. Finally, extensive experiments demonstrate that the proposed DGDNet can address the problem of missing modalities effectively and achieves state-of-the-art performance. The code is available at https://github.com/shicaiwei123/TGRS-DGDNet. Shicai Wei, Yang Luo 0001, Chunbo Luo |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | MSH-Net: Modality-Shared Hallucination With Joint Adaptation Distillation for Remote Sensing Image Classification Using Missing ModalitiesabstractLearning based multimodal data has attracted increasing interest in the remote sensing community owing to its robust performance. Although it is preferable to collect multiple modalities for training, not all of them are available in practical scenarios due to the restriction of imaging conditions. Therefore, how to assist the model inference with missing modalities is significant for multimodal remote sensing image processing. In this work, we propose a general framework called modality-shared hallucination network (MSH-Net) to address this issue by reconstructing complete modality-shared features from the incomplete inference modalities. Compared to conventional privilege modality hallucination methods, MSH-Net does not only help preserve the cross-modal interactions for model inference, but also scales well with the increasing number of missing modalities. We further develop a novel joint adaptation distillation (JAD) method that guides the hallucination model to learn the modality-shared knowledge from the multimodal model by matching the joint probability distributions between representation and groundtruth. This overcomes the representation heterogeneity caused by the discrepancy between inputs and structures of multimodal and hallucination model, while preserving the decision boundaries refined by multimodal cues. Finally, extensive experiments conducted on four common modality combinations demonstrate that the proposed MSH-Net can effectively address the problem of missing modalities and achieve state-of-the-art performance. Code is available at: https://github.com/shicaiwei123/MSHNet. Shicai Wei, Yang Luo 0001, Xiaoguang Ma, Peng Ren 0001, Chunbo Luo |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | A Benchmark Dataset of Endoscopic Images and Novel Deep Learning Method to Detect Intestinal Metaplasia and Gastritis AtrophyabstractEndoscopy has been routinely used to diagnose stomach diseases including intestinal metaplasia (IM) and gastritis atrophy (GA). Such routine examination usually demands highly skilled radiologists to focus on a single patient with substantial time, causing the following two key challenges: 1) the dependency on the radiologist's experience leading to inconsistent diagnosis results across different radiologists; 2) limited examination efficiency due to the demanding time and energy consumption to the radiologist. This paper proposes to address these two issues in endoscopy using novel machine learning method in three-folds. Firstly, we build a novel and relatively big endoscopy dataset of 21,420 images from the widely used White Light Imaging (WLI) endoscopy and more recent Linked Color Imaging (LCI) endoscopy, which were annotated by experienced radiologists and validated with biopsy results, presenting a benchmark dataset. Secondly, we propose a novel machine learning model inspired by the human visual system, named as local attention grouping, to effectively extract key visual features, which is further improved by learning from multiple randomly selected regional images via ensemble learning. Such a method avoids the significant problem in the deep learning methods that decrease the resolution of original images to reduce the size of input samples, which would remove smaller lesions in endoscopy images. Finally, we propose a dual transfer learning strategy to train the model with co-distributed features between WLI and LCI images to further improve the performance. The experiment results, measured by accuracy, specificity, sensitivity, positive detection rate and negative detection rate, on IM are 99.18 %, 98.90 %, 99.45 %, 99.45 %, 98.91 %, respectively, and on GA are 97.12 %, 95.34 %, 98.90 %, 98.86 %, 95.50 %, respectively, achieving state of the art performance that outperforms current mainstream deep learning models. Yan Ou, Zhiqian Chen, Wenjian Sun, Yang Luo 0001, Chunbo Luo |
IEEE J. Biomed. Health Informatics | 6 |
| 2022 | Deep Log-Normal Label Distribution Learning for Pneumoconiosis Staging on Chest RadiographsabstractPneumoconiosis staging has been a challenging task for deep neural networks due to the stage ambiguity in early pneumoconiosis. In this article, we propose a deep log-normal label distribution learning method named DLN-LDL for pneumo-coniosis staging by exploring the intrinsic stage distribution pat-terns of pneumoconiosis. DLN-LDL effectively prevents the deep network from overfitting features in ambiguous chest radiographs that are irrelevant to the stage to which they belong by replacing the one-hot labels with log-normally distributed vectors. The experiments on our collected pneumoconiosis dataset confirm that the proposed DLN-LDL algorithm outperforms other classical methods in terms of Accuracy, Precision, Sensitivity, Specificity, F1-score and Area Under the Curve. Wenjian Sun, Dongsheng Wu, Yang Luo 0001, Lu Liu 0001, Hongjing Zhang, Houjun Zheng, Jiang Shen, Chunbo Luo |
CBMS | 3 |
| 2022 | A transductive learning method to leverage graph structure for few-shot learning
Zijian Liu 0004, Yang Luo 0001, Chunbo Luo |
Pattern Recognit. Lett. | 3 |
| 2022 | Distributed UAV Swarm Formation and Collision Avoidance Strategies Over Fixed and Switching TopologiesabstractThis article proposes a controlling framework for multiple unmanned aerial vehicles (UAVs) to integrate the modes of formation flight and swarm deployment over fixed and switching topologies. Formation strategies enable UAVs to enjoy key collective benefits including reduced energy consumption, but the shape of the formation and each UAV's freedom are significantly restrained. Swarm strategies are thus proposed to maximize each UAV's freedom following simple yet powerful rules. This article investigates the integration and switch between these two strategies, considering the deployment environment factors, such as poor network conditions and unknown and often highly mobile obstacles. We design a distributed formation controller to guide multiple UAVs in orderless states to swiftly reach an intended formation. Inspired by starling birds and similar biological creatures, a distributed collision avoidance controller is proposed to avoid unknown and mobile obstacles. We further illustrated the stability of the controllers over both fixed and switching topologies. The experimental results confirm the effectiveness of the framework. Chunbo Luo, Yang Luo 0001, Ke Li 0001 |
IEEE Trans. Cybern. | 3 |
| 2022 | A Fully Deep Learning Paradigm for Pneumoconiosis Staging on Chest RadiographsabstractPneumoconiosis staging has been a very challenging task, both for certified radiologists and computer-aided detection algorithms. Although deep learning has shown proven advantages in the detection of pneumoconiosis, it remains challenging in pneumoconiosis staging due to the stage ambiguity of pneumoconiosis and noisy samples caused by misdiagnosis when they are used in training deep learning models. In this article, we propose a fully deep learning pneumoconiosis staging paradigm that comprises a segmentation procedure and a staging procedure. The segmentation procedure extracts lung fields in chest radiographs through an Asymmetric Encoder-Decoder Network (AED-Net) that can mitigate the domain shift between multiple datasets. The staging procedure classifies the lung fields into four stages through our proposed deep log-normal label distribution learning and focal staging loss. The two cascaded procedures can effectively solve the problem of model overfitting caused by stage ambiguity and noisy labels of pneumoconiosis. Besides, we collect a clinical chest radiograph dataset of pneumoconiosis from the certified radiologist's diagnostic reports. The experimental results on this novel pneumoconiosis dataset confirm that the proposed deep pneumoconiosis staging paradigm achieves an Accuracy of 90.4%, a Precision of 84.8%, a Sensitivity of 78.4%, a Specificity of 95.6%, an F1-score of 80.9% and an Area Under the Curve (AUC) of 96%. In particular, we achieve 68.4% Precision, 76.5% Sensitivity, 95% Specificity, 72.2% F1-score and 89% AUC on the early pneumoconiosis 'stage-1'. Wenjian Sun, Dongsheng Wu, Yang Luo 0001, Lu Liu 0001, Hongjing Zhang, Houjun Zheng, Jiang Shen, Chunbo Luo |
IEEE J. Biomed. Health Informatics | 3 |
| 2021 | An Adaptive Multi-Scale and Multi-Level Features Fusion Network with Perceptual Loss for Change DetectionabstractChange detection plays a vital role in monitoring and analyzing temporal changes in Earth observation tasks. This paper proposes a novel adaptive multi-scale and multi-level features fusion network for change detection in very-high-resolution bi-temporal remote sensing images. The proposed approach has three advantages. Firstly, it excels in abstracting high-level representations empowered by a highly effective feature extraction module. Secondly, an elaborate feature fusion module incorporated with the channel and spatial attention mechanism is proposed to provide efficient fusion strategies for multi-scale and multi-level features from bi-temporal images and multiple convolutional layers. Finally, a novel perceptual auxiliary component is designed to capture the perceptual loss of the global perceptual and structural differences and address the optimization problem caused by only using per-pixel loss function in change detection. Comprehensive experiments on two benchmark datasets confirm that our proposed framework outperforms state-of-the-art algorithms in both quantitative assessment and visual interpretation. Jialang Xu, Yang Luo 0001, Xinyue Chen 0009, Chunbo Luo |
ICASSP | 2 |