Qi Zhao 0037

dblp:05/490-37 · DBLP profile ↗
← Back
39ranked-venue papers
11as first author
33since 2021 · last 2026
0000-0002-3508-027XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 5 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 8 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 ViTaL: A multimodality dataset and benchmark for multi-pathological ovarian tumor recognition
abstract
Ovarian tumor, as a common gynecological disease, can rapidly deteriorate into serious health crises when undetected early, thus posing significant threats to the health of women. Deep neural networks have the potential to identify ovarian tumors, thereby reducing mortality rates, but limited public datasets hinder its progress. To address this gap, we introduce a vital ovarian tumor pathological recognition dataset called ViTaL that contains V isual, T abular and L inguistic modality data of 496 patients across six pathological categories. The ViTaL dataset comprises three subsets corresponding to different patient data modalities: visual data from 2216 two-dimensional ultrasound images, tabular data from medical examinations of 496 patients, and linguistic data from ultrasound reports of 496 patients. It is insufficient to merely distinguish between benign and malignant ovarian tumors in clinical practice. To enable multi-pathology classification of ovarian tumor, we propose a ViTaL-Net based on the Triplet Hierarchical Offset Attention Mechanism (THOAM) to minimize the loss incurred during feature fusion of multi-modal data. This mechanism could effectively enhance the relevance and complementarity between information from different modalities. ViTaL-Net serves as a benchmark for the task of multi-pathology, multi-modality classification of ovarian tumors. In our comprehensive experiments, the proposed method exhibits satisfactory performance, achieving accuracies exceeding 90 % on the two most common pathological types of ovarian tumors and an overall performance of 85 %. Our dataset and code are available at https://github.com/GGbond-study/vitalnet .
Lijiang Chen, Guangxia Cui, Wenpei Bai, Shuchang Lyu, Qi Zhao 0037
Expert Syst. Appl.8
2026 SDP-GS: Sparse-view Gaussian splatting via segmentation-aware depth priors
Qi Zhao 0037, Yangyan Deng, Hong Zhang 0018, Yifan Yang 0003, Ding Yuan 0001
Neurocomputing1
2026 Unsupervised cross-domain semantic segmentation on multi-modality ovarian tumor ultrasound data
Shuchang Lyu, Qi Zhao 0037, Wenpei Bai, Linghan Cai, Guangxia Cui, Lijiang Chen, Huiyu Zhou 0001
Pattern Recognit.2
2026 LM2CNet: Enhancing monocular 3D visual grounding with language guided multi-modality coupling network
Qi Zhao 0037, Shuchang Lyu, Longhao Zou
Pattern Recognit. Lett.2
2026 The Teacher-Student Interactive Cycle: Joint Optimization With Inner-Loop Self-Distillation in Prompted Foundation Models for Efficient Semantic Segmentation
abstract
In the field of semantic segmentation, the high computational cost of deep models poses a major barrier to deployment on edge devices. Among various efficiency-oriented methods, knowledge distillation has emerged as a promising technique for transferring knowledge from large models to lightweight networks. However, current knowledge distillation methods for efficient semantic segmentation still face two key challenges: (1) they often rely on large offline pre-trained teacher networks that remain fixed during training, and (2) they lack joint optimization mechanisms that enable effective teacher-student interaction in pixel-wise dense prediction. As a result, mutual learning strategies originally designed for image-level classification often fail to capture the fine-grained consistency required for semantic segmentation. To address these two challenges, we propose a novel training framework termed Teacher-Student Interactive Cycle (TSIC), which performs efficient semantic segmentation. Specifically, TSIC integrates a lightweight student network into a prompt-based foundation model as a prompted segmentor to assist an online-trained teacher. The student provides coarse mask prompts to guide the teacher, while the teacher offers fine-grained supervision through posterior probabilities and intermediate feature maps. This loop enables joint online optimization without relying on offline pre-trained teachers and fosters effective bidirectional communication. Extensive experiments conducted on several benchmark datasets, including Cityscapes, Pascal VOC, CamVid, and ADE20k, demonstrate the effectiveness of TSIC. Compared to previous methods, TSIC achieves superior segmentation mIoU in most scenarios. Our code will be made publicly available at https://github.com/CV-ShuchangLyu/TSIC.
Qi Zhao 0037, Shuchang Lyu, Longhao Zou, Dingding Yao, Chenguang Yang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 CCMamba: A Multi-Head Criss-Cross Mamba for Pavement Crack Segmentation
abstract
Automatic and accurate detection of pavement cracks on highways and main roads is very important for intelligent pavement maintenance and traffic safety. However, due to the irregular shapes and uncertain characteristics of cracks, current methods face serious challenges on crack segmentation task. To solve the above issues, we propose a Multi-Head Criss-Cross Mamba (CCMamba) to efficiently enhance crack segmentation performance. Based on the sufficient details retained in low-level features and semantic information with strong discriminative ability in high-level features, CCMamba utilizes high-level features to extract crack information from low-level features. First, because of the feature representation capacity of latent state in State Space Model (SSM), we introduce a Multi-Head Latent State Module (Multi-Head LSM) to criss-cross study high-level local features and generate multiple Dynamic Convolution kernels. Second, these Dynamic Convolution kernels are applied to low-level features in a convolutional way, filtering crack information from massive background interferences. Third, the horizontal and vertical features output from Dynamic Convolution layers are fused with head attention mechanism, producing crack sensitive features. Finally, we use the obtained multi-scale features to predict segmentation masks. Comprehensive experiments on five public datasets, Crack500, GAPs384, CFD, CrackVision12K and CPRID, are conducted and our CCMamba achieves state-of-the-art (SOTA) performances compared to current crack segmentation methods. Meanwhile, ablation studies, visualization analysis and real scene testing also validate the effectiveness of CCMamba. Codes of this paper are public available athttps://github.com/cv516Buaa/BinghaoLiu/tree/main/CCMamba
Binghao Liu, Qi Zhao 0037, Hongbo Xie, Hong Zhang 0018
IEEE Trans. Intell. Transp. Syst.2
2025 GLDNN: A Group-Level Dynamic Neural Network via Knowledge Distillation for Image Recognition
Qi Zhao 0037, Ahmed Oluwatoyin, Shuchang Lyu, Alhassan Kamara
ICIG (3)2
2025 ROME is Forged in Adversity: Robust Distilled Datasets via Information Bottleneck
abstract
Dataset Distillation (DD) compresses large datasets into smaller, synthetic subsets, enabling models trained on them to achieve performance comparable to those trained on the full data. However, these models remain vulnerable to adversarial attacks, limiting their use in safety-critical applications. While adversarial robustness has been extensively studied in related fields, research on improving DD robustness is still limited. To address this, we propose ROME, a novel method that enhances the adversarial RObustness of DD by leveraging the InforMation BottlenEck (IB) principle. ROME includes two components: a performance-aligned term to preserve accuracy and a robustness-aligned term to improve robustness by aligning feature distributions between synthetic and perturbed images. Furthermore, we introduce the Improved Robustness Ratio (I-RR), a refined metric to better evaluate DD robustness. Extensive experiments on CIFAR-10 and CIFAR-100 demonstrate that ROME outperforms existing DD methods in adversarial robustness, achieving maximum I-RR improvements of nearly 40% under white-box attacks and nearly 35% under black-box attacks. Our code is available at https://github.com/zhouzhengqd/ROME.
Zheng Zhou 0007, Wenquan Feng, Shuchang Lyu, Qi Zhao 0037
ICML5
2025 HFDNet: High-Frequency Divergence Attention Network for Underwater Segmentation
abstract
Currently, most underwater operations are conducted in deep water, and there is usually insufficient illumination in these areas. At this time, the local texture features of some objects are highly similar in images, and it is difficult to distinguish the inter-class boundaries. This typically results in poor performance of the current semantic segmentation models of terrestrial images in underwater scenes. Taking advantage of the general characteristic that high-frequency regions are more likely to correspond to semantic segmentation boundaries, we introduce the high-frequency Divergence Attention Network (HFDNet), a semantic segmentation model based on transformer. HFDNet extracts its frequency distribution by analyzing the frequency domain of the feature map, and then calculates the relative spectral magnitude of each component by comparing its frequency amplitude against the average amplitude within its local neighborhood in the frequency domain. The local frequency map can be incorporated into the attention matrix as a weighting factor to realize the divergence of attention to the surrounding areas, which improves the attention to the high-frequency areas. This operation can enhance the model’s focus on the object boundary region and local neigh-borhood categories for each component. Therefore, our model can alleviate the problem of determining the object boundary caused by insufficient light in underwater image segmentation, and enhance the ability to segment objects with similar local features under low light conditions. We conduct comprehensive experiments on three underwater segmentation datasets: Caveseg, SUIM and UWS. The results show that our HFDNet achieves state-of-the-art (SOTA) performance on the testing datasets. The source code is available at https://github.com/cv516Buaa/HongboXie/tree/main/HFDNet.
Hongbo Xie, Qi Zhao 0037, Binghao Liu
IROS2
2025 Look in different views: Multi-scheme regression guided cell instance segmentation
Yunmeng Huang, Wenquan Feng, Shuchang Lyu, Qi Zhao 0037, Lijiang Chen
Knowl. Based Syst.5
2025 CLIP-TNseg: A Multi-Modal Hybrid Framework for Thyroid Nodule Segmentation in Ultrasound Images
abstract
Thyroid nodule segmentation in ultrasound images is crucial for accurate diagnosis and treatment planning. However, existing methods struggle with segmentation accuracy, interpretability, and generalization. This letter proposes CLIP-TNseg, a novel framework that integrates a multimodal large model with a neural network architecture to address these challenges. We innovatively divide visual features into coarse-grained and fine-grained components, leveraging textual integration with coarse-grained features for enhanced semantic understanding. Specifically, the Coarse-grained Branch extracts high-level semantic features from a frozen CLIP model, while the Fine-grained Branch refines spatial details using U-Net-style residual blocks. Extensive experiments on the newly collected PKTN dataset and other public datasets demonstrate the competitive performance of CLIP-TNseg. Additional ablation experiments confirm the critical contribution of textual inputs, particularly highlighting the effectiveness of our carefully designed textual prompts compared to fixed or absent textual information.
Boxiong Wei, Yalong Jiang, Liquan Mao, Qi Zhao 0037
IEEE Signal Process. Lett.5
2025 A Voronoi Density-Based Locally Unique Network for Fine-Grained Multi-Label Classification
abstract
Multi-label image classification aims to classify all categories in images simultaneously. When current multi-label classification methods meet fine-grained objects in a single image, the extreme inter-class similarity and over-prediction problems are two major challenges that hinder model performance. To solve the above two problems, we propose Voronoi density based Locally Unique Network (VoLUNet). First, due to high correlation between predictions of different classes, following the Kolmogorov-Arnold Network (KAN), we design the Weak Inter-class Correlation Classifier (WIC-Classifier) to replace linear weights setting in MLP architecture, promoting the potential of fine-grained discrimination. Second, we propose a Local Non-Maximum Suppression (Local-NMS) loss to multi-label classification model, predicting only one unique class with high prediction value for each local region. Third, different classes may have different pixel proportions and Local-NMS loss will be imbalanced for diverse fine-grained classes, we design the Voronoi Density based Superpixel Module (VDSM) to balance the quantities of local feature vectors with different classes. Finally, comprehensive experiments are conducted on four datasets, TreeSatAI, GeoLifeCLEF, FothemNet ShipRSImageNet, and our VoLUNet can significantly improve the classification performance compared to current state-of-the-art models. Codes of this paper are public available at https://github.com/cv516Buaa/BinghaoLiu/tree/main/VoLUNet.
Binghao Liu, Qi Zhao 0037, Lijiang Chen
IEEE Trans. Circuits Syst. Video Technol.2
2025 A Masked Reference Token Supervision-Based Iterative Visual-Language Framework for Robust Visual Grounding
abstract
Visual Grounding (VG) has become a prominent task in recent years, achieving significant advancements with the development of detection and vision transformers. However, existing VG methods struggle to handle the effects of inaccurate or irrelevant textual descriptions, tending to generate false-alarm objects. Moreover, existing methods fail to capture fine-grained features, accurate localization, and comprehensive context understanding from the whole image and textual descriptions. To address these issues, we propose an Iterative Robust Visual Grounding (IR-VG) framework with Multi-stage False-alarm Sensitive Decoder (MFSD) to prevent the generation of false-alarm objects when presented with inaccurate expressions. The framework introduces Masked Reference based Centerpoint Supervision (MRCS) and Iterative Multi-level Vision-language Fusion (IMVF) for enhancing the accuracy of localization and better visual-language alignment. To investigate the elements that affect VG robustness further, we release a robust VG benchmark with 24,000 instances and we also provide a detailed classification of false-alarm according to different parts of speech. Extensive experiments on existing state-of-the-art (SOTA) VG methods and foundation models have proven that it is difficult to handle the robustness of VG by existing models. Even foundation models, which have been pre-trained with a large amount of data, have difficulty to understand inaccurate language descriptions. Our IR-VG can handle false-alarm issues in robust VG well and achieve new SOTA results on the newly proposed robust VG datasets. Ablation studies and visualization experiments demonstrate the effectiveness of the proposed components. Moreover, the proposed framework is also verified effective on five regular VG datasets. Codes and models will be publicly athttps://github.com/cv516Buaa/IR-VG.
Wenquan Feng, Shuchang Lyu, Xiangtai Li, Binghao Liu, Qi Zhao 0037
IEEE Trans. Circuits Syst. Video Technol.7
2025 Unsupervised Domain Adaptation for VHR Urban Scene Segmentation via Prompted Foundation Model-Based Hybrid Training Joint-Optimized Network
abstract
Unsupervised Domain Adaptation for Remote Sensing Semantic Segmentation (UDA-RSSeg) is to adapt a model trained on the source domain data to the target domain samples, thereby minimizing the need for annotated data across diverse remote sensing scenes. In urban planning and monitoring, the task of UDA-RSSeg on Very-High-Resolution (VHR) images has garnered significant research interest. While recent deep learning techniques have demonstrated huge success in tackling the UDA-RSSeg task for VHR urban scenes, a persistent challenge in addressing the domain shift issue remains. Specifically, there are two primary problems: (1) severe inconsistencies in feature representation across diverse domains, characterized by notably differing data distributions, and (2) the domain gap problem due to the representation bias of the source domain patterns when translating features to predictive logits. To solve these problems, we propose a prompted foundation model based hybrid training joint-optimized network (PFM-JONet) for UDA-RSSeg on VHR urban scene. Our approach integrates the notable “Segment Anything Model” (SAM) as prompted foundation model to leverage its robust generalized representation capabilities, thereby alleviating feature inconsistencies. Based on the feature extracted by SAM-Encoder, we introduce a mapping decoder designed to convert SAM-Encoder features into predictive logits. Additionally, a prompted segmentor is employed to generate class-agnostic maps, which guide the mapping decoder’s feature representations. To efficiently optimize the entire network in an end-to-end manner, we design a hybrid training scheme that integrates feature-level and logits-level adversarial training strategies alongside a self-training mechanism. This scheme enhances the model from diverse, compatible perspectives. To evaluate the performance of our proposed PFM-JONet, we conduct extensive experiments on urban scene benchmark datasets, including ISPRS (Potsdam/Vaihingen) and CITY-OSM (Paris/Chicago). On ISPRS dataset, PFM-JONet surpasses previous SOTA methods by 1.60% in mean IoU value across four adaptation tasks. For CITY-OSM’s adaptation task, it outperforms SOTA by 4.84% in mean IoU value. These results demonstrate the effectiveness of our method. Furthermore, visualization and analysis reinforce the method’s interpretability. The code of this paper is available at https://github.com/CV-ShuchangLyu/PFM-JONet.
Shuchang Lyu, Qi Zhao 0037, Yaxuan Sun, Yiwei He, Guangbiao Wang, Jinchang Ren, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.2
2025 Learn From Past to Future: Exploiting Self-Training and Curriculum Learning in Remote Sensing Class-Incremental Semantic Segmentation
abstract
Class-incremental semantic segmentation focuses on updating the segmentation model with only new-class samples. Catastrophic forgetting and background shift are the two prevalent challenges. We identify two additional issues in remote sensing data that worsen these problems: significant class distribution variability and error accumulation-induced model degradation. To solve these three problems, we propose a new Self-Training and Curriculum Learning Guided Dynamic Refined Network (STCL-DRNet). First, we introduce a self-training auxiliary branch to complement the frozen last-step model, integrating cross-step knowledge to mitigate rapid forgetting. Then, a gradient-oriented Dynamic Refined Loss is proposed to assess under-learned classes and mitigate class imbalance. Furthermore, class-balanced curriculum learning is embedded to alleviate performance degradation throughout incremental training. Extensive experiments on benchmark datasets, including DeepGlobe, iSAID, ISPRS Potsdam, and Vaihingen, demonstrate that the proposed STCL-DRNet achieves state-of-the-art (SOTA) performance. In the 1-1s setting of the DeepGlobe dataset, STCL-DRNet exceeds previous SOTA methods by 11.6% in mIoU. For the iSAID 10-1s setting, it outperforms the previous SOTA by 12.76% in mIoU. As for ISPRS Potsdam and Vaihingen, our STCL-DRNet surpasses the SOTA by 5%-8% in all settings. Visualization and analysis further validate its interpretability. Our code is available at https://github.com/cv516Buaa/STCL-DRNet.
Ruimin Ren, Hongbo Zhao 0001, Shuchang Lyu, Guangbiao Wang, Qi Zhao 0037, Jinchang Ren
IEEE Trans. Geosci. Remote. Sens.7
2025 Online Self-Training Driven Attention-Guided Self-Mimicking Network for Semantic Segmentation
abstract
In the realm of semantic segmentation tasks, knowledge distillation (KD) has emerged as a prominent strategy, leveraging the transfer of mature knowledge from large teacher networks to enhance the performance of smaller student networks. However, existing methods often rely heavily on high-quality yet cumbersome teacher networks, leading to a complex training process. To address this challenge, we introduce a novel approach termed self-training driven attention-guided self-mimicking online ensemble network. Our proposed method begins by employing intermediate channel-joint attention maps to guide image augmentation. Both the original and augmented images are then input into the networks. Leveraging intermediate feature maps and predictive predictions generated from the two images, we employ KD to uncover invariant features. To further harness representation potential through learning from credible predictions, we introduce a self-training mechanism. This mechanism utilizes an exponential moving average (EMA)-teacher network constructed using the exponential moving average technique to generate feature maps and predicted posterior probabilities. The knowledge of the EMA-teacher is subsequently transferred to the student network through distillation. Extensive experiments and visualization analyses conducted on multiple benchmark datasets, including Cityscapes, Pascal VOC, CamVid, and ADE20k, validate the effectiveness of self-training driven attention-guided self-mimicking network (ST-ASMNet). The interpretability of our method is further validated through visualization and analysis. Our code will be publicly available.
Shuchang Lyu, Qi Zhao 0037, Hong Zhang 0018, Chenguang Yang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 Optimization for Efficient Federated Learning: Joint Scheduling in Vehicular Edge Computing Networks
abstract
In the context of rapid urban informatization, numerous vehicular devices have undertaken the responsibilities of local storage and data processing, with Federated Learning (FL) assuming a pivotal role within Vehicular Edge Computing Networks (VECNs). However, disparities in data quality and resources among vehicles may pose challenges to the efficiency of FL. To this end, we investigate the client selection and resource allocation issues specific to Unmanned Aerial Vehicle (UAV)-assisted vehicles within the domain of FL. Firstly, we construct a dynamic interactive reputation model where UAVs evaluate and select client vehicles based on factors like performance and capability, effectively filtering out high-quality data sources and enhancing the system’s ability to resist malicious node attacks. Secondly, we formulate a joint optimization problem to design a scheduling strategy that efficiently manages computational resources and communication capabilities, thus controlling latency and reducing energy consumption resulting from local model training. Additionally, we propose an asynchronous parallel Deep Deterministic Policy Gradient (APDDPG) algorithm with shared experience replay, aimed at enhancing the stability of global model convergence. Simulation results reveal that our proposed model and algorithm can more effectively resist attacks from malicious nodes and more fully utilize resources compared to other approaches, ultimately achieving efficient FL.
Changming Zou, Hongbo Zhao 0001, Liwei Geng, Qi Zhao 0037, Dawei Wang 0001
GLOBECOM4
2024 Self-Training and Curriculum Learning Guided Dynamic Refined Network for Remote Sensing Class-Incremental Semantic Segmentation
abstract
Class-incremental semantic segmentation aims to update the segmentation model with training samples containing only novel categories. Within this domain, catastrophic forgetting is a common challenge. In remote sensing scenes, images always have large discrepancies caused by a large variety of geographical objects. Therefore, besides catastrophic forgetting, there exists two additional primary challenges persist. The first one is a huge imbalance in image categories while the second one is error accumulation during multiple incremental training steps. To solve these three problems, we propose a new Self-Training and Curriculum Learning Guided Dynamic Refined Network (STCL-DRNet). Specifically, we first design a self-training-based branch to ease the tendency of catastrophic forgetting. We then design a dynamic refined loss to mitigate the uneven category distribution. Through further embedding class-balanced curriculum learning, we can alleviate the performance drop from noisy accumulation. Extensive experiments on benchmark datasets, DeepGLobe and iSAID, prove that the proposed STCL-DRNet achieves new SOTA performance. Visualization and analysis further substantiate the interpretability.
Hongbo Zhao 0001, Ruimin Ren, Shuchang Lyu, Binghao Liu, Qi Zhao 0037
IGARSS8
2024 OVGNet: A Unified Visual-Linguistic Framework for Open-Vocabulary Robotic Grasping
abstract
Recognizing and grasping novel-category objects remains a crucial yet challenging problem in real-world robotic applications. Despite its significance, limited research has been conducted in this specific domain. To address this, we seamlessly propose a novel framework that integrates open-vocabulary learning into the domain of robotic grasping, empowering robots with the capability to adeptly handle novel objects. Our contributions are threefold. Firstly, we present a large-scale benchmark dataset specifically tailored for evaluating the performance of open-vocabulary grasping tasks. Secondly, we propose a unified visual-linguistic framework that serves as a guide for robots in successfully grasping both base and novel objects. Thirdly, we introduce two alignment modules designed to enhance visual-linguistic perception in the robotic grasping process. Extensive experiments validate the efficacy and utility of our approach. Notably, our framework achieves an average accuracy of 71.2% and 64.4% on base and novel categories in our new dataset, respectively. Our code and dataset are available at https://github.com/cv516Buaa/OVGNet.
Qi Zhao 0037, Shuchang Lyu, Yujing Ma, Chenguang Yang 0001
IROS2
2024 OV-VG: A benchmark for open-vocabulary visual grounding
abstract
Open-vocabulary learning has emerged as a cutting-edge research area, particularly in light of the widespread adoption of vision-based foundational models. Its primary objective is to comprehend novel concepts that are not encompassed within a predefined vocabulary. One key facet of this endeavor is Visual Grounding (VG), which entails locating a specific region within an image based on a corresponding language description. While current foundational models excel at various visual language tasks, there is a noticeable absence of models specifically tailored for open-vocabulary visual grounding (OV-VG). This research endeavor introduces novel and challenging OV tasks, namely Open-Vocabulary Visual Grounding (OV-VG) and Open-Vocabulary Phrase Localization (OV-PL). The overarching aim is to establish connections between language descriptions and the localization of novel objects. To facilitate this, we have curated a comprehensive annotated benchmark, encompassing 7272 OV-VG images (comprising 10,000 instances) and 1000 OV-PL images. In our pursuit of addressing these challenges, we delved into various baseline methodologies rooted in existing open-vocabulary object detection (OV-D), VG, and phrase localization (PL) frameworks. Surprisingly, we discovered that state-of-the-art (SOTA) methods often falter in diverse scenarios. Consequently, we developed a novel framework that integrates two critical components: Text-Image Query Selection (TIQS) and Language-Guided Feature Attention (LGFA). These modules are designed to bolster the recognition of novel categories and enhance the alignment between visual and linguistic information. Extensive experiments demonstrate the efficacy of our proposed framework, which consistently attains SOTA performance across the OV-VG task. Additionally, ablation studies provide further evidence of the effectiveness of our innovative models. Codes and datasets will be made publicly available at https://github.com/cv516Buaa/OV-VG.
Wenquan Feng, Xiangtai Li, Shuchang Lyu, Binghao Liu, Lijiang Chen, Qi Zhao 0037
Neurocomputing8
2024 SWIN-TOD: Smooth Wasserstein Distance and Instance-Level Neighboring Enhancement for Remote Sensing Tiny Object Detection
abstract
The advancement of deep neural network has propelled the widespread application of remote sensing target detection. However, compared to natural scenes, remote sensing targets possess inherent characteristics such as weak features and small scale, leading to a significant performance gap in traditional detection methods. To address these challenges, we undertake a systematic analysis of existing approaches, focusing on two key aspects: inadequate extraction of discriminative features and inappropriate regression measurement metrics. To tackle the first issue, an instance-level neighboring enhancement network (INEN) is proposed, enhancing the network’s feature extraction capability through inter-object feature aggregation. To address the second issue, a novel metric, smooth Wasserstein loss (SWL), is devised. Building upon these principles, a new tiny object detection (TOD) network for remote sensing images is developed. Extensive experiments on AI-TOD v1/v2 and DOTA v2 remote sensing tiny target detection datasets demonstrate that our approach achieves state-of-the-art (SOTA) performance. Codes are available athttps://github.com/sevenwgb/SWIN-TOD.
Guangbiao Wang, Hongbo Zhao 0001, Shuchang Lyu, Qing Chang 0003, Wenquan Feng, Qi Zhao 0037, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.7
2023 Enhancing Spatial Consistency and Class-Level Diversity for Segmenting Fine-Grained Objects
Qi Zhao 0037, Binghao Liu, Shuchang Lyu, Yifan Yang 0003
ICONIP (11)1
2023 Feature reconstruction and metric based network for few-shot object detection
abstract
In the object detection task, deep learning-based methods always need a large amount of annotated training data. However, annotating a large number of images is labor-intensive. In order to reduce the dependency of expensive annotations, we propose a novel end-to-end feature reconstruction and metric based network for few-shot object detection (FM-FSOD). FM-FSOD integrates metric learning and meta-learning to tackle the few-shot object detection task. FM-FSOD is a class-agnostic detection model that can accurately recognize novel categories without fine-tuning on novel categories. Specifically, to quickly learn the characteristics of novel categories, we propose a meta-representation module (MR module) to learn from intra-class mean prototypes and acquire the ability to reconstruct high-level features with the meta-learning method. To further conduct the similarity of features between support prototypes and query ROI features, we propose Pearson metric module (PR module), which serves as a classifier. Compared with the previous standard cosine distance module, the PR module enables the model to acquire robust ability for large bias features. We have conducted extensive experiments on benchmark datasets FSOD, MS COCO, and PASCAL VOC to demonstrate the feasibility and efficiency of our model. Comparing with the previous methods, FM-FSOD obtains comparable results.
Wenquan Feng, Shuchang Lyu, Qi Zhao 0037
Comput. Vis. Image Underst.4
2023 Learn by Oneself: Exploiting Weight-Sharing Potential in Knowledge Distillation Guided Ensemble Network
abstract
Recent CNNs (convolutional neural networks) have become more and more compact. The elegant structure design highly improves the performance of CNNs. With the development of knowledge distillation technique, the performance of CNNs gets further improved. However, existing knowledge distillation guided methods either rely on offline pretrained high-quality large teacher models or online heavy training burden. To solve the above problems, we propose a feature-sharing and weight-sharing based ensemble network (training framework) guided by knowledge distillation (EKD-FWSNet) to make baseline models stronger in terms of representation ability with less training computation and memory cost involved. Specifically, motivated by getting rid of the dependence of offline pretrained teacher model, we design an end-to-end online training scheme to optimize EKD-FWSNet. Motivated by decreasing the online training burden, we only introduce one auxiliary classmate branch to construct multiple forward branches, which will then be integrated as ensemble teacher to guide baseline model. Compared to previous online ensemble training frameworks, EKD-FWSNet can provide diverse output predictions without relying on increasing auxiliary classmate branches. Motivated by maximizing the optimization power of EKD-FWSNet, we exploit the representation potential of weight-sharing blocks and design efficient knowledge distillation mechanism in EKD-FWSNet. Extensive comparison experiments and visualization analysis on benchmark datasets (CIFAR-10/100, tiny-ImageNet, CUB-200 and ImageNet) show that self-learned EKD-FWSNet can boost the performance of baseline models by large margin, which has obvious superiority compared to previous related methods. Extensive analysis also proves the interpretability of EKD-FWSNet. Our code is available at https://github.com/cv516Buaa/EKD-FWSNet.
Qi Zhao 0037, Shuchang Lyu, Lijiang Chen, Binghao Liu, Ting-Bing Xu, Wenquan Feng
IEEE Trans. Circuits Syst. Video Technol.1
2023 MGML: Multigranularity Multilevel Feature Ensemble Network for Remote Sensing Scene Classification
Qi Zhao 0037, Shuchang Lyu, Yujing Ma, Lijiang Chen
IEEE Trans. Neural Networks Learn. Syst.1
2022 A Similarity Distillation Guided Feature Refinement Network for Few-Shot Semantic Segmentation
abstract
Few-shot semantic segmentation is a challenging task of predicting object categories in pixel-wise with only few annotated samples. Existing methods mainly have two problems, which are representation inconsistency of query and support images and semantic-level feature insufficient. To tackle the two problems, we propose a similarity distillation guided feature refinement network (SD-FRNet). Specifically, we first use support label to generate support similarity feature map and coarse prediction of query image. Then, we use this coarse prediction to generate query similarity feature. To compensate feature representation inconsistency, we conduct knowledge distillation mechanism to align similarity features of query and support images. To enrich semantic-level feature, we further design a feature refinement module, which achieves high-quality segmentation. Extensive experiments show the effectiveness of SD-FRNet. On benchmark datasets, PASCAL-5iand COCO-20i, our proposed SD-FRNet outperform the previous SOTA (state-of-the-art) results.
Shuchang Lyu, Binghao Liu, Lijiang Chen, Qi Zhao 0037
ICIP4
2022 Using Guided Self-Attention with Local Information for Polyp Segmentation
Linghan Cai, Meijing Wu, Lijiang Chen, Wenpei Bai, Shuchang Lyu, Qi Zhao 0037
MICCAI (4)7
2022 ML-FDA: Meta-Learning via Feature Distribution Alignment for Few-Shot Learning
abstract
Computer vision tasks suffer from the high cost of collecting large amounts of labeled data. Few-shot Learning (FSL) is a dominant approach to solve this problem because it provides an insight to learn the knowledge of novel categories with few training samples. In FSL task, Meta-learning and metric learning have achieved impressive results. However, the performance of this task is still limited by large intra-class variance and small inter-class distance caused by limited number of few samples. To solve this problem, In this paper, we propose a new method, which integrates meta-learning and metric learning techniques. Specifically, we first propose a feature representation module (FR) to construct representative support class prototypes and query features. Then, we design bias loss to minimize the bias between support and query samples. Furthermore, we design an intra-class loss to minimize the distance between query class prototype and each query sample. We denote this model as ML-FDA and validate it on standard few-shot classification benchmark datasets (MiniImageNet, CIFAR-FS, FC100). The results show that our method improves the performance over other same paradigm methods and achieves the best performance on most benchmarks. The ablation study and visulization analysis also demonstrate the effectiveness of our method.
Binghao Liu, Shuchang Lyu, Lijiang Chen, Qi Zhao 0037, Wenquan Feng
VCIP5
2022 Limited text speech synthesis with electroglottograph based on Bi-LSTM and modified Tacotron-2
abstract
Abstract This paper proposes a framework of applying only the EGG signal for speech synthesis in the limited categories of contents scenario. EGG is a sort of physiological signal which can reflect the trends of the vocal cord movement. Note that EGG’s different acquisition method contrasted with speech signals, we exploit its application in speech synthesis under the following two scenarios. (1) To synthesize speeches under high noise circumstances, where clean speech signals are unavailable. (2) To enable dumb people who retain vocal cord vibration to speak again. Our study consists of two stages, EGG to text and text to speech. The first is a text content recognition model based on Bi-LSTM, which converts each EGG signal sample into the corresponding text with a limited class of contents. This model achieves 91.12% accuracy on the validation set in a 20-class content recognition experiment. Then the second step synthesizes speeches with the corresponding text and the EGG signal. Based on modified Tacotron-2, our model gains the Mel cepstral distortion (MCD) of 5.877 and the mean opinion score (MOS) of 3.87, which is comparable with the state-of-the-art performance and achieves an improvement by 0.42 and a relatively smaller model size than the origin Tacotron-2. Considering to introduce the characteristics of speakers contained in EGG to the final synthesized speech, we put forward a fine-grained fundamental frequency modification method, which adjusts the fundamental frequency according to EGG signals and achieves a lower MCD of 5.781 and a higher MOS of 3.94 than that without modification.
Lijiang Chen, Xia Mao, Qi Zhao 0037
Appl. Intell.5
2022 A feature consistency driven attention erasing network for fine-grained image retrieval
Qi Zhao 0037, Shuchang Lyu, Binghao Liu, Yifan Yang 0003
Pattern Recognit.1
2022 Semantic Segmentation With Attention Mechanism for Remote Sensing Images
abstract
Semantic segmentation for high-resolution remote sensing images is one of the most significant tasks in the field of remote sensing applications. Remote sensing images contain substantial detailed information of ground objects, such as shape, location, and texture. Therefore, these objects make the images exhibit large intraclass variance and small interclass variance, which makes it very difficult to be recognized. In this study, an end-to-end attention-based semantic segmentation network (SSAtNet) is proposed. A pyramid attention pooling module is proposed to introduce the attention mechanism into the multiscale module for adaptive features refinement. To correct the detailed information, the pooling index correction module integrates pooling index maps from the encoder with high-level feature maps, which can help recover the fine-grained features. In the encoder phase, a more effective ResNet-101 backbone is designed to capture detailed features. What is more, a series of data augmentation methods are proposed to enhance the model’s robustness. The proposed model is compared with several previous advanced networks and achieves the state of the art on the ISPRS Vaihingen dataset. The experiment results prove the effectiveness of the SSAtNet.
Qi Zhao 0037, Hong Zhang 0018
IEEE Trans. Geosci. Remote. Sens.1
2022 Embedded Self-Distillation in Compact Multibranch Ensemble Network for Remote Sensing Scene Classification
abstract
Remote sensing (RS) image scene classification task faces many challenges due to the interference from different characteristics of different geographical elements. To solve this problem, we propose a multi-branch ensemble network to enhance the feature representation ability by fusing features in final output logits and intermediate feature maps. However, simply adding branches will increase the complexity of models and decline the inference efficiency. On this issue, we embed self-distillation (SD) method to transfer knowledge from ensemble network to main-branch in it. Through optimizing with SD, main-branch will have close performance as ensemble network. During inference, we can cut other branches to simplify the whole model. In this paper, we first design compact multi-branch ensemble network, which can be trained in an end-to-end manner. Then, we insert SD method on output logits and feature maps. Compared to previous methods, our proposed architecture (ESD-MBENet) performs strongly on classification accuracy with compact design. Extensive experiments are applied on three benchmark RS datasets AID, NWPU-RESISC45 and UC-Merced with three classic baseline models, VGG16, ResNet50 and DenseNet121. Results prove that our proposed ESD-MBENet can achieve better accuracy than previous state-of-the-art (SOTA) complex models. Moreover, abundant visualization analysis make our method more convincing and interpretable.
Qi Zhao 0037, Yujing Ma, Shuchang Lyu, Lijiang Chen
IEEE Trans. Geosci. Remote. Sens.1
2021 Make Baseline Model Stronger: Embedded Knowledge Distillation in Weight-Sharing Based Ensemble Network
Shuchang Lyu, Qi Zhao 0037, Yujing Ma, Lijiang Chen
BMVC2
2019 RSNet: A Compact Relative Squeezing Net for Image Recognition
abstract
Convolutional neural networks(CNN) are showing powerful performance on image recognition tasks. However, when CNN is applied to mobile devices, with limited computing and memory resource, it requires more compact design to maintain a relatively high performance. In this paper, we propose Relative Squeezing Net(RSNet) that provides technical insight into CNN structure for designing a compact model. In an endeavor to improve CondenseNet, we introduce Relative-Squeezing bottleneck where output is weighted percentage of input channels. The design of our bottleneck can transmit diverse and most useful features at all stages. We also employ multiple compression layers to constrain the output channels of feature maps which can eliminate superfluous feature maps and transmit powerful representations to next layers. We evaluate our model on two benchmark datasets; CIFAR and ImageNet. Experimental results show that RSNet achieves state-of-the-art results with less parameters and FLOPs and is more efficient than compact architectures such as CondenseNet, MobileNet and ShuffleNet.
Qi Zhao 0037, Nauman Raoof, Shuchang Lyu, Boxue Zhang, Wenquan Feng
VCIP1
2019 Interpretable Relative Squeezing bottleneck design for compact convolutional neural networks model
abstract
Convolutional neural networks (CNN) are mainly used for image recognition tasks. However, some huge models are infeasible for mobile devices because of limited computing and memory resources. In this paper, feature maps of DenseNet and CondenseNet are visualized. It could be observed that there are some feature channels in locked state and some have similar distribution property, which could be compressed further. Thus, in this work, a novel architecture — RSNet is introduced to improve the computing efficiency of CNNs. This paper proposes Relative-Squeezing (RS) bottleneck design, where the output is the weighted percentage of input channels. Besides, RSNet also contains multiple compression layers and learned group convolutions (LGCs). By eliminating superfluous feature maps, relative squeezing and compression layers only transmit the most significant features to the next layer. Less parameters are employed and much computation is saved. The proposed model is evaluated on three benchmark datasets: CIFAR-10, CIFAR-100 and ImageNet. Experiment results show that RSNet performs better with less parameters and FLOPs, compared to the state-of-the-art baseline, including CondenseNet, MobileNet and ShuffleNet.
Qi Zhao 0037, Boxue Zhang, Shuchang Lyu, Nauman Raoof, Wenquan Feng
Image Vis. Comput.1
2018 AlphaMEX: A smarter global pooling method for convolutional neural networks
abstract
Deep convolutional neural networks have achieved great success on image classification. A series of feature extractors learned from CNN have been used in many computer vision tasks. Global pooling layer plays a very important role in deep convolutional neural networks. It is found that the input feature-maps of global pooling become sparse, as the increasing use of Batch Normalization and ReLU layer combination, which makes the original global pooling low efficiency. In this paper, we proposed a novel end-to-end trainable global pooling operator AlphaMEX Global Pool for convolutional neural network. A nonlinear smooth log-mean-exp function is designed, called AlphaMEX, to extract features effectively and make networks smarter. Compared to the original global pooling layer, our proposed method can improve classification accuracy without increasing any layers or too much redundant parameters. Experimental results on CIFAR-10/CIFAR100, SVHN and ImageNet demonstrate the effectiveness of the proposed method. The AlphaMEX-ResNet outperforms original ResNet-110 by 8.3% on CIFAR10+, and the top-1 error rate of AlphaMEX-DenseNet (k = 12) reaches 5.03% which outperforms original DenseNet (k = 12) by 4.0%.
Boxue Zhang, Qi Zhao 0037, Wenquan Feng, Shuchang Lyu
Neurocomputing2
2018 Multiactivation Pooling Method in Convolutional Neural Networks for Image Recognition
abstract
Convolutional neural networks (CNNs) are becoming more and more popular today. CNNs now have become a popular feature extractor applying to image processing, big data processing, fog computing, etc. CNNs usually consist of several basic units like convolutional unit, pooling unit, activation unit, and so on. In CNNs, conventional pooling methods refer to 2×2 max‐pooling and average‐pooling, which are applied after the convolutional or ReLU layers. In this paper, we propose a Multiactivation Pooling (MAP) Method to make the CNNs more accurate on classification tasks without increasing depth and trainable parameters. We add more convolutional layers before one pooling layer and expand the pooling region to 4×4, 8×8, 16×16, and even larger. When doing large‐scale subsampling, we pick top‐k activation, sum up them, and constrain them by a hyperparameter σ . We pick VGG, ALL‐CNN, and DenseNets as our baseline models and evaluate our proposed MAP method on benchmark datasets: CIFAR‐10, CIFAR‐100, SVHN, and ImageNet. The classification results are competitive.
Qi Zhao 0037, Shuchang Lyu, Boxue Zhang, Wenquan Feng
Wirel. Commun. Mob. Comput.1
2015 Cooperative spectrum sensing via relay-assisted random broadcast in cognitive smartphone networks
Qi Zhao 0037, Daqiang Zhang 0001, Mina Shim, Changqing Yin
Multim. Syst.1
2013 Sensing-Throughput Tradeoff of Relay-Assisted Random Broadcast Based Cognitive Radio Networks
abstract
Random broadcast has been advocated as an effective approach to message sharing among the secondary users (SUs) in the distributed cognitive radio networks (CRNs). However, it is always assumed that if a message is received, it can be received by all SUs in CRNs. The real performance of message sharing and the scheme designing in practice are still unclear. In this paper, we consider a cluster-based distributed CRN, in which the broadcasts of SUs are restricted within clusters and as a result there is no message sharing among clusters. To overcome this, we propose a relay-assisted random broadcast scheme, and introduce the message sharing probability to quantify its performance. We then employ the proposed scheme for cooperative spectrum sensing, and study the sensing-throughput tradeoff problem. Simulations show that the proposed scheme achieves substantial performance improvement on message sharing. Moreover, when given total frame duration, both the sensing performance optimization and the average throughput maximization can be achieved by properly distributing the durations of the primary signal detection, the message sharing via random broadcast and the data transmission.
Qi Zhao 0037
VTC Spring2