EDBT 2026 Demo / reviewers in the wild / expert
Shuchang Lyu
dblp:223/5344
· DBLP profile ↗
32ranked-venue papers
5as first author
28since 2021 · last 2026
0000-0001-9769-7083ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 3 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ViTaL: A multimodality dataset and benchmark for multi-pathological ovarian tumor recognitionabstractOvarian tumor, as a common gynecological disease, can rapidly deteriorate into serious health crises when undetected early, thus posing significant threats to the health of women. Deep neural networks have the potential to identify ovarian tumors, thereby reducing mortality rates, but limited public datasets hinder its progress. To address this gap, we introduce a vital ovarian tumor pathological recognition dataset called ViTaL that contains V isual, T abular and L inguistic modality data of 496 patients across six pathological categories. The ViTaL dataset comprises three subsets corresponding to different patient data modalities: visual data from 2216 two-dimensional ultrasound images, tabular data from medical examinations of 496 patients, and linguistic data from ultrasound reports of 496 patients. It is insufficient to merely distinguish between benign and malignant ovarian tumors in clinical practice. To enable multi-pathology classification of ovarian tumor, we propose a ViTaL-Net based on the Triplet Hierarchical Offset Attention Mechanism (THOAM) to minimize the loss incurred during feature fusion of multi-modal data. This mechanism could effectively enhance the relevance and complementarity between information from different modalities. ViTaL-Net serves as a benchmark for the task of multi-pathology, multi-modality classification of ovarian tumors. In our comprehensive experiments, the proposed method exhibits satisfactory performance, achieving accuracies exceeding 90 % on the two most common pathological types of ovarian tumors and an overall performance of 85 %. Our dataset and code are available at https://github.com/GGbond-study/vitalnet . Lijiang Chen, Guangxia Cui, Wenpei Bai, Shuchang Lyu, Qi Zhao 0037 |
Expert Syst. Appl. | 6 |
| 2026 | Unsupervised cross-domain semantic segmentation on multi-modality ovarian tumor ultrasound data
Shuchang Lyu, Qi Zhao 0037, Wenpei Bai, Linghan Cai, Guangxia Cui, Lijiang Chen, Huiyu Zhou 0001 |
Pattern Recognit. | 1 |
| 2026 | LM2CNet: Enhancing monocular 3D visual grounding with language guided multi-modality coupling network
Qi Zhao 0037, Shuchang Lyu, Longhao Zou |
Pattern Recognit. Lett. | 3 |
| 2026 | The Teacher-Student Interactive Cycle: Joint Optimization With Inner-Loop Self-Distillation in Prompted Foundation Models for Efficient Semantic SegmentationabstractIn the field of semantic segmentation, the high computational cost of deep models poses a major barrier to deployment on edge devices. Among various efficiency-oriented methods, knowledge distillation has emerged as a promising technique for transferring knowledge from large models to lightweight networks. However, current knowledge distillation methods for efficient semantic segmentation still face two key challenges: (1) they often rely on large offline pre-trained teacher networks that remain fixed during training, and (2) they lack joint optimization mechanisms that enable effective teacher-student interaction in pixel-wise dense prediction. As a result, mutual learning strategies originally designed for image-level classification often fail to capture the fine-grained consistency required for semantic segmentation. To address these two challenges, we propose a novel training framework termed Teacher-Student Interactive Cycle (TSIC), which performs efficient semantic segmentation. Specifically, TSIC integrates a lightweight student network into a prompt-based foundation model as a prompted segmentor to assist an online-trained teacher. The student provides coarse mask prompts to guide the teacher, while the teacher offers fine-grained supervision through posterior probabilities and intermediate feature maps. This loop enables joint online optimization without relying on offline pre-trained teachers and fosters effective bidirectional communication. Extensive experiments conducted on several benchmark datasets, including Cityscapes, Pascal VOC, CamVid, and ADE20k, demonstrate the effectiveness of TSIC. Compared to previous methods, TSIC achieves superior segmentation mIoU in most scenarios. Our code will be made publicly available at https://github.com/CV-ShuchangLyu/TSIC. Qi Zhao 0037, Shuchang Lyu, Longhao Zou, Dingding Yao, Chenguang Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | GLDNN: A Group-Level Dynamic Neural Network via Knowledge Distillation for Image Recognition
Qi Zhao 0037, Ahmed Oluwatoyin, Shuchang Lyu, Alhassan Kamara |
ICIG (3) | 4 |
| 2025 | ROME is Forged in Adversity: Robust Distilled Datasets via Information BottleneckabstractDataset Distillation (DD) compresses large datasets into smaller, synthetic subsets, enabling models trained on them to achieve performance comparable to those trained on the full data. However, these models remain vulnerable to adversarial attacks, limiting their use in safety-critical applications. While adversarial robustness has been extensively studied in related fields, research on improving DD robustness is still limited. To address this, we propose ROME, a novel method that enhances the adversarial RObustness of DD by leveraging the InforMation BottlenEck (IB) principle. ROME includes two components: a performance-aligned term to preserve accuracy and a robustness-aligned term to improve robustness by aligning feature distributions between synthetic and perturbed images. Furthermore, we introduce the Improved Robustness Ratio (I-RR), a refined metric to better evaluate DD robustness. Extensive experiments on CIFAR-10 and CIFAR-100 demonstrate that ROME outperforms existing DD methods in adversarial robustness, achieving maximum I-RR improvements of nearly 40% under white-box attacks and nearly 35% under black-box attacks. Our code is available at https://github.com/zhouzhengqd/ROME. Zheng Zhou 0007, Wenquan Feng, Shuchang Lyu, Qi Zhao 0037 |
ICML | 4 |
| 2025 | RAIDX: A Retrieval-Augmented Generation and GRPO Reinforcement Learning Framework for Explainable Deepfake DetectionabstractThe rapid advancement of AI-generation models has enabled the creation of hyperrealistic imagery, posing ethical risks through widespread misinformation. Current deepfake detection methods, categorized as face-specific detectors or general AI-generated detectors, lack transparency by framing detection as a classification task without explaining decisions. While several LLM-based approaches offer explainability, they suffer from coarse-grained analyses and dependency on labor-intensive annotations. This paper introduces RAIDX (Retrieval-Augmented Image Deepfake Detection and Explainability), a novel deepfake detection framework integrating Retrieval-Augmented Generation (RAG) and Group Relative Policy Optimization (GRPO) to enhance detection accuracy and decision explainability. Specifically, RAIDX leverages RAG to incorporate external knowledge for improved detection accuracy and employs GRPO to autonomously generate fine-grained textual explanations and saliency maps, eliminating the need for extensive manual annotations. Experiments on multiple benchmarks demonstrate RAIDX's effectiveness in identifying real or fake, and providing interpretable rationales in both textual descriptions and saliency maps, achieving state-of-the-art detection performance while advancing transparency in deepfake identification. RAIDX represents the first unified framework to synergize RAG and GRPO, addressing critical gaps in accuracy and explainability. Our code and models will be publicly available. Zhenglin Huang, Haiquan Wen, Yiwei He, Shuchang Lyu, Baoyuan Wu |
ACM Multimedia | 5 |
| 2025 | Look in different views: Multi-scheme regression guided cell instance segmentation
Yunmeng Huang, Wenquan Feng, Shuchang Lyu, Qi Zhao 0037, Lijiang Chen |
Knowl. Based Syst. | 3 |
| 2025 | A Masked Reference Token Supervision-Based Iterative Visual-Language Framework for Robust Visual GroundingabstractVisual Grounding (VG) has become a prominent task in recent years, achieving significant advancements with the development of detection and vision transformers. However, existing VG methods struggle to handle the effects of inaccurate or irrelevant textual descriptions, tending to generate false-alarm objects. Moreover, existing methods fail to capture fine-grained features, accurate localization, and comprehensive context understanding from the whole image and textual descriptions. To address these issues, we propose an Iterative Robust Visual Grounding (IR-VG) framework with Multi-stage False-alarm Sensitive Decoder (MFSD) to prevent the generation of false-alarm objects when presented with inaccurate expressions. The framework introduces Masked Reference based Centerpoint Supervision (MRCS) and Iterative Multi-level Vision-language Fusion (IMVF) for enhancing the accuracy of localization and better visual-language alignment. To investigate the elements that affect VG robustness further, we release a robust VG benchmark with 24,000 instances and we also provide a detailed classification of false-alarm according to different parts of speech. Extensive experiments on existing state-of-the-art (SOTA) VG methods and foundation models have proven that it is difficult to handle the robustness of VG by existing models. Even foundation models, which have been pre-trained with a large amount of data, have difficulty to understand inaccurate language descriptions. Our IR-VG can handle false-alarm issues in robust VG well and achieve new SOTA results on the newly proposed robust VG datasets. Ablation studies and visualization experiments demonstrate the effectiveness of the proposed components. Moreover, the proposed framework is also verified effective on five regular VG datasets. Codes and models will be publicly athttps://github.com/cv516Buaa/IR-VG. Wenquan Feng, Shuchang Lyu, Xiangtai Li, Binghao Liu, Qi Zhao 0037 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Unsupervised Domain Adaptation for VHR Urban Scene Segmentation via Prompted Foundation Model-Based Hybrid Training Joint-Optimized NetworkabstractUnsupervised Domain Adaptation for Remote Sensing Semantic Segmentation (UDA-RSSeg) is to adapt a model trained on the source domain data to the target domain samples, thereby minimizing the need for annotated data across diverse remote sensing scenes. In urban planning and monitoring, the task of UDA-RSSeg on Very-High-Resolution (VHR) images has garnered significant research interest. While recent deep learning techniques have demonstrated huge success in tackling the UDA-RSSeg task for VHR urban scenes, a persistent challenge in addressing the domain shift issue remains. Specifically, there are two primary problems: (1) severe inconsistencies in feature representation across diverse domains, characterized by notably differing data distributions, and (2) the domain gap problem due to the representation bias of the source domain patterns when translating features to predictive logits. To solve these problems, we propose a prompted foundation model based hybrid training joint-optimized network (PFM-JONet) for UDA-RSSeg on VHR urban scene. Our approach integrates the notable “Segment Anything Model” (SAM) as prompted foundation model to leverage its robust generalized representation capabilities, thereby alleviating feature inconsistencies. Based on the feature extracted by SAM-Encoder, we introduce a mapping decoder designed to convert SAM-Encoder features into predictive logits. Additionally, a prompted segmentor is employed to generate class-agnostic maps, which guide the mapping decoder’s feature representations. To efficiently optimize the entire network in an end-to-end manner, we design a hybrid training scheme that integrates feature-level and logits-level adversarial training strategies alongside a self-training mechanism. This scheme enhances the model from diverse, compatible perspectives. To evaluate the performance of our proposed PFM-JONet, we conduct extensive experiments on urban scene benchmark datasets, including ISPRS (Potsdam/Vaihingen) and CITY-OSM (Paris/Chicago). On ISPRS dataset, PFM-JONet surpasses previous SOTA methods by 1.60% in mean IoU value across four adaptation tasks. For CITY-OSM’s adaptation task, it outperforms SOTA by 4.84% in mean IoU value. These results demonstrate the effectiveness of our method. Furthermore, visualization and analysis reinforce the method’s interpretability. The code of this paper is available at https://github.com/CV-ShuchangLyu/PFM-JONet. Shuchang Lyu, Qi Zhao 0037, Yaxuan Sun, Yiwei He, Guangbiao Wang, Jinchang Ren, Zhenwei Shi 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Learn From Past to Future: Exploiting Self-Training and Curriculum Learning in Remote Sensing Class-Incremental Semantic SegmentationabstractClass-incremental semantic segmentation focuses on updating the segmentation model with only new-class samples. Catastrophic forgetting and background shift are the two prevalent challenges. We identify two additional issues in remote sensing data that worsen these problems: significant class distribution variability and error accumulation-induced model degradation. To solve these three problems, we propose a new Self-Training and Curriculum Learning Guided Dynamic Refined Network (STCL-DRNet). First, we introduce a self-training auxiliary branch to complement the frozen last-step model, integrating cross-step knowledge to mitigate rapid forgetting. Then, a gradient-oriented Dynamic Refined Loss is proposed to assess under-learned classes and mitigate class imbalance. Furthermore, class-balanced curriculum learning is embedded to alleviate performance degradation throughout incremental training. Extensive experiments on benchmark datasets, including DeepGlobe, iSAID, ISPRS Potsdam, and Vaihingen, demonstrate that the proposed STCL-DRNet achieves state-of-the-art (SOTA) performance. In the 1-1s setting of the DeepGlobe dataset, STCL-DRNet exceeds previous SOTA methods by 11.6% in mIoU. For the iSAID 10-1s setting, it outperforms the previous SOTA by 12.76% in mIoU. As for ISPRS Potsdam and Vaihingen, our STCL-DRNet surpasses the SOTA by 5%-8% in all settings. Visualization and analysis further validate its interpretability. Our code is available at https://github.com/cv516Buaa/STCL-DRNet. Ruimin Ren, Hongbo Zhao 0001, Shuchang Lyu, Guangbiao Wang, Qi Zhao 0037, Jinchang Ren |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | High-Precision Wave-Parameter Perception via Spatiotemporal Coupling of Sea-Clutter Imagery and Ship-Motion ResponsesabstractAccurate detection of near-wave field parameters is crucial for safe navigation and efficient offshore operations. However, estimating precise wave parameter remains challenging due to strong nonlinear wave dynamics and inherent measurement uncertainties. Most of existing deep learning methods relying on single modality data face limitations in insufficient accuracy. Motivated by advances in multimodal fusion algorithms, we propose a novel maritime multimodal fusion inversion model, MR-FuNet. The proposed model integrates a Convolutional Neural Network (CNN) and Bidirectional Long Short-Term Memory (BiLSTM) in parallel, enabling effective fusion of spatial information from X-band radar sea clutter images and temporal patterns from ship motion data at the feature level. A multi-dimensional attention strategy was employed to enable the model to dynamically calibrate and integrate heterogeneous information across modalities and feature domains which can significantly improve the accuracy of inversion for significant wave height and characteristic wave period. To address the lack of comprehensive and high-quality public datasets in this research area, this study constructs and releases a large-scale and multimodal datasetRadar Images and Ship Motion Dataset(RSD). RSD is a large-scale multimodal dataset generated via numerical simulation and covers 99 representative sea states. It provides a valuable benchmark for future research. Extensive experiments validate the effectiveness of the proposed model, demonstrating substantial performance gains across various metrics. Compared to traditional Artificial Neural Networks (ANN), the proposed multimodal fusion model achieves notable reductions in RMSE by 61.6% and 62.8% for the characteristic period and significant wave height inversion tasks, respectively. Dataset is available at https://github.com/felixfelixXu/Radar-Images-and-Ship-Motion-Dataset. Guangbiao Wang, Zihang Xu, Limin Huang, Shuchang Lyu, Jingjun Li, Linzhou Tang, Zhenwei Shi 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Online Self-Training Driven Attention-Guided Self-Mimicking Network for Semantic SegmentationabstractIn the realm of semantic segmentation tasks, knowledge distillation (KD) has emerged as a prominent strategy, leveraging the transfer of mature knowledge from large teacher networks to enhance the performance of smaller student networks. However, existing methods often rely heavily on high-quality yet cumbersome teacher networks, leading to a complex training process. To address this challenge, we introduce a novel approach termed self-training driven attention-guided self-mimicking online ensemble network. Our proposed method begins by employing intermediate channel-joint attention maps to guide image augmentation. Both the original and augmented images are then input into the networks. Leveraging intermediate feature maps and predictive predictions generated from the two images, we employ KD to uncover invariant features. To further harness representation potential through learning from credible predictions, we introduce a self-training mechanism. This mechanism utilizes an exponential moving average (EMA)-teacher network constructed using the exponential moving average technique to generate feature maps and predicted posterior probabilities. The knowledge of the EMA-teacher is subsequently transferred to the student network through distillation. Extensive experiments and visualization analyses conducted on multiple benchmark datasets, including Cityscapes, Pascal VOC, CamVid, and ADE20k, validate the effectiveness of self-training driven attention-guided self-mimicking network (ST-ASMNet). The interpretability of our method is further validated through visualization and analysis. Our code will be publicly available. Shuchang Lyu, Qi Zhao 0037, Hong Zhang 0018, Chenguang Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Self-Training and Curriculum Learning Guided Dynamic Refined Network for Remote Sensing Class-Incremental Semantic SegmentationabstractClass-incremental semantic segmentation aims to update the segmentation model with training samples containing only novel categories. Within this domain, catastrophic forgetting is a common challenge. In remote sensing scenes, images always have large discrepancies caused by a large variety of geographical objects. Therefore, besides catastrophic forgetting, there exists two additional primary challenges persist. The first one is a huge imbalance in image categories while the second one is error accumulation during multiple incremental training steps. To solve these three problems, we propose a new Self-Training and Curriculum Learning Guided Dynamic Refined Network (STCL-DRNet). Specifically, we first design a self-training-based branch to ease the tendency of catastrophic forgetting. We then design a dynamic refined loss to mitigate the uneven category distribution. Through further embedding class-balanced curriculum learning, we can alleviate the performance drop from noisy accumulation. Extensive experiments on benchmark datasets, DeepGLobe and iSAID, prove that the proposed STCL-DRNet achieves new SOTA performance. Visualization and analysis further substantiate the interpretability. Hongbo Zhao 0001, Ruimin Ren, Shuchang Lyu, Binghao Liu, Qi Zhao 0037 |
IGARSS | 3 |
| 2024 | OVGNet: A Unified Visual-Linguistic Framework for Open-Vocabulary Robotic GraspingabstractRecognizing and grasping novel-category objects remains a crucial yet challenging problem in real-world robotic applications. Despite its significance, limited research has been conducted in this specific domain. To address this, we seamlessly propose a novel framework that integrates open-vocabulary learning into the domain of robotic grasping, empowering robots with the capability to adeptly handle novel objects. Our contributions are threefold. Firstly, we present a large-scale benchmark dataset specifically tailored for evaluating the performance of open-vocabulary grasping tasks. Secondly, we propose a unified visual-linguistic framework that serves as a guide for robots in successfully grasping both base and novel objects. Thirdly, we introduce two alignment modules designed to enhance visual-linguistic perception in the robotic grasping process. Extensive experiments validate the efficacy and utility of our approach. Notably, our framework achieves an average accuracy of 71.2% and 64.4% on base and novel categories in our new dataset, respectively. Our code and dataset are available at https://github.com/cv516Buaa/OVGNet. Qi Zhao 0037, Shuchang Lyu, Yujing Ma, Chenguang Yang 0001 |
IROS | 3 |
| 2024 | OV-VG: A benchmark for open-vocabulary visual groundingabstractOpen-vocabulary learning has emerged as a cutting-edge research area, particularly in light of the widespread adoption of vision-based foundational models. Its primary objective is to comprehend novel concepts that are not encompassed within a predefined vocabulary. One key facet of this endeavor is Visual Grounding (VG), which entails locating a specific region within an image based on a corresponding language description. While current foundational models excel at various visual language tasks, there is a noticeable absence of models specifically tailored for open-vocabulary visual grounding (OV-VG). This research endeavor introduces novel and challenging OV tasks, namely Open-Vocabulary Visual Grounding (OV-VG) and Open-Vocabulary Phrase Localization (OV-PL). The overarching aim is to establish connections between language descriptions and the localization of novel objects. To facilitate this, we have curated a comprehensive annotated benchmark, encompassing 7272 OV-VG images (comprising 10,000 instances) and 1000 OV-PL images. In our pursuit of addressing these challenges, we delved into various baseline methodologies rooted in existing open-vocabulary object detection (OV-D), VG, and phrase localization (PL) frameworks. Surprisingly, we discovered that state-of-the-art (SOTA) methods often falter in diverse scenarios. Consequently, we developed a novel framework that integrates two critical components: Text-Image Query Selection (TIQS) and Language-Guided Feature Attention (LGFA). These modules are designed to bolster the recognition of novel categories and enhance the alignment between visual and linguistic information. Extensive experiments demonstrate the efficacy of our proposed framework, which consistently attains SOTA performance across the OV-VG task. Additionally, ablation studies provide further evidence of the effectiveness of our innovative models. Codes and datasets will be made publicly available at https://github.com/cv516Buaa/OV-VG. Wenquan Feng, Xiangtai Li, Shuchang Lyu, Binghao Liu, Lijiang Chen, Qi Zhao 0037 |
Neurocomputing | 5 |
| 2024 | WEA-DINO: An Improved DINO With Word Embedding Alignment for Remote Scene Zero-Shot Object DetectionabstractRemote sensing scene zero-shot object detection aims to detect and recognize both seen and unseen catrgories of landscape elements with the guidance of the word embeddings. In this task, two primary challenges are identified. Firstly, there exists considerable variability within categories of landscape elements, causing a misalignment between visual features and word embeddings, particularly noticeable for unseen categories. Secondly, existing detection models struggle to provide accurate localization predictions, greatly impacting overall performance. To address these two issues, we propose WEA-DINO (Word Embedding Alignment-DINO). Based on the original DINO structure, our WEA-DINO-Head is specifically designed to align the hidden features of “matching queries” with word embedding features, effectively addressing the misalignment issue between visual features and word embeddings. Furthermore, aligning the hidden features of “denoising queries” with word embedding features enables the translation of localization capabilities from known categories to previously unseen ones. Through extensive experimentation on the DIOR benchmark dataset, our method demonstrates state-of-the-art performance. The code is available at https://github.com/cv516Buaa/WEA-DINO. Guangbiao Wang, Hongbo Zhao 0001, Qing Chang 0003, Shuchang Lyu, Huojin Chen |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | SWIN-TOD: Smooth Wasserstein Distance and Instance-Level Neighboring Enhancement for Remote Sensing Tiny Object DetectionabstractThe advancement of deep neural network has propelled the widespread application of remote sensing target detection. However, compared to natural scenes, remote sensing targets possess inherent characteristics such as weak features and small scale, leading to a significant performance gap in traditional detection methods. To address these challenges, we undertake a systematic analysis of existing approaches, focusing on two key aspects: inadequate extraction of discriminative features and inappropriate regression measurement metrics. To tackle the first issue, an instance-level neighboring enhancement network (INEN) is proposed, enhancing the network’s feature extraction capability through inter-object feature aggregation. To address the second issue, a novel metric, smooth Wasserstein loss (SWL), is devised. Building upon these principles, a new tiny object detection (TOD) network for remote sensing images is developed. Extensive experiments on AI-TOD v1/v2 and DOTA v2 remote sensing tiny target detection datasets demonstrate that our approach achieves state-of-the-art (SOTA) performance. Codes are available athttps://github.com/sevenwgb/SWIN-TOD. Guangbiao Wang, Hongbo Zhao 0001, Shuchang Lyu, Qing Chang 0003, Wenquan Feng, Qi Zhao 0037, Zhenwei Shi 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Enhancing Spatial Consistency and Class-Level Diversity for Segmenting Fine-Grained Objects
Qi Zhao 0037, Binghao Liu, Shuchang Lyu, Yifan Yang 0003 |
ICONIP (11) | 3 |
| 2023 | Feature reconstruction and metric based network for few-shot object detectionabstractIn the object detection task, deep learning-based methods always need a large amount of annotated training data. However, annotating a large number of images is labor-intensive. In order to reduce the dependency of expensive annotations, we propose a novel end-to-end feature reconstruction and metric based network for few-shot object detection (FM-FSOD). FM-FSOD integrates metric learning and meta-learning to tackle the few-shot object detection task. FM-FSOD is a class-agnostic detection model that can accurately recognize novel categories without fine-tuning on novel categories. Specifically, to quickly learn the characteristics of novel categories, we propose a meta-representation module (MR module) to learn from intra-class mean prototypes and acquire the ability to reconstruct high-level features with the meta-learning method. To further conduct the similarity of features between support prototypes and query ROI features, we propose Pearson metric module (PR module), which serves as a classifier. Compared with the previous standard cosine distance module, the PR module enables the model to acquire robust ability for large bias features. We have conducted extensive experiments on benchmark datasets FSOD, MS COCO, and PASCAL VOC to demonstrate the feasibility and efficiency of our model. Comparing with the previous methods, FM-FSOD obtains comparable results. Wenquan Feng, Shuchang Lyu, Qi Zhao 0037 |
Comput. Vis. Image Underst. | 3 |
| 2023 | Learn by Oneself: Exploiting Weight-Sharing Potential in Knowledge Distillation Guided Ensemble NetworkabstractRecent CNNs (convolutional neural networks) have become more and more compact. The elegant structure design highly improves the performance of CNNs. With the development of knowledge distillation technique, the performance of CNNs gets further improved. However, existing knowledge distillation guided methods either rely on offline pretrained high-quality large teacher models or online heavy training burden. To solve the above problems, we propose a feature-sharing and weight-sharing based ensemble network (training framework) guided by knowledge distillation (EKD-FWSNet) to make baseline models stronger in terms of representation ability with less training computation and memory cost involved. Specifically, motivated by getting rid of the dependence of offline pretrained teacher model, we design an end-to-end online training scheme to optimize EKD-FWSNet. Motivated by decreasing the online training burden, we only introduce one auxiliary classmate branch to construct multiple forward branches, which will then be integrated as ensemble teacher to guide baseline model. Compared to previous online ensemble training frameworks, EKD-FWSNet can provide diverse output predictions without relying on increasing auxiliary classmate branches. Motivated by maximizing the optimization power of EKD-FWSNet, we exploit the representation potential of weight-sharing blocks and design efficient knowledge distillation mechanism in EKD-FWSNet. Extensive comparison experiments and visualization analysis on benchmark datasets (CIFAR-10/100, tiny-ImageNet, CUB-200 and ImageNet) show that self-learned EKD-FWSNet can boost the performance of baseline models by large margin, which has obvious superiority compared to previous related methods. Extensive analysis also proves the interpretability of EKD-FWSNet. Our code is available at https://github.com/cv516Buaa/EKD-FWSNet. Qi Zhao 0037, Shuchang Lyu, Lijiang Chen, Binghao Liu, Ting-Bing Xu, Wenquan Feng |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | MGML: Multigranularity Multilevel Feature Ensemble Network for Remote Sensing Scene Classification
Qi Zhao 0037, Shuchang Lyu, Yujing Ma, Lijiang Chen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | A Similarity Distillation Guided Feature Refinement Network for Few-Shot Semantic SegmentationabstractFew-shot semantic segmentation is a challenging task of predicting object categories in pixel-wise with only few annotated samples. Existing methods mainly have two problems, which are representation inconsistency of query and support images and semantic-level feature insufficient. To tackle the two problems, we propose a similarity distillation guided feature refinement network (SD-FRNet). Specifically, we first use support label to generate support similarity feature map and coarse prediction of query image. Then, we use this coarse prediction to generate query similarity feature. To compensate feature representation inconsistency, we conduct knowledge distillation mechanism to align similarity features of query and support images. To enrich semantic-level feature, we further design a feature refinement module, which achieves high-quality segmentation. Extensive experiments show the effectiveness of SD-FRNet. On benchmark datasets, PASCAL-5iand COCO-20i, our proposed SD-FRNet outperform the previous SOTA (state-of-the-art) results. Shuchang Lyu, Binghao Liu, Lijiang Chen, Qi Zhao 0037 |
ICIP | 1 |
| 2022 | Using Guided Self-Attention with Local Information for Polyp Segmentation
Linghan Cai, Meijing Wu, Lijiang Chen, Wenpei Bai, Shuchang Lyu, Qi Zhao 0037 |
MICCAI (4) | 6 |
| 2022 | ML-FDA: Meta-Learning via Feature Distribution Alignment for Few-Shot LearningabstractComputer vision tasks suffer from the high cost of collecting large amounts of labeled data. Few-shot Learning (FSL) is a dominant approach to solve this problem because it provides an insight to learn the knowledge of novel categories with few training samples. In FSL task, Meta-learning and metric learning have achieved impressive results. However, the performance of this task is still limited by large intra-class variance and small inter-class distance caused by limited number of few samples. To solve this problem, In this paper, we propose a new method, which integrates meta-learning and metric learning techniques. Specifically, we first propose a feature representation module (FR) to construct representative support class prototypes and query features. Then, we design bias loss to minimize the bias between support and query samples. Furthermore, we design an intra-class loss to minimize the distance between query class prototype and each query sample. We denote this model as ML-FDA and validate it on standard few-shot classification benchmark datasets (MiniImageNet, CIFAR-FS, FC100). The results show that our method improves the performance over other same paradigm methods and achieves the best performance on most benchmarks. The ablation study and visulization analysis also demonstrate the effectiveness of our method. Binghao Liu, Shuchang Lyu, Lijiang Chen, Qi Zhao 0037, Wenquan Feng |
VCIP | 3 |
| 2022 | A feature consistency driven attention erasing network for fine-grained image retrieval
Qi Zhao 0037, Shuchang Lyu, Binghao Liu, Yifan Yang 0003 |
Pattern Recognit. | 3 |
| 2022 | Embedded Self-Distillation in Compact Multibranch Ensemble Network for Remote Sensing Scene ClassificationabstractRemote sensing (RS) image scene classification task faces many challenges due to the interference from different characteristics of different geographical elements. To solve this problem, we propose a multi-branch ensemble network to enhance the feature representation ability by fusing features in final output logits and intermediate feature maps. However, simply adding branches will increase the complexity of models and decline the inference efficiency. On this issue, we embed self-distillation (SD) method to transfer knowledge from ensemble network to main-branch in it. Through optimizing with SD, main-branch will have close performance as ensemble network. During inference, we can cut other branches to simplify the whole model. In this paper, we first design compact multi-branch ensemble network, which can be trained in an end-to-end manner. Then, we insert SD method on output logits and feature maps. Compared to previous methods, our proposed architecture (ESD-MBENet) performs strongly on classification accuracy with compact design. Extensive experiments are applied on three benchmark RS datasets AID, NWPU-RESISC45 and UC-Merced with three classic baseline models, VGG16, ResNet50 and DenseNet121. Results prove that our proposed ESD-MBENet can achieve better accuracy than previous state-of-the-art (SOTA) complex models. Moreover, abundant visualization analysis make our method more convincing and interpretable. Qi Zhao 0037, Yujing Ma, Shuchang Lyu, Lijiang Chen |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Make Baseline Model Stronger: Embedded Knowledge Distillation in Weight-Sharing Based Ensemble Network
Shuchang Lyu, Qi Zhao 0037, Yujing Ma, Lijiang Chen |
BMVC | 1 |
| 2019 | RSNet: A Compact Relative Squeezing Net for Image RecognitionabstractConvolutional neural networks(CNN) are showing powerful performance on image recognition tasks. However, when CNN is applied to mobile devices, with limited computing and memory resource, it requires more compact design to maintain a relatively high performance. In this paper, we propose Relative Squeezing Net(RSNet) that provides technical insight into CNN structure for designing a compact model. In an endeavor to improve CondenseNet, we introduce Relative-Squeezing bottleneck where output is weighted percentage of input channels. The design of our bottleneck can transmit diverse and most useful features at all stages. We also employ multiple compression layers to constrain the output channels of feature maps which can eliminate superfluous feature maps and transmit powerful representations to next layers. We evaluate our model on two benchmark datasets; CIFAR and ImageNet. Experimental results show that RSNet achieves state-of-the-art results with less parameters and FLOPs and is more efficient than compact architectures such as CondenseNet, MobileNet and ShuffleNet. Qi Zhao 0037, Nauman Raoof, Shuchang Lyu, Boxue Zhang, Wenquan Feng |
VCIP | 3 |
| 2019 | Interpretable Relative Squeezing bottleneck design for compact convolutional neural networks modelabstractConvolutional neural networks (CNN) are mainly used for image recognition tasks. However, some huge models are infeasible for mobile devices because of limited computing and memory resources. In this paper, feature maps of DenseNet and CondenseNet are visualized. It could be observed that there are some feature channels in locked state and some have similar distribution property, which could be compressed further. Thus, in this work, a novel architecture — RSNet is introduced to improve the computing efficiency of CNNs. This paper proposes Relative-Squeezing (RS) bottleneck design, where the output is the weighted percentage of input channels. Besides, RSNet also contains multiple compression layers and learned group convolutions (LGCs). By eliminating superfluous feature maps, relative squeezing and compression layers only transmit the most significant features to the next layer. Less parameters are employed and much computation is saved. The proposed model is evaluated on three benchmark datasets: CIFAR-10, CIFAR-100 and ImageNet. Experiment results show that RSNet performs better with less parameters and FLOPs, compared to the state-of-the-art baseline, including CondenseNet, MobileNet and ShuffleNet. Qi Zhao 0037, Boxue Zhang, Shuchang Lyu, Nauman Raoof, Wenquan Feng |
Image Vis. Comput. | 4 |
| 2018 | AlphaMEX: A smarter global pooling method for convolutional neural networksabstractDeep convolutional neural networks have achieved great success on image classification. A series of feature extractors learned from CNN have been used in many computer vision tasks. Global pooling layer plays a very important role in deep convolutional neural networks. It is found that the input feature-maps of global pooling become sparse, as the increasing use of Batch Normalization and ReLU layer combination, which makes the original global pooling low efficiency. In this paper, we proposed a novel end-to-end trainable global pooling operator AlphaMEX Global Pool for convolutional neural network. A nonlinear smooth log-mean-exp function is designed, called AlphaMEX, to extract features effectively and make networks smarter. Compared to the original global pooling layer, our proposed method can improve classification accuracy without increasing any layers or too much redundant parameters. Experimental results on CIFAR-10/CIFAR100, SVHN and ImageNet demonstrate the effectiveness of the proposed method. The AlphaMEX-ResNet outperforms original ResNet-110 by 8.3% on CIFAR10+, and the top-1 error rate of AlphaMEX-DenseNet (k = 12) reaches 5.03% which outperforms original DenseNet (k = 12) by 4.0%. Boxue Zhang, Qi Zhao 0037, Wenquan Feng, Shuchang Lyu |
Neurocomputing | 4 |
| 2018 | Multiactivation Pooling Method in Convolutional Neural Networks for Image RecognitionabstractConvolutional neural networks (CNNs) are becoming more and more popular today. CNNs now have become a popular feature extractor applying to image processing, big data processing, fog computing, etc. CNNs usually consist of several basic units like convolutional unit, pooling unit, activation unit, and so on. In CNNs, conventional pooling methods refer to 2×2 max‐pooling and average‐pooling, which are applied after the convolutional or ReLU layers. In this paper, we propose a Multiactivation Pooling (MAP) Method to make the CNNs more accurate on classification tasks without increasing depth and trainable parameters. We add more convolutional layers before one pooling layer and expand the pooling region to 4×4, 8×8, 16×16, and even larger. When doing large‐scale subsampling, we pick top‐k activation, sum up them, and constrain them by a hyperparameter σ . We pick VGG, ALL‐CNN, and DenseNets as our baseline models and evaluate our proposed MAP method on benchmark datasets: CIFAR‐10, CIFAR‐100, SVHN, and ImageNet. The classification results are competitive. Qi Zhao 0037, Shuchang Lyu, Boxue Zhang, Wenquan Feng |
Wirel. Commun. Mob. Comput. | 2 |