Xin Zhang 0092

dblp:76/1584-92 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
20since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Breaking Class Barriers: Efficient Dataset Distillation via Inter-Class Feature Compensator
abstract
Dataset distillation has emerged as a technique aiming to condense informative features from large, natural datasets into a compact and synthetic form. While recent advancements have refined this technique, its performance is bottlenecked by the prevailing class-specific synthesis paradigm. Under this paradigm, synthetic data is optimized exclusively for a pre-assigned one-hot label, creating an implicit class barrier in feature condensation. This leads to inefficient utilization of the distillation budget and oversight of inter-class feature distributions, which ultimately limits the effectiveness and efficiency, as demonstrated in our analysis. To overcome these constraints, this paper presents the Inter-class Feature Compensator (INFER), an innovative distillation approach that transcends the class-specific data-label framework widely utilized in current dataset distillation methods. Specifically, INFER leverages a Universal Feature Compensator (UFC) to enhance feature integration across classes, enabling the generation of multiple additional synthetic instances from a single UFC input. This significantly improves the efficiency of the distillation budget. Moreover, INFER enriches inter-class interactions during the distillation, thereby enhancing the effectiveness and generalizability of the distilled data. By allowing for the linear interpolation of labels similar to those in the original dataset, INFER meticulously optimizes the synthetic data and dramatically reduces the size of soft labels in the synthetic dataset to almost zero, establishing a new benchmark for efficiency and effectiveness in dataset distillation. In practice, INFER demonstrates state-of-the-art performance across benchmark datasets. For instance, in the $\texttt{ipc} = 50$ setting on ImageNet-1k with the same compression level, it outperforms SRe2L by 34.5\% using ResNet18. Codes are available at https://github.com/zhangxin-xd/UFC.
Xin Zhang 0092, Jiawei Du 0002, Ping Liu 0004, Joey Tianyi Zhou
ICLR1
2025 AEGIS: Authenticity Evaluation Benchmark for AI-Generated Video Sequences
abstract
Recent advances in AI-generated content have fueled the rise of highly realistic synthetic videos, posing severe risks to societal trust and digital integrity. Existing benchmarks for video authenticity detection typically suffer from limited realism, insufficient scale, and inadequate complexity, failing to effectively evaluate modern vision-language models against sophisticated forgeries. To address this critical gap, we introduce AEGIS, a novel large-scale benchmark explicitly targeting the detection of hyper-realistic and semantically nuanced AI-generated videos. AEGIS comprises over 10,000 rigorously curated real and synthetic videos generated by diverse, state-of-the-art generative models, including Stable Video Diffusion, CogVideoX-5B, KLing, and Sora, encompassing open-source and proprietary architectures. In particular, AEGIS features specially constructed challenging subsets enhanced with GPT-4o-refined prompts, creating unprecedentedly realistic scenarios for rigorous robustness evaluation. Furthermore, we provide multimodal annotations spanning Semantic-Authenticity Descriptions, Motion Features, and Low-level Visual Features, facilitating authenticity detection and supporting downstream tasks such as multimodal fusion and forgery localization. Extensive experiments using advanced vision-language models demonstrate limited detection capabilities on the most challenging subsets of AEGIS, highlighting the dataset's unique complexity and realism beyond the current generalization capabilities of existing models. In essence, AEGIS establishes an indispensable evaluation benchmark, fundamentally advancing research toward developing genuinely robust, reliable, and broadly generalizable video authenticity detection methodologies capable of addressing real-world forgery threats. Our dataset is avaliable on https://huggingface.co/datasets/Clarifiedfish/AEGIS.
Xin Zhang 0092, Joey Tianyi Zhou
ACM Multimedia2
2025 TF-ATM: Training-Free Adaptive Token Merging
Xin Zhang 0092, Weiying Xie, Yunsong Li 0001, Tianlin Hui, Jitao Ma, Leyuan Fang
ACM Multimedia1
2025 Beyond Modality Collapse: Representation Blending for Multimodal Dataset Distillation
abstract
Multimodal Dataset Distillation (MDD) seeks to condense large-scale image-text datasets into compact surrogates while retaining their effectiveness for cross-modal learning. Despite recent progress, existing MDD approaches often suffer from ***Modality Collapse***, characterized by over-concentrated intra-modal representations and enlarged distributional gap across modalities. In this paper, at the first time, we identify this issue as stemming from a fundamental conflict between the over-compression behavior inherent in dataset distillation and the cross-modal supervision imposed by contrastive objectives. To alleviate modality collapse, we introduce **RepBlend**, a novel MDD framework that weakens overdominant cross-modal supervision via representation blending, thereby significantly enhancing intra-modal diversity. Additionally, we observe that current MDD methods impose asymmetric supervision across modalities, resulting in biased optimization. To address this, we propose symmetric projection trajectory matching, which synchronizes the optimization dynamics using modality-specific projection heads, thereby promoting balanced supervision and enhancing cross-modal alignment. Experiments on Flickr-30K and MS-COCO show that RepBlend consistently outperforms prior state-of-the-art MDD methods, achieving significant gains in retrieval performance (e.g., +9.4 IR@10, +6.3 TR@10 under the 100-pair setting) and offering up to 6.7$\times$ distillation speedup.
Xin Zhang 0092, Ziruo Zhang, Jiawei Du 0002, Zuozhu Liu, Joey Tianyi Zhou
NeurIPS1
2025 M³amba: CLIP-Driven Mamba Model for Multi-Modal Remote Sensing Classification
abstract
Multi-modal fusion holds great promise for integrating information from different modalities. However, due to a lack of consideration for modal consistency, existing multi-modal fusion methods in the field of remote sensing still face challenges of incomplete semantic information and low computational efficiency in their fusion designs. Inspired by the observation that the visual language pre-training model CLIP can effectively extract strong semantic information from visual features, we propose M3amba, a novel end-to-end CLIP-driven Mamba model for multi-modal fusion to address these challenges. Specifically, we introduce CLIP-driven modality-specific adapters in the fusion architecture to avoid the bias of understanding specific domains caused by direct inference, making the original CLIP encoder modality-specific perception. This unified framework enables minimal training to achieve a comprehensive semantic understanding of different modalities, thereby guiding cross-modal feature fusion. To further enhance the consistent association between modality mappings, a multi-modal Mamba fusion architecture with linear complexity and a cross-attention module Cross-SS2D are designed, which fully considers effective and efficient information interaction to achieve complete fusion. Extensive experiments have shown that M3amba has an average performance improvement of at least 5.98% compared with the state-of-the-art methods in multi-modal hyperspectral image classification tasks in the remote sensing field, while also demonstrating excellent training efficiency, achieving a double improvement in accuracy and efficiency. The code is released athttps://github.com/kaka-Cao/M3amba.
Mingxiang Cao, Weiying Xie, Xin Zhang 0092, Kai Jiang 0001, Jie Lei 0001, Yunsong Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 Dual-Depth Unified Joint Optimization: Adaptive Curvature-Based Compression
abstract
Model compression methods such as pruning and quantization have been proposed to facilitate the deployment of convolutional neural networks (CNNs) on resource-constrained devices. Existing methods aim to combine the two for simultaneous improvement in compression ratio and runtime efficiency. However, most of the joint methods adopt linear tandem structures. Due to the lack of a unified framework, different optimization directions result in suboptimal solutions, especially when the compression ratio is extremely high. In this paper, we propose a novel adaptive curvature-based compression (ACC) method, which achieves a dual-depth unified joint optimization of pruning and quantization. In the first depth, we unify the pruning and quantization criteria using mean curvature, which leverages the discrete nature of image data and the continuum theory of differential geometry. In the second depth, we replace the traditional training process in the joint pruning-quantization method with curvature-aware knowledge distillation (CKD), unifying the two-stage approach into a simple but powerful parallel step. Our method is effective and interpretable by utilizing inherent properties to promote the understanding of information distribution and the importance of feature maps. Extensive experiments on multiple advanced benchmarks and diverse downstream task datasets have validated the superiority and generalizability of our ACC. Notably, we can achieve a 1.05% Top-1 accuracy improvement over the baseline under an extreme compression ratio of 454.55×, outperforming existing state-of-the-art (SOTA) methods.
Yunsong Li 0001, Xin Zhang 0092, Weiying Xie, Daixun Li, Hangyu Ye, Leyuan Fang
IEEE Trans. Circuits Syst. Video Technol.2
2025 RS-IML: Federated Intrinsic Mask Learning on Remote Sensing Image Interpretation
abstract
Interpreting remote sensing (RS) images plays a crucial role in numerous applications such as environmental monitoring, urban planning, agricultural management and disaster assessment. Nevertheless, remote sensing data is frequently dispersed among various organizations. Privacy concerns and data-sharing limitations make it difficult to utilize large-scale datasets within a centralized training framework. Federated learning (FL) provides a promising approach by facilitating collaborative model training across decentralized data sources, eliminating the need for data centralization. However, the application of FL in RS scenarios is challenging due to the resource-limited edge nodes cannot meet the high demand for computation and memory resources to train deep learning models. Neural network lightweighting techniques have the potential to enhance model efficiency, but existing methods still present significant challenges, such as reliance on initial training of dense models during lightweighting and potential performance degradation after lightweighting. To address these challenges, we propose RS-IML, a novel FL framework for RS Image Interpretation based on Intrinsic Mask Learning without training dense models. RS-IML comprises three key components. First, we introduce the intrinsic dimension of objective landscape that the neural network is projected onto a low-dimensional subnetwork to lightweight neural networks without training dense models. Second, we propose intrinsic dimension parameter averaging to aggregate inconsistent local intrinsic models while suppressing the adverse effects of non-intrinsic parameters between local models. Third, we fine-tune the specific parameters of local intrinsic models to mitigate global intrinsic noise for better performance. Extensive experiments demonstrate the effectiveness of our proposed RS-IML. It achieves a significant improvement in model efficiency during lightweighting compared to existing methods while obtaining superior accuracy.
Hangyu Ye, Weiying Xie, Xin Zhang 0092, Yibing Lu, Yunsong Li 0001, Leyuan Fang
IEEE Trans. Geosci. Remote. Sens.3
2024 Spanning Training Progress: Temporal Dual-Depth Scoring (TDDS) for Enhanced Dataset Pruning
abstract
Dataset pruning aims to construct a coreset capable of achieving performance comparable to the original, full dataset. Most existing dataset pruning methods rely on snapshot-based criteria to identify representative samples, often resulting in poor generalization across various pruning and cross-architecture scenarios. Recent studies have addressed this issue by expanding the scope of training dynamics considered, including factors such as forgetting event and probability change, typically using an averaging approach. However, these works struggle to integrate a broader range of training dynamics without overlooking well-generalized samples, which may not be sufficiently highlighted in an averaging manner. In this study, we propose a novel dataset pruning method termed as Temporal Dual-Depth Scoring (TDDS), to tackle this problem. TDDS utilizes a dual-depth strategy to achieve a balance between incorporating extensive training dynamics and identifying representative samples for dataset pruning. In the first depth, we estimate the series of each sample's individual contributions spanning the training progress, ensuring comprehensive integration of training dynamics. In the second depth, we focus on the variability of the sample-wise contributions identified in the first depth to highlight well- generalized samples. Extensive experiments conducted on CIFAR and ImageNet datasets verify the superiority of TDDS over previous SOTA methods. Specifically on CIFAR-100, our method achieves 54.51% accuracy with only 10% training data, surpassing baselines methods by more than 12.69%. Our codes are available at https://github.com/zhangxin-xd/Dataset-Pruning-TDDS.
Xin Zhang 0092, Jiawei Du 0002, Yunsong Li 0001, Weiying Xie, Joey Tianyi Zhou
CVPR1
2024 Evolution-aware VAriance (EVA) Coreset Selection for Medical Image Classification
abstract
In the medical field, managing high-dimensional massive medical imaging data and performing reliable medical analysis from it is a critical challenge, especially in resource-limited environments such as remote medical facilities and mobile devices. This necessitates effective dataset compression techniques to reduce storage, transmission, and computational cost. However, existing coreset selection methods are primarily designed for natural image datasets, and exhibit doubtful effectiveness when applied to medical image datasets due to challenges such as intra-class variation and inter-class similarity. In this paper, we propose a novel coreset selection strategy termed as Evolution-aware VAriance (EVA), which captures the evolutionary process of model training through a dual-window approach and reflects the fluctuation of sample importance more precisely through variance measurement. Extensive experiments on medical image datasets demonstrate the effectiveness of our strategy over previous SOTA methods, especially at high compression rates. EVA achieves 98.27% accuracy with only 10% training data, compared to 97.20% for the full training set. None of the compared baseline methods can exceed Random at 5% selection rate, while EVA outperforms Random by 5.61%, showcasing its potential for efficient medical image analysis.
Yuxin Hong, Xiao Zhang 0006, Xin Zhang 0092, Joey Tianyi Zhou
ACM Multimedia3
2024 Diversity-Driven Synthesis: Enhancing Dataset Distillation through Directed Weight Adjustment
abstract
The sharp increase in data-related expenses has motivated research into condensing datasets while retaining the most informative features. Dataset distillation has thus recently come to the fore. This paradigm generates synthetic datasets that are representative enough to replace the original dataset in training a neural network. To avoid redundancy in these synthetic datasets, it is crucial that each element contains unique features and remains diverse from others during the synthesis stage. In this paper, we provide a thorough theoretical and empirical analysis of diversity within synthesized datasets. We argue that enhancing diversity can improve the parallelizable yet isolated synthesizing approach. Specifically, we introduce a novel method that employs dynamic and directed weight adjustment techniques to modulate the synthesis process, thereby maximizing the representativeness and diversity of each synthetic instance. Our method ensures that each batch of synthetic data mirrors the characteristics of a large, varying subset of the original dataset. Extensive experiments across multiple datasets, including CIFAR, Tiny-ImageNet, and ImageNet-1K, demonstrate the superior performance of our method, highlighting its effectiveness in producing diverse and representative synthetic datasets with minimal computational expense. Our code is available at https://github.com/AngusDujw/Diversity-Driven-Synthesis.
Jiawei Du 0002, Xin Zhang 0092, Wenxin Huang, Joey Tianyi Zhou
NeurIPS2
2024 Markov-PQ: Joint Pruning-Quantization via Learnable Markov Chain
abstract
Various network compression methods, such as pruning and quantization, have been proposed to synergistically reduce resource requirements. However, existing joint compression works are based on black-box optimization and do not interpret the interaction mechanism between these two compression techniques, leading to a slow and unstable convergence of compression strategy. To address this issue, we present Markov-PQ, the first interpretable pruning-quantization co-compression framework using a Markov Chain. In Markov-PQ, the joint strategy search is modeled as a Markov Chain and decoupled with Bayes Rule into pruning and quantization strategy searching. Specifically, the quantization state accounts for the co-compression state from the last time and is updated by a learnable transition probability matrix. To ensure differentiability, we design a forward-hard and backward-soft quantization. The pruning state is influenced not only by the last co-compression state but also by the concurrent quantization state. In addition, to perceive the current layer-wise bit sensitivity and alleviate the long-tail problem, a complexity-aware regularizer is devised to re-evaluate the filter importance. Extensive experiments demonstrate the superiority of Markov-PQ. For example, with an accuracy loss of only 0.33%, we can achieve a$56.12\times $acceleration for ResNet-18 on ImageNet2012.
Yunsong Li 0001, Xin Zhang 0092, Weiying Xie, Leyuan Fang, Jiawei Du 0002
IEEE Trans. Circuits Syst. Video Technol.2
2024 Distribution-Aware Interactive Attention Network and Large-Scale Cloud Recognition Benchmark on FY-4A Satellite Image
abstract
Accurate cloud recognition and warning are crucial for various applications, including in-flight support, weather forecasting, and climate research. However, recent deep learning algorithms have predominantly focused on detecting cloud regions in satellite imagery, with insufficient attention to the specificity required for accurate cloud recognition. This limitation inspired us to develop the novel FY-4A-Himawari-8 (FYH) dataset, which includes nine distinct cloud categories and uses precise domain adaptation methods to align 70419 image-label pairs (including 110000 train/5500 test$100\times 100$size images) in terms of projection, temporal resolution, and spatial resolution, thereby facilitating the training of supervised deep learning networks. Given the complexity and diversity of cloud formations, we have thoroughly analyzed the challenges inherent to cloud recognition tasks, examining the intricate characteristics and distribution of the data. To effectively address these challenges, we designed a distribution-aware interactive-attention network (DIAnet), which preserves pixel-level details through a high-resolution branch and a parallel multiresolution cross-branch. We also integrated a distribution-aware loss (DAL) to mitigate the imbalance across cloud categories. An interactive attention module (IAM) further enhances the robustness of feature extraction combined with spatial and channel information. Empirical evaluations on the FYH dataset demonstrate that our method outperforms other cloud recognition networks, achieving superior performance in terms of mean intersection over union (mIoU). The code for implementing DIAnet is available athttps://github.com/icey-zhang/DIAnet.
Jie Lei 0001, Weiying Xie, Kai Jiang 0001, Xin Zhang 0092, Mingxiang Cao, Yunsong Li 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Block-Wise Partner Learning for Model Compression
abstract
Despite the great potential of convolutional neural networks (CNNs) in various tasks, the resource-hungry nature greatly hinders their wide deployment in cost-sensitive and low-powered scenarios, especially applications in remote sensing. Existing model pruning approaches, implemented by a "subtraction" operation, impose a performance ceiling on the slimmed model. Self-knowledge distillation (Self-KD) resorts to auxiliary networks that are only active in the training phase for performance improvement. However, the knowledge is holistic and crude, and the learning-based knowledge transfer is mediate and lossy. Here, we propose a novel model-compression method, termed block-wise partner learning (BPL), which comprises "extension" and "fusion" operations and liberates the compressed model from the bondage of baseline. Different from the Self-KD, the proposed BPL creates a partner for each block for performance enhancement in training. For the model to absorb more diverse information, a diversity loss (DL) is designed to evaluate the difference between the original block and the partner. Besides, the partner is fused equivalently instead of being discarded directly. After training, we can simply adopt the fused compressed model that contains the enhancement information of partners but with fewer parameters and less inference cost. As validated using the UC Merced land-use, NWPU-RESISC45, and RSD46-WHU datasets, the BPL demonstrates superiority over other compared model-compression approaches. For example, it attains a substantial floating-point operations (FLOPs) reduction of 73.97% with only 0.24 accuracy (ACC.) loss for ResNet-50 on the UC Merced land-use dataset. The code is available at https://github.com/zhangxin-xd/BPL.
Xin Zhang 0092, Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Kai Jiang 0001, Leyuan Fang, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Filter Pruning via Learned Representation Median in the Frequency Domain
abstract
In this article, we propose a novel filter pruning method for deep learning networks by calculating the learned representation median (RM) in frequency domain (LRMF). In contrast to the existing filter pruning methods that remove relatively unimportant filters in the spatial domain, our newly proposed approach emphasizes the removal of absolutely unimportant filters in the frequency domain. Through extensive experiments, we observed that the criterion for "relative unimportance" cannot be generalized well and that the discrete cosine transform (DCT) domain can eliminate redundancy and emphasize low-frequency representation, which is consistent with the human visual system. Based on these important observations, our LRMF calculates the learned RM in the frequency domain and removes its corresponding filter, since it is absolutely unimportant at each layer. Thanks to this, the time-consuming fine-tuning process is not required in LRMF. The results show that LRMF outperforms state-of-the-art pruning methods. For example, with ResNet110 on CIFAR-10, it achieves a 52.3% FLOPs reduction with an improvement of 0.04% in Top-1 accuracy. With VGG16 on CIFAR-100, it reduces FLOPs by 35.9% while increasing accuracy by 0.5%. On ImageNet, ResNet18 and ResNet50 are accelerated by 53.3% and 52.7% with only 1.76% and 0.8% accuracy loss, respectively. The code is based on PyTorch and is available at https://github.com/zhangxin-xd/LRMF.
Xin Zhang 0092, Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Qian Du 0001
IEEE Trans. Cybern.1
2023 A Model-Driven Deep Mixture Network for Robust Hyperspectral Anomaly Detection
abstract
Hyperspectral anomaly detection (HAD) aims to identify samples with unknown atypical spectra from the background. Deep learning (DL)-based methods, particularly autoencoders (AEs), have proven effective in uncovering the underlying profiles for HAD. However, in real-world applications of hyperspectral images (HSIs), complex background land-covers and anomaly corruptions are common, leading to two issues: 1) A low-dimensional manifold characterized by DL-based HAD methods can only reveal a few underlying variation factors of the background distribution and cannot capture the complex structures behind land-covers of all categories. 2) DL-based HAD methods trained on anomaly-contaminated HSIs tend to overfit specific anomalies, resulting in poor background characterization. To tackle these issues, this study presents a novel and robust framework for HAD called Model-Driven Deep Mixture Network (MDMN) that combines the strengths of model-driven and data-driven approaches while emphasizing interpretability. By assuming that the background, consisting of various land-covers, arises from a mixture of low-dimensional manifolds, the MDMN incorporates a novel deep mixture module to comprehensively characterize the background. This module utilizes a low-dimensional manifold learned by an AE to represent a specific category of background land-covers. To mitigate the impact of anomaly corruptions, the MDMN incorporates a convex relaxation of a sparse constraint, which helps prevent overfitting anomalies. Extensive experimental results demonstrate that the proposed MDMN offers more satisfactory and robust detection performance.
Yunsong Li 0001, Kai Jiang 0001, Weiying Xie, Jie Lei 0001, Xin Zhang 0092, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 Co-Compression via Superior Gene for Remote Sensing Scene Classification
abstract
Convolutional neural networks (CNNs) have been successfully employed in remote sensing image classification because of their robust feature representation for different visual tasks and powerful graphics processing units (GPUs). The attendant problem is that high computational cost and high memory footprint hindering the application of CNNs for remote sensing applications in resource- and time-sensitive situations. Based on practical deployment requirements, we pioneer a pruning-quantization joint learning model compression method for remote sensing image classification, called co-compression via superior gene (CC-SG). An enhanced evolution algorithm (EEA) is adopted as the agent to search a “superior gene,” and immediately following, a director receives the “superior gene” and gives a compression mask and a resource constraint feedback to the agent. The network is eventually compressed and fine-tuned according to the optimal compression mask. Specifically, we introduce gene age and progressive shrinkage mutation rate to EEA and design a fitness function that balances accuracy and resource constraints. As validated using the UC Merced land-use and NWPU-RESISC45 datasets, the proposed CC-SG demonstrated superiority over other compared model compression approaches. For example, CC-SG attained substantial bit operations (BOPs) compression ratio of 40.04 with 0.956% accuracy increase for VGG-16 on UC Merced land-use dataset and 40.00 with 0.203% accuracy increase for ResNet-56 on NWPU-RESISC45 dataset. The code is available athttps://github.com/fanxxxxyi/CC-SG.
Weiying Xie, Xiaoyi Fan 0002, Xin Zhang 0092, Yunsong Li 0001, Min Sheng, Leyuan Fang
IEEE Trans. Geosci. Remote. Sens.3
2023 REAF: Remembering Enhancement and Entropy-Based Asymptotic Forgetting for Filter Pruning
abstract
Neurologically, filter pruning is a procedure of forgetting and remembering recovering. Prevailing methods directly forget less important information from an unrobust baseline at first and expect to minimize the performance sacrifice. However, unsaturated base remembering imposes a ceiling on the slimmed model leading to suboptimal performance. And significantly forgetting at first would cause unrecoverable information loss. Here, we design a novel filter pruning paradigm termed Remembering Enhancement and Entropy-based Asymptotic Forgetting (REAF). Inspired by robustness theory, we first enhance remembering by over-parameterizing baseline with fusible compensatory convolutions which liberates pruned model from the bondage of baseline at no inference cost. Then the collateral implication between original and compensatory filters necessitates a bilateral-collaborated pruning criterion. Specifically, only when the filter has the largest intra-branch distance and its compensatory counterpart has the strongest remembering enhancement power, they are preserved. Further, Ebbinghaus curve-based asymptotic forgetting is proposed to protect the pruned model from unstable learning. The number of pruned filters is increasing asymptotically in the training procedure, which enables the remembering of pretrained weights gradually to be concentrated in the remaining filters. Extensive experiments demonstrate the superiority of REAF over many state-of-the-art (SOTA) methods. For example, REAF removes 47.55% FLOPs and 42.98% parameters of ResNet-50 only with 0.98% TOP-1 accuracy loss on ImageNet. The code is available at https://github.com/zhangxin-xd/REAF.
Xin Zhang 0092, Weiying Xie, Yunsong Li 0001, Kai Jiang 0001, Leyuan Fang
IEEE Trans. Image Process.1
2022 Algorithm/Hardware Codesign for Real-Time On-Satellite CNN-Based Ship Detection in SAR Imagery
abstract
Recently, the convolutional neural network (CNN)-based approach for on-satellite ship detection in synthetic aperture radar (SAR) images has received increasing attention since it does not rely on predefined imagery features and distributions that are required in conventional detection methods. To achieve high detection accuracy, most of the existing CNN-based methods leverage complex off-the-shelf CNN models for optical imagery. Unfortunately, this usually leads to expensive computational cost, which is hard to process in real time using resource-constrained devices deployed in the harsh satellite environment. In this article, we propose OSCAR-RT, the first end-to-end algorithm/hardware codesign framework for real-time on-satellite CNN-based SAR ship detection, which can simultaneously produce an accurate and hardware-friendly CNN model and an ultraefficient field-programmable gate array (FPGA)-based hardware accelerator that can be deployed on satellites. With the real-time on-satellite processing speed in mind, we start from a state-of-the-art compact CNN model for optical imagery. To eliminate the sharp decrease in the detection accuracy for SAR imagery, we analyze the discrepancy between the SAR domain and optical domain and propose to adapt the model by adjusting the output feature size to better detect relatively smaller objects in SAR imagery. To improve the detection speed, we propose to develop a fully pipelined interlayer streaming accelerator architecture, where all the layers of the CNN model can be concurrently processed using on-chip FPGA resources. To achieve this architecture, we first propose a hardware-guided, progressive, and structural pruning strategy, which is guided by our modeled hardware metrics and applies state-of-the-art coarse-grained and fine-grained filter pruning as well as mixed-precision quantization techniques. Moreover, to improve the reusability and portability of the hardware accelerator design, we develop a library of highly optimized CNN components in high-level synthesis, together with their performance and resource models. Finally, we map the pruned CNN model onto these hardware library components in a fully pipelined interlayer streaming fashion, by adjusting their parallelism factors to balance the execution of each layer and fit into the resource constraint. Experimental results using the adapted MobileNetV1, MobileNetV2, and SqueezeNet models on the widely used SAR ship detection dataset (SSDD) demonstrate the effectiveness of OSCAR-RT; for the MobileNetV1 model, it achieves an average precision of 94%, a detection speed of 652 frames/s on the Xilinx VC709 FPGA evaluation board while consuming about 5.8-W power.
Geng Yang 0001, Jie Lei 0001, Weiying Xie, Zhenman Fang, Yunsong Li 0001, Xin Zhang 0092
IEEE Trans. Geosci. Remote. Sens.7
2022 Rank-Aware Generative Adversarial Network for Hyperspectral Band Selection
abstract
Traditional clustering-based band selection (BS) methods treat each band as individuals, and selection is conducted by enlarging the difference between clusters, which leads to the loss of band interaction and information saliency evaluation. In this article, we propose a BS method named rank-aware generative adversarial network (R-GAN) to address these problems. First, centralized reference feature extraction (FE) with GAN aids R-GAN to combine interpretability and interband relevance. Then, the reference feature is refined with the saliency estimation provided by the rank-aware strategy. According to data characteristics, there are two versions of rank computation including tensor and matrix. Finally, the structural similarity index measurement (SSIM) maps the saliency to the original data space to obtain the final BS result. Extensive comparison experiments with popular existing BS approaches on five hyperspectral images (HSIs) datasets show that the proposed R-GAN can address spectral saliency effectively and select more informative band subsets, which outperforms other competitors for both detection and classification tasks. For example, on the SD-1 dataset, the ten bands selected by R-GAN achieve 0.982 ± 0.003 with an improvement of 13.7% in the area under the curve (AUC) value of anomaly detection performance. The peaked accuracy surpasses the baseline by 0.46% for the classification on the PaviaU dataset.
Xin Zhang 0092, Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Qian Du 0001, Geng Yang 0001
IEEE Trans. Geosci. Remote. Sens.1
2021 Weakly Supervised Low-Rank Representation for Hyperspectral Anomaly Detection
abstract
In this article, we propose a weakly supervised low-rank representation (WSLRR) method for hyperspectral anomaly detection (HAD), which formulates deep learning-based HAD into a low-lank optimization problem not only characterizing the complex and diverse background in real HSIs but also obtaining relatively strong supervision information. Different from the existing unsupervised and supervised methods, we first model the background in a weakly supervised manner, which achieves better performance without prior information and is not restrained by richly correct annotation. Considering reconstruction biases introduced by the weakly supervised estimation, LRR is an effective method for further exploring the intricate background structures. Instead of directly applying the conventional LRR approaches, a dictionary-based LRR, including both observed training data and hidden learned data drawn by the background estimation model, is proposed. Finally, the derived low-rank part and sparse part and the result of the initial detection work together to achieve anomaly detection. Comparative analyses validate that the proposed WSLRR method presents superior detection performance compared with the state-of-the-art methods.
Weiying Xie, Xin Zhang 0092, Yunsong Li 0001, Jie Lei 0001, Jiaojiao Li 0001, Qian Du 0001
IEEE Trans. Cybern.2