VLDB 2026 Research / reviewers in the wild / expert
Yimin Fu
dblp:332/5961
· DBLP profile ↗
13ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0003-2070-995XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semantic-driven representation disentanglement for robust open-set retrieval of cross-modal remote sensing images
Lizhuo Liu, Peiyuan Ma, Yimin Fu, Zhunga Liu |
Neurocomputing | 3 |
| 2026 | Scattering-guided class-irrelevant filtering for adversarially robust SAR automatic target recognition
Zhunga Liu, Jialin Lyu, Yimin Fu |
Signal Process. | 3 |
| 2026 | Enhancing the Transferability of Adversarial Attacks for Remote Sensing Scene Classification via Asymmetric Frequency MixingabstractAdversarial attacks serve as an efficient approach to reveal potential weaknesses of deep neural networks (DNNs). In practice, models are often deployed in black-box settings, where access to parameters and architectures is restricted. This imposes requirements on the transferability of adversarial examples crafted on the source model, which is typically enhanced by simulating diverse input patterns through input transformations during gradient computation. However, given the lack of a clear foreground-background distinction, scene classification in remote sensing images typically relies on global contextual understanding, which exhibits substantial inter-model variation. In addition, the diversity of geospatial structures further amplifies the discrepancies across different model architectures. These factors jointly impede the enhancement of adversarial transferability due to overfitting to model-specific information. To address this challenge, we propose asymmetric frequency mixing (AFM) to improve adversarial transferability for remote sensing scene classification. Specifically, context patterns are simulated by selectively employing an asymmetrical mixing operation in the frequency domain. For images of the same scene category, low-frequency components decomposed by the discrete wavelet transform are mixed to enrich semantic-related diversity. Meanwhile, high-frequency components are mixed in a class-agnostic fashion to suppress model-specific bias toward background texture information. Moreover, a block shuffling operation is imposed on the reconstructed images to reduce inter-model discrepancies in capturing geospatial structures. Extensive experiments on the UCMerced Land Use and SIRI-WHU datasets demonstrate that the proposed method achieves state-of-the-art attack performance across different model architectures. Lizhuo Liu, Yuefeng Bai, Yimin Fu |
IEEE Signal Process. Lett. | 3 |
| 2026 | Hierarchical Structure Dependency Whitening for Single-Domain Generalized Infrared Small Target DetectionabstractExisting infrared small target detection (IRSTD) methods typically assume that training and testing data share the same distribution. However, this assumption often fails in real-world applications due to environmental and sensor-induced variations, resulting in significant performance degradation caused by domain shifts. Besides, the inherently low signal-to-clutter ratio of targets in infrared images further impedes the extraction of underlying target information, increasing the risk of overfitting to domain-specific patterns. This severely constrains the generalizability of knowledge learned from source domains, particularly when training is confined to a single source domain due to the high cost of data annotation. To solve this problem, we propose hierarchical structure dependency whitening (HSDW) for single-domain generalized IRSTD. Specifically, we characterize domain discrepancies in infrared images as differences in structural information. Building upon this point, we employ feature whitening to mitigate the dependency on domain-specific structure information, whose distribution is diversely simulated by a dual-branch nonlinear transformation module. Moreover, we adopt a hierarchical suppression mechanism to alleviate the structural biases across multiple decoding stages, thereby facilitating more generalized target understanding across domains. Extensive experiments on three public IRSTD datasets demonstrate that our method achieves state-of-the-art performance. Lizhuo Liu, Songbo Wang, Yimin Fu |
IEEE Signal Process. Lett. | 3 |
| 2026 | Adaptive Mixture-of-Experts Distillation for Cross-Satellite Generalizable Incremental Remote Sensing Scene ClassificationabstractIncremental learning aims to continuously acquire new knowledge from data streams while maintaining previously learned knowledge. Existing incremental learning methods typically assume that the training (source domain) and testing (target domain) data are identically distributed. However, differences in sensor parameters and imaging conditions inevitably lead to distribution gaps between data collected from different satellites (domains). The ensuing domain shift problem substantially impairs the generalization of continuously learned knowledge from source domains to unseen ones. To tackle this problem, we propose adaptive mixture-of-experts distillation (AMoED) for cross-satellite generalizable incremental remote sensing scene classification (CSGIRSSC). Specifically, AMoED adopts a high-level semantic learning pipeline, in which new knowledge is acquired through the coordinated guidance of multiple domain-specific experts, rather than directly from raw data. This pipeline prevents the model from being exposed to large volumes of newly emerging data, thereby alleviating the erasure of previous knowledge when adapting to new data distributions. Besides, the adaptive mixture of domain-specific experts facilitates the formation of universal class concepts, which exhibit strong generalizability across different domains. During the learning process, an equi-partite subset is constructed for knowledge acquisition and consolidation, accompanied by a shallow style-mixing operation to mitigate the interference of domain discrepancies. Extensive experiments are conducted on four remote sensing scene classification datasets, and the proposed method consistently achieves state-of-the-art performance across various scenarios and settings. The code is released at https://github.com/fuyimin96/AMoED. Yimin Fu, Runqing Yang, Zhunga Liu, Michael Kwok-Po Ng |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Reason and Discovery: A New Paradigm for Open Set RecognitionabstractOpen set recognition (OSR) effectively enhances the reliability of pattern recognition systems by accurately identifying samples of unknown classes. However, the decision-making process in most existing OSR methods adheres to an ill-considered pipeline, where classification probabilities are inferred directly from overall feature representations, neglecting the reasoning about inherent relations. Besides, the handling of identified unknown samples is typically restricted to the assignment of a generic "unknown" class label but fails to explore underlying category information. To tackle the above challenges, we propose a new paradigm for OSR, entitled Reason and Discovery (RAD), which comprises two main modules: the Reason Module and the Discovery Module. Specifically, in the Reason Module, the distinction between known and unknown is performed from the perspective of reasoning the matching relations between topological information and appearance characteristics of discriminative regions. Then, the mixture and recombination of relation representations across classes are employed to provide diverse estimations of unknown distribution, thereby recalibrating OSR decision boundaries. Moreover, in the Discovery Module, the identified unknown samples are semantically grouped through a biased deep clustering process for discovering novel category information. Experimental results on various datasets indicate that the proposed method can achieve outstanding OSR performance and good novel category discovery efficacy. Yimin Fu, Zhunga Liu, Jialin Lyu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | A Unified SAM-Guided Self-Prompt Learning Framework for Infrared Small Target DetectionabstractInfrared small target detection (ISTD) aims to precisely capture the location and morphology of small targets under all-weather conditions. Compared with generic objects, infrared targets in remote fields of view are smaller in size and exhibit lower signal-to-clutter ratios. This poses a significant challenge in simultaneously preserving low-level target details and understanding high-level contextual semantics, forcing a trade-off between reducing miss detection and suppressing false alarms. In addition, most existing ISTD methods are designed for specific target types under certain infrared platforms, rather than as a unified framework broadly applicable across diverse infrared sensing scenarios. To address these challenges, we propose a unified self-prompt learning framework for ISTD under the guidance of the Segment Anything Model (SAM). Specifically, the model is incorporated with SAM in the encoding stage through a consult-guide manner, adapting the general knowledge to facilitate task-specific contextual understanding. Then, shallow-layer features are employed to generate self-derived prompts, which bidirectionally interact with encoded latent representations to complement subtle low-level details. Moreover, the semantic inconsistency during resolution recovery is mitigated by integrating a mutual calibration module into skip connections, ensuring coherent spatial-semantic fusion. Extensive experiments are conducted on four public ISTD datasets, and the results demonstrate that the proposed method consistently achieves superior performance across different infrared sensing platforms and target types. The code is released at https://github.com/fuyimin96/SAM-SPL. Yimin Fu, Jialin Lyu, Peiyuan Ma, Zhunga Liu, Michael Kwok-Po Ng |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Logit prototype learning with active multimodal representation for robust open-set recognition
Yimin Fu, Zhunga Liu |
Sci. China Inf. Sci. | 1 |
| 2024 | Transferable Adversarial Attacks for Remote Sensing Object Recognition via Spatial- Frequency Co-TransformationabstractAdversarial attacks serve as an efficient approach to investigating model robustness, providing insights into internal weaknesses. In real-world applications, the model deployment typically adheres to a black-box setting, necessitating the transferability of adversarial examples crafted on a source model to others. Attack methods in the general computer vision field often employ global input transformations in individual spatial or frequency domains to boost adversarial transferability. However, the recognition of remote sensing objects primarily relies on target-related discriminative regions, whose determination exhibits significant model specificity. Besides, the coupling between objects and background further exacerbates the gap between models. Consequently, the transferability of adversarial examples is limited due to overfitting to the source model. To tackle this problem, we propose a spatial-frequency co-transformation (SFCoT) to improve adversarial transferability for remote sensing object recognition. Specifically, the input image is decomposed into blocks and components in the spatial and frequency domains, respectively. Then, a selective frequency transformation (SFT) is performed on the low-frequency components to narrow intermodel gaps. Subsequently, modular spatial transformations (MSTs) are adopted in blocks to enhance target-related diversity. Incorporating transformations across domains effectively mitigates the overfitting to model-specific information, leading to better adversarial transferability. Extensive experiments have been conducted on FGSCR-42 and MTARSI datasets, and the results demonstrate that the proposed method achieves state-of-the-art performance across various model architectures. The code will be released athttps://github.com/fuyimin96/SFCoT. Yimin Fu, Zhunga Liu, Jialin Lyu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Matching Intentions for Discourse Parsing in Multi-party Dialogues
Tiezheng Mao, Jialing Fu, Osamu Yoshie, Yimin Fu, Zhuyun Li |
IEA/AIE (2) | 4 |
| 2023 | Orientational Distribution Learning With Hierarchical Spatial Attention for Open Set RecognitionabstractOpen set recognition (OSR) aims to correctly recognize the known classes and reject the unknown classes for increasing the reliability of the recognition system. The distance-based loss is often employed in deep neural networks-based OSR methods to constrain the latent representation of known classes. However, the optimization is usually conducted using the nondirectional euclidean distance in a single feature space without considering the potential impact of spatial distribution. To address this problem, we propose orientational distribution learning (ODL) with hierarchical spatial attention for OSR. In ODL, the spatial distribution of feature representation is optimized orientationally to increase the discriminability of decision boundaries for open set recognition. Then, a hierarchical spatial attention mechanism is proposed to assist ODL to capture the global distribution dependencies in the feature space based on spatial relationships. Moreover, a composite feature space is constructed to integrate the features from different layers and different mapping approaches, and it can well enrich the representation information. Finally, a decision-level fusion method is developed to combine the composite feature space and the naive feature space for producing a more comprehensive classification result. The effectiveness of ODL has been demonstrated on various benchmark datasets, and ODL achieves state-of-the-art performance. Zhunga Liu, Yimin Fu, Quan Pan 0001, Zuowei Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Progressive Learning Vision Transformer for Open Set Recognition of Fine-Grained Objects in Remote Sensing ImagesabstractOpen set recognition (OSR) aims to classify known classes and recognize unknown classes simultaneously. Existing OSR methods have primarily focused on learning decision boundaries based on overall feature representations, and have achieved good performance on various coarse-grained image datasets. However, the overall feature representations of objects in fine-grained image datasets are highly similar, making it difficult to distinguish between known and unknown classes by overall feature-based decision boundaries. To address this problem, we propose a progressive learning vision transformer (PLViT) with a coarse-to-fine optimization strategy. In PLViT, the overall feature representations are first optimized in the distance space to learn the initial decision boundaries. Then, a context-aware patch selection module is designed to locate the discriminative part regions. Afterwards, the multi-layer representations of each selected patch are aggregated according to the self-attention weights, and input into the last transformer layer to extract local feature representations. Finally, overall and local feature representations are adaptively fused and optimized in the angular space to further refine the decision boundaries. Experimental results on four fine-grained remote sensing object recognition datasets show that PLViT outperforms state-of-the-art methods. Yimin Fu, Zhunga Liu, Zuowei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Adaptive Open Set Recognition with Multi-modal Joint Metric Learning
Yimin Fu, Zhunga Liu, Yanbo Yang 0001, Hua Lan |
PRCV (1) | 1 |