Yimin Xu

dblp:168/7699 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
13since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2027 Unified-width adaptive dynamic network for all-in-one image restoration
Yimin Xu, Chunmei Yuan, Yunshan Zhong, Fei Chao 0001
Inf. Sci.1
2026 HL-CMR: Hypergraph Learning for Cross-Modal Retrieval
abstract
Cross-modal retrieval is a fundamental task in multimedia understanding, aimed at querying samples with similar semantics in one modality (e.g., text) using another modality (e.g., image). Existing methods merely focus on point-to-point comparisons between individual samples, while overlooking the widely present many-to-many structural relationships in real-world scenarios. However, the many-to-many relationships formed by multiple samples sharing similar semantics are crucial for effectively achieving semantic alignment and accurately constructing shared semantic representations. To address this, we propose a novel hypergraph-based cross-modal retrieval approach, which explicitly establishes many-to-many associations between multiple samples using a label-driven hypergraph construction mechanism, combined with differentiated hyperedge weighting. Additionally, to avoid the limitation of information interaction direction imposed by traditional unidirectional cross-attention mechanisms, we design a bidirectional cross-attention structure, with image and text as separate query sources, to achieve symmetric semantic enhancement between modalities. The resulting joint image-text representations are then mapped as hypergraph vertices, further enhancing the model's ability to align cross-modal semantics. Since constructing a global hypergraph on a large-scale sample set would incur high computational cost, we introduce global label co-occurrence frequency to supervise the batch-level hypergraph construction, enhancing the local graph's ability to capture global semantics. Experimental results show that our model outperforms existing state-of-the-art methods on three benchmark cross-modal retrieval datasets.
Yimin Xu, Rui Zhou 0001
WWW3
2026 Dynamic trajectory diffusion model for all-in-one image restoration
Yimin Xu, Yunshan Zhong, Fei Chao 0001
Expert Syst. Appl.1
2025 Initial Results of a W/D-Band Millimeter-Wave Radiometer on Unmanned Aerial Vehicles (UAVs)
abstract
In recent years, the application of passive millimeter-wave imaging (PMWI) in target detection, reconnaissance, and surveillance has drawn our attention again, especially passive interferometric millimeter-wave imaging (PIMI). However, PIMI suffers with high hardware and signal processing complexity, and passive interferometric millimeter-wave sensors (PIMSs) are very expensive. In this letter, a W/D-band millimeter-wave radiometer on unmanned aerial vehicles (UAVs) is first developed for target detection, reconnaissance, and surveillance, especially for ship detection. The W/D-band millimeter-wave radiometer has the advantages of low hardware and signal processing complexity and low cost. The W/D-band millimeter-wave radiometer is introduced, and a proof prototype of the W/D-band millimeter-wave radiometer is manufactured. Numerical simulations are performed to analyze the brightness temperature (TB) characteristic of the metallic ships. Outdoor experiments are performed to assess the feasibility of the metallic target detected by the W/D-band millimeter-wave radiometer on a UAV. Initial experimental results have demonstrated that metallic targets can be detected by the W/D-band millimeter-wave radiometer from UAVs, as expected.
Qi Yang 0002, Hailiang Lu 0001, Yimin Xu, Junqi Fu, Wenchao Zheng 0001, Xiaokang Mei, Hongqiang Wang 0001
IEEE Geosci. Remote. Sens. Lett.3
2025 ARF: Arbitrary Routing Framework for All-in-One Image Restoration
abstract
All-in-one image restoration methods, as opposed to conventional image restoration methods, reconstruct images impaired by various degradations within a unified model, eliminating the need for separate network parameters for each task. However, current all-in-one image restoration approaches tackle various types of image degradation using an identical underlying model, neglecting the inherent variability in complexity across different image restoration tasks, resulting in inefficient allocation of computational resources. To address this limitation, this article introduces the arbitrary routing framework (ARF), designed to effectively assess the difficulty of image restoration tasks and identify the most suitable network structure based on these complexities. This framework can be integrated with existing all-in-one image restoration models, enabling efficient inference by activating various proportions of the entire network, that is subnetworks, based on their task-specific complexities. More specifically, the ARF comprises two principal components: 1) the arbitrary routing backbone (ARB) and 2) a task-specific neural architecture search (T-NAS). The ARB incorporates a routing layer between consecutive convolutional groups, offering a wide array of potential subnetwork configurations while adding only negligible extra parameters. Concurrently, T-NAS autonomously identifies the most effective subnetworks for each image restoration task, optimizing both performance and efficiency through an efficiency-aware reward function. Comprehensive experiments across various image restoration tasks demonstrate that the ARF significantly improves performance metrics, that is, an increase of 0.31 in reconstruction PSNR, while also achieving a notable reduction in computational demands by 37.1% compared with the benchmark AirNet method. The code has been made available in the supplementary materials.
Yimin Xu, Nanxi Gao, Yunshan Zhong, Fei Chao 0001, Rongrong Ji
IEEE Trans. Cybern.1
2024 An efficient blur kernel estimation method for blind image Super-Resolution
Yimin Xu, Nanxi Gao, Fei Chao 0001, Rongrong Ji
Pattern Recognit.1
2024 Shadow-aware dynamic convolution for shadow removal
Yimin Xu, Mingbao Lin, Fei Chao 0001, Rongrong Ji
Pattern Recognit.1
2022 DDoS Attack Detection Combining Time Series-based Multi-dimensional Sketch and Machine Learning
abstract
Machine learning-based DDoS attack detection methods are mostly implemented at the packet level with expensive computational time costs, and the space cost of those sketch-based detection methods is uncertain. This paper proposes a two-stage DDoS attack detection algorithm combining time series-based multi-dimensional sketch and machine learning technologies. Besides packet numbers, total lengths, and protocols, we construct the time series-based multi-dimensional sketch with limited space cost by storing elephant flow information with the Boyer-Moore voting algorithm and hash index. For the first stage of detection, we adopt CNN to generate sketch-level DDoS attack detection results from the time series-based multi-dimensional sketch. For the sketch with potential DDoS attacks, we use RNN with flow information extracted from the sketch to implement flow-level DDoS attack detection in the second stage. Experimental results show that not only is the detection accuracy of our proposed method much close to that of packet-level DDoS attack detection methods based on machine learning, but also the computational time cost of our method is much smaller with regard to the number of machine learning operations.
Yanchao Sun, Yuanfeng Han, Mingsong Chen 0001, Shui Yu 0001, Yimin Xu
APNOMS6
2022 Hyperspectral Image Classification With Multiattention Fusion Network
abstract
Hyperspectral image (HSI) has hundreds of continuous bands that contain a lot of redundant information. Besides, a spatial patch of a hyperspectral cube often contains some pixels different from the center pixel category, which are usually called interference pixels. The existence of such interference pixels has a negative effect on extracting more discriminative information. Therefore, in this letter, a multiattention fusion network (MAFN) for HSI classification is proposed. Compared with the current state-of-the-art methods, MAFN uses band attention module (BAM) and spatial attention module (SAM), respectively, to alleviate the influence of redundant bands and interfering pixels. In this way, MAFN realizes feature reuse and obtains complementary information from different levels by combining multiattention and multilevel fusion mechanisms, which can extract more representative features. Experiments were conducted on two public HSI data sets to demonstrate the effectiveness of MAFN. Our source code is available athttps://github.com/Li-ZK/MAFN-2021.
Zhaokui Li, Xiaodan Zhao, Yimin Xu, Wei Li 0032, Lin Zhai, Zhuoqun Fang, Xiangbin Shi
IEEE Geosci. Remote. Sens. Lett.3
2022 Small-Scale Linguistic Steganalysis for Multi-Concealed Scenarios
abstract
Recently, due to the considerable feature expression ability of neural networks, deep linguistic steganalysis methods have been greatly developed. However, there are still two issues that need to be ameliorated. First, the prevailing linguistic steganalysis methods rely heavily on massive training data, which is labor-intensive and time-consuming. Second, these methods implement steganalysis only in different weak-concealed scenarios, the stego texts in each of which have only a single language style and payload. But in practice, the intercepted network samples are probably the mixture of the stego texts that possess different language styles and payloads, in which the semantic spatial distribution may be more chaotic than that in weak-concealed scenarios, thus making steganalysis more difficult. To address the above issues, a novel linguistic steganalysis method is proposed in this letter. First, the pre-trained BERT language model is constructed as an embedder to compensate for the shortage of data. Then, in addition to learning local and global semantic features, a feature interaction module is designed for exploring mutual effects between them. Furthermore, besides the typical cross-entropy loss, triplet loss is also introduced for the model training. In this way, the proposed method can refine more comprehensive and discriminative deep features in the intricate semantic space. The performance of the proposed method is compared with the representative linguistic steganalysis methods on datasets of different scales, and the experimental results reveal the superiority of the proposed method.
Yimin Xu, Tengyun Zhao, Ping Zhong 0003
IEEE Signal Process. Lett.1
2022 Deep Cross-Domain Few-Shot Learning for Hyperspectral Image Classification
abstract
One of the challenges in hyperspectral image (HSI) classification is that there are limited labeled samples to train a classifier for very high-dimensional data. In practical applications, we often encounter an HSI domain (called target domain) with very few labeled data, while another HSI domain (called source domain) may have enough labeled data. Classes between the two domains may not be the same. This article attempts to use source class data to help classify the target classes, including the same and new unseen classes. To address this classification paradigm, a meta-learning paradigm for few-shot learning (FSL) is usually adopted. However, existing FSL methods do not account for domain shift between source and target domain. To solve the FSL problem under domain shift, a novel deep cross-domain few-shot learning (DCFSL) method is proposed. For the first time, DCFSL tackles FSL and domain adaptation issues in a unified framework. Specifically, a conditional adversarial domain adaptation strategy is utilized to overcome domain shift, which can achieve domain distribution alignment. In addition, FSL is executed in source and target classes at the same time, which can not only discover transferable knowledge in the source classes but also learn a discriminative embedding model to the target classes. Experiments conducted on four public HSI data sets demonstrate that DCFSL outperforms the existing FSL methods and deep learning methods for HSI classification. Our source code is available athttps://github.com/Li-ZK/DCFSL-2021.
Zhaokui Li, Yushi Chen 0002, Yimin Xu, Wei Li 0032, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Dual-Channel Residual Network for Hyperspectral Image Classification With Noisy Labels
abstract
Hyperspectral image (HSI) classification has drawn increasing attention recently. However, it suffers from noisy labels that may occur during field surveys due to a lack of prior information or human mistakes. To address this issue, this article proposes a novel dual-channel residual network (DCRN) to resolve HSI classification with noisy labels. Currently, the influence of noisy labels is reduced by simply detecting and removing those anomalous samples. Different from such a specifically designed noise cleansing method, DCRN is easy to implement but highly effective. It enhances its model robustness to noisy labels to a great extent by employing a novel dual-channel structure and a noise-robust loss function. In this way, DCRN can mitigate influence from noisy labels while fully utilizing useful information from mislabeled samples for augmented training. Experiments are conducted on several hyperspectral data sets with manually generated noisy labels to demonstrate its excellent performance. The code is available athttps://github.com/Li-ZK/DCRN-2021.
Yimin Xu, Zhaokui Li, Wei Li 0032, Qian Du 0001, Cuiwei Liu, Zhuoqun Fang, Lin Zhai
IEEE Trans. Geosci. Remote. Sens.1
2021 Robust multiview feature selection via view weighted
Ping Zhong 0003, Yimin Xu, Liran Yang
Multim. Tools Appl.3