Yanxia Wu 0001

dblp:88/5327-1 · DBLP profile ↗
← Back
20ranked-venue papers
1as first author
17since 2021 · last 2025
0000-0001-8384-9234ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Systems, architecture and hardware · 5 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 SAG-KeyNet: Scale-Adaptive Keypoint Gaussian Heatmap Regression Network for Oriented SAR Ship Detection
abstract
Ship detection in synthetic aperture radar (SAR) images is crucial for maritime monitoring. Anchor-based deep learning methods are prone to performance degradation due to redundant anchors. In contrast, keypoint-based anchor-free methods typically rely on Gaussian kernels with fixed standard deviations to generate heatmaps for detecting keypoints, reducing accuracy in multi-scale ship detection. To this end, we propose the scale-adaptive keypoint Gaussian heatmap regression network (SAG-KeyNet), an anchor-free framework for oriented SAR ship detection. The proposed scale-adaptive Gaussian heatmap regression (SAGHR) dynamically adjusts the Gaussian kernel’s standard deviation, ensuring precise keypoint localization across varying ship scales. Additionally, the ViTAE with enhanced deformation (ViTAED) backbone is utilized to extract ship features, enhancing adaptability to variations in ship posture and shape. Finally, the large kernel attention-enhanced PAFPN (LKA-PAFPN) is designed to improve multi-scale feature fusion across spatial and channel dimensions. Extensive experimental results on three datasets demonstrate the effectiveness and generalization of SAG-KeyNet. Notably, an average precision (AP50) of 73.8% is achieved on the RSDD-SAR dataset for inshore ship detection.
Xu Wang 0041, Yanxia Wu 0001, Ye Yuan 0011, Zhirou Ma
ICME3
2025 TriDet-MLLM: Triple-Feature Fusion Prompt Learning for AI-Generated Image Detection
abstract
The proliferation and advancement of generative AI tools blur the boundaries between authentic and AI-synthesized imagery, sparking widespread public concern about visual content authenticity. While existing AI-generated image detection approaches have demonstrated notable achievements, Multimodal Large Language Models (MLLMs) remain relatively underutilized in this field. As MLLMs show strong abilities in understanding image features and can provide high-quality textual analysis, substantial potential within MLLM itself remains well explored. In this work, we propose a novel triple-feature fusion prompt learning framework to effectively stimulate the potential of the MLLMs in detecting AI-generated images. We build a prompt structure that allows MLLM to analyze the inauthenticity of AI-generated images from different perspectives, both locally and globally, and ask questions on the generated image dataset to obtain open-ended answers. We then design a triple fusion encoder that combines semantic features, structured knowledge representation and fine-grained modifications. Through a hierarchical feature induction mechanism, we were able to construct multidimensional feature sets with interpretive properties. Then, we compose a set of prompts based on the obtained feature sets to guide the MLLM to discriminate the input images under global and partial strategies. The test result shows prominent improvement compared to direct inquiry approaches, suggesting that MLLMs demonstrate substantial potential in no need for fine-tuning in AI-generated image detection. This work will drive future research toward prompt-based methods, further expanding the capabilities of MLLMs.
Rongsheng Li, Yanxia Wu 0001, Qiao Tian 0002, Shang Feng
MMAsia5
2025 Thread-sensitive fuzzing for concurrency bug detection
Yanxia Wu 0001, Jibin Dong, Ruize Hong
Comput. Secur.3
2025 Recursive Hybrid Compression for Sparse Matrix-Vector Multiplication on GPU
abstract
ABSTRACT Sparse Matrix‐Vector Multiplication (SpMV) is a fundamental operation in scientific computing, machine learning, and data analysis. The performance of SpMV on GPUs is crucial for accelerating various applications. However, the efficiency of SpMV on GPUs is significantly affected by irregular memory access patterns, high memory bandwidth requirements, and insufficient exploitation of parallelism. In this paper, we propose a Recursive Hybrid Compression (RHC) method to address these challenges. RHC begins by splitting the initial matrix into two portions: an Ellpack (ELL) portion and a Coordinate (COO) portion. This partitioning is followed by further recursive division of the COO portion into additional ELL and COO portions, continuing this process until predefined termination criteria, based on a percentage threshold of the number of nonzero elements, are met. Additionally, we introduce a dynamic partitioning method to determine the optimal threshold for partitioning the matrix into ELL and COO portions based on the distribution of nonzero elements and the memory footprint. We develop the RHC algorithm to fully exploit the advantages of the ELL kernel on GPUs and achieve high thread‐level parallelism. We evaluated our proposed method on two different NVIDIA GPUs: the GeForce RTX 2080 Ti and the A100, using a set of sparse matrices from the SuiteSparse Matrix Collection. We compare RHC with NVIDIA's cuSPARSE library and three state‐of‐the‐art methods: SELLP, MergeBase, and BalanceCSR. RHC achieves average speedups of 2.13, 1.13, 1.87, and 1.27 over cuSPARSE, SELLP, MergeBase, and BalanceCSR, respectively.
Zhixiang Zhao, Yanxia Wu 0001, Guoyin Zhang, Yiqing Yang, Ruize Hong
Concurr. Comput. Pract. Exp.2
2025 Bi-granularity balance learning for long-tailed image classification
Ning Ren, Xiaosong Li 0003, Yanxia Wu 0001
Comput. Vis. Image Underst.3
2024 Unpaired image despeckling based on adversarial speckle generation
abstract
Speckle suppression is a critical step in a coherent imaging system. Currently, deep learning-based despeckling methods can be categorized as supervised and self-supervised. The former is unsuitable for this task due to the unavailability of speckle-free images, while the latter may fall short in despeckling performance due to the lack of prior knowledge constraints. To this end, we propose an unpaired image despeckling method based on adversarial speckle generation (UID-ASG), eliminating the need for manual design of image and speckle priors.UID-ASG emphasizes the joint distribution of speckled and clean images. It integrates self-supervised blind spot learning with adversarial learning to precisely simulate speckle distribution and utilizes unpaired speckle-clean image pairs for training. By merging the advantages of supervised and self-supervised methods, UID-ASG significantly enhances its despeckling performance. Experimental results demonstrate that UID-ASG outperforms other state-of-the-art methods in despeckling effectiveness.
Xu Wang 0041, Yanxia Wu 0001, Ye Yuan 0011
ICME2
2024 Explicitly learning augmentation invariance for image classification by Consistent Augmentation
Xiaosong Li 0003, Yanxia Wu 0001, Chuheng Tang, Lidan Zhang
Eng. Appl. Artif. Intell.2
2024 Balanced self-distillation for long-tailed recognition
Ning Ren, Xiaosong Li 0003, Yanxia Wu 0001
Knowl. Based Syst.3
2024 Split-bucket partition (SBP): a novel execution model for top-K and selection algorithms on GPUs
abstract
Abstract Top-K and selection operations are critical in data processing and analysis, and their efficient implementation on GPUs is increasingly important due to the growing demands of data analysis. Existing methods, primarily relying on the bucket partition execution model, encounter challenges such as uneven bucket distribution and latency in merging processes. To address these issues, we introduce a novel Split-Bucket Partition (SBP) execution model that specifically addresses these challenges. Additionally, we propose task and control flow optimizations targeted at top-K and selection algorithms, which further contribute to performance improvements. Our optimized algorithms significantly outperform existing approaches, delivering performance gains of up to $$2.3$$ 2.3 times and $$1.6$$ 1.6 times for different bucket partitioning rules. Our algorithms show robust performance improvements in non-uniform data scenarios, with gains ranging from $$1.9$$ 1.9 times to $$15.5$$ 15.5 times. However, it should be noted that the SBP model has limitations related to shared memory and register utilization, potentially impacting performance. Tests on TU102 and A100 GPU architectures validate the effectiveness of our approach, achieving a maximum speedup of $$2.9$$ 2.9 times. The study suggests that while the SBP model is effective for top-K and selection algorithms, it also holds promise for other computational tasks, setting the stage for future research.
Yiqing Yang, Guoyin Zhang, Yanxia Wu 0001, Zhixiang Zhao
J. Supercomput.3
2024 Block-wise dynamic mixed-precision for sparse matrix-vector multiplication on GPUs
abstract
Abstract Sparse matrix-vector multiplication (SpMV) plays a critical role in a wide range of linear algebra computations, particularly in scientific and engineering disciplines. However, the irregular memory access patterns, extensive memory usage, high bandwidth requirements, and underutilization of parallelism hinder the computational efficiency of SpMV on GPUs. In this paper, we propose a novel approach called block-wise dynamic mixed-precision (BDMP) to address these challenges. Our methodology involves partitioning the original matrix into uniformly sized blocks, with each block’s size determined by considering architectural characteristics and accuracy requirements. Additionally, we dynamically assign precision to each block using a precision selection method that takes into account the value distribution of the original sparse matrix. We develop two distinct SpMV computation algorithms for BDMP: BDMP-PBP (Precision-based partitioning) and BDMP-TCKI (Tailored compression and kernel implementation). BDMP-PBP partitions the matrix into two independent matrices for separate computations based on block precision, offering flexibility for integration with other optimization techniques. Meanwhile, BDMP-TCKI focuses on achieving significant thread-level parallelism and memory utilization by tailoring an appropriate compressed storage format and kernel implementation for each block. We compare BDMP with NVIDIA’s cuSPARSE library and three state-of-the-art SpMV methods, including SELLP, MergeBase, and BalanceCSR, using matrices from the University of Florida’s SuiteSparse dataset collection. BDMP-PBP and BDMP-TCKI show average speedups up to 2.64 $$\times $$ × and 2.91 $$\times $$ × on Turing RTX 2080Ti, and up to 2.99 $$\times $$ × and 3.22 $$\times $$ × on Ampere A100. The results demonstrate that BDMP enables the optimization of computation speed without compromising the precision necessary for reliable results.
Zhixiang Zhao, Guoyin Zhang, Yanxia Wu 0001, Ruize Hong, Yiqing Yang
J. Supercomput.3
2023 A lightweight bus passenger detection model based on YOLOv5
abstract
Abstract The bus passenger detection algorithm is a key component of a public transportation bus management system. The detection techniques based on the convolutional neural network have been widely used in bus passenger detection. However, they require high memory and computational requirements, which hinder the deployment of bus passenger detectors in the bus system. In this paper, a lightweight bus passenger detection model based on YOLOv5 is introduced. To make the model more lightweight, the inner and outer cross‐stage bottleneck modules, called ICB and OCB, respectively, are proposed. The proposed module reduces the quantity of parameter and floating point operations and increases the detection speed. In addition, the neighbour feature attention pooling is adopted to improve detection accuracy. The performance of the lightweight model on the bus passenger dataset is empirically demonstrated. The experiment results demonstrate that the proposed model is lightweight and efficient. Compared lightweight YOLOv5n with the original algorithm, the model weight is reduced by 31% to 2.6M, and the detection speed is increased by 6% to 40FPS without an accuracy drop.
Xiaosong Li 0003, Yanxia Wu 0001, Lidan Zhang, Ruize Hong
IET Image Process.2
2023 SGDAT: An optimization method for binary neural networks
Gu Shan, Guoyin Zhang, Jia Chengwei, Yanxia Wu 0001
Neurocomputing4
2023 Improving generalization of convolutional neural network through contrastive augmentation
Xiaosong Li 0003, Yanxia Wu 0001, Chuheng Tang, Lidan Zhang
Knowl. Based Syst.2
2023 Segmentation-Guided Semantic-Aware Self-Supervised Denoising for SAR Image
abstract
Synthetic aperture radar (SAR) images often suffer from speckle noise, which can degrade their visual quality and affect downstream applications. Recently, supervised and self-supervised methods have been proposed by virtue of deep learning with synthetic “noisy-clean” image pairs or only real noisy images as training data, respectively. Among these methods, by avoiding artifact problems in real SAR image denoising, self-supervised methods solve the domain gap problem of supervised methods and hence have attracted significant attention. However, existing self-supervised denoising methods essentially rely on pixel information of images and ignore corresponding semantic information, which makes them challenging to remove speckle noise while retaining detailed features. To this end, we propose a segmentation-guided semantic-aware self-supervised denoising method for SAR images, namely SARDeSeg, where a segmentation network is incorporated with a denoising network and guides it to learn and be aware of the semantic information of the input noisy SAR images. Additionally, a wavelet transform-based connector is introduced to efficiently transmit semantic information between the denoising network and the segmentation network, together with an edge-aware smoothing loss to improve speckle noise suppression while preserving edge features. Experimental results demonstrate that the proposed SARDeSeg outperforms state-of-the-art denoising methods for SAR images, particularly in preserving detailed edge features.
Ye Yuan 0011, Yanxia Wu 0001, Pengming Feng, Yulei Wu
IEEE Trans. Geosci. Remote. Sens.2
2022 Neighbour feature attention-based pooling
Xiaosong Li 0003, Yanxia Wu 0001, Chuheng Tang, Lidan Zhang
Neurocomputing2
2022 A Practical Solution for SAR Despeckling With Adversarial Learning Generated Speckled-to-Speckled Images
abstract
In this letter, we aim to address a synthetic aperture radar (SAR) despeckling problem with the necessity of neither clean (speckle-free) SAR images nor independent speckled image pairs from the same scene, and a practical solution for SAR despeckling (PSD) is proposed. First, an adversarial learning framework is designed to generate speckled-to-speckled (S2S) image pairs from the same scene in the situation where only single speckled SAR images are available. Then, the S2S SAR image pairs are employed to train a modified despeckling Nested-UNet model using the Noise2Noise (N2N) strategy. Moreover, an iterative version of the PSD method (PSDi) is also presented. Experiments are conducted on both synthetic speckled and real SAR data to demonstrate the superiority of the proposed methods compared with several state-of-the-art methods. The results show that our methods can reach a good tradeoff between feature preservation and speckle suppression.
Ye Yuan 0011, Jian Guan 0001, Pengming Feng, Yanxia Wu 0001
IEEE Geosci. Remote. Sens. Lett.4
2021 Self-Calibrated Convolutional Neural Network for SAR Image Despeckling
abstract
Synthetic aperture radar (SAR) images are contaminated by speckle noise, which has largely limited its practical applications. Recently, convolutional neural networks (CNNs) have indicated good potential for various image processing tasks. In this paper, we propose a self-calibrated convolutional neural network for SAR image despeckling, called SAR-SCCNN. To enlarge the receptive field of the network, downsampling and dilated convolutions are employed in each self-calibrated block. Also, the contextual information from spaces with different scales is extracted and concentrated to obtain accurate despeckled images. Experiments on synthetic speckled and real SAR data are conducted to perform the subjective visual assessment of image quality and objective evaluation. Results show that our proposed method can effectively suppress speckle noise and preserve detailed features.
Ye Yuan 0011, Yanxia Wu 0001, Richard Jiang 0001
IGARSS3
2019 Joint Convolutional Neural Network for Small-Scale Ship Classification in SAR Images
abstract
Ship classification using synthetic aperture radar (SAR) imagery is a challenge problem in maritime surveillance. Because of the scale limitation of ship targets in SAR image, convolutional neural networks (CNNs) can not achieve similar performance as for natural image classification. In this paper, we propose a joint CNNs framework for small-scale ship targets classification in SAR image, where a generator and a classifier are jointly connected. The generator can reconstruct the small-scale low-resolution (LR) images to large-scale super-resolution (SR) images, and the classifier is used for ship classification. A novel joint loss optimization strategy is introduced to solve the problem, where an MSE-based content loss is employed to generate high quality SR images, and a classification loss is applied to enable the generator and the classifier to be trained in a joint way. Experiments are conducted to demonstrate the superior performance of our proposed method, as compared with the state-of-the-art methods.
Yanxia Wu 0001, Ye Yuan 0011, Jian Guan 0001, Libo Yin, Pengming Feng
IGARSS1
2013 An Improved FPGAs-Based Loop Pipeline Scheduling Algorithm for Reconfigurable Compiler
Yanxia Wu 0001, Guoyin Zhang, Tianxiang Sui
APPT2
2009 32-bit floating-point FPGA gaussian elimination
abstract
The well-known Gaussian elimination (with partial pivoting) is a widely-used algorithm, one of traditional methods for solving dense linear systems of equations (LSEs). This paper presents a hardware-optimized variant of Gaussian elimination and its 32-bit ANSI/IEEE Std 754-1985 floating-point implementation on a Xilinx Virtex-5 FPGA with highly efficient design. The logic of the traditional algorithm is changed in order to make use of parallelism in hardware. According to this change the proposed hardware architecture can accomplish the solution very fast. Its average running time for n×n 32-bit floating-point matrices with uniformly distributed entries equals around n2(clock cycles) as opposed to n3 in software. Meanwhile, an open source library FPLibrary, which provides parameterizable pipelined floating-point operators, is used in the design. In realization, the design is finally integrated in an developed prototype system to accelerate the general purpose processor's work with the data exchanging through PCI-express between host and FPGA with DMA access method. Furthermore, by means of Strasson's algorithm, large LSEs also can be solved based on multiple FPGAs' co-work. The whole implementation placed and routed in the xc5vlx110t-3 FPGA with the applicability for solving LSE at most dimension 22, can be clocked with a frequency of up to 200MHz and computes the solution in 5.39 ¼s on average, providing a speed-up of up to almost 15 times over an equivalent software implementation on a Pentium IV 2.6GHz CPU. To the best of authors' knowledge, there has been no previous work on floating-point LSEs solving hardware and its implementation used as an application function unit in reconfigurable computing system.
Bowei Zhang 0002, Guochang Gu, Yanxia Wu 0001
FPGA4