EDBT 2026 Demo / reviewers in the wild / expert
Feng Xu 0008
dblp:03/2611-8
· DBLP profile ↗
46ranked-venue papers
1as first author
35since 2021 · last 2026
0000-0002-2175-9943ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 11 · 10 since 2021Computer networks · 3 · 1 since 2021Security and privacy · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CLIP-driven feature disambiguation and cross-modal synergy for few-shot semantic segmentation
Shangjing Chen, Feng Xu 0008, Xin Lyu 0001, Dafa Wang, Xin Li 0090 |
Expert Syst. Appl. | 2 |
| 2026 | TranCAD: Transforming tabular data into color images for deep semi-supervised anomaly detection
Yucong Huang, Feng Xu 0008, Xin Lyu 0001, Zhennan Xu |
Expert Syst. Appl. | 2 |
| 2026 | HFRW: High Fidelity and Robust Watermarking Using Deep Reinforcement LearningabstractDeep learning-based watermarking technology has made significant success due to its excellent robustness and convenient traceability. However, previous studies rarely considered users’ demands for high-fidelity images and file size growth rates, especially in scenarios involving the watermarking of a large number of high-resolution images. This paper proposes a high fidelity and robust watermarking using deep reinforcement learning. Specifically, we propose a self-optimization module that utilizes a dueling Q-learning network to select the optimal watermark embedding patch, which minimizes the quality impact after adding a watermark to the image. Additionally, we incorporate a convolutional block attention module (CBAM) into encoder and decoder networks to enhance feature learning in both frequency and spatial dimensions. Experimental results demonstrate that our method significantly enhances image fidelity, achieving PSNR values above 54dB across different datasets. The method also exhibits good robustness against common image manipulations, including both non-geometric and geometric attacks. Furthermore, our method achieves an FSVR value below 0.5, reducing file size inflation by a factor of 34 compared to existing advanced methods. This makes it very friendly for storing a large number of high-resolution watermarked images. Ping Ping, Ruixuan Jiang, Bobiao Guo, Feng Xu 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | A spectrum-enhanced attention model for semantic segmentation of remote sensing imagesabstractSemantic segmentation of remote sensing images (RSIs) is essential for applications such as environmental monitoring, urban planning, and disaster management. Convolutional Neural Networks (CNNs) and their variants struggle to capture comprehensive spectral context for learning discriminative representations. In this paper, we propose a Spectrum-Enhanced Network (SPENet) that leverages the Frequency Transformer Block (FTB) to capture rich spectral context. FTB integrates Spectrum-Enhanced Attention (SEA) with Multi-Head Frequency Self-Attention (MH-FSA), incorporating more informative contextual cues. Specifically, SEA aggregates spectral statistics through covariance matrix normalization before applying channel-wise attention. By projecting feature maps onto the frequency domain, MH-FSA provides the network with a broader context, extending beyond the low-frequency focus of standard self-attention mechanisms. Extensive experiments on the ISPRS Potsdam and LoveDA datasets show that SPENet significantly outperforms state-of-the-art methods. Besides, the proposed SEA module notably rises average F1-score/overall accuracy/mean insert over union wiht more than 2.5/2.6%/2.3%, as demonstrated by ablation study. Xin Li 0090, Feng Xu 0008, Feifei Tao, Xin Lyu 0001, Jianyi Zhong, André Kaup |
ICASSP | 2 |
| 2025 | Adaptive Frequency Threshold Pooling for Mitigating Aliasing in Few-Shot SegmentationabstractFeature diversity and interference from complex backgrounds pose substantial challenges to the few-shot segmentation (FSS). Convolutional Neural Networks, as the mainstream backbone for feature extraction, violate the sampling theorem during downsampling, leading to aliasing that causes blurred and distorted class feature boundaries. Such prototype features are likely to induce false responses of query features during the guidance process. To address this problem, we propose an innovative module called Adaptive Frequency Threshold Pooling (AFTP) that can be placed between the encoder stages. AFTP aims to selectively filter out high-frequency components that are prone to aliasing. It dynamically calculates frequency thresholds according to the spectral energy distribution, determining which frequency components are essential for maintaining feature integrity. Furthermore, it incorporates resolution-influenced threshold adjustment, enabling it to adapt to different feature resolutions. Experiments demonstrate that our module effectively enhances the performance of the FSS models on the PASCAL-5iand COCO-20idatasets. Shangjing Chen, Feng Xu 0008, Xin Lyu 0001, Xin Li 0090 |
ICME | 2 |
| 2025 | Multi-view aggregation and multi-relation alignment for few-shot fine-grained recognition
Feng Xu 0008, Xin Lyu 0001, Xin Li 0090, Shangjing Chen |
Expert Syst. Appl. | 2 |
| 2025 | Detail retention and enhancement for camouflaged object detectionabstractThe objective of camouflaged object detection is to segment objects that share similar visual attributes with their backgrounds. Existing methods leverage the consistency and variability of the object to pinpoint its location; however, when presented with samples exhibiting intricate edges, the model tends to overlook critical details. Toward this issue, we analyzed the reasons from the perspective of processing basic features and found that roughly compressing the feature channel could lead to the loss of fine features at the initial stage of the pipeline. Therefore, we introduce a multi-scale feature preprocessing approach named Soft Channel Compression prior to feature fusion, which retains key cues by enhancing the visibility of fine-grained features. We then utilize self-attention to fuse the results of multi-scale feature sampling, synthesizing contextual information between extended details and the main object. Furthermore, we propose the Hierarchical Receptive Field Complementarity module, using a divide-and-conquer strategy to enhance discriminative features with different scales in object details. This facilitates the integration of information across various ranges by employing a hierarchical manner to expand the receptive field, thereby achieving complementary information between local areas. Extensive experimental results on benchmark datasets demonstrate that the proposed method outperforms other state-of-the-art methods. Shangjing Chen, Feng Xu 0008, Xin Li 0090, Xin Lyu 0001 |
Neurocomputing | 2 |
| 2025 | Knowledge-driven prototype refinement for few-shot fine-grained recognition
Feng Xu 0008, Xin Lyu 0001, Xin Li 0090, Shangjing Chen |
Knowl. Based Syst. | 2 |
| 2025 | Knowledge-guided distribution alignment for cross-domain few-shot learning
Feng Xu 0008, Xin Lyu 0001, Xin Li 0090, Shangjing Chen |
Knowl. Based Syst. | 2 |
| 2025 | Foster noisy label learning by exploiting noise-induced distortion in foreground localization
Feng Xu 0008, Xin Lyu 0001, Xin Li 0090 |
Neural Networks | 2 |
| 2025 | Highly Robust and Diverse Coverless Image Steganography Against Passive and Active SteganalysisabstractTo avoid the pixel modification traces left by steganography from being detected by passive steganalysis, and to prevent the hidden data from being destroyed by active steganalysis attacks, Coverless Image Steganography (CIS) that does not modify pixels has attracted widespread attention. However, most existing CIS methods are limited in their maximum capacity due to insufficient diversity in their hash sequences. In addition, these methods struggle to maintain high robustness against both geometric and non-geometric attacks simultaneously. To address these two issues, a new coverless image steganography method is proposed to enhance CIS methods’ applicability, security, and robustness in highly insecure networks. During the hiding process, hash sequences are generated by a SHA-256 algorithm that integrates inter-block and inter-channel fusion, providing higher diversity than other CIS methods. Consequently, the proposed CIS method achieves higher capacity on publicly available datasets. During the extraction process, an evaluation metric that combines visual and histogram similarity is designed to improve the accuracy of inverse image retrieval. The experimental results demonstrate that the proposed CIS method achieves capacity increases of 26.73% and 38.34% over other CIS methods on the VOC and COCO datasets, respectively. Moreover, this method exhibits nearly 100% robustness against common active steganalysis. Bobiao Guo, Ping Ping, Feng Xu 0008 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2025 | A Euclidean Affinity-Augmented Hyperbolic Neural Network for Semantic Segmentation of Remote Sensing ImagesabstractSemantic segmentation of remote sensing images (RSIs) plays a pivotal role in advancing geospatial analyses and applications across diverse fields, such as urban planning and environmental monitoring. Traditional learning paradigms predominantly utilize Euclidean spaces for feature extraction. This approach can introduce spatial distortions when representing objects, as Euclidean architectures typically focus on locality and are optimized for grid data, not always yielding optimal geometrical representations for data structured in non-Euclidean spaces. To address these problems, we propose EAAHNet, the first fully hyperbolic neural network designed for semantic segmentation of RSIs. EAAHNet employs the Lorentz model to reformalize conventional Euclidean-based neural network operations, ensuring the preservation of hyperbolic properties. Furthermore, to account for the inherently Euclidean nature of ground objects, we propose a Euclidean affinity-augmented hyperbolic attention module (EAAHAM) that enriches contextual dependencies through an attention fusion manner. This enhancement significantly improves the network’s capacity to discern pixel-wise semantics. Extensive experiments conducted on the ISPRS Vaihingen, ISPRS Potsdam, and LoveDA datasets demonstrate EAAHNet’s superior performance over several state-of-the-art methods. Additionally, the ablation study verifies the impacts of EAAHAM. Xin Li 0090, Feng Xu 0008, Fan Liu 0003, Xin Lyu 0001, Hongmin Gao 0001, Jun Zhou 0011, André Kaup |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | A Frequency Decoupling Network for Semantic Segmentation of Remote Sensing ImagesabstractSemantic segmentation of remote sensing images (RSIs) is vital for numerous geospatial applications, including land-use mapping, urban planning, and environmental monitoring. Traditional neural networks for semantic segmentation primarily focus on learning in the spatial domain, which often results in suboptimal performance due to the complexity of RSIs that exhibit diverse and intricate structures. To address this problem, we propose a novel frequency decoupling network (FDNet) that enhances feature representation by independently refining high-frequency and low-frequency components in the frequency domain. FDNet introduces three core components: a sparse-aware spectral enhancement module (SSEM) that optimizes spectral feature learning by compressing redundant information while highlighting informative spectral bands, a frequency decoupling attention module (FDAM) that precisely distinguishes and enhances high-frequency and low-frequency features and an attentive frequency context module (AFCM) that integrates SSEM and FDAM into a cohesive framework for enriched spectral context modeling. Extensive experiments conducted on four benchmark datasets demonstrate that FDNet outperforms several state-of-the-art methods, achieving superior segmentation accuracy and robustness across various terrains and imaging conditions. Ablation experiments further confirm the impacts of SSEM, FDAM, and AFCM. Xin Li 0090, Feng Xu 0008, Anzhu Yu, Xin Lyu 0001, Hongmin Gao 0001, Jun Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | IGFNet: An Interactive-Guided Fusion Network for Hyperspectral PansharpeningabstractHyperspectral pansharpening is an efficient approach to obtaining high-resolution hyperspectral images (HR-HSIs) by fusing low-resolution hyperspectral images (LR-HSIs) with high-resolution panchromatic images (HR-PANs). However, the spatial and spectral distortions in reconstructed HR-HSIs are almost inevitable due to the modal gap between LR-HSIs and HR-PANs. Therefore, the performance of multi-source features fusion largely hinges on the ability to extract and align heterogeneous features across modalities. Most of the existing methods focus on integrating decoupled spatial and spectral information from different sources directly, which poses a dual challenge in aligning both spatial and spectral features effectively. To address the issues above-mentioned, a novel method named Interactive-Guided Fusion Network (IGFNet) is proposed, which is built upon a multi-stage progressive fusion framework. A high-resolution branch (HR) is introduced to interactively guide the alignment between cross-modal features, by fusing up-sampled HSI and PAN as a joint spatial-spectral guidance signal. Furthermore, the alignment is progressively conducted across stages, narrowing the modality gap and enhancing the representation of HR spatial-spectral feature. Additionally, we designed parameter-free spatial, spectral, and spatial-spectral attention mechanisms to extract global and local features effectively. Extensive experiments on reduced-resolution and full-resolution datasets demonstrate that IGFNet outperforms state-of-the-art across various metrics. Specifically, with a scaling factor of 4 on the Pavia University dataset, our method achieves a 2.09% relative improvement in PSNR, a 0.99% relative increase in SSIM, a 2.6% relative reduction in SAM, while reducing the parameter count by compared to the baseline MDA-Net. Zhennan Xu, Xin Lyu 0001, Feng Xu 0008, Xin Li 0090, Yucong Huang, Caifeng Wu, Yiwei Fang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Robust Reversible Watermarking With Invisible Distortion Against VAE Watermark RemovalabstractOrthogonal Moment-based Robust Reversible Watermarking (OM-RRW) is crucial for intellectual property protection, providing the dual benefits of robustness and reversibility. However, OM-RRW embeds watermarks into visually sensitive global low-frequency features, which easily leads to ring-like distortions that expose watermark locations, making them vulnerable to removal through image inpainting. To address this issue, this paper makes the first attempt to introduce an innovative strategy to eliminate these visible distortions, thereby overcoming OM-RRW's inherent limitations. The strategy innovates on two fronts: first, it customizes varying embedding step sizes based on the stability differences of moment values to minimize distortion; second, it designs a texture-aware adaptive basis function fine-tuning strategy. This strategy adjusts the representation capability of the basis functions in different regions based on the human eye's sensitivity to various texture areas, helping to avoid visible ring-like distortions. The performance of the proposed method is evaluated using Polar Harmonic Transform (PHT) moments, comprising three moments that exhibit remarkable performance in existing OM-RRW methods. Extensive experiments show that the proposed method can embed 128-bit watermarks with no visible distortions while minimizing the loss of robustness. In addition, this paper finds that OM-RRW demonstrates satisfactory robustness against VAE watermark removal attacks. Bobiao Guo, Ping Ping, Fan Liu 0003, Feng Xu 0008 |
IEEE Trans. Image Process. | 4 |
| 2024 | Locality-Enhanced Transformer for Semantic Segmentation of High-Resolution Remote Sensing ImagesabstractTransformers have emerged as a transformative tool in various computer vision tasks, excelling at capturing long-range dependencies. Their potential applicability and scalability in the interpretation of high-resolution remote sensing images (HRRSIs) have thus garnered substantial interest. However, unlike natural images, HRRSIs present intricate scenes characterized by scale variations and diverse appearances. These challenges underscore the importance of enabling networks to effectively assimilate both local intricacies and global context. In this letter, we introduce LETFormer, a semantic segmentation transformer. LETFormer balances capturing longrange dependencies with preserving local details through its unique LETFormer block, featuring an anchor token. This token aggregates localized contextual information within a designated window and promotes meaningful interactions among anchor tokens. With a mask transformer decoder, LETFormer gains ample contextual cues for precise semantic mask prediction. Empirical findings based on evaluations using the ISPRS Potsdam and LoveDA benchmarks unequivocally establish LETFormer’s superiority over state-of-the-art models. Additionally, we analyze the parameter size and floating-point operations per second (FLOPs) of LETFormer. Xin Li 0090, Feng Xu 0008, Runliang Xia, Nan Xu 0008, Fan Liu 0003, Chi Yuan, Qian Huang 0008, Xin Lyu 0001 |
ICASSP | 2 |
| 2024 | FreqFormer: A Frequency Transformer for Semantic Segmentation of Remote Sensing ImagesabstractSemantic segmentation of remote sensing images (RSIs) is vital for geospatial intelligence. However, traditional methods face challenges with mixed pixels and complex land cover types. Convolutional neural networks and transformers have led the field of RSI semantic segmentation by learning visual features in the spatial domain, but they often overlook the rich spectral features which can be well-described in the frequency domain, resulting in inadequate context modeling. In this paper, we present FreqFormer, a frequency transformer that enhances semantic segmentation by incorporating both spectral and spatial information through a devised frequency attention (FA) module. FA refines representations in the frequency domain through two parallel branches. Specifically, the high-frequency branch (HFB) utilizes a convolution layer with a Canny kernel to preserve high-frequency details, followed by multi-head self-attention to model high-frequency context. Followed by an element summation, high-frequency and low-frequency contexts are aggregated. Then, the formed FreqFormer block is sequentially deployed in the encoder stage with patch merging for spatial contraction. As for the decoder, the mask transformer decoder applies a scalar product to predict patch-wise semantics before upsampling. In experiments, FreqFormer outperforms state-of-the-art models on the ISPRS Potsdam and LoveDA datasets, demonstrating significant improvements in numerical evaluations. The integration of HFB significantly boosts the model’s ability to capture fine details, highlighting its potential for geospatial analysis. Xin Li 0090, Feng Xu 0008, Fan Liu 0003, Yiwei Fang, Xin Lyu 0001, Jun Zhou 0001 |
MMAsia | 2 |
| 2024 | SigCo: Eliminate the inter-class competition via sigmoid for learning with noisy labels
Feng Xu 0008, Xin Lyu 0001, Xin Li 0090 |
Knowl. Based Syst. | 2 |
| 2024 | AAFormer: Attention-Attended Transformer for Semantic Segmentation of Remote Sensing ImagesabstractThe rapid advancements in remote sensing technology have enabled the widespread availability of fine-resolution remote sensing images (RSIs), offering rich spatial details and semantics. Despite the applicability and scalability of transformers in semantic segmentation of RSIs by learning pairwise contextual affinity, they inevitably introduce irrelevant context, hindering accurate inference of patch semantics. To address this, we propose a novel multi-head attention-attended module (AAM) that refines the multi-head self-attention mechanism. The AAM filters out irrelevant context while highlighting informative ones by considering the relevance between self-attention maps and the query vector. The AAM generates an attention gate to complement contextual affinity and emphasize the useful ones with a higher weight simultaneously. Leveraging multi-head AAM as the core unit, we construct a lightweight attention-attended transformer block (ATB). Subsequently, we devise AAFormer, a pure transformer with a mask transformer decoder, for achieving semantic segmentation of RSIs. We extensively evaluate our approach on the ISPRS Potsdam and LoveDA datasets, demonstrating compelling performance compared to mainstream methods. Additionally, we conduct evaluations to analyze the effects of AAM. Xin Li 0090, Feng Xu 0008, Linyang Li, Nan Xu 0008, Fan Liu 0003, Chi Yuan, Xin Lyu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | A Cross-Domain Coupling Network for Semantic Segmentation of Remote Sensing ImagesabstractSemantic segmentation of remote sensing images (RSIs) is critical for various applications, including urban planning, agriculture, and disaster management. Existing methods often fail to capture fine-grained textures and periodic patterns in RSIs, leading to suboptimal results in complex terrains. To address these challenges, we propose a cross-domain coupling network (CDCNet) that leverages both domain-specific extraction and cross-domain coupling (CDC) to enrich contextual cues for semantic inference. Our CDCNet integrates a CDC layer within the encoder-decoder architecture to simultaneously refine representations in the frequency and spatial domains. This approach effectively models fine-grained textures and periodic patterns in the frequency domain, as well as edges, shapes, and broad structural elements in the spatial domain. Extensive experiments on the ISPRS Potsdam and LoveDA datasets demonstrate the superiority of CDCNet over several state-of-the-art methods. Ablation studies confirm the significant impact of the CDC layer, validating the effectiveness of our approach in handling RSIs. Xin Li 0090, Feng Xu 0008, Feifei Tao, Hongmin Gao 0001, Fan Liu 0003, Xin Lyu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | A Frequency Domain Feature-Guided Network for Semantic Segmentation of Remote Sensing ImagesabstractSemantic segmentation of Remote Sensing Images (RSIs) entails assigning semantic labels to each pixel accurately. RSIs are rich in spatial and spectral data, revealing diverse material and object characteristics. Yet, current RSI-focused computer vision models struggle with significant intra-class variation and inter-class resemblance due to limited spectral data usage. We propose the Frequency Domain Feature-Guided Network (FFGNet) for RSI semantic segmentation, influenced by digital signal processing theories. FFGNet initially generates frequency domain features via patch partitioning and 2D discrete cosine transformation. Our Frequency Enhancement Attention module (FEA) then distinguishes and intensifies frequency components to retain detailed information. These enhanced features are integrated with the Spatial-Spectral Attention (SSA) for enriched spectral signals. In the inference phase, these features are upsampled and combined with decoded features, emphasizing spectral details. Additionally, our novel loss function combines frequency and cross-entropy losses. Experiments on LoveDA and ISPRS Potsdam datasets demonstrate FFGNet's effectiveness, surpassing other mainstream models. An ablation study further validates our dual-guidance design. Xin Li 0090, Feng Xu 0008, Hongmin Gao 0001, Fan Liu 0003, Xin Lyu 0001 |
IEEE Signal Process. Lett. | 2 |
| 2024 | IMIH: Imperceptible Medical Image Hiding for Secure HealthcareabstractMedical images play a crucial role in doctors' clinical diagnosis and treatment. However, the transmission and sharing of such private information raises security concerns. To address this issue, image hiding is used as an effective technique to protect images. To achieve large hiding capacity, lossless recovery and anti-steganalysis, we propose a novel two-stage medical image-hiding method in this paper. In the first stage, a QR code for the patient diagnosis information (PDI) is generated and embedded into a secret medical image using reversible data hiding. In the second stage, the secret medical image containing PDI is hidden in a natural target image. A kind of lossless compression technique named soft compression is innovatively introduced in two hiding stages, to ensure that the reconstructed secret medical image and PDI are exactly identical to the original ones. Moreover, an adaptiven-LSB model is proposed to improve the stego image quality. Extensive experimental results show that our method achieves a PSNR of over 40dB for the stego image at 2 BPP while recovering the PDI and secret medical image with 100% accuracy on the DIV2K, COCO and ImageNet datasets. It outperforms other state-of-the-art methods in terms of hiding invisibility, recovery accuracy and security. Ping Ping, Pan Wei, Deyin Fu, Bobiao Guo, Olano Teah Bloh, Feng Xu 0008 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2024 | Semantic Segmentation of Remote Sensing Images by Interactive Representation Refinement and Geometric Prior-Guided InferenceabstractHigh spatial resolution remote sensing images (HRRSIs) contain intricate details and varied spectral distributions, making their semantic segmentation a challenging task. To address this problem, it is crucial to adequately capture both local and global contexts to reduce semantic ambiguity. While self-attention modules in vision transformers capture long-range context, they tend to sacrifice local details. In this article, we propose a geometric prior-guided interactive network (GPINet), a hybrid network that refines features across encoder and decoder stages. First of all, a dual branch structure encoder with local-global interaction modules (LGIMs) is designed to fully exploit local and global contexts for feature refinement. Unlike commonly used skip connections or concatenations, the LGIMs bilaterally couple and exchange CNN features with transformer features by lossless transformation and elaborating cross-attention. Moreover, we introduce a geometric prior generation module (GPGM) that iteratively updates the randomly initialized geometric prior. Subsequently, the geometric priors are stored and used to guide feature recovery. Finally, a weighted summation is applied to the upsampled decoded features and geometric priors. By comprehensively capturing contexts and enabling lossless decoding and deterministic inference, GPINet allows the network to learn discriminative representations for accurately specifying pixel-level semantics. Experiments on three benchmark datasets demonstrate the superiority of the proposed GPINet over state-of-the-art methods. Furthermore, we validate the effectiveness of geometric priors and compare the model sizes. Xin Li 0090, Feng Xu 0008, Fan Liu 0003, Xin Lyu 0001, Jun Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Hiding Multiple Images into a Single Image Using Up-SamplingabstractThe goal of multiple-image hiding is to hide several secret images within another carrier image without significantly changing its appearance, and then perfectly reconstruct all of the secret images. The challenge is to ensure that the stego-image has great visual quality and can resist various steganalysis under the premise of hiding as much information as possible in one image. To address this issue, the majority of known image-hiding methods focus on hiding images using compression techniques. In this article, we present a novel multiple-image hiding method based on up-sampling and reversible color transformation. First, the interpolation algorithm up-samples the carrier image, so that the attribute of similar neighboring pixel values in the up-sampled image can significantly improve the effect of image hiding. The embedding procedure is then performed using the proposed Euclidean Distance (ED)-based block matching and reversible color transformation, which decreases the chance of local blurring in the stego-image. Experimental results show that the proposed method surpasses existing advanced methods by achieving an average of 33 dB and 28 dB of PSNR for the stego-image with a hiding capacity 2 BPP and 8 BPP, and obtaining 100% reconstructing accuracy for all secret images. It also has a high level of resistance to steganalysis and a strong robustness against various image-processing attacks. Ping Ping, Bobiao Guo, Olano Teah Bloh, Yingchi Mao, Feng Xu 0008 |
IEEE Trans. Multim. | 5 |
| 2024 | AISM: An Adaptable Image Steganography Model With User CustomizationabstractIn the field of image steganography, quality, security, and capacity emerge as three crucial aspects for ensuring the security of images stored on cloud servers. However, many existing methods fail to strike a balance on these three aspects according to the various requirements of image users. To solve this issue, we propose an Adaptable Image Steganography Model (AISM) capable of customizing suitable steganography strategies for different user requirements. Initially, AISM customizes appropriate down-sampling methods, ratios, and up-sampling ratios for secret/cover images based on requirements, followed by the image sampling process. Following that, an embedding algorithm based on pixel-value coding is proposed, which maps pixel values from [0, 255] to [-9, 9] and then replaces the high-frequency sub-band coefficients of the up-sampled image. During the embedding process, no auxiliary information is generated, which is a key aspect for user-friendliness. Extensive experimental results demonstrate that our method is capable of customizing satisfactory steganography strategies for various user requirements. Moreover, our method outperforms many state-of-the-art methods in terms of quality and security, i.e., lossless recovery of secret images, PSNR of 40 dB for stego-images, evasion from detection by six classic steganalysis tools. Bobiao Guo, Ping Ping, Olano Teah Bloh, Feng Xu 0008 |
IEEE Trans. Serv. Comput. | 5 |
| 2023 | Asymmetric exponential loss function for crack segmentation
Fan Liu 0003, Delong Chen, Chunmei Shen, Feng Xu 0008 |
Multim. Syst. | 5 |
| 2023 | MEP-3M: A large-scale multi-modal E-commerce product dataset
Fan Liu 0003, Delong Chen, Xiaoyu Du 0002, Ruizhuo Gao, Feng Xu 0008 |
Pattern Recognit. | 5 |
| 2023 | Mobility-Aware Proactive Flow Setup in Software-Defined Mobile Edge NetworksabstractThe software-defined network (SDN) enabled mobile edge network greatly facilitates network resource management and promotes many emerging applications. However, user mobility may cause the SDN controller to set flow rules frequently, introduce additional flow setup latency, cause delay jitter, and undermine latency-sensitive services. Proactive flow setup is an effective way to eliminate flow setup latency, but existing work fails to maximize the flow setup hit ratio, a metric for evaluating the quality of proactive flow setup decisions, which is critical for latency-sensitive services. In this paper, we study how to proactively set flow rules to maximize the flow setup hit ratio under limited available network resources to eliminate the flow setup latency as much as possible. Then, we formalize the proactive flow setup problem as two integer linear programming problems under two typical routing strategies, default routing and dynamic routing. Both problems are proved to be NP-hard. To tackle these two problems, we propose a linear programming-based polynomial-time approximation algorithm for the default routing case and a greedy-based heuristic algorithm for the dynamic routing case. Extensive trace-driven experimental and simulation results verify that our algorithms can improve the flow setup hit ratio by up to 30.99% compared to existing solutions. Yue Zeng 0002, Bin Tang 0002, Sanglu Lu, Feng Xu 0008, Song Guo 0001, Zhihao Qu |
IEEE Trans. Commun. | 5 |
| 2023 | Hyperspectral Target Detection via Spectral Aggregation and Separation Network With Target Band Random MaskabstractHyperspectral target detection (HTD) is a pixel-wise detection method based on limited prior targets and spectral differences, which has been widely studied and applied in many fields. Recently, deep learning (DL) plays an important role in hyperspectral imagery (HSI) processing. However, for HTD, the severe lack of class-balanced training sets is an enormous challenge. Meanwhile, it is difficult to suppress backgrounds while highlighting targets through the deep network. To address these issues, we propose a spectral aggregation and separation network (SASN) with a target band random mask (TBRM) for HTD in this paper. For the training sets of SASN, a multifarious representative background selection strategy (MRBS) is first proposed to obtain a multifarious and representative background training set. Next, aiming at the notorious class imbalance, a data augmentation (DA) method, TBRM, is proposed to generate adequate target training set by repeating randomly zero-masking the spectral bands of a prior target. Subsequently, in the training of SASN, residual connection and squeeze-and-excitation (SE) channel attention mechanism are applied to fully extract high discriminative features and nonlinear ones in the spectra. Besides, to better separate the targets and backgrounds, a triplet-soft loss function is presented, which makes the training in the direction of spectral separation of background samples from both the prior target and target samples. During testing, the trained SASN distinguishes the spectral similarities and differences simultaneously for highlighting targets and suppressing backgrounds. Moreover, extensive experimental results validate that the proposed method has superior detection performances, background suppression capacity, and separability compared with ten cutting-edge HTD algorithms on six benchmark HSI datasets. Hongmin Gao 0001, Zhonghao Chen, Feng Xu 0008, Danfeng Hong, Bing Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | A Synergistical Attention Model for Semantic Segmentation of Remote Sensing ImagesabstractIn remotely sensed images, high intraclass variance and interclass similarity are ubiquitous due to complex scenes and objects with multivariate features, making semantic segmentation a challenging task. Deep convolutional neural networks can solve this problem by modeling the context of features and improving their discriminability. However, current learning paradigms model the feature affinity in spatial dimension and channel dimension separately and then fuse them in a sequential or parallel manner, leading to suboptimal performance. In this study, we first analyze this problem practically and summarize it as attention bias that reduces the capability of network in distinguishing weak and discretely distributed objects from wide-range objects with internal connectivity, when modeled only in spatial or channel domain. To jointly model both spatial and channel affinity, we design a synergistic attention module (SAM), which allows for channelwise affinity extraction while preserving spatial details. In addition, we propose a synergistic attention perception neural network (SAPNet) for the semantic segmentation of remote sensing images. The hierarchical-embedded synergistic attention perception module aggregates SAM-refined features and decoded features. As a result, SAPNet enriches inference clues with desired spatial and channel details. Experiments on three benchmark datasets show that SAPNet is competitive in accuracy and adaptability compared with state-of-the-art methods. The experiments also validate the hypothesis of attention bias and the efficiency of SAM. Xin Li 0090, Feng Xu 0008, Fan Liu 0003, Xin Lyu 0001, Zhennan Xu, Jun Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | A review of driver fatigue detection and its advances on the use of RGB-D camera and deep learning
Fan Liu 0003, Delong Chen, Jun Zhou 0001, Feng Xu 0008 |
Eng. Appl. Artif. Intell. | 4 |
| 2022 | Feature difference for single-shot object detectionabstractAbstract The one‐stage detectors achieve a good trade‐off between performance and latency, owing to the plain architecture and divergent learning mechanism for classification and localization. However, the two sub‐tasks require features with various inherency with which to generate inconsistent detections, fettering detectors. In this study, the misalignment is deeply analyzed via kernel density estimation (KDE) for the first time. Moreover, to address the misalignment, a plug‐and‐play detection head, named Diff‐Head, is devised and embedded in one‐stage detectors. Concretely, the authors merge parallel branches into a semi‐parallel structure, establishing the correlation between classification and regression. In the regression branch, a feature difference module (FDM) gets rid of the features that favour classification by subtracting salient object features from the original feature map, and position encoding (PE) modules enhance the absolute position information. The flexibility and efficiency of the detection head are retained. Experiments on Pascal visual object classes (VOC) and MS COCO demonstrate that Diff‐Head is effective and achieves competitive performance with state‐of‐the‐art detectors. Meanwhile, the amount of parameters is reduced at least 30% and 83.0% average precision (AP) is achieved on Pascal VOC. The analyses of consistency and error show that Diff‐Head has better localization and the capability of mitigating the misalignment. Feng Xu 0008, Xin Lyu 0001, Xin Li 0090, Xinyuan Wang 0002, Caifeng Wu |
IET Image Process. | 2 |
| 2022 | Self-Supervised Music Motion Synchronization Learning for Music-Driven Conducting Motion Generation
Fan Liu 0003, Delong Chen, Ruizhi Zhou, Sai Yang, Feng Xu 0008 |
J. Comput. Sci. Technol. | 5 |
| 2022 | Hybridizing Euclidean and Hyperbolic Similarities for Attentively Refining Representations in Semantic Segmentation of Remote Sensing ImagesabstractAttention mechanisms have revolutionized the semantic segmentation network in interpreting remotely sensed images (RSIs) due to their amazing ability in establishing contextual dependencies. Nevertheless, due to the complex scenes and diverse objects in RSIs, a variety of details and correlations are not available in Euclidean space. Therefore, a similarity-hybrid attention module (SHAM) is devised to attentively learn the hyperbolic and Euclidean attention maps between any two positions, followed by a weighted element-wise summation. The hybrid attention maps posses latent geometric properties of both Euclidean and hyperboloid. Taking commonly-used fully convolutional network (FCN) as baseline, HAENet that embeds SHAM, is presented. Experiments on ISPRS Potsdam and DeepGlobe benchmarks reveal its superiority to comparative methods. In addition, the ablation study validates the effectiveness of SHAM compared to other attention modules. Xin Li 0090, Feng Xu 0008, Fan Liu 0003, Runliang Xia, Linyang Li, Zhennan Xu, Xin Lyu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2021 | Semantic Labeling of Very High-Resolution Imagery by Leveraging Contextual Information with Optimized Non-Local Neural NetworkabstractSemantic labeling of very high-resolution (VHR) aerial imagery attempts to assign a specific label to every pixel. Resorting to the significant progress of semantic segmentation for natural images, the non-local neural networks express superior capability in capturing multi-scale contextual information, which is helpful for accurately labeling raw images. However, the computational complexity of non-local neural networks is criticized. In this paper, the optimized non-local block (ONN) is devised and embedded in the encoder stage to reduce the computational complexity in capturing contextual information. Besides, the DUpsampling is introduced to further boost the efficiency of decoder. Finally, the extensive experiments, including numerical metrics and visual inspection comparisons, are conducted on ISPRS benchmarks. The results indicate that the proposed method performs better than several SOTA methods both in accuracy and efficiency. Xin Li 0090, Feng Xu 0008, Xin Lyu 0001, Liancheng Zhao, Xinyuan Wang 0002 |
IGARSS | 2 |
| 2019 | Single sample face recognition via BoF using multistage KNN collaborative codingabstractIn this paper, we propose a multistage KNN collaborative coding based Bag-of-Feature (MKCC-BoF) method to address SSPP problem, which tries to weaken the semantic gap between facial features and facial identification. First, local descriptors are extracted from the single training face images and a visual dictionary is obtained offline by clustering a large set of descriptors with K-means. Then, we design a multistage KNN collaborative coding scheme to project local features into the semantic space, which is much more efficient than the most commonly used non-negative sparse coding algorithm in face recognition. To describe the spatial information as well as reduce the feature dimension, the encoded features are then pooled on spatial pyramid cells by max-pooling, which generates a histogram of visual words to represent a face image. Finally, a SVM classifier based on linear kernel is trained with the concatenated features from pooling results. Experimental results on three public face databases show that the proposed MKCC-BoF is much superior to those specially designed methods for SSPP problem. Moreover, it also has great robustness to expression, illumination, occlusion and, time variation. Fan Liu 0003, Sai Yang, Yuhua Ding, Feng Xu 0008 |
Multim. Tools Appl. | 4 |
| 2018 | Designing permutation-substitution image encryption networks with Henon map
Ping Ping, Feng Xu 0008, Yingchi Mao, Zhijian Wang 0002 |
Neurocomputing | 2 |
| 2018 | Design of image cipher using life-like cellular automata and chaotic map
Ping Ping, Jinjie Wu, Yingchi Mao, Feng Xu 0008, Jinyang Fan |
Signal Process. | 4 |
| 2017 | POSTER: Neural Network-based Graph Embedding for Malicious Accounts DetectionabstractWe present a neural network based graph embedding method for detecting malicious accounts at Alipay, one of the world's leading mobile payment platform. Our method adaptively learns discriminative embeddings from an account-device graph based on two fundamental weaknesses of attackers, i.e. device aggregation and activity aggregation. Experiments show that our method achieves outstanding precision-recall curve compared with existing methods. Chaochao Chen 0001, Jun Zhou 0011, Xiaolong Li 0005, Feng Xu 0008 |
CCS | 5 |
| 2016 | A multi-phase sparse probability framework via entropy minimization for single sample face recognitionabstractIn this paper, we propose a robust probability based sparse method to solve single sample face recognition, which harvests the advantages of both local and global representation. Different from previous sparse representation methods that generate sparse coefficients by l1, we produce sparse class probability distribution by proposing a multi-phase sparse probability (MSP) framework. To create class probability distribution, we divide each face image into many local blocks and vote based on the classification results of all blocks. For classifying each block, we propose local similarity assumption that makes many conventional methods feasible to SSPP problem. Moreover, we also propose a heuristic multiphase class selection scheme to solve the entropy minimization problem, which finally provides a higher classification confidence from the global perspective. Experimental results on three popular databases show that our approach not only generalizes well to SSPP problem but also has strong robustness to expression, illumination, occlusion and time variation. Fan Liu 0003, Jinhui Tang 0001, Yan Song 0005, Qian Huang 0008, Feng Xu 0008 |
ICIP | 5 |
| 2014 | Pilot Power Optimization for AF Relaying Using Maximum Likelihood Channel EstimationabstractBit error rates (BERs) for amplify-and-forward (AF) relaying systems with two different pilot-symbol- aided channel estimation methods, disintegrated channel estimation (DCE) and cascaded channel estimation (CCE), are derived in Rayleigh fading channels. Based on these BERs, the pilot powers at the source and at the relay are optimized when their total transmitting powers are fixed. Numerical results show that the optimized system has a better performance than other conventional nonoptimized allocation systems. They also show that the optimal pilot power in variable gain is nearly the same as that in fixed gain for similar system settings. Kezhi Wang, Yunfei Chen 0001, Mohamed-Slim Alouini, Feng Xu 0008 |
VTC Fall | 4 |
| 2014 | Polynomial-approximation-based locally optimum detector for signals with symmetric alpha stable noiseabstractRational approximation to the non‐linear score function used in the locally optimum detector is derived for signals corrupted by the impulsive symmetric alpha stable noise. The new approximation uses a third‐order polynomial in the numerator and a fourth‐order polynomial in the denominator, compared with the existing approximation that uses a first‐order polynomial in the numerator and a second‐order polynomial in the denominator. The parameters of the polynomials are derived using non‐linear least squares curve fitting. The relationships between the polynomial parameters and the value of the characteristic exponent are also obtained. Numerical results show that the proposed approximation has superior accuracy to the existing approximations. The proposed new approximation is then applied to the locally optimum detector by replacing the score function in the decision variable. Numerical results show that the proposed detector, optimised according to a curve‐fitting approach, outperforms previous approximations of the locally optimum detector, also optimised according to a curve‐fitting approach, for binary phase shift keying signals and in some cases for on–off keying signals. Yunfei Chen 0001, Feng Xu 0008, Jiming Chen 0001 |
IET Commun. | 2 |
| 2014 | Image encryption based on non-affine and balanced cellular automata
Ping Ping, Feng Xu 0008, Zhijian Wang 0002 |
Signal Process. | 2 |
| 2014 | BER and Optimal Power Allocation for Amplify-and-Forward Relaying Using Pilot-Aided Maximum Likelihood EstimationabstractBit error rate (BER) and outage probability for amplify-and-forward (AF) relaying systems with two different channel estimation methods, disintegrated channel estimation and cascaded channel estimation, using pilot-aided maximum likelihood method in slowly fading Rayleigh channels are derived. Based on the BERs, the optimal values of pilot power under the total transmitting power constraints at the source and the optimal values of pilot power under the total transmitting power constraints at the relay are obtained, separately. Moreover, the optimal power allocation between the pilot power at the source, the pilot power at the relay, the data power at the source and the data power at the relay are obtained when their total transmitting power is fixed. Numerical results show that the derived BER expressions match with the simulation results. They also show that the proposed systems with optimal power allocation outperform the conventional systems without power allocation under the same other conditions. In some cases, the gain could be as large as several dB's in effective signal-to-noise ratio. Kezhi Wang, Yunfei Chen 0001, Mohamed-Slim Alouini, Feng Xu 0008 |
IEEE Trans. Commun. | 4 |
| 2011 | A new identity-based threshold ring signature schemeabstractAiming at the security problems existed in most identity-based threshold ring signature (ITRS) schemes, including poor security and inefficient in verifying, a new ITRS scheme is proposed. We significantly enhance the security of signature key by using bilinear pairings and secret sharing techniques. Our analysis indicates that the proposed scheme achieves unconditional anonymity and unforgeability. A security problem (ring member-changing attack) also can be solved. Compared to the congener schemes, ours is much more efficient in signing process. Feng Xu 0008 |
SMC | 1 |
| 2006 | The Service-Oriented Data Integration Platform for Water Resources Management
Zhijian Wang 0002, Feng Xu 0008 |
APWeb | 3 |