EDBT 2026 Demo / reviewers in the wild / expert
Yan Zhang 0108
dblp:04/3348-108
· DBLP profile ↗
36ranked-venue papers
15as first author
36since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021Computer networks · 8 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Kalman filter scheduling for 6TiSCH network with traffic adaptation optimized for bursty traffic
Yan Zhang 0108, Yan Han 0002 |
Ad Hoc Networks | 1 |
| 2026 | HICID: A hierarchical classification framework for intrusion detection under extreme class imbalance
Yan Han 0002, Kaikang Zheng, Yan Zhang 0108 |
Comput. Networks | 4 |
| 2026 | Self-learned generalized scheduling for 6TiSCH networks: An AI for science method
Yan Zhang 0108, Haopeng Huang, Yan Han 0002 |
Comput. Networks | 1 |
| 2026 | A lightweight granular perception feature pyramid network with context-awareness for small traffic sign detection
Yan Zhang 0108, Dengfeng Bi, Yan Han 0002, Minghang Zhao |
Expert Syst. Appl. | 1 |
| 2026 | FedAKD: A Lightweight Federated Adaptive Knowledge Distillation Method for Fault Diagnosis
Yan Han 0002, Yan Zhang 0108 |
IEEE Internet Things J. | 4 |
| 2026 | Joint Optimization of Task Offloading and Resource Allocation in Smart Factories via Cascaded Dual-Branch Network-Based Deep Reinforcement Learning ApproachabstractMulti-access Edge Computing (MEC) is pivotal for smart factories, yet a significant challenge remains in efficiently offloading heterogeneous tasks, which requires balancing deterministic tasks with strict time windows and nondeterministic tasks that demand minimized latency. Traditional optimization methods and standard Deep Reinforcement Learning (DRL) algorithms often struggle to address the inherent hybrid discrete-continuous action space in joint task offloading and resource allocation. To address this challenge, a Cascaded Dual-Branch Network-based Deep Reinforcement Learning (CDBN-DRL) method is proposed. First, a dual-branch network that combines convolutional and self-attention mechanisms is constructed for the robust extraction of environmental state features. Then, the extracted feature vector is fed into a cascaded two-layer DRL architecture, where an upper-layer Double Deep Q-Network (DDQN) handles discrete offloading decisions, while a lower-layer Soft Actor-Critic (SAC) network manages continuous resource allocation based on the upper-layer decisions. This cascaded structure is designed to effectively manage and co-optimize the hybrid discrete-continuous action space. Experimental results indicate that the CDBN-DRL method achieves notable performance, achieving completion rates of 99.4% for deterministic tasks and 98.3% for nondeterministic tasks while outperforming existing state-of-the-art offloading schemes in terms of latency, power consumption, and generalization capability. Songsong Mu, Yan Han 0002, Wendi Nie, Yan Zhang 0108 |
IEEE Internet Things J. | 5 |
| 2026 | A cluster aggregation-based personalized federated learning framework for wind turbine gearbox fault diagnosis under model heterogeneity
Yan Han 0002, Zhiyao Liu, Yan Zhang 0108, Bin Yong |
Knowl. Based Syst. | 5 |
| 2026 | CFDNet: An Interpretable Causal Filtering Disentanglement Domain Generalization Network for Fault Diagnosis Under Unseen ConditionsabstractDomain generalization-based fault diagnosis (DGFD) has emerged as a promising approach for addressing mechanical fault diagnosis under unseen working conditions. However, mainstream DGFD methods based on statistical dependencies typically model only explicit relationships between temporal data and labels. They often fail to uncover implicit causal connections across different conditions, impairing the reliability and interpretability of the diagnosis results. To address this issue, an interpretable causal filtering and disentanglement domain generalization network is proposed. Specifically, discrete wavelet prior knowledge is embedded in the model to expand the signal from the time domain to the wavelet space, a causal filter is then designed to mine and filter causal features in the wavelet domain. Furthermore, a causal clustering loss and a noncausal discrimination loss are introduced to guide the network in disentangling causal information from latent representations within the causal feature space, thereby enhancing causal disentanglement. Extensive generalization experiments conducted on four machines demonstrate that the proposed method achieves superior generalization performance and interpretability. Sipeng Lv, Yan Han 0002, Yan Zhang 0108 |
IEEE Trans. Ind. Informatics | 4 |
| 2025 | DichotomyIR: Universal Image Reconstruction via Dichotomy Classification and Uncertainty Elimination
Yan Zhang 0108, Shiwen He, Lin Yuan 0002, Jiaxu Leng, Xinbo Gao 0001 |
ACM Multimedia | 1 |
| 2025 | MLEP: Multi-granularity Local Entropy Patterns for Generalized AI-generated Image DetectionabstractAdvances in image generation technologies have raised growing concerns about their potential misuse, particularly in producing misinformation and deepfakes. This creates an urgent demand for effective methods to detect AI-generated images (AIGIs). While progress has been made, achieving reliable performance across diverse generative models and scenarios remains challenging due to the absence of source-invariant features and the limited generalization of existing approaches. In this study, we investigate the potential of using image entropy as a discriminative cue for AIGI detection and propose Multi-granularity Local Entropy Patterns (MLEP), a set of feature maps computed based on Shannon entropy from shuffled small patches at multiple image scales. MLEP effectively captures pixel dependencies across scales and dimensions while disrupting semantic content, thereby reducing potential content bias. Based on MLEP, we can easily build a robust CNN-based classifier capable of detecting AIGIs with enhanced reliability. Extensive experiments in an open-world setting, involving images synthesized by 32 distinct generative models, demonstrate that our approach achieves substantial improvements over state-of-the-art methods in both accuracy and generalization. Our code and models are available at https://www.github.com/fkeufss/MLEP/. Lin Yuan 0002, Xiaowan Li, Yan Zhang 0108, Jiawei Zhang 0011, Xinbo Gao 0001 |
NeurIPS | 3 |
| 2025 | See as You Desire: Scale-Adaptive Face Super-Resolution for Varying Low ResolutionsabstractFace super-resolution (FSR) is critical for bolstering intelligent security in Internet of Things (IoT) systems. Recent deep learning-driven FSR algorithms have attained remarkable progress. However, they always require separate model training and optimization for each scaling factor or input resolution, leading to inefficiency and impracticality. To overcome these limitations, we propose SAFNet, an innovative framework tailored for scale-adaptive FSR with arbitrary input resolution. SAFNet integrates scale information into representation learning to enable adaptive feature extraction and introduces dual-embedding attention to boost adaptive feature reconstruction. It leverages facial self-similarity and spatial-frequency collaboration to achieve precise scale-aware SR representations. This is attained through three key modules: 1) the scale adaption guidance unit (SAGU); 2) the scale-aware nonlocal self-similarity (SNLS) module; and 3) the spatial-frequency interactive modulation (SFIM) module. SAGU imports scaling factors using frequency encoding, SNLS exploits self-similarity to enrich feature representations, and SFIM incorporates spatial and frequency information to predict target pixel values adaptively. Comprehensive evaluations across four benchmark datasets reveal that SAFNet outperforms the second-best compared state-of-the-art (SOTA) method by about 0.2 dB/0.007 in PSNR/SSIM ($\times 4$on CelebA) with reduced 18.68%/42.64% computational complexity/time cost. This demonstrates SAFNet’s effectiveness and superiority, showcasing its potential as a promising solution for scale and input resolution adaptation challenges in FSR. The code will be available athttps://github.com/ICVIPLab/SAFNet. Yan Zhang 0108, Lin Yuan 0002, Xinbo Gao 0001 |
IEEE Internet Things J. | 2 |
| 2025 | Fed-MWFP: Lightweight federated learning with interpretable multiple wavelet fusion network for fault diagnosis under variable operating conditions
Yan Zhang 0108, Haitao Kong, Yan Han 0002 |
Knowl. Based Syst. | 1 |
| 2025 | SANet: Face super-resolution based on self-similarity prior and attention integration
Yan Zhang 0108, Lin Yuan 0002, Xinbo Gao 0001 |
Pattern Recognit. | 2 |
| 2025 | Deepfake Detection Leveraging Self-Blended Artifacts Guided by Facial Embedding DiscrepancyabstractCurrent deepfake detection methods commonly use data augmentation and authenticity-content disentanglement to extract more generalized features for detection tasks. However, these methods rely exclusively on low-level spatial artifacts to distinguish real from fake images, which presents significant challenges in accurately capturing the rich forgery cues. Deepfakes create discrepancies between forged and original facial features within the face-recognition (FR) embedding space, which can serve as an additional cue for detection. To better exploit the artifacts in deepfake images, we propose a novel detection method that enhances the detector’s perception capability by incorporating not only the real and fake samples during training, but also the visual residual between real and fake images. Meanwhile, we integrate the discrepancy in facial embedding between the real and fake samples into the training procedure of artifact extraction, serving as a guidance signal with strong knowledge provided by the pretrained face recognition model. Specialized distillation loss along with additional cross-entropy losses are designed to enhance the detection capability. Experiments on multiple benchmarks demonstrate the superiority of the proposed approach in deepfake detection over literature methods. Shuodi Wang, Lin Yuan 0002, Yan Zhang 0108, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | S²CMamba: A Mamba-Based Pansharpening Model Incorporating Spatial and Spectral ConsistencyabstractPrevailing pan-sharpening methods tackle the ill-posed challenge of reconstructing high-resolution multispectral (HRMS) images from low-resolution multispectral (LRMS) and panchromatic (PAN) inputs. This ill-posed nature introduces distortions such as blurred spatial edges and spectral color deviations, which compromise the fidelity of the reconstructed image. To address this, we propose S2CMamba, a dual-branch framework that leverages Mamba’s efficient contextual modeling and enforces spatial and spectral consistency through tailored priors, effectively alleviating the ill-posed nature of the pan-sharpening. The spatial context branch utilizes the Windowed Spatial Local Mamba (WSLM) for local details and the Global Spatial Interaction Mamba (GSIM) for long-range structures. Within WSLM, a Manifold Preservation (MP) constraint is proposed to align HRMS features with the low-dimensional manifold consistency of PAN and LRMS, thereby mitigating high-dimensional distortions and enhancing spatial consistency. Meanwhile, the spectral branch integrates multi-scale feature extraction and designed Spectral Context Mamba (SCM) to capture spectral context. Moreover, the spatial and spectral properties of multispectral images are investigated, and the wavelet transforms are introduced for better accomplish consistency. By incorporating contextual information, S2CMamba reduces the ambiguity of the solution space, while wavelet-based consistency constraints and designed MP prior further alleviate the ill-posed nature. Extensive experiments on benchmark datasets demonstrate that S2CMamba surpasses state-of-the-art methods, validating the efficacy of this approach in addressing the ill-posed nature of pan-sharpening tasks. Yan Zhang 0108, Yaohui Song, Qingyan Duan, Xinbo Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Jointly RS Image Deblurring and Super-Resolution With Adjustable-Kernel and Multi-Domain AttentionabstractRemote sensing (RS) image deblurring and super-resolution (SR) are common tasks in computer vision that aim at restoring RS image detail and spatial scale, respectively. However, real-world RS images often suffer from a complex combination of global low-resolution (LR) degeneration and local blurring degeneration. Although carefully designed deblurring and SR models perform well on these two tasks individually, a unified model that performs jointly RS image deblurring and SR (JRSIDSR) task is still challenging due to the vital dilemma of reconstructing the global and local degeneration simultaneously. In addition, existing methods struggle to capture the interrelationship between deblurring and SR processes, leading to suboptimal results. To tackle these issues, we give a unified theoretical analysis of RS images’ spatial and blur degeneration processes and propose a dual-branch parallel network named adjustable-kernel and multi-domain network (AKMD-Net) for the JRSIDSR task. AKMD-Net consists of two main branches: deblurring and SR branches. In the deblurring branch, we design a pixel-adjustable kernel block (PAKB) to estimate the local and spatial-varying blur kernels. In the SR branch, a multi-domain attention block (MDAB) is proposed to capture the global contextual information enhanced with high-frequency details. Furthermore, we develop an adaptive feature fusion (AFF) module to model the contextual relationships between the deblurring and SR branches. Finally, we design an adaptive Wiener loss (AW Loss) to depress the prior noise in the reconstructed images. Extensive experiments demonstrate that the proposed AKMD-Net achieves state-of-the-art (SOTA) quantitative and qualitative performance on commonly used RS image datasets. The source code is publicly available at:https://github.com/zpc456/AKMD-Net. Yan Zhang 0108, Chengxiao Zeng, Bin Xiao 0002, Xinbo Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Multi-scale and dynamic snake convolution-based YOLOv9 for steel surface defect detection
Junhua Chen 0003, Weilin Jin, Xueda Huang, Yan Zhang 0108 |
J. Supercomput. | 5 |
| 2024 | Structure-Aware in-Air Handwritten Text Recognition with Graph-Guided Cross-Modality TranslatorabstractIn-air handwriting as a new human-computer interaction way plays an important role in many virtual/mixed-reality applications. Existing methods for in-air handwritten text recognition (IAHTR) typically directly process handwriting trajectories with deep neural networks. However, those methods all simply learn discriminative patterns by modelling low-level relationships between adjacent points of trajectories, while completely ignoring the inherent geometric structures of characters. Instead, we propose a novel Graph-guided Cross-modality Translator for IAHTR, which further explicitly exploits the geometric structures of characters for guiding the decoding of trajectories via graph-guided cross-modality attention mechanism without introducing extra annotation costs. Experiments on benchmarks IAHEW-UCAS2016 & IAM-OnDB show that our method has achieved state-of-the-art performance for handwritten text recognition. Yuyan Chen, Ji Gan, Jiaxu Leng, Yan Zhang 0108, Xinbo Gao 0001 |
ICASSP | 5 |
| 2024 | 6TiSCH IIoT network: A review
Yan Zhang 0108, Haopeng Huang, Yan Han 0002 |
Comput. Networks | 1 |
| 2024 | DsP-YOLO: An anchor-free network with DsPAN for small object detection of multiscale defects
Yan Zhang 0108, Yan Han 0002, Minghang Zhao |
Expert Syst. Appl. | 1 |
| 2024 | AMCW-DFFNSA: An interpretable deep feature fusion network for noise-robust machinery fault diagnosis
Yan Han 0002, Sipeng Lv, Yan Zhang 0108 |
Knowl. Based Syst. | 4 |
| 2024 | MiC: Image-text Matching in Circles with cross-modal generative knowledge enhancement
Xiao Pu 0002, Lin Yuan 0002, Yan Zhang 0108, Liping Jing, Xinbo Gao 0001 |
Knowl. Based Syst. | 4 |
| 2024 | Invertible Image Obfuscation for Facial Privacy Protection via Secure FlowabstractThis paper presents a fresh paradigm for protecting facial privacy via an invertible image obfuscation framework that incorporates multiple characteristics including anonymity, diversity, reversibility, security, and lightweight all at once. We name the framework PRO-Face S, an acronym for Privacy-preserving Reversible Obfuscation of Face images via Secure flow. The core of the proposed framework is a flow-based generative model (or invertible neural network), which takes as input a face image along with its pre-obfuscated form, and outputs the privacy-protected image that visually mirrors the pre-obfuscated one. The pre-obfuscation applied can be in various forms with different types and strengths. The invertibility of the flow-based model ensures that the original image can be easily recovered from the protected image in high fidelity. An elaborate secret key mechanism is devised to securely guide the mutual transformations of privacy protection and image recovery, such that the correct recovery is only possible upon the availability of the correct secret, pre-specified by the user in the protection stage. Two modes of wrong recovery are investigated to deal with malicious recovery attempts in different scenarios. Finally, extensive experiments conducted on multiple image datasets demonstrate the superiority of the proposed framework over state-of-the-art methods. Lin Yuan 0002, Xiao Pu 0002, Yan Zhang 0108, Jiaxu Leng, Tao Wu 0003, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | PLGNet: Prior-Guided Local and Global Interactive Hybrid Network for Face Super-ResolutionabstractRecent CNN-driven face super-resolution (FSR) technologies have achieved excellent breakthroughs by incorporating facial prior knowledge. However, most of them suffer from some obvious limitations. They always estimate facial priors from input low-resolution (LR) faces or coarsely enhanced LR faces, obtaining unfaithful priors that cannot be adequately exploited. This may bring noticeable artifacts to the target results, especially for large scaling factors, deteriorating the fidelity and naturalness and generating suboptimal reconstructed results. In this paper, we propose a two-stage prior-guided FSR approach to learn facial prior knowledge from the optimal SR results of stage one and explore the complementarity between priors to further guide more accurate reconstruction in stage two. Specifically, we develop an efficient local and global interactive hybrid network incorporating facial semantic and geometric priors for more discriminative results. To reach this, we devise a multiscale interconnected symmetric encoder-decoder architecture composed of Prior Interaction-Integration Modules (PIIMs), the Coarse-to-fine Feature Refinement Module (CFRM), and Feature Aggregation Modulation Modules (FAMMs). The encoder concentrates on hierarchically extracting multiscale features. The CFRM is devised to explore the potential correlations between the encoder and the decoder and further guide the refinement and reinforcement of the encoded features. The decoder aims to take full advantage of informative multiscale encoded features to reconstruct high-quality SR representations. Comprehensive evaluation and visualization results on four benchmark datasets demonstrate the superiority of the proposed PLGNet over current state-of-the-art methods. The source code of PLGNet will be available at https://github.com/lil808/PLGNet.git. Yan Zhang 0108, Lin Yuan 0002, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | TCDM: Effective Large-Factor Image Super-Resolution via Texture Consistency DiffusionabstractRecently, remote sensing super-resolution (SR) tasks have been widely studied and achieved remarkable performance. However, due to the complex texture and serious image degeneration, the conventional methods (e.g. CNN-based, GAN-based) cannot reconstruct high-resolution (HR) remote sensing images with a large SR factor (≥ ×8). In this paper, we model the large-factor super-resolution (LFSR) task as a referenced diffusion process and explore how to embed pixel-wise constraint into the popular diffusion model. Following this motivation, we propose the first diffusion-based LFSR method named texture consistency diffusion model (TCDM) for remote sensing images. Specifically, we build a novel conditional truncated noise generator (CTNG) in TCDM to simultaneously generate the expectation of posterior probabilityp(xt-1|xt) and the truncated noise image. With the predicted truncated noise image, sampling an SR image using CTNG saves nearly 90% processing time compared to the naive diffusion model. Additionally, we design a new denoising process named texture consistency diffusion (TC-diffusion) to explicitly embed pixel-wise constraints into the LFSR diffusion model during the training stage. Universal experiments on five commonly used remote sensing datasets demonstrate that the proposed TCDM surpasses the latest SR methods by a large margin and reports new SOTA results on several evaluation metrics. Additionally, the proposed method demonstrates impressive visual quality on reconstructed remote sensing image texture and details. Yan Zhang 0108, Hanqi Liu, Xinbo Gao 0001, Guangyao Shi, Jianan Jiang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | PRO-Face C: Privacy-Preserving Recognition of Obfuscated Face via Feature CompensationabstractThe advancement of face recognition technology has delivered substantial societal advantages. However, it has also raised global privacy concerns due to the ubiquitous collection and potential misuse of individuals’ facial data. This presents a notable paradox: while there is a societal demand for a robust face recognition ecosystem to ensure public security and convenience, an increasing number of individuals are hesitant to release their facial data. Numerous studies have endeavored to find such a utility-privacy trade-off, yet many struggle with the dilemma of prioritizing one at the expense of the other. In response to this challenge, this paper proposes PRO-Face C, a novel paradigm for privacy-preserving recognition of obfuscated faces via a dedicated feature compensation mechanism, aimed at optimizing the equilibrium between privacy preservation and utility maximization. The proposed approach is characterized by a specialized client-server architecture: the client transmits only obfuscated images to the server, which then performs identity recognition using a pre-trained model in conjunction with a suite of privacy-free complementary features. This framework facilitates accurate face identification while safeguarding the original facial appearance from explicit disclosure. Furthermore, the obfuscated image retains its visualization capability, crucial for image preview functionalities. To ensure the desired properties, we have developed an identity-guided feature compensation mechanism, complemented by several privacy-enhancing techniques. Extensive experiments conducted across multiple face datasets underscore the effectiveness of the proposed approach in diverse scenarios. Lin Yuan 0002, Xiao Pu 0002, Yan Zhang 0108, Yushu Zhang 0001, Xinbo Gao 0001, Touradj Ebrahimi |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Contextual Learning in Fourier Complex Field for VHR Remote Sensing ImagesabstractVery high-resolution (VHR) remote sensing (RS) image classification is the fundamental task for RS image analysis and understanding. Recently, Transformer-based models demonstrated outstanding potential for learning high-order contextual relationships from natural images with general resolution ( pixels) and achieved remarkable results on general image classification tasks. However, the complexity of the naive Transformer grows quadratically with the increase in image size, which prevents Transformer-based models from VHR RS image ( pixels) classification and other computationally expensive downstream tasks. To this end, we propose to decompose the expensive self-attention (SA) into real and imaginary parts via discrete Fourier transform (DFT) and, therefore, propose an efficient complex SA (CSA) mechanism. Benefiting from the conjugated symmetric property of DFT, CSA is capable to model the high-order contextual information with less than half computations of naive SA. To overcome the gradient explosion in Fourier complex field, we replace the Softmax function with the carefully designed Logmax function to normalize the attention map of CSA and stabilize the gradient propagation. By stacking various layers of CSA blocks, we propose the Fourier complex Transformer (FCT) model to learn global contextual information from VHR aerial images following the hierarchical manners. Universal experiments conducted on commonly used RS classification datasets demonstrate the effectiveness and efficiency of FCT, especially on VHR RS images. The source code of FCT will be available at https://github.com/Gao-xiyuan/FCT. Yan Zhang 0108, Xiyuan Gao, Qingyan Duan, Jiaxu Leng, Xiao Pu 0002, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | DecomFormer: Decompose Self-Attention Via Fourier Transform for VHR Aerial Image Scene ClassificationabstractVery high-resolution (VHR) aerial image scene classification is an essential task for aerial image understanding. Although transformer-based models have demonstrated strong ability in natural image classification, transformer-based methods on VHR aerial image tasks are still lack of concern because the complexity of self-attention in the transformer grows quadratically with the image resolution. To address this issue, we decompose the self-attention via Fourier Transform and propose a novel Fourier self-attention (FSA) mechanism. Based on FSA, we design a highly efficient network named DecomFormer, which learns contextual relationships in the real part and imaginary part of the Fourier field, respectively. Theoretically, the DecomFormer reduces the complexity of the naive self-attention mechanism from O(n2) to O(nlog(n)). Universal experiments on public VHR aerial image classification benchmarks demonstrated the DecomFormer’s efficiency, especially on images with very high-resolution. Yan Zhang 0108, Xiyuan Gao, Xiao Pu 0002, Xinbo Gao 0001 |
ICASSP | 1 |
| 2023 | FCIR: Rethink Aerial Image Super Resolution with Fourier AnalysisabstractRecent years, deep-learning-based methods achieve remarkable improvements on the super-resolution (SR) task. However, recovering high-quality (HQ) texture from the low-quality (LQ) aerial image is still challenging due to the limited contextual modeling ability of current deep-learning methods as well as the sharp artificial texture of aerial images. In this paper, we rethink aerial image super resolution (AISR) task with the perspective of Fourier analysis. Firstly, we build the Fourier Global Convolution (FGC) inspired by the convolution theorem of the Fourier Transform to extract the shadow features. Then, following the Gabor Transform, a carefully designed oriented Texture Contextual Block (OTCB) is proposed to enhance the oriented texture representation. By stacking FGC and OTCB, we propose a simple but effective straight-forward network named Fourier Consistency Image Reconstruction Model (FCIR) to restore HQ aerial image. Moreover, we design a gradient consistency loss (GC Loss) to enhance the quality of reconstructed high-frequency details. Compared with very recent state-of-the-art super-resolution methods, experimental results demonstrate promising SR performance boosts from FCIR on 3 typical aerial image datasets. Yan Zhang 0108, Jianan Jiang, Xiao Pu 0002, Xinbo Gao 0001 |
ICASSP | 1 |
| 2023 | PSGAN: Revisit the binary discriminator and an alternative for face frontalization
Qingyan Duan, Lei Zhang 0038, Yan Zhang 0108, Xinbo Gao 0001 |
Neurocomputing | 3 |
| 2023 | Where to look: Multi-granularity occlusion aware for video person re-identification
Jiaxu Leng, Xinbo Gao 0001, Yan Zhang 0108, Ye Wang 0006, Mengjingcheng Mo |
Neurocomputing | 4 |
| 2023 | Intelligent Fault Identification for Industrial Internet of Things via Prototype-Guided Partial Domain Adaptation With Momentum WeightabstractPartial domain adaptation (PDA) for fault identification has been widely researched to help construct self-monitoring systems in the era of the Industrial Internet of Things (IIoT). However, the existing PDA fault identification methods neglect the influence of uncertainty of the target domain on the identification performance. To solve this problem, this work developed a prototype-guided PDA method with momentum weight for fault diagnosis. Specifically, to reduce the risk of ruling out the outlier by the output of a classifier or a discriminator, a classwise selectively source weighting strategy that follows the number of the target pseudo labels is proposed. The target instances’ pseudo labels, which are obtained by calculating the distance between the target instance and the source prototypes, are irrelevant to the classifier and discriminator. Furthermore, the momentum algorithm, by which the historical weights information could be retained, is employed in the source weights calculation procedure to alleviate the fluctuation and more closely to the global optimal. Experiments demonstrated the effectiveness and superiority of the developed method. Yan Han 0002, Yan Zhang 0108 |
IEEE Internet Things J. | 5 |
| 2023 | Target-Aware Transformer TrackingabstractObject tracking is aimed at locating a specific object in the image sequence, such as pedestrians, vehicles, and so on. The existing algorithms based on siamese neural network predict the target through similarity matching. Although these algorithms have achieved satisfactory performance, in the process of similarity calculation between template image and search image, only local information is often concerned, which makes the algorithms difficult to obtain the optimal solution. To deal with the abovementioned problems, we propose a model based on Transformer, named TaTrack. Specifically, we first use the encoders to enhance the features. Then, the dependency between template features and search features is established through the target-aware module. Finally, we utilize the classification regression network to locate the target, and use the classification score to adapt to update the template image. Experiments show that our model can achieve great performance on GOT-10k, LaSOT, and TrackingNet datasets. Yuhui Zheng, Yan Zhang 0108, Bin Xiao 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Boundary-Aware Bias Loss for Transformer-Based Aerial Image Segmentation ModelabstractInspired by the tremendous success of the transformer-based model in natural language processing (NLP), many efforts introduce the transformer-based model into the image processing tasks. However, naive transformer models have to down-sample the image resolution to satisfy computational restrictions, thus discarding the local information, which is catastrophic for high-performance remote sensing image segmentation. Hence, this paper proposes a novel trainable boundary-aware bias loss function to enhance transformer-based models of extracting local information. On the Challenging ISPRS Potsdam dataset, two representative transformer-based models achieve remarkable performance improvements, proving the effectiveness of the proposed method. Yan Zhang 0108, Siqi Liu 0008, Bo Hu 0008, Xinbo Gao 0001 |
ICASSP | 1 |
| 2022 | Sampling-invariant fully metric learning for few-shot object detection
Jiaxu Leng, Taiyue Chen, Xinbo Gao 0001, Mengjingcheng Mo, Yongtao Yu, Yan Zhang 0108 |
Neurocomputing | 6 |
| 2022 | DHT: Deformable Hybrid Transformer for Aerial Image SegmentationabstractDue to the strong ability to model global information, the transformer-based methods have shown remarkable improvements in image segmentation tasks. However, the self-attention mechanism in the transformer is computationally expensive and relies on pre-trained parameters. Moreover, the transformer method is weak in modeling local information, which is unfavorable for accurately segmenting objects from high-resolution aerial images. To this end, an efficient deformable orientational self-attention (DoA) is proposed to simultaneously extract the global information and the local information. Besides, for parameter efficiency, we design a depthwise channel self-attention (DcA) to model the contextual information among channels. Combining with the DoA and DcA, we propose the deformable hybrid transformer (DHT) to perform high-quality object segmentation on aerial images. Experiments on ISPRS Potsdam dataset and WHU building dataset illustrate that the proposed DHT can not only achieve state-of-the-art (SOTA) results but also markedly reduce the dependence of the transformer on pre-trained parameters. Yan Zhang 0108, Xiyuan Gao, Qingyan Duan, Lin Yuan 0002, Xinbo Gao 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |