Xiaofeng Zhang 0006

dblp:61/3976-6 · DBLP profile ↗
← Back
29ranked-venue papers
7as first author
29since 2021 · last 2026
0000-0002-7185-4682ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 5 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 D3ToM: Decider-Guided Dynamic Token Merging for Accelerating Diffusion MLLMs
abstract
Diffusion-based multimodal large language models (Diffusion MLLMs) have recently demonstrated impressive non-autoregressive generative capabilities across vision-and-language tasks. However, Diffusion MLLMs exhibit substantially slower inference than autoregressive models: Each denoising step employs full bidirectional self-attention over the entire sequence, resulting in cubic decoding complexity that becomes computationally impractical with thousands of visual tokens. To address this challenge, we propose D³ToM, a Decider-guided dynamic token merging method that dynamically merges redundant visual tokens at different denoising steps to accelerate inference in Diffusion MLLMs. At each denoising step, D³ToM uses decider tokens—the tokens generated in the previous denoising step—to build an importance map over all visual tokens. Then it maintains a proportion of the most salient tokens and merges the remainder through similarity-based aggregation. This plug-and-play module integrates into a single transformer layer, physically shortening the visual token sequence for all subsequent layers without altering model parameters. Moreover, D³ToM employs a merge ratio that dynamically varies with each denoising step, aligns with the native decoding process of Diffusion MLLMs, achieving superior performance under equivalent computational budgets. Extensive experiments show that D³ToM accelerates inference while preserving competitive performance.
Shuochen Chang, Xiaofeng Zhang 0006, Qingyang Liu 0008, Li Niu 0002
AAAI2
2026 Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to Intervention
abstract
Shuochen Chang, Tong Bai, Xiaofeng Zhang, Qianli Ma, Qingyang Liu, Zhaohe Liao, Yibo Miao, Li Niu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Shuochen Chang, Tong Bai, Xiaofeng Zhang 0006, Qianli Ma 0008, Qingyang Liu 0008, Zhaohe Liao, Yibo Miao, Li Niu 0002
ACL (1)3
2026 ART: Attention Replacement Technique to Improve Factuality in LLMs
abstract
Hallucination in large language models (LLMs) continues to be a significant issue, particularly in tasks like question answering, where models often generate plausible yet incorrect or irrelevant information.Although various methods have been proposed to mitigate hallucinations, the relationship between attention patterns and hallucinations has not been fully explored.In this paper, we analyze the distribution of attention scores across each layer and attention head of LLMs, revealing a common and intriguing phenomenon: Shallow layers of LLMs primarily rely on uniform attention patterns, where the model distributes its attention evenly across the entire sequence.This uniform attention pattern can lead to hallucinations, as the model fails to focus on the most relevant information.To mitigate this issue, we propose a trainingfree method called Attention Replacement Technique (ART), which replaces these uniform attention patterns in the shallow layers with local attention patterns.This change directs the model to focus more on the relevant contexts, thus reducing hallucinations.Through extensive experiments, ART demonstrates significant reductions in hallucinations across multiple LLM architectures, proving its effectiveness and generalizability without requiring fine-tuning or additional training data.
Ziqin Luo, Yihao Quan, Xiaofeng Zhang 0006, Xiaosong Yuan, Chen Shen 0003
ACL (1)3
2026 Reasoning Fails Where Step Flow Breaks
abstract
Xiaoyu Xu, Yulan Pan, Xiaosong Yuan, Zhihong Shen, Minghao Su, Yuanhao Su, Xiaofeng Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yulan Pan, Xiaosong Yuan, Zhihong Shen, Minghao Su, Yuanhao Su, Xiaofeng Zhang 0006
ACL (1)7
2026 SeaRAG: Reducing Hallucination in Retrieval-Augmented Generation via Statement-Entity Adaptive Ranking
Xiaosong Yuan, Xiaofeng Zhang 0006, Yijia Zhang 0003, Ying Wang 0009
WWW2
2026 CCSFusion: A Hierarchical Semantic Chain-of-Thought Reasoning Architecture for Infrared-Visible Image Fusion and Captioning
abstract
Infrared-Visible Image Fusion (IVIF) aims to generate a single, information-rich image for downstream tasks. However, prevailing methods exhibit two key limitations. First, many approaches lack explicit hierarchical semantic decoupling, failing to effectively integrate semantic features across different levels, which restricts their ability to capture complex scene structures. Second, task-driven fusion frameworks typically adopt a cascaded design, with unidirectional supervision provided by geometry-centric downstream tasks like detection. This architecture not only limits mutual reinforcement between the fusion and task networks, but also creates a ”supervision bottleneck” by lacking interaction with the linguistic modality that captures richer scene relationships. To tackle these challenges, we propose CCSFusion, the first framework that leverages Chain-of-Thought captioning as supervision, redirecting IVIF optimization from narrow geometric accuracy to multimodal scene comprehension. It establishes a mutually reinforcing coupling between the fusion network and the captioning task. Specifically, we introduce a Segmentation Mask Calibration Unit (SMCU) to refine coarse semantic priors, providing precise pixel-level guidance. Subsequently, the calibrated features are fed into Chained Semantic Fusion Module (CSFM) which explicitly decomposes the semantic priors into three hierarchical levels, and then feeds them into the Hierarchical Semantic Attention module. Finally, a bidirectional knowledge distillation mechanism transfers the reasoning ability of the teacher network to the student. Experiments show that CCSFusion achieves superior fusion performance and generates more semantically coherent images for high-level cognitive tasks. The code is available at: https://github.com/Snaillms/CCSFusion.
Miaoshan Lin, Guoheng Huang, Jietao Yang, Jiehao Zheng, Xiaochen Yuan, Yan Li 0122, Xiaofeng Zhang 0006, Kim Fung Tsang, Chi-Man Pun
IEEE Internet Things J.7
2026 MilleniaGuard: An Event-Driven Edge-AI and AIGC-Based IoT System for Ancient Mural Monitoring and Restoration
abstract
This paper addresses the challenges of automatic monitoring and restoration in ancient mural conservation, aiming to enhance the efficiency and quality of heritage preservation. Traditional manual inspection is time-consuming and often misses early damage, while existing digital restoration models struggle with consistent restoration, especially for large-scale damage. To address these issues, we propose an Internet of things (IoT)-based solution combining event-driven edge intelligence and artificial intelligence generated content (AIGC) techniques. A fine-tuned EdgeSAM model, using a Conv-adapter, enables efficient damage segmentation at the edge; an event-driven mechanism reduces resource consumption; and a LoRA-tuned PowerPaint model, aided by Blip2 and Qwen, provides effective restoration of large damaged areas. Cloud-side processing utilizes AIGC techniques to restore damaged mural areas, ensuring high-quality restoration while minimizing communication demands. Experimental results demonstrate that the proposed method achieves accurate damage monitoring on resource-constrained edge devices and generates diverse, contextually appropriate restoration results on cloud servers, providing a deployment-oriented feasibility validation under simulated temporal degradation and real hardware constraints.
Zishan Xu, Jiansen Zhang, Wei Chen 0036, Xiaofeng Zhang 0006, Jueting Liu, Zehua Wang 0001, F. Richard Yu, Victor C. M. Leung
IEEE Internet Things J.4
2026 What drives attention sinks? A study of massive activations and rotational positional encoding in large vision-language models
Xiaofeng Zhang 0006, Yuanchao Zhu, Chaochen Gu, Hao Cheng 0004, Kaijie Wu 0002
Inf. Process. Manag.1
2026 QWNet: A quaternion wavelet network for spatial-frequency aware multi-modal image fusion
Jietao Yang, Miaoshan Lin, Guoheng Huang, Xuhang Chen 0002, Xiaofeng Zhang 0006, Xiaochen Yuan, Chi-Man Pun, Bingo Wing-Kuen Ling
Neural Networks5
2025 Shallow Focus, Deep Fixes: Enhancing Shallow Layers Vision Attention Sinks to Alleviate Hallucination in LVLMs
abstract
Xiaofeng Zhang, Yihao Quan, Chen Shen, Chaochen Gu, Xiaosong Yuan, Shaotian Yan, Jiawei Cao, Hao Cheng, Kaijie Wu, Jieping Ye. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Xiaofeng Zhang 0006, Yihao Quan, Chen Shen 0003, Chaochen Gu, Xiaosong Yuan, Shaotian Yan, Hao Cheng 0004, Kaijie Wu 0002, Jieping Ye
EMNLP1
2025 Improving Complex Reasoning with Dynamic Prompt Corruption: A Soft Prompt Optimization Approach
abstract
Prompt Tuning (PT) has emerged as a promising Parameter-Efficient Fine-Tuning (PEFT) approach by appending trainable continuous prompt vectors to the input, maintaining competitive performance with significantly fewer trainable parameters. While PT has shown effectiveness in enhancing task performance, particularly for classification tasks, its application to complex reasoning tasks has been largely overlooked. Our investigation reveals that PT provides limited improvement and may even degrade performance in reasoning tasks. This phenomenon suggests that soft prompts can positively impact certain instances while negatively affecting others, particularly during the latter stages of reasoning. To address these challenges, we propose a novel method called Dynamic Prompt Corruption (DPC), which seeks to optimize the use of soft prompts in reasoning tasks. DPC dynamically adjusts the influence of soft prompts based on their impact on the reasoning process. Specifically, it involves two key components: Dynamic Trigger and Dynamic Corruption. Dynamic Trigger measures the influence of soft prompts, determining whether their impact is beneficial or detrimental. Dynamic Corruption mitigates the negative effects of soft prompts by selectively masking key tokens that interfere with the reasoning process. We validate our approach through extensive experiments on various large language models (LLMs) and reasoning tasks, including GSM8K, MATH, and AQuA. The results demonstrate that Dynamic Prompt Corruption consistently improves the performance of LLMs, achieving 4\%-8\% accuracy gains compared to standard prompt tuning. These findings highlight the effectiveness of our approach and its potential to enhance complex reasoning in LLMs.
Sinan Fan, Liang Xie 0003, Chen Shen 0003, Ge Teng, Xiaosong Yuan, Xiaofeng Zhang 0006, Chenxi Huang 0004, Wenxiao Wang 0001, Xiaofei He 0001, Jieping Ye
ICLR6
2025 EFDTR: Learnable Elliptical Fourier Descriptor Transformer for Instance Segmentation
abstract
Polygon-based object representations efficiently model object boundaries but are limited by high optimization complexity, which hinders their adoption compared to more flexible pixel-based methods. In this paper, we introduce a novel vertex regression loss grounded in Fourier elliptic descriptors, which removes the need for rasterization or heuristic approximations and resolves ambiguities in boundary point assignment through frequency-domain matching. To advance polygon-based instance segmentation, we further propose EFDTR (\textbf{E}lliptical \textbf{F}ourier \textbf{D}escriptor \textbf{Tr}ansformer), an end-to-end learnable framework that leverages the expressiveness of Fourier-based representations. The model achieves precise contour predictions through a two-stage approach: the first stage predicts elliptical Fourier descriptors for global contour modeling, while the second stage refines contours for fine-grained accuracy. Experimental results on the COCO dataset show that EFDTR outperforms existing polygon-based methods, offering a promising alternative to pixel-based approaches. Code is available at \url{https://github.com/chrisclear3/EFDTR}.
Chaochen Gu, Hao Cheng 0004, Xiaofeng Zhang 0006, Kaijie Wu 0002, Changsheng Lu
ICML4
2025 LensNet: An End-to-End Learning Framework for Empirical Point Spread Function Modeling and Lensless Imaging Reconstruction
abstract
Lensless imaging stands out as a promising alternative to conventional lens-based systems, particularly in scenarios demanding ultracompact form factors and cost-effective architectures. However, such systems are fundamentally governed by the Point Spread Function (PSF), which dictates how a point source contributes to the final captured signal. Traditional lensless techniques often require explicit calibrations and extensive pre-processing, relying on static or approximate PSF models. These rigid strategies can result in limited adaptability to real-world challenges, including noise, system imperfections, and dynamic scene variations, thus impeding high-fidelity reconstruction. In this paper, we propose LensNet, an end-to-end deep learning framework that integrates spatial-domain and frequency-domain representations in a unified pipeline. Central to our approach is a learnable Coded Mask Simulator (CMS) that enables dynamic, data-driven estimation of the PSF during training, effectively mitigating the shortcomings of fixed or sparsely calibrated kernels. By embedding a Wiener filtering component, LensNet refines global structure and restores fine-scale details, thus alleviating the dependency on multiple handcrafted pre-processing steps. Extensive experiments demonstrate LensNet's robust performance and superior reconstruction quality compared to state-of-the-art methods, particularly in preserving high-frequency details and attenuating noise. The proposed framework establishes a novel convergence between physics-based modeling and data-driven learning, paving the way for more accurate, flexible, and practical lensless imaging solutions for applications ranging from miniature sensors to medical diagnostics. The link of code is https://github.com/baijiesong/Lensnet.
Jiesong Bai, Yuhao Yin, Yihang Dong, Xiaofeng Zhang 0006, Chi-Man Pun, Xuhang Chen 0002
IJCAI4
2025 Longitudinal MRI-Clinical Multimodal Fusion for pCR Prediction in Breast Cancer
Dingrui Ma, Hao Cheng 0004, Xiaofeng Zhang 0006, Kaijie Wu 0002, Chaochen Gu, Xin-Ping Guan
MICCAI (15)6
2025 MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language Models
abstract
Hallucinations pose a significant challenge in Large Vision Language Models (LVLMs), with misalignment between multimodal features identified as a key contributing factor. This paper reveals the negative impact of the long-term decay in Rotary Position Encoding (RoPE), used for positional modeling in LVLMs, on multimodal alignment. Concretely, under long-term decay, instruction tokens exhibit uneven perception of image tokens located at different positions within the two-dimensional space: prioritizing image tokens from the bottom-right region since in the one-dimensional sequence, these tokens are positionally closer to the instruction tokens. This biased perception leads to insufficient image-instruction interaction and suboptimal multimodal alignment. We refer to this phenomenon as ''image alignment bias.'' To enhance instruction's perception of image tokens at different spatial locations, we propose MCA-LLaVA, based on Manhattan distance, which extends the long-term decay to a two-dimensional, multi-directional spatial decay. MCA-LLaVA integrates the one-dimensional sequence order and two-dimensional spatial position of image tokens for positional modeling, mitigating hallucinations by alleviating image alignment bias. Experimental results of MCA-LLaVA across various hallucination and general benchmarks demonstrate its effectiveness and generality. The code can be accessed in https://github.com/ErikZ719/MCA-LLaVA.
Qiyan Zhao, Xiaofeng Zhang 0006, Yun Xing 0001, Xiaosong Yuan, Sinan Fan, Xuhang Chen 0002, Dahan Wang, Xu-Yao Zhang
ACM Multimedia2
2025 From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks
abstract
Xiaofeng Zhang, Yihao Quan, Chen Shen, Xiaosong Yuan, Shaotian Yan, Liang Xie, Wenxiao Wang, Chaochen Gu, Hao Tang, Jieping Ye. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Xiaofeng Zhang 0006, Yihao Quan, Chen Shen 0003, Xiaosong Yuan, Shaotian Yan, Liang Xie 0003, Wenxiao Wang 0001, Chaochen Gu, Hao Tang 0005, Jieping Ye
NAACL (Long Papers)1
2025 High-Fidelity Document Stain Removal via A Large-Scale Real-World Dataset and A Memory-Augmented Transformer
abstract
Document images are often degraded by various stains, significantly impacting their readability and hindering downstream applications such as document digitization and analysis. The absence of a comprehensive stained document dataset has limited the effectiveness of existing document enhancement methods in removing stains while preserving fine-grained details. To address this challenge, we construct StainDoc, the first large-scale, high-resolution (2145 x 2245) dataset specifically designed for document stain removal. StainDoc comprises over 5,000 pairs of stained and clean document images across multiple scenes. This dataset encompasses a diverse range of stain types, severities, and document backgrounds, facilitating robust training and evaluation of document stain removal algorithms. Furthermore, we propose StainRestorer, a Transformer-based document stain removal approach. StainRestorer employs a memory-augmented Transformer architecture that captures hierarchical stain representations at part, instance, and semantic levels via the DocMemory module. The Stain Removal Transformer (SRTransformer) leverages these feature representations through a dual attention mechanism: an enhanced spatial attention with an expanded receptive field, and a channel attention captures channel-wise feature importance. This combination enables precise stain removal while preserving document content integrity. Extensive experiments demonstrate StainRestorer's superior performance over state-of-the-art methods on the Stain-Doc dataset and its variants StainDoc.Mark and Stain-Doc.Seal, establishing a new benchmark for document stain removal. Our work highlights the potential of memory-augmented Transformers for this task and contributes a valuable dataset to advance future research.
Mingxian Li, Yingtie Lei, Xiaofeng Zhang 0006, Yihang Dong, Zimeng Li 0001, Xuhang Chen 0002
WACV4
2025 Simignore: Exploring and enhancing multimodal large model complex reasoning via similarity computation
Xiaofeng Zhang 0006, Fanshuo Zeng, Chaochen Gu
Neural Networks1
2025 MuralAgent: Enhancing Ancient Mural Outpainting with RAG-Based Texts and Multimodal Integration
abstract
In the context of the digital age, utilizing cutting-edge technology for the digitization and creative expansion of ancient murals is crucial, aimed at preserving and passing on cultural heritage. Existing image outpainting techniques suffer from a lack of semantic guidance. This article introduces MuralAgent, a multimodal model based on Retrieval-Augmented Generation (RAG) technology. It precisely extracts key information from mural images and integrates it with a constructed ancient texts knowledge base to ensure the cultural and semantic consistency of the expanded images. Moreover, fine-tuning the Stable Diffusion model ensures the fidelity of the generated image styles. Specifically, this study involves constructing an ancient texts knowledge base for accurate matching, designing specific prompts for GPT-4V(ision) to extract key information, and innovatively expanding artworks through Stable Diffusion, providing a novel way for the public to reinterpret ancient murals.
Zishan Xu, Xiaofeng Zhang 0006, Wei Chen 0036, Jueting Liu, Zehua Wang 0001, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.2
2025 Wakeup-Darkness: When Multimodal Meets Unsupervised Low-Light Image Enhancement
abstract
Low-light image enhancement is a crucial visual task, and many unsupervised methods overlook the degradation of visible information in low-light scenes, adversely affecting the fusion of complementary information and hindering the generation of satisfactory results. To address this, we introduce Wakeup-Darkness, a multimodal enhancement framework that innovatively enriches user interaction through voice and textual commands. This approach signifies a technical leap and represents a paradigm shift in user engagement. We introduce a Cross-Modal Feature Fusion (CMFF) that synergizes semantic and depth context with low-light enhancement operations. Moreover, we propose a Gated Residual Block (GRB) and a channel-aware Look-Up Table (LUT) to adjust the intensity distribution of each channel. Crucially, the proposed Wakeup-Darkness scheme demonstrates remarkable generalization in unsupervised scenarios. The source code can be accessed from https://github.com/zhangbaijin/Wakeup-Dakness .
Xiaofeng Zhang 0006, Zishan Xu, Hao Tang 0005, Chaochen Gu, Wei Chen 0036, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.1
2024 DOPRA: Decoding Over-accumulation Penalization and Re-allocation in Specific Weighting Layer
abstract
In this work, we introduce DOPRA, a novel approach designed to mitigate hallucinations in multi-modal large language models (MLLMs). Unlike existing solutions that typically involve costly supplementary training data or the integration of external knowledge sources, DOPRA innovatively addresses hallucinations by decoding specific weighted layer penalties and redistribution, offering an economical and effective solution without additional resources. DOPRA is grounded in unique insights into the intrinsic mechanisms controlling hallucinations within MLLMs, especially the models' tendency to over-rely on a subset of summary tokens in the self-attention matrix, neglecting critical image-related information. This phenomenon is particularly pronounced in certain strata. To counteract this over-reliance, DOPRA employs a strategy of weighted overlay penalties and redistribution in specific layers, such as the 12th layer, during the decoding process. Furthermore, DOPRA includes a retrospective allocation process that re-examines the sequence of generated tokens, allowing the algorithm to reallocate token selection to better align with the actual image content, thereby reducing the incidence of hallucinatory descriptions in auto-generated captions. Overall, DOPRA represents a significant step forward in improving the output quality of MLLMs by systematically reducing hallucinations through targeted adjustments during the decoding process.
Jinfeng Wei, Xiaofeng Zhang 0006
ACM Multimedia2
2024 Instance-adaptive Zero-shot Chain-of-Thought Prompting
abstract
Zero-shot Chain-of-Thought (CoT) prompting emerges as a simple and effective strategy for enhancing the performance of large language models (LLMs) in real-world reasoning tasks. Nonetheless, the efficacy of a singular, task-level prompt uniformly applied across the whole of instances is inherently limited since one prompt cannot be a good partner for all, a more appropriate approach should consider the interaction between the prompt and each instance meticulously. This work introduces an instance-adaptive prompting algorithm as an alternative zero-shot CoT reasoning scheme by adaptively differentiating good and bad prompts. Concretely, we first employ analysis on LLMs through the lens of information flow to detect the mechanism under zero-shot CoT reasoning, in which we discover that information flows from question to prompt and question to rationale jointly influence the reasoning results most. We notice that a better zero-shot CoT reasoning needs the prompt to obtain semantic information from the question then the rationale aggregates sufficient information from the question directly and via the prompt indirectly. On the contrary, lacking any of those would probably lead to a bad one. Stem from that, we further propose an instance-adaptive prompting strategy (IAP) for zero-shot CoT reasoning. Experiments conducted with LLaMA-2, LLaMA-3, and Qwen on math, logic, and commonsense reasoning tasks (e.g., GSM8K, MMLU, Causal Judgement) obtain consistent improvement, demonstrating that the instance-adaptive zero-shot CoT prompting performs better than other task-level methods with some curated prompts or sophisticated procedures, showing the significance of our findings in the zero-shot CoT reasoning mechanism.
Xiaosong Yuan, Chen Shen 0003, Shaotian Yan, Xiaofeng Zhang 0006, Liang Xie 0003, Wenxiao Wang 0001, Renchu Guan, Ying Wang 0009, Jieping Ye
NeurIPS4
2024 LL-Diff: Low-Light Image Enhancement Utilizing Langevin Sampling Diffusion
abstract
In this paper, we propose a new algorithm called LL-Diff, which is innovative compared to traditional augmentation methods in that it introduces the sampling method of Langevin dynamics. This sampling approach simulates the motion of particles in complex environments and can better handle noise and details in low-light conditions. We also incorporate a causal attention mechanism to achieve causality and address the issue of confounding effects. This attention mechanism enables us to better capture local information while avoiding over-enhancement. We have conducted experiments on the LOL-V1 and LOL-V2 datasets, and the results show that LL-Diff significantly improves computational speed and several evaluation metrics, demonstrating the superiority and effectiveness of our method for low-light image enhancement tasks. The code will be released on GitHub when the paper has been accepted.
Boren Ding, Xiaofeng Zhang 0006, Zekun Yu, Zheng Hui
Int. J. Pattern Recognit. Artif. Intell.2
2024 Shadclips: When Parameter-Efficient Fine-Tuning with Multimodal Meets Shadow Removal
abstract
Segment Anything Model (SAM), an advanced universal image segmentation model trained on an expansive visual dataset, has set a new benchmark in image segmentation and computer vision. However, it faced challenges when it came to distinguishing between shadows and their backgrounds. To address this, we proposed ShadClips, which consists of SAM-optimizer and SONet. It has dramatically enhanced SAM’s ability to segment shadow images, differentiating between the background and both soft and hard shadows adeptly. Due to its dependence on pixel point inputs, the SAM-Optimizer interface could do better. This method presents challenges, especially when dealing with long, extended shadows. To make the user experience more intuitive and effective, we incorporated the capabilities of CLIPs. Therefore, simple text descriptions like “A photo of a shadow” can be used to guide the SAM-Optimizer, allowing it to select the most relevant shadow mask from SAM’s comprehensive category list. Meanwhile, we introduce SONet to shadow removal. A large number of experiments on ISTD/SRD prove that the proposed method is effective and satisfactory. The source code of the ShadClips can be accessed from https://github.com/zhangbaijin/SAM-helps-Shadow .
Xiaofeng Zhang 0006, Chaochen Gu, Zishan Xu, Hao Tang 0005, Hao Cheng 0004, Kaijie Wu 0002, Shanying Zhu
Int. J. Pattern Recognit. Artif. Intell.1
2024 Efficient Remote-Sensing Segmentation With Generative Adversarial Transformer
abstract
Most deep-learning methods that achieve high segmentation accuracy require deep network architectures that are too heavy and complex to run on embedded devices with limited storage and memory space. To address this issue, this letter proposes an efficient generative adversarial transformer (GATrans) for achieving high-precision semantic segmentation while maintaining an extremely efficient size. The framework utilizes a global transformer network (GTNet) as the generator, efficiently extracting multilevel features through residual connections. GTNet employs global transformer blocks with progressively linear computational complexity to reassign global features based on a learnable similarity function. To focus on object- and pixel-level information, the GATrans optimizes the objective function by combining structural similarity losses. We validate the effectiveness of our approach through extensive experiments on the Vaihingen dataset, achieving an average$F1$score of 90.17% and an overall accuracy (OA) of 91.92%. Codes are available athttps://github.com/qiuluyi/GATrans.
Luyi Qiu, Dayu Yu, Xiaofeng Zhang 0006
IEEE Geosci. Remote. Sens. Lett.3
2023 SpA-Former:An Effective and lightweight Transformer for image shadow removal
abstract
In this paper, we propose an Effective and lightweight Transformer for image shadow detection and removal named SpA-Former to recover a shadow-free image from a single shaded image. In contrast to conventional methods that require two stages for shadow detection and then shadow removal, the SpA-Former is a one-stage network capable of learning the mapping function between shadows and no shadows, and does not require a separate shadow detection. SpA-Former is composed of Transformer encoder and CNN decoder, where the CNN decoder contains the GAN network. In the Transformer encoding stage, Gated Feed-Forward Network(GFFN) is devised to control the information flow. In the CNN decoding stage, Two-wheel RNN joint spatial attention(TWRNN) and Fourier transform residual block (FTR) are designed to achieve satisfactory results in shadow removal. The combination of Transformer and CNN is able to feed global features from the Vision Transformer encoder into CNN to enhance the global perception of CNN branches, taking into account the complementarity of local features and the global. The SpA-Former's inference speed is 0.0459s, and the final Parameters and FLOPS are only 0.47MB and 15G, achieving the current lightweight of image shadow removal. The source code of MemoryNet can be obtained from https://github.com/zhangbaijin/SpA-Former-shadow-removal
Xiaofeng Zhang 0006, Yudi Zhao, Chaochen Gu, Changsheng Lu, Shanying Zhu
IJCNN1
2023 A Semantics-Geometry Framework for Road Extraction From Remote Sensing Images
abstract
Road extraction from remote sensing images in very high resolution is important for autonomous driving and road planning. Compared with large-scale objects, roads are smaller, winding, and likely to be covered by buildings’ shadows, causing deep convolutional neural networks (DCNNs) to be difficult to identify roads. The paper proposes a semantics-geometry framework (SGNet) with a two-branch backbone, i.e., semantics-dominant branch and geometry-dominant branch. The semantics-dominant branch inputs images to predict dense semantic features, and the geometry-dominant branch takes images to generate sparse boundary features. Then, dense semantic features and boundary details generated by two branches are adaptively fused. Further, by utilizing affinity between neighborhood pixels, a feature refinement module is proposed to refine textures and road details. We evaluate the SGNet on the Ottawa road dataset. Experiments show that the SGNet outperforms other competitors on the road extraction task. Codes is available at https://github.com/qiuluyi/SGNet.
Luyi Qiu, Dayu Yu, Xiaofeng Zhang 0006
IEEE Geosci. Remote. Sens. Lett.4
2022 AFFNet: Attention Mechanism Network Based on Fusion Feature for Image Cloud Removal
abstract
Clouds frequently affect optical remote sensing pictures throughout the gathering process, resulting in low-resolution images that affect judgment and subsequent use of ground data. Because of the thick cloud cover, the ground surface information below is entirely incorrect. This kind of end-to-end image problem should not be dismissed as a simple task of image inpainting or image translation. Therefore, this paper proposes a multi-head self-attention module based on the encoding–decoding generative adversarial network, considering the redundant information of the deep network, furthermore this paper introduces Ghost convolution to effectively solve the influence of redundant feature maps in the network on the increase of time consumption and parameters. The method in this paper can solve the problem of cloud occlusion. By considering spatial information, it can better complete the prediction of cloud removal. It can reduce the amount of network calculations and parameters while maintaining the effect. In addition, Feature Fusion Module is proposed to integrate high-level features with low-level features, so that the network can extract enough feature information and better supplement the details to complete the cloud removal. The method in this paper has achieved excellent results on the RICE1 and RICE2 datasets.
Runhan Shen, Xiaofeng Zhang 0006, Yonggang Xiang
Int. J. Pattern Recognit. Artif. Intell.2
2021 RGAN: Rethinking generative adversarial networks for cloud removal
abstract
Optical remote sensing imagery is at the core of many Earth observation activities. Many applications take use of the satellite data's regular, consistent, and global-scale characteristics, such as farmland monitoring, climate change assessment, land-cover, and land-use categorization, and catastrophe assessment. Optical remote sensing images, on the other hand, are frequently impacted by clouds during the collection process, resulting in reduced image clarity, which impairs feature assessment and future usage, and heavy cloud blockage renders the surface information below totally useless. In this paper, We propose a soft attention recurrent neural module based on an encoder-decoder network, which can solve the cloud occlusion problem. We also propose an adaptive padding convolution at the end of the decoder by taking into account the spatial information, which results in better declouding predictions, and our network achieves good results on the RICE1 and RICE2 data sets.
Xinyu Ran, Xiaofeng Zhang 0006
Int. J. Intell. Syst.3