EDBT 2026 Demo / reviewers in the wild / expert
Lefei Zhang
dblp:28/10770
· DBLP profile ↗
209ranked-venue papers
14as first author
118since 2021 · last 2026
0000-0003-0542-2280ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 117 · 7 first-author · 66 since 2021Graphics, computer vision, multimedia, augmented reality and games · 81 · 3 first-author · 52 since 2021Applied, interdisciplinary, general and emerging computing · 48 · 5 first-author · 25 since 2021Databases, data management, data science and information retrieval · 12 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OFL-SAM2: Prompt SAM2 with Online Few-shot Learner for Efficient Medical Image SegmentationabstractThe Segment Anything Model 2 (SAM2) has demonstrated remarkable promptable visual segmentation capabilities in video data, showing potential for extension to medical image segmentation (MIS) tasks involving 3D volumes and temporally correlated 2D image sequences. However, adapting SAM2 to MIS presents several challenges, including the need for extensive annotated medical data for fine-tuning and high-quality manual prompts, which are both labor-intensive and require intervention from medical experts. To address these challenges, we introduce OFL-SAM2, a prompt-free SAM2 framework for label-efficient MIS. Our core idea is to leverage limited annotated samples to train a lightweight mapping network that captures medical knowledge and transforms generic image features into target features, thereby providing additional discriminative target representations for each frame and eliminating the need for manual prompts. Crucially, the mapping network supports online parameter update during inference, enhancing the model’s generalization across test sequences. Technically, we introduce two key components: (1) an online few-shot learner that trains the mapping network to generate target features using limited data, and (2) an adaptive fusion module that dynamically integrates the target features with the memory-attention features generated by frozen SAM2, leading to accurate and robust target representation. Extensive experiments on three diverse MIS datasets demonstrate that OFL-SAM2 achieves state-of-the-art performance with limited training data. Meng Lan, Lefei Zhang |
AAAI | 2 |
| 2026 | S5: Scalable Semi-Supervised Semantic Segmentation in Remote SensingabstractSemi-supervised semantic segmentation (S4) has advanced remote sensing (RS) analysis by leveraging unlabeled data through pseudo-labeling and consistency learning. However, existing S4 studies often rely on small-scale datasets and models, limiting their practical applicability. To address this, we propose S5, the first scalable framework for semi-supervised semantic segmentation in RS, which unlocks the potential of vast unlabeled Earth observation data typically underutilized due to costly pixel-level annotations. Built upon existing large-scale RS datasets, S5 introduces a data selection strategy that integrates entropy-based filtering and diversity expansion, resulting in the RS4P-1M dataset. Using this dataset, we systematically scale up S4 into a new pretraining paradigm, S4 pre-training (S4P), to pretrain RS foundation models (RSFMs) of varying sizes on this extensive corpus, significantly boosting their performance on land cover segmentation and object detection tasks. Furthermore, during fine-tuning, we incorporate a Mixture-of-Experts (MoE)-based multi-dataset fine-tuning approach, which enables efficient adaptation to multiple RS benchmarks with fewer parameters. This approach improves the generalization and versatility of RSFMs across diverse RS benchmarks. The resulting RSFMs achieve state-of-the-art performance across all benchmarks, underscoring the viability of scaling semi-supervised learning for RS applications. Di Wang 0023, Jing Zhang 0037, Lefei Zhang |
AAAI | 4 |
| 2026 | RS2-SAM2: Customized SAM2 for Referring Remote Sensing Image SegmentationabstractReferring Remote Sensing Image Segmentation (RRSIS) aims to segment target objects in remote sensing (RS) images based on textual descriptions. Although Segment Anything Model 2 (SAM2) has shown remarkable performance in various segmentation tasks, its application to RRSIS presents several challenges, including understanding the text-described RS scenes and generating effective prompts from text. To address these issues, we propose RS2-SAM2, a novel framework that adapts SAM2 to RRSIS by aligning the adapted RS features and textual features while providing pseudo-mask-based dense prompts. Specifically, we employ a union encoder to jointly encode the visual and textual inputs, generating aligned visual and text embeddings as well as multimodal class tokens. A bidirectional hierarchical fusion module is introduced to adapt SAM2 to RS scenes and align adapted visual features with the visually enhanced text embeddings, improving the model's interpretation of text-described RS scenes. To provide precise target cues for SAM2, we design a mask prompt generator, which takes the visual embeddings and class tokens as input and produces a pseudo-mask as the dense prompt of SAM2. Experimental results on several RRSIS benchmarks demonstrate that RS2-SAM2 achieves state-of-the-art performance. Fu Rong, Meng Lan, Qian Zhang 0009, Lefei Zhang |
AAAI | 4 |
| 2026 | Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch ScenariosabstractSpeculative decoding accelerates LLM inference by utilizing otherwise idle computational resources during memory-to-chip data transfer. Current speculative decoding methods typically assume a considerable amount of available computing power, then generate a complex and massive draft tree using a small autoregressive language model to improve overall prediction accuracy. However, methods like batching have been widely applied in mainstream model inference systems as a superior alternative to speculative decoding, as they compress the available idle computing power. Therefore, performing speculative decoding with low verification resources and low scheduling costs has become an important research problem. We believe that more capable models that allow for parallel generation on draft sequences are what we truly need. Recognizing the fundamental nature of draft models to only generate sequences of limited length, we propose SpecFormer, a novel architecture that integrates unidirectional and bidirectional attention mechanisms. SpecFormer combines the autoregressive model’s ability to extract information from the entire input sequence with the parallel generation benefits of non-autoregressive models. This design eliminates the reliance on large prefix trees and achieves consistent acceleration, even in large-batch scenarios. Through lossless speculative decoding experiments across models of various scales, we demonstrate that SpecFormer sets a new standard for scaling LLM inference with lower training demands and reduced computational costs. Luohe Shi, Zuchao Li, Lefei Zhang, Baoyuan Qi, Guoming Liu, Hai Zhao 0001 |
AAAI | 3 |
| 2026 | Any2RSI: Controllable Remote Sensing Text-to-Image Generation via Any Control and Enriched DescriptionabstractRecent advances in controllable text-to-image (T2I) generation have achieved impressive results in natural images, but remote sensing (RS) T2I remains challenging due to the unique nature of geospatial data. Existing methods struggle to integrate diverse spatial controls and model complex spatial relationships, often failing to maintain semantic consistency with typically vague or incomplete textual descriptions. Moreover, limited by small-scale, low-quality datasets, these models produce outputs with inconsistent layouts and unrealistic content. To address these issues, we propose Any2RSI, a flexible framework for controllable RS T2I generation. It features a Cross-Modal Multi-Control Adapter that extracts modality-agnostic embeddings from heterogeneous spatial inputs, enabling precise spatial guidance. To compensate for sparse or ambiguous text prompts, we introduce a VLM-Empowered Enriched Description Generation module that enhances input descriptions with cross-modal semantics for more coherent image generation. Furthermore, we present RST2I-110K, a new large-scale dataset with over 115,000 high-quality RS image-text pairs across diverse scenes, alleviating data scarcity in this domain. Extensive experiments show that Any2RSI achieves state-of-the-art performance on both existing and new datasets, improving the realism and structural accuracy of generated RS imagery. Xu Zhang 0044, Lefei Zhang |
AAAI | 3 |
| 2026 | Anchor-Guided Discriminative Subspace Alignment and Clustering for Cross-Scene Hyperspectral ImageryabstractCross-scene hyperspectral image (HSI) recognition aims to assign a unique label to each pixel in the target scene by transferring knowledge from the source scene. Existing methods primarily rely on fully labeled source data and either partially labeled or unlabeled target data. No prior work has addressed the more challenging scenario of cross-scene recognition without label guidance in both scenes. To bridge this gap, we present the first study on cross-scene HSI clustering, proposing an anchor-guided discriminative subspace alignment and clustering (ADSAC) framework that follows a well-structured three-step learning paradigm to effectively mitigate distribution shifts. Specifically, we first develop an anchor-promoted graph learning (APGL) model to efficiently derive accurate clustering labels for the source scene by leveraging anchor-based structural information. Next, we propose a discriminative cross-scene subspace alignment (DCSA) model to improve feature discriminability and reduce distribution discrepancies. Finally, labels of the target scene are inferred after source clustering and cross-scene alignment. To solve the formulated models, we design tailored optimization algorithms to ensure high-quality learning. Extensive experiments demonstrate the superiority of the proposed framework over state-of-the-art methods. Yongshan Zhang, Xinxin Wang 0003, Lefei Zhang, Zhihua Cai |
AAAI | 4 |
| 2026 | ClearAIR: A Human-Visual-Perception-Inspired All-in-One Image RestorationabstractRecently, All-in-One image restoration (AiOIR) has advanced significantly, offering promising solutions for complex real-world degradations. However, most existing approaches heavily rely on degradation-specific representation learning, which can lead to oversmoothing and artifacts in the restored images. To address this limitation, we propose ClearAIR, a novel AiOIR framework inspired by human visual perception and designed with a hierarchical restoration strategy in a coarse-to-fine manner. First, leveraging the global priority characteristic of early human visual perception, we employ an image quality assessment model to evaluate the overall image structure and degradation level. Next, we introduce a Semantic Guidance Unit to provide coarse semantic region guidance and a Task Identifier to predict local degradation types, enabling a more informed characterization of local degradation patterns. Finally, aiming at the challenge of local detail restoration, we propose an Internal Clue Reuse Mechanism that deeply mines the internal information of the image in a self-supervised manner to enhance the model’s capacity for fine-detail recovery. Experimental results demonstrate that ClearAIR achieves superior restoration performance across diverse synthetic and real-world datasets. Xu Zhang 0044, Huan Zhang 0008, Guoli Wang 0004, Qian Zhang 0009, Lefei Zhang |
AAAI | 5 |
| 2026 | Vista-LLM: Decoupled Query-Guided Visual Token Pruning for Efficient Long-Video Large Language ModelsabstractLong-video understanding is bottlenecked by the high cost of processing massive visual tokens.Current reduction strategies often rely on static allocation or inefficient in-network selection that disrupts optimized attention kernels.In this paper, we introduce Vista-LLM, a decoupled framework for query-guided visual token pruning.By filtering redundancy prior to inference with minimal overhead, Vista-LLM ensures full compatibility with Flash Attention.Our method employs a coarse-tofine pipeline: (1) Query-Guided Dynamic Budgeting for adaptive temporal allocation; (2) a lightweight Semantic Scout for fine-grained, query-specific selection; and (3) Structure-Aware Compensation to preserve global context.Extensive experiments on benchmarks like Video-MME and MLVU demonstrate a significantly improved Pareto frontier.Notably, on LLaVA-OneVision, Vista-LLM reduces visual tokens by 90% and accelerates inference while retaining over 98% of baseline performance on average, effectively filtering visual noise.Our code is available at https: //github.com/lizhenyu-123/Vista-LLM. Zuchao Li, Ping Wang 0028, Lefei Zhang, Haojun Ai |
ACL (1) | 4 |
| 2026 | From AR to Diffusion: Efficiently Adapting Large Language Models with Strictly Causal and Elastic HorizonsabstractDiffusion models promise efficient parallel text generation but rely on bidirectional attention, creating a structural mismatch with pre-trained Autoregressive (AR) models.This incompatibility precludes reusing robust AR priors, necessitating prohibitive pre-training from scratch.To bridge this gap, we propose FLUID, a framework that efficiently adapts AR backbones to the diffusion paradigm.By enforcing Strictly Causal Alignment, FLUID enables seamless initialization from standard GPT-style checkpoints, circumventing the need for massive pre-training.Furthermore, we introduce Elastic Horizons, an entropy-driven mechanism that dynamically modulates denoising strides based on local information density rather than fixed schedules.Experiments demonstrate that FLUID achieves state-of-theart performance while reducing training costs by orders of magnitude, effectively reconciling established AR foundations with efficient parallel generation. Teng Xiao, Zuchao Li, Lefei Zhang |
ACL (1) | 4 |
| 2026 | Task prior attention network for multi-task learning of dense prediction
Yangyang Xu 0001, Lefei Zhang, Bo Du 0001 |
Sci. China Inf. Sci. | 3 |
| 2026 | Universal Facial Landmark Detection by Landmark-Clustering Relation-Reasoning Transformer
Jun Wan 0005, Yuanzhi Yao, Jiaxing Huang 0001, Xiaoying Ding, Lefei Zhang, Yongsheng Gao 0001, Dacheng Tao |
Int. J. Comput. Vis. | 5 |
| 2026 | MambaFPN: A SSM-based feature pyramid network for object detection
Le Liang, Cheng Wang 0048, Lefei Zhang |
Neural Networks | 3 |
| 2026 | Generalized and group spherical linear interpolation for token-level context compression
Jinhao Tian, Zuchao Li, Meng-Jia Shen, Lefei Zhang |
Neural Networks | 4 |
| 2026 | Efficient and Effective Weight-Ensembling Mixture of Experts for Multi-Task Model MergingabstractMulti-task learning (MTL) leverages a shared model to accomplish multiple tasks and facilitate knowledge transfer. Recent research on task arithmetic-based MTL demonstrates that merging the parameters of independently fine-tuned models can effectively achieve MTL. However, existing merging methods primarily seek a static optimal solution within the original model parameter space, which often results in performance degradation due to the inherent diversity among tasks and potential interferences. To address this challenge, in this paper, we propose a Weight-Ensembling Mixture of Experts (WEMoE) method for multi-task model merging. Specifically, we first identify critical (or sensitive) modules by analyzing parameter variations in core modules of Transformer-based models before and after fine-tuning. Then, our WEMoE statically merges non-critical modules while transforming critical modules into a mixture-of-experts (MoE) structure. During inference, expert modules in the MoE are dynamically merged based on input samples, enabling a more flexible and adaptive merging approach. Building on WEMoE, we further introduce an efficient-and-effective WEMoE (E-WEMoE) method, whose core mechanism involves eliminating non-essential elements in the critical modules of WEMoE and implementing shared routing across multiple MoE modules, thereby significantly reducing both the trainable parameters, the overall parameter count, and computational overhead of the merged model by WEMoE. Experimental results across various architectures and tasks demonstrate that both WEMoE and E-WEMoE outperform state-of-the-art (SOTA) model merging methods in terms of MTL performance, generalization, and robustness. Li Shen 0008, Anke Tang, Enneng Yang, Guibing Guo, Yong Luo 0002, Lefei Zhang, Xiaochun Cao, Bo Du 0001, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Zero-Shot Sparse Mixture of Low-Rank Experts Construction From Pre-Trained Foundation ModelsabstractDeep model training on extensive datasets is increasingly becoming cost-prohibitive, prompting the widespread adoption of deep model fusion techniques to leverage knowledge from pre-existing models. From simple weight averaging to more sophisticated methods like AdaMerging, model fusion effectively improves model performance and accelerates the development of new models. However, potential interference between parameters of individual models and the lack of interpretability in the fusion progress remain significant challenges. Existing methods often try to resolve the parameter interference issue by evaluating attributes of parameters, such as their magnitude or sign, or by parameter pruning. In this study, we begin by examining the fine-tuning of linear layers through the lens of subspace analysis and explicitly define parameter interference as an optimization problem to shed light on this subject. Subsequently, we introduce an innovative approach to model fusion called zero-shot Sparse MIxture of Low-rank Experts (SMILE) construction, which allows for the upscaling of source models into an MoE model without extra data or further training. Our approach relies on the observation that fine-tuning mostly keeps the important parts from the pre-training, but it uses less significant or unused areas to adapt to new tasks. Additionally, the issue of parameter interference, which is intrinsically challenging in the original parameter space, can be managed by expanding the dimensions. We conduct extensive experiments across diverse scenarios, such as image classification and text generation tasks, using full fine-tuning and LoRA fine-tuning, and we apply our method to large language models (CLIP models, Flan-T5 models, and Mistral-7B models), highlighting the adaptability and scalability of SMILE. For full fine-tuned models, about 50% additional parameters can achieve around 98-99% of the performance of eight individual fine-tuned ViT models, while for LoRA fine-tuned Flan-T5 models, maintaining 99% performance with only 2% extra parameters. Code is available athttps://github.com/tanganke/fusion_bench. Anke Tang, Li Shen 0008, Yong Luo 0002, Shuai Xie, Han Hu 0003, Lefei Zhang, Bo Du 0001, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Perceive-IR: Learning to Perceive Degradation Better for All-in-One Image RestorationabstractExisting All-in-One image restoration methods often fail to perceive degradation types and severity levels simultaneously, overlooking the importance of fine-grained quality perception. Moreover, these methods often utilize highly customized backbones, which hinder their adaptability and integration into more advanced restoration networks. To address these limitations, we propose Perceive-IR, a novel backbone-agnostic All-in-One image restoration framework designed for fine-grained quality control across various degradation types and severity levels. Its modular structure allows core components to function independently of specific backbones, enabling seamless integration into advanced restoration models without significant modifications. Specifically, Perceive-IR operates in two key stages: 1) multi-level quality-driven prompt learning stage, where a fine-grained quality perceiver is meticulously trained to discern three-tier quality levels by optimizing the alignment between prompts and images within the CLIP perception space. This stage ensures a nuanced understanding of image quality, laying the groundwork for subsequent restoration; 2) restoration stage, where the quality perceiver is seamlessly integrated with a difficulty-adaptive perceptual loss, forming a quality-aware learning strategy. This strategy not only dynamically differentiates sample learning difficulty but also achieves fine-grained quality control by driving the restored image toward the ground truth while pulling it away from both low- and medium-quality samples. Furthermore, Perceive-IR incorporates a Semantic Guidance Module (SGM) and Compact Feature Extraction (CFE). The SGM leverages semantic information from pre-trained vision models to provide high-level contextual guidance, while the CFE focuses on extracting degradation-specific features, ensuring accurate handling of diverse image degradations. Extensive experiments demonstrate that Perceive-IR not only surpasses state-of-the-art methods but also generalizes reliably to zero-shot real-world and unknown degraded scenes, while adapting seamlessly to different backbone networks. This versatility underscores the framework's robustness and backbone-agnostic design. Project page at https://house-yuyu.github.io/Perceive-IR/. Xu Zhang 0044, Jiaqi Ma 0002, Guoli Wang 0004, Qian Zhang 0009, Huan Zhang 0008, Lefei Zhang |
IEEE Trans. Image Process. | 6 |
| 2026 | Leveraging Multi-Text Joint Prompts in SAM for Robust Medical Image SegmentationabstractThe Segment Anything Model (SAM) has attracted considerable attention due to its impressive performance and demonstrates potential in medical image segmentation. Compared to SAM's native point andbounding box prompts, text prompts offer a simpler and more efficient alternative in the medical field, yet this approach remains relatively underexplored. In this paper, we propose a SAM-based framework that integrates a pre-trained vision-language model to generate referring prompts, with SAM handling the segmentation task. The outputs from multimodal models such as CLIP serve as input to SAM's prompt encoder. A critical challenge stems from the inherent complexity of medical text descriptions: they typically encompass anatomical characteristics, imaging modalities, and diagnostic priorities, resulting in information redundancy and semantic ambiguity. To address this, we propose a text decomposition-recomposition strategy. First, clinical narratives are parsed into atomic semantic units (appearance, location, pathology, and so on). These elements are then recombined into optimized text expressions. We employ a cross-attention module among multiple texts to interact with the joint features, ensuring that the model focuses on features corresponding to effective descriptions. To validate the effectiveness of our method, we conducted experiments on several datasets. Compared to the native SAM based on geometric prompts, our model shows improved performance and usability. Xu Zhang 0044, Huangxuan Zhao, Lefei Zhang, Yuan Xiong |
IEEE J. Biomed. Health Informatics | 3 |
| 2026 | Proto-Former: Unified Facial Landmark Detection by Prototype TransformerabstractRecent advances in deep learning have significantly improved facial landmark detection. However, existing facial landmark detection datasets often define different numbers of landmarks, and most mainstream methods can only be trained on a single dataset. This limits the model generalization to different datasets and hinders the development of a unified model. To address this issue, we propose Proto-Former, a unified, adaptive, end-to-end facial landmark detection framework that explicitly enhances dataset-specific facial structural representations (i.e., prototype). Proto-Former overcomes the limitations of single-dataset training by enabling joint training across multiple datasets within a unified architecture. Specifically, Proto-Former comprises two key components: an Adaptive Prototype-Aware Encoder (APAE) that performs adaptive feature extraction and learns prototype representations, and a Progressive Prototype-Aware Decoder (PPAD) that refines these prototypes to generate prompts that guide the model's attention to key facial regions. Furthermore, we introduce a novel Prototype-Aware (PA) loss, which achieves optimal path finding by constraining the selection weights of prototype experts. This loss function effectively resolves the problem of prototype expert addressing instability during multi-dataset training, alleviates gradient conflicts, and enables the extraction of more accurate facial structure features. Extensive experiments on widely used benchmark datasets demonstrate that our Proto-Former achieves superior performance compared to existing state-of-the-art methods. The code is publicly available at:https://github.com/Husk021118/Proto-Former. Shengkai Hu, Haozhe Qi, Jun Wan 0005, Jiaxing Huang 0001, Lefei Zhang, Dacheng Tao |
IEEE Trans. Multim. | 5 |
| 2025 | Imitate Before Detect: Aligning Machine Stylistic Preference for Machine-Revised Text DetectionabstractLarge Language Models (LLMs) have revolutionized text generation, making detecting machine-generated text increasingly challenging. Although past methods have achieved good performance on detecting pure machine-generated text, those detectors have poor performance on distinguishing machine-revised text (rewriting, expansion, and polishing), which can have only minor changes from its original human prompt. As the content of text may originate from human prompts, detecting machine-revised text often involves identifying distinctive machine styles, e.g., worded favored by LLMs. However, existing methods struggle to detect machine-style phrasing hidden within the content contributed by humans. We propose the “Imitate Before Detect” (ImBD) approach, which first imitates the machine-style token distribution, and then compares the distribution of the text to be tested with the machine-style distribution to determine whether the text has been machine-revised. To this end, we introduce Style Preference Optimization (SPO), which aligns a scoring LLM model to the preference of text styles generated by machines. The aligned scoring model is then used to calculate the style-conditional probability curvature (Style-CPC), quantifying the log probability difference between the original and conditionally sampled texts for effective detection. We conduct extensive comparisons across various scenarios, encompassing text revisions by six LLMs, four distinct text domains, and three machine revision types. Compared to existing state-of-the-art methods, our method yields a 13% increase in AUC for detecting text revised by open-source LLMs, and improves performance by 5% and 19% for detecting GPT-3.5 and GPT-4o revised text, respectively. Notably, our method surpasses the commercially trained GPT-Zero with just 1,000 samples and five minutes of SPO, demonstrating its efficiency and effectiveness. Xiaoye Zhu, Yiwen Yuan, Chak Tou Leong, Zuchao Li, Tang Long, Chenyu Yan, Guanghao Mei, Lefei Zhang |
AAAI | 14 |
| 2025 | SongSong: A Time Phonograph for Chinese SongCi Music from Thousand of Years AwayabstractRecently, there have been significant advancements in music generation. However, existing models primarily focus on creating modern pop songs, making it challenging to produce ancient music with distinct rhythms and styles, such as ancient Chinese SongCi. In this paper, we introduce SongSong, the first music generation model capable of restoring Chinese SongCi to our knowledge. Our model first predicts the melody from the input SongCi, then separately generates the singing voice and accompaniment based on that melody, and finally combines all elements to create the final piece of music. Additionally, to address the lack of ancient music datasets, we create OpenSongSong, a comprehensive dataset of ancient Chinese SongCi music, featuring 29.9 hours of compositions by various renowned SongCi music masters. To assess SongSong's proficiency in performing SongCi, we randomly select 85 SongCi sentences that were not part of the training set for evaluation against SongSong and music generation platforms such as Suno and SkyMusic. The subjective and objective outcomes indicate that our proposed model achieves leading performance in generating high-quality SongCi music. Jiliang Hu 0001, Jiajia Li 0005, Ziyi Pan, Zuchao Li, Ping Wang 0028, Lefei Zhang |
AAAI | 7 |
| 2025 | ScaleMatch: Multi-scale Consistency Enhancement for Semi-supervised Semantic SegmentationabstractSemi-supervised learning improves semantic segmentation performance by leveraging unlabeled data, thereby significantly reducing labeling costs. Previous semi-supervised semantic segmentation (S4) methods explored perturbations at the image level but neglected to adequately utilize multi-scale information. When labeled information is insufficient, the scale variation between different objects makes learning instances with extreme scales even more difficult. To address this issue, we propose ScaleMatch, which aims to learn scale-invariant features by obtaining a mixed dual-scale pseudo-label and scale consistency learning. Specifically, the cross-scale interaction fusion (CIF) module enforces interactive information across different scaled-views, allowing for more reliable pseudo-label generation. More importantly, ScaleMatch introduces variable scale branches to utilize scale-invariant supervision. It consists of image-level scale variation consistency (ISVC) and feature-level scale variation consistency (FSVC). Consequently, our ScaleMatch enhances the model's generalization under scale variation, outperforming existing state-of-the-art methods on both the Pascal VOC and Cityscapes datasets under various partition protocols. Lefei Zhang |
AAAI | 2 |
| 2025 | KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional EmbeddingabstractLarge language models (LLMs) based on Transformer Decoders have become the preferred choice for conversational generative AI.Despite the overall superiority of the Decoder architecture, the gradually increasing Key-Value (KV) cache during inference has emerged as a primary efficiency bottleneck, both in aspects of memory consumption and data transfer bandwidth limitations.To address these challenges, we propose a paradigm called KV-Latent.By down-sampling the Key-Value vector dimensions into a latent space, we can significantly reduce the KV Cache footprint and improve inference speed, only with a small amount of extra training, less than 1% of pretraining takes.Besides, we enhanced the stability of Rotary Positional Embedding applied on lower-dimensional vectors by modifying its frequency sampling mechanism, avoiding noise introduced by higher frequencies while retaining position attenuation.Our experiments, including both models with Grouped Query Attention and those without, have yielded satisfactory results.Finally, we conducted comparative experiments to study the impact of separately reducing Key and Value components on model's performance.Our approach allows for the construction of more efficient language model systems, and opens the new possibility on KV Cache saving and efficient LLMs.Our code is available at https://github.com/ShiLuohe/KV- Latent. Luohe Shi, Zuchao Li, Lefei Zhang, Baoyuan Qi, Guoming Liu, Hai Zhao 0001 |
ACL (1) | 3 |
| 2025 | SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep LayersabstractZicong Tang, Shi Luohe, Zuchao Li, Baoyuan Qi, Liu Guoming, Lefei Zhang, Ping Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zicong Tang, Luohe Shi, Zuchao Li, Baoyuan Qi, Guoming Liu, Lefei Zhang, Ping Wang 0028 |
ACL (1) | 6 |
| 2025 | Segment First or Comprehend First? Explore the Limit of Unsupervised Word Segmentation with Large Language ModelsabstractWord segmentation stands as a cornerstone of Natural Language Processing (NLP).Based on the concept of "comprehend first, segment later", we propose a new framework to explore the limit of unsupervised word segmentation with Large Language Models (LLMs) and evaluate the semantic understanding capabilities of LLMs based on word segmentation.We employ current mainstream LLMs to perform word segmentation across multiple languages to assess LLMs' "comprehension".Our findings reveal that LLMs are capable of following simple prompts to segment raw text into words.There is a trend suggesting that models with more parameters tend to perform better on multiple languages.Additionally, we introduce a novel unsupervised method, termed LLACA (Large Language Model-Inspired Aho-Corasick Automaton).Leveraging the advanced pattern recognition capabilities of Aho-Corasick automata, LLACA innovatively combines these with the deep insights of well-pretrained LLMs.This approach not only enables the construction of a dynamic n-gram model that adjusts based on contextual information but also integrates the nuanced understanding of LLMs, offering significant improvements over traditional methods. Zihong Zhang, Liqi He, Zuchao Li, Lefei Zhang, Hai Zhao 0001, Bo Du 0001 |
ACL (1) | 4 |
| 2025 | Intention Analysis Makes LLMs A Good Jailbreak DefenderabstractAligning large language models (LLMs) with human values, particularly when facing complex and stealthy jailbreak attacks, presents a formidable challenge. Unfortunately, existing methods often overlook this intrinsic nature of jailbreaks, which limits their effectiveness in such complex scenarios. In this study, we present a simple yet highly effective defense strategy, i.e., Intention Analysis (IA). IA works by triggering LLMs’ inherent self-correct and improve ability through a two-stage process: 1) analyzing the essential intention of the user input, and 2) providing final policy-aligned responses based on the first round conversation. Notably,IA is an inference-only method, thus could enhance LLM safety without compromising their helpfulness. Extensive experiments on varying jailbreak benchmarks across a wide range of LLMs show that IA could consistently and significantly reduce the harmfulness in responses (averagely -48.2% attack success rate). Encouragingly, with our IA, Vicuna-7B even outperforms GPT-3.5 regarding attack success rate. We empirically demonstrate that, to some extent, IA is robust to errors in generated intentions. Further analyses reveal the underlying principle of IA: suppressing LLM’s tendency to follow jailbreak prompts, thereby enhancing safety. Yuqi Zhang 0002, Liang Ding 0006, Lefei Zhang, Dacheng Tao |
COLING | 3 |
| 2025 | ToM: Leveraging Tree-oriented MapReduce for Long-Context Reasoning in Large Language ModelsabstractLarge Language Models (LLMs), constrained by limited context windows, often face significant performance degradation when reasoning over long contexts.To address this, Retrieval-Augmented Generation (RAG) retrieves and reasons over chunks but frequently sacrifices logical coherence due to its reliance on similarity-based rankings.Similarly, divideand-conquer frameworks (DCF) split documents into small chunks for independent reasoning and aggregation.While effective for local reasoning, DCF struggles to capture longrange dependencies and risks inducing conflicts by processing chunks in isolation.To overcome these limitations, we propose ToM, a novel Tree-oriented MapReduce framework for long-context reasoning.ToM leverages the inherent hierarchical structure of long documents (e.g., main headings and subheadings) by constructing a DocTree through hierarchical semantic parsing and performing bottom-up aggregation.Using a Tree MapReduce approach, ToM enables recursive reasoning: in the Map step, rationales are generated at child nodes; in the Reduce step, these rationales are aggregated across sibling nodes to resolve conflicts or reach consensus at parent nodes.Experimental results on 70B+ LLMs show that ToM significantly outperforms existing divide-andconquer frameworks and retrieval-augmented generation methods, achieving better logical coherence and long-context reasoning.Our code is available at https://github.com/gjn12- 31/ToM. Jiani Guo, Zuchao Li, Jie Wu 0001, Qianren Wang, Yun Li 0011, Lefei Zhang, Hai Zhao 0001, Yujiu Yang 0001 |
EMNLP | 6 |
| 2025 | MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object SegmentationabstractReferring video object segmentation (RVOS) aims to segment objects in a video according to textual descriptions, which requires the integration of multimodal information and temporal dynamics perception. The Segment Anything Model 2 (SAM 2) has shown great effectiveness across various video segmentation tasks. However, its application to offline RVOS is challenged by the translation of the text into effective prompts and a lack of global context awareness. In this paper, we propose a novel RVOS framework, termed MPG-SAM 2, to address these challenges. Specifically, MPG-SAM 2 employs a unified multimodal encoder to jointly encode video and textual features, generating semantically aligned video and text embeddings, along with multimodal class tokens. A mask prior generator utilizes the video embeddings and class tokens to create pseudo masks of target objects and global context. These masks are fed into the prompt encoder as dense prompts along with multimodal class tokens as sparse prompts to generate accurate prompts for SAM 2. To provide the online SAM 2 with a global view, we introduce a hierarchical global-historical aggregator, which allows SAM 2 to aggregate global and historical information of target objects at both pixel and object levels, enhancing the target representation and temporal consistency. Extensive experiments on several RVOS benchmarks demonstrate the superiority of MPG-SAM 2 and the effectiveness of our proposed modules. The code is available at https://github.com/rongfu-dsb/MPG-SAM2. Fu Rong, Meng Lan, Qian Zhang 0009, Lefei Zhang |
ICCV | 4 |
| 2025 | What Limits Bidirectional Model's Generative Capabilities? A Uni-Bi-Directional Mixture-of-Expert Method For Bidirectional Fine-tuningabstractLarge Language Models (LLMs) excel in generation tasks, yet their causal attention mechanisms limit performance in embedding tasks. While bidirectional modeling may enhance embeddings, naively fine-tuning unidirectional models bidirectionally severely degrades generative performance. To investigate this trade-off, we analyze attention weights as dependence indicators and find that bidirectional fine-tuning increases subsequent dependence, impairing unidirectional generation. Through systematic Transformer module evaluations, we discover the FFN layer is least affected by such dependence. Leveraging this discovery, we propose UBMoE-LLM, a novel Uni-Bi-directional Mixture-of-Experts LLM, which integrates the original unidirectional FFN with a bidirectionally fine-tuned FFN via unsupervised contrastive learning. This MoE-based approach enhances embedding performance while preserving robust generation. Extensive experiments across diverse datasets and model scales validate our attention dependence metric and demonstrate UBMoE-LLM’s superior generative quality and reduced hallucination. Code is available at: https://github.com/heiyonghua/ubmoe_llm. Zuchao Li, Yonghua Hei, Qiwei Li 0002, Lefei Zhang, Ping Wang 0028, Hai Zhao 0001, Baoyuan Qi, Guoming Liu |
ICML | 4 |
| 2025 | The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward HackingabstractThis work identifies the *Energy Loss Phenomenon* in Reinforcement Learning from Human Feedback (RLHF) and its connection to reward hacking. Specifically, energy loss in the final layer of a Large Language Model (LLM) gradually increases during the RL process, with an *excessive* increase in energy loss characterizing reward hacking. Beyond empirical analysis, we further provide a theoretical foundation by proving that, under mild conditions, the increased energy loss reduces the upper bound of contextual relevance in LLMs, which is a critical aspect of reward hacking as the reduced contextual relevance typically indicates overfitting to reward model-favored patterns in RL. To address this issue, we propose an *Energy loss-aware PPO algorithm (EPPO)* which penalizes the increase in energy loss in the LLM's final layer during reward calculation to prevent excessive energy loss, thereby mitigating reward hacking. We theoretically show that EPPO can be conceptually interpreted as an entropy-regularized RL algorithm, which provides deeper insights into its effectiveness. Extensive experiments across various LLMs and tasks demonstrate the commonality of the energy loss phenomenon, as well as the effectiveness of EPPO in mitigating reward hacking and improving RLHF performance. Yuchun Miao, Sen Zhang 0006, Liang Ding 0006, Yuqi Zhang 0002, Lefei Zhang, Dacheng Tao |
ICML | 5 |
| 2025 | HSRMamba: Contextual Spatial-Spectral State Space Model for Single Hyperspectral Image Super-ResolutionabstractMamba has demonstrated exceptional performance in visual tasks due to its powerful global modeling capabilities and linear computational complexity, offering considerable potential in hyperspectral image super-resolution (HSISR). However, in HSISR, Mamba faces challenges as transforming images into 1D sequences neglects the spatial-spectral structural relationships between locally adjacent pixels, and its performance is highly sensitive to input order, which affects the restoration of both spatial and spectral details. In this paper, we propose HSRMamba, a contextual spatial-spectral modeling state space model for HSISR, to address these issues both locally and globally. Specifically, a local spatial-spectral partitioning mechanism is designed to establish patch-wise causal relationships among adjacent pixels in 3D features, mitigating the local forgetting issue. Furthermore, a global spectral reordering strategy based on spectral similarity is employed to enhance the causal representation of similar pixels across both spatial and spectral dimensions. Finally, experimental results demonstrate our HSRMamba outperforms the state-of-the-art methods in quantitative quality and visual results. Code is available at: https://github.com/Tomchenshi/HSRMamba. Shi Chen 0010, Lefei Zhang, Liangpei Zhang 0001 |
IJCAI | 2 |
| 2025 | Spatial-Spectral Similarity-Guided Fusion Network for PansharpeningabstractPansharpening fuses lower-resolution multispectral (LRMS) images with high-resolution panchromatic (PAN) images to generate high-resolution multispectral (HRMS) images that preserves both spatial and spectral information. Most deep pansharpening methods face challenges in cross-modal feature extraction and fusion, as well as in exploring the similarities between the fused image and both PAN and LRMS images. In this paper, we propose a spatial-spectral similarity-guided fusion network (S3FNet) for pansharpening. This architecture is composed of three parts. Specifically, a shallow feature extraction layer learns initial spatial, spectral and fused features from PAN and LRMS images. Then, a multi-branch asymmetric encoder, consisting of spatial, spectral and fusion branches, generates corresponding high-level features at different scales. A multi-scale reconstruction decoder, equipped with a well-designed cross-feature multi-head attention fusion block, processes the intermediate feature maps to generate HRMS images. To ensure HRMS images retain maximum spatial and spectral information, a similarity-constrained loss is defined for network training. Extensive experiments demonstrate the effectiveness of our S3FNet over state-of-the-art methods. The code is released at https://github.com/ZhangYongshan/S3FNet. Jiazhuang Xiong, Yongshan Zhang, Xinxin Wang 0003, Lefei Zhang |
IJCAI | 4 |
| 2025 | Deep Multi-Level Contrastive Clustering for Multi-Modal Remote Sensing Images
Yongshan Zhang, Xinxin Wang 0003, Lefei Zhang |
ACM Multimedia | 4 |
| 2025 | Multimodal Decomposed Distillation with Instance Alignment and Uncertainty Compensation for Thermal Object DetectionabstractRGB-Thermal images leverage complementary optical and thermal modalities to identify objects. While achieving superior performance, the reliance on multimodal fusion inherently limits inference efficiency and adaptability to harsh RGB-failure environments. In this work, we propose a multimodal decomposed distillation framework to develop robust thermal-only detectors by transferring knowledge from multimodal teachers. Unlike conventional one-to-one distillation, we decouple the tasks of simultaneously mimicking RGB-T teacher representations and preserving thermal-specific student feature integrity into dual branches to avoid intrinsic semantic conflicts. Specifically, we present channel-adaptive prompt learning for cross-modal decomposition and a frequency-guided dynamic module for decomposed knowledge integration. The dual-branch architecture employs asymmetric training objectives to ensure effective cross-modal knowledge transfer while preserving the integrity of thermal information. Furthermore, to exploit finer-grained instance knowledge across both feature and prediction levels, we introduce a customized instance alignment distillation to enhance the local discriminability in feature pyramids, and propose an uncertainty-aware logit distillation to compensate for ambiguous predictions in detection heads. Experiments on three datasets validate the effectiveness of our framework in boosting thermal-based detectors. Code is released at https://github.com/lyf0801/DecomKD. Yanfeng Liu, Lefei Zhang |
ACM Multimedia | 2 |
| 2025 | Label Drop for Multi-Aspect Relation Modeling in Universal Information ExtractionabstractLu Yang, Jiajia Li, En Ci, Lefei Zhang, Zuchao Li, Ping Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Lu Yang 0008, Jiajia Li 0005, En Ci, Lefei Zhang, Zuchao Li, Ping Wang 0028 |
NAACL (Long Papers) | 4 |
| 2025 | Merging on the Fly Without Retraining: A Sequential Approach to Scalable Continual Model MergingabstractDeep model merging represents an emerging research direction that combines multiple fine-tuned models to harness their specialized capabilities across different tasks and domains. Current model merging techniques focus on merging all available models simultaneously, with weight interpolation-based methods being the predominant approach. However, these conventional approaches are not well-suited for scenarios where models become available sequentially, and they often suffer from high memory requirements and potential interference between tasks. In this study, we propose a training-free projection-based continual merging method that processes models sequentially through orthogonal projections of weight matrices and adaptive scaling mechanisms. Our method operates by projecting new parameter updates onto subspaces orthogonal to existing merged parameter updates while using an adaptive scaling mechanism to maintain stable parameter distances, enabling efficient sequential integration of task-specific knowledge. Our approach maintains constant memory complexity to the number of models, minimizes interference between tasks through orthogonal projections, and retains the performance of previously merged models through adaptive task vector scaling. Extensive experiments on CLIP-ViT models demonstrate that our method achieves a 5-8% average accuracy improvement while maintaining robust performance in different task orderings. Code is publicly available at https://github.com/tanganke/opcm . Anke Tang, Enneng Yang, Li Shen 0008, Yong Luo 0002, Han Hu 0003, Lefei Zhang, Bo Du 0001, Dacheng Tao |
NeurIPS | 6 |
| 2025 | DM-PCL: Text-Driven Dual-Modal Prototype Consistency Learning for Weakly-Supervised Few-Shot Part Segmentation
Mengya Han, Yong Luo 0002, Han Hu 0003, Zengmao Wang, Lefei Zhang, Bo Du 0001, Ling-Yu Duan, Dacheng Tao |
Int. J. Comput. Vis. | 5 |
| 2025 | Mamba-driven hierarchical temporal multimodal alignment for referring video object segmentation
Le Liang, Lefei Zhang |
Neurocomputing | 2 |
| 2025 | FusionBench: A Unified Library and Comprehensive Benchmark for Deep Model FusionabstractDeep model fusion is an emerging technique that unifies the predictions or parameters of several deep neural networks into a single better-performing model in a cost-effective and data-efficient manner. Although a variety of deep model fusion techniques have been introduced, their evaluations tend to be inconsistent and often inadequate to validate their effectiveness and robustness. We present FusionBench, the first benchmark and a unified library designed specifically for deep model fusion. Our benchmark consists of multiple tasks, each with different settings of models and datasets. This variety allows us to compare fusion methods across different scenarios and model scales. Additionally, FusionBench serves as a unified library for easy implementation and testing of new fusion techniques. FusionBench is open source and actively maintained, with community contributions encouraged. Anke Tang, Li Shen 0008, Yong Luo 0002, Enneng Yang, Han Hu 0003, Lefei Zhang, Bo Du 0001, Dacheng Tao |
J. Mach. Learn. Res. | 6 |
| 2025 | Sandbox: safeguarded multi-label learning through safe optimal transport
Lefei Zhang, Geng Yu, Jiangchao Yao, Yew-Soon Ong, Ivor W. Tsang, James T. Kwok |
Mach. Learn. | 1 |
| 2025 | CFI-Former: Efficient lane detection by multi-granularity perceptual query attention transformer
Rong Gao 0001, Siqi Hu, Lingyu Yan, Lefei Zhang, Jia Wu 0001 |
Neural Networks | 4 |
| 2025 | Dual selective fusion transformer network for hyperspectral image classification
Yichu Xu, Di Wang 0023, Lefei Zhang, Liangpei Zhang 0001 |
Neural Networks | 3 |
| 2025 | HyperSIGMA: Hyperspectral Intelligence Comprehension Foundation ModelabstractAccurate hyperspectral image (HSI) interpretation is critical for providing valuable insights into various earth observation-related applications such as urban planning, precision agriculture, and environmental monitoring. However, existing HSI processing methods are predominantly task-specific and scene-dependent, which severely limits their ability to transfer knowledge across tasks and scenes, thereby reducing the practicality in real-world applications. To address these challenges, we present HyperSIGMA, a vision transformer-based foundation model that unifies HSI interpretation across tasks and scenes, scalable to over one billion parameters. To overcome the spectral and spatial redundancy inherent in HSIs, we introduce a novel sparse sampling attention (SSA) mechanism, which effectively promotes the learning of diverse contextual features and serves as the basic block of HyperSIGMA. HyperSIGMA integrates spatial and spectral features using a specially designed spectral enhancement module. In addition, we construct a large-scale hyperspectral dataset, HyperGlobal-450K, for pre-training, which contains about 450 K hyperspectral images, significantly surpassing existing datasets in scale. Extensive experiments on various high-level and low-level HSI tasks demonstrate HyperSIGMA's versatility and superior representational capability compared to current state-of-the-art methods. Moreover, HyperSIGMA shows significant advantages in scalability, robustness, cross-modal transferring capability, real-world applicability, and computational efficiency. Di Wang 0023, Meiqi Hu, Yuchun Miao, Jiaqi Yang 0005, Yichu Xu, Xiaolei Qin, Jiaqi Ma 0002, Chenxing Li, Chuan Fu, Hongruixuan Chen, Chengxi Han, Naoto Yokoya, Jing Zhang 0037, Minqiang Xu, Lefei Zhang, Chen Wu 0003, Bo Du 0001, Dacheng Tao, Liangpei Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 18 |
| 2025 | Constraint Boundary Wandering Framework: Enhancing Constrained Optimization With Deep Neural NetworksabstractConstrained optimization problems are pervasive in various fields, and while conventional techniques offer solutions, they often struggle with scalability. Leveraging the power of deep neural networks (DNNs) in optimization, we present a novel learning-based approach, the Constraint Boundary Wandering Framework (CBWF), to address these challenges. Our contributions include introducing a boundary wandering strategy inspired by the active-set method, enhancing equality constraint feasibility, and treating the Lipschitz constant as a learnable parameter. Additionally, we evaluate the regularization term, illustrating that the nonsmooth L2 norm yields superior results. Extensive testing on synthetic datasets and the ACOPT dataset demonstrates CBWF's superiority, outperforming existing deep learning-based solvers in terms of both objective and constraint loss. Shixiang Chen, Li Shen 0008, Lefei Zhang, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Cyclic Cross-Modality Interaction for Hyperspectral and Multispectral Image FusionabstractIntegrating low-resolution hyperspectral images with high-resolution multispectral images is an effective approach to derive high-resolution hyperspectral images. Recently, numerous deep learning-based approaches have been employed to model the mapping relationships for the fusion directly. However, these methods often neglect the spectral characteristics and fail to facilitate comprehensive interactions among global features from heterogeneous modalities. In this paper, we propose a novel cyclic Transformer based on the cross-modality spatial-spectral interaction, exploiting diverse interaction modes to explore the similarity and complementarity among cross-modality features. Specifically, we design a cyclic interactive architecture to fully exploit the abundant spectral prior information in low-resolution hyperspectral images and the rich spatial prior information in high-resolution multispectral images. By incorporating spatial and spectral priors into the attention mechanisms in Transformer modules, we explore the long-range dependency information within the cross-modality features. Furthermore, to enhance interaction among features from different modalities, we devise the cross-modality adaptive interaction mechanisms in both spatial and spectral dimensions to facilitate information reciprocity between different modalities. Extensive experiments demonstrate that the proposed approach outperforms the state-of-the-art fusion methods both quantitatively and visually. The code is available athttps://github.com/Tomchenshi/CYformer. Shi Chen 0010, Lefei Zhang, Liangpei Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Structured Anchor Learning for Large-Scale Hyperspectral Image Projected ClusteringabstractHyperspectral image (HSI) clustering has attracted increasing attention in recent years, because it doesn’t rely on labeled pixels. However, it is a challenging task due to the complex spectral-spatial structure. The emergence of large-scale HSIs introduces a new challenge in terms of heightened computational complexity. To address the above challenges, in this paper, we propose a structured anchor projected clustering (SAPC) model for large-scale HSIs. Specifically, we exploit spatial information reflecting in the generated superpixels to perform denoising and generate anchors. Based on the preprocessing, we simultaneously learn a pixel-anchor graph and an anchor-anchor graph in a projected feature space. Meanwhile, the rank-constraint is imposed on the Laplacian matrix related to the anchor-anchor graph. To uncover the clustering structure, we design a clustering inference strategy to propagate clustering labels from anchors to pixels based on the dual graphs. Additionally, we propose an efficient optimization strategy for the formulated SAPC model with linear time complexity in terms of the number of pixels. Since the anchor-anchor graph is with much smaller size, it is high efficient to obtain the structured anchors with pseudo labels. Thus, the clustering process is significantly accelerated. Extensive experiments on multiple large-scale HSI datasets demonstrates the superiority of our SAPC over the state-of-the-art methods. The source code is released athttps://github.com/ZhangYongshan/SAPC. Guozhu Jiang, Yongshan Zhang, Xinxin Wang 0003, Xinwei Jiang, Lefei Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Elastic Graph Fusion Subspace Clustering for Large Hyperspectral ImageabstractHyperspectral image (HSI) clustering is challenging to partition pixels into different clusters due to the complex spatial distribution and high-correlated spectrum. Subspace clustering is a representative learning paradigm and has shown competitive performance in HSIs. Most existing methods ignore potential spatial or structural information and show difficulties in dealing with large-scale HSIs. In this paper, we propose an elastic graph fusion subspace clustering (EGFSC) framework that can flexibly incorporate spectral, spatial and structural information for large HSI clustering. Instead of performing pixel-level learning, superpixel-level learning is conducted according to the generated superpixels to lessen computation burden and memory cost. To explore structural information in two perspectives, a superpixel graph and a band graph are constructed based on the superpixel features. Considering the incompatible sizes of the two graphs, we present three effective dual graph fusion strategies to fuse them in different ways. With these graph fusion strategies, EGFSC is able to improve clustering performance by simultaneously considering spatial and structural information. To solve the proposed framework, we present a closed-form solution for easy implementation. Experiments demonstrate that the proposed EGFSC obtains 70.08%, 75.76%, 87.28% and 77.23% clustering accuracies on the four HSI datasets and outperforms the state-of-the-art methods. The source code is released athttps://github.com/ZhangYongshan/EGFSC. Yongshan Zhang, Xinxin Wang 0003, Xinwei Jiang, Lefei Zhang, Bo Du 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Dual-Perspective Alignment Learning for Multimodal Remote Sensing Object DetectionabstractRecently, anchor-based detectors can achieve decent performance in multimodal remote sensing scenarios, whereas their anchor-free counterparts fail to reach comparable results. To remedy this problem, we first comprehensively investigate the misalignment issues in multimodal features and detection heads, and present a dual-perspective alignment learning (DPAL) framework for multimodal remote sensing object detection. Particularly, we design a cross-modal alignment module (CMAM), which utilizes the multiscale dilation strategy and differentiable alignment function with channel-wise modulation for cross-modal feature integration. Additionally, to cope with the misalignment problem in regression and classification heads, we propose a task-head alignment module (THAM). It presents a novel pseudo-anchor mechanism, introduces a semi-fixed offset generation strategy to capture task-variant sampling coordinates, and ultimately deploys an offset knowledge transfer mechanism with deformable alignment for anchor-free detection heads. Extensive experiments on four multimodal object detection datasets show impressive results of the proposed DPAL framework. The project code is released at https://github.com/lyf0801/DPAL. Yanfeng Liu, Chaojun Yao, Lefei Zhang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Structured Cross-Resolution Distillation for Remote Sensing Salient Object DetectionabstractExisting salient object detection methods for optical remote sensing images achieve superior results with high-resolution inputs but exhibit significant degradation in low-resolution conditions. To bridge this resolution discrepancy, we propose a structured cross-resolution knowledge distillation (SCRKD) framework designed for severely low-resolution inputs. It leverages high-resolution models as teachers to guide low-resolution students through three synergistic distillation mechanisms: 1) multiview correlation distillation (MVCD); 2) multiscale feature distillation (MSFD); and 3) decoupled saliency distillation (DSD). In addition, we present cascaded SCRKD that progressively refines structured knowledge in a multistage manner, achieving further performance boosts. Experiments on three datasets indicate that SCRKD surpasses 13 state-of-the-art methods across various cross-resolution settings. Besides, our framework based on three distinct baselines validates its model-agnostic nature. This work provides an efficient solution for low-resolution salient object detection. Code is available at:https://github.com/lyf0801/SCRKD Yanfeng Liu, Lefei Zhang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Hierarchical Heterogeneous Geometric Foreground Perception Network for Remote Sensing Object DetectionabstractRecently, deep learning-based remote sensing object detection (RSOD) has been widely explored and obtained remarkable performance. However, most existing multiscale feature extraction methods neglect exploring the interfering representation of different hierarchical features in the backbone, which is crucial for learning more discriminative features. Moreover, feature pyramid network (FPN) and its variants have difficulty in effectively perceiving the pose and salient information of remote sensing objects, leading to reduced detection accuracy. To address these issues, we propose a hierarchical heterogeneous geometric foreground perception network (HHGFP-Net) for RSOD. Specifically, a hierarchical heterogeneous receptive-field module (HHRM) is proposed to reward and penalize the feature information of the corresponding levels according to the differences between the shallow and deep feature layers in the backbone, improving discriminative feature ability. Furthermore, a geometric foreground perception FPN (GFP-FPN) is developed to refine geometric shapes and enhance foreground contents, providing more precise feature representations for objects, particularly small objects. Experimental results on four challenging RSOD datasets demonstrate that our HHGFP-Net achieves state-of-the-art performance. Codes are available at:https://github.com/YyLinkWorld/HHGFP-Net. Yang Liu 0420, Lefei Zhang, Jun Wan 0005 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Dynamic-Routing 3D-Fusion Network for Remote Sensing Image Haze RemovalabstractRecently, U-shaped neural networks (U-Net) and Full resolution convolutional neural networks (F-Net) have been extensively explored for remote sensing image haze removal, achieving excellent performance. However, downsampling in U-Net leads to significant loss of high-frequency information, while F-Net fails to satisfy the large receptive field demand of remote sensing images, resulting in suboptimal dehazing results for both architectures. Moreover, most existing haze removal methods neglect exploring the correlation between spatial and channel information in feature fusion, which is crucial for restoring image texture details and colors. To address these issues, we propose a Dynamic-Routing 3D-Fusion Network (DR3DF-Net), comprising a Dynamic Routing Features Framework (DRFF) and a 3D Perceptual Feature Fusion (3D-PFF) module. Specifically, the DRFF utilizes a Self-generated Constrained Feature Routing (SCFR) mechanism to learn the most representative features extracted from U-Net, F-Net, and their fused features to enhance clear image reconstruction. Furthermore, the 3D-PFF module enhances interaction between spatial and channel information of multiple features, assigning pixel-level weights for feature fusion, improving dehazed image texture details and colors. Experiments on challenging benchmark datasets demonstrate our DR3DF-Net outperforms several state-of-the-art haze removal methods. The source code is available at https://github.com/lslyttx/DR3DF-Net. Shuanglong Li, Bo Du 0001, Lefei Zhang, Lyuyang Tong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | PTDNet: Progressive Temporal Difference Network With Global Guidance for Building Change DetectionabstractAs the spatial resolution of satellite remote sensing increases, high-resolution remote sensing (HRRS) images provide richer surface structure details, also make noise and artifacts more obvious than before. Therefore, the building change detection (BCD) task is affected by irrelevant features caused by explicit noise and artifacts, resulting in incoherent semantics and rough boundaries of buildings, which leads to fragmented change detection results and blurred boundaries between background and changes. To address this, we propose the Progressive Temporal Difference Network (PTDNet). PTDNet employs an interleaved CNN-Transformer encoder to enhance semantic structural correlation, supplemented by a Global Information Supplement Module (GISM) for semantic alignment. The Progressive Temporal Difference Module (PTDM) then suppresses artifacts and reinforces change semantic coherence through multi-stage temporal difference fusion. Finally, a Change Guidance Module (CGM) with deep semantic-guided attention refines change boundary. During this process, multi-scale features are effectively aggregated layer by layer to produce the final binary prediction map. We conducted comparative experiments with other state-of-the-art (SOTA) methods on three public BCD datasets, namely WHU, LEVIR and SYSU, and the results show that our PTDNet achieves the highest F1-scores of 93.07%, 91.80% and 83.46% respectively. The code for this work is publicly available at https://github.com/Caijiaqi85/PTDNet-CD. Jiaqi Cai, Zhiwei Ye, Lefei Zhang, Mi Wang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | TPOV-Seg: Textually Enhanced Prompt Tuning of Vision-Language Models for Open-Vocabulary Remote Sensing Semantic SegmentationabstractRemote sensing semantic segmentation faces significant challenges in open-world scenarios due to domain gaps and the presence of unseen categories in the test datasets. Open-vocabulary semantic segmentation (OVSS) based on vision-language models (VLMs) has emerged as a promising paradigm for remote sensing imagery interpretation, which enables adaptation to new datasets with arbitrary semantic categories. However, current OVSS approaches often struggle to achieve fine-grained pixel-level localization and classification for unseen categories when relying solely on fixed textual prompts and pretrained VLM encoders. The model’s generalization capability is further hindered by insufficiently fine-grained and adaptive textual representations. To address these limitations, we propose TPOV-Seg, a textually enhanced prompt tuning approach for OVSS Specifically, a remote sensing-specific Text TempLator (TTL) is introduced to enrich textual prompts and semantic representations for land cover categories by incorporating synonymous vocabulary combinations. To efficiently align the text encoder with remote sensing characteristics, a Lightweight Text-aware Prompt Tuning (LTP-Tuning) strategy is proposed for contextual modeling of word embeddings adaptation. Furthermore, a Textual-Guided Channel-Aware Aggregator (TGCA) is developed to promote inter-channel feature interaction and facilitate semantic modeling, leveraging Grouped Cross-Channel Transformers and linear Transformers under the guidance of enhanced textual features from TTL. Extensive experiments on five large-scale remote sensing segmentation datasets demonstrate that TPOV-Seg outperforms existing methods in OVSS tasks, showing strong discriminative ability for unseen categories while maintaining robust cross-domain generalization. The source codes will be available at https://github.com/zxk688/TPOVSeg. Chufeng Zhou, Lefei Zhang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | UniRS: Toward Unified Multitask Fine-Tuning for Remote Sensing Foundation ModelabstractWhile the recent advancements highlight the significant potential of the remote sensing foundation model in addressing Earth observation tasks, they require fine-tuning with task-specific data to transfer to various applications. Fine-tuning for each task greatly limits the generalizability of foundation models. The process directs the models toward the task-relevant information while less focus on the universal feature extraction for multiple tasks, which is considered the primary advantage of foundation models. Moreover, it also overlooks the complementary relationships among tasks, and significantly increases the computational complexity. However, simply training a model with multiple tasks also struggles to find effective representations of features. RSI interpretation tasks exhibit substantial finegrained differences, with each task potentially favoring features at different scales. Additionally, using the same features for all tasks can cause mutual interference, and the lack of multi-task datasets in remote sensing society hampers the generalization and sharing of visual features across tasks. To address these issues, we propose a multi-task fine-tuning framework,UniRS. To facilitate cross-scale feature interactions, we introduce the Cross-scale Generic Feature Interaction (CGFI) mechanism, which integrates multi-scale features and uses Generic Feature Query (GFQ) to achieve more generalized representations. Additionally, we propose the Task-specific Low-rank Mixture of Experts (TLMoE) Decomposition. General features are fed into multiple low-rank experts, each capable of capturing different patterns of the features. By dynamically combining the outputs of these experts, features for each sample and each task are encoded dynamically. Furthermore, we have expanded the existing datasets and developed MOTA, a new dataset annotated for three tasks: semantic segmentation, rotated object detection, and multi-label scene classification. Extensive experiments demonstrate the effectiveness of the well-designedUniRS. The code and dataset will be released at https://github.com/Duckyee728/UniRS. Zhiyu Zheng, Jianan He, Lefei Zhang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | DSMT: Dual-Stage Multiscale Transformer for Hyperspectral Snapshot Compressive ImagingabstractSnapshot compressive imaging (SCI) compresses a 3D hyperspectral image (HSI) into a 2D measurement, significantly improving imaging efficiency while preserving the spatial and spectral information inherent in HSI. However, reconstructing high-quality HSIs from compressed measurements remains a core challenge due to the complexity of the inverse problem. Transformer-based methods have recently shown promising performance in HSI reconstruction. Nonetheless, effectively capturing local information, long-range dependencies, and multi-scale features within a reasonable computational cost remains a significant challenge. In this paper, we propose a dual-stage multiscale Transformer (DSMT) tailored for HSI reconstruction, which adopts a coarse-to-fine framework to enhance reconstruction accuracy and network generalization. Specifically, we design a novel U-Net architecture with a dual-branch encoder, where two separate branches process distinct features and are fused to achieve more refined reconstruction results. Full-scale skip connections are introduced to strengthen feature fusion across different stages. To further improve performance, we develop a novel self-attention mechanism called dual-window multiscale multi-head self-attention (DWM-MSA). By utilizing two differently sized windows, DWM-MSA captures long-range dependencies and local information at multiple scales, significantly boosting reconstruction quality. Additionally, we introduce a hybrid positional embedding method, conditional/relative positional embedding (CRPE), which dynamically models both spatial and spectral dependencies, effectively enhancing the Transformer's capacity for HSI reconstruction. Extensive quantitative and qualitative experiments on both the simulated and the real data are conducted to demonstrate the superior performance, stability, and generalization ability of our DSMT. Code of this project is at https://github.com/chenx2000/DSMT. Fulin Luo, Xi Chen 0087, Tan Guo, Xiuwen Gong, Lefei Zhang, Ce Zhu |
IEEE Trans. Image Process. | 5 |
| 2025 | Fine-Grained Image Captioning by Ranking Diffusion TransformerabstractThe CLIP visual feature-based image captioning models have developed rapidly and achieved remarkable results. However, existing models still struggle to produce descriptive and discriminative captions because they insufficiently exploit fine-grained visual cues and fail to model complex vision-language alignment. To address these limitations, we propose a Ranking Diffusion Transformer (RDT), which integrates a Ranking Visual Encoder (RVE) and a Ranking Loss (RL) for fine-grained image captioning. The RVE introduces a novel ranking attention mechanism that effectively mines diverse and discriminative visual information from CLIP features. Meanwhile, the RL leverages the ranking of generated caption quality as a global semantic supervisory signal, thereby enhancing the diffusion process and strengthening vision-language semantic alignment. We show that by collaborating RVE and RL via the novel RDT-and by gradually adding and removing noise in the diffusion process-more discriminative visual features are learned and precisely aligned with the language features. Experimental results on popular benchmark datasets demonstrate that our proposed RDT surpasses existing state-of-the-art image captioning models in the literature. The code is publicly available at: https://github.com/junwan2014/RDT. Jun Wan 0005, Min Gan, Lefei Zhang, Jie Zhou 0009, Jun Liu 0036, Bo Du 0001, C. L. Philip Chen |
IEEE Trans. Image Process. | 3 |
| 2025 | Enhancing Perception of Key Changes in Remote Sensing Image Change CaptioningabstractRecently, while significant progress has been made in remote sensing image change captioning, existing methods fail to filter out areas unrelated to actual changes, making models susceptible to irrelevant features. In this article, we propose a novel multimodal model for remote sensing image change captioning, guided by Key Change Features and Instruction-tuned (KCFI). This model aims to fully leverage the intrinsic knowledge of large language models through visual instructions and enhance the effectiveness and accuracy of change features using pixel-level change detection tasks. Specifically, KCFI includes a ViTs encoder for extracting bi-temporal remote sensing image features, a key feature perceiver for identifying critical change areas, a pixel-level change detection decoder to constrain key change features, and an instruction-tuned decoder based on a large language model. Moreover, to ensure that change captioning and change detection tasks are jointly optimized, we employ a dynamic weight-averaging strategy to balance the losses between the two tasks. We also explore various feature combinations for visual fine-tuning instructions and demonstrate that using only key change features to guide the large language model is the optimal choice. To validate the effectiveness of our approach, we compare it against several state-of-the-art change captioning methods on the LEVIR-CC dataset, achieving the best performance. Our code will be available at https://github.com/yangcong356/KCFI.git. Zuchao Li, Hongzan Jiao, Zhi Gao 0005, Lefei Zhang |
IEEE Trans. Image Process. | 5 |
| 2025 | UniUIR: Considering Underwater Image Restoration as an All-in-One LearnerabstractExisting underwater image restoration (UIR) methods generally only handle color distortion or jointly address color and haze issues, but they often overlook the more complex degradations that can occur in underwater scenes. To address this limitation, we propose a Universal Underwater Image Restoration method, termed as UniUIR, considering the complex scenario of real-world underwater mixed distortions as an all-in-one manner. To disentangle degradation-specific effects and capture their inter-correlations, we propose the Mamba Mixture-of-Experts module (MMoEM). Each expert specializes in distinct aspects of degradation, while gating mechanism dynamically routes features to appropriate experts. This design enables collaborative prior extraction and preserves global context, all within linear computational complexity. Building upon this foundation, to enhance degradation representation and address the task conflicts that arise when handling multiple types of degradation, we introduce the spatial-frequency prior generator. This module extracts degradation prior information in both spatial and frequency domains, and adaptively selects the most appropriate task-specific prompts based on image content, thereby improving the accuracy of image restoration. Finally, to more effectively address complex, region-dependent distortions in UIR task, we incorporate depth information derived from a large-scale pre-trained depth prediction model, thereby enabling the network to perceive and leverage depth variations across different image regions to handle localized degradation. Extensive experiments demonstrate that UniUIR can produce more attractive results across qualitative and quantitative comparisons, and shows strong generalization than state-of-the-art methods. Project page at https://house-yuyu.github.io/UniUIR. Xu Zhang 0044, Huan Zhang 0008, Guoli Wang 0004, Qian Zhang 0009, Lefei Zhang, Bo Du 0001 |
IEEE Trans. Image Process. | 5 |
| 2024 | A Novel Energy Based Model Mechanism for Multi-Modal Aspect-Based Sentiment AnalysisabstractMulti-modal aspect-based sentiment analysis (MABSA) has recently attracted increasing attention. The span-based extraction methods, such as FSUIE, demonstrate strong performance in sentiment analysis due to their joint modeling of input sequences and target labels. However, previous methods still have certain limitations: (i) They ignore the difference in the focus of visual information between different analysis targets (aspect or sentiment). (ii) Combining features from uni-modal encoders directly may not be sufficient to eliminate the modal gap and can cause difficulties in capturing the image-text pairwise relevance. (iii) Existing span-based methods for MABSA ignore the pairwise relevance of target span boundaries. To tackle these limitations, we propose a novel framework called DQPSA. Specifically, our model contains a Prompt as Dual Query (PDQ) module that uses the prompt as both a visual query and a language query to extract prompt-aware visual information and strengthen the pairwise relevance between visual information and the analysis target. Additionally, we introduce an Energy-based Pairwise Expert (EPE) module that models the boundaries pairing of the analysis target from the perspective of an Energy-based Model. This expert predicts aspect or sentiment span based on pairwise stability. Experiments on three widely used benchmarks demonstrate that DQPSA outperforms previous approaches and achieves a new state-of-the-art performance. The code will be released at https://github.com/pengts/DQPSA. Tianshuo Peng, Zuchao Li, Ping Wang 0028, Lefei Zhang, Hai Zhao 0001 |
AAAI | 4 |
| 2024 | A Coin Has Two Sides: A Novel Detector-Corrector Framework for Chinese Spelling CorrectionabstractChinese Spelling Correction (CSC) stands as a foundational Natural Language Processing (NLP) task, which primarily focuses on the correction of erroneous characters in Chinese texts. Certain existing methodologies opt to disentangle the error correction process, employing an additional error detector to pinpoint error positions. However, owing to the inherent performance limitations of error detector, precision and recall are like two sides of the coin which can not be both facing up simultaneously. Furthermore, it is also worth investigating how the error position information can be judiciously applied to assist the error correction. In this paper, we introduce a novel approach based on error detector-corrector framework. Our detector is designed to yield two error detection results, each characterized by high precision and recall. Given that the occurrence of errors is context-dependent and detection outcomes may be less precise, we incorporate the error detection results into the CSC task using an innovative feature fusion strategy and a selective masking strategy. Empirical experiments conducted on mainstream CSC datasets substantiate the efficacy of our proposed method. Xiangke Zeng, Zuchao Li, Lefei Zhang, Ping Wang 0028, Hongqiu Wu, Hai Zhao 0001 |
ECAI | 3 |
| 2024 | VHASR: A Multimodal Speech Recognition System With Vision HotwordsabstractThe image-based multimodal automatic speech recognition (ASR) model enhances speech recognition performance by incorporating audio-related image.However, some works suggest that introducing image information to model does not help improving ASR performance.In this paper, we propose a novel approach effectively utilizing audio-related image information and set up VHASR, a multimodal speech recognition system that uses vision as hotwords to strengthen the model's speech recognition capability.Our system utilizes a dual-stream architecture, which firstly transcribes the text on the two streams separately, and then combines the outputs.We evaluate the proposed model on four datasets: Flickr8k, ADE20k, COCO, and OpenImages.The experimental results show that VHASR can effectively utilize key information in images to enhance the model's speech recognition ability.Its performance not only surpasses unimodal ASR, but also achieves SOTA among existing image-based multimodal ASR. 1 Jiliang Hu 0001, Zuchao Li, Ping Wang 0028, Haojun Ai, Lefei Zhang, Hai Zhao 0001 |
EMNLP | 5 |
| 2024 | Sparse is Enough in Fine-tuning Pre-trained Large Language ModelsabstractWith the prevalence of pre-training-fine-tuning paradigm, how to efficiently adapt the pre-trained model to the downstream tasks has been an intriguing issue. $\textbf{P}$arameter-$\textbf{E}$fficient $\textbf{F}$ine-$\textbf{T}$uning(PEFT) methods have been proposed for low-cost adaptation. Although PEFT has demonstrated effectiveness and been widely applied, the underlying principles are still unclear. In this paper, we adopt the PAC-Bayesian generalization error bound, viewing pre-training as a shift of prior distribution which leads to a tighter bound for generalization error. We validate this shift from the perspectives of oscillations in the loss landscape and the quasi-sparsity in gradient distribution. Based on this, we propose a gradient-based sparse fine-tuning algorithm, named $\textbf{S}$parse $\textbf{I}$ncrement $\textbf{F}$ine-$\textbf{T}$uning(SIFT), and validate its effectiveness on a range of tasks including the GLUE Benchmark and Instruction-tuning. The code is accessible at https://github.com/song-wx/SIFT/. Weixi Song, Zuchao Li, Lefei Zhang, Hai Zhao 0001, Bo Du 0001 |
ICML | 3 |
| 2024 | Merging Multi-Task Models via Weight-Ensembling Mixture of ExpertsabstractMerging various task-specific Transformer-based vision models trained on different tasks into a single unified model can execute all the tasks concurrently. Previous methods, exemplified by task arithmetic, have been proven to be both effective and scalable. Existing methods have primarily focused on seeking a static optimal solution within the original model parameter space. A notable challenge is mitigating the interference between parameters of different models, which can substantially deteriorate performance. In this paper, we propose to merge most of the parameters while upscaling the MLP of the Transformer layers to a weight-ensembling mixture of experts (MoE) module, which can dynamically integrate shared and task-specific knowledge based on the input, thereby providing a more flexible solution that can adapt to the specific needs of each instance. Our key insight is that by identifying and separating shared knowledge and task-specific knowledge, and then dynamically integrating them, we can mitigate the parameter interference problem to a great extent. We conduct the conventional multi-task model merging experiments and evaluate the generalization and robustness of our method. The results demonstrate the effectiveness of our method and provide a comprehensive understanding of our method. The code is available at https://github.com/tanganke/weight-ensembling_MoE Anke Tang, Li Shen 0008, Yong Luo 0002, Lefei Zhang, Dacheng Tao |
ICML | 5 |
| 2024 | Eliminating the Cross-Domain Misalignment in Text-guided Image Inpainting
Muqi Huang, Yong Luo 0002, Lefei Zhang |
IJCAI | 4 |
| 2024 | MAdapter: A Better Interaction Between Image and Language for Medical Image Segmentation
Xu Zhang 0044, Bo Ni, Lefei Zhang |
MICCAI (9) | 4 |
| 2024 | Multi-modal Auto-regressive Modeling via Visual TokensabstractLarge Language Models (LLMs), benefiting from the auto-regressive modelling approach performed on massive unannotated texts corpora, demonstrates powerful perceptual and reasoning capabilities. However, as for extending auto-regressive modelling to multi-modal scenarios to build Large Multi-modal Models (LMMs), there lies a great difficulty that the image information is processed in the LMM as continuous visual embeddings, which cannot obtain discrete supervised labels for classification. In this paper, we successfully perform multi-modal auto-regressive modeling with a unified objective for the first time. Specifically, we propose the concept of visual tokens, which maps the visual features to probability distributions over LLM's vocabulary, providing supervision information for visual modelling. We further explore the distribution of visual features in the semantic space within LMM and the possibility of using text embeddings to represent visual information. Experimental results and ablation studies on 5 VQA tasks and 4 benchmark toolkits validate the powerful performance of our proposed approach. Tianshuo Peng, Zuchao Li, Lefei Zhang, Hai Zhao 0001, Ping Wang 0028, Bo Du 0001 |
ACM Multimedia | 3 |
| 2024 | Driving Scene Understanding with Traffic Scene-Assisted Topology Graph TransformerabstractDriving scene topology reasoning aims to understand the objects present in the current road scene and model their topology relationships to provide guidance information for downstream tasks. Previous approaches fail to adequately facilitate interactions among traffic objects and neglect to incorporate scene information into topology reasoning, thus limiting the comprehensive exploration of potential correlations among objects and diminishing the practical significance of the reasoning results. Besides, the lack of constraints on lane direction may introduce erroneous guidance information and lead to a decrease in topology prediction accuracy. In this paper, we propose a novel topology reasoning framework, dubbed TSTGT, to address these issues. Specifically, we design a divide-and-conquer topology graph Transformer to respectively infer the lane-lane and lane-traffic topology relationships, which can effectively aggregate the local and global object information in the driving scene and facilitate the topology relationship learning. Additionally, a traffic scene-assisted reasoning module is devised and combined with the topology graph Transformer to enhance the practical significance of lane-traffic topology. In terms of lane detection, we develop a point-wise matching strategy to infer lane centerlines with correct directions, thereby improving the topology reasoning accuracy. Extensive experimental results on Openlane-V2 benchmark validate the superiority of our TSTGT over state-of-the-art methods and the effectiveness of our proposed modules. The code is available at https://github.com/rongfu-dsb/TSTGT. Fu Rong, Wenjin Peng, Meng Lan, Qian Zhang 0009, Lefei Zhang |
ACM Multimedia | 5 |
| 2024 | InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward ModelingabstractDespite the success of reinforcement learning from human feedback (RLHF) in aligning language models with human values, reward hacking, also termed reward overoptimization, remains a critical challenge. This issue primarily arises from reward misgeneralization, where reward models (RMs) compute reward using spurious features that are irrelevant to human preferences. In this work, we tackle this problem from an information-theoretic perspective and propose a framework for reward modeling, namely InfoRM, by introducing a variational information bottleneck objective to filter out irrelevant information.
Notably, we further identify a correlation between overoptimization and outliers in the IB latent space of InfoRM, establishing it as a promising tool for detecting reward overoptimization.
Inspired by this finding, we propose the Cluster Separation Index (CSI), which quantifies deviations in the IB latent space, as an indicator of reward overoptimization to facilitate the development of online mitigation strategies. Extensive experiments on a wide range of settings and RM scales (70M, 440M, 1.4B, and 7B) demonstrate the effectiveness of InfoRM. Further analyses reveal that InfoRM's overoptimization detection mechanism is not only effective but also robust across a broad range of datasets, signifying a notable advancement in the field of RLHF. The code will be released upon acceptance. Yuchun Miao, Sen Zhang 0006, Liang Ding 0006, Rong Bao, Lefei Zhang, Dacheng Tao |
NeurIPS | 5 |
| 2024 | MMSite: A Multi-modal Framework for the Identification of Active Sites in ProteinsabstractThe accurate identification of active sites in proteins is essential for the advancement of life sciences and pharmaceutical development, as these sites are of critical importance for enzyme activity and drug design. Recent advancements in protein language models (PLMs), trained on extensive datasets of amino acid sequences, have significantly improved our understanding of proteins. However, compared to the abundant protein sequence data, functional annotations, especially precise per-residue annotations, are scarce, which limits the performance of PLMs. On the other hand, textual descriptions of proteins, which could be annotated by human experts or a pretrained protein sequence-to-text model, provide meaningful context that could assist in the functional annotations, such as the localization of active sites. This motivates us to construct a $\textbf{ProT}$ein-$\textbf{A}$ttribute text $\textbf{D}$ataset ($\textbf{ProTAD}$), comprising over 570,000 pairs of protein sequences and multi-attribute textual descriptions. Based on this dataset, we propose $\textbf{MMSite}$, a multi-modal framework that improves the performance of PLMs to identify active sites by leveraging biomedical language models (BLMs). In particular, we incorporate manual prompting and design a MACross module to deal with the multi-attribute characteristics of textual descriptions. MMSite is a two-stage ("First Align, Then Fuse") framework: first aligns the textual modality with the sequential modality through soft-label alignment, and then identifies active sites via multi-modal fusion. Experimental results demonstrate that MMSite achieves state-of-the-art performance compared to existing protein representation learning methods. The dataset and code implementation are available at https://github.com/Gift-OYS/MMSite. Song Ouyang, Huiyu Cai, Yong Luo 0002, Kehua Su, Lefei Zhang, Bo Du 0001 |
NeurIPS | 5 |
| 2024 | Reference Trustable Decoding: A Training-Free Augmentation Paradigm for Large Language ModelsabstractLarge language models (LLMs) have rapidly advanced and demonstrated impressive capabilities. In-Context Learning (ICL) and Parameter-Efficient Fine-Tuning (PEFT) are currently two mainstream methods for augmenting LLMs to downstream tasks. ICL typically constructs a few-shot learning scenario, either manually or by setting up a Retrieval-Augmented Generation (RAG) system, helping models quickly grasp domain knowledge or question-answering patterns without changing model parameters. However, this approach involves trade-offs, such as slower inference speed and increased space occupancy. PEFT assists the model in adapting to tasks through minimal parameter modifications, but the training process still demands high hardware requirements, even with a small number of parameters involved. To address these challenges, we propose Reference Trustable Decoding (RTD), a paradigm that allows models to quickly adapt to new tasks without fine-tuning, maintaining low inference costs. RTD constructs a reference datastore from the provided training examples and optimizes the LLM's final vocabulary distribution by flexibly selecting suitable references based on the input, resulting in more trustable responses and enabling the model to adapt to downstream tasks at a low cost. Experimental evaluations on various LLMs using different benchmarks demonstrate that RTD establishes a new paradigm for augmenting models to downstream tasks. Furthermore, our method exhibits strong orthogonality with traditional methods, allowing for concurrent usage. Our code can be found at https://github.com/ShiLuohe/ReferenceTrustableDecoding. Luohe Shi, Yao Yao 0008, Zuchao Li, Lefei Zhang, Hai Zhao 0001 |
NeurIPS | 4 |
| 2024 | Centroid-Centered Modeling for Efficient Vision Transformer Pre-Training
Xin Yan 0008, Zuchao Li, Lefei Zhang |
PRCV (4) | 3 |
| 2024 | Convex-Concave Tensor Robust Principal Component Analysis
Youfa Liu, Bo Du 0001, Yongyong Chen, Lefei Zhang, Mingming Gong, Dacheng Tao |
Int. J. Comput. Vis. | 4 |
| 2024 | DSQNet: Enhancing Change Detection Network via Deep Semantics Query for Remote Sensing ImagesabstractChange detection from remote sensing (RS) images has made significant progress in many applications including environmental protection and agricultural monitoring. Recently, RS change detection algorithms mainly focus on feature interaction between the bitemporal images. However, there is a challenge in existing methods regarding how to focus more attention on local prominent features and specifically enhance their salience during the interaction of representations. For this issue, this letter presents a change detection network DSQNet that marks multiple regions of interest using query vectors with deep semantics and explicitly searches the heterogeneous features for focused interaction. Specifically, DSQNet uses bidirectional matching query (BMQ) module to effectively perceive local relevant features by feature space query for cross matching and adequately enhance the context relationship within specific regions during the interaction, which helps the model better learn the idea of potential change. Moreover, to make the query vectors and the feature representation aligned in the semantic space, we further propose the deep semantic adjustment feature pyramid (DSP) module. It realizes interlayer feature adjustment from the inside out in the pyramid and enables query vectors to represent extremely rich semantics, improving query efficiency. Experimental results on four benchmark datasets show that DSQNet achieves better performance than other advanced change detection networks. The codes are available athttps://github.com/QinZelin/DSQNet/. Zilin Qin, Lefei Zhang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Enhancing the Semi-Supervised Semantic Segmentation With Prototype-Based Supervision for Remote Sensing ImagesabstractWhile image semantic segmentation is a fundamental and well-studied task in remote sensing (RS) society, it usually depends on large amounts of pixel-level annotations. RS image semi-supervised semantic segmentation (RSIS4) tries to improve performance by exploring the unlabeled data, thus significantly reducing the label costs. The core idea of RSIS4 is to transfer the prior information from the labeled to unlabeled pixels, which is commonly achieved by considering the confident part of the softmax prediction as pseudolabels for further supervised learning. However, such pixel-level instruction could inevitably involve uncertainty (e.g., noise and error) due to the extremely limited annotated data at the initialization. To address this issue, in this letter, we employ the prototypes, which contain inbuilt resistance to potentially inaccurate pixels, to bring substantial supervision directly from the embedded feature space. Specifically, we project deep features into the embedding space to generate prototypes, each of which can be regarded as the category-level feature representation of a certain semantic category. These prototypes are then used to perform the pixelwise classification, with the advantage of capturing the global similarity throughout the whole pixels within the category. Moreover, to ensure accurate prototypes, we further introduce pixel-prototype contrast to better explore the discriminative category-level feature embedding. By integrating the guidance from the above pixel-level and category-level feature representations, the proposed algorithm obtains high-quality pseudolabels and extracts effective features. Extensive experiments on four RS image segmentation datasets have demonstrated the effectiveness of the proposed method. The code is available athttps://github.com/Duckyee728/PCSSS.git. Zhiyu Zheng, Lefei Zhang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Bidirectional correlation-driven inter-frame interaction Transformer for referring video object segmentation
Meng Lan, Fu Rong, Zuchao Li, Wei Yu 0009, Lefei Zhang |
Pattern Recognit. | 5 |
| 2024 | Robust multiple subspaces transfer for heterogeneous domain adaptation
Youfa Liu, Bo Du 0001, Yongyong Chen, Lefei Zhang |
Pattern Recognit. | 4 |
| 2024 | Multi-Task Learning With Multi-Query Transformer for Dense PredictionabstractPrevious multi-task dense prediction studies developed complex pipelines such as multi-modal distillations in multiple stages or searching for task relational contexts for each task. The core insight beyond these methods is to maximize the mutual effects of each task. Inspired by the recent query-based Transformers, we propose a simple pipeline named Multi-Query Transformer (MQTransformer) that is equipped with multiple queries from different tasks to facilitate the reasoning among multiple tasks and simplify the cross-task interaction pipeline. Instead of modeling the dense per-pixel context among different tasks, we seek a task-specific proxy to perform cross-task reasoning via multiple queries where each query encodes the task-related context. The MQTransformer is composed of three key components: shared encoder, cross-task query attention module and shared decoder. We first model each task with a task-relevant query. Then both the task-specific feature output by the feature extractor and the task-relevant query are fed into the shared encoder, thus encoding the task-relevant query from the task-specific feature. Secondly, we design a cross-task query attention module to reason the dependencies among multiple task-relevant queries; this enables the module to only focus on the query-level interaction. Finally, we use a shared decoder to gradually refine the image features with the reasoned query features from different tasks. Extensive experiment results on two dense prediction datasets (NYUD-v2 and PASCAL-Context) show that the proposed method is an effective approach and achieves state-of-the-art results. Code and models are available athttps://github.com/yangyangxu0/MQTransformer. Yangyang Xu 0001, Xiangtai Li, Haobo Yuan, Lefei Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Language Query-Based Transformer With Multiscale Cross-Modal Alignment for Visual Grounding on Remote Sensing ImagesabstractVisual grounding for remote sensing images (RSVG) aims to localize the referred objects in the remote sensing (RS) images according to a language expression. Existing methods tend to align visual and text features followed by concatenation and then employ a fusion Transformer to learn a token representation for final target localization. However, simple fusion Transformer structure fails to sufficiently learn the location representation of referred object from the multi-modal features. Inspired by the detection Transformer, in this paper, we propose a novel language query based Transformer framework for RSVG termed LQVG. Specifically, we adopt the extracted sentence-level text features as the queries, called language queries, to retrieve and aggregate representation information of the referred object from the multi-scale visual features in the Transformer decoder. The language queries are then converted into object embeddings for final coordinate prediction of referred object. Besides, a multi-scale cross-modal alignment module is devised before the multimodal Transformer to enhance the semantic correlation between the visual and text features, thus facilitating the cross-modal decoding process to generate more precise object representation. Moreover, a new RSVG dataset named RSVG-HR is built to evaluate the performance of the RSVG approaches on very high-resolution remote sensing images with inconspicuous objects. Experimental results on two benchmark datasets demonstrate that our proposed method significantly surpasses the comparison methods and achieves state-of-the-art performance. The dataset and code are available at https://github.com/LANMNG/LQVG. Meng Lan, Fu Rong, Hongzan Jiao, Zhi Gao 0005, Lefei Zhang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | MambaHSI: Spatial-Spectral Mamba for Hyperspectral Image ClassificationabstractTransformer has been extensively explored for hyperspectral image (HSI) classification. However, transformer poses challenges in terms of speed and memory usage because of its quadratic computational complexity. Recently, the Mamba model has emerged as a promising approach, which has strong long-distance modeling capabilities while maintaining a linear computational complexity. However, representing the HSI is challenging for the Mamba due to the requirement for an integrated spatial and spectral understanding. To remedy these drawbacks, we propose a novel HSI classification model based on a Mamba model, named MambaHSI, which can simultaneously model long-range interaction of the whole image and integrate spatial and spectral information in an adaptive manner. Specifically, we design a spatial Mamba block (SpaMB) to model the long-range interaction of the whole image at the pixel-level. Then, we propose a spectral Mamba block (SpeMB) to split the spectral vector into multiple groups, mine the relations across different spectral groups, and extract spectral features. Finally, we propose a spatial-spectral fusion module (SSFM) to adaptively integrate spatial and spectral features of a HSI. To our best knowledge, this is the first image-level HSI classification model based on the Mamba. We conduct extensive experiments on four diverse HSI datasets. The results demonstrate the effectiveness and superiority of the proposed model for HSI classification. This reveals the great potential of Mamba to be the next-generation backbone for HSI models. Codes are available athttps://github.com/li-yapeng/MambaHSI. Yong Luo 0002, Lefei Zhang, Zengmao Wang, Bo Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | SDST: Self-Supervised Double-Structure Transformer for Hyperspectral Images ClusteringabstractDue to the lack of labeled information and the high spectral variability in high-dimensional hyperspectral images (HSI), HSI clustering has emerged as an effective unsupervised approach for HSI information extraction and classification. Deep clustering methods have achieved significant success in unsupervised HSI classification (HSIC) and have gained increasing attention. However, these methods have limitations in terms of robustness, adaptability, and feature representation when dealing with complex large-scale HSI datasets. Therefore, this paper introduced a novel Self-supervised Double-Structure Transformer (SDST) approach for hyperspectral image clustering. Specifically, in our approach, we designed a shared Autoformer structure based on autoencoder to learn the global properties of HSI data by fusing the multi-level features from autoencoder with Transformer. Furthermore, we proposed a siamese Dual-Former Graph Module with superpixel-level features for fewer nodes, which reveals long-dependency graph convolutional features, resulting in more precise graph structure features. By constructing graph with long-dependencies, this module significantly preserves the properties of global dependencies, while focusing on the local features of each superpixel to better represent the fine-grained local details. Finally, we designed a Joint Optimization Module to jointly optimize the double-structure model composed of the shared Autoformer Module and the siamese Dual-Former Graph Module. To validate the effectiveness of the proposed SDST method, we conducted a series of experiments on the Salinas, Botswana, Indian Pines, and Houston2013 datasets. The proposed SDST achieves competitive clustering accuracies compared with the state-of-the-art clustering methods. Codes: https://github.com/YiLiu1999/SDST. Fulin Luo, Yi Liu 0038, Tan Guo, Lefei Zhang, Bo Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | AGMS: Adversarial Sample Generation-Based Multiscale Siamese Network for Hyperspectral Target DetectionabstractHyperspectral target detection (HTD) has been a critical issue in the field of Earth observation, with widespread applications in both military and civilian domains. However, existing deep learning-based HTD methods are hindered due to insufficient and low-quality prior training samples, as well as inadequate background suppression capabilities. To address these issues, this article proposes an adversarial sample generation-based multiscale Siamese network (AGMS) for HTD. First, the AGMS utilizes the idea of generative adversarial learning based on the prior few targets and diverse backgrounds to generate adversarial target-background sample pairs, thereby producing high-quality training samples, which enhances the distinctiveness between the target and background samples by adversarially training the generator to produce the target/background samples. In addition, to further highlight the targets and suppress the backgrounds, a difference amplification loss and an adaptive weighted binary cross-entropy loss are proposed. Finally, a multiscale convolutional Siamese network model is designed to explore the generated spectral information at multiple levels and achieve target detection through contrastive learning. Numerous experimental results on four real HSI datasets verify the superiority of the AGMS in comparison to many classical and recently proposed HTD methods. The codes are available athttps://github.com/ShissHAN/AGMS. Fulin Luo, Shanshan Shi, Tan Guo, Yanni Dong, Lefei Zhang, Bo Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Advancing Data-Efficient Exploitation for Semi-Supervised Remote Sensing Images Semantic SegmentationabstractTo reduce the dependence of remote sensing (RS) image semantic segmentation models on extensive pixel-level annotated images, this paper aims to address the issue of insufficient exploitation of RS images potential within existing semi-supervised learning methods, introducing a novel semi-supervised RS image semantic segmentation method. Specifically, for unlabeled samples, the multi-perturbation dynamic consistency (MDC) is proposed to align multiple predictions from diverse data augmentations, MDC leverages a dynamic decay threshold instead of fixed thresholds to learn more reliable information, enriching the perturbation space and assisting the segmentation model in acquiring more discriminative feature representations. Furthermore, considering the rich contextual information in RS images, the class prototype memory (CPM) derived from labeled samples is maintained during the training stage, which is leveraged to guide the refinement of predictions from segmentation model at the inference stage. Extensive experiments are conducted on six RS image semantic segmentation datasets, including DFC22, iSAID, MER, MSL, GID-15, and Vaihingen. The experimental results demonstrate the superiority of the proposed method. The code is available at https://github.com/lvliang6879/MCSS. Lefei Zhang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Spatial Focused Bitemporal Interactive Network for Remote Sensing Image Change DetectionabstractRecently, transformers have been widely explored in remote sensing image change detection (RSICD) and achieved remarkable performance. However, most existing transformer-based change detection methods overlook exploring the spatiotemporal relationships between bitemporal images at the features within the same layer, which is crucial for learning discriminative features to perceive changes. Moreover, no explicit spatial constraint has been imposed on the final fused bitemporal features, leading to reduced detection performance on small targets. To address these issues, we propose a spatial focused bitemporal interactive network (SFBI-Net) for RSICD. Specifically, a bitemporal spatiotemporal interactive (BSI) module is proposed, which performs global interactions on bitemporal features at the same network layer and supplements local information to obtain spatiotemporal relationships of bitemporal features for discriminative representation. Furthermore, a spatial focus diversity loss (SFD-Loss) is developed to maximize bitemporal features in the spatial dimension and further enhance the feature representation of change areas, especially small target areas. The experimental results on challenging benchmark datasets demonstrate the superiority of our SFBI-Net. The source code is available at (https://github.com/Mryao-yuan/SFBI-Net). Lefei Zhang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | S³ANet: Spatial-Spectral Self-Attention Learning Network for Defending Against Adversarial Attacks in Hyperspectral Image ClassificationabstractDeep neural networks have demonstrated impressive capabilities in hyperspectral image (HSI) classification tasks. However, they are highly vulnerable to adversarial attacks, raising significant security concerns, especially in the remote sensing community. Even subtle adversarial perturbations that are imperceptible to human observers can mislead deep learning (DL) models and result in incorrect predictions. Therefore, ensuring the robustness of DL models has become a critical focus in addressing security-related remote sensing tasks. Considerable progress has been made in defending against adversarial attacks in HSI classification. Nevertheless, existing methods primarily concentrate on spatial relationships between pixels while overlooking the valuable spectral information present in HSI. Besides, these methods are usually limited to a specific scale and cannot accommodate the precise classification demands for ground objects with various scales. To address these limitations, we propose a spatial–spectral self-attention learning network (S3ANet) for defending against adversarial attacks in HSI classification. Our S3ANet incorporates a pyramid spatial attention learning module to effectively capture spatial dependency at multiple scales. In addition, it utilizes a global spectral transformer to establish correlations between pixels in the spectral dimension. By employing the defense method of spatial–spectral fusion, our model can effectively address adversarial attacks from a comprehensive perspective, seamlessly integrating both spatial and spectral information. Extensive experiments conducted on four benchmark HSI datasets illustrate that the proposed S3ANet achieves competitive performance compared to state-of-the-art methods when faced with adversarial attacks. The code is available online athttps://github.com/YichuXu/S3ANet. Yichu Xu, Yonghao Xu, Hongzan Jiao, Zhi Gao 0005, Lefei Zhang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Bootstrapping Interactive Image-Text Alignment for Remote Sensing Image CaptioningabstractRecently, remote sensing image captioning has gained significant attention in the remote sensing community. Due to the significant differences in spatial resolution of remote sensing images, existing methods in this field have predominantly concentrated on the fine-grained extraction of remote sensing image features, but they cannot effectively handle the semantic consistency between visual features and textual features. To efficiently align the image-text, we propose a novel two-stage vision-language pre-training-based approach to bootstrap interactive image-text alignment for remote sensing image captioning, called BITA, which relies on the design of a lightweight interactive Fourier Transformer to better align remote sensing image-text features. The Fourier layer in the interactive Fourier Transformer is capable of extracting multi-scale features of remote sensing images in the frequency domain, thereby reducing the redundancy of remote sensing visual features. Specifically, the first stage involves preliminary alignment through image-text contrastive learning, which aligns the learned multi-scale remote sensing features from the interactive Fourier Transformer with textual features. In the second stage, the interactive Fourier Transformer connects the frozen image encoder with a large language model. Then, prefix causal language modeling is utilized to guide the text generation process using visual features. Ultimately, across the UCM-caption, RSICD, and NWPU-caption datasets, the experimental results clearly demonstrate that BITA outperforms other advanced comparative approaches. The code is available at https://github.com/yangcong356/BITA. Zuchao Li, Lefei Zhang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | MGIMM: Multi-Granularity Instruction Multimodal Model for Attribute-Guided Remote Sensing Image Detailed DescriptionabstractGenerating detailed textual descriptions of remote sensing images is challenging because it requires capturing both global and local visual information. The complexity of backgrounds and the scale variations among targets make it difficult to align visual regions with corresponding textual attributes. Furthermore, large multimodal models, while effective in general scenarios, struggle in remote sensing due to their lack of specialized knowledge and regional awareness. To address these issues, this article proposes an attribute-guided multi-granularity instruction multimodal model (MGIMM) for remote sensing image detailed description. MGIMM guides the multimodal model to learn the consistency between visual regions and corresponding text attributes (such as object names, colors, and shapes) through region-level instruction tuning. Then, with the multimodal model aligned on region attribute, guided by multigrain visual features, MGIMM fully perceives both region-level and global image information, utilizing large language models for comprehensive descriptions of remote sensing images. Due to the lack of a standard benchmark for generating detailed descriptions of remote sensing images, we construct a dataset featuring 38320 region-attribute pairs and 23463 image-detailed description pairs. Compared with various advanced methods on this dataset, the results demonstrate the effectiveness of MGIMM’s region-attribute-guided learning approach. The code is available athttps://github.com/yangcong356/MGIMM.git. Zuchao Li, Lefei Zhang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Dual Graph Learning Affinity Propagation for Multimodal Remote Sensing Image ClusteringabstractMultimodal remote sensing image recognition aims to identify a category of land cover for every pixel with consistency and complementary information provided by different modalities. Most existing methods perform land cover recognition in a supervised manner with explicit label guidance. It is challenging to perform recognition without label guidance due to the complex spatial distribution and modality incompatibility, especially for large-scale data. In this article, we propose a dual graph learning affinity propagation (DGLAP) method for multimodal remote sensing image clustering. Based on the consistent spatial distribution from local regions, the proposed method learns an$N \times M$consensus anchor graph from N denoised pixels and M anchors by adaptive weighting different modalities along with projection learning. Meanwhile, an optimal$M \times M$compressed consensus anchor graph is learned from the updated anchors in different modalities with diverse adaptive contributions and connectivity constraint. Since$M \ll N$, clustering results can be efficiently obtained according to affinity propagation from the pseudolabeled anchors to the pixels without additional steps. An alternating optimization algorithm is devised to solve the proposed formulation. This is the first attempt to propose a ultraefficient graph-based clustering method with linear time complexity$\mathcal {O}(N)$and low time cost for large-scale multimodal remote sensing data. Extensive experiments on three datasets demonstrate the superiority of the proposed method over the state-of-the-art methods in both efficacy and efficiency. The code is released athttps://github.com/ZhangYongshan/DGLAP. Yongshan Zhang, Shuaikang Yan, Xinwei Jiang, Lefei Zhang, Zhihua Cai, Jun Li 0009 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Cross-Scope Spatial-Spectral Information Aggregation for Hyperspectral Image Super-ResolutionabstractHyperspectral image super-resolution has attained widespread prominence to enhance the spatial resolution of hyperspectral images. However, convolution-based methods have encountered challenges in harnessing the global spatial-spectral information. The prevailing transformer-based methods have not adequately captured the long-range dependencies in both spectral and spatial dimensions. To alleviate this issue, we propose a novel cross-scope spatial-spectral Transformer (CST) to efficiently investigate long-range spatial and spectral similarities for single hyperspectral image super-resolution. Specifically, we devise cross-attention mechanisms in spatial and spectral dimensions to comprehensively model the long-range spatial-spectral characteristics. By integrating global information into the rectangle-window self-attention, we first design a cross-scope spatial self-attention to facilitate long-range spatial interactions. Then, by leveraging appropriately characteristic spatial-spectral features, we construct a cross-scope spectral self-attention to effectively capture the intrinsic correlations among global spectral bands. Finally, we elaborate a concise feed-forward neural network to enhance the feature representation capacity in the Transformer structure. Extensive experiments over three hyperspectral datasets demonstrate that the proposed CST is superior to other state-of-the-art methods both quantitatively and visually. The code is available at https://github.com/Tomchenshi/CST.git. Shi Chen 0010, Lefei Zhang, Liangpei Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2024 | Tracking With Saliency Region TransformerabstractTransformers show a great impact on visual tracking thanks to their powerful representation learning capabilities. As the capacity of the model grows, the speed of the tracker tends to decrease gradually. Our work focuses on dealing with massively redundant information in tracking sequences with the Saliency Region Tracker (SRTrack). SRTrack is a heuristic two-stage tracker consisting of a lightweight tracking stage and a saliency stage. The former can handle simple tracking sequences while the latter is designed to perform delicate tracking on challenging frames with more discriminative features. However, the two-stage design leads to feature extrapolation, creating inconsistencies between training and inference features. In order to mitigate this problem, we develop an attention scaling factor that guarantees model robustness while yielding a slight performance gain. Our SRTrack achieves a state-of-the-art 0.699 AUC running at 61 FPS on LaSOT. Several experiments on large benchmarks demonstrate the high efficiency and accuracy of SRTrack. Tianpeng Liu, Jing Li 0055, Jia Wu 0001, Lefei Zhang, Jun Wan 0005, Lezhi Lian |
IEEE Trans. Image Process. | 4 |
| 2024 | Fast Projected Fuzzy Clustering With Anchor Guidance for Multimodal Remote Sensing ImageryabstractMultimodal remote sensing image recognition is a popular research topic in the field of remote sensing. This recognition task is mostly solved by supervised learning methods that heavily rely on manually labeled data. When the labels are absent, the recognition is challenging for the large data size, complex land-cover distribution and large modality spectrum variation. In this paper, a novel unsupervised method, named fast projected fuzzy clustering with anchor guidance (FPFC), is proposed for multimodal remote sensing imagery. Specifically, according to the spatial distribution of land covers, meaningful superpixels are obtained for denoising and generating high-quality anchor. The denoised data and anchors are projected into the optimal subspace to jointly learn the shared anchor graph as well as the shared anchor membership matrix from different modalities in an adaptively weighted manner to accelerate the clustering process. Finally, the shared anchor graph and shared anchor membership matrix are combined to derive clustering labels for all pixels. An effective alternating optimization algorithm is designed to solve the proposed formulation. This is the first attempt to propose a soft clustering method for large-scale multimodal remote sensing data. Experiments show that the proposed FPFC achieves 81.34%, 55.43% and 93.34% clustering accuracies on the three datasets and outperforms the state-of-the-art methods. The source code is released at https://github.com/ZhangYongshan/FPFC. Yongshan Zhang, Shuaikang Yan, Lefei Zhang, Bo Du 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | MICCF: A Mutual Information Constrained Clustering Framework for Learning Clustering-Oriented Feature RepresentationsabstractDeep clustering is a crucial task in machine learning and data mining that focuses on acquiring feature representations conducive to clustering. Previous research relies on self-supervised representation learning for general feature representations, such features may not be optimally suited for downstream clustering tasks. In this article, we introduce MICCF, a framework designed to bridge this gap and enhance clustering performance. MICCF enhances feature representations by combining mutual information constraints at different levels and employs an auxiliary alignment mutual information module for learning clustering-oriented features. To be specific, we propose a dual mutual information constraints module, incorporating minimal mutual information constraints at the feature level and maximal mutual information constraints at the instance level. This reduction in feature redundancy encourages the neural network to extract more discriminative features, while maximization ensures more unbiased and robust representations. To obtain clustering-oriented representations, the auxiliary alignment mutual information module utilizes pseudo-labels to maximize mutual information through a multi-classifier network, aligning features with the clustering task. The main network and the auxiliary module work in synergy to jointly optimize feature representations that are well-suited for the clustering task. We validate the effectiveness of our method through extensive experiments on six benchmark datasets. The results indicate that our method performs well in most scenarios, particularly on fine-grained datasets, where our approach effectively distinguishes subtle differences between closely related categories. Notably, our approach achieved a remarkable accuracy of 96.4% on the ImageNet-10 dataset, surpassing other comparison methods. The code is available at https://github.com/Li-Hyn/MICCF.git . Hongyu Li 0004, Lefei Zhang, Kehua Su, Wei Yu 0009 |
ACM Trans. Knowl. Discov. Data | 2 |
| 2023 | Learning to Learn Better for Video Object SegmentationabstractRecently, the joint learning framework (JOINT) integrates matching based transductive reasoning and online inductive learning to achieve accurate and robust semi-supervised video object segmentation (SVOS). However, using the mask embedding as the label to guide the generation of target features in the two branches may result in inadequate target representation and degrade the performance. Besides, how to reasonably fuse the target features in the two different branches rather than simply adding them together to avoid the adverse effect of one dominant branch has not been investigated. In this paper, we propose a novel framework that emphasizes Learning to Learn Better (LLB) target features for SVOS, termed LLB, where we design the discriminative label generation module (DLGM) and the adaptive fusion module to address these issues. Technically, the DLGM takes the background-filtered frame instead of the target mask as input and adopts a lightweight encoder to generate the target features, which serves as the label of the online few-shot learner and the value of the decoder in the transformer to guide the two branches to learn more discriminative target representation. The adaptive fusion module maintains a learnable gate for each branch, which reweighs the element-wise feature representation and allows an adaptive amount of target information in each branch flowing to the fused target feature, thus preventing one branch from being dominant and making the target feature more robust to distractor. Extensive experiments on public benchmarks show that our proposed LLB method achieves state-of-the-art performance. Meng Lan, Jing Zhang 0037, Lefei Zhang, Dacheng Tao |
AAAI | 3 |
| 2023 | Dual Mutual Information Constraints for Discriminative ClusteringabstractDeep clustering is a fundamental task in machine learning and data mining that aims at learning clustering-oriented feature representations. In previous studies, most of deep clustering methods follow the idea of self-supervised representation learning by maximizing the consistency of all similar instance pairs while ignoring the effect of feature redundancy on clustering performance. In this paper, to address the above issue, we design a dual mutual information constrained clustering method named DMICC which is based on deep contrastive clustering architecture, in which the dual mutual information constraints are particularly employed with solid theoretical guarantees and experimental validations. Specifically, at the feature level, we reduce the redundancy among features by minimizing the mutual information across all the dimensionalities to encourage the neural network to extract more discriminative features. At the instance level, we maximize the mutual information of the similar instance pairs to obtain more unbiased and robust representations. The dual mutual information constraints happen simultaneously and thus complement each other to jointly optimize better features that are suitable for the clustering task. We also prove that our adopted mutual information constraints are superior in feature extraction, and the proposed dual mutual information constraints are clearly bounded and thus solvable. Extensive experiments on five benchmark datasets show that our proposed approach outperforms most other clustering algorithms. The code is available at https://github.com/Li-Hyn/DMICC. Hongyu Li 0004, Lefei Zhang, Kehua Su |
AAAI | 2 |
| 2023 | DeMT: Deformable Mixer Transformer for Multi-Task Learning of Dense PredictionabstractConvolution neural networks (CNNs) and Transformers have their own advantages and both have been widely used for dense prediction in multi-task learning (MTL). Most of the current studies on MTL solely rely on CNN or Transformer. In this work, we present a novel MTL model by combining both merits of deformable CNN and query-based Transformer for multi-task learning of dense prediction. Our method, named DeMT, is based on a simple and effective encoder-decoder architecture (i.e., deformable mixer encoder and task-aware transformer decoder). First, the deformable mixer encoder contains two types of operators: the channel-aware mixing operator leveraged to allow communication among different channels (i.e., efficient channel location mixing), and the spatial-aware deformable operator with deformable convolution applied to efficiently sample more informative spatial locations (i.e., deformed features). Second, the task-aware transformer decoder consists of the task interaction block and task query block. The former is applied to capture task interaction features via self-attention. The latter leverages the deformed features and task-interacted features to generate the corresponding task-specific feature through a query-based Transformer for corresponding task predictions. Extensive experiments on two dense image prediction datasets, NYUD-v2 and PASCAL-Context, demonstrate that our model uses fewer GFLOPs and significantly outperforms current Transformer- and CNN-based competitive models on a variety of metrics. The code is available at https://github.com/yangyangxu0/DeMT. Lefei Zhang |
AAAI | 3 |
| 2023 | FSUIE: A Novel Fuzzy Span Mechanism for Universal Information ExtractionabstractUniversal Information Extraction (UIE) has been introduced as a unified framework for various Information Extraction (IE) tasks and has achieved widespread success.Despite this, UIE models have limitations.For example, they rely heavily on span boundaries in the data during training, which does not reflect the reality of span annotation challenges.Slight adjustments to positions can also meet requirements.Additionally, UIE models lack attention to the limited span length feature in IE.To address these deficiencies, we propose the Fuzzy Span Universal Information Extraction (FSUIE) framework.Specifically, our contribution consists of two concepts: fuzzy span loss and fuzzy span attention.Our experimental results on a series of main IE tasks show significant improvement compared to the baseline, especially in terms of fast convergence and strong performance with small amounts of data and training epochs.These results demonstrate the effectiveness and generalization of FSUIE in different tasks, settings, and scenarios. Tianshuo Peng, Zuchao Li, Lefei Zhang, Bo Du 0001, Hai Zhao 0001 |
ACL (1) | 3 |
| 2023 | DDS2M: Self-Supervised Denoising Diffusion Spatio-Spectral Model for Hyperspectral Image RestorationabstractDiffusion models have recently received a surge of interest due to their impressive performance for image restoration, especially in terms of noise robustness. However, existing diffusion-based methods are trained on a large amount of training data and perform very well in-distribution, but can be quite susceptible to distribution shift. This is especially inappropriate for data-starved hyperspectral image (HSI) restoration. To tackle this problem, this work puts forth a self-supervised diffusion model for HSI restoration, namely Denoising Diffusion Spatio-Spectral Model (DDS2M), which works by inferring the parameters of the proposed Variational Spatio-Spectral Module (VS2M) during the reverse diffusion process, solely using the degraded HSI without any extra training data. In VS2M, a variational inference-based loss function is customized to enable the untrained spatial and spectral networks to learn the posterior distribution, which serves as the transitions of the sampling chain to help reverse the diffusion process. Benefiting from its self-supervised nature and the diffusion process, DDS2M enjoys stronger generalization ability to various HSIs compared to existing diffusion-based methods and superior robustness to noise compared to existing HSI restoration methods. Extensive experiments on HSI denoising, noisy HSI completion and super-resolution on a variety of HSIs demonstrate DDS2M’s superiority over the existing task-specific state-of-the-arts. Code is available at: https://github.com/miaoyuchun/DDS2M. Yuchun Miao, Lefei Zhang, Liangpei Zhang 0001, Dacheng Tao |
ICCV | 2 |
| 2023 | Multi-Task Learning with Knowledge Distillation for Dense PredictionabstractWhile multi-task learning (MTL) has become an attractive topic, its training usually poses more difficulties than the single-task case. How to successfully apply knowledge distillation into MTL to improve training efficiency and model performance is still a challenging problem. In this paper, we introduce a new knowledge distillation procedure with an alternative match for MTL of dense prediction based on two simple design principles. First, for memory and training efficiency, we use a single strong multitask model as a teacher during training instead of multiple teachers, as widely adopted in existing studies. Second, we employ a less sensitive Cauchy-Schwarz (CS) divergence instead of the Kullback–Leibler (KL) divergence and propose a CS distillation loss accordingly. With the less sensitive divergence, our knowledge distillation with an alternative match is applied for capturing inter-task and intratask information between the teacher model and the student model of each task, thereby learning more "dark knowledge" for effective distillation. We conducted extensive experiments on dense prediction datasets, including NYUD-v2 and PASCAL-Context, for multiple vision tasks, such as semantic segmentation, human parts segmentation, depth estimation, surface normal estimation, and boundary detection. The results show that our proposed method decidedly improves model performance and the practical inference efficiency. Lefei Zhang |
ICCV | 3 |
| 2023 | Bidirectional Looking with A Novel Double Exponential Moving Average to Adaptive and Non-adaptive Momentum OptimizersabstractOptimizer is an essential component for the success of deep learning, which guides the neural network to update the parameters according to the loss on the training set. SGD and Adam are two classical and effective optimizers on which researchers have proposed many variants, such as SGDM and RAdam. In this paper, we innovatively combine the backward-looking and forward-looking aspects of the optimizer algorithm and propose a novel Admeta (**A** **D**ouble exponential **M**oving averag**E** **T**o **A**daptive and non-adaptive momentum) optimizer framework. For backward-looking part, we propose a DEMA variant scheme, which is motivated by a metric in the stock market, to replace the common exponential moving average scheme. While in the forward-looking part, we present a dynamic lookahead strategy which asymptotically approaches a set value, maintaining its speed at early stage and high convergence performance at final stage. Based on this idea, we provide two optimizer implementations, AdmetaR and AdmetaS, the former based on RAdam and the latter based on SGDM. Through extensive experiments on diverse tasks, we find that the proposed Admeta optimizer outperforms our base optimizers and shows advantages over recently proposed competitive optimizers. We also provide theoretical proof of these two algorithms, which verifies the convergence of our proposed Admeta. Yineng Chen, Zuchao Li, Lefei Zhang, Bo Du 0001, Hai Zhao 0001 |
ICML | 3 |
| 2023 | iRe2f: Rethinking Effective Refinement in Language Structure Prediction via Efficient Iterative Retrospecting and ReasoningabstractRefinement plays a critical role in language structure prediction, a process that deals with complex situations such as structural edge interdependencies. Since language structure prediction usually modeled as graph parsing, typical refinement methods involve taking an initial parsing graph as input and refining it using language input and other relevant information. Intuitively, a refinement component, i.e., refiner, should be lightweight and efficient, as it is only responsible for correcting faults in the initial graph. However, current refiners add a significant burden to the parsing process due to their reliance on time-consuming encoding-decoding procedure on the language input and graph. To make the refiner more practical for real-world applications, this paper proposes a lightweight but effective iterative refinement framework, iRe^2f, based on iterative retrospecting and reasoning without involving the re-encoding process on the graph. iRe^2f iteratively refine the parsing graph based on interaction between graph and sequence and efficiently learns the shortcut to update the sequence and graph representations in each iteration. The shortcut is calculated based on the graph representation in the latest iteration. iRe^2f reduces the number of refinement parameters by 90% compared to the previous smallest refiner. Experiments on a variety of language structure prediction tasks show that iRe^2f performs comparably or better than current state-of-the-art refiners, with a significant increase in efficiency. Zuchao Li, Xingyi Guo, Letian Peng, Lefei Zhang, Hai Zhao 0001 |
IJCAI | 4 |
| 2023 | Cross-Domain Facial Expression Recognition via Disentangling Identity RepresentationabstractMost existing cross-domain facial expression recognition (FER) works require target domain data to assist the model in analyzing distribution shifts to overcome negative effects. However, it is often hard to obtain expression images of the target domain in practical applications. Moreover, existing methods suffer from the interference of identity information, thus limiting the discriminative ability of the expression features. We exploit the idea of domain generalization (DG) and propose a representation disentanglement model to address the above problems. Specifically, we learn three independent potential subspaces corresponding to the domain, expression, and identity information from facial images. Meanwhile, the extracted expression and identity features are recovered as Fourier phase information reconstructed images, thereby ensuring that the high-level semantics of images remain unchanged after disentangling the domain information. Our proposed method can disentangle expression features from expression-irrelevant ones (i.e., identity and domain features). Therefore, the learned expression features exhibit sufficient domain invariance and discriminative ability. We conduct experiments with different settings on multiple benchmark datasets, and the results show that our method achieves superior performance compared with state-of-the-art methods. Tong Liu 0039, Jing Li 0055, Jia Wu 0001, Lefei Zhang, Jun Wan 0005 |
IJCAI | 4 |
| 2023 | Multi-Scale Depth-Aware Unsupervised Domain Adaption in Semantic SegmentationabstractUnsupervised domain adaptation (UDA) for semantic segmentation aims to transfer the domain-invariant knowledge from the labeled source domain to the unlabeled target domain. Leveraging highly relevant tasks as auxiliary tasks has become a common approach to UDA tasks because it contributes to the mutual promotion of the tasks. However, when applying task interactions on a single scale, the model fails to perceive the overall context of the image. To address this issue, we propose a multi-scale depth-aware (Mti-DA) method for domain adaption. In particular, we use the channel attention mechanism to distill the task features and then fuse them with other task features as a complement. The semantic features will better perceive the shape and edges of the objects when they are enhanced by the depth features. We perform task interaction on every scale to deliver the full potential of multi-task learning. Exploiting the depth perception on each scale in the source domain to guide the target domain contributes to enhanced segmentation performance because the complementary relationships of different tasks in the target domain are available. Extensive experiments on two benchmarks (GTA5 to Cityscapes and SYNTHIA to Cityscapes) demonstrate that Mti-DA achieves state-of-the-art performance. Congying Xing, Lefei Zhang |
IJCNN | 2 |
| 2023 | Fused pyramid attention network for single image super-resolutionabstractAbstract In image super‐resolution, deep neural networks with various attention mechanisms have achieved noticeable performance in recent years, for example, channel attention and layer attention. Although many researchers have achieved good super‐resolution results with only a certain style of attention, the divergence and the complementarity focused by multiple attention mechanisms are ignored. In addition, most of these methods fail to utilize the diverse information from multi‐scale features. To efficiently manipulate the above rich information, this paper strives to combine multi‐scale structure and multi‐attention schemes in architecture and module levels for super‐resolution. Especially, in the architecture level, a fused pyramid attention network is developed to extract deep features with the multi‐scale context information from multiple different sizes of receptive field recurrently with skip connections. For the module level, a fused pyramid attention module is designed to fuse the two attention mechanisms to further refine the deep features with fine‐grained information. Compared with the common fusion strategy, the adopted feature fusion structure can maintain better structural information while establishing long‐range dependency. Extensive experimental results demonstrate that the proposed network achieves favorable performance quantitatively and visually. Shi Chen 0010, Xiuping Bi, Lefei Zhang |
IET Image Process. | 3 |
| 2023 | Restoration and enhancement on low exposure raw images by joint demosaicing and denoising
Jiaqi Ma 0002, Guoli Wang 0004, Lefei Zhang, Qian Zhang 0009 |
Neural Networks | 3 |
| 2023 | MSDformer: Multiscale Deformable Transformer for Hyperspectral Image Super-ResolutionabstractDeep learning-based hyperspectral image super-resolution (SR) methods have achieved remarkable success, which can improve the spatial resolution of hyperspectral images with abundant spectral information. However, most of them utilize 2D or 3D convolutions to extract local features while ignoring the rich global spatial-spectral information. In this paper, we propose a novel method called the Multi-Scale Deformable Transformer (MSDformer) for single hyperspectral image super-resolution (SR). The proposed method incorporates the strengths of the convolutional neural network for local spatial-spectral information and the Transformer structure for global spatial-spectral information. Specifically, a multi-scale spectral attention module based on dilated convolution is designed to extract local multi-scale spatial-spectral information, which leverages shared module parameters to exploit the intrinsic spatial redundancy and spectral attention mechanism to accentuate the subtle differences between different spectral groups. Then a deformable convolution-based Transformer module is proposed to further extract the global spatial-spectral information from the local multi-scale features of the previous stage, which can explore the diverse long-range dependencies among all spectral bands. Extensive experiments on three hyperspectral datasets demonstrate that the proposed method achieves excellent SR performance and outperforms the state-of-the-art methods in terms of quantitative quality and visual results. The code is available at https://github.com/Tomchenshi/MSDformer.git. Shi Chen 0010, Lefei Zhang, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Partial Siamese With Multiscale Bi-Codec Networks for Remote Sensing Image Haze RemovalabstractRecently, the U-Shaped networks has been widely explored in remote sensing image dehazing and obtained promising performance. However, most of the existing dehazing methods based on U-Shaped framework lack the reconstruction constraints of haze areas, which is particularly important to restore haze-free images. Moreover, their encoding and decoding layers cannot effectively fuse multi-scale features, resulting in deviations in the color and texture of the dehazing image. To address these issues, in this paper, we propose a Partial Siamese with Multiscale Bi-codec Dehazing Network (PSMB-Net) which is mainly composed of a Partial Siamese Framework (PSF) and a Multiscale Bi-codec Information Fusion (MBIF) module. Specifically, the PSF is proposed to create dehazing prior information to guide the network to build Siamese constraints and achieve improved dehazing results. Furthermore, we design a MBIF module which can enhance feature extraction, and the multi-scale information is used to improve the reconstruction ability of the network for the color and texture of the dehazing image. Experimental results on challenging benchmark datasets demonstrate the superiority of our PSMB-Net over state-of-the-art image dehazing methods. The source code is available at https://github.com/thislzm/PSMB-Net. Zhiming Luo, Bo Du 0001, Wen Yang 0001, Jun Wan 0005, Lefei Zhang |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2023 | Monocular Road Planar Parallax EstimationabstractEstimating the 3D structure of the drivable surface and surrounding environment is a crucial task for assisted and autonomous driving. It is commonly solved either by using 3D sensors such as LiDAR or directly predicting the depth of points via deep learning. However, the former is expensive, and the latter lacks the use of geometry information for the scene. In this paper, instead of following existing methodologies, we propose Road Planar Parallax Attention Network (RPANet), a new deep neural network for 3D sensing from monocular image sequences based on planar parallax, which takes full advantage of the omnipresent road plane geometry in driving scenes. RPANet takes a pair of images aligned by the homography of the road plane as input and outputs a γ map (the ratio of height to depth) for 3D reconstruction. The γ map has the potential to construct a two-dimensional transformation between two consecutive frames. It implies planar parallax and can be combined with the road plane serving as a reference to estimate the 3D structure by warping the consecutive frames. Furthermore, we introduce a novel cross-attention module to make the network better perceive the displacements caused by planar parallax. To verify the effectiveness of our method, we sample data from the Waymo Open Dataset and construct annotations related to planar parallax. Comprehensive experiments are conducted on the sampled dataset to demonstrate the 3D reconstruction accuracy of our approach in challenging scenarios. Haobo Yuan, Wei Sui, Jiafeng Xie, Lefei Zhang, Qian Zhang 0009 |
IEEE Trans. Image Process. | 5 |
| 2022 | Siamese Network with Interactive Transformer for Video Object SegmentationabstractSemi-supervised video object segmentation (VOS) refers to segmenting the target object in remaining frames given its annotation in the first frame, which has been actively studied in recent years. The key challenge lies in finding effective ways to exploit the spatio-temporal context of past frames to help learn discriminative target representation of current frame. In this paper, we propose a novel Siamese network with a specifically designed interactive transformer, called SITVOS, to enable effective context propagation from historical to current frames. Technically, we use the transformer encoder and decoder to handle the past frames and current frame separately, i.e., the encoder encodes robust spatio-temporal context of target object from the past frames, while the decoder takes the feature embedding of current frame as the query to retrieve the target from the encoder output. To further enhance the target representation, a feature interaction module (FIM) is devised to promote the information flow between the encoder and decoder. Moreover, we employ the Siamese architecture to extract backbone features of both past and current frames, which enables feature reuse and is more efficient than existing methods. Experimental results on three challenging benchmarks validate the superiority of SITVOS over state-of-the-art methods. Code is available at https://github.com/LANMNG/SITVOS. Meng Lan, Jing Zhang 0037, Fengxiang He, Lefei Zhang |
AAAI | 4 |
| 2022 | PolyphonicFormer: Unified Query Learning for Depth-Aware Video Panoptic Segmentation
Haobo Yuan, Xiangtai Li, Jing Zhang 0037, Yunhai Tong, Lefei Zhang, Dacheng Tao |
ECCV (27) | 7 |
| 2022 | ELMformer: Efficient Raw Image Restoration with a Locally Multiplicative TransformerabstractIn order to get raw images of high quality for downstream Image Signal Process (ISP), in this paper we present an Efficient Locally Multiplicative Transformer called ELMformer for raw image restoration. ELMformer contains two core designs especially for raw images whose primitive attribute is single-channel. The first design is a Bi-directional Fusion Projection (BFP) module, where we consider both the color characteristics of raw images and spatial structure of single-channel. The second one is that we propose a Locally Multiplicative Self-Attention (L-MSA) scheme to effectively deliver information from the local space to relevant parts. ELMformer can efficiently reduce the computational consumption and perform well on raw image restoration tasks. Enhanced by these two core designs, ELMformer achieves the highest performance and keeps the lowest FLOPs on raw denoising and raw deblurring benchmarks compared with state-of-the-arts. Extensive experiments demonstrate the superiority and generalization ability of ELMformer. On SIDD benchmark, our method has even better denoising performance than ISP-based methods which need huge amount of additional sRGB training images. Jiaqi Ma 0002, Shengyuan Yan, Lefei Zhang, Guoli Wang 0004, Qian Zhang 0009 |
ACM Multimedia | 3 |
| 2022 | Atrous Pyramid Transformer with Spectral Convolution for Image InpaintingabstractOwing to the ability of extracting features of images on long-range dependencies naturally, transformer is possible to reconstruct the damaged areas of images with the information from the uncorrupted regions globally. In this paper, we propose a two-stage framework based on a novel atrous pyramid transformer (APT) for image inpainting that recovers the structure and texture of an image progressively. Specifically, the patches of APT blocks are embedded in an atrous pyramid manner to explicitly enhance the correlation for both inter-and intra-windows to restore the high-level semantic structures of images more precisely, which could be served as a guide map for the second phase. Subsequently, a dual spectral transform convolution (DSTC) module is further designed to work together with APT to infer the low-level features of the generated areas. The DSTC module decouples the image signal into high frequency and low frequency for capturing texture information with a global view. Experiments on the CelebA-HQ, Paris StreetView, and Places2 demonstrate the superiority of the proposed approach. Muqi Huang, Lefei Zhang |
ACM Multimedia | 2 |
| 2022 | GT-MUST: Gated Try-on by Learning the Mannequin-Specific TransformationabstractGiven the mannequin (i.e., reference person) and target garment, the virtual try-on (VTON) task aims at dressing the mannequin in the provided garment automatically, having attracted increasing attention in recent years. Previous works usually conduct the garment deformation under the guidance of ''shape''. However, ''shape-only transformation'' ignores the local structures and results in unnatural distortions. To address this issue, we propose a Gated Try-on method by learning the ManneqUin-Specific Transformation (GT-MUST). Technically, we implement GT-MUST as a three-stage deep neural model. First, GT-MUST learns the ''mannequin-specific transformation'' with a ''take-off'' mechanism, which recovers the warped clothes of the mannequin to its original in-shop state. Then, the learned ''mannequin-specific transformation'' is inverted and utilized to help generate the mannequin-specific warped state for a target garment. Finally, a special gate is employed to better combine the mannequin-specific warped garment with the mannequin. GT-MUST benefits from learning to solve a much easier ''take-off'' task to obtain the mannequin-specific information than the common ''try-on'' task, since flat in-shop garments usually have less variation in shape than those clothed on the body. Experiments on the fashion dataset demonstrate that GT-MUST outperforms the state-of-the-art virtual try-on methods. The code is available at https://github.com/wangning-001/GT-MUST. Ning Wang 0031, Jing Zhang 0037, Lefei Zhang, Dacheng Tao |
ACM Multimedia | 3 |
| 2022 | LocalDrop: A Hybrid Regularization for Deep Neural NetworksabstractIn neural networks, developing regularization algorithms to settle overfitting is one of the major study areas. We propose a new approach for the regularization of neural networks by the local Rademacher complexity called LocalDrop. A new regularization function for both fully-connected networks (FCNs) and convolutional neural networks (CNNs), including drop rates and weight matrices, has been developed based on the proposed upper bound of the local Rademacher complexity by the strict mathematical deduction. The analyses of dropout in FCNs and DropBlock in CNNs with keep rate matrices in different layers are also included in the complexity analyses. With the new regularization function, we establish a two-stage procedure to obtain the optimal keep rate matrix and weight matrix to realize the whole training model. Extensive experiments have been conducted to demonstrate the effectiveness of LocalDrop in different models by comparing it with several algorithms and the effects of different hyperparameters on the final performances. Ziqing Lu, Chang Xu 0002, Bo Du 0001, Takashi Ishida 0001, Lefei Zhang, Masashi Sugiyama |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Stagewise Unsupervised Domain Adaptation With Adversarial Self-Training for Road Segmentation of Remote-Sensing ImagesabstractRoad segmentation from remote-sensing images is a challenging task with wide ranges of application potentials. Deep neural networks have advanced this field by leveraging the power of large-scale labeled data, which, however, are extremely expensive and time-consuming to acquire. One solution is to use cheap available data to train a model and deploy it to directly process the data from a specific application domain. Nevertheless, the well-known domain shift (DS) issue prevents the trained model from generalizing well on the target domain. In this article, we propose a novel stagewise domain adaptation model called RoadDA to address the DS issue in this field. In the first stage, RoadDA adapts the target domain features to align with the source ones via generative adversarial networks (GANs)-based interdomain adaptation. Specifically, a feature pyramid fusion module is devised to avoid information loss of long and thin roads and learn discriminative and robust features. Besides, to address the intradomain discrepancy in the target domain, in the second stage, we propose an adversarial self-training method. We generate the pseudo labels of the target domain using the trained generator and divide it to labeled easy split and unlabeled hard split based on the road confidence scores. The features of hard split are adapted to align with the easy ones using adversarial learning and the intradomain adaptation process is repeated to progressively improve the segmentation performance. Experiment results on two benchmarks demonstrate that RoadDA can efficiently reduce the domain gap and outperforms state-of-the-art methods. The code is available athttps://github.com/LANMNG/RoadDA. Lefei Zhang, Meng Lan, Jing Zhang 0037, Dacheng Tao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Addressing Domain Gap via Content Invariant Representation for Semantic SegmentationabstractThe problem of unsupervised domain adaptation in semantic segmentation is a major challenge for numerous computer vision tasks because acquiring pixel-level labels is time-consuming with expensive human labor. A large gap exists among data distributions in different domains, which will cause severe performance loss when a model trained with synthetic data is generalized to real data. Hence, we propose a novel domain adaptation approach, called Content Invariant Representation Network, to narrow the domain gap between the source (S) and target (T) domains. The previous works developed a network to directly transfer the knowledge from the S to T. On the contrary, the proposed method aims to progressively reduce the gap between S and T on the basis of a Content Invariant Representation (CIR). CIR is an intermediate domain (I) sharing invariant content with S and having similar data distribution to T. Then, an Ancillary Classifier Module (ACM) is designed to focus on pixel-level details and generate attention-aware results. ACM adaptively assigns different weights to pixels according to their domain offsets, thereby reducing local domain gaps. The global domain gap between CIR and T is also narrowed by enforcing local alignments. Last, we perform self-supervised training in the pseudo-labeled target domain to further fit the distribution of the real data. Comprehensive experiments on two domain adaptation tasks, that is, GTAV → Cityscapes and SYNTHIA → Cityscapes, clearly demonstrate the superiority of our method compared with state-of-the-art methods. Lefei Zhang, Qian Zhang 0009 |
AAAI | 2 |
| 2021 | DSP: Dual Soft-Paste for Unsupervised Domain Adaptive Semantic SegmentationabstractUnsupervised domain adaptation (UDA) for semantic segmentation aims to adapt a segmentation model trained on the labeled source domain to the unlabeled target domain. Existing methods try to learn domain invariant features while suffering from large domain gaps that make it difficult to correctly align discrepant features, especially in the initial training phase. To address this issue, we propose a novel Dual Soft-Paste (DSP) method in this paper. Specifically, DSP selects some classes from a source domain image using a long-tail class first sampling strategy and softly pastes the corresponding image patch on both the source and target training images with a fusion weight. Technically, we adopt the mean teacher framework for domain adaptation, where the pasted source and target images go through the student network while the original target image goes through the teacher network. Output-level alignment is carried out by aligning the probability maps of the target fused image from both networks using a weighted cross-entropy loss. In addition, feature-level alignment is carried out by aligning the feature maps of the source and target images from student network using a weighted maximum mean discrepancy loss. DSP facilitates the model learning domain-invariant features from the intermediate domains, leading to faster convergence and better performance. Experiments on two challenging benchmarks demonstrate the superiority of DSP over state-of-the-art methods. Code is available at https://github.com/GaoLii/DSP. Jing Zhang 0037, Lefei Zhang, Dacheng Tao |
ACM Multimedia | 3 |
| 2021 | Discriminative subspace matrix factorization for multiview data clustering
Jiaqi Ma 0002, Yipeng Zhang 0001, Lefei Zhang |
Pattern Recognit. | 3 |
| 2021 | Nonlocal Low-Rank Tensor Completion for Visual DataabstractIn this paper, we propose a novel nonlocal patch tensor-based visual data completion algorithm and analyze its potential problems. Our algorithm consists of two steps: the first step is initializing the image with triangulation-based linear interpolation and the second step is grouping similar nonlocal patches as a tensor then applying the proposed tensor completion technique. Specifically, with treating a group of patch matrices as a tensor, we impose the low-rank constraint on the tensor through the recently proposed tensor nuclear norm. Moreover, we observe that after the first interpolation step, the image gets blurred and, thus, the similar patches we have found may not exactly match the reference. We name the problem "Patch Mismatch," and then in order to avoid the error caused by it, we further decompose the patch tensor into a low-rank tensor and a sparse tensor, which means the accepted horizontal strips in mismatched patches. Furthermore, our theoretical analysis shows that the error caused by Patch Mismatch can be decomposed into two components, one of which can be bounded by a reasonable assumption named local patch similarity, and the other part is lower than that using matrix completion. Extensive experimental results on real-world datasets verify our method's superiority to the state-of-the-art tensor-based image inpainting methods. Lefei Zhang, Liangchen Song, Bo Du 0001, Yipeng Zhang 0001 |
IEEE Trans. Cybern. | 1 |
| 2021 | Dynamic Selection Network for Image InpaintingabstractImage inpainting is a challenging computer vision task that aims to fill in missing regions of corrupted images with realistic contents. With the development of convolutional neural networks, many deep learning models have been proposed to solve image inpainting issues by learning information from a large amount of data. In particular, existing algorithms usually follow an encoding and decoding network architecture in which some operations with standard schemes are employed, such as static convolution, which only considers pixels with fixed grids, and the monotonous normalization style (e.g., batch normalization). However, these techniques are not well-suited for the image inpainting task because the random corrupted regions in the input images tend to mislead the inpainting process and generate unreasonable content. In this paper, we propose a novel dynamic selection network (DSNet) to solve this problem in image inpainting tasks. The principal idea of the proposed DSNet is to distinguish the corrupted region from the valid ones throughout the entire network architecture, which may help make full use of the information in the known area. Specifically, the proposed DSNet has two novel dynamic selection modules, namely, the validness migratable convolution (VMC) and regional composite normalization (RCN) modules, which share a dynamic selection mechanism that helps utilize valid pixels better. By replacing vanilla convolution with the VMC module, spatial sampling locations are dynamically selected in the convolution phase, resulting in a more flexible feature extraction process. Besides, the RCN module not only combines several normalization methods but also normalizes the feature regions selectively. Therefore, the proposed DSNet can illustrate realistic and fine-detailed images by adaptively selecting features and normalization styles. Experimental results on three public datasets show that our proposed method outperforms state-of-the-art methods both quantitatively and qualitatively. Ning Wang 0031, Yipeng Zhang 0001, Lefei Zhang |
IEEE Trans. Image Process. | 3 |
| 2021 | Incorporating Distribution Matching into Uncertainty for Multiple Kernel Active LearningabstractDue to the lack of the labeled data and the complex structures of various data, it is very hard to learn the uncertainty and representativeness accurately in active learning. In this paper, we propose a multiple kernel active learning framework that incorporates a group regularizer of distribution information into the estimation of uncertainty. The proposed method takes the advantage of multiple kernel learning to learn the kernel space in which the complex structures can be well captured by kernel weights. Meanwhile, we have developed an efficient optimization algorithm to solve the proposed method. Experimental results on twelve UCI benchmark data sets and eight subsets of ImageNet show that the proposed method outperforms several state-of-the-art active learning methods. Moreover, we also have applied the proposed method to multiple feature scenario on Caltech101, and the promising results are also obtained compared with single feature scenario. Zengmao Wang, Bo Du 0001, Weiping Tu, Lefei Zhang, Dacheng Tao |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2020 | Recurrent Feature Reasoning for Image InpaintingabstractExisting inpainting methods have achieved promising performance for recovering regular or small image defects. However, filling in large continuous holes remains difficult due to the lack of constraints for the hole center. In this paper, we devise a Recurrent Feature Reasoning (RFR) network which is mainly constructed by a plug-and-play Recurrent Feature Reasoning module and a Knowledge Consistent Attention (KCA) module. Analogous to how humans solve puzzles (i.e., first solve the easier parts and then use the results as additional information to solve difficult parts), the RFR module recurrently infers the hole boundaries of the convolutional feature maps and then uses them as clues for further inference. The module progressively strengthens the constraints for the hole center and the results become explicit. To capture information from distant places in the feature map for RFR, we further develop KCA and incorporate it in RFR. Empirically, we first compare the proposed RFR-Net with existing backbones, demonstrating that RFR-Net is more efficient (e.g., a 4% SSIM improvement for the same model size). We then place the network in the context of the current state-of-the-art, where it exhibits improved performance. The corresponding source code is available at: https://github.com/jingyuanli001/RFR-Inpainting. Ning Wang 0031, Lefei Zhang, Bo Du 0001, Dacheng Tao |
CVPR | 3 |
| 2020 | E3SN: Efficient End-to-End Siamese Network for Video Object SegmentationabstractIn the semi-supervised video object segmentation (VOS) field, SiamMask has achieved competitive accuracy and the fastest running speed. However, the two-stage training procedure requires additional manual intervention, and using only single-level features does not maximize the rich hierarchical feature information. This paper proposes an efficient end-to-end Siamese network for VOS. In particular, a supervised sampling strategy is designed to optimize the training procedure. Such an optimization facilitates the training of the entire model in an end-to-end manner. Moreover, a multilevel feature aggregation module is developed to enhance feature representability and improve segmentation accuracy. Experimental results on DAVIS2016 and DAVIS2017 datasets show that the proposed approach outperforms the SiamMask in accuracy with similar FPS. Moreover, this approach also achieves good accuracy-speed trade-off compared with that of other state-of-the-art VOS algorithms. Meng Lan, Yipeng Zhang 0001, Qinning Xu, Lefei Zhang |
IJCAI | 4 |
| 2020 | Parallel DNN Inference Framework Leveraging a Compact RISC-V ISA-based Multi-core SystemabstractRISC-V is an open-source instruction set and now has been examined as a universal standard to unify the heterogeneous platforms. However, current research focuses primarily on the design and fabrication of general-purpose processors based on RISC-V, despite the fact that in the era of IoT (Internet of Things), the fusion of heterogeneous platforms should also take application-specific processors into account. Accordingly, this paper proposes a collaborative RISC-V multi-core system for Deep Neural Network (DNN) accelerators. To the best of our knowledge, this is the first time that a multi-core scheduling architecture for DNN acceleration is formulated and RISC-V is explored as the ISA of a multi-core system to bridge the gap between the memory and the DNN Processor in order to increase the entire system throughput. The experiment realizes a four-stage design of the RISC-V core, and further reveals that a multi-core design along with an appropriate scheduling algorithm can efficiently decrease the runtime and elevate the throughput. Moreover, the experiment also provides us with a constructive suggestion regarding the ideal proportion of the cores to Process Engines (PE), which provides us with significant assistance in building highly efficient AI System-on-Chips (SoCs) in resource-aware situations. Yipeng Zhang 0001, Bo Du 0001, Lefei Zhang, Jia Wu 0001 |
KDD | 3 |
| 2020 | Robust learning with imperfect privileged information
Bo Du 0001, Chang Xu 0002, Yipeng Zhang 0001, Lefei Zhang, Dacheng Tao |
Artif. Intell. | 5 |
| 2020 | Learning selection channels for image steganalysis in spatial domain
Weixiang Ren, Liming Zhai, Ju Jia, Lina Wang 0001, Lefei Zhang |
Neurocomputing | 5 |
| 2020 | Global context based automatic road segmentation via dilated convolutional neural network
Meng Lan, Yipeng Zhang 0001, Lefei Zhang, Bo Du 0001 |
Inf. Sci. | 3 |
| 2020 | Robust face alignment by cascaded regression and de-occlusion
Jun Wan 0005, Jing Li 0055, Zhihui Lai 0001, Bo Du 0001, Lefei Zhang |
Neural Networks | 5 |
| 2020 | Transferable heterogeneous feature subspace learning for JPEG mismatched steganalysis
Ju Jia, Liming Zhai, Weixiang Ren, Lina Wang 0001, Yanzhen Ren, Lefei Zhang |
Pattern Recognit. | 6 |
| 2020 | Unsupervised domain adaptive re-identification: Theory and practice
Liangchen Song, Cheng Wang 0048, Lefei Zhang, Bo Du 0001, Qian Zhang 0009, Chang Huang, Xinggang Wang |
Pattern Recognit. | 3 |
| 2020 | Multistage attention network for image inpainting
Ning Wang 0031, Sihan Ma, Yipeng Zhang 0001, Lefei Zhang |
Pattern Recognit. | 5 |
| 2020 | Data-augmented matched subspace detector for hyperspectral subpixel target detection
Mingzhi Dong, Ziyu Wang 0003, Lianru Gao, Lefei Zhang, Jing-Hao Xue |
Pattern Recognit. | 5 |
| 2020 | Dimensionality Reduction With Enhanced Hybrid-Graph Discriminant Learning for Hyperspectral Image ClassificationabstractDimensionality reduction (DR) is an important way of improving the classification accuracy of a hyperspectral image (HSI). Graph learning, which can effectively reveal the intrinsic relationships of data, has been widely used in the case of HSIs. However, most of them are based on a simple graph to represent the binary relationships of data. An HSI contains complex high-order relationships among different samples. Therefore, in this article, we propose a hybrid-graph learning method to reveal the complex high-order relationships of the HSI, termed enhanced hybrid-graph discriminant learning (EHGDL). In EHGDL, an intraclass hypergraph and an interclass hypergraph are constructed to analyze the complex multiple relationships of a HSI. Then, a supervised locality graph is applied to reveal the binary relationships of a HSI which can form the complementarity of a hypergraph. Simultaneously, we also construct a weighted neighborhood margin model to boost the difference of samples from different classes. Finally, we design a DR model based on the intraclass hypergraph, the interclass hypergraph, the supervised locality graph, and the weighted neighborhood margin to improve the compactness of the intraclass samples and the separability of the interclass samples, and an optimal projection matrix can be achieved to extract the low-dimensional embedding features of the HSI. To demonstrate the effectiveness of the proposed method, experiments have been conducted on the Indian Pines, PaviaU, and HoustonU data sets. The experimental results show that EHGDL can generate better classification performance compared with some related DR methods. As a result, EHGDL can better reveal the complex intrinsic relationships of a HSI by the complementarity of different characteristics and enhance the discriminant performance of land-cover types. Fulin Luo, Liangpei Zhang 0001, Bo Du 0001, Lefei Zhang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2020 | Multiscale Dynamic Graph Convolutional Network for Hyperspectral Image ClassificationabstractConvolutional neural network (CNN) has demonstrated impressive ability to represent hyperspectral images and to achieve promising results in hyperspectral image classification. However, traditional CNN models can only operate convolution on regular square image regions with fixed size and weights, and thus, they cannot universally adapt to the distinct local regions with various object distributions and geometric appearances. Therefore, their classification performances are still to be improved, especially in class boundaries. To alleviate this shortcoming, we consider employing the recently proposed graph convolutional network (GCN) for hyperspectral image classification, as it can conduct the convolution on arbitrarily structured non-Euclidean data and is applicable to the irregular image regions represented by graph topological information. Different from the commonly used GCN models that work on a fixed graph, we enable the graph to be dynamically updated along with the graph convolution process so that these two steps can be benefited from each other to gradually produce the discriminative embedded features as well as a refined graph. Moreover, to comprehensively deploy the multiscale information inherited by hyperspectral images, we establish multiple input graphs with different neighborhood scales to extensively exploit the diversified spectral-spatial correlations at multiple scales. Therefore, our method is termed multiscale dynamic GCN (MDGCN). The experimental results on three typical benchmark data sets firmly demonstrate the superiority of the proposed MDGCN to other state-of-the-art methods in both qualitative and quantitative aspects. Sheng Wan, Chen Gong 0002, Ping Zhong 0001, Bo Du 0001, Lefei Zhang, Jian Yang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2020 | Homologous Component Analysis for Domain AdaptationabstractCovariate shift assumption based domain adaptation approaches usually utilize only one common transformation to align marginal distributions and make conditional distributions preserved. However, one common transformation may cause loss of useful information, such as variances and neighborhood relationship in both source and target domain. To address this problem, we propose a novel method called homologous component analysis (HCA) where we try to find two totally different but homologous transformations to align distributions with side information and make conditional distributions preserved. As it is hard to find a closed form solution to the corresponding optimization problem, we solve them by means of the alternating direction minimizing method (ADMM) in the context of Stiefel manifolds. We also provide a generalization error bound for domain adaptation in semi-supervised case and two transformations can help to decrease this upper bound more than only one common transformation does. Extensive experiments on synthetic and real data show the effectiveness of the proposed method by comparing its classification accuracy with the state-of-the-art methods and numerical evidence on chordal distance and Frobenius distance shows that resulting optimal transformations are different. Youfa Liu, Weiping Tu, Bo Du 0001, Lefei Zhang, Dacheng Tao |
IEEE Trans. Image Process. | 4 |
| 2020 | Learning From Synthetic Images via Active Pseudo-LabelingabstractSynthetic visual data refers to the data automatically rendered by the mature computer graphic algorithms. With the rapid development of these techniques, we can now collect photo-realistic synthetic images with accurate pixel-level annotations without much effort. However, due to the domain gaps between synthetic data and real data, in terms of not only visual appearance but also label distribution, directly applying models trained on synthetic images to real ones can hardly yield satisfactory performance. Since the collection of accurate labels for real images is very laborious and time-consuming, developing algorithms which can learn from synthetic images is of great significance. In this paper, we propose a novel framework, namely Active Pseudo-Labeling (APL), to reduce the domain gaps between synthetic images and real images. In APL framework, we first predict pseudo-labels for the unlabeled real images in the target domain by actively adapting the style of the real images to source domain. Specifically, the style of real images is adjusted via a novel task guided generative model, and then pseudo-labels are predicted for these actively adapted images. Lastly, we fine-tune the source-trained model in the pseudo-labeled target domain, which helps to fit the distribution of the real data. Experiments on both semantic segmentation and object detection tasks with several challenging benchmark data sets demonstrate the priority of our proposed method compared to the existing state-of-the-art approaches. Liangchen Song, Yonghao Xu, Lefei Zhang, Bo Du 0001, Qian Zhang 0009, Xinggang Wang |
IEEE Trans. Image Process. | 3 |
| 2019 | Self-Ensembling Attention Networks: Addressing Domain Shift for Semantic SegmentationabstractRecent years have witnessed the great success of deep learning models in semantic segmentation. Nevertheless, these models may not generalize well to unseen image domains due to the phenomenon of domain shift. Since pixel-level annotations are laborious to collect, developing algorithms which can adapt labeled data from source domain to target domain is of great significance. To this end, we propose self-ensembling attention networks to reduce the domain gap between different datasets. To the best of our knowledge, the proposed method is the first attempt to introduce selfensembling model to domain adaptation for semantic segmentation, which provides a different view on how to learn domain-invariant features. Besides, since different regions in the image usually correspond to different levels of domain gap, we introduce the attention mechanism into the proposed framework to generate attention-aware features, which are further utilized to guide the calculation of consistency loss in the target domain. Experiments on two benchmark datasets demonstrate that the proposed framework can yield competitive performance compared with the state of the art methods. Yonghao Xu, Bo Du 0001, Lefei Zhang, Qian Zhang 0009, Guoli Wang 0004, Liangpei Zhang 0001 |
AAAI | 3 |
| 2019 | Leveraging Ratings and Reviews with Gating Mechanism for RecommendationabstractRecommender system plays an important role to provide people with personalized information based on their history records. However, it is still a challenge to capture the preference of users accurately due to the sparsity of rating data and the heterogeneity of review data. In this paper, we propose a hybrid deep collaborative filtering model that jointly learns latent representations from ratings and reviews. Specifically, the model learns the rating feature and textual feature based on ratings and reviews simultaneously. Two embedding layers are employed to learn rating feature for users and items based on the user and item interactions, and two attention-based GRU networks learn context-aware representation from user and item reviews. Then a gating mechanism is used to leverage contributions from rating feature and textual feature. Experimental results on six real-world datasets demonstrate the superior performance of the proposed method over several state-of-the-art methods. Moreover, the keywords in reviews can be highlighted to interpret the predictions with the attention mechanism. Haifeng Xia, Zengmao Wang, Bo Du 0001, Lefei Zhang, Gang Chun |
CIKM | 4 |
| 2019 | Fast Spatio-Temporal Residual Network for Video Super-ResolutionabstractRecently, deep learning based video super-resolution (SR) methods have achieved promising performance. To simultaneously exploit the spatial and temporal information of videos, employing 3-dimensional (3D) convolutions is a natural approach. However, straight utilizing 3D convolutions may lead to an excessively high computational complexity which restricts the depth of video SR models and thus undermine the performance. In this paper, we present a novel fast spatio-temporal residual network (FSTRN) to adopt 3D convolutions for the video SR task in order to enhance the performance while maintaining a low computational load. Specifically, we propose a fast spatio-temporal residual block (FRB) that divide each 3D filter to the product of two 3D filters, which have considerably lower dimensions. Furthermore, we design a cross-space residual learning that directly links the low-resolution space and the high-resolution space, which can greatly relieve the computational burden on the feature fusion and up-scaling parts. Extensive evaluations and comparisons on benchmark datasets validate the strengths of the proposed approach and demonstrate that the proposed network significantly outperforms the current state-of-the-art methods. Sheng Li 0001, Fengxiang He, Bo Du 0001, Lefei Zhang, Yonghao Xu, Dacheng Tao |
CVPR | 4 |
| 2019 | Progressive Reconstruction of Visual Structure for Image InpaintingabstractInpainting methods aim to restore missing parts of corrupted images and play a critical role in many computer vision applications, such as object removal and image restoration. Although existing methods perform well on images with small holes, restoring large holes remains elusive. To address this issue, this paper proposes a Progressive Reconstruction of Visual Structure (PRVS) network that progressively reconstructs the structures and the associated visual feature. Specifically, we design a novel Visual Structure Reconstruction (VSR) layer to entangle reconstructions of the visual structure and visual feature, which benefits each other by sharing parameters. We repeatedly stack four VSR layers in both encoding and decoding stages of a U-Net like architecture to form the generator of a generative adversarial network (GAN) for restoring images with either small or large holes. We prove the generalization error upper bound of the PRVS network is O(1\sqrt(N)), which theoretically guarantees its performance. Extensive empirical evaluations and comparisons on Places2, Paris Street View and CelebA datasets validate the strengths of the proposed approach and demonstrate that the model outperforms current state-of-the-art methods. The source code package is available at https://github.com/jingyuanli001/PRVS-Image-Inpainting. Fengxiang He, Lefei Zhang, Bo Du 0001, Dacheng Tao |
ICCV | 3 |
| 2019 | Self-Paced Subspace ClusteringabstractSubspace clustering aims to segment data sampled from a union of subspaces in visual data tasks. Structured Sparse Subspace Clustering (SSSC) model is a unified optimization framework, which proves successful in learning both the self representation of the data and their subspace segmentation. However, SSSC involves solving non-convex subproblems and hence it may be stuck into bad local minima such that clustering performance degrades. In this paper, we propose a self-paced subspace clustering algorithm to tackle this problem, which learns subspace segmentation of data by progressing from 'easy' to 'complex' examples under a novel self-paced regularizer. Experiments on the real-world human face datasets verify the effectiveness of the proposed algorithm. Youfa Liu, Bo Du 0001, Lefei Zhang |
ICME | 3 |
| 2019 | Pseudo Supervised Matrix Factorization in Discriminative SubspaceabstractNon-negative Matrix Factorization (NMF) and spectral clustering have been proved to be efficient and effective for data clustering tasks and have been applied to various real-world scenes. However, there are still some drawbacks in traditional methods: (1) most existing algorithms only consider high-dimensional data directly while neglect the intrinsic data structure in the low-dimensional subspace; (2) the pseudo-information got in the optimization process is not relevant to most spectral clustering and manifold regularization methods. In this paper, a novel unsupervised matrix factorization method, Pseudo Supervised Matrix Factorization (PSMF), is proposed for data clustering. The main contributions are threefold: (1) to cluster in the discriminant subspace, Linear Discriminant Analysis (LDA) combines with NMF to become a unified framework; (2) we propose a pseudo supervised manifold regularization term which utilizes the pseudo-information to instruct the regularization term in order to find subspace that discriminates different classes; (3) an efficient optimization algorithm is designed to solve the proposed problem with proved convergence. Extensive experiments on multiple benchmark datasets illustrate that the proposed model outperforms other state-of-the-art clustering algorithms. Jiaqi Ma 0002, Yipeng Zhang 0001, Lefei Zhang, Bo Du 0001, Dapeng Tao |
IJCAI | 3 |
| 2019 | MUSICAL: Multi-Scale Image Contextual Attention Learning for InpaintingabstractWe study the task of image inpainting, where an image with missing region is recovered with plausible context. Recent approaches based on deep neural networks have exhibited potential for producing elegant detail and are able to take advantage of background information, which gives texture information about missing region in the image. These methods often perform pixel/patch level replacement on the deep feature maps of missing region and therefore enable the generated content to have similar texture as background region. However, this kind of replacement is a local strategy and often performs poorly when the background information is misleading. To this end, in this study, we propose to use a multi-scale image contextual attention learning (MUSICAL) strategy that helps to flexibly handle richer background information while avoid to misuse of it. However, such strategy may not promising in generating context of reasonable style. To address this issue, both of the style loss and the perceptual loss are introduced into the proposed method to achieve the style consistency of the generated image. Furthermore, we have also noticed that replacing some of the down sampling layers in the baseline network with the stride 1 dilated convolution layers is beneficial for producing sharper and fine-detailed results. Experiments on the Paris Street View, Places, and CelebA datasets indicate the superior performance of our approach compares to the state-of-the-arts. Ning Wang 0031, Lefei Zhang, Bo Du 0001 |
IJCAI | 3 |
| 2019 | Accelerated Inference Framework of Sparse Neural Network Based on Nested Bitmask StructureabstractIn order to satisfy the ever-growing demand for high-performance processors for neural networks, the state-of-the-art processing units tend to use application-oriented circuits to replace Processing Engine (PE) on the GPU under circumstances where low-power solutions are required. The application-oriented PE is fully optimized in terms of the circuit architecture and eliminates incorrect data dependency and instructional redundancy. In this paper, we propose a novel encoding approach on a sparse neural network after pruning. We partition the weight matrix into numerous blocks and use a low-rank binary map to represent the validation of these blocks. Furthermore, the elements in each nonzero block are also encoded into two submatrices: one is the binary stream discriminating the zero/nonzero position, while the other is the pure nonzero elements stored in the FIFO. In the experimental part, we implement a well pre-trained sparse neural network on the Xilinx FPGA VC707. Experimental results show that our algorithm outperforms the other benchmarks. Our approach has successfully optimized the throughput and the energy efficiency to deal with a single frame. Accordingly, we contend that Nested Bitmask Neural Network (NBNN), is an efficient neural network structure with only minor accuracy loss on the SoC system. Yipeng Zhang 0001, Bo Du 0001, Lefei Zhang, Rongchun Li, Yong Dou |
IJCAI | 3 |
| 2019 | On combining active and transfer learning for medical data classificationabstractThis study presents a novel algorithm which combines active learning (AL) and transfer learning for medical data classification. The main idea of the proposed algorithm is iteratively querying a small number of informative unlabelled target samples, and, at the same time, removing the source samples which do not fit with the posterior probability distributions in the target domain, so as to combine the basic idea of AL with transfer learning. The experimental results obtained in the classification of the datasets from the University of California Irvine (UCI) Machine Learning Repository and The Cancer Imaging Archive (TCIA) confirm the effectiveness of the proposed algorithm. Bo Du 0001, Zengmao Wang, Lefei Zhang |
IET Comput. Vis. | 5 |
| 2019 | Hyperspectral image unsupervised classification by robust manifold matrix factorization
Lefei Zhang, Liangpei Zhang 0001, Bo Du 0001, Jane You, Dacheng Tao |
Inf. Sci. | 1 |
| 2019 | A robust dimensionality reduction and matrix factorization framework for data clustering
Lefei Zhang, Bo Du 0001 |
Pattern Recognit. Lett. | 2 |
| 2019 | MSDH: Matched subspace detector with heterogeneous noise
Lefei Zhang, Lianru Gao, Jing-Hao Xue |
Pattern Recognit. Lett. | 2 |
| 2019 | Robust Graph-Based Semisupervised Learning for Noisy Labeled Data via Maximum Correntropy CriterionabstractSemisupervised learning (SSL) methods have been proved to be effective at solving the labeled samples shortage problem by using a large number of unlabeled samples together with a small number of labeled samples. However, many traditional SSL methods may not be robust with too much labeling noisy data. To address this issue, in this paper, we propose a robust graph-based SSL method based on maximum correntropy criterion to learn a robust and strong generalization model. In detail, the graph-based SSL framework is improved by imposing supervised information on the regularizer, which can strengthen the constraint on labels, thus ensuring that the predicted labels of each cluster are close to the true labels. Furthermore, the maximum correntropy criterion is introduced into the graph-based SSL framework to suppress labeling noise. Extensive image classification experiments prove the generalization and robustness of the proposed SSL method. Bo Du 0001, Zengmao Wang, Lefei Zhang, Dacheng Tao |
IEEE Trans. Cybern. | 4 |
| 2019 | Feature Learning Using Spatial-Spectral Hypergraph Discriminant Analysis for Hyperspectral ImageabstractHyperspectral image (HSI) contains a large number of spatial-spectral information, which will make the traditional classification methods face an enormous challenge to discriminate the types of land-cover. Feature learning is very effective to improve the classification performances. However, the current feature learning approaches are mostly based on a simple intrinsic structure. To represent the complex intrinsic spatial-spectral of HSI, a novel feature learning algorithm, termed spatial-spectral hypergraph discriminant analysis (SSHGDA), has been proposed on the basis of spatial-spectral information, discriminant information, and hypergraph learning. SSHGDA constructs a reconstruction between-class scatter matrix, a weighted within-class scatter matrix, an intraclass spatial-spectral hypergraph, and an interclass spatial-spectral hypergraph to represent the intrinsic properties of HSI. Then, in low-dimensional space, a feature learning model is designed to compact the intraclass information and separate the interclass information. With this model, an optimal projection matrix can be obtained to extract the spatial-spectral features of HSI. SSHGDA can effectively reveal the complex spatial-spectral structures of HSI and enhance the discriminating power of features for land-cover classification. Experimental results on the Indian Pines and PaviaU HSI data sets show that SSHGDA can achieve better classification accuracies in comparison with some state-of-the-art methods. Fulin Luo, Bo Du 0001, Liangpei Zhang 0001, Lefei Zhang, Dacheng Tao |
IEEE Trans. Cybern. | 4 |
| 2019 | On Gleaning Knowledge From Cross Domains by Sparse Subspace Correlation Analysis for Hyperspectral Image ClassificationabstractDespite the availability of an increasing amount of remote sensing images, problems still arise in that the knowledge from existing images is underutilized and the collection of reference knowledge for each newly obtained image is expensive. Recently, an attractive solution called “transfer learning” has received increasing attention in the remote sensing field, by transferring knowledge from source domains to help improve the learning procedure in the target domain. In this paper, we propose a sparse subspace correlation analysis-based supervised classification (SSCA-SC) method for transfer learning in hyperspectral remote sensing image classification, which is not restricted by the data dimensionality or the data acquisition sensors. Specifically, we first propose a sparse subspace correlation analysis (SSCA) method to simultaneously learn the optimal projection matrices for heterogeneous domains into a common subspace and obtain sparse reconstruction coefficients over a shared self-expressive dictionary in the derived subspace. In order to fully utilize the label information to improve the class separability, the SSCA-SC framework learns more discriminative representations for the input data by training a corresponding SSCA model for each class. As a result, the projected data belonging to the same class are maximally correlated and represented well, while those from different classes will have a low correlation. Another advantage of the SSCA-SC framework lies in the fact that it not only learns new representations for the data from different domains but it also designs a discriminative and robust classifier that properly adapts to the new representation. The proposed method was tested with three hyperspectral remote sensing data sets, and the experimental results confirm the effectiveness and reliability of the proposed SSCA-SC method. Liangpei Zhang 0001, Bo Du 0001, Lefei Zhang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2019 | Tracking Objects From Satellite Videos: A Velocity Feature Based Correlation FilterabstractSatellite video target tracking is a new topic in the remote sensing field, which refers to tracking moving objects of interest from satellite video in real time. The target of interest usually occupies only a few pixels in a satellite video image, even when the train is long. Thus, satellite video target tracking still faces new challenges compared with traditional visual tracking, including the detection of low-resolution targets, features with less representation, and targets with an extremely similar background. Little research has been done on satellite video target tracking, and little is known about whether or not the existing tracking algorithms can still work on the satellite video data. This paper, for the first time, intensively investigated 13 typical trackers in traditional visual tracking. The experimental results suggest that most of the state-of-the-art tracking algorithms mainly rely on luminance, color features, or convolutional features, and they fail to track satellite video targets due to their inadequate representation features. To overcome this difficulty, we propose a velocity correlation filter (VCF) algorithm, which employs both a velocity feature and an inertia mechanism (IM) to construct a specific kernel correlation filter for the satellite video target tracking. The velocity feature has a high discriminative ability to detect moving targets in satellite videos, and the IM can prevent model drift adaptively. Experimental results on three real satellite video data sets show that the VCF outperforms state-of-the-art tracking methods with regard to precision and success plots while running at over 100 frames per second. Jia Shao, Bo Du 0001, Chen Wu 0003, Lefei Zhang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2019 | Can We Track Targets From Space? A Hybrid Kernel Correlation Filter Tracker for Satellite VideoabstractDespite the great success of correlation filter-based trackers in visual tracking, it is questionable whether they can still perform on the satellite video data, acquired by a satellite or space station very high above the earth. The difficulty lies in that the targets usually occupy only a few pixels compared with the image size of over one million pixels and almost melt into the similar background. Since correlation filter models strongly depend on the quality of features and the spatial layout of the tracked object, they would probably fail on satellite video tracking tasks. In this paper, we propose a hybrid kernel correlation filter (HKCF) tracker employing two complementary features adaptively in a ridge regression framework. One feature is the optical flow that can detect variation pixels of the target. The other one is the histogram of oriented gradient that can capture the contour and texture information in the target, and an adaptive fusion strategy is proposed to employ the strengths of both features in different satellite videos. Quantitative evaluations are performed on six real satellite video data sets. The results show that our approach outperforms state-of-the-art tracking methods while running at more than 100 frames/s. Jia Shao, Bo Du 0001, Chen Wu 0003, Lefei Zhang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2018 | Nonlocal Patch Based t-SVD for Image Inpainting: Algorithm and Error AnalysisabstractIn this paper, we propose a novel image inpainting framework consisting of an interpolation step and a low-rank tensor completion step. More specifically, we first initial the image with triangulation-based linear interpolation, and then we find similar patches for each missing-entry centered patch. Treating a group of patch matrices as a tensor, we employ the recently proposed effective t-SVD tensor completion algorithm with a warm start strategy to inpaint it. We observe that the interpolation step is such a rough initialization that the similar patch we found may not exactly match with the reference, so we name the problem as Patch Mismatch and analyse the error caused by it thoroughly. Our theoretical analysis shows that the error caused by Patch Mismatch can be decomposed into two components, one of which can be bounded by a reasonable assumption named local patch similarity, and another part is lower than that using matrix. Experiments on real images verify our method's superiority to the state-of-the-art inpainting methods. Liangchen Song, Bo Du 0001, Lefei Zhang, Liangpei Zhang 0001, Jia Wu 0001, Xuelong Li 0001 |
AAAI | 3 |
| 2018 | Independent Feature and Label Components for Multi-label ClassificationabstractInvestigating correlation between example features and example labels is essential to solve classification problems. However, identification and calculation of the correlation between features and labels can be rather difficult for high-dimensional multi-label data. Both feature embedding and label embedding have been developed to tackle this challenge, and a shared subspace for both labels and features are usually learned by existing embedding methods to simultaneously reduce dimensionality of features and labels. In contrast, this paper suggests to learn separated subspaces for features and labels by maximizing the independence between components in each subspace and maximizing the correlation between these two subspaces. The learned independent label components indicates fundamental combinations of labels in multi-label datasets, which thus helps to reveals the correlation between labels. On the other hand, the learned independent feature components lead to a compact representation of example features. The connections between the proposed algorithm and existing embedding methods have been discussed. Experimental results on real-world multi-label datasets demonstrate the necessity of exploring independence components from multi-label data and the effectiveness of the proposed algorithm. Yongjian Zhong, Chang Xu 0002, Bo Du 0001, Lefei Zhang |
ICDM | 4 |
| 2018 | TLR: Transfer Latent Representation for Unsupervised Domain AdaptationabstractDomain adaptation refers to the process of learning prediction models in a target domain by making use of data from a source domain. Many classic methods solve the domain adaptation problem by establishing a common latent space, which may cause the loss of many important properties across both domains. In this manuscript, we develop a novel method, transfer latent representation (TLR), to learn a better latent space. Specifically, we design an objective function based on a simple linear autoencoder to derive the latent representations of both domains. The encoder in the autoencoder aims to project the data of both domains into a robust latent space. Besides, the decoder imposes an additional constraint to reconstruct the original data, which can preserve the common properties of both domains and reduce the noise that causes domain shift. Experiments on cross-domain tasks demonstrate the advantages of TLR over competing methods. Bo Du 0001, Jia Wu 0001, Lefei Zhang, Ruimin Hu, Xuelong Li 0001 |
ICME | 4 |
| 2018 | Discriminant Spatial-Spectral Hypergraph Learning for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) contains a large number of spatial-spectral information, which will make the traditional classification methods face an enormous challenge to discriminate the types of land-cover. Feature learning is very effective to improve the classification performances. However, the current feature learning approaches are most based on a simple intrinsic structure. To represent the complex intrinsic spatial-spectral of HSI, a novel feature learning algorithm, termed discriminant spatial-spectral hypergraph learning (DSSHL), has been proposed on the basis of spatial-spectral information and hypergraph learning. DSSHL constructs an intraclass spatial-spectral hypergraph and an interclass spatial-spectral hypergraph to represent the intrinsic properties of HSI. Then, a feature learning model is designed to compact the intraclass information and separate the interclass information. DSSHL can effectively reveal the complex spatial-spectral structures of HSI for land-cover classification. Experimental results on the Salinas HSI data set shows that DSSHL can achieve better classification accuracies in comparison with some state-of-the-art methods. Fulin Luo, Liangpei Zhang 0001, Bo Du 0001, Lefei Zhang, Yanni Dong |
IGARSS | 4 |
| 2018 | R-SVM+: Robust Learning with Privileged InformationabstractIn practice, the circumstance that training and test data are clean is not always satisfied. The performance of existing methods in the learning using privileged information (LUPI) paradigm may be seriously challenged, due to the lack of clear strategies to address potential noises in the data. This paper proposes a novel Robust SVM+ (RSVM+) algorithm based on a rigorous theoretical analysis. Under the SVM+ framework in the LUPI paradigm, we study the lower bound of perturbations of both example feature data and privileged feature data, which will mislead the model to make wrong decisions. By maximizing the lower bound, tolerance of the learned model over perturbations will be increased. Accordingly, a novel regularization function is introduced to upgrade a variant form of SVM+. The objective function of RSVM+ is transformed into a quadratic programming problem, which can be efficiently optimized using off-the-shelf solvers. Experiments on real-world datasets demonstrate the necessity of studying robust SVM+ and the effectiveness of the proposed algorithm. Bo Du 0001, Chang Xu 0002, Yipeng Zhang 0001, Lefei Zhang, Dacheng Tao |
IJCAI | 5 |
| 2018 | Self-Representative Manifold Concept Factorization with Adaptive Neighbors for ClusteringabstractMatrix Factorization based methods, e.g., the Concept Factorization (CF) and Nonnegative Matrix Factorization (NMF), have been proved to be efficient and effective for data clustering tasks. In recent years, various graph extensions of CF and NMF have been proposed to explore intrinsic geometrical structure of data for the purpose of better clustering performance. However, many methods build the affinity matrix used in the manifold structure directly based on the input data. Therefore, the clustering results are highly sensitive to the input data. To further improve the clustering performance, we propose a novel manifold concept factorization model with adaptive neighbor structure to learn a better affinity matrix and clustering indicator matrix at the same time. Technically, the proposed model constructs the affinity matrix by assigning the adaptive and optimal neighbors to each point based on the local distance of the learned new representation of the original data with itself as a dictionary. Our experimental results present superior performance over the state-of-the-art alternatives on numerous datasets. Sihan Ma, Lefei Zhang, Wenbin Hu 0001, Yipeng Zhang 0001, Jia Wu 0001, Xuelong Li 0001 |
IJCAI | 2 |
| 2018 | Discriminatively guided filtering (DGF) for hyperspectral image classification
Ziyu Wang 0003, Huafeng Hu, Lefei Zhang, Jing-Hao Xue |
Neurocomputing | 3 |
| 2018 | Simultaneous Spectral-Spatial Feature Selection and Extraction for Hyperspectral ImagesabstractIn hyperspectral remote sensing data mining, it is important to take into account of both spectral and spatial information, such as the spectral signature, texture feature, and morphological property, to improve the performances, e.g., the image classification accuracy. In a feature representation point of view, a nature approach to handle this situation is to concatenate the spectral and spatial features into a single but high dimensional vector and then apply a certain dimension reduction technique directly on that concatenated vector before feed it into the subsequent classifier. However, multiple features from various domains definitely have different physical meanings and statistical properties, and thus such concatenation has not efficiently explore the complementary properties among different features, which should benefit for boost the feature discriminability. Furthermore, it is also difficult to interpret the transformed results of the concatenated vector. Consequently, finding a physically meaningful consensus low dimensional feature representation of original multiple features is still a challenging task. In order to address these issues, we propose a novel feature learning framework, i.e., the simultaneous spectral-spatial feature selection and extraction algorithm, for hyperspectral images spectral-spatial feature representation and classification. Specifically, the proposed method learns a latent low dimensional subspace by projecting the spectral-spatial feature into a common feature space, where the complementary information has been effectively exploited, and simultaneously, only the most significant original features have been transformed. Encouraging experimental results on three public available hyperspectral remote sensing datasets confirm that our proposed method is effective and efficient. Lefei Zhang, Qian Zhang 0009, Bo Du 0001, Xin Huang 0002, Yuan Yan Tang, Dacheng Tao |
IEEE Trans. Cybern. | 1 |
| 2017 | Robust Manifold Matrix Factorization for Joint Clustering and Feature ExtractionabstractLow-rank matrix approximation has been widely used for data subspace clustering and feature representation in many computer vision and pattern recognition applications. However, in order to enhance the discriminability, most of the matrix approximation based feature extraction algorithms usually generate the cluster labels by certain clustering algorithm (e.g., the kmeans) and then perform the matrix approximation guided by such label information. In addition, the noises and outliers in the dataset with large reconstruction errors will easily dominate the objective function by the conventional ℓ2-norm based squared residue minimization. In this paper, we propose a novel clustering and feature extraction algorithm based on an unified low-rank matrix factorization framework, which suggests that the observed data matrix can be approximated by the production of projection matrix and low dimensional representation, among which the low-dimensional representation can be approximated by the cluster indicator and latent feature matrix simultaneously. Furthermore, we have proposed using the ℓ2,1-norm and integrating the manifold regularization to further promote the proposed model. A novel Augmented Lagrangian Method (ALM) based procedure is designed to effectively and efficiently seek the optimal solution of the problem. The experimental results in both clustering and feature extraction perspectives demonstrate the superior performance of the proposed method. Lefei Zhang, Qian Zhang 0009, Bo Du 0001, Dacheng Tao, Jane You |
AAAI | 1 |
| 2017 | On Gleaning Knowledge from Multiple Domains for Active LearningabstractHow can a doctor diagnose new diseases with little historical knowledge, which are emerging over time? Active learning is a promising way to address the problem by querying the most informative samples. Since the diagnosed cases for new disease are very limited, gleaning knowledge from other domains (classical prescriptions) to prevent the bias of active leaning would be vital for accurate diagnosis. In this paper, a framework that attempts to glean knowledge from multiple domains for active learning by querying the most uncertain and representative samples from the target domain and calculating the importance weights for re-weighting the source data in a single unified formulation is proposed. The weights are optimized by both a supervised classifier and distribution matching between the source domain and target domain with maximum mean discrepancy. Besides, a multiple domains active learning method is designed based on the proposed framework as an example. The proposed method is verified with newsgroups and handwritten digits data recognition tasks, where it outperforms the state-of-the-art methods. Zengmao Wang, Bo Du 0001, Lefei Zhang, Liangpei Zhang 0001, Ruimin Hu, Dacheng Tao |
IJCAI | 3 |
| 2017 | Adaptive Manifold Regularized Matrix Factorization for Data ClusteringabstractData clustering is the task to group the data samples into certain clusters based on the relationships of samples and structures hidden in data, and it is a fundamental and important topic in data mining and machine learning areas. In the literature, the spectral clustering is one of the most popular approaches and has many variants in recent years. However, the performance of spectral clustering is determined by the affinity matrix, which is always computed by a predefined model (e.g., Gaussian kernel function) with carefully tuned parameters combination, and may far from optimal in practice. In this paper, we propose to consider the observed data clustering as a robust matrix factorization point of view, and learn an affinity matrix simultaneously to regularize the proposed matrix factorization. The solution of the proposed adaptive manifold regularized matrix factorization (AMRMF) is reached by a novel Augmented Lagrangian Multiplier (ALM) based algorithm. The experimental results on standard clustering datasets demonstrate the superior performance over the exist alternatives. Lefei Zhang, Qian Zhang 0009, Bo Du 0001, Jane You, Dacheng Tao |
IJCAI | 1 |
| 2017 | Multi-class active learning: A hybrid informative and representative criterion inspired approachabstractLabeling each instance in a large-scale data set is extremely labor- and time-consuming. One way to alleviate this problem is active learning, which aims to discover the most valuable instances for labeling to construct a powerful classifier with low generalization error. Considering both informativeness and representativeness provides a promising way to design a practical active learning. However, most existing active learning methods select instances favoring either informativeness or representativeness. Meanwhile, many are designed based on the binary class, so that they may present suboptimal solutions on the data sets with multiple classes. In this paper, a hybrid informative and representative criterion based multi-class active learning approach is proposed. We combine the informativeness and representativeness into one formulation, which can be solved under a unified framework. The informativeness is measured by the margin minimum while the representative information is measured by the maximum mean discrepancy. By minimizing the loss risk, we generalize the loss risk minimization principle to the multi-class active learning setting. Hence, the proposed method is not only suitable to the binary class but also the multiple classes. We conduct our experiments on twelve benchmark UCI data sets, and the experimental results demonstrate that the proposed method performs better than some state-of-the-art methods. Zengmao Wang, Bo Du 0001, Lefei Zhang |
IJCNN | 3 |
| 2017 | Stochastic Decorrelation Constraint Regularized Auto-Encoder for Visual Recognition
Fengling Mao, Wei Xiong 0008, Bo Du 0001, Lefei Zhang |
MMM (2) | 4 |
| 2017 | LAM3L: Locally adaptive maximum margin metric learning for visual data classification
Yanni Dong, Bo Du 0001, Lefei Zhang, Liangpei Zhang 0001, Dacheng Tao |
Neurocomputing | 3 |
| 2017 | GPU Parallel Implementation of Isometric Mapping for Hyperspectral ClassificationabstractManifold learning algorithms such as the isometric mapping (ISOMAP) algorithm have been widely used in the analysis of hyperspectral images (HSIs), for both visualization and dimension reduction. As advanced versions of the traditional linear projection techniques, the manifold learning algorithms find the low-dimensional feature representation by nonlinear mapping, which can better preserve the local structure of the original data and thus benefit the data analysis. However, the high computational complexity of the manifold learning algorithms hinders their application in HSI processing. Although there are a few parallel implementations of manifold learning approaches that are available in the remote sensing community, they have not been designed to accelerate the eigen-decomposition process, which is actually the most time-consuming part of the manifold learning algorithms. In this letter, as a case study, we discuss the graphics processing unit parallel implementation of the ISOMAP algorithm. In particular, we focus on the eigen-decomposition process and verify the applicability of the proposed method by validating the embedding vectors and the subsequent classification accuracies. The experimental results obtained on different HSI data sets show an excellent speedup performance and consistent classification accuracy compared with the serial implementation. Liangpei Zhang 0001, Lefei Zhang, Bo Du 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2017 | Combining local and global: Rich and robust feature pooling for visual recognition
Wei Xiong 0008, Lefei Zhang, Bo Du 0001, Dacheng Tao |
Pattern Recognit. | 2 |
| 2017 | Real-time tracking based on weighted compressive tracking and a cognitive memory model
Bo Du 0001, Chen Wu 0003, Lefei Zhang, Liangpei Zhang 0001 |
Signal Process. | 4 |
| 2017 | Exploring Representativeness and Informativeness for Active LearningabstractHow can we find a general way to choose the most suitable samples for training a classifier? Even with very limited prior information? Active learning, which can be regarded as an iterative optimization procedure, plays a key role to construct a refined training set to improve the classification performance in a variety of applications, such as text analysis, image recognition, social network modeling, etc. Although combining representativeness and informativeness of samples has been proven promising for active sampling, state-of-the-art methods perform well under certain data structures. Then can we find a way to fuse the two active sampling criteria without any assumption on data? This paper proposes a general active learning framework that effectively fuses the two criteria. Inspired by a two-sample discrepancy problem, triple measures are elaborately designed to guarantee that the query samples not only possess the representativeness of the unlabeled data but also reveal the diversity of the labeled data. Any appropriate similarity measure can be employed to construct the triple measures. Meanwhile, an uncertain measure is leveraged to generate the informativeness criterion, which can be carried out in different ways. Rooted in this framework, a practical active learning algorithm is proposed, which exploits a radial basis function together with the estimated probabilities to construct the triple measures and a modified best-versus-second-best strategy to construct the uncertain measure, respectively. Experimental results on benchmark datasets demonstrate that our algorithm consistently achieves superior performance over the state-of-the-art active learning algorithms. Bo Du 0001, Zengmao Wang, Lefei Zhang, Liangpei Zhang 0001, Wei Liu 0005, Jialie Shen 0001, Dacheng Tao |
IEEE Trans. Cybern. | 3 |
| 2017 | Stacked Convolutional Denoising Auto-Encoders for Feature RepresentationabstractDeep networks have achieved excellent performance in learning representation from visual data. However, the supervised deep models like convolutional neural network require large quantities of labeled data, which are very expensive to obtain. To solve this problem, this paper proposes an unsupervised deep network, called the stacked convolutional denoising auto-encoders, which can map images to hierarchical representations without any label information. The network, optimized by layer-wise training, is constructed by stacking layers of denoising auto-encoders in a convolutional way. In each layer, high dimensional feature maps are generated by convolving features of the lower layer with kernels learned by a denoising auto-encoder. The auto-encoder is trained on patches extracted from feature maps in the lower layer to learn robust feature detectors. To better train the large network, a layer-wise whitening technique is introduced into the model. Before each convolutional layer, a whitening layer is embedded to sphere the input data. By layers of mapping, raw images are transformed into high-level feature representations which would boost the performance of the subsequent support vector machine classifier. The proposed algorithm is evaluated by extensive experimentations and demonstrates superior classification performance to state-of-the-art unsupervised networks. Bo Du 0001, Wei Xiong 0008, Jia Wu 0001, Lefei Zhang, Liangpei Zhang 0001, Dacheng Tao |
IEEE Trans. Cybern. | 4 |
| 2017 | Dimensionality Reduction and Classification of Hyperspectral Images Using Ensemble Discriminative Local Metric LearningabstractThe high-dimensional data space of hyperspectral images (HSIs) often result in ill-conditioned formulations, which finally leads to many of the high-dimensional feature spaces being empty and the useful data existing primarily in a subspace. To avoid these problems, we use distance metric learning for dimensionality reduction. The goal of distance metric learning is to incorporate abundant discriminative information by reducing the dimensionality of the data. Considering that global metric learning is not appropriate for all training samples, this paper proposes an ensemble discriminative local metric learning (EDLML) algorithm for HSI analysis. The EDLML algorithm learns robust local metrics from both the training samples and the relative neighborhood of them and considers the different local discriminative distance metrics by dealing with the data region by region. It aims to learn a subspace to keep all the samples in the same class are as near as possible, while those from different classes are separated. The learned local metrics are then used to build an ensemble metric. Experiments on a number of different hyperspectral data sets confirm the effectiveness of the proposed EDLML algorithm compared with that of the other dimension reduction methods. Yanni Dong, Bo Du 0001, Liangpei Zhang 0001, Lefei Zhang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2017 | A Sparse and Low-Rank Near-Isometric Linear Embedding Method for Feature Extraction in Hyperspectral Imagery ClassificationabstractA sparse and low-rank near-isometric linear embedding (SLRNILE) method has been proposed to make dimensionality reduction and extract proper features for hyperspectral imagery (HSI) classification. The SLRNILE stands on the theory of the John-Lindenstrauss lemma, and tries to estimate a sparse and low-rank projection matrix that satisfies the restricted isometric property (RIP) condition on all secants of the HSI data. The RIP condition guarantees that the desired linear mapping near-isometrically preserves nearest neighbor points of all HSI pixels. Seeking the desired mapping is then modeled into minimizing a Lagrange multipliers formulation. The alternating direction method of multipliers framework is utilized to solve the above convex program, and column generation techniques are adopted to alleviate the computation memory burden during the optimization procedure. Five experiments on three widely used HSI data sets are designed to completely test the performance of SLRNILE, and experimental results are compared against those of six state-of-the-art feature extraction methods, including principal component analysis, Laplacian eigenmaps, locality preserving projections, neighborhood preserving embedding, sparse nonnegative matrix underapproximation, and random projections. The results show that SLRNILE performs best among all the seven methods, and its computational time is longest of all but still bearable for regular users. Therefore, the SLRNILE can be a good choice for feature extraction in HSI classification. Weiwei Sun 0005, Gang Yang 0006, Bo Du 0001, Lefei Zhang, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2017 | A Novel Semisupervised Active-Learning Algorithm for Hyperspectral Image ClassificationabstractLess training samples are a challenging problem in hyperspectral image classification. Active learning and semisupervised learning are two promising techniques to address the problem. Active learning solves the problem by improving the quality of the training samples, while semisupervised learning solves the problem by increasing the quantity of the training samples. However, they pay too much attention to the discriminative information in the unlabeled data, leading to information bias to train supervised models, and much more effort to label samples. Therefore, a method to discover representativeness and discriminativeness by semisupervised active learning is proposed. It takes advantages of both active learning and semisupervised learning. The representativeness and discriminativeness are discovered with a labeling process based on a supervised clustering technique and classification results. Specifically, the supervised clustering results can discover important structural information in the unlabeled data, and the classification results are also highly confidential in the active-learning process. With these clustering results and classification results, we can assign pseudolabels to the unlabeled data. Meanwhile, the unlabeled samples that cannot be assigned with pseudolabels with high confidence at each iteration are regarded as candidates in active learning. The methodology is validated on four hyperspectral data sets. Significant improvements in classification accuracy are achieved by the proposed method with respect to the state-of-the-art methods. Zengmao Wang, Bo Du 0001, Lefei Zhang, Liangpei Zhang 0001, Xiuping Jia |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2017 | Robust and Discriminative Labeling for Multi-Label Active Learning Based on Maximum Correntropy CriterionabstractMulti-label learning draws great interests in many real world applications. It is a highly costly task to assign many labels by the oracle for one instance. Meanwhile, it is also hard to build a good model without diagnosing discriminative labels. Can we reduce the label costs and improve the ability to train a good model for multi-label learning simultaneously? Active learning addresses the less training samples problem by querying the most valuable samples to achieve a better performance with little costs. In multi-label active learning, some researches have been done for querying the relevant labels with less training samples or querying all labels without diagnosing the discriminative information. They all cannot effectively handle the outlier labels for the measurement of uncertainty. Since maximum correntropy criterion (MCC) provides a robust analysis for outliers in many machine learning and data mining algorithms, in this paper, we derive a robust multi-label active learning algorithm based on an MCC by merging uncertainty and representativeness, and propose an efficient alternating optimization method to solve it. With MCC, our method can eliminate the influence of outlier labels that are not discriminative to measure the uncertainty. To make further improvement on the ability of information measurement, we merge uncertainty and representativeness with the prediction labels of unknown data. It cannot only enhance the uncertainty but also improve the similarity measurement of multi-label data with labels information. Experiments on benchmark multi-label data sets have shown a superior performance than the state-of-the-art methods. Bo Du 0001, Zengmao Wang, Lefei Zhang, Liangpei Zhang 0001, Dacheng Tao |
IEEE Trans. Image Process. | 3 |
| 2017 | Robust Dual Clustering with Adaptive Manifold RegularizationabstractIn recent years, various data clustering algorithms have been proposed in the data mining and engineering communities. However, there are still drawbacks in traditional clustering methods which are worth to be further investigated, such as clustering for the high dimensional data, learning an ideal affinity matrix which optimally reveals the global data structure, discovering the intrinsic geometrical and discriminative properties of the data space, and reducing the noises influence brings by the complex data input. In this paper, we propose a novel clustering algorithm called robust dual clustering with adaptive manifold regularization (RDC), which simultaneously performs dual matrix factorization tasks with the target of an identical cluster indicator in both of the original and projected feature spaces, respectively. Among which, the$l_{2,1}$-norm is used instead of the conventional$l_{2}$-norm to measure the loss, which helps to improve the model robustness by relieving the influences by the noises and outliers. In order to better consider the intrinsic geometrical and discriminative data structure, we incorporate the manifold regularization term on the cluster indicator by using a particularly learned affinity matrix which is more suitable for the clustering task. Moreover, a novel augmented lagrangian method (ALM) based procedure is designed to effectively and efficiently seek the optimal solution of the proposed RDC optimization. Numerous experiments on the representative data sets demonstrate the superior performance of the proposed method compares to the existing clustering algorithms. Nengwen Zhao, Lefei Zhang, Bo Du 0001, Qian Zhang 0009, Jane You, Dacheng Tao |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2017 | PLTD: Patch-Based Low-Rank Tensor Decomposition for Hyperspectral ImagesabstractRecent years has witnessed growing interest in hyperspectral image (HSI) processing. In practice, however, HSIs always suffer from huge data size and mass of redundant information, which hinder their application in many cases. HSI compression is a straightforward way of relieving these problems. However, most of the conventional image encoding algorithms mainly focus on the spatial dimensions, and they need not consider the redundancy in the spectral dimension. In this paper, we propose a novel HSI compression and reconstruction algorithm via patch-based low-rank tensor decomposition (PLTD). Instead of processing the HSI separately by spectral channel or by pixel, we represent each local patch of the HSI as a third-order tensor. Then, the similar tensor patches are grouped by clustering to form a fourth-order tensor per cluster. Since the grouped tensor is assumed to be redundant, each cluster can be approximately decomposed to a coefficient tensor and three dictionary matrices, which leads to a low-rank tensor representation of both the spatial and spectral modes. The reconstructed HSI can then be simply obtained by the product of the coefficient tensor and dictionary matrices per cluster. In this way, the proposed PLTD algorithm simultaneously removes the redundancy in both the spatial and spectral domains in a unified framework. The extensive experimental results on various public HSI datasets demonstrate that the proposed method outperforms the traditional image compression approaches and other tensor-based methods. Bo Du 0001, Mengfei Zhang, Lefei Zhang, Ruimin Hu, Dacheng Tao |
IEEE Trans. Multim. | 3 |
| 2016 | Multi-label Active Learning Based on Maximum Correntropy Criterion: Towards Robust and Discriminative Labeling
Zengmao Wang, Bo Du 0001, Lefei Zhang, Liangpei Zhang 0001, Dacheng Tao |
ECCV (3) | 3 |
| 2016 | Regularizing Deep Convolutional Neural Networks with a Structured Decorrelation ConstraintabstractDeep convolutional networks have achieved successful performance in data mining field. However, training large networks still remains a challenge, as the training data may be insufficient and the model can easily get overfitted. Hence the training process is usually combined with a model regularization. Typical regularizers include weight decay, Dropout, etc. In this paper, we propose a novel regularizer, named Structured Decorrelation Constraint (SDC), which is applied to the activations of the hidden layers to prevent overfitting and achieve better generalization. SDC impels the network to learn structured representations by grouping the hidden units and encouraging the units within the same group to have strong connections during the training procedure. Meanwhile, it forces the units in different groups to learn non-redundant representations by minimizing the cross-covariance between them. Compared with Dropout, SDC reduces the co-adaptions between the hidden units in an explicit way. Besides, we propose a novel approach called Reg-Conv that can help SDC to regularize the complex convolutional layers. Experiments on extensive datasets show that SDC significantly reduces overfitting and yields very meaningful improvements on classification performance (on CIFAR-10 6.22% accuracy promotion and on CIFAR-100 9.63% promotion). Wei Xiong 0008, Bo Du 0001, Lefei Zhang, Ruimin Hu, Dacheng Tao |
ICDM | 3 |
| 2016 | Multiview clustering based on Robust and Regularized Matrix ApproximationabstractPattern recognition tasks such as the data classification and clustering usually can be represented by the perspective of multiple views or feature spaces. Obviously, the accuracy of the classification and clustering should be greatly improved if we carefully consider the discriminabilities from multiple views and explore the complementary information among them. However, multiple features also bring new challenges to handle them. In the literature, many existed multiview feature learning methods dealt with different views equally, thus they couldn't optimally utilize the complementary property of them. On the other hand, the matrix factorization based clustering algorithms usually adopt the conventional ℓ2-norm based squared residue minimization to measure the loss, which is easily influenced by the outliers and noises from the multiple sources of input. In this paper, we propose a novel multiview data clustering algorithm based on the matrix factorization to relieve the above issues. The basic idea of the proposed Robust and Regularized Matrix Approximation (RRMA) is that the observed data matrix could be low-rank approximated by a cluster centroid matrix and a cluster indicator matrix, respectively, and the major contributions of our work lie in the introduction of the robust ℓ2,1-norm and ensemble manifold regularization to regularize the matrix factorization and make the model more discriminative for multiview data clustering. We properly adjust the importance of different views by assigning a set of trainable weights on the views. Moreover, we propose an efficient solution featured with impactful updating rules to seek the local optimal parameters. Encouraging experimental results on numerous public multiview datasets demonstrate the superiority of our model compared to some state-of-the-art methods. Jiameng Pu, Qian Zhang 0009, Lefei Zhang, Bo Du 0001, Jane You |
ICPR | 3 |
| 2016 | A quantum-behaved particle swarm optimization for hyperspectral endmember extractionabstractIn this paper, endmember extraction algorithm is described as a combinatorial optimization problem. A novel quantum-behaved particle swarm optimization (QPSO) approach which employs quantum-behaved particle swarm optimization to find endmembers with good performance is proposed. As far as our knowledge, it is the first time that quantum-behaved particle swarm optimization is introduced into hyperspectral endmember extraction. In order to follow the law of particle movement, a high dimensional particles definition is proposed. The proposed algorithm was tested and evaluated by both synthetic and real hyperspectral data sets. Experimental results indicate that the proposed method get a better result compared to the algorithms of vertex component analysis (VCA), N-FINDR and discrete particle swarm optimization (D-PSO). Mingming Xu 0001, Liangpei Zhang 0001, Bo Du 0001, Lefei Zhang, Yuxiang Zhang 0001 |
IGARSS | 4 |
| 2016 | Denoising auto-encoders toward robust unsupervised feature representationabstractDeep networks like the convolutional neural network and its variants usually learn hierarchical features from labeled images, which is very expensive to obtain. How can we find an unsupervised way to effectively extract deep and abstract features from images without annotations? Even from large qualities of images with noise? In this paper, we propose a robust deep neural network, named as stacked convolutional denoising auto-encoders (SCDAE), which can map raw images to hierarchical representations in an unsupervised manner. Our network is elaborately designed to fit for the visual recognition tasks. It is established by stacking the denoising auto-encoders. Unlike the prior works, in the training phase, the auto-encoders are trained patch-wisely so that the latent features can be applied to powerful regularizers for better representation; in the inference phase, the denoising auto-encoders are stacked convolutionally, hence the generated feature maps in the higher layers can preserve the coherent structures within the features in the lower layers. To achieve better performance, we apply whitening to each layer to sphere the input features. Our network is evaluated on the challenging image datasets MNIST, CIFAR-10 and STL-10 and demonstrates superior performance to the state-of-the-art unsupervised networks. Wei Xiong 0008, Bo Du 0001, Lefei Zhang, Liangpei Zhang 0001, Dacheng Tao |
IJCNN | 3 |
| 2016 | Sparse tensor discriminative locality alignment for gait recognitionabstractGait recognition is a rising biometric technology which aims to distinguish people purely through the analysis of the way they walk, while the problem is that the dimensionality of the gait data is too high, so it is necessary to carry on dimensionality reduction task. Up to date, in the area of computer vision and pattern recognition, various dimensionality reduction algorithms have been employed for gait data, including the conventional vector representation based methods principal components analysis (PCA) and, locality preserving projection (LPP), and the recently proposed multi-linear subspace learning based approaches such as multilinear principal component analysis (MPCA). In this paper, inspired by the advantages of the tensor representation and manifold learning, we propose a novel sparse tensor discriminative locality alignment for human gait feature representation and dimensionality reduction algorithm, and subsequently apply the refined feature for gait recognition by a lazy classifier of the KNN. The proposed method adopts sparse multi-way projection based on the high-order version of discriminative locality alignment, by which the class separability is enhanced and the potential model overfitting is simultaneously avoided. Extensive experiments on the University of South Florida (USF) HumanID Gait Database show that the proposed method achieves better recognition rate compared with some existing classical dimensionality reduction algorithms. Nengwen Zhao, Lefei Zhang, Bo Du 0001, Liangpei Zhang 0001, Dacheng Tao, Jane You |
IJCNN | 2 |
| 2016 | Hyperspectral signal unmixing based on constrained non-negative matrix factorization approach
Bo Du 0001, Nan Wang 0016, Lefei Zhang, Dacheng Tao, Lifu Zhang 0002 |
Neurocomputing | 4 |
| 2016 | A minimal Munsell value error based laser printer model
Juhua Liu, Hai Su, Wenbin Hu 0001, Lefei Zhang, Dacheng Tao |
Neurocomputing | 4 |
| 2016 | A batch-mode active learning framework by querying discriminative and representative samples for hyperspectral image classification
Zengmao Wang, Bo Du 0001, Lefei Zhang, Liangpei Zhang 0001 |
Neurocomputing | 3 |
| 2016 | An image-based endmember bundle extraction algorithm using reconstruction error for hyperspectral imagery
Mingming Xu 0001, Liangpei Zhang 0001, Bo Du 0001, Lefei Zhang |
Neurocomputing | 4 |
| 2016 | Hierarchical feature learning with dropout k-means for hyperspectral image classification
Fan Zhang 0006, Bo Du 0001, Liangpei Zhang 0001, Lefei Zhang |
Neurocomputing | 4 |
| 2016 | Adaptive Laplacian Eigenmap-Based Dimension Reduction for Ocean Target DiscriminationabstractIt is well known that polarimetric synthetic aperture radar (PolSAR) backscattering features are highly influenced by the variation of incidence angle (VIA), which usually hampers the classification of most grazing-angle-sensitive targets, such as land and ocean targets. To relieve this issue, various feature extraction approaches have been suggested to enhance the class discriminability while reducing the observed feature dimensionality. The Laplacian eigenmap-based dimension reduction (DR) has been proven to be an effective way to deal with VIA problems, provided that the manifold parameters [e.g., the heat kernel (HK)] have been optimally sought, which is often difficult in practice. In this letter, an adaptive Laplacian eigenmap-based DR method is presented to find a learned subspace where the local geometry with discriminative prior knowledge is preserved as much as possible while near optimal HK and scale factor parameters are automatically identified. The learned feature representation is then employed for the subsequent classification. The improved Laplacian eigenmap algorithm was validated by three uninhabited-aerial-vehicle-synthetic-aperture-radar L-band PolSAR images from the Gulf Deepwater Horizon oil spill, which were clearly impacted by the VIA phenomenon. The experimental results showed that the proposed algorithm works well in ocean target discrimination compared with the current common methods. Lei Shi 0005, Lefei Zhang, Lingli Zhao, Liangpei Zhang 0001, Pingxiang Li, Dan Wu 0003 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2016 | Regularized set-to-set distance metric learning for hyperspectral image classification
Jiangtao Peng, Lefei Zhang, Luoqing Li |
Pattern Recognit. Lett. | 2 |
| 2016 | A spectral-spatial based local summation anomaly detection method for hyperspectral images
Bo Du 0001, Rui Zhao 0003, Liangpei Zhang 0001, Lefei Zhang |
Signal Process. | 4 |
| 2016 | A scene change detection framework for multi-temporal very high resolution remote sensing images
Chen Wu 0003, Lefei Zhang, Liangpei Zhang 0001 |
Signal Process. | 2 |
| 2016 | Support Tensor Machines for Classification of Hyperspectral Remote Sensing ImageryabstractIn recent years, the support vector machines (SVMs) have been very successful in remote sensing image classification, particularly when dealing with high-dimensional data and limited training samples. Nevertheless, the vector-based feature alignment of the SVM can lead to an information loss in representation of hyperspectral images, which intrinsically have a tensor-based data structure. In this paper, a new multiclass support tensor machine (STM) is specifically developed for hyperspectral image classification. Our newly proposed STM processes the hyperspectral image as a data cube and then identifies the information classes in tensor space. The multiclass STM is developed from a set of binary STM classifiers using the one-against-one parallel strategy. As a part of our tensor-based processing chain, a multilinear principal component analysis (MPCA) is used for preprocessing, in order to reduce the tensorial data redundancy and, at the same time, preserve the tensorial structure information in sparse and high-order subspaces. As a result, the contributions of this work are twofold: a new multiclass STM model for hyperspectral image classification is developed, and a tensorial image interpretation framework is constructed, which provides a system consisting of tensor-based feature representation, feature extraction, and classification. Experiments with four hyperspectral data sets, covering agricultural and urban areas, are conducted to validate the effectiveness of the proposed framework. Our experimental results show that the proposed STM and MPCA-STM can achieve better results than traditional SVM-based classifiers. Xin Huang 0002, Lefei Zhang, Liangpei Zhang 0001, Antonio Plaza, Jón Atli Benediktsson |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2016 | Multidomain Subspace Classification for Hyperspectral ImagesabstractHyperspectral imaging offers new opportunities for pattern recognition tasks in the remote sensing community through its improved discrimination in the spectral domain. However, such advanced image processing also brings new challenges due to the high data dimensionality in both the spatial and spectral domains. To relieve this issue, in this paper, we present a novel multidomain subspace (MDS) feature representation and classification method for hyperspectral images. The proposed method is based on a patch alignment framework. In order to optimally combine the feature representations from the various domains and simultaneously enhance the subspace discriminability, we incorporate the supervised label information into each domain and further generalize the framework to a multidomain version. Furthermore, we develop an iterative approach to alternately optimize the MDS objective function by considering it as two subconvex optimizations. The classification performance on three standard hyperspectral remote sensing images confirms the superiority of the proposed MDS algorithm over the state-of-the-art subspace learning methods. Liangpei Zhang 0001, Xiaojie Zhu, Lefei Zhang, Bo Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2016 | Beyond Background Feature Extraction: An Anomaly Detection Algorithm Inspired by Slowly Varying Signal AnalysisabstractBackground feature extraction is an important step in hyperspectral anomaly detection. However, the lack of prior information about anomaly targets and the complex spectral mixture result in a challenge for robust background feature extraction. Can we solve the anomaly detection problem other than with background feature extraction? Relative to anomalies, the background spectral signal is usually stable and slowly varying. In view of this point, slowly varying background analysis is introduced into anomaly detection in this paper. The desired background signals are obtained through a generalized eigenvalue decomposition problem based on the original data and the differential image. The extracted signals are then combined with a Mahalanobis distance metric to construct the detection estimation. Different data processing procedures and signal extraction patterns are respectively formulated to construct different versions of the slowly varying background-signal-based detector. The performances of the proposed methods were validated on both synthetic and real hyperspectral data. The experimental results reveal that the proposed methods outperform the state-of-the-art anomaly detectors, with superior receiver operating characteristic (ROC) curves, area-under-ROC values, and background-target separation. The sensitivity of the relevant parameters was also analyzed in an experimental analysis. Rui Zhao 0003, Bo Du 0001, Liangpei Zhang 0001, Lefei Zhang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2015 | Batch Mode Active Learning for Geographical Image Classification
Zengmao Wang, Bo Du 0001, Lefei Zhang, Wenbin Hu 0001, Dacheng Tao, Liangpei Zhang 0001 |
APWeb | 3 |
| 2015 | R2FP: Rich and Robust Feature Pooling for Mining Visual DataabstractThe human visual system proves smart in extracting both global and local features. Can we design a similar way for unsupervised feature learning? In this paper, we propose anovel pooling method within an unsupervised feature learningframework, named Rich and Robust Feature Pooling (R2FP), to better explore rich and robust representation from sparsefeature maps of the input data. Both local and global poolingstrategies are further considered to instantiate such a methodand intensively studied. The former selects the most conductivefeatures in the sub-region and summarizes the joint distributionof the selected features, while the latter is utilized to extractmultiple resolutions of features and fuse the features witha feature balancing kernel for rich representation. Extensiveexperiments on several image recognition tasks demonstratethe superiority of the proposed techniques. Wei Xiong 0008, Bo Du 0001, Lefei Zhang, Ruimin Hu, Wei Bian 0003, Jialie Shen 0001, Dacheng Tao |
ICDM | 3 |
| 2015 | MMFE: Multitask Multiview Feature EmbeddingabstractIn data mining and pattern recognition area, the learned objects are often represented by the multiple features from various of views. How to learn an efficient and effective feature embedding for the subsequent learning tasks? In this paper, we address this issue by providing a novel multi-task multiview feature embedding (MMFE) framework. The MMFE algorithm is based on the idea of low-rank approximation, which suggests that the observed multiview feature matrix is approximately represented by the low-dimensional feature embedding multiplied by a projection matrix. In order to fully consider the particular role of each view to the multiview feature embedding, we simultaneously suggest the multitask learning scheme and ensemble manifold regularization into the MMFE algorithm to seek the optimal projection. Since the objection function of MMFE is multi-variable and non-convex, we further provide an iterative optimization procedure to find the available solution. Two real world experiments show that the proposed method outperforms single-task-based as well as state-of-the-art multiview feature embedding methods for the classification problem. Qian Zhang 0009, Lefei Zhang, Bo Du 0001, Wei Bian 0003, Dacheng Tao |
ICDM | 2 |
| 2015 | Local decision maximum margin metric learning for hyperspectral target detectionabstractDetecting certain targets from hyperspectral images (HSIs) is of great interest for both civilian and military applications, with the aim being to detect and identify target pixels based on specific spectral signatures. However, the classical algorithms are generally dependent on the specific statistical hypothesis test, and the algorithms may only perform well with certain assumptions. Therefore, in this paper, a novel metric-learning-based target detection framework, named local decision maximum margin metric learning (LDM3L), is proposed for HSI target detection. The proposed method can better separate the target samples from background ones, without the need for certain assumptions. The experimental results demonstrate that the proposed method outperforms both the state-of-the-art target detection algorithms and the other classical metric learning methods. Yanni Dong, Bo Du 0001, Lefei Zhang, Liangpei Zhang 0001 |
IGARSS | 3 |
| 2015 | Compression of hyperspectral remote sensing images by tensor approach
Lefei Zhang, Liangpei Zhang 0001, Dacheng Tao, Xin Huang 0002, Bo Du 0001 |
Neurocomputing | 1 |
| 2015 | Ensemble manifold regularized sparse low-rank approximation for multiview feature embedding
Lefei Zhang, Qian Zhang 0009, Liangpei Zhang 0001, Dacheng Tao, Xin Huang 0002, Bo Du 0001 |
Pattern Recognit. | 1 |
| 2015 | A hypothesis independent subpixel target detector for hyperspectral Images
Bo Du 0001, Yuxiang Zhang 0001, Liangpei Zhang 0001, Lefei Zhang |
Signal Process. | 4 |
| 2015 | A sparse and discriminative tensor to vector projection for human gait feature representation
Lefei Zhang, Liangpei Zhang 0001, Dacheng Tao, Bo Du 0001 |
Signal Process. | 1 |
| 2014 | Hyperspectral biological images compression based on multiway tensor projectionabstractSince the hyperspectral images (HSI) could provide much more useful discriminative information that cannot be obtained by the conventional imaging techniques, the hyper-spectral imaging technology was widely used in remote sensing area and recently used in many other aspects, such as the biological images recognition. However, most of the time, the size of hyperspectral data is so large that to process these data is both time-consuming and space-consuming. In this paper, a multiway tensor projection (MTP) algorithm is proposed as an extension to the conventional PCA for hyperspectral data compression and reconstruction. Technologically speaking, MTP carries out a tensor data compression in all the modes simultaneously to seek a projection matrix along each order to make sure that the projected core tensor can preserve most of the information present in the original tensor. Since the MTP algorithm uses the arbitrary order tensor as the input, it can preserve the structure information not only among the rows and columns but also among the spectral channels as much as possible and without vectorization. Numerous experiments on hyperspectral biological databases show that the MTP algorithm has better compression performance than PCA in many aspects. Bo Du 0001, Mengfei Zhang, Lefei Zhang, Xuelong Li 0001 |
ICME | 3 |
| 2014 | Local Patch Discriminative Metric Learning for Hyperspectral Image Feature ExtractionabstractIn hyperspectral image (HSI) classification, feature extraction is one important step. Traditional methods, e.g., principal component analysis (PCA) and locality preserving projection, usually neglect the information of within-class similarity and between-class dissimilarity, which is helpful to the improvement of classification. On the other hand, most of these methods, e.g., PCA and linear discriminative analysis, consider that the HSI data lie on a low-dimensional manifold or each class is on a submanifold. However, some class data of HSI may lie on a multimanifold. To avoid these problems, we propose a method for feature extraction in HSIs, assuming that a local region resides on a submainfold. In our method, we deal with the data region by region by taking into account the different discriminative locality information. Then, under the metric learning framework, a robust distance metric is learned. It aims to learn a subspace in which the samples in the same class are as near as possible while the samples in different classes are as far as possible. Encouraging experimental results on two available hyperspectral data sets indicate that our proposed algorithm outperforms many existing feature extract methods for HSI classification. Qian Zhang 0009, Lefei Zhang, Lubin Weng |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2014 | Sparse Transfer Manifold Embedding for Hyperspectral Target DetectionabstractTarget detection is one of the most important applications in hyperspectral remote sensing image analysis. However, the state-of-the-art machine-learning-based algorithms for hyperspectral target detection cannot perform well when the training samples, especially for the target samples, are limited in number. This is because the training data and test data are drawn from different distributions in practice and given a small-size training set in a high-dimensional space, traditional learning models without the sparse constraint face the over-fitting problem. Therefore, in this paper, we introduce a novel feature extraction algorithm named sparse transfer manifold embedding (STME), which can effectively and efficiently encode the discriminative information from limited training data and the sample distribution information from unlimited test data to find a low-dimensional feature embedding by a sparse transformation. Technically speaking, STME is particularly designed for hyperspectral target detection by introducing sparse and transfer constraints. As a result of this, it can avoid over-fitting when only very few training samples are provided. The proposed feature extraction algorithm was applied to extensive experiments to detect targets of interest, and STME showed the outstanding detection performance on most of the hyperspectral datasets. Lefei Zhang, Liangpei Zhang 0001, Dacheng Tao, Xin Huang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2014 | Hyperspectral Remote Sensing Image Subpixel Target Detection Based on Supervised Metric LearningabstractThe detection and identification of target pixels such as certain minerals and man-made objects from hyperspectral remote sensing images is of great interest for both civilian and military applications. However, due to the restriction in the spatial resolution of most airborne or satellite hyperspectral sensors, the targets often appear as subpixels in the hyperspectral image (HSI). The observed spectral feature of the desired target pixel (positive sample) is therefore a mixed signature of the reference target spectrum and the background pixels spectra (negative samples), which belong to various land cover classes. In this paper, we propose a novel supervised metric learning (SML) algorithm, which can effectively learn a distance metric for hyperspectral target detection, by which target pixels are easily detected in positive space while the background pixels are pushed into negative space as far as possible. The proposed SML algorithm first maximizes the distance between the positive and negative samples by an objective function of the supervised distance maximization. Then, by considering the variety of the background spectral features, we put a similarity propagation constraint into the SML to simultaneously link the target pixels with positive samples, as well as the background pixels with negative samples, which helps to reject false alarms in the target detection. Finally, a manifold smoothness regularization is imposed on the positive samples to preserve their local geometry in the obtained metric. Based on the public data sets of mineral detection in an Airborne Visible/Infrared Imaging Spectrometer image and fabric and vehicle detection in a Hyperspectral Mapper image, quantitative comparisons of several HSI target detection methods, as well as some state-of-the-art metric learning algorithms, were performed. All the experimental results demonstrate the effectiveness of the proposed SML algorithm for hyperspectral target detection. Lefei Zhang, Liangpei Zhang 0001, Dacheng Tao, Xin Huang 0002, Bo Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2013 | Supervised Graph Embedding for Polarimetric SAR Image ClassificationabstractThis letter introduces an efficiency-manifold-learning-based supervised graph embedding (SGE) algorithm for polarimetric synthetic aperture radar (POLSAR) image classification. We use a linear dimensionality reduction technology named SGE to obtain a low-dimensional subspace which can preserve the discriminative information from training samples. Various POLSAR decomposition features are stacked into the input feature cube in the original high-dimensional feature space. The SGE is then implemented to project the input feature into the learned subspace for subsequent classification. The suggested method is validated by the full polarimetric airborne SAR system EMISAR, in Foulum, Denmark. The experiments show that the SGE presents a favorable classification accuracy and the valid components of the multifeature cube are also distinguished. Lei Shi 0005, Lefei Zhang, Jie Yang 0040, Liangpei Zhang 0001, Pingxiang Li |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2013 | Tensor Discriminative Locality Alignment for Hyperspectral Image Spectral-Spatial Feature ExtractionabstractIn this paper, we propose a method for the dimensionality reduction (DR) of spectral-spatial features in hyperspectral images (HSIs), under the umbrella of multilinear algebra, i.e., the algebra of tensors. The proposed approach is a tensor extension of conventional supervised manifold-learning-based DR. In particular, we define a tensor organization scheme for representing a pixel's spectral-spatial feature and develop tensor discriminative locality alignment (TDLA) for removing redundant information for subsequent classification. The optimal solution of TDLA is obtained by alternately optimizing each mode of the input tensors. The methods are tested on three public real HSI data sets collected by hyperspectral digital imagery collection experiment, reflective optics system imaging spectrometer, and airborne visible/infrared imaging spectrometer. The classification results show significant improvements in classification accuracies while using a small number of features. Liangpei Zhang 0001, Lefei Zhang, Dacheng Tao, Xin Huang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2012 | On Combining Multiple Features for Hyperspectral Remote Sensing Image ClassificationabstractIn hyperspectral remote sensing image classification, multiple features, e.g., spectral, texture, and shape features, are employed to represent pixels from different perspectives. It has been widely acknowledged that properly combining multiple features always results in good classification performance. In this paper, we introduce the patch alignment framework to linearly combine multiple features in the optimal way and obtain a unified low-dimensional representation of these multiple features for subsequent classification. Each feature has its particular contribution to the unified representation determined by simultaneously optimizing the weights in the objective function. This scheme considers the specific statistical properties of each feature to achieve a physically meaningful unified low-dimensional representation of multiple features. Experiments on the classification of the hyperspectral digital imagery collection experiment and reflective optics system imaging spectrometer hyperspectral data sets suggest that this scheme is effective. Lefei Zhang, Liangpei Zhang 0001, Dacheng Tao, Xin Huang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2011 | A Multifeature Tensor for Remote-Sensing Target RecognitionabstractIn remote-sensing image target recognition, the target or background object is usually transformed to a feature vector, such as a spectral feature vector. However, this kind of vector represents only one pixel of a remote-sensing image that considers the spectral information but ignores the spatial relationship of neighboring pixels (i.e., the local texture and structure). In this letter, we propose a new way to represent an image object as a multifeature tensor that encodes both the spectral and textural information (Gabor function) and then apply the support tensor machine for target recognition. A range of experiments demonstrates that the effectiveness of the proposed method can deliver a high and correct recognition rate with a small number of training samples. Lefei Zhang, Liangpei Zhang 0001, Dacheng Tao, Xin Huang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 1 |