VLDB 2026 Research / reviewers in the wild / expert
Junyu Gao 0001
dblp:153/4522-1
· DBLP profile ↗
86ranked-venue papers
11as first author
76since 2021 · last 2026
0000-0001-6000-8168ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 8 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 2 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 23 · 2 first-author · 21 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring Efficient Open-Vocabulary Segmentation in the Remote SensingabstractOpen-Vocabulary Remote Sensing Image Segmentation (OVRSIS), an emerging task that adapts Open-Vocabulary Segmentation (OVS) to the remote sensing (RS) domain, remains underexplored due to the absence of a unified evaluation benchmark and the domain gap between natural and RS images. To bridge these gaps, we first establish a standardized OVRSIS benchmark (OVRSISBench) based on widely-used RS segmentation datasets, enabling consistent evaluation across methods. Using this benchmark, we comprehensively evaluate several representative OVS/OVRSIS models and reveal their limitations when directly applied to remote sensing scenarios. Building on these insights, we propose RSKT-Seg, a novel open-vocabulary segmentation framework tailored for remote sensing. RSKT-Seg integrates three key components: (1) a Multi-Directional Cost Map Aggregation (RS-CMA) module that captures rotation-invariant visual cues by computing vision-language cosine similarities across multiple directions; (2) an Efficient Cost Map Fusion (RS-Fusion) transformer, which jointly models spatial and semantic dependencies with a lightweight dimensionality reduction strategy; and (3) a Remote Sensing Knowledge Transfer (RS-Transfer) module that injects pre-trained knowledge and facilitates domain adaptation via enhanced upsampling. Extensive experiments on the benchmark show that RSKT-Seg consistently outperforms strong OVS baselines by +3.8 mIoU and +5.9 mACC, while achieving 2× faster inference through efficient aggregation. Bingyu Li 0002, Haocheng Dong, Da Zhang 0010, Zhiyuan Zhao 0005, Hao Sun 0038, Junyu Gao 0001 |
AAAI | 6 |
| 2026 | Reasoning via Implicit Self-supervised Emergence for Instruction SegmentationabstractWe challenge the assumption that complex instruction-guided segmentation tasks necessitate equally complex and explicit supervision. This paper introduces RISE (Reasoning via Implicit Self-supervised Emergence), a framework that learns intricate compositional reasoning, spanning spatial relations to world knowledge, without a single ground-truth mask. To achieve this, RISE employs reinforcement learning with GRPO guided by a single, strikingly simple reward: the semantic alignment score between the textual instruction and the predicted image region. Our primary discovery is the implicit emergence of a high-quality chain-of-thought process from this minimalist signal. Within a structured format, the model autonomously learns to understand instructions by accessing its latent knowledge, inferring spatial relationships—capabilities inherent in its architecture but unlocked by our simple objective. Remarkably, our emergent reasoning yields highly competitive results: RISE achieves 58.7 gIoU on the ReasonSeg benchmark, on par with methods using geometric rewards. Furthermore, we show extreme data efficiency: a variant trained on only 2,000 ImageNet-label pairs establishes a new state-of-the-art for annotation-free referring segmentation with 79.6 cIoU on RefCOCO. Lichang Yang, Yuyu Jia, Junyu Gao 0001, Weiping Ni, Junzheng Wu, Qi Wang 0009 |
AAAI | 4 |
| 2026 | FusAD: Time-Frequency Fusion with Adaptive Denoising for General Time Series AnalysisabstractTime series analysis plays a vital role in fields such as finance, healthcare, industry, and meteorology, underpinning key tasks including classification, forecasting, and anomaly detection. Although deep learning models have achieved remarkable progress in these areas in recent years, constructing an efficient, multi-task compatible, and generalizable unified framework for time series analysis remains a significant challenge. Existing approaches are often tailored to single tasks or specific data types, making it difficult to simultaneously handle multi-task modeling and effectively integrate information across diverse time series types. Moreover, real-world data are often affected by noise, complex frequency components, and multi-scale dynamic patterns, which further complicate robust feature extraction and analysis. To ameliorate these challenges, we propose FusAD, a unified analysis framework designed for diverse time series tasks. FusAD features an adaptive time-frequency fusion mechanism, integrating both Fourier and Wavelet transforms to efficiently capture global-local and multi-scale dynamic features. With an adaptive denoising mechanism, FusAD automatically senses and filters various types of noise, highlighting crucial sequence variations and enabling robust feature extraction in complex environments. In addition, the framework integrates a general information fusion and decoding structure, combined with masked pre-training, to promote efficient learning and transfer of multi-granularity representations. Extensive experiments demonstrate that FusAD consistently outperforms state-of-the-art models on mainstream time series benchmarks for classification, forecasting, and anomaly detection tasks, while maintaining high efficiency and scalability. Code is available at https://github.com/zhangda1018/FusAD. Da Zhang 0010, Bingyu Li 0002, Zhiyuan Zhao 0005, Feiping Nie 0001, Junyu Gao 0001, Xuelong Li 0001 |
ICDE | 5 |
| 2026 | Multimodal Graph Conditioned Diffusion Model for Video CaptioningabstractVideo captioning aims to describe the content of a given video with condensed natural language sentences. Such a captioning task is full of challenges since the high requirements for visual-textual relevance and multimodal fusion understanding. Previous works primarily focus on visual content modeling, often overlooking the rich semantic correlations between visual and textual modalities, which results in incomplete understanding of the multimodal context and suboptimal caption accuracy. In this paper, we propose a multimodal graph conditioned diffusion model for video captioning, named MGCDVc. The idea behind our model is to incorporate graph-based relational reasoning with diffusion-based generative modeling to jointly model cross-modal relationships and capture latent semantic structure. Specifically, we learn a set of latent concept anchors to bridge the visual and textual modality nodes, enabling the construction of a weighted multimodal graph. Then we introduce the graph conditioned diffusion strategy which generates the textual semantic nodes and associated edges under the graph structure awareness condition. Furthermore, a soft pruning mechanism is designed to filter out low-quality nodes, thus further refining the generated multimodal graph to provide more accurate semantic structural guidance for caption generation. Experimental results on several popular datasets demonstrate that our model achieves better performance in video captioning task. Benhui Zhang 0001, Junyu Gao 0001, Yuan Yuan 0001 |
WWW | 2 |
| 2026 | Vision and acoustic emission multi-modal learning for aircraft crack monitoring
Kang Liu 0014, Ruiyao Huang, Gang Miao, Ruiyuan Wang, Junyu Gao 0001, Ju Huang, Xuelong Li 0001 |
Adv. Eng. Informatics | 7 |
| 2026 | Cross-attention multi-scale state space model for remaining useful life prediction of aircraft engines
Da Zhang 0010, Bingyu Li 0002, Zhiyuan Zhao 0005, Junyu Gao 0001, Xuelong Li 0001 |
Adv. Eng. Informatics | 5 |
| 2026 | Quantum-inspired interpretable deep learning architecture for text sentiment analysis
Bingyu Li 0002, Da Zhang 0010, Zhiyuan Zhao 0005, Yuan Yuan 0001, Junyu Gao 0001, Xuelong Li 0001 |
Neural Networks | 5 |
| 2026 | Dynamic proxy domain generalizes the crowd localization by better binary segmentation
Junyu Gao 0001, Da Zhang 0010, Qiyu Wang, Zhiyuan Zhao 0005, Xuelong Li 0001 |
Pattern Recognit. | 1 |
| 2026 | Distantly supervised reinforcement localization for real-world object distribution estimation
Haojie Guo, Junyu Gao 0001, Yuan Yuan 0001 |
Pattern Recognit. | 2 |
| 2026 | Fully PolSAR image reconstruction for enhanced land cover mapping
Junyu Gao 0001, Yuan Yuan 0001 |
Pattern Recognit. | 2 |
| 2026 | Adaptive momentum weight averaging reduces initialization noise
Jia Wan 0001, Ziquan Liu, Junyu Gao 0001, Antoni B. Chan |
Pattern Recognit. | 3 |
| 2026 | A benchmark For multi-lingual vision-language learning in remote sensing image captioning
Qi Wang 0009, Junyu Gao 0001, Weiping Ni, Junzheng Wu |
Pattern Recognit. | 4 |
| 2026 | Hierarchical textual-visual guidance for referring remote sensing segmentation
Qi Wang 0009, Yuan Yuan 0001, Junyu Gao 0001 |
Pattern Recognit. | 5 |
| 2026 | Balancing Optimization Strategies and Practical Goals: An Efficient Scene Text DetectorabstractScene text reading is a crucial task for scene understanding. Text detection, as a fundamental task in scene text reading, has recently garnered significant attention. Among various approaches, segmentation-based methods stand out for their flexible pixel-level prediction capabilities. However, two main issues remain. 1) These methods treat all text instances as a pixel set during training, causing the features of large-scale instances to dominate the model optimization process. As a result, the optimization deviates from the instance-level objectives. 2) Segmentation methods filter candidates based on pixel-level class scores, whereas what is needed is an evaluation of whether an instance is text, which also deviates from the original goals. To address these issues, we propose an Instance-Equal Feature Guide Module (IEFGM), a Cross-Level Feature Interaction Module (CLIFM), and a Pixel-Instance Fusion Discriminator (PIFD) to balance optimization strategies with practical goals. The IEFGM introduces instance-level features and positional information, guiding the model to treat instances of different scales equally at the feature level. The CLIFM encourages feature interaction across different levels, enabling the model to recognize text from various perspectives. Unlike existing methods that filter candidates using pixel-level results, the PIFD integrates both instance-level and pixel-level information to identify candidate regions, aligning with the original goals of text detection. A series of ablation studies demonstrates the effectiveness of the proposed modules. Extensive experiments across six datasets from different scenes demonstrate that our method outperforms existing state-of-the-art approaches. Xu Han 0019, Chuang Yang 0003, Junyu Gao 0001, Qi Wang 0009 |
IEEE Trans. Multim. | 3 |
| 2025 | Enhancing Low-Rank Adaptation with Recoverability-Based Reinforcement Pruning for Object CountingabstractObject counting is crucial for understanding the distribution of objects in different scenarios. Recently, many object counting networks have been designed to be more complex to achieve marginal improvements, leading to excessive time spent on model design. With the development of large models (LMs), various visual tasks can be accomplished by transferring pre-trained weights from LMs and fine-tuning them. However, tens of millions of training data make the pre-training parameters of LMs not entirely necessary. Moreover, if unnecessary parameters in the large model are not removed, it may lead to decreased performance on the tasks to be transferred. Motivated by this, this paper proposes an Enhancing low-Rank adaptation with Recoverability-based Reinforcement Pruning (E3RP) method to balance the complexity of large model and the accuracy of counting tasks. Firstly, we design a new reward mechanism based on the feature similarity of large model before and after globally unstructured pruning of specific parameters. Additionally, we propose a Patch Query Flip Attention (PQFA) mechanism to align multi-scale features through bidirectional interaction of features. Finally, the parameters of large model are pruned utilizing the pruning rate autonomously determined by the reinforcement learning network, and the large model is fine-tuned to counting tasks by a simple decoding head. Extensive experiments on four cross-scenario datasets demonstrate that the proposed method can remove redundant network parameters while ensuring network performance, with a maximum reduction of up to 63%. Haojie Guo, Junyu Gao 0001, Yuan Yuan 0001 |
AAAI | 2 |
| 2025 | LLMs Caught in the Crossfire: Malware Requests and Jailbreak ChallengesabstractThe widespread adoption of Large Language Models (LLMs) has heightened concerns about their security, particularly their vulnerability to jailbreak attacks that leverage crafted prompts to generate malicious outputs.While prior research has been conducted on general security capabilities of LLMs, their specific susceptibility to jailbreak attacks in code generation remains largely unexplored.To fill this gap, we propose MalwareBench, a benchmark dataset containing 3,520 jailbreaking prompts for malicious code-generation, designed to evaluate LLM robustness against such threats.Mal-wareBench is based on 320 manually crafted malicious code generation requirements, covering 11 jailbreak methods and 29 code functionality categories.Experiments show that mainstream LLMs exhibit limited ability to reject malicious code-generation requirements, and the combination of multiple jailbreak methods further reduces the model's security capabilities: specifically, the average rejection rate for malicious content is 60.93%, dropping to 39.92% when combined with jailbreak attack algorithms.Our work highlights that the code security capabilities of LLMs still pose significant challenges. Zhiyuan Zhao 0005, Junyu Gao 0001, Xuelong Li 0001 |
ACL (1) | 5 |
| 2025 | Unity in Diversity: Video Editing via Gradient-Latent PurificationabstractRecently, text-driven video editing methods that optimize target latent representations have garnered significant attention and demonstrated promising results. However, these methods rely on self-supervised objectives to compute the gradients needed for updating latent representations, which inevitably introduces gradient noise, compromising content generation quality. Additionally, it is challenging to determine the optimal stopping point for the editing process, making it difficult to achieve an optimal solution for the latent representation. To address these issues, we propose a unified gradient-latent purification framework that collects gradient and latent information across different stages to identify effective and concordant update directions. We design a local coordinate system construction method based on feature decomposition, enabling short-term gradients and final-stage latents to be reprojected onto new axes. Then, we employ tailored coefficient regularization terms to effectively aggregate the decomposed information. Additionally, a temporal smoothing axis extension strategy is developed to enhance the temporal coherence of the generated content. Extensive experiments demonstrate that our proposed method outperforms state-of-the-art methods across various editing tasks, delivering superior editing performance. Project page is available in https://unityin-diversity-editing.github.io. Junyu Gao 0001, Yufan Hu |
CVPR | 1 |
| 2025 | Scale Efficient Training for Large DatasetsabstractThe rapid growth of dataset scales has been a key driver in advancing deep learning research. However, as dataset scale increases, the training process becomes increasingly inefficient due to the presence of low-value samples, including excessive redundant samples, overly challenging samples, and inefficient easy samples that contribute little to model improvement. To address this challenge, we propose Scale Efficient Training (SeTa) for large datasets, a dynamic sample pruning approach that losslessly reduces training time. To remove low-value samples, SeTa first performs random pruning to eliminate redundant samples, then clusters the remaining samples according to their learning difficulty measured by loss. Building upon this clustering, a sliding window strategy is employed to progressively remove both overly challenging and inefficient easy clusters following an easy-to-hard curriculum. We conduct extensive experiments on large-scale synthetic datasets, including ToCa, SS1M, and ST+MJ, each containing over 3 million samples. SeTa reduces training costs by up to 50% while maintaining or improving performance, with minimal degradation even at 70% cost reduction. Furthermore, experiments on various scale real datasets across various backbones (CNNs, Transformers, and Mambas) and diverse tasks (instruction tuning, multi-view stereo, geo-localization, composed image retrieval, referring image segmentation) demonstrate the powerful effectiveness and universality of our approach. Code is available at https://github.com/mrazhou/SeTa. Junyu Gao 0001, Qi Wang 0009 |
CVPR | 2 |
| 2025 | KAID: Knowledge-Aware Interactive Distillation for Vision-Language ModelsabstractVision-Language Models (VLMs) such as CLIP have demonstrated outstanding performance in cross-modal tasks, but the prohibitive computational cost hinders practical deployment. Although Knowledge Distillation (KD) provides a promising compression paradigm, most existing methods rely heavily on feature imitation and contrastive relations without explicit fine-grained alignment. Additionally, they do not fully leverage the multimodal interaction knowledge from the teacher model, restricting cross-modal semantic alignment. To address these challenges, we propose KAID, a Knowledge-Aware Interactive Distillation method for VLMs. Specifically, we first pretrain a large CLIP teacher model with domain few-shot labels and store text features as category vectors. Then, an Image Feature Matching (IFM) module is introduced to calculate the feature distribution of teacher-student models with improved cosine similarity, which achieves hierarchical knowledge transfer from global to local levels and enhances the fine-grained perception of student model through joint optimization. Moreover, a Pixel-Wise Alignment (PWA) module is constructed between the teacher's text features and the student's image features, employing a cross-modal attention mechanism to establish semantic associations, while a Text-guided Pixel alignment Loss function (TPloss) is concurrently designed to enhance the student's comprehension capabilities. Ultimately, the well-trained student model is used for inference. Extensive experiments on 11 datasets validate the effectiveness of our method. Specifically, our method achieves average improvements of 2.14% and 2.40% on the base and new classes across these datasets. Da Zhang 0010, Bingyu Li 0002, Zhiyuan Zhao 0005, Junyu Gao 0001, Xuelong Li 0001 |
ACM Multimedia | 5 |
| 2025 | Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model SecurityabstractThe rapid advancement of multimodal large language models (MLLMs) has led to breakthroughs in various applications, yet their security remains a critical challenge. One pressing issue involves unsafe image-query pairs-jailbreak inputs specifically designed to bypass security constraints and elicit unintended responses from MLLMs. Compared to general multimodal data, such unsafe inputs are relatively sparse, which limits the diversity and richness of training samples available for developing robust defense models. Meanwhile, existing guardrail-type methods rely on external modules to enforce security constraints but fail to address intrinsic vulnerabilities within MLLMs. Traditional supervised fine-tuning (SFT), on the other hand, often over-refuses harmless inputs, compromising general performance. Given these challenges, we propose Secure Tug-of-War (SecTOW), an innovative iterative defense-attack training method to enhance the security of MLLMs. SecTOW consists of two modules: a defender and an auxiliary attacker, both trained iteratively using reinforcement learning (GRPO). During the iterative process, the attacker identifies security vulnerabilities in the defense model and expands jailbreak data. The expanded data are then used to train the defender, enabling it to address identified security vulnerabilities. We also design reward mechanisms used for GRPO to simplify the use of response labels, reducing dependence on complex generative labels and enabling the efficient use of synthetic data. Additionally, a quality monitoring mechanism is used to mitigate the defender's over-refusal of harmless inputs and ensure the diversity of the jailbreak data generated by the attacker. Experimental results on safety-specific and general benchmarks demonstrate that SecTOW significantly improves security while preserving general performance. Warning: This paper contains offensive and unsafe content. Muzhi Dai, Zhiyuan Zhao 0005, Junyu Gao 0001, Hao Sun 0038, Xuelong Li 0001 |
ACM Multimedia | 4 |
| 2025 | From Captions to Rewards (CaReVL): Leveraging Large Language Model Experts for Enhanced Reward Modeling in Large Vision-Language ModelsabstractAligning large vision-language models (LVLMs) with human preferences is challenging due to the scarcity of fine-grained, high-quality, and multimodal preference data without human annotations. Existing methods relying on direct distillation often struggle with low-confidence data, leading to suboptimal performance. To address this, we propose (CaReVL), a novel method for preference reward modeling by reliably using both high- and low-confidence data. First, a cluster of auxiliary expert models (textual reward models) innovatively leverages image captions as weak supervision signals to filter high-confidence data. The high-confidence data are then used to fine-tune the LVLM. Second, low-confidence data are used to generate diverse preference samples using the fine-tuned LVLM. These samples are then scored and selected to construct reliable chosen-rejected pairs for further training. (CaReVL) achieves performance improvements over traditional distillation-based methods on VL-RewardBench and MLLM-as-a-Judge benchmark, demonstrating its effectiveness. Muzhi Dai, Jiashuo Sun, Zhiyuan Zhao 0005, Junyu Gao 0001, Xuelong Li 0001 |
ACM Multimedia | 6 |
| 2025 | StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic SegmentationabstractMultimodal semantic segmentation shows significant potential for enhancing segmentation accuracy in complex scenes. However, current methods often incorporate specialized feature fusion modules tailored to specific modalities, thereby restricting input flexibility and increasing the number of training parameters. To address these challenges, we propose StitchFusion, a straightforward yet effective modal fusion framework that integrates large-scale pre-trained models directly as encoders and feature fusers. This approach facilitates comprehensive multi-modal and multi-scale feature fusion, accommodating any visual modal inputs. Specifically, our framework achieves modal integration during encoding by sharing multi-modal visual information. To enhance information exchange across modalities, we introduce a multi-directional Modality Adapter module (MoA) to enable cross-modal information transfer during encoding. By leveraging MoA to propagate multi-scale information across pre-trained encoders during the encoding process, StitchFusion achieves multi-modal visual information integration during encoding. Extensive comparative experiments demonstrate that our model achieves state-of-the-art performance on four multi-modal segmentation datasets with minimal additional parameters. Furthermore, the experimental integration of MoA with existing Feature Fusion Modules (FFMs) highlights their complementary nature. Our anonymous code is https://anonymous.4open.science/r/StitchFusion_V2-E777 Bingyu Li 0002, Da Zhang 0010, Zhiyuan Zhao 0005, Junyu Gao 0001, Xuelong Li 0001 |
ACM Multimedia | 4 |
| 2025 | SVGen: Interpretable Vector Graphics Generation with Large Language ModelsabstractScalable Vector Graphics (SVG) has become an indispensable technology in front-end development and UI/UX design, due to its inherent advantages in scalability, editability, and rendering efficiency. In the creation of vector graphics, while expressing creative concepts is straightforward, translating them into precise digital artworks is often challenging and time-consuming. To overcome this technical bottleneck and achieve intelligent conversion from concept to final product, we have constructed SVG-1M, a large-scale dataset of high-quality SVG samples with paired textual descriptions. Through innovative data augmentation and annotation processes, we built precisely aligned ''Text instruction-SVG code'' training pairs, with a subset enhanced by Chain-of-Thought (CoT) annotations. This provides rich semantic supervision signals for model learning. Based on this dataset, we propose SVGen, an end-to-end generative model capable of directly converting natural language descriptions into SVG code. This design addresses the challenges of generating semantically accurate vector graphics while preserving complete structural information. We explored various training strategies and introduced a progressive curriculum learning approach, optimized with reinforcement learning algorithms. Notably, this study innovatively applies the CoT paradigm to vector graphics generation, effectively enhancing both the accuracy and interpretability of SVG synthesis. Experimental validation demonstrates that SVGen exhibits significant advantages over general large models in terms of SVG generation quality, while also surpassing optimization-based rendering methods in generation efficiency. The proposed method enables intelligent conversion between natural language and vector graphics, enabling novel workflows like real-time AI-assisted design iteration. Code, model, and data is released at: https://github.com/gitcat-404/SVGen Zhiyuan Zhao 0005, Yuandong Liu, Da Zhang 0010, Junyu Gao 0001, Hao Sun 0038, Xuelong Li 0001 |
ACM Multimedia | 5 |
| 2025 | PUO-Bench: A Panel Understanding and Operation Benchmark with A Privacy-Preserving FrameworkabstractRecent advancements in Vision-Language Models (VLMs) have enabled GUI agents to leverage visual features for interface understanding and operation in the digital world. However, limited research has addressed the interpretation and interaction with control panels in real-world settings. To bridge this gap, we propose the Panel Understanding and Operation (PUO) benchmark, comprising annotated panel images from appliances and associated vision-language instruction pairs. Experimental results on the benchmark demonstrate significant performance disparities between zero-shot and fine-tuned VLMs, revealing the lack of PUO-specific capabilities in existing language models. Furthermore, we introduce a Privacy-Preserving Framework (PPF) to address privacy concerns in cloud-based panel parsing and reasoning. PPF employs a dual-stage architecture, performing panel understanding on edge devices while delegating complex reasoning to cloud-based LLMs. Although this design introduces a performance trade-off due to edge model limitations, it eliminates the transmission of raw visual data, thereby mitigating privacy risks. Overall, this work provides foundational resources and methodologies for advancing interactive human-machine systems and robotic field in panel-centric applications. Wei Lin 0018, Yiwei Zhou, Zhiyuan Zhao 0005, Junyu Gao 0001, Antoni B. Chan, Xuelong Li 0001 |
NeurIPS | 6 |
| 2025 | Building extraction from remote sensing images with deep learning: A survey on vision techniques
Yuan Yuan 0001, Junyu Gao 0001 |
Comput. Vis. Image Underst. | 3 |
| 2025 | H3T: Hierarchical Transferable Transformer with TokenMix for Unsupervised Domain Adaptation
Yihua Ren, Junyu Gao 0001, Yuan Yuan 0001 |
Expert Syst. Appl. | 2 |
| 2025 | Memory-enhanced hierarchical transformer for video paragraph captioning
Benhui Zhang 0001, Junyu Gao 0001, Yuan Yuan 0001 |
Neurocomputing | 2 |
| 2025 | U3M: Unbiased Multiscale Modal Fusion Model for Multimodal Semantic Segmentation
Bingyu Li 0002, Da Zhang 0010, Zhiyuan Zhao 0005, Junyu Gao 0001, Xuelong Li 0001 |
Pattern Recognit. | 4 |
| 2025 | Distance-aware network for physical-world object distribution estimation and counting
Yuan Yuan 0001, Haojie Guo, Junyu Gao 0001 |
Pattern Recognit. | 3 |
| 2025 | Combining SAM With Limited Data for Change Detection in Remote SensingabstractChange detection is a critical task in the remote sensing image (RSI) analysis, widely used in fields such as land cover change and urban planning. With the introduction of foundational models like SAM in computer vision (CV) tasks, their advantages in zero-shot and interactive segmentation have enabled rapid application across diverse visual scenarios. Current research in change detection focuses on designing learnable plug-in modules and fine-tuning foundational models using large annotated data. However, constructing comprehensive datasets and designing effective additional modules pose significant challenges, leading to high costs. To address these issues, we propose a model named Meta-CD for remote sensing change detection (RSCD) with limited data. By introducing a simple fine-tuning module, this model is trained on limited datasets and quickly adapts to change detection tasks. Specifically, we integrate an additional CNN as an adapter with the foundational model FastSAM. Initially, we freeze the parameters of FastSAM and train only the parameters of the introduced adapter and decoder to generate change confidence maps. Subsequently, to enhance the quality of change detection, we introduce a novel pixel-level binarization module that learns the threshold for each pixel in the original image. This module combines the thresholds with the confidence maps to output binary change detection maps, filtering out invalid change pixels. Experimental results demonstrate that our method outperforms other competing approaches on limited datasets and has great zero-shot learning ability. Our code is available at Meta-CD. Junyu Gao 0001, Da Zhang 0010, Lichen Ning, Zhiyuan Zhao 0005, Xuelong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Dual-View Classifier Evolution for Generalized Remote Sensing Few-Shot SegmentationabstractAdvancements in few-shot segmentation (FSS) for remote sensing images have significantly improved the ability to binarization parse novel classes using only a few supports. Generalized few-shot segmentation (GFSS), a challenging and practical task, has recently attracted research attention. It involves recognizing base and novel classes while segmenting multiple categories in a query. Most GFSS methods adopt a two-stage approach: base classifier training and novel classifier registering. However, they encounter two key challenges: the data scale disparity between base and novel classes and significant intraclass variation in remote sensing images. In this article, we present a dual-view classifier evolution (DiCE) method. Our approach utilizes the well-trained base classifier to allocate attention within the novel classifier, effectively addressing the disparities between the two. Simultaneously, it fosters context-driven interactions between the query and the classifier, tailoring sample-specific classifiers to mitigate intraclass variations. Furthermore, we propose a binocular hybrid training (BHT) mechanism that integrates normal base training with episodic training, endowing the model with the ability to adapt to few-shot tasks. Extensive experiments on the iSAID-$5^{i}$dataset demonstrate the superior performance of DiCE. Yuyu Jia, Junyu Gao 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Entity-Guided Attention Twisting Network for Referring Remote Sensing Image SegmentationabstractReferring Remote Sensing Image Segmentation (RRSIS) aims to establish pixel-level interpretation of specific regions queried by textual expressions, bridging textual semantics and intelligent analysis of remote sensing imagery. In contrast to natural scenarios, the intricate backgrounds in remote sensing scenarios result in low target-background contrast, often leading to semantic dispersion in segmented regions. Furthermore, conventional cross-attention-based referring image segmentation (RIS) methods struggle to bridge the modal gap, hindering fine-grained alignment between linguistic descriptions and geographical features. To overcome these challenges, we present a pioneering Entity-Guided Attention Twisting Network (Enti-TwistNet) for RRSIS. Our framework first introduces a SAM-inspired Entity Guidance (SEG) module that extracts spatially constrained entity prompts through a self-reasoning mask generation mechanism, constructing a comprehensive entity-visual-text tri-modal information cube. Subsequently, during cross-modal interaction, we propose a Dual-phase Attention-Twisting (DAT) mechanism: (1) initially sequential channel-wise scanning to facilitate cross-modal semantic propagation; (2) Subsequently, twist attention to the spatial dimension, integrating entity guidance to enhance the representation of irregular geographic boundaries. Extensive experiments on two widely used benchmarks, RefSegRS and RRSIS-D, demonstrate that Enti-TwistNet achieves significant performance improvements over existing state-of-the-art models. Yuyu Jia, Junyu Gao 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Dual-Stage Prior-Driven Diffusion Model for Remote Sensing Spectral Super-ResolutionabstractSpectral super-resolution (SSR) is a key technology for generating high spatial resolution hyperspectral images (HSIs). However, deep learning approaches for SSR, especially generative models like diffusion, often rely heavily on large training datasets. Furthermore, their stochastic generation process can compromise the precise spectral fidelity required in remote sensing. To address this limitation, we propose a dual-stage prior-driven diffusion model (DPDM), for SSR tasks. DPDM comprises two modules: the prior-driven diffusion module (PDM) and the spectral refinement module (SRM). PDM replaces the conventional pure noise input with a structured prior, which we term the prior-informed noise (PIN). This PIN is deterministically generated by projecting the input multispectral image (MSI) onto a spectral basis, which is extracted from derived from the spectral response function. By initializing the reverse process with this information-rich starting point, our model significantly reduces its dependence on large training datasets and inherently enforces spectral consistency. The SRM is subsequently introduced to specifically target and correct residual artifacts and coarse features from the initial stage. Employing a hierarchical multi-scale architecture and a single-sample optimization framework, the SRM meticulously restores fine-grained details while suppressing noise in the PDM’s output. By integrating these two modules, DPDM progressively enhances both spatial and spectral fidelity. Extensive experimental results demonstrate that DPDM achieves competitive performance in SSR tasks. Zengyi Li, Junyu Gao 0001, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Cross-Resolution Change Detection in Remote Sensing via Unequal Relationships From a Frequency PerspectiveabstractCross-resolution change detection (CRCD) identifies changes between bitemporal images with different resolutions, which provide better adaptability to real-world applications than conventional change detection (CD). Existing CRCD methods first align the resolution of different temporal images, then employ the Siamese network. Experiments conducted from a frequency-domain perspective validate that resize operations disrupt the data distribution and result in performance degradation of the Siamese network. Further experiments reveal the resolution-invariant temporal and spatial unequal relationship between bitemporal images. Specifically, spatial specificity information within a specific temporal domain is more critical for CRCD, i.e., high-frequency components in a specific temporal domain are closely related to change label. And this unequal relationship exhibits invariance in resolution. On this basis, we propose the Fourier and wavelet transform-based inequality Siamese network (FWISN) to address the performance degradation observed in Siamese networks on CRCD, leveraging the inequality between bitemporal images to improve network performance. FWISN includes a frequency reconstruction (FRC) stage, in which high-frequency components of a given temporal-domain image are extracted and reconstructed using our proposed high-frequency attention (HFA) module. We further propose the wavelet transform-based frequency learning block (WFB), which enhances high-frequency features and is integrated into both the encoder (WFB-E) and the decoder (WFB-D) of the network. The experiments demonstrate state-of-the-art performance, compared with methods specifically designed for cross-resolution tasks, FWISN achieving$F1$/intersection over union (IoU) improvements of 2.15/3.51, 2.50/4.59, and 1.93/1.73 on the LEVIR-CD ($4 \times $), SV-CD ($8 \times $), and DE-CD ($3.3\times $) tests, respectively. Furthermore, in the continuous CRCD task, FWISN achieves$F1$/IoU improvements of 8.39/11.24 on the LEVIR-CD ($8 \times $) test. Our code will be public onhttps://github.com/blacksheep182/NN_FWISN Lichen Ning, Qi Wang 0009, Junyu Gao 0001, Xuelong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Embedding Generalized Semantic Knowledge Into Few-Shot Remote Sensing SegmentationabstractFew-shot segmentation (FSS) for remote sensing (RS) imagery leverages supporting information from limited annotated samples to achieve query segmentation of novel classes. Previous efforts are dedicated to mining segmentation-guiding visual cues from a constrained set of support samples. However, they still struggle to address the pronounced intra-class differences in RS images, as sparse visual cues make it challenging to establish robust class-specific representations. In this article, we propose a holistic semantic embedding (HSE) approach that effectively harnesses general semantic knowledge, i.e., class description (CD) embeddings. Instead of the naive combination of CD embeddings and visual features for segmentation decoding, we investigate embedding the general semantic knowledge during the feature extraction stage. Specifically, in HSE, a spatial dense interaction (SDI) module allows the interaction of visual support features with CD embeddings along the spatial dimension via self-attention. Furthermore, a global content modulation (GCM) module efficiently augments the global information of the target category in both support and query features, thanks to the transformative fusion of visual features and CD embeddings. These two components holistically synergize CD embeddings and visual cues, constructing a robust class-specific representation. Through extensive experiments on the standard FSS benchmark, the proposed HSE approach demonstrates superior performance compared to peer work, setting a new state-of-the-art. Qi Wang 0009, Yuyu Jia, Wei Huang 0068, Junyu Gao 0001, Qiang Li 0042 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Real-Time Text Detection With Similar Mask in Traffic, Industrial, and Natural ScenesabstractTexts on the intelligent transportation scene include mass information. Fully harnessing this information is one of the critical drivers for advancing intelligent transportation. Unlike the general scene, detecting text in transportation has extra demand, such as a fast inference speed, except for high accuracy. Most existing real-time text detection methods are based on the shrink mask, which loses some geometry semantic information and needs complex post-processing. In addition, the previous method usually focuses on correct output, which ignores feature correction and lacks guidance during the intermediate process. To this end, we propose an efficient multi-scene text detector that contains an effective text representation similar mask (SM) and a feature correction module (FCM). Unlike previous methods, the former aims to preserve the geometric information of the instances as much as possible. Its post-progressing saves 50% of the time, accurately and efficiently reconstructing text contours. The latter encourages false positive features to move away from the positive feature center, optimizing the predictions from the feature level. Some ablation studies demonstrate the efficiency of the SM and the effectiveness of the FCM. Moreover, the deficiency of existing traffic datasets (such as the low-quality annotation or closed source data unavailability) motivated us to collect and annotate a traffic text dataset, which introduces motion blur. In addition, to validate the scene robustness of the SM-Net, we conduct experiments on traffic, industrial, and natural scene datasets. Extensive experiments verify it achieves (SOTA) performance on several benchmarks. The code and dataset are available at:https://github.com/fengmulin/SMNet. Xu Han 0019, Junyu Gao 0001, Chuang Yang 0003, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | SignEye: Traffic Sign Interpretation From Vehicle First-Person ViewabstractTraffic signs play a key role in assisting autonomous driving systems (ADS) by enabling the assessment of vehicle behavior in compliance with traffic regulations and providing navigation instructions. However, current works are limited to basic sign understanding without considering the egocentric vehicle’s spatial position, which fails to support further regulation assessment and direction navigation. Following the above issues, we introduce a new task: traffic sign interpretation from the vehicle’s first-person view, referred to asTSI-FPV. Meanwhile, we develop a traffic guidance assistant (TGA) scenario application to re-explore the role of traffic signs in ADS as a complement to popular autonomous technologies (such as obstacle perception). Notably, TGA is not a replacement for electronic map navigation; rather, TGA can be an automatic tool for updating it and complementing it in situations such as offline conditions or temporary sign adjustments. Lastly, a spatial and semantic logic-aware stepwise reasoning pipeline (SignEye) is constructed to achieve the TSI-FPV and TGA, and an application-specific dataset (Traffic-CN) is built. Experiments show that TSI-FPV and TGA are achievable via our SignEye trained on Traffic-CN. The results also demonstrate that the TGA can provide complementary information to ADS beyond existing popular autonomous technologies. Chuang Yang 0003, Xu Han 0019, Tao Han 0002, Yuejiao Su, Junyu Gao 0001, Hongyuan Zhang 0001, Yi Wang 0068, Lap-Pui Chau |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Focus Entirety and Perceive Environment for Arbitrary-Shaped Text DetectionabstractDue to the diversity of scene text in aspects such as font, color, shape, and size, accurately and efficiently detecting text is still a formidable challenge. Among the various detection approaches, segmentation-based approaches have emerged as prominent contenders owing to their flexible pixel-level predictions. However, these methods typically model text instances in a bottom-up manner, which is highly susceptible to noise. In addition, the prediction of pixels is isolated without introducing pixel-feature interaction, which also influences the detection performance. To alleviate these problems, we propose a multi-information level arbitrary-shaped text detector consisting of a focus entirety module (FEM) and a perceive environment module (PEM). The former extracts instance-level features and adopts a top-down scheme to model texts to reduce the influence of noises. Specifically, it assigns consistent entirety information to pixels within the same instance to improve their cohesion. In addition, it emphasizes the scale information, enabling the model to distinguish varying scale texts effectively. The latter extracts region-level information and encourages the model to focus on the distribution of positive samples in the vicinity of a pixel, which perceives environment information. It treats the kernel pixels as positive samples and helps the model differentiate text and kernel features. Extensive experiments demonstrate the FEM's ability to efficiently support the model in handling different scale texts and confirm the PEM can assist in perceiving pixels more accurately by focusing on pixel vicinities. Comparisons show the proposed model outperforms existing state-of-the-art approaches on four public datasets. Xu Han 0019, Junyu Gao 0001, Chuang Yang 0003, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Multim. | 2 |
| 2025 | Spotlight Text Detector: Spotlight on Candidate Regions Like a CameraabstractThe irregular contour representation is one of the tough challenges in scene text detection. Although segmentation-based methods have achieved significant progress with the help of flexible pixel prediction, the overlap of geographically close texts hinders detecting them separately. To alleviate this problem, some shrink-based methods predict text kernels and expand them to restructure texts. However, the text kernel is an artificial object with incomplete semantic features that are prone to incorrect or missing detection. In addition, different from the general objects, the geometry features (aspect ratio, scale, and shape) of scene texts vary significantly, which makes it difficult to detect them accurately. To consider the above problems, we propose an effective spotlight text detector (STD), which consists of a spotlight calibration module (SCM) and a multivariate information extraction module (MIEM). The former concentrates efforts on the candidate kernel, like a camera focus on the target. It obtains candidate features through a mapping filter and calibrates them precisely to eliminate some false positive samples. The latter designs different shape schemes to explore multiple geometric features for scene texts. It helps extract various spatial relationships to improve the model's ability to recognize kernel regions. Ablation studies prove the effectiveness of the designed SCM and MIEM. Extensive experiments verify that our STD is superior to existing state-of-the-art methods on various datasets, including ICDAR2015, CTW1500, MSRA-TD500, and Total-Text. Xu Han 0019, Junyu Gao 0001, Chuang Yang 0003, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Multim. | 2 |
| 2025 | Like Humans to Few-Shot Learning Through Knowledge Permeation of Visual and LanguageabstractFew-shot learning aims to generalize the recognizer from seen categories to an entirely novel scenario. With only a few support samples, several advanced methods initially introduce class names as prior knowledge for identifying novel classes. However, obstacles still impede achieving a comprehensive understanding of how to harness the mutual advantages of visual and textual knowledge. In this paper, we set out to fill this gap via a coherent Bidirectional Knowledge Permeation strategy called BiKop, which is grounded in human intuition: a class name description offers a moregeneralrepresentation, whereas an image captures thespecificityof individuals. BiKop primarily establishes a hierarchical joint general-specific representation through bidirectional knowledge permeation. On the other hand, considering the bias of joint representation towards the base set, we disentangle base-class-relevant semantics during training, thereby alleviating the suppression of potential novel-class-relevant information. Experiments on four challenging benchmarks demonstrate the remarkable superiority of BiKop, particularly outperforming previous methods by a substantial margin in the 1-shot setting (improving the accuracy by 7.58% onminiImageNet). Yuyu Jia, Junyu Gao 0001, Qiang Li 0042, Qi Wang 0009 |
IEEE Trans. Multim. | 3 |
| 2024 | Combating Data Imbalances in Federated Semi-supervised Learning with Dual RegulatorsabstractFederated learning has become a popular method to learn from decentralized heterogeneous data. Federated semi-supervised learning (FSSL) emerges to train models from a small fraction of labeled data due to label scarcity on decentralized clients. Existing FSSL methods assume independent and identically distributed (IID) labeled data across clients and consistent class distribution between labeled and unlabeled data within a client. This work studies a more practical and challenging scenario of FSSL, where data distribution is different not only across clients but also within a client between labeled and unlabeled data. To address this challenge, we propose a novel FSSL framework with dual regulators, FedDure. FedDure lifts the previous assumption with a coarse-grained regulator (C-reg) and a fine-grained regulator (F-reg): C-reg regularizes the updating of the local model by tracking the learning effect on labeled data distribution; F-reg learns an adaptive weighting scheme tailored for unlabeled instances in each client. We further formulate the client model training as bi-level optimization that adaptively optimizes the model in the client with two regulators. Theoretically, we show the convergence guarantee of the dual regulators. Empirically, we demonstrate that FedDure is superior to the existing methods across a wide range of settings, notably by more than 11% on CIFAR-10 and CINIC-10 datasets. Sikai Bai, Weiming Zhuang, Jie Zhang 0076, Shuai Yi, Junyu Gao 0001 |
AAAI | 9 |
| 2024 | A Descriptive Basketball Highlight Dataset for Automatic Commentary GenerationabstractThe emergence of video captioning makes it possible to automatically generate natural language description for a given video. However, generating detailed video descriptions that incorporate domain-specific information remains an unsolved challenge, holding significant research and application value, particularly in domains such as sports commentary generation. Moreover, sports event commentary goes beyond being a mere game report, it involves entertaining, metaphorical, and emotional descriptions. To promote the field of sports commentary automatic generation, in this paper, we introduce a novel dataset, the Basketball Highlight Commentary (BH-Commentary), comprising approximately 4K basketball highlight videos with groundtruth commentaries from professional commentators. In addition, we propose an end-to-end framework as a benchmark for basketball highlight commentary generation task, in which a lightweight and effective prompt strategy is designed to enhance alignment fusion among visual and textual features. Experimental results on the BH-Commentary dataset demonstrate the validity of the dataset and the effectiveness of the proposed benchmark for sports highlight commentary generation. Benhui Zhang 0001, Junyu Gao 0001, Yuan Yuan 0001 |
ACM Multimedia | 2 |
| 2024 | Multivariate time series classification with crucial timestamps guidance
Da Zhang 0010, Junyu Gao 0001, Xuelong Li 0001 |
Expert Syst. Appl. | 2 |
| 2024 | RRTrN: A lightweight and effective backbone for scene text recognition
Junyu Gao 0001, Yuan Yuan 0001, Qi Wang 0009 |
Expert Syst. Appl. | 2 |
| 2024 | Audio-visual representation learning for anomaly events detection in crowds
Junyu Gao 0001, Hao Yang 0046, Maoguo Gong, Xuelong Li 0001 |
Neurocomputing | 1 |
| 2024 | Center-enhanced video captioning model with multimodal semantic alignment
Benhui Zhang 0001, Junyu Gao 0001, Yuan Yuan 0001 |
Neural Networks | 2 |
| 2024 | SSIR: Spatial shuffle multi-head self-attention for Single Image Super-ResolutionabstractBenefiting from the development of deep convolutional neural networks , CNN-based single-image super-resolution methods have achieved remarkable reconstruction results. However, the limited perceptual field of the convolutional kernel and the use of static weights in the inference process limit the performance of CNN-based methods. Recently, a few vision transformer-based image super-resolution methods have achieved excellent performance compared to CNN-based methods. These methods contain many parameters and require vast amounts of GPU memory for training. In this paper, we propose a spatial shuffle multi-head self-attention for single-image super-resolution that can significantly model long-range pixel dependencies without additional computational consumption. A local perception module is also proposed to combine convolutional neural networks’ local connectivity and translational invariance. Reconstruction results on five popular benchmarks show that the proposed method outperforms existing methods in both reconstruction accuracy and visual performance. The proposed method matches the performance of transformed-based methods but requires an inferior number of transformer blocks, which reduces the number of parameters by 40%, GPU memory by 30%, and inference time by 30% compared to transformer-based methods. Junyu Gao 0001, Donghu Deng, Xuelong Li 0001 |
Pattern Recognit. | 2 |
| 2024 | NWPU-MOC: A Benchmark for Fine-Grained Multicategory Object Counting in Aerial ImagesabstractObject counting is a hot topic in computer vision, which aims to estimate the number of objects in a given image. However, most methods only count objects of a single category for an image, which cannot be applied to scenes that need to count objects with multiple categories simultaneously, especially in aerial scenes. To this end, this paper introduces a Multi-category Object Counting (MOC) task to estimate the numbers of different objects (cars, buildings, ships,etc.) in an aerial image. Considering the absence of a dataset for this task, a large-scale Dataset (NWPU-MOC) is collected, consisting of 3,416 scenes with a resolution of 1024 × 1024 pixels, and well-annotated using 14 fine-grained object categories. Besides, each scene contains RGB and Near Infrared (NIR) images, of which the NIR spectrum can provide richer characterization information compared with only the RGB spectrum. Based on NWPU-MOC, the paper presents a multi-spectrum, multi-category object counting framework, which employs a dual-attention module to fuse the features of RGB and NIR and subsequently regress multi-channel density maps corresponding to each object category. In addition, to modeling the dependency between different channels in the density map with each object category, a spatial contrast loss is designed as a penalty for overlapping predictions at the same spatial position. Experimental results demonstrate that the proposed method achieves state-of-the-art performance compared with some mainstream counting algorithms. The dataset, code and models are publicly available at https://github.com/lyongo/NWPU-MOC. Junyu Gao 0001, Xuelong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Balanced Density Regression Network for Remote Sensing Object CountingabstractCounting objects in remote sensing is crucial for analyzing their distribution in images. Compared to surveillance perspectives, counting dense objects in remote sensing images is more challenging due to the smaller sizes of these targets. Recently, many methods utilize Gaussian convolution regression to estimate the count of dense objects in remote sensing images. However, most methods ignore the issue of regression imbalance inherent in Gaussian distribution, which is caused by the numerical differences in the center and edge regions. To tackle this challenge, we propose a Balanced Density Regression Network (BDRNet) to mitigate regression inaccuracies in Gaussian distributions due to numerical variances. Different from other methods, we divide the regression problem into two steps: first focusing on the regions of interest, then achieving precise regression. BDRNet consists of an Adaptive Kernel Weighting Attention (AKWA) mechanism and a Pixel-wise Occupancy Prediction (PwOE) module. Firstly, AKWA is designed to acquire accurate semantic feature information, which is obtained by learning the weights of dilated convolutions with different sizes of receptive fields. Secondly, the PwOE module applies Gaussian position embeddings to point labels to constrain the network to focus on the object region without increasing annotation cost. Finally, the integration of pixel-wise occupancy prediction features and kernel weighting features forms multi-layer cross-attention mechanisms, facilitating channel-level feature interaction and improving density regression predictions. Thus, the center and edge regions of the Gaussian kernel are treated equally, and the regression is balanced. Additionally, Extensive experiments on diverse datasets validate the effectiveness of the method, resulting in preferable performance. The code is available at: https://github.com/HotChieh/BDRNet. Haojie Guo, Junyu Gao 0001, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Contrastive Tokens and Label Activation for Remote Sensing Weakly Supervised Semantic SegmentationabstractIn recent years, there has been remarkable progress in Weakly Supervised Semantic Segmentation (WSSS), with Vision Transformer (ViT) architectures emerging as a natural fit for such tasks due to their inherent ability to leverage global attention for comprehensive object information perception. However, directly applying ViT to WSSS tasks can introduce challenges. The characteristics of ViT can lead to an over-smoothing problem, particularly in dense scenes of remote sensing images, significantly compromising the effectiveness of Class Activation Maps (CAM) and posing challenges for segmentation. Moreover, existing methods often adopt multi-stage strategies, adding complexity and reducing training efficiency. To overcome these challenges, a comprehensive framework CTFA (Contrastive Token and Foreground Activation) based on the ViT architecture for WSSS of remote sensing images is presented. Our proposed method includes a Contrastive Token Learning Module (CTLM), incorporating both patch-wise and class-wise token learning to enhance model performance. In patch-wise learning, we leverage the semantic diversity preserved in intermediate layers of ViT and derive a relation matrix from these layers and employ it to supervise the final output tokens, thereby improving the quality of CAM. In class-wise learning, we ensure the consistency of representation between global and local tokens, revealing more entire object regions. Additionally, by activating foreground features in the generated pseudo label using a dual-branch decoder, we further promote the improvement of CAM generation. Our approach demonstrates outstanding results across three well-established datasets, providing a more efficient and streamlined solution for WSSS. Code will be available at: https://github.com/ZaiyiHu/CTFA. Zaiyi Hu, Junyu Gao 0001, Yuan Yuan 0001, Xuelong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Enhancing Unimodal Features Matters: A Multimodal Framework for Building ExtractionabstractIn recent years, deep learning and multi-modal data have substantially propelled the development of building extraction models. However, prevailing multi-modal methods are difficult to cope with two challenges: 1) modal laziness: the training error is minimized before the model has learned extensive uni-modal patterns; 2) modal imbalance: the backpropagation process is easily dominated by a certain modality. As a result, the uni-modal features learning is insufficient, leading to limited performance of the model when dealing with the intricate foreground and background contexts surrounding the buildings. In this paper, we deal with this problem from the perspective of algorithm and model evaluation. At the algorithmic level, we propose a Uni-modal Feature Enhancement (UFE) framework. Specifically, UFE is model-agnostic, comprising two distinct components: Adaptive Gradient Enhancement (AGE) for modal laziness and Consistency Constraint Loss (CCL) for modal imbalance. AGE dynamically modulates the original gradient by monitoring the representation effects of uni-modal features and multi-modal fusion features. CCL imposes mutual constraints on diverse modal branches at the semantic level to reconcile the optimization process. At the model evaluation level, a new metric, named Uni-modal Utilization Ratio (UUR), is presented to assess models through the learning efficacy of uni-modal features. The experimental results including the variants of UUR on two building extraction datasets demonstrate a substantial performance improvement by UFE. Moreover, UFE also exhibits its adaptability when integrated with various model components and its generalization on other multi-modal image-related tasks. Junyu Gao 0001, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Alignment and Fusion Using Distinct Sensor Data for Multimodal Aerial Scene ClassificationabstractMultimodal sensors offer a wealth of rich and diverse data, which is helpful for classify similar and complex aerial scenes. However, the heterogeneity of the data collected from different sensors brings a great challenge for alignment and fusion. For this, we present a multimodal aerial scene classification approach for extracting distinct modal information representations, realizing alignment and fusion of semantic information at both the data and feature levels. Firstly, an Adaptive Zero-Crossing Rate (AZCR) module is proposed to convert the sequential data into images, achieving alignment at the data level. This module is proficient at extracting temporal and frequency domain features from sequential data through adaptive parameter adjustments. Secondly, we propose a Multi-Modal Alignment and Fusion (MMAF) module to facilitate the alignment and fusion of distinct data, thereby achieving comprehensive modality integration at the feature level. Finally, the multimodal alignment loss function is designed to assess the alignment outcomes and constrain the training process. Our approach has been proven effective in accurately classifying aerial scenes, as demonstrated by the results of our experiments on two public datasets. The proposed method achieves 81.32% and 59.80% F1 score on the ADVANCE and URFC datasets. Junyu Gao 0001, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Integrating SAM With Feature Interaction for Remote Sensing Change DetectionabstractVision foundation models (VFMs) have rapidly gained application across various visual scenarios due to their robust universality and generalization capabilities. However, when directly applied to remote sensing images (RSIs), their performance often falls short owing to the unique inherent imaging characteristics. Moreover, these models typically suffer from inadequate feature extraction capabilities and unclear boundary detection because of the lack of specialized knowledge in the remote sensing (RS) field. To ameliorate these issues, we propose SFCD-Net, a novel network integratingSAM withfeature interaction for RSchangedetection. To be specific, we first introduce a parameter-efficient fine-tuning (PEFT) method that allows the model to learn domain-specific knowledge, thereby enhancing its fine-grained feature extraction capability. Second, an innovative bitemporal feature interaction (BFI) module is designed to improve the model’s sensitivity to changes. Finally, we use the boundary loss function (BLF) to enhance the model’s ability to process boundary details, thereby improving its performance in recognizing boundaries and small targets. Through a series of ablation studies and comparative experiments, we demonstrate that the proposed SFCD-Net significantly improves model adaptability in RS tasks under limited computational resources, outperforming existing models. Da Zhang 0010, Lichen Ning, Zhiyuan Zhao 0005, Junyu Gao 0001, Xuelong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Single-Stream Extractor Network With Contrastive Pre-Training for Remote-Sensing Change CaptioningabstractRemote sensing (RS) image change captioning is a visual semantic understanding task that has received increasing attention. The change captioning methods are required to understand the visual information of the images and capture the most significant difference between them, then describe it in natural language. Most existing methods mainly focus on improving the difference feature encoder or language decoder, while ignoring the visual feature extractor. The current feature extractors suffer from several issues, including 1) domain gap between pre-training on single temporal natural images and downstream bi-temporal RS task, 2) limited difference feature modeling in the implicit single-stream network, and 3) high computational costs caused by extracting features for each temporal phase image under the dual-stream extractor. To address these issues, we propose a Single-stream Extractor Network (SEN). It consists of a single-stream extractor pre-trained on bi-temporal RS images using contrastive learning to mitigate the domain gap and high computational cost. Additionally, to improve feature modeling for difference information, we propose a shallow feature embedding (SFE) module and a cross attention guided difference (CAGD) module, which enhance the representation of temporal features and extract the difference features explicitly. Extensive experiments and visualizations demonstrate the effectiveness and advanced performance of SEN. The code and model weights are available at https://github.com/mrazhou/SEN. Junyu Gao 0001, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | FF-LPD: A Real-Time Frame-by-Frame License Plate Detector With Knowledge Distillation and Feature PropagationabstractWith the increasing availability of cameras in vehicles, obtaining license plate (LP) information via on-board cameras has become feasible in traffic scenarios. LPs play a pivotal role in vehicle identification, making automatic LP detection (ALPD) a crucial area within traffic analysis. Recent advancements in deep learning have spurred a surge of studies in ALPD. However, the computational limitations of on-board devices hinder the performance of real-time ALPD systems for moving vehicles. Therefore, we propose a real-time frame-by-frame LP detector focusing on real-time accurate LP detection. Specifically, video frames are categorized into keyframes and non-keyframes. Keyframes are processed by a deeper network (high-level stream), while non-keyframes are handled by a lightweight network (low-level stream), significantly enhancing efficiency. To achieve accurate detection, we design a knowledge distillation strategy to boost the performance of low-level stream and a feature propagation method to introduce the temporal clues in video LP detection. Our contributions are: (1) A real-time frame-by-frame LP detector for video LP detection is proposed, achieving a competitive performance with popular one-stage LP detectors. (2) A simple feature-based knowledge distillation strategy is introduced to improve the low-level stream performance. (3) A spatial-temporal attention feature propagation method is designed to refine the features from non-keyframes guided by the memory features from keyframes, leveraging the inherent temporal correlation in videos. The ablation studies show the effectiveness of knowledge distillation strategy and feature propagation method. Haoxuan Ding, Junyu Gao 0001, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Image Process. | 2 |
| 2024 | An End-to-End Contrastive License Plate DetectorabstractAs a unique identity of vehicle, License Plate (LP) facilitates the intelligent transportation in many fields, such as traffic enforcement, intelligent transportation dispatching, etc. Recently, the LP detectors are trained by supervised learning which is directly guided by manual annotations and lacks the use of visual knowledge in image content, limiting the further development of detection performance. Inspired by the contrast and comparison in perception of human beings, a contrastive learning method is introduced into license plate detection task and we propose an end-to-end Contrastive License Plate Detector (CLPD). In CLPD, a special contrastive triad for contrastive learning is designed which aims to decouple the foregrounds and backgrounds. Based on this triad, a contrastive learning branch is introduced into the license plate detection pipeline to prompt the feature expression ability of backbone and extracting more discriminative features for detection. This contrastive learning branch is jointly trained with supervised learning branch for detection and it is only used in training, keeping the efficiency in inference. The experiment results show that the proposed CLPD improves the detection accuracy compared to baselines and other license plate detectors significantly on three datasets. The ablation studies further explore the potential of CLPD. In addition, the proposed CLPD has generalization to improve the performance on different baselines. And the visualization results in latent space verify our proposed CLPD aggregates features tightly and extracts discriminative features effectively. Haoxuan Ding, Junyu Gao 0001, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Text kernel calculation for arbitrary shape text detection
Xu Han 0019, Junyu Gao 0001, Yuan Yuan 0001, Qi Wang 0009 |
Vis. Comput. | 2 |
| 2023 | NAS-Kernel: Learning Suitable Gaussian Kernel for Remote-Sensing Object CountingabstractThe purpose of object counting is to estimate the number of specific kinds of objects in a given image. In remote sensing imagery, challenges arise in object counting due to issues like scale variations and complex backgrounds. Existing density map-based object counting methods have achieved satisfactory performance in some general scenarios (i.e., crowd counting and vehicle counting) and have become the mainstream methods. These density map-based counting methods use a fixed Gaussian kernel in the density map generation stage, thus they are not well adapted to the challenges such as scale variations present in remote sensing scenes. In this letter, we propose to use the strategy of neural architecture search (NAS-Kernel) to select appropriate Gaussian kernels corresponding to objects of different scales in the Gaussian density map generation stage. NAS-Kernel is a plug-and-play algorithm that can be used in other density map-based counting methods. In addition, a contextual path aggregation feature fusion strategy is proposed to fuse multi-scale feature information. The ablation experiments verify that the proposed method can significantly improve the performance of baseline. Experimental results on the four sub-datasets of RSOC show that the proposed method achieves state-of-the-art performance. On the Building sub-dataset, the proposed method achieves 18% and 12% lower Mean Absolute Error (MAE) and Root Mean Square Error (RMSE) than the existing methods. Junyu Gao 0001, Xuelong Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | Boosting One-Stage License Plate Detector via Self-Constrained Contrastive AggregationabstractScene Text Detection (STD) has applied in many fields successfully. One of the important applications of STD is License Plate Detection (LPD). As a unique identity of vehicle, License Plate (LP) facilitates the intelligent transportation in many fields, such as traffic enforcement, intelligent transportation dispatching, etc. However, there are many scene texts similar to LPs causing misjudgment of LP detector. To alleviate these disturbances, more discriminative features are necessary. In latent feature space, discriminative features should aggregate into a tight cluster to widen decision boundary. We assume three perspectives about how to aggregate features and boost feature expression. From these assumptions, a special contrastive triad is designed. Then, we propose a Self-Constrained Contrastive Aggregation (SCCA) method to lead the feature aggregation in latent space and boost the feature expression of backbone. The proposed SCCA is jointly trained with supervised learning for detection to improve the detection performance. The experiments show that our proposed SCCA prompts the baseline significantly and exceeds recent LP detectors, reaching 99.7 on both F1-score and AP on UFPR-ALPR dataset. Meanwhile, we compare the self-constrained contrastive learning with vanilla contrastive learning in experiments and visualize their LP features. The results show that our proposed SCCA reaches better performance and verifies our assumptions are reasonable. Haoxuan Ding, Junyu Gao 0001, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | LGNet: Location-Guided Network for Road Extraction From Satellite ImagesabstractRoad connectivity is vital in road extraction for accurate vehicle navigation. However, the segmentation-based methods fail to model the connectivity resulting in broken road segments. Therefore, we propose a Location-Guided Network (LGNet) for promoting connectivity performance in a very effective and efficient way. Specifically, an auxiliary Road Location Prediction (RLP) task is designed to obtain global road connectivity information, which improves the performance of road segmentation. The RLP can predict the location coordinates of the whole roads with row anchors and column anchors. By aggregating the global location context to the segmentation branch with a location-guided decoder (LG-Decoder), the features can finally capture the connectivity of each road segment. Overall, LGNet has the following advantages: 1) The proposed RLP and LCG can plug into any encoder-decoder network and achieve an impressive performance. 2) High computational efficiency. In comparison with the multi-branch method, our proposed LGNet requires about 6× fewer GFLOPs. 3) The superior road connectivity performance. A series of experiments are conducted on two road extraction data sets (SpaceNet and DeepGlobe), confirming the effectiveness of the LGNet. Jingtao Hu, Junyu Gao 0001, Yuan Yuan 0001, Jocelyn Chanussot, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Exploring Hard Samples in Multiview for Few-Shot Remote Sensing Scene ClassificationabstractFew-shot remote sensing scene classification is of high practical value in real situations where data are scarce and annotated costly. The few-shot learner needs to identify new categories with limited examples, and the core issue of this assignment is how to prompt the model to learn transferable knowledge from a large-scale base dataset. Although current approaches based on transfer learning or meta-learning have achieved significant performance on this task, there are still two problems to be addressed: (i) as an essential characteristic of remote sensing images, spatial rotation insensitivity surprisingly remains largely unexplored; (ii) the high distribution uncertainty of hard samples reduces the discriminative power of the model decision boundary. Stimulated by these, we propose a corresponding end-to-end framework termed a Hard Sample Learning (HSL) and Multi-view Integration (MI) Network (HSL-MINet). First, the MI module contains a pretext task introduced to guide the knowledge transfer, and a multiview-attention mechanism used to extract correlational information across different rotation views of images. Second, aiming at increasing the discrimination of the model decision boundary, the HSL module is designed to evaluate and select hard samples via a class-wise adaptive threshold strategy, and then decrease the uncertainty of their feature distributions by a devised triplet loss. Extensive evaluations on NWPU-RESISC45, WHU-RS19, and UCM datasets show that the effectiveness of our HSL-MINet surpasses the former state-of-the-art approaches. Yuyu Jia, Junyu Gao 0001, Wei Huang 0068, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Holistic Mutual Representation Enhancement for Few-Shot Remote Sensing SegmentationabstractFew-shot segmentation endeavors to utilize a minimal amount of annotated samples (support) to guide the segmentation of unseen objects (query). Previous techniques primarily employ asupport-to-queryparadigm, neglecting to sufficiently leverage the mutual representation between query and support images, which leaves models suffering from intra-class variations and background interference in remote sensing images. This paper proposes a Holistic Mutual Representation Enhancement (HMRE) method to bridge these gaps. First, a Dual Activation (DA) module is devised to establish information symmetry between the two branches and forms the foundation for mutual representation enhancement. Subsequently, the holistic mutual enhancement is jointly constructed by the Global Semantic (GS) and Spatial Dense (SD) mutual enhancement modules. In the prediction stage for segmentation, we integrate the enhanced mutual representation into the Mutual-Fusion Decoder to activate the homologous object regions bidirectionally. To expedite the replication of investigation in this task, we further create a corresponding benchmark Flood-3i. The whole dataset is attainable at https://drive.google.com/drive/folders/1FMAKf2sszoFKjq0UrUmSLnJDbwQSpfxR. Extensive experiments on two benchmarks iSAID-5i and Flood-3i demonstrate the superiority of our proposed method, which also sets a new state-of-the-art. Yuyu Jia, Junyu Gao 0001, Wei Huang 0068, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Crowd Localization From Gaussian Mixture Scoped Knowledge and Scoped TeacherabstractCrowd localization is to predict each instance head position in crowd scenarios. Since the distance of pedestrians being to the camera are variant, there exists tremendous gaps among scales of instances within an image, which is called the intrinsic scale shift. The core reason of intrinsic scale shift being one of the most essential issues in crowd localization is that it is ubiquitous in crowd scenes and makes scale distribution chaotic. To this end, the paper concentrates on access to tackle the chaos of the scale distribution incurred by intrinsic scale shift. We propose Gaussian Mixture Scope (GMS) to regularize the chaotic scale distribution. Concretely, the GMS utilizes a Gaussian mixture distribution to adapt to scale distribution and decouples the mixture model into sub-normal distributions to regularize the chaos within the sub-distributions. Then, an alignment is introduced to regularize the chaos among sub-distributions. However, despite that GMS is effective in regularizing the data distribution, it amounts to dislodging the hard samples in training set, which incurs overfitting. We assert that it is blamed on the block of transferring the latent knowledge exploited by GMS from data to model. Therefore, a Scoped Teacher playing a role of bridge in knowledge transform is proposed. What' s more, the consistency regularization is also introduced to implement knowledge transform. To that effect, the further constraints are deployed on Scoped Teacher to derive feature consistence between teacher and student end. With proposed GMS and Scoped Teacher implemented on four mainstream datasets of crowd localization, the extensive experiments demonstrate the superiority of our work. Moreover, comparing with existing crowd locators, our work achieves state-of-the-art via F1-measure comprehensively on four datasets. Junyu Gao 0001, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Image Process. | 2 |
| 2023 | Domain-Adaptive Crowd Counting via High-Quality Image Translation and Density ReconstructionabstractRecently, crowd counting using supervised learning achieves a remarkable improvement. Nevertheless, most counters rely on a large amount of manually labeled data. With the release of synthetic crowd data, a potential alternative is transferring knowledge from them to real data without any manual label. However, there is no method to effectively suppress domain gaps and output elaborate density maps during the transferring. To remedy the above problems, this article proposes a domain-adaptive crowd counting (DACC) framework, which consists of a high-quality image translation and density map reconstruction. To be specific, the former focuses on translating synthetic data to realistic images, which prompts the translation quality by segregating domain-shared/independent features and designing content-aware consistency loss. The latter aims at generating pseudo labels on real scenes to improve the prediction quality. Next, we retrain a final counter using these pseudo labels. Adaptation experiments on six real-world datasets demonstrate that the proposed method outperforms the state-of-the-art methods. Junyu Gao 0001, Tao Han 0002, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Scale-Prior Deformable Convolution for Exemplar-Guided Class-Agnostic Counting
Wei Lin 0018, Xinzhu Ma, Junyu Gao 0001, Lingbo Liu, Shinan Liu, Shuai Yi, Antoni B. Chan |
BMVC | 4 |
| 2022 | DR.VIC: Decomposition and Reasoning for Video Individual CountingabstractPedestrian counting is a fundamental tool for under-standing pedestrian patterns and crowd flow analysis. Existing works (e.g., image-level pedestrian counting, cross-line crowd counting et al.) either only focus on the image-level counting or are constrained to the manual annotation of lines. In this work, we propose to conduct the pedes-trian counting from a new perspective - Video Individual Counting (VIC), which counts the total number of individual pedestrians in the given video (a person is only counted once). Instead of relying on the Multiple Object Tracking (MOT) techniques, we propose to solve the problem by decomposing all pedestrians into the initial pedestrians who existed in the first frame and the new pedestrians with separate identities in each following frame. Then, an end-to-end Decomposition and Reasoning Network (DRNet) is designed to predict the initial pedestrian count with the density estimation method and reason the new pedestrian's count of each frame with the differentiable optimal transport. Extensive experiments are conducted on two datasets with congested pedestrians and diverse scenes, demonstrating the effectiveness of our method over baselines with great superiority in counting the individual pedestrians. Code: https://github.com/taohan10200/DRNet. Tao Han 0002, Lei Bai 0001, Junyu Gao 0001, Qi Wang 0009, Wanli Ouyang |
CVPR | 3 |
| 2022 | Congested crowd instance localization with dilated convolutional swin transformer
Junyu Gao 0001, Maoguo Gong, Xuelong Li 0001 |
Neurocomputing | 1 |
| 2022 | Density-Aware Curriculum Learning for Crowd CountingabstractRecently, crowd counting draws much attention on account of its significant meaning in congestion control, public safety, and ecological surveys. Although the performance is improved dramatically due to the development of deep learning, the scales of these networks also become larger and more complex. Moreover, a large model also entails more time to train for better performance. To tackle these problems, this article first constructs a lightweight model, which is composed of an image feature encoder and a simple but effective decoder, called the pixel shuffle decoder (PSD). PSD ends with a pixel shuffle operator, which can display more density information without increasing the number of convolutional layers. Second, a density-aware curriculum learning (DCL) training strategy is designed to fully tap the potential of crowd counting models. DCL gives each predicted pixel a weight to determine its predicting difficulty and provides guidance on obtaining better generalization. Experimental results exhibit that PSD can achieve outstanding performance on most mainstream datasets while training under the DCL training framework. Besides, we also conduct some experiments about adopting DCL on existing typical crowd counters, and the results show that they all obtain new better performance than before, which further validates the effectiveness of our method. Qi Wang 0009, Wei Lin 0018, Junyu Gao 0001, Xuelong Li 0001 |
IEEE Trans. Cybern. | 3 |
| 2022 | Video Crowd Localization With Multifocus Gaussian Neighborhood Attention and a Large-Scale BenchmarkabstractVideo crowd localization is a crucial yet challenging task, which aims to estimate exact locations of human heads in the given crowded videos. To model spatial-temporal dependencies of human mobility, we propose a multi-focus Gaussian neighborhood attention (GNA), which can effectively exploit long-range correspondences while maintaining the spatial topological structure of the input videos. In particular, our GNA can also capture the scale variation of human heads well using the equipped multi-focus mechanism. Based on the multi-focus GNA, we develop a unified neural network called GNANet to accurately locate head centers in video clips by fully aggregating spatial-temporal information via a scene modeling module and a context cross-attention module. Moreover, to facilitate future researches in this field, we introduce a large-scale crowd video benchmark named VSCrowd (https://github.com/HopLee6/VSCrowd), which consists of 60K+ frames captured in various surveillance scenes and 2M+ head annotations. Finally, we conduct extensive experiments on three datasets including our VSCrowd, and the experiment results show that the proposed method is capable to achieve state-of-the-art performance for both video crowd localization and counting. Haopeng Li 0001, Lingbo Liu, Shinan Liu, Junyu Gao 0001, Bin Zhao 0001, Rui Zhang 0003 |
IEEE Trans. Image Process. | 5 |
| 2022 | Neuron Linear Transformation: Modeling the Domain Shift for Crowd CountingabstractCross-domain crowd counting (CDCC) is a hot topic due to its importance in public safety. The purpose of CDCC is to alleviate the domain shift between the source and target domain. Recently, typical methods attempt to extract domain-invariant features via image translation and adversarial learning. When it comes to specific tasks, we find that the domain shifts are reflected in model parameters' differences. To describe the domain gap directly at the parameter level, we propose a neuron linear transformation (NLT) method, exploiting domain factor and bias weights to learn the domain shift. Specifically, for a specific neuron of a source model, NLT exploits few labeled target data to learn domain shift parameters. Finally, the target neuron is generated via a linear transformation. Extensive experiments and analysis on six real-world data sets validate that NLT achieves top performance compared with other domain adaptation methods. An ablation study also shows that the NLT is robust and more effective than supervised and fine-tune training. Code is available at https://github.com/taohan10200/NLT. Qi Wang 0009, Tao Han 0002, Junyu Gao 0001, Yuan Yuan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Multitask Attention Network for Lane Detection and FittingabstractMany CNN-based segmentation methods have been applied in lane marking detection recently and gain excellent success for a strong ability in modeling semantic information. Although the accuracy of lane line prediction is getting better and better, lane markings' localization ability is relatively weak, especially when the lane marking point is remote. Traditional lane detection methods usually utilize highly specialized handcrafted features and carefully designed postprocessing to detect the lanes. However, these methods are based on strong assumptions and, thus, are prone to scalability. In this work, we propose a novel multitask method that: 1) integrates the ability to model semantic information of CNN and the strong localization ability provided by handcrafted features and 2) predicts the position of vanishing line. A novel lane fitting method based on vanishing line prediction is also proposed for sharp curves and nonflat road in this article. By integrating segmentation, specialized handcrafted features, and fitting, the accuracy of location and the convergence speed of networks are improved. Extensive experimental results on four-lane marking detection data sets show that our method achieves state-of-the-art performance. Qi Wang 0009, Tao Han 0002, Zequn Qin, Junyu Gao 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Multi-Domain Synchronous Refinement Network for Unsupervised Cross-Domain Person Re-IdentificationabstractUnsupervised cross-domain person re-identification (re-ID) is a challenging task, because it is an open-set problem with completely unknown person identities in the target domain. Existing methods attempt to tackle the challenge by transferring image style across domains or generating pseudo labels in the target domain, whereas the valuable information in multiple domains (i.e., source domain, style-transferred data, and target domain) is not taken fully into consideration. To this end, we propose a novel multi-domain synchronous refinement (MDSR) network, where valuable knowledge from multiple domains is sufficiently exploited and refined to enforce the discriminative ability of the model. MDSR network contains two complementary modules dedicated to source-to-target domain adaptation and style-transferred data to the target domain adaptation, respectively. The domain adaptive knowledge from two modules is aggregated in the final stage. Extensive experiments verify our method achieves significant improvements over the state-of-the-art approaches on multiple unsupervised domain adaptative person re-ID tasks. Sikai Bai, Junyu Gao 0001, Qi Wang 0009, Xuelong Li 0001 |
ICME | 2 |
| 2021 | Pixel-Wise Crowd Understanding via Synthetic Data
Qi Wang 0009, Junyu Gao 0001, Wei Lin 0018, Yuan Yuan 0001 |
Int. J. Comput. Vis. | 2 |
| 2021 | Learning to detect anomaly events in crowd scenes from synthetic data
Wei Lin 0018, Junyu Gao 0001, Qi Wang 0009, Xuelong Li 0001 |
Neurocomputing | 2 |
| 2021 | NWPU-Crowd: A Large-Scale Benchmark for Crowd Counting and LocalizationabstractIn the last decade, crowd counting and localization attract much attention of researchers due to its wide-spread applications, including crowd monitoring, public safety, space design, etc. Many convolutional neural networks (CNN) are designed for tackling this task. However, currently released datasets are so small-scale that they can not meet the needs of the supervised CNN-based algorithms. To remedy this problem, we construct a large-scale congested crowd counting and localization dataset, NWPU-Crowd, consisting of 5,109 images, in a total of 2,133,375 annotated heads with points and boxes. Compared with other real-world datasets, it contains various illumination scenes and has the largest density range ( 0 ∼ 20,033). Besides, a benchmark website is developed for impartially evaluating the different methods, which allows researchers to submit the results of the test set. Based on the proposed dataset, we further describe the data characteristics, evaluate the performance of some mainstream state-of-the-art (SOTA) methods, and analyze the new problems that arise on the new data. What's more, the benchmark is deployed at https://www.crowdbenchmark.com/, and the dataset/code/models/results are available at https://gjy3035.github.io/NWPU-Crowd-Sample-Code/. Qi Wang 0009, Junyu Gao 0001, Wei Lin 0018, Xuelong Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Feature-Aware Adaptation and Density Alignment for Crowd Counting in Video SurveillanceabstractWith the development of deep neural networks, the performance of crowd counting and pixel-wise density estimation is continually being refreshed. Despite this, there are still two challenging problems in this field: 1) current supervised learning needs a large amount of training data, but collecting and annotating them is difficult and 2) existing methods cannot generalize well to the unseen domain. A recently released synthetic crowd dataset alleviates these two problems. However, the domain gap between the real-world data and synthetic images decreases the models' performance. To reduce the gap, in this article, we propose a domain-adaptation-style crowd counting method, which can effectively adapt the model from synthetic data to the specific real-world scenes. It consists of multilevel feature-aware adaptation (MFA) and structured density map alignment (SDA). To be specific, MFA boosts the model to extract domain-invariant features from multiple layers. SDA guarantees the network outputs fine density maps with a reasonable distribution on the real domain. Finally, we evaluate the proposed method on four mainstream surveillance crowd datasets, Shanghai Tech Part B, WorldExpo'10, Mall, and UCSD. Extensive experiments are evidence that our approach outperforms the state-of-the-art methods for the same cross-domain counting problem. Junyu Gao 0001, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Cybern. | 1 |
| 2020 | Focus on Semantic Consistency for Cross-Domain Crowd UnderstandingabstractFor pixel-level crowd understanding, it is time-consuming and laborious in data collection and annotation. Some domain adaptation algorithms try to liberate it by training models with synthetic data, and the results in some recent works have proved the feasibility. However, we found that a mass of estimation errors in the background areas impede the performance of the existing methods. In this paper, we propose a domain adaptation method to eliminate it. According to the semantic consistency, a similar distribution in deep layer's features of the synthetic and real-world crowd area, we first introduce a semantic extractor to effectively distinguish crowd and background in high-level semantic information. Besides, to further enhance the adapted model, we adopt adversarial learning to align features in the semantic space. Experiments on three representative real datasets show that the proposed domain adaptation scheme achieves the state-of-the-art for cross-domain counting problems. Tao Han 0002, Junyu Gao 0001, Yuan Yuan 0001, Qi Wang 0009 |
ICASSP | 2 |
| 2020 | Pixel-Level Self-Paced Learning For Super-ResolutionabstractRecently, lots of deep networks are proposed to improve the quality of predicted super-resolution (SR) images, due to its widespread use in several image-based fields. However, with these networks being constructed deeper and deeper, they also cost much longer time for training, which may guide the learners to local optimization. To tackle this problem, this paper designs a training strategy named Pixel-level Self-Paced Learning (PSPL) to accelerate the convergence velocity of SISR models. PSPL imitating self-paced learning gives each pixel in the predicted SR image and its corresponding pixel in ground truth an attention weight, to guide the model to a better region in parameter space. Extensive experiments proved that PSPL could speed up the training of SISR models, and prompt several existing models to obtain new better results. Furthermore, the source code is available at https://github.com/Elin24/PSPL. Wei Lin 0018, Junyu Gao 0001, Qi Wang 0009, Xuelong Li 0001 |
ICASSP | 2 |
| 2020 | Unsupervised Semantic Aggregation and Deformable Template Matching for Semi-Supervised LearningabstractUnlabeled data learning has attracted considerable attention recently. However, it is still elusive to extract the expected high-level semantic feature with mere unsupervised learning. In the meantime, semi-supervised learning (SSL) demonstrates a promising future in leveraging few samples. In this paper, we combine both to propose an Unsupervised Semantic Aggregation and Deformable Template Matching (USADTM) framework for SSL, which strives to improve the classification performance with few labeled data and then reduce the cost in data annotating. Specifically, unsupervised semantic aggregation based on Triplet Mutual Information (T-MI) loss is explored to generate semantic labels for unlabeled data. Then the semantic labels are aligned to the actual class by the supervision of labeled data. Furthermore, a feature pool that stores the labeled samples is dynamically updated to assign proxy labels for unlabeled data, which are used as targets for cross-entropy minimization. Extensive experiments and analysis across four standard semi-supervised learning benchmarks validate that USADTM achieves top performance (e.g., 90.46% accuracy on CIFAR-10 with 40 labels and 95.20% accuracy with 250 labels). The code is released at https://github.com/taohan10200/USADTM. Tao Han 0002, Junyu Gao 0001, Yuan Yuan 0001, Qi Wang 0009 |
NeurIPS | 2 |
| 2020 | PCC Net: Perspective Crowd Counting via Spatial Convolutional NetworkabstractCrowd counting from a single image is a challenging task due to high appearance similarity, perspective changes, and severe congestion. Many methods only focus on the local appearance features and they cannot handle the aforementioned challenges. In order to tackle them, we propose a perspective crowd counting network (PCC Net), which consists of three parts: 1) density map estimation (DME) focuses on learning very local features of density map estimation; 2) random high-level density classification (R-HDC) extracts global features to predict the coarse density labels of random patches in images; and 3) fore-/background segmentation (FBS) encodes mid-level features to segments the foreground and background. Besides, the Down, Up, Left, and Right (DULR) module is embedded in PCC Net to encode the perspective changes on four directions (DULR). The proposed PCC Net is verified on five mainstream datasets, which achieves the state-of-the-art performance on the one and attains the competitive results on the other four datasets. The source code is available at https://github.com/gjy3035/PCC-Net. Junyu Gao 0001, Qi Wang 0009, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Learning From Synthetic Data for Crowd Counting in the WildabstractRecently, counting the number of people for crowd scenes is a hot topic because of its widespread applications (e.g. video surveillance, public security). It is a difficult task in the wild: changeable environment, large-range number of people cause the current methods can not work well. In addition, due to the scarce data, many methods suffer from over-fitting to a different extent. To remedy the above two problems, firstly, we develop a data collector and labeler, which can generate the synthetic crowd scenes and simultaneously annotate them without any manpower. Based on it, we build a large-scale, diverse synthetic dataset. Secondly, we propose two schemes that exploit the synthetic data to boost the performance of crowd counting in the wild: 1) pretrain a crowd counter on the synthetic data, then finetune it using the real data, which significantly prompts the model's performance on real data; 2) propose a crowd counting method via domain adaptation, which can free humans from heavy data annotations. Extensive experiments show that the first method achieves the state-of-the-art performance on four real datasets, and the second outperforms our baselines. The dataset and source code are available at https://gjy3035.github.io/GCC-CL/. Qi Wang 0009, Junyu Gao 0001, Wei Lin 0018, Yuan Yuan 0001 |
CVPR | 2 |
| 2019 | SCAR: Spatial-/channel-wise attention regression networks for crowd counting
Junyu Gao 0001, Qi Wang 0009, Yuan Yuan 0001 |
Neurocomputing | 1 |
| 2019 | Weakly Supervised Adversarial Domain Adaptation for Semantic Segmentation in Urban ScenesabstractSemantic segmentation, a pixel-level vision task, is rapidly developed by using convolutional neural networks (CNNs). Training CNNs requires a large amount of labeled data, but manually annotating data is difficult. For emancipating manpower, in recent years, some synthetic datasets are released. However, they are still different from real scenes, which causes that training a model on the synthetic data (source domain) cannot achieve a good performance on real urban scenes (target domain). In this paper, we propose a weakly supervised adversarial domain adaptation to improve the segmentation performance from synthetic data to real scenes, which consists of three deep neural networks. A detection and segmentation (DS) model focuses on detecting objects and predicting segmentation map; a pixel-level domain classifier (PDC) tries to distinguish the image features from which domains; and an object-level domain classifier (ODC) discriminates the objects from which domains and predicts object classes. PDC and ODC are treated as the discriminators, and DS is considered as the generator. By the adversarial learning, DS is supposed to learn domain-invariant features. In experiments, our proposed method yields the new record of mIoU metric in the same problem. Qi Wang 0009, Junyu Gao 0001, Xuelong Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | A Joint Convolutional Neural Networks and Context Transfer for Street Scenes LabelingabstractStreet scene understanding is an essential task for autonomous driving. One important step toward this direction is scene labeling, which annotates each pixel in the images with a correct class label. Although many approaches have been developed, there are still some weak points. First, many methods are based on the hand-crafted features whose image representation ability is limited. Second, they cannot label foreground objects accurately due to the data set bias. Third, in the refinement stage, the traditional Markov random filed inference is prone to over smoothness. For improving the above problems, this paper proposes a joint method of priori convolutional neural networks at superpixel level (called as “priori s-CNNs”) and soft restricted context transfer. Our contributions are threefold: 1) a priori s-CNNs model that learns priori location information at superpixel level is proposed to describe various objects discriminatingly; 2) a hierarchical data augmentation method is presented to alleviate data set bias in the priori s-CNNs training stage, which improves foreground objects labeling significantly; and 3) a soft restricted MRF energy function is defined to improve the priori s-CNNs model's labeling performance and reduce the over smoothness at the same time. The proposed approach is verified on CamVid data set (11 classes) and SIFT Flow Street data set (16 classes) and achieves a competitive performance. Qi Wang 0009, Junyu Gao 0001, Yuan Yuan 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2018 | Embedding Structured Contour and Location Prior in Siamesed Fully Convolutional Networks for Road DetectionabstractRoad detection from the perspective of moving vehicles is a challenging issue in autonomous driving. Recently, many deep learning methods spring up for this task, because they can extract high-level local features to find road regions from raw RGB data, such as convolutional neural networks and fully convolutional networks (FCNs). However, how to detect the boundary of road accurately is still an intractable problem. In this paper, we propose siamesed FCNs (named “s-FCN-loc”), which is able to consider RGB-channel images, semantic contours, and location priors simultaneously to segment the road region elaborately. To be specific, the s-FCN-loc has two streams to process the original RGB images and contour maps, respectively. At the same time, the location prior is directly appended to the siamesed FCN to promote the final detection performance. Our contributions are threefold: 1) An s-FCN-loc is proposed that learns more discriminative features of road boundaries than the original FCN to detect more accurate road regions. 2) Location prior is viewed as a type of feature map and directly appended to the final feature map in s-FCN-loc to promote the detection performance effectively, which is easier than other traditional methods, namely, different priors for different inputs (image patches). 3) The convergent speed of training s-FCN-loc model is 30% faster than the original FCN because of the guidance of highly structured contours. The proposed approach is evaluated on the KITTI road detection benchmark and one-class road detection data set, and achieves a competitive result with the state of the arts. Qi Wang 0009, Junyu Gao 0001, Yuan Yuan 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2017 | Embedding structured contour and location prior in siamesed fully convolutional networks for road detectionabstractRoad detection from the perspective of moving vehicles is a challenging issue in autonomous driving. Recently, many deep learning methods spring up for this task because they can extract high-level local features to find road regions from raw RGB data, such as Convolutional Neural Networks (CNN) and Fully Convolutional Networks (FCN). However, how to detect the boundary of road accurately is still an intractable problem. In this paper, we propose a siamesed fully convolutional network (named as “s-FCN-loc”) based on VGG-net architecture, which is able to consider RGB-channel, semantic contour and location prior simultaneously to segment road region elaborately. To be specific, the s-FCN-loc has two streams to process original RGB images and contour maps respectively. At the same time, the location prior is directly appended to the last feature map to promote the final detection performance. Experiments demonstrate that the proposed s-FCN-loc can learn more discriminative features of road boundaries and converge 30% faster than the original FCN during the training stage. Finally, the proposed approach is evaluated on KITTI road detection benchmark, and achieves a competitive result. Junyu Gao 0001, Qi Wang 0009, Yuan Yuan 0001 |
ICRA | 1 |