VLDB 2026 Research / reviewers in the wild / expert
Qingjie Liu 0001
dblp:72/10584
· DBLP profile ↗
89ranked-venue papers
4as first author
61since 2021 · last 2026
0000-0002-5181-6451ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 50 · 3 first-author · 38 since 2021Applied, interdisciplinary, general and emerging computing · 31 · 1 first-author · 16 since 2021Artificial intelligence and machine learning · 27 · 2 first-author · 23 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reference patch momentum distillation for open-vocabulary semantic segmentation
Qingjie Liu 0001, Di Huang 0001 |
Sci. China Inf. Sci. | 3 |
| 2026 | A3Bench: an audience-aligned multilingual benchmark for video audience insights understanding
Yiming Lei 0001, Guozhen Peng, Zeming Liu, Hui Qiu, Haitao Leng, Shaoguo Liu, Tingting Gao, Qingjie Liu 0001, Annan Li, Yunhong Wang 0001 |
Frontiers Comput. Sci. | 8 |
| 2026 | LightOcc: Lightweight Spatial Embedding for Efficient Vision-based 3D Occupancy Prediction
Jinqing Zhang, Yanan Zhang 0005, Qingjie Liu 0001, Yunhong Wang 0001 |
Int. J. Comput. Vis. | 4 |
| 2026 | Learn more, forget less: A gradient-Aware data selection approach for LLM
Zeming Liu, Yibai Liu, Zheming Song, Qingjie Liu 0001, Guangxu Chen, Yunhong Wang 0001 |
Signal Process. | 7 |
| 2026 | FACT: A Simple and Efficient Framework for Active FinetuningabstractThe main goal of active finetuning is to improve a pretrained model's performance on a specific task or domain by finetuning it with carefully selected informative or challenging data. Previous research has predominantly focused on the active aspect (i.e., data selection) while uniformly employing full finetuning for model adaptation, which inevitably distorts pretrained features due to distribution shift. This issue becomes particularly pronounced when the model size is large relative to the finetuning data quantity, leading to heightened overfitting risks. To address this critical gap, we formally outline the FiAF task that emphasizes systematic exploration of finetuning methodologies in active learning. We propose FACT, a three-phase hierarchical finetuning framework featuring both efficiency and simplicity, specifically designed for active finetuning scenarios. Our comprehensive experiments span: 1) Three major dataset categories encompassing classic (CIFAR10, CIFAR100, ImageNet-1k), imbalanced (CIFAR10-LT, CIFAR100-LT), and fine-grained (StanfordCars, FGVCAircraft) image classification datasets, each evaluated under 3-5 distinct sampling ratios; 2) Diverse pretrained architectures including Convolutional Neural Network (ConvNeXt), Vision Transformer (ViT), and Vision LSTM (ViL) networks; 3) A systematic investigation of frozen feature augmentation (FroFA) strategies. 4) A comprehensive and rigorous analysis of efficiency and generalizability. The results demonstrate significant improvements with strong generalization and robustness. Notably, under low sampling ratios, our framework achieves remarkable performance gains of over 20% on the ViT model for CIFAR10, CIFAR100, and ImageNet-1k benchmarks. This systematic approach establishes new state-of-the-art performance while maintaining parameter efficiency, proving particularly effective when labeled data is scarce. Wenshuai Xu, You Song, Yuzhuo Cui, Minjie Ren, Qingjie Liu 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | Unveiling the Knowledge of CLIP for Training-Free Open-Vocabulary Semantic SegmentationabstractTraining-free open-vocabulary semantic segmentation aims to explore the potential of frozen vision-language models (VLM) for segmentation tasks. Recent works reform the inference process of CLIP and utilize the features from the final layer to reconstruct dense representations for segmentation, demonstrating promising performance. However, the final layer tends to prioritize global components over local representations, leading to suboptimal robustness and effectiveness of existing methods. In this paper, we propose CLIPSeg, a novel training-free framework that fully exploits the diverse knowledge across layers in CLIP for dense predictions. Our study unveils two key discoveries: Firstly, the features in the middle layers exhibit high locality awareness and feature coherence compared to the final layer, based on which we propose the coherence enhanced residual attention module that generates semantic-aware attention. Secondly, despite not being directly aligned with the text, the deep layers capture valid local semantics that complement those in the final layer. Leveraging this insight, we introduce the deep semantic integration module to boost the patch semantics in the final block. Experiments conducted on 9 segmentation benchmarks with various CLIP models demonstrate that CLIPSeg consistently outperforms all training-free methods by substantial margins, e.g., a 7.8 % improvement in average mIoU for CLIP with a ViT-L backbone, and competes with learning-based counterparts in generalizing to novel concepts in an efficient way. Guodong Wang 0006, Qingjie Liu 0001, Di Huang 0001 |
AAAI | 4 |
| 2025 | GeoBEV: Learning Geometric BEV Representation for Multi-view 3D Object DetectionabstractBird's-Eye-View (BEV) representation has emerged as a mainstream paradigm for multi-view 3D object detection, demonstrating impressive perceptual capabilities. However, existing methods overlook the geometric quality of BEV representation, leaving it in a low-resolution state and failing to restore the authentic geometric information of the scene. In this paper, we identify the drawbacks of previous approaches that limit the geometric quality of BEV representation and propose Radial-Cartesian BEV Sampling (RC-Sampling), which outperforms other feature transformation methods in efficiently generating high-resolution dense BEV representation to restore fine-grained geometric information. Additionally, we design a novel In-Box Label to substitute the traditional depth label generated from the LiDAR points. This label reflects the actual geometric structure of objects rather than just their surfaces, injecting real-world geometric information into the BEV representation. In conjunction with the In-Box Label, Centroid-Aware Inner Loss (CAI Loss) is developed to capture the inner geometric structure of objects. Finally, we integrate the aforementioned modules into a novel multi-view 3D object detector, dubbed GeoBEV, which achieves a state-of-the-art result of 66.2% NDS on the nuScenes test set. Jinqing Zhang, Yanan Zhang 0005, Yunlong Qi, Zehua Fu, Qingjie Liu 0001, Yunhong Wang 0001 |
AAAI | 5 |
| 2025 | GODBench: A Benchmark for Multimodal Large Language Models in Video Comment ArtabstractVideo Comment Art enhances user engagement by providing creative content that conveys humor, satire, or emotional resonance, requiring a nuanced and comprehensive grasp of cultural and contextual subtleties. Although Multimodal Large Language Models (MLLMs) and Chain-of-Thought (CoT) have demonstrated strong reasoning abilities in STEM tasks (e.g. mathematics and coding), they still struggle to generate creative expressions such as resonant jokes and insightful satire. Moreover, existing benchmarks are constrained by their limited modalities and insufficient categories, hindering the exploration of comprehensive creativity in video-based Comment Art creation. To address these limitations, we introduce GODBench, a novel benchmark that integrates video and text modalities to systematically evaluate MLLMs’ abilities to compose Comment Art. Furthermore, inspired by the propagation patterns of waves in physics, we propose Ripple of Thought (RoT), a multi-step reasoning framework designed to enhance the creativity of MLLMs. Extensive experiments on GODBench reveal that existing MLLMs and CoT methods still face significant challenges in understanding and generating creative video comments. In contrast, RoT provides an effective approach to improving creative composing, highlighting its potential to drive meaningful advancements in MLLM-based creativity. Yiming Lei 0001, Zeming Liu, Haitao Leng, Shaoguo Liu, Tingting Gao, Qingjie Liu 0001, Yunhong Wang 0001 |
ACL (1) | 7 |
| 2025 | SPMTrack: Spatio-Temporal Parameter-Efficient Fine-Tuning with Mixture of Experts for Scalable Visual TrackingabstractMost state-of-the-art trackers adopt one-stream paradigm, using a single Vision Transformer for joint feature extraction and relation modeling of template and search region images. However, relation modeling between different image patches exhibits significant variations. For instance, background regions dominated by target-irrelevant information require reduced attention allocation, while foreground, particularly boundary areas, need to be be em-phasized. A single model may not effectively handle all kinds of relation modeling simultaneously. In this paper, we propose a novel tracker called SPMTrack based on mixture-of-experts tailored for visual tracking task (TMoE), combining the capability of multiple experts to handle diverse relation modeling more flexibly. Benefiting from TMoE, we extend relation modeling from image pairs to spatio-temporal context, further improving tracking accuracy with minimal increase in model parameters. Moreover, we employ TMoE as a parameter-efficient fine-tuning method, substantially reducing trainable parameters, which enables us to train SPMTrack of varying scales efficiently and preserve the generalization ability of pretrained models to achieve superior performance. We conduct experiments on seven datasets, and experimental results demonstrate that our method significantly outperforms current state-of-the-art trackers. The source code is available at https://github.com/WenRuiCai/SPMTrack. Qingjie Liu 0001, Yunhong Wang 0001 |
CVPR | 2 |
| 2025 | SeriesBench: A Benchmark for Narrative-Driven Drama Series UnderstandingabstractWith the rapid development of Multi-modal Large Language Models (MLLMs), an increasing number of benchmarks have been established to evaluate the video understanding capabilities of these models. However, these benchmarks focus on standalone videos and only assess "visual elements" like human actions and object states. In reality, contemporary videos often encompass complex and continuous narratives, typically presented as a series. To address this challenge, we propose SeriesBench, a benchmark consisting of 105 carefully curated narrative-driven series, covering 28 specialized tasks that require deep narrative understanding to solve. Specifically, we first select a diverse set of drama series spanning various genres. Then, we introduce a novel long-span narrative annotation method, combined with a full-information transformation approach to convert manual annotations into diverse task formats. To further enhance the model’s capacity for detailed analysis of plot structures and character relationships within series, we propose a novel narrative reasoning framework, PC-DCoT. Extensive results on SeriesBench indicate that existing MLLMs still face significant challenges in understanding narrative-driven series, while PC-DCoT enables these MLLMs to achieve performance improvements. Overall, our SeriesBench and PC-DCoT highlight the critical necessity of advancing model capabilities for understanding narrative-driven series, guiding future MLLMs development. SeriesBench is publicly available at https://github.com/zackhxn/SeriesBench-CVPR2025. Yiming Lei 0001, Zeming Liu, Haitao Leng, Shaoguo Liu, Tingting Gao, Qingjie Liu 0001, Yunhong Wang 0001 |
CVPR | 7 |
| 2025 | SkeletonMix: A Mixup-Based Data Augmentation Framework for Skeleton-Based Action RecognitionabstractSkeleton-based human action recognition has received widespread attention for its robustness to changes in the background and appearance of actors compared to the RGB modality. Data augmentation is widely used to explicitly regularize the model to prevent overfitting, especially when the number of labeled samples is scarce. However, compared to various augmentation methods available for the RGB modality, there are fewer works on the skeleton modality, especially a model-agnostic augmentation method that can be easily integrated into multiple models. We address the problem by proposing a comprehensive data augmentation framework named SkeletonMix, which contains a pair sample selection module for mixup and random augmentations tailored for skeleton modality. SkeletonMix is a non-learning framework and can be applied to different models in a plug-and-play manner. Extensive experiments on NTU RGB+D, NTU RGB+D 120 and PKU-MMD datasets demonstrate the effectiveness of our proposed framework under limited labeled training data. Our proposed method boosts the performance by a maximum of 7.5% on scarce training data setup (5% of training data). Zongye Zhang 0002, Huanyu Zhou, Qingjie Liu 0001, Yunhong Wang 0001 |
ICASSP | 3 |
| 2025 | OpenRSD: Towards Open-Prompts for Object Detection in Remote Sensing Images
Ziyue Huang 0001, Yongchao Feng, Qingjie Liu 0001, Yunhong Wang 0001 |
ICCV | 5 |
| 2025 | Towards Robust and Controllable Text-to-Motion via Masked Autoregressive DiffusionabstractGenerating 3D human motion from text descriptions remains challenging due to the diverse and complex nature of human motion. While existing methods excel within the training distribution, they often struggle with out-of-distribution motions, limiting their applicability in real-world scenarios. Existing VQVAE-based methods often fail to represent novel motions faithfully using discrete tokens, which hampers their ability to generalize beyond seen data. Meanwhile, diffusion-based methods operating on continuous representations often lack fine-grained control over individual frames. To address these challenges, we propose a robust motion generation framework MoMADiff, which combines masked modeling with diffusion processes to generate motion using frame-level continuous representations. Our model supports flexible user-provided keyframe specification, enabling precise control over both spatial and temporal aspects of motion synthesis. MoMADiff demonstrates strong generalization capability on novel text-to-motion datasets with sparse keyframes as motion prompts. Extensive experiments on two held-out datasets and two standard benchmarks show that our method consistently outperforms state-of-the-art models in motion quality, instruction fidelity, and keyframe adherence. The code is available at: https://github.com/zzysteve/MoMADiff Zongye Zhang 0002, Bohan Kong, Qingjie Liu 0001, Yunhong Wang 0001 |
ACM Multimedia | 3 |
| 2025 | AttriPrompt: Dynamic Prompt Composition Learning for CLIPabstractThe evolution of prompt learning methodologies has driven exploration of deeper prompt designs to enhance model performance. However, current deep text prompting approaches suffer from two critical limitations: Over-reliance on constrastive learning objectives that prioritize high-level semantic alignment, neglecting fine-grained feature optimization; Static prompts across all input categories, preventing content-aware adaptation. To address these limitations, we propose AttriPrompt-a novel framework that enhances and refines textual semantic representations by leveraging the intermediate-layer features of CLIP's vision encoder. We designed an Attribute Retrieval module that first clusters visual features from each layer. The aggregated visual features retrieve semantically similar prompts from a prompt pool, which are then concatenated to the input of every layer in the text encoder. Leveraging hierarchical visual information embedded in prompted text features, we introduce Dual-stream Contrastive Learning to realize fine-grained alignment. Furthermore, we introduce a Self-Regularization mechanism by applying explicit regularization constraints between the prompted and non-prompted text features to prevent overfitting on limited training data. Extensive experiments across three benchmarks demonstrate AttriPrompt's superiority over state-of-the-art methods, achieving up to 7.37% improvement in the base-to-novel setting. The observed strength of our method in cross-domain knowledge transfer positions vision-language pre-trained models as more viable solutions for real-world implementation. Qiqi Zhan, Qingjie Liu 0001, Yunhong Wang 0001 |
ACM Multimedia | 3 |
| 2025 | De-Simplifying Pseudo Labels to Enhancing Domain Adaptive Object DetectionabstractDespite its significant success, object detection in traffic and transportation scenarios requires time-consuming and laborious efforts in acquiring high-quality labeled data. Therefore, Unsupervised Domain Adaptation (UDA) for object detection has recently gained increasing research attention. UDA for object detection has been dominated by domain alignment methods, which achieve top performance. Recently, self-labeling methods have gained popularity due to their simplicity and efficiency. In this paper, we investigate the limitations that prevent self-labeling detectors from achieving commensurate performance with domain alignment methods. Specifically, we identify the high proportion of simple samples during training, i.e., the simple-label bias, as the central cause. We propose a novel approach called De-Simplifying Pseudo Labels (DeSimPL) to mitigate the issue. DeSimPL utilizes an instance-level memory bank to implement an innovative pseudo label updating strategy. Then, adversarial samples are introduced during training to enhance the proportion. Furthermore, we propose an adaptive weighted loss to avoid the model suffering from an abundance of false positive pseudo labels in the late training period. Experimental results demonstrate that DeSimPL effectively reduces the proportion of simple samples during training, leading to a significant performance improvement for self-labeling detectors. Extensive experiments conducted on four benchmarks validate our analysis and conclusions. Zehua Fu, Jiaqi Zhou 0016, Qingjie Liu 0001, Yunhong Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | SkeletonX: Data-Efficient Skeleton-Based Action Recognition via Cross-Sample Feature AggregationabstractWhile current skeleton action recognition models demonstrate impressive performance on large-scale datasets, their adaptation to new application scenarios remains challenging. These challenges are particularly pronounced when facing new action categories, diverse performers, and varied skeleton layouts, leading to significant performance degeneration. Additionally, the high cost and difficulty of collecting skeleton data make large-scale data collection impractical. This paper studies one-shot and limited-scale learning settings to enable efficient adaptation with minimal data. Existing approaches often overlook the rich mutual information between labeled samples, resulting in sub-optimal performance in low-data scenarios. To boost the utility of labeled data, we identify the variability among performers and the commonality within each action as two key attributes. We present SkeletonX, a lightweight training pipeline that integrates seamlessly with existing GCN-based skeleton action recognizers, promoting effective training under limited labeled data. First, we propose a tailored sample pair construction strategy on two key attributes to form and aggregate sample pairs. Next, we develop a concise and effective feature aggregation module to process these pairs. Extensive experiments are conducted on NTU RGB+D, NTU RGB+D 120, and PKU-MMD with various GCN backbones, demonstrating that the pipeline effectively improves performance when trained from scratch with limited data. Moreover, it surpasses previous state-of-the-art methods in the one-shot setting, with only 1/10 of the parameters and much fewer FLOPs. Zongye Zhang 0002, Qingjie Liu 0001, Yunhong Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Multi-Grained Contrastive Learning for Text-Supervised Open-Vocabulary Semantic SegmentationabstractLearning open-vocabulary semantic segmentation (OVSS) from text supervision has recently received increasing attention for its promising potential in real-world applications. However, only with image-level supervision, it struggles to achieve dense and robust cross-modal alignment and thus limits pixel-level predictions. In this article, we present a novel approach to this task with M ulti- G rained C ross-modal C ontrastive L earning, named MGCCL. Specifically, unlike current solutions restricted by coarse image/object-text alignment, MGCCL constructs pseudo multi-granular semantic correspondences at the object-, part-, and pixel-level and collaborates with hard sampling strategies to conduct cross-modal contrastive learning, significantly facilitating fine-grained alignment. Further, we develop an adaptive semantic unit which flexibly harnesses the learned multi-grained cross-modal alignment capabilities to effectively mitigate the under- and over-segmentation issues arising from the per-group and per-pixel units. Extensive experiments over a broad suite of eight segmentation benchmarks show that our approach delivers significant advancements over state-of-the-art counterparts, demonstrating its effectiveness. Pu Ge, Guodong Wang 0006, Qingjie Liu 0001, Di Huang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | CtxMIM: Context-Enhanced Masked Image Modeling for Remote Sensing Image UnderstandingabstractLearning representations through self-supervision on unlabeled data has proven highly effective for understanding diverse images. However, remote sensing images often have complex and densely populated scenes with multiple land objects and no clear foreground objects. This intrinsic property generates high object density, resulting in false positive pairs or missing contextual information in self-supervised learning. To address these problems, we propose a context-enhanced masked image modeling (CtxMIM) method, a simple yet efficient MIM-based self-supervised learning for remote sensing image understanding. CtxMIM formulates original image patches as a reconstructive template and employs a Siamese framework to operate on two sets of image patches. A context-enhanced generative branch is introduced to provide contextual information through context consistency constraints in the reconstruction. With the simple and elegant design, CtxMIM encourages the pretraining model to learn object-level or pixel-level features on a large-scale dataset without specific temporal or geographical constraints. Finally, extensive experiments show that features learned by CtxMIM outperform fully supervised and state-of-the-art self-supervised learning methods on various downstream tasks, including land cover classification, semantic segmentation, object detection, and instance segmentation. These results demonstrate that CtxMIM learns impressive remote sensing representations with high generalization and transferability. Qingjie Liu 0001, Yunhong Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | HIPTrack: Visual Tracking with Historical PromptsabstractTrackers that follow Siamese paradigm utilize similarity matching between template and search region features for tracking. Many methods have been explored to enhance tracking performance by incorporating tracking history to better handle scenarios involving target appearance variations such as deformation and occlusion. However, the utilization of historical information in existing methods is insufficient and incomprehensive, which typically requires repetitive training and introduces a large amount of computation. In this paper, we show that by providing a tracker that follows Siamese paradigm with precise and updated historical information, a significant performance improvement can be achieved with completely unchanged parameters. Based on this, we propose a historical prompt network that uses refined historical foreground masks and historical visual features of the target to provide comprehensive and precise prompts for the tracker. We build a novel tracker called HIPTrack based on the historical prompt network, which achieves considerable performance improvements without the need to retrain the entire model. We conduct experiments on seven datasets and experimental results demonstrate that our method surpasses the current state-of-the-art trackers on LaSOT, LaSOText, GOT-10k and NfS. Furthermore, the historical prompt network can seamlessly integrate as a plug-and-play module into existing trackers, providing performance enhancements. The source code is available at https://github.com/WenRuiCai/HIPTrack. Qingjie Liu 0001, Yunhong Wang 0001 |
CVPR | 2 |
| 2024 | ActiveDC: Distribution Calibration for Active FinetuningabstractThe pretraining-finetuning paradigm has gained popularity in various computer vision tasks. In this paradigm, the emergence of active finetuning arises due to the abundance of large-scale data and costly annotation requirements. Active finetuning involves selecting a subset of data from an unlabeled pool for annotation, facilitating subsequent finetuning. However, the use of a limited number of training samples can lead to a biased distribution, potentially resulting in model overfitting. In this paper, we propose a new method called ActiveDC for the active finetuning tasks. Firstly, we select samples for annotation by optimizing the distribution similarity between the subset to be selected and the entire unlabeled pool in continuous space. Secondly, we calibrate the distribution of the selected samples by exploiting implicit category information in the unlabeled pool. The feature visualization provides an intuitive sense of the effectiveness of our method to distribution calibration. We conducted extensive experiments on three image classification datasets with different sampling ratios. The results indicate that ActiveDC consistently outperforms the baseline performance in all image classification tasks. The improvement is particularly significant when the sampling ratio is low, with performance gains of up to 10%. Our code will be publicly available. Wenshuai Xu, Jinzhou Meng, Qingjie Liu 0001, Yunhong Wang 0001 |
CVPR | 5 |
| 2024 | MutDet: Mutually Optimizing Pre-training for Remote Sensing Object Detection
Ziyue Huang 0001, Yongchao Feng, Qingjie Liu 0001, Yunhong Wang 0001 |
ECCV (12) | 3 |
| 2024 | FSD-BEV: Foreground Self-distillation for Multi-view 3D Object Detection
Jinqing Zhang, Yanan Zhang 0005, Qingjie Liu 0001, Baohui Wang, Yunhong Wang 0001 |
ECCV (8) | 4 |
| 2024 | Read, Spell and Repeat: Scene Text Recognition with Vision-Language Circular RefinementabstractScene Text Recognition (STR) has long been considered an important yet challenging task in the field of computer vision. Recent works have demonstrated that utilizing language information is effective for the visually difficult images, like ones with occultation or blurring. However, the use of language information sometimes leads to the over-correction problem. For out-of-vocabulary samples (e.g. "hou" and "0x4a"), some methods have tended to be biased to language side and over-corrected (e.g. over-correct "hou" to "hot"). This imbalance of vision and language has limited the usage of models in practical scenarios, yet it is rarely occurs for human. To address this issue, we rethink the human’s recognition process and propose a model behaving in the order of "Read, Spell and Repeat". It refines the recognition process circularly with vision and language information. With this mechanism, our model integrates vision and language information in a more effective manner, achieving higher accuracy with less parameters compared to baseline and competitive performance with SOTA methods in the standard benchmarks. Taiwei Zhang, Weixin Li 0001, Qingjie Liu 0001, Yunhong Wang 0001 |
ICASSP | 4 |
| 2024 | Towards Generalizable Referring Image Segmentation Via Target Prompt And Visual CoherenceabstractReferring image segmentation (RIS) aims to segment objects in an image conditioning on free-form text descriptions. Despite the overwhelming progress, it still remains challenging for current approaches to perform well on cases with various text expressions or with unseen visual entities, limiting its further application. In this paper, we present a novel RIS approach, which substantially improves the generalization ability by addressing the two dilemmas mentioned above. Specially, to deal with unconstrained texts, we propose to boost a given expression with an explicit and crucial prompt, which complements the expression in a unified context, facilitating target capturing in the presence of linguistic style changes. Furthermore, we introduce a multi-modal fusion aggregation module with visual guidance from a powerful pretrained model to leverage spatial relations and pixel coherences to handle the incomplete target masks and false positive irregular clumps which often appear on unseen visual entities. Extensive experiments are conducted in the zero-shot cross-dataset settings and the proposed approach achieves consistent gains compared to the state-of-the-art, e.g., $4.15 \%$, $5.45 \%$, and $4.64 \%$ mIoU increase on RefCOCO, RefCOCO+ and ReferIt respectively, demonstrating its effectiveness. Pu Ge, Shichao Fan, Qingjie Liu 0001, Di Huang 0001, Yunhong Wang 0001 |
ICIP | 5 |
| 2024 | Semantic Enhanced Few-Shot Object DetectionabstractFew-shot object detection (FSOD), which aims to detect novel objects with limited annotated instances, has made significant progress in recent years. However, existing methods still suffer from biased representations, especially for novel classes in extremely low-shot scenarios. During fine-tuning, a novel class may exploit knowledge from similar base classes to construct its own feature distribution, leading to classification confusion and performance degradation. To address these challenges, we propose a fine-tuning based FSOD framework that utilizes semantic embeddings for better detection. In our proposed method, we align the visual features with class name embeddings and replace the linear classifier with our semantic similarity classifier. Our method trains each region proposal to converge to the corresponding class embedding. Furthermore, we introduce a multimodal feature fusion to augment the vision-language communication, enabling a novel class to draw support explicitly from well-trained similar base classes. To prevent class confusion, we propose a semantic-aware max-margin loss, which adaptively applies a margin beyond similar classes. As a result, our method allows each novel class to construct a compact feature space without being confused with similar base classes. Extensive experiments on Pascal VOC and MS COCO demonstrate the superiority of our method. Yingjie Gao 0001, Qingjie Liu 0001, Yunhong Wang 0001 |
ICIP | 3 |
| 2024 | Fast Textile Pilling Classification Based on a Lightweight Network and 3D Point CloudsabstractPoint clouds have demonstrated extensive application prospects in various fields, including research related to the evaluation of textile pilling. We collect 3D point cloud data in the actual test environment of textiles, which has been organized and named the TextileNet dataset. To the best of our knowledge, it is the first publicly available 3D point cloud dataset in the field of textile pilling assessment. Based on the Non-parametric Network for 3D point cloud analysis (Point-NN), we construct a Few-parameter Network called Point-FN for experiments on the TextileNet dataset. Experimental results indicate that under conditions with a parameter count of only 0.5M and FLOPs of 1.7G, Point-FN achieves an Overall Accuracy (OA) of 91.1% and a Mean per-class Accuracy (MA) of 93.0%. Moreover, under the testing conditions of a single RTX 2080Ti GPU, Point-FN demonstrates an inference speed of 164 FPS. Testing results on other publicly available datasets also validate the competitive performance of Point-FN. The proposed TextileNet dataset will be publicly available. Yizhou Jin, Qingjie Liu 0001, Di Huang 0001, Yunhong Wang 0001 |
ICME | 6 |
| 2024 | DSD-DA: Distillation-based Source Debiasing for Domain Adaptive Object DetectionabstractThough feature-alignment based Domain Adaptive Object Detection (DAOD) methods have achieved remarkable progress, they ignore the source bias issue, i.e., the detector tends to acquire more source-specific knowledge, impeding its generalization capabilities in the target domain. Furthermore, these methods face a more formidable challenge in achieving consistent classification and localization in the target domain compared to the source domain. To overcome these challenges, we propose a novel Distillation-based Source Debiasing (DSD) framework for DAOD, which can distill domain-agnostic knowledge from a pre-trained teacher model, improving the detector’s performance on both domains. In addition, we design a Target-Relevant Object Localization Network (TROLN), which can mine target-related localization information from source and target-style mixed data. Accordingly, we present a Domain-aware Consistency Enhancing (DCE) strategy, in which these information are formulated into a new localization representation to further refine classification scores in the testing stage, achieving a harmonization between classification and localization. Extensive experiments have been conducted to manifest the effectiveness of this method, which consistently improves the strong baseline by large margins, outperforming existing alignment-based works. Yongchao Feng, Yingjie Gao 0001, Ziyue Huang 0001, Yanan Zhang 0005, Qingjie Liu 0001, Yunhong Wang 0001 |
ICML | 6 |
| 2024 | Prior knowledge-guided multilevel graph neural network for tumor risk prediction and interpretation via multi-omics data integrationabstractThe interrelation and complementary nature of multi-omics data can provide valuable insights into the intricate molecular mechanisms underlying diseases. However, challenges such as limited sample size, high data dimensionality and differences in omics modalities pose significant obstacles to fully harnessing the potential of these data. The prior knowledge such as gene regulatory network and pathway information harbors useful gene-gene interaction and gene functional module information. To effectively integrate multi-omics data and make full use of the prior knowledge, here, we propose a Multilevel-graph neural network (GNN): a hierarchically designed deep learning algorithm that sequentially leverages multi-omics data, gene regulatory networks and pathway information to extract features and enhance accuracy in predicting survival risk. Our method achieved better accuracy compared with existing methods. Furthermore, key factors nonlinearly associated with the tumor pathogenesis are prioritized by employing two interpretation algorithms (i.e. GNN-Explainer and IGscore) for neural networks, at gene and pathway level, respectively. The top genes and pathways exhibit strong associations with disease in survival analyses, many of which such as SEC61G and CYP27B1 are previously reported in the literature. Hongxi Yan, Dawei Weng, Dongguo Li, Wenji Ma, Qingjie Liu 0001 |
Briefings Bioinform. | 6 |
| 2024 | Learning group interaction for sports video understanding from a perspective of athlete
Zehua Fu, Qingjie Liu 0001, Yunhong Wang 0001, Xunxun Chen |
Frontiers Comput. Sci. | 3 |
| 2024 | An Empirical Study on Multi-domain Robust Semantic Segmentation
Pu Ge, Qingjie Liu 0001, Shichao Fan, Yunhong Wang 0001 |
Int. J. Comput. Vis. | 3 |
| 2024 | Refining and reweighting pseudo labels for weakly supervised object detection
Yongchao Feng, Qingjie Liu 0001, Yunhong Wang 0001 |
Neurocomputing | 4 |
| 2024 | Improving Object Detection From Remote Sensing Images via Self-Supervised Adaptive Fusion NetworksabstractDetecting objects in remote sensing images is essential for intelligent interpretation. Although deep neural networks have made significant progress in recent years, they often struggle with complex backgrounds in remote sensing images, which can lead to inaccurate detection. To tackle this problem, a self-supervised adaptive fusion network (SSAFN) has been developed. The SSAFN includes an adaptive fusion module (AFM) and a self-supervised task module (SSTM). The AFM mainly fuses the deep semantic information to the shallow features with appropriate weights to enhance the semantic information of the shallow features. The SSTM is mainly to constrain the AFM through self-supervised tasks to fulfill the function similar to the attention mechanism: to make the AFM enhance the target feature representation and suppress the background information. The SSAFN reduces the impact of complex backgrounds on object representation, resulting in better detection results for various types of objects such as buildings, ships, and more. The proposed method has been tested on various datasets and has not only improved the detection accuracy for different types of objects but also enhanced the performance of popular object detection algorithms. Qiu Lu, Tao Xu 0021, Jiwen Dong, Qingjie Liu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Generic Knowledge Boosted Pretraining for Remote Sensing ImagesabstractDeep learning models are essential for scene classification, change detection, land cover segmentation, and other remote sensing image understanding tasks. Most backbones of existing remote sensing deep learning models are typically initialized by pre-trained weights obtained from ImageNet pre-training (IMP). However, domain gaps exist between remote sensing images and natural images (e.g., ImageNet), making deep learning models initialized by pre-trained weights of IMP perform poorly for remote sensing image understanding. Although some pre-training methods are studied in the remote sensing community, current remote sensing pre-training methods face the problem of vague generalization by only using remote sensing images. In this paper, we propose a novel remote sensing pre-training framework, Generic Knowledge Boosted Remote Sensing Pre-training (GeRSP), to learn robust representations from remote sensing and natural images for remote sensing understanding tasks. GeRSP contains two pre-training branches: (1) A self-supervised pre-training branch is adopted to learn domain-related representations from unlabeled remote sensing images. (2) A supervised pre-training branch is integrated into GeRSP for general knowledge learning from labeled natural images. Moreover, GeRSP combines two pre-training branches using a teacher-student architecture to simultaneously learn representations with general and special knowledge, which generates a powerful pre-trained model for deep learning model initialization. Finally, we evaluate GeRSP and other remote sensing pre-training methods on three downstream tasks,i.e., object detection, semantic segmentation, and scene classification. The extensive experimental results consistently demonstrate that GeRSP can effectively learn robust representations in a unified manner, improving the performance of remote sensing downstream tasks. Code and pre-trained models: https://github.com/floatingstarZ/GeRSP. Ziyue Huang 0001, Yuan Gong 0004, Qingjie Liu 0001, Yunhong Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | A Sparse Sharing Multitask Framework for Building Footprint Extraction From Remote Sensing Imagery Following the Dual Lottery Ticket HypothesisabstractBuilding footprint extraction from high-resolution remote sensing imagery is significant for urban planning, change detection, disaster management, and other applications. Recently, researchers have found that the edge features of buildings are crucial in extracting building footprints, and multitask deep learning is used to share edge feature information. However, these multitask deep learning frameworks adopt a hard sharing approach, which cannot avoid the adverse effects caused by the differences between different tasks, resulting in the problem of blurred edges and building boundaries. To address this issue, this article proposes a dual lottery ticket hypothesis (DLTH) and sparse sharing-based multitask deep learning framework, dual sparse sharing architecture (DSSA), to transmit the edge information in the edge detection to the building footprint extraction by sharing partial parameters. First, the subnetworks of building footprint extraction and edge detection are constructed according to the sparse rate and parameter sharing rate to control the dependencies between the subnetworks. Second, given the difference in the importance of the two tasks, a cosine unequal-scaled alternating training strategy is proposed to strengthen and weaken the transmission of edge information periodically. Third, following the DLTH, the loss function with${L}3$/2 regularization constraint is used to promote the information transmission and parameter conversion of the subnetwork by using the global information. Finally, aiming at the edge of building footprint extraction results, a pixel-based evaluation index, edge extraction accuracy (${\mathrm {EEA}}^{(n)})$, is designed by morphological erosion to better evaluate the integrity of the edge of building footprint extraction results. The experiments conducted on a self-annotated dataset and two public datasets (i.e., WHU Aerial Imagery dataset and Massachusetts Building dataset) show that DSSA can achieve better edge effects than the baseline and show excellent generalization ability. Huaqiao Xing, Junwu Xiang, Qingjie Liu 0001, Yunhong Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | HiT: Building Mapping With Hierarchical TransformersabstractDeep learning-based methods have been extensively explored for automatic building mapping from high-resolution remote sensing images over recent years. While most building mapping models produce vector polygons of buildings for geographic and mapping systems, dominant methods typically decompose polygonal building extraction in some sub-problems, including segmentation, polygonization, and regularization, leading to complex inference procedures, low accuracy, and poor generalization. In this paper, we propose a simple and novel building mapping method with Hierarchical Transformers, called HiT, improving polygonal building mapping quality from high-resolution remote sensing images. HiT builds on a two-stage detection architecture by adding a polygon head parallel to classification and bounding box regression heads. HiT simultaneously outputs building bounding boxes and vector polygons, which is fully end-to-end trainable. The polygon head formulates a building polygon as serialized vertices with the bidirectional characteristic, a simple and elegant polygon representation avoiding the start or end vertex hypothesis. Under this new perspective, the polygon head adopts a transformer encoder-decoder architecture to predict serialized vertices supervised by the designed bidirectional polygon loss. Furthermore, a hierarchical attention mechanism combined with convolution operation is introduced in the encoder of the polygon head, providing more geometric structures of building polygons at vertex and edge levels. Comprehensive experiments on two benchmarks (the CrowdAI and Inria datasets) demonstrate that our method achieves a new state-of-the-art in terms of instance segmentation and polygonal metrics compared with state-of-the-art methods. Moreover, qualitative results verify the superiority and effectiveness of our model under complex scenes. Qingjie Liu 0001, Yunhong Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Pixel-Level Domain Adaptation: A New Perspective for Enhancing Weakly Supervised Semantic SegmentationabstractRecent attention has been devoted to the pursuit of learning semantic segmentation models exclusively from image tags, a paradigm known as image-level Weakly Supervised Semantic Segmentation (WSSS). Existing attempts adopt the Class Activation Maps (CAMs) as priors to mine object regions yet observe the imbalanced activation issue, where only the most discriminative object parts are located. In this paper, we argue that the distribution discrepancy between the discriminative and the non-discriminative parts of objects prevents the model from producing complete and precise pseudo masks as ground truths. For this purpose, we propose a Pixel-Level Domain Adaptation (PLDA) method to encourage the model in learning pixel-wise domain-invariant features. Specifically, a multi-head domain classifier trained adversarially with the feature extraction is introduced to promote the emergence of pixel features that are invariant with respect to the shift between the source (i.e., the discriminative object parts) and the target (i.e., the non-discriminative object parts) domains. In addition, we come up with a Confident Pseudo-Supervision strategy to guarantee the discriminative ability of each pixel for the segmentation task, which serves as a complement to the intra-image domain adversarial training. Our method is conceptually simple, intuitive and can be easily integrated into existing WSSS methods. Taking several strong baseline models as instances, we experimentally demonstrate the effectiveness of our approach under a wide range of settings. Ye Du 0002, Zehua Fu, Qingjie Liu 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | YOLC: You Only Look Clusters for Tiny Object Detection in Aerial ImagesabstractDetecting objects from aerial images poses significant challenges due to the following factors: 1) Aerial images typically have very large sizes, generally with millions or even hundreds of millions of pixels, while computational resources are limited. 2) Small object size leads to insufficient information for effective detection. 3) Non-uniform object distribution leads to computational resource wastage. To address these issues, we propose YOLC (You Only Look Clusters), an efficient and effective framework that builds on an anchor-free object detector, CenterNet. To overcome the challenges posed by large-scale images and non-uniform object distribution, we introduce a Local Scale Module (LSM) that adaptively searches cluster regions for zooming in for accurate detection. Additionally, we modify the regression loss using Gaussian Wasserstein distance (GWD) to obtain high-quality bounding boxes. Deformable convolution and refinement methods are employed in the detection head to enhance the detection of small objects. We perform extensive experiments on two aerial image datasets, including Visdrone2019 and UAVDT, to demonstrate the effectiveness and superiority of our proposed approach. Guangshuai Gao, Ziyue Huang 0001, Qingjie Liu 0001, Yunhong Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Improving Multi-Person Pose Tracking With a Confidence NetworkabstractHuman pose estimation and tracking are fundamental tasks for understanding human behaviors in videos. Existing top-down framework-based methods usually perform three-stage tasks: human detection, pose estimation and tracking. Although promising results have been achieved, these methods rely heavily on high-performance detectors and may fail to track persons who are occluded or miss-detected. To overcome these problems, in this paper, we develop a novel keypoint confidence network and a tracking pipeline to improve human detection and pose estimation in top-down approaches. Specifically, the keypoint confidence network is designed to determine whether each keypoint is occluded, and it is incorporated into the pose estimation module. In the tracking pipeline, we propose the Bboxrevision module to reduce missing detection and the ID-retrieve module to correct lost trajectories, improving the performance of the detection stage. Experimental results show that our approach is universal in human detection and pose estimation, achieving state-of-the-art performance on both PoseTrack 2017 and 2018 datasets. Zehua Fu, Wenhang Zuo, Qingjie Liu 0001, Yunhong Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Learning Discriminative Representations for Skeleton Based Action RecognitionabstractHuman action recognition aims at classifying the category of human action from a segment of a video. Recently, people have dived into designing GCN-based models to extract features from skeletons for performing this task, because skeleton representations are much more efficient and robust than other modalities such as RGB frames. However, when employing the skeleton data, some important clues like related items are also discarded. It results in some ambiguous actions that are hard to be distinguished and tend to be misclassified. To alleviate this problem, we propose an auxiliary feature refinement head (FR Head), which consists of spatial-temporal decoupling and contrastive feature refinement, to obtain discriminative representations of skeletons. Ambiguous samples are dynamically discovered and calibrated in the feature space. Furthermore, FR Head could be imposed on different stages of GCNs to build a multi-level refinement for stronger supervision. Extensive experiments are conducted on NTU RGB+D, NTU RGB+D 120, and NW-UCLA datasets. Our proposed models obtain competitive results from state-of-the-art methods and can help to discriminate those ambiguous samples. Codes are available at https://github.com/zhysora/FR-Head. Huanyu Zhou, Qingjie Liu 0001, Yunhong Wang 0001 |
CVPR | 2 |
| 2023 | BISVP: Building Footprint Extraction Via Bidirectional Serialized Vertex PredictionabstractExtracting building footprints from remote sensing images has been attracting extensive attention recently. Dominant approaches address this challenging problem by generating vectorized building masks with cumbersome refinement stages, which limits the application of such methods. In this paper, we introduce a new refinement-free and end-to-end building footprint extraction method, which is conceptually intuitive, simple, and effective. Our method, termed as BiSVP, represents a building instance with ordered vertices and formulates the building footprint extraction as predicting the serialized vertices directly in a bidirectional fashion. Moreover, we propose a cross-scale feature fusion (CSFF) module to facilitate high resolution and rich semantic feature learning, which is essential for the dense building vertex prediction task. Without bells and whistles, our BiSVP outperforms state-of-the-art methods by considerable margins on three building instance segmentation benchmarks, clearly demonstrating its superiority. The code and datasets will be made public available. Ye Du 0002, Qingjie Liu 0001, Yunhong Wang 0001 |
ICASSP | 4 |
| 2023 | SA-BEV: Generating Semantic-Aware Bird's-Eye-View Feature for Multi-view 3D Object DetectionabstractRecently, the pure camera-based Bird’s-Eye-View (BEV) perception provides a feasible solution for economical autonomous driving. However, the existing BEV-based multiview 3D detectors generally transform all image features into BEV features, without considering the problem that the large proportion of background information may submerge the object information. In this paper, we propose Semantic-Aware BEV Pooling (SA-BEVPool), which can filter out background information according to the semantic segmentation of image features and transform image features into semantic-aware BEV features. Accordingly, we propose BEV-Paste, an effective data augmentation strategy that closely matches with semantic-aware BEV feature. In addition, we design a Multi-Scale Cross-Task (MSCT) head, which combines task-specific and cross-task information to predict depth distribution and semantic segmentation more accurately, further improving the quality of semantic-aware BEV feature. Finally, we integrate the above modules into a novel multi-view 3D object detection framework, namely SA-BEV Experiments on nuScenes show that SA-BEV achieves state-of-the-art performance. Code has been available at https://github.com/mengtan00/SA-BEV.git. Jinqing Zhang, Yanan Zhang 0005, Qingjie Liu 0001, Yunhong Wang 0001 |
ICCV | 3 |
| 2023 | Transbuilding: An End-to-End Polygonal Building Extraction with TransformersabstractIn this paper, we propose a simple yet powerful network, called TransBuilding, for high-quality polygonal building extraction from remote sensing images. Unlike many previous methods that vectorize building masks through mask refinement and fitting or vertex prediction and assembling, our approach predicts the building vertex sequence with a vertex transformer (termed as VertexFormer) branch without any additional processing. The VertexFormer branch represents a polygon as a Bi-directional Ring without start or end vertex hypothesis, which leads to a simple and elegant representation of polygons avoiding ambiguous of defining the start vertex in polygons. Furthermore, three self-attention modules in row-wise, column-wise, and vertex-wise are integrated in parallel together to better capture geometric structures of building polygons. We graft the VertexFormer module onto the standard Faster RCNN detector and train the model end-to-endly using the novel Bi-Ring loss developed by the new perspective of Bi-directional Ring. Extensive experiments on the benchmark CrowdAI dataset demonstrate that our method outperforms state-of-the-art methods by considerable margins. Weiming Zhang 0001, Qingjie Liu 0001, Wei Wang 0115, Yunhong Wang 0001 |
ICIP | 2 |
| 2023 | TERNformer: Topology-Enhanced Road Network Extraction by Exploring Local ConnectivityabstractRemote-sensing images provide us with rich information for extracting road networks. However, there are still great challenges ahead, such as occlusions caused by trees and shadows, and complex topology. In this work, we focus on the topology of road networks. Inspired by the observation that road networks are composed of road fragments in a bottom-up way and the breaks between fragments tend to be connected within a local area, we propose a Topology-Enhanced Road Network extraction (termed TERNformer) method by exploring local connectivity. First, a transformer-based network is built for road feature extraction to capture long-range context. Furthermore, we propose parallel depth-wise separable dilated convolution blocks (DSDB) to extract local information within different ranges. Thereafter, a minimum spanning tree-based local structure exploring block (LSEB) is built to enhance the topology of the road network. Finally, a simple but effective shortest-path-based method is used to refine the road network connectivity within a local threshold. Experiments conducted on two datasets demonstrate the superiority of TERNformer. TERNformer outperforms the state-of-the-art methods on CityScale dataset with the best topology performance. The result on DeepGlobe dataset improves 4.83% APLS to state-of-the-art methods. Qingjie Liu 0001, Wei Wang 0115, Yunhong Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | LgNet: A Local-Global Network for Action Recognition and BeyondabstractThis work addresses the task of action recognition in video sequences. In real world applications, this task is quite challenging due to the complex background of video content, the similarities between different types of actions, the dependence on a large amount of annotated data, and so on. Most of the existing methods fail to distinguish similar actions with the same static appearance and motion pattern. We attempt to address this issue from the perspective of a local-global view, considering videos as combinations of a set of action units (local semantic information) and their relations along temporal dimension (global relation information). To achieve this end, we propose a novel Local-global Networks (LgNet) to enhance recognition of similar action. Besides, we propose an end-to-end training method to decrease the reliance on annotated data. It combines self-supervised learning and supervised learning, which not only enables the model to learn video representations from a large number unannotated data but also avoids subsequent finetuning. The proposed training method can be flexibly equipped to a wide array of vision tasks. Experiments on several benchmark datasets show that our proposed model and training method achieve state-of-the-art performance. Jiaqi Zhou 0016, Zehua Fu, Qiuyu Huang, Qingjie Liu 0001, Yunhong Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2022 | D3: Duplicate Detection Decontaminator for Multi-Athlete Tracking in Sports Videos
Zehua Fu, Qingjie Liu 0001, Yunhong Wang 0001, Xunxun Chen |
ACCV (7) | 3 |
| 2022 | Reading Chinese in Natural Scenes with a Bag-of-Radicals Prior
Qingjie Liu 0001, Jiaxin Chen 0002, Yunhong Wang 0001 |
BMVC | 2 |
| 2022 | Weakly Supervised Semantic Segmentation by Pixel-to-Prototype ContrastabstractThough image-level weakly supervised semantic seg-mentation (WSSS) has achieved great progress with Class Activation Maps (CAMs) as the cornerstone, the large su-pervision gap between classification and segmentation still hampers the model to generate more complete and precise pseudo masks for segmentation. In this study, we propose weakly-supervised pixel-to-prototype contrast that can provide pixel-level supervisory signals to narrow the gap. Guided by two intuitive priors, our method is executed across different views and within per single view of an image, aiming to impose cross-view feature semantic consistency regularization and facilitate intra(inter)-class compactness(dispersion) of the feature space. Our method can be seamlessly incorporated into existing WSSS models with-out any changes to the base networks and does not incur any extra inference burden. Extensive experiments manifest that our method consistently improves two strong baselines by large margins, demonstrating the effectiveness. Specifically, built on top of SEAM, we improve the initial seed mIoU on PASCAL VOC 2012 from 55.4% to 61.5%. Moreover, armed with our method, we increase the segmentation mIoU of EPS from 70.8% to 73.6%, achieving new state-of-the-art. Ye Du 0002, Zehua Fu, Qingjie Liu 0001, Yunhong Wang 0001 |
CVPR | 3 |
| 2022 | Robust Logo Detection Across Large Style Variations
Qingjie Liu 0001 |
ICANN (1) | 2 |
| 2022 | Visual Grounding with TransformersabstractIn this paper, we propose a transformer based approach for visual grounding. Unlike existing proposal-and-rank frameworks that rely heavily on pretrained object detectors or proposal-free frameworks that upgrade an off-the-shelf one-stage detector by fusing textual embeddings, our approach is built on top of a transformer encoder-decoder and is independent of any pretrained detectors or word embedding models. Termed as VGTR – Visual Grounding with TRansformers, our approach is designed to learn semantic-discriminative visual features under the guidance of the textual description without harming their location ability. This information flow enables our VGTR to have a strong capability in capturing context-level semantics of both vision and language modalities, rendering us to aggregate accurate visual clues implied by the description to locate the interested object instance. Experiments show that our method outperforms state-of-the-art proposal-free approaches by a considerable margin on four benchmarks. Ye Du 0002, Zehua Fu, Qingjie Liu 0001, Yunhong Wang 0001 |
ICME | 3 |
| 2022 | PanFormer: A Transformer Based Model for Pan-SharpeningabstractPan-sharpening aims at producing a high-resolution (HR) multi-spectral (MS) image from a low-resolution (LR) multi-spectral (MS) image and its corresponding panchromatic (PAN) image acquired by a same satellite. Inspired by a new fashion in recent deep learning community, we propose a novel Transformer based model for pan-sharpening. We explore the potential of Transformer in image feature extraction and fusion. Following the successful development of vision transformers, we design a two-stream network with the self-attention to extract the modality-specific features from the PAN and MS modalities and apply a cross-attention module to merge the spectral and spatial features. The pan-sharpened image is produced from the enhanced fused features. Extensive experiments on GaoFen-2 and WorldView-3 images demonstrate that our Transformer based model achieves impressive results and outperforms many existing CNN based methods, which shows the great potential of introducing Transformer to the pan-sharpening task. Codes are available at https://github.com/zhysora/PanFormer. Huanyu Zhou, Qingjie Liu 0001, Yunhong Wang 0001 |
ICME | 2 |
| 2022 | SparseTT: Visual Tracking with Sparse TransformersabstractTransformers have been successfully applied to the visual tracking task and significantly promote tracking performance. The self-attention mechanism designed to model long-range dependencies is the key to the success of Transformers. However, self-attention lacks focusing on the most relevant information in the search regions, making it easy to be distracted by background. In this paper, we relieve this issue with a sparse attention mechanism by focusing the most relevant information in the search regions, which enables a much accurate tracking. Furthermore, we introduce a double-head predictor to boost the accuracy of foreground-background classification and regression of target bounding boxes, which further improve the tracking performance. Extensive experiments show that, without bells and whistles, our method significantly outperforms the state-of-the-art approaches on LaSOT, GOT-10k, TrackingNet, and UAV123, while running at 40 FPS. Notably, the training time of our method is reduced by 75% compared to that of TransT. The source code and models are available at https://github.com/fzh0917/SparseTT. Zhihong Fu, Zehua Fu, Qingjie Liu 0001, Yunhong Wang 0001 |
IJCAI | 3 |
| 2022 | Exploring Effective Knowledge Transfer for Few-shot Object DetectionabstractRecently, few-shot object detection(FSOD) has received much attention from the community, and many methods are proposed to address this problem from a knowledge transfer perspective. Though promising results have been achieved, these methods fail to achieve shot-stable:methods that excel in low-shot regimes are likely to struggle in high-shot regimes, and vice versa. We believe this is because the primary challenge of FSOD changes when the number of shots varies. In the low-shot regime, the primary challenge is the lack of inner-class variation. In the high-shot regime, as the variance approaches the real one, the main hindrance to the performance comes from misalignment between learned and true distributions. However, these two distinct issues remain unsolved in most existing FSOD methods. In this paper, we propose to overcome these challenges by exploiting rich knowledge the model has learned and effectively transferring them to the novel classes. For the low-shot regime, we propose a distribution calibration method to deal with the lack of inner-class variation problem. Meanwhile, a shift compensation method is proposed to compensate for possible distribution shift during fine-tuning. For the high-shot regime, we propose to use the knowledge learned from ImageNet as guidance for the feature learning in the fine-tuning stage, which will implicitly align the distributions of the novel classes. Although targeted toward different regimes, these two strategies can work together to further improve the FSOD performance. Experiments on both the VOC and COCO benchmarks show that our proposed method can significantly outperform the baseline method and produce competitive results in both low-shot settings(shot<5) and high-shot settings(shot>=5). Code is available at https://github.com/JulioZhao97/EffTrans_Fsdet.git. Qingjie Liu 0001, Yunhong Wang 0001 |
ACM Multimedia | 2 |
| 2022 | Sparse Relation Graph for Group Activity RecognitionabstractModeling relations between actors is critical for understanding group activities of dynamic scenes. Existing Group Activity Recognition (GAR) methods usually build strong connection in each actor pair. However, not all the connetions are necessary because not all actors are visible or related to each other. Based on this observation, we provide a Sparse Relation Graph (SRG) for GAR, in which the key relations are focused to mine more discriminative features. Then a graph convolutional network is designed for automatically learning the key relations. Extensive experiments on two popular group activity datasets, the Volleyball dataset and the Collective Activity dataset, demonstrate the effectiveness of our method. Especially in the Volleyball dataset, SRG can get better performance with less but delicate information. Zehua Fu, Qingjie Liu 0001, Yunhong Wang 0001, Xunxun Chen |
MMSP | 3 |
| 2022 | Learning from Future: A Novel Self-Training Framework for Semantic SegmentationabstractSelf-training has shown great potential in semi-supervised learning. Its core idea is to use the model learned on labeled data to generate pseudo-labels for unlabeled samples, and in turn teach itself. To obtain valid supervision, active attempts typically employ a momentum teacher for pseudo-label prediction yet observe the confirmation bias issue, where the incorrect predictions may provide wrong supervision signals and get accumulated in the training process. The primary cause of such a drawback is that the prevailing self-training framework acts as guiding the current state with previous knowledge because the teacher is updated with the past student only. To alleviate this problem, we propose a novel self-training strategy, which allows the model to learn from the future. Concretely, at each training step, we first virtually optimize the student (i.e., caching the gradients without applying them to the model weights), then update the teacher with the virtual future student, and finally ask the teacher to produce pseudo-labels for the current student as the guidance. In this way, we manage to improve the quality of pseudo-labels and thus boost the performance. We also develop two variants of our future-self-training (FST) framework through peeping at the future both deeply (FST-D) and widely (FST-W). Taking the tasks of unsupervised domain adaptive semantic segmentation and semi-supervised semantic segmentation as the instances, we experimentally demonstrate the effectiveness and superiority of our approach under a wide range of settings. Code is available at https://github.com/usr922/FST. Ye Du 0002, Yujun Shen, Jingjing Fei, Wei Li 0314, Rui Zhao 0001, Zehua Fu, Qingjie Liu 0001 |
NeurIPS | 9 |
| 2022 | PSGCNet: A Pyramidal Scale and Global Context Guided Network for Dense Object Counting in Remote-Sensing ImagesabstractObject counting, which aims to count the accurate number of object instances in images, has been attracting more and more attention. However, challenges such as large-scale variation, complex background interference, and nonuniform density distribution greatly limit the counting accuracy, particularly striking in remote-sensing imagery. To mitigate the above issues, this article proposes a novel framework for dense object counting in remote-sensing images, which incorporates a pyramidal scale module (PSM) and a global context module (GCM), dubbed PSGCNet, where PSM is used to adaptively capture multi-scale information and GCM is to guide the model to select suitable scales generated from PSM. Moreover, a reliable supervision manner improved from Bayesian and counting loss (BCL) is utilized to learn the density probability and then compute the count expectation at each annotation. It can relieve nonuniform density distribution to a certain extent. Extensive experiments on four remote-sensing counting datasets demonstrate the effectiveness of the proposed method and its superiority compared with state of the arts. Additionally, experiments extended on four commonly used crowd counting datasets further validate the generalization ability of the model. Code is available athttps://github.com/gaoguangshuai/psgcnet. Guangshuai Gao, Qingjie Liu 0001, Yunhong Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | MRDet: A Multihead Network for Accurate Rotated Object Detection in Aerial ImagesabstractObjects in aerial images usually have arbitrary orientations and are densely located over the ground, making them extremely challenge to be detected. Many of the recent developed methods attempt to solve these issues by estimating an extra orientation parameter and placing dense anchors, which will result in high model complexity and computational costs. In this article, we propose an arbitrary-oriented region proposal network (AO-RPN) to generate oriented proposals transformed from horizontal anchors. The AO-RPN is very efficient with only a few amounts of parameters increase than the original RPN. Furthermore, to obtain accurate bounding boxes, we decouple the detection task into multiple subtasks and propose a multihead network to accomplish them. Each head is specially designed to learn the features optimal for the corresponding task, which allows our network to detect objects accurately. We name it multihead rotated object detector (MRDet). We evaluate the performance of the proposed MRDet on two challenging benchmarks, i.e., DOTA and HRSC2016, and compare it with several state-of-the-art methods. Our method achieves very promising results, which clearly demonstrates its effectiveness. Code has been available athttps://github.com/qinr/MRDet. Ran Qin, Qingjie Liu 0001, Guangshuai Gao, Di Huang 0001, Yunhong Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Unsupervised Cycle-Consistent Generative Adversarial Networks for Pan SharpeningabstractDeep learning-based pan sharpening has received significant research interest in recent years. Most of the existing methods fall into the supervised learning framework in which they downsample the multispectral (MS) and panchromatic (PAN) images and regard the original MS images as ground truths to form training samples based on Wald’s protocol. Although impressive performance could be achieved, they have difficulties when generalizing to the original full-scale images due to the scale gap, which makes them lack of practicability. In this article, we propose an unsupervised generative adversarial framework that learns from the full-scale images without the ground truths to alleviate this problem. We first extract the modality-specific features from the PAN and MS images with a two-stream generator, perform fusion in the feature domain, and then reconstruct the pan-sharpened images. Furthermore, we introduce a novel hybrid loss based on the cycle-consistency and adversarial scheme to improve the performance. Comparison experiments with the state-of-the-art methods are conducted on GaoFen-2 (GF-2) and WorldView-3 satellites. Results demonstrate that the proposed method can greatly improve the pan-sharpening performance on the full-scale images, which clearly shows its practical value. Codes are available athttps://github.com/zhysora/UCGAN. Huanyu Zhou, Qingjie Liu 0001, Dawei Weng, Yunhong Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | STMTrack: Template-Free Visual Tracking With Space-Time Memory NetworksabstractBoosting performance of the offline trained siamese trackers is getting harder nowadays since the fixed information of the template cropped from the first frame has been almost thoroughly mined, but they are poorly capable of resisting target appearance changes. Existing trackers with template updating mechanisms rely on time-consuming numerical optimization and complex hand-designed strategies to achieve competitive performance, hindering them from real-time tracking and practical applications. In this paper, we propose a novel tracking framework built on top of a space-time memory network that is competent to make full use of historical information related to the target for better adapting to appearance variations during tracking. Specifically, a novel memory mechanism is introduced, which stores the historical information of the target to guide the tracker to focus on the most informative regions in the current frame. Furthermore, the pixel-level similarity computation of the memory network enables our tracker to generate much more accurate bounding boxes of the target. Extensive experiments and comparisons with many competitive trackers on challenging large-scale benchmarks, OTB-2015, TrackingNet, GOT-10k, LaSOT, UAV123, and VOT2018, show that, without bells and whistles, our tracker outperforms all previous state-of-the-art real-time methods while running at 37 FPS. The code is available at https: //github.com/fzh0917/STMTrack. Zhihong Fu, Qingjie Liu 0001, Zehua Fu, Yunhong Wang 0001 |
CVPR | 2 |
| 2021 | Co-Saliency Detection With Co-Attention Fully Convolutional NetworkabstractCo-saliency detection aims to detect common salient objects from a group of relevant images. Some attempts have been made with the Fully Convolutional Network (FCN) framework and achieve satisfactory detection results. However, due to stacking convolution layers and pooling operation, the boundary details tend to be lost. In addition, existing models often utilize the extracted features without discrimination, leading to redundancy in representation since actually not all features are helpful to the final prediction and some even bring distraction. In this paper, we propose a co-attention module embedded FCN framework, called as Co-Attention FCN (CA-FCN). Specifically, the co-attention module is plugged into the high-level convolution layers of FCN, which can assign larger attention weights on the common salient objects and smaller ones on the background and uncommon distractors to boost final detection performance. Extensive experiments on three popular co-saliency benchmark datasets demonstrate the superiority of the proposed CA-FCN, which outperforms state-of-the-arts in most cases. Besides, the effectiveness of our new co-attention module is also validated with ablation studies. Guangshuai Gao, Wenting Zhao 0007, Qingjie Liu 0001, Yunhong Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Counting From Sky: A Large-Scale Data Set for Remote Sensing Object Counting and a Benchmark MethodabstractObject counting, whose aim is to estimate the number of objects from a given image, is an important and challenging computation task. Significant efforts have been devoted to addressing this problem and achieved great progress, yet counting the number of ground objects from remote sensing images is barely studied. In this article, we are interested in counting dense objects from remote sensing images. Compared with object counting in a natural scene, this task is challenging in the following factors: large-scale variation, complex cluttered background, and orientation arbitrariness. More importantly, the scarcity of data severely limits the development of research in this field. To address these issues, we first construct a large-scale object counting data set with remote sensing images, which contains four important geographic objects: buildings, crowded ships in harbors, and large vehicles and small vehicles in parking lots. We then benchmark the data set by designing a novel neural network that can generate a density map of an input image. The proposed network consists of three parts, namely attention module, scale pyramid module, and deformable convolution module (DCM) to attack the aforementioned challenging factors. Extensive experiments are performed on the proposed data set and one crowd counting data set, which demonstrates the challenges of the proposed data set and the superiority and effectiveness of our method compared with state-of-the-art methods. Guangshuai Gao, Qingjie Liu 0001, Yunhong Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | PSGAN: A Generative Adversarial Network for Remote Sensing Image Pan-SharpeningabstractThis article addresses the problem of remote sensing image pan-sharpening from the perspective of generative adversarial learning. We propose a novel deep neural network-based method named pansharpening GAN (PSGAN). To the best of our knowledge, this is one of the first attempts at producing high-quality pan-sharpened images with generative adversarial networks (GANs). The PSGAN consists of two components: a generative network (i.e., generator) and a discriminative network (i.e., discriminator). The generator is designed to accept panchromatic (PAN) and multispectral (MS) images as inputs and maps them to the desired high-resolution (HR) MS images, and the discriminator implements the adversarial training strategy for generating higher fidelity pan-sharpened images. In this article, we evaluate several architectures and designs, namely, two-stream input, stacking input, batch normalization layer, and attention mechanism to find the optimal solution for pan-sharpening. Extensive experiments on QuickBird, GaoFen-2, and WorldView-2 satellite images demonstrate that the proposed PSGANs not only are effective in generating high-quality HR MS images and superior to state-of-the-art methods but also generalize well to full-scale images. Qingjie Liu 0001, Huanyu Zhou, Qizhi Xu, Yunhong Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Counting Dense Objects in Remote Sensing ImagesabstractEstimating accurate number of interested objects from a given image is a challenging yet important task. Significant efforts have been made to address this problem and achieve great progress, yet counting number of ground objects from remote sensing images is barely studied. In this paper, we are interested in counting dense objects from remote sensing images. Compared with object counting in natural scene, this task is challenging in following factors: large scale variation, complex cluttered background and orientation arbitrariness. More importantly, the scarcity of data severely limits the development of research in this field. To address these issues, we first construct a large-scale object counting dataset based on remote sensing images, which contains four kinds of objects: buildings, crowded ships in harbor, large-vehicles and small-vehicles in parking lot. We then benchmark the dataset by designing a novel neural network which can generate density map of an input image. The proposed network consists of three parts namely convolution block attention module (CBAM), scale pyramid module (SPM) and deformable convolution module (DCM). Experiments on the proposed dataset and comparisons with state of the art methods demonstrate the challenging of the proposed dataset, and superiority and effectiveness of our method. Guangshuai Gao, Qingjie Liu 0001, Yunhong Wang 0001 |
ICASSP | 2 |
| 2020 | Unsupervised Conditional Disentangle Network For Image DehazingabstractImage dehazing aims to restore the blurry image information caused by the ambiguities of unknown scene radiance and transmission. Instead of using paired images or depth information, we propose an Unsupervised Conditional Disentangle Network (UCDN) using unpaired dataset. Our approach enforces the constraint by introducing physical-based disentanglement. Unlike other unsupervised dehazing models, our approach adapts the multi-concentration of fog and outperforms on the dataset with different concentrations. Extensive experiments on synthesized dataset demonstrate that our approach can surpass state-of-the-arts. Meanwhile, through benchmarking on our collected natural hazy dataset, our approach can generate more perceptually appealing dehazing results. Yizhou Jin, Guangshuai Gao, Qingjie Liu 0001, Yunhong Wang 0001 |
ICIP | 3 |
| 2020 | Pan-Sharpening with a CNN-Based Two Stage Ratio Enhancement MethodabstractWe propose a hybrid method combining the deep learning technique and the ratio enhancement (RE) method for pansharpening. The intuition behind is to utilize the deep learning technique to synthesize a panchromatic (PAN) image for the RE method to reduce the spectral distortion while keeping the spatial details. The method consists of two stages. First, the CNN synthesizer is optimized to generate the downsampled PAN image to guarantee the network have a good initialization. Second, CNN is integrated into the RE method and supervised by the ground truth multi-spectral (MS) to produce an ideal synthesized PAN for the RE method. We conduct experiments on various datasets and compare with widely used methods to demonstrate the superiority of the proposed method. Huanyu Zhou, Qingjie Liu 0001, Qizhi Xu, Yunhong Wang 0001 |
IGARSS | 2 |
| 2020 | Building Detection via Complementary Convolutional Features of Remote Sensing Images
Zeshan Lu, Kun Liu 0022, Zhen Liu 0032, Jiwen Dong, Qingjie Liu 0001, Tao Xu 0021 |
PRCV (1) | 6 |
| 2020 | From W-Net to CDGAN: Bitemporal Change Detection via Deep Learning TechniquesabstractTraditional change detection methods usually follow the image differencing, change feature extraction, and classification framework, and their performance is limited by such simple image domain differencing and also the hand-crafted features. Recently, the success of deep convolutional neural networks (CNNs) has widely spread across the whole field of computer vision for their powerful representation abilities. Therefore, in this article, we address the remote sensing image change detection problem with deep learning techniques. We first propose an end-to-end dual-branch architecture, termed the W-Net, with each branch taking as input one of the two bitemporal images as in the traditional change detection models. In this way, CNN features with more powerful representative abilities can be obtained to boost the final detection performance. In addition, W-Net performs differencing in the feature domain rather than in the traditional image domain, which greatly alleviates loss of useful information for determining the changes. Furthermore, by reformulating change detection as an image translation problem, we apply the recently popular generative adversarial network (GAN) in which our W-Net serves as the generator, leading to a new GAN architecture for change detection which we call CDGAN. To train our networks and also facilitate future research, we construct a large scale data set by collecting images from Google Earth and provide carefully manually annotated ground truths. Experiments show that our proposed methods can provide fine-grained change detection results superior to the existing state-of-the-art baselines. Bin Hou, Qingjie Liu 0001, Yunhong Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | 5M-Building: A Large-Scale High-Resolution Building Dataset with CNN Based Detection AnalysisabstractBuilding detection in remote sensing images plays an important role in applications such as urban management and urban planning. Recently, convolutional neural network (CNN) based methods which benefits from the popularity of large-scale datasets have achieved good performance for object detection. To our best knowledge, there is no large-scale remote sensing image dataset specially build for building detection. Existing building datasets are in small size and lack of diversity, which hinder the development of building detection. In this paper, we present a large-scale high-resolution building dataset named 5M-Building after the number of samples in the dataset. The dataset consists of more than 10 thousand images all collected from GaoFen-2 with a spatial resolution of 0.8 meter. We also present a baseline for the dataset by evaluating three state of the art CNN based detectors. The experiments demonstrate that it is great challenge to accurately detect various buildings from remote sensing images. We hope the 5M-Building dataset will facilitate the research on building detection. Zeshan Lu, Tao Xu 0021, Kun Liu 0022, Zhen Liu 0032, Feipeng Zhou, Qingjie Liu 0001 |
ICTAI | 6 |
| 2019 | Image Spectral Data Classification Using Pixel-Purity Kernel Graph Cuts and Support Vector Machines: A Case Study of Vegetation Identification in Indian Pine Experimental AreaabstractSalt and pepper phenomenon of pixel-based images classification, has a major negative impacts on the accuracy of Imaging Spectra classification. Various kernel-based methods, such as Kernel graph cuts (KGC) and support vector machine (SVM), are used to solve the nonlinear problems by mapping the original nonlinear data into higher dimensional space. Four experiment schemes, including Original-Pixel-SVM (OPSVM), Original-PKGC-SVM (OPKGCSVM), PCA-PKGC-SVM (PPKGCSVM), and MNF-PKGC-SVM (MPKGCSVM), are designed to class AVRIS in India Pine of USA for comparison of classification User Accuracy (UA), Producer Accuracy (PA), Overall Accuracy (OA) and Kappa index quantitatively. The average UAs of MPKGCSVM, PPKGCSVM, OPKGCSVM and OPSVM are 91.92%, 83.09%, 84.51% and 75.55%, the average PAs are 95.33%, 91.47%, 88.03% and 87.78% their OAs are 93.57%,88.99%,85.35% and 82.36%, their Kappa indexes are 0.92,0.85,0.83 and 0.79 respectively. From MPKGCSVM to OPSVM, the OA and Kappa indexes are improved 11.21% and 0.13 respectively. Therefore, PKGC reduce the salt-and-pepper effects of classification obviously, and improve the accuracy and robustness greatly. Besides, dimensionality reduction pre-processing before PKGA of HSI with Minimum noise fraction can enhance the performance of final classification than other transformations. Weijie Jia, Qingjie Liu 0001, Fengxian Miao |
IGARSS | 3 |
| 2019 | A Robust Multi-Athlete Tracking Algorithm by Exploiting Discriminant Features and Long-Term Dependencies
Nan Ran, Longteng Kong, Yunhong Wang 0001, Qingjie Liu 0001 |
MMM (1) | 4 |
| 2018 | Hough Transform Guided Deep Feature Extraction for Dense Building Detection in Remote Sensing ImagesabstractDetecting dense buildings without elevation information is an important and challenging task in remote sensing applications. In this paper, we present a novel cascaded deep neural network architecture, incorporating multi -stage region proposal detection and Hough transform to obtain better mid-level semantic information for man-made objects. This proposed network can be trained end-to-end by multi-loss jointly. We train and test it on a large building dataset collected from Google Earth, including buildings from urban, suburban and rural areas. Experiments demonstrate great robustness and superiority of our method to various buildings over other convolutional neural network (CNN) based detection methods. Qingpeng Li, Yunhong Wang 0001, Qingjie Liu 0001, Wei Wang 0115 |
ICASSP | 3 |
| 2018 | Psgan: A Generative Adversarial Network for Remote Sensing Image Pan-SharpeningabstractRemote sensing image fusion (also known as pan-sharpening) aims to generate a high resolution multi -spectral image from inputs of a high spatial resolution single band panchromatic (PAN) image and a low spatial resolution multi-spectral (MS) image. In this paper, we propose PSGAN, a generative adversarial network (GAN) for remote sensing image pansharpening. To the best of our knowledge, this is the first attempt at producing high quality pan-sharpened images with GANs. The PSGAN consists of two parts. Firstly, a two-stream fusion architecture is designed to generate the desired high resolution multi -spectral images, then a fully convolutional network serving as a discriminator is applied to distinct “real” or “pan-sharpened” MS images. Experiments on images acquired by Quickbird and GaoFen-1 satellites demonstrate that the proposed PSGAN can fuse PAN and MS images effectively and significantly improve the results over the state of the art traditional and CNN based pan-sharpening methods. Yunhong Wang 0001, Qingjie Liu 0001 |
ICIP | 3 |
| 2018 | Hierarchical Region Based Convolution Neural Network for Multiscale Object Detection in Remote Sensing ImagesabstractIn this paper, we propose a novel Faster R-CNN based method to detect multiscale objects in very high resolution optical remote sensing images. Firstly, a pre-trained CNN is used to extract features from an input image; and then a set of object candidates are generated. To efficiently detect objects with various scales, we design a hierarchical selective filtering (HSF) layer to map features in different scales to the same scale space. The HSF layer can be applied on both region proposal and the subsequent detection network. More importantly, it can be plugged into Faster R-CNN network without modifying its architecture, meanwhile boosting the performance on detecting objects with varying scales. The proposed model can be trained in an end-to-end manner. We test our network on three datasets containing different multiscale objects, including airplanes, ships and buildings, which are collected from Google Earth images and GaoFen-2 images. Experiments demonstrate high precision and robustness of our method. Qingpeng Li, Lichao Mou, Kaiyu Jiang, Qingjie Liu 0001, Yunhong Wang 0001, Xiao Xiang Zhu 0001 |
IGARSS | 4 |
| 2018 | Remote Sensing Image Fusion Based on Two-Stream Fusion Network
Yunhong Wang 0001, Qingjie Liu 0001 |
MMM (1) | 3 |
| 2018 | Road Extraction by Deep Residual U-NetabstractRoad extraction from aerial images has been a hot research topic in the field of remote sensing image analysis. In this letter, a semantic segmentation neural network, which combines the strengths of residual learning and U-Net, is proposed for road area extraction. The network is built with residual units and has similar architecture to that of U-Net. The benefits of this model are twofold: first, residual units ease training of deep networks. Second, the rich skip connections within the network could facilitate information propagation, allowing us to design networks with fewer parameters, however, better performance. We test our network on a public road data set and compare it with U-Net and other two state-of-the-art deep-learning-based road extraction methods. The proposed approach outperforms all the comparing methods, which demonstrates its superiority over recently developed state of the arts. Qingjie Liu 0001, Yunhong Wang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2018 | HSF-Net: Multiscale Deep Feature Embedding for Ship Detection in Optical Remote Sensing ImageryabstractShip detection is an important and challenging task in remote sensing applications. Most methods utilize specially designed hand-crafted features to detect ships, and they usually work well only on one scale, which lack generalization and impractical to identify ships with various scales from multiresolution images. In this paper, we propose a novel deep feature-based method to detect ships in very high-resolution optical remote sensing images. In our method, a regional proposal network is used to generate ship candidates from feature maps produced by a deep convolutional neural network. To efficiently detect ships with various scales, a hierarchical selective filtering layer is proposed to map features in different scales to the same scale space. The proposed method is an end-to-end network that can detect both inshore and offshore ships ranging from dozens of pixels to thousands. We test our network on a large ship data set which will be released in the future, consisting of Google Earth images, GaoFen-2 images, and unmanned aerial vehicle data. Experiments demonstrate high precision and robustness of our method. Further experiments on aerial images show its good generalization to unseen scenes. Qingpeng Li, Lichao Mou, Qingjie Liu 0001, Yunhong Wang 0001, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2017 | Feature map pooling for cross-view gait recognition based on silhouette sequence imagesabstractIn this paper, we develop a novel convolutional neural network based approach to extract and aggregate useful information from gait silhouette sequence images instead of simply representing the gait process by averaging silhouette images. The network takes a pair of arbitrary length sequence images as inputs and extracts features for each silhouette independently. Then a feature map pooling strategy is adopted to aggregate sequence features. Subsequently, a network which is similar to Siamese network is designed to perform recognition. The proposed network is simple and easy to implement and can be trained in an end-to-end manner Cross-view gait recognition experiments are conducted on OU-ISIR large population dataset. The results demonstrate that our network can extract and aggregate features from silhouette sequence effectively. It also achieves significant equal error rates and comparable identification rates when compared with the state of the art. Yunhong Wang 0001, Zheng Liu 0014, Qingjie Liu 0001, Di Huang 0001 |
IJCB | 4 |
| 2017 | Visual and textual sentiment analysis using deep fusion convolutional neural networksabstractSentiment analysis is attracting more and more attentions and has become a very hot research topic due to its potential applications in personalized recommendation, opinion mining, etc. Most of the existing methods are based on either textual or visual data and can not achieve satisfactory results, as it is very hard to extract sufficient information from only one single modality data. Inspired by the observation that there exists strong semantic correlation between visual and textual data in social medias, we propose an end-to-end deep fusion convolutional neural network to jointly learn textual and visual sentiment representations from training examples. The two modality information are fused together in a pooling layer and fed into fully-connected layers to predict the sentiment polarity. We evaluate the proposed approach on two widely used data sets. Results show that our method achieves promising result compared with the state-of-the-art methods which clearly demonstrate its competency. Xingyue Chen, Yunhong Wang 0001, Qingjie Liu 0001 |
ICIP | 3 |
| 2017 | Change Detection Based on Deep Features and Low RankabstractIn this letter, we address the problem of change detection for remote sensing images from the perspective of visual saliency computation. The proposed method incorporates low-rank-based saliency computation and deep feature representation. First, multilevel convolutional neural network (CNN) features are extracted for superpixels generated using SLIC, in which a fixed-size CNN feature can be formed to represent each superpixel. Then, low-rank decomposition is applied to the change features of the two input images to generate saliency maps that indicate change probabilities of each pixel. Finally, binarized change map can be obtained with a simple threshold. To deal with scale variations, a multiscale fusion strategy is employed to produce more reliable detection results. Extensive experiments on Google Earth and GF-2 images demonstrate the feasibility and effectiveness of the proposed method. Bin Hou, Yunhong Wang 0001, Qingjie Liu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2016 | A genetic-optimized multi-angle normalized cross correlation SIFT for automatic remote sensing registrationabstractA new method of remote sensing image registration is proposed to reduce the adverse effects on the image registration caused by the rotation transform, based on genetic-optimized multi-angle normalized cross correlation (GMNCC) to get more matched scale invariant feature transform (SIFT) feature points. The GMNCC can detect the angle-offsets (AOs) of object images to reference image by determining the maximum of correlation coefficients between the two images with genetic algorithm, and complete the rotation offset correction. Then SIFT is used to extract the feature points and feature matching, which is subsequently refined by RANSAC to eliminate the false matched control points. Two Woldview-2 (WV2) images of Beijing Olympic Forest Park were used for testing the GMNCC-SIFT registration. GMNCC detects more accurate rotation angle-offsets than multi-angle normalized cross correlation (MANCC), reduces the detection process from 3000s to 2500s, and gives more matched points than simple SIFT to improve registration accuracy. Qingjie Liu 0001, Linhai Jing, Fengxian Miao |
IGARSS | 2 |
| 2016 | Efficient sky segmentation approach for small UAV autonomous obstacles avoidance in cluttered environmentabstractNowadays, Unmanned Air Systems (UAS) or Unmanned Air Vehicles (UAV) plays an important role in different critical missions. UAVs must have the ability to perform different kinds of missions that may be, it is impossible to be performed by the human operator. Moreover, UAVs can reach various areas in different weather and environmental conditions. As a result, UAV should be equipped with an efficient autonomous obstacle avoidance system. In this paper, we are going to introduce an efficient approach for sky segmentation in a cluttered environment that is considered as a vital step for UAV autonomous obstacle avoidance. From experimental results, the proposed sky segmentation approach gives promising results for significantly sky segmentation and preserving the potential obstacles. Ahmed S. Mashaly, Yunhong Wang 0001, Qingjie Liu 0001 |
IGARSS | 3 |
| 2016 | CNN based suburban building detection using monocular high resolution Google Earth imagesabstractThis paper proposes a deep convolutional neural networks (CNNs) based method to automatically detect suburban buildings from high resolution Google Earth imagery. Traditional methods based on low-level hand-engineered features or mid-level bag of features have great limitations in complex environment, especially in suburban areas. Inspired by the astounding achievement of CNNs in object recognition and detection, we develop a novel method to detect buildings in cluttered images which consists of three main steps. Firstly, a multi-scale saliency computation is employed to extract built-up areas and a sliding windows approach is applied to generate candidate regions. Then, a CNN is applied to classify the regions. Finally, an improved non maximum suppression is used to remove false buildings. We test our method on a collection of very challenging Google Earth images and achieve 89% precision, which shows robustness and efficiency of our method. Qinchuan Zhang, Yunhong Wang 0001, Qingjie Liu 0001, Wei Wang 0115 |
IGARSS | 3 |
| 2015 | Object-based feature extraction and semi-supervised classification for urban change detection using high-resolution remote sensing imagesabstractThis paper presents a novel approach for urban change detection of high resolution (HR) remote sensing images. To overcome deficiency of traditional pixel-based methods and better annotate HR images, object-based strategies are adopted. Firstly change vector analysis (CVA) and local binary patterns (LBP) are utilized to extract the object-specific features based on the image-objects acquired by multitemporal segmentation. Then sparse representation is further exploited to characterize highly effective sparse features. Finally, the final change map is obtained by support vector machine (SVM) with the pseudotraining set acquired by expectation maximization (EM). Comparative experiments demonstrate the effectiveness of the proposed method. Bin Hou, Qingjie Liu 0001, Yunhong Wang 0001 |
IGARSS | 2 |
| 2015 | A new region growing-based segmentation method for high resolution remote sensing imageryabstractIn this article, a newsegmentation method based on traditional region growing (RG) is proposed for high resolutionremote sensing imagery. This method takes regional minima from horizontal and vertical gradient maps of the image as seeds for the following region growing processing. The new method consists of several steps as follows: (1)deriving a morphological gradient map from the input multispectral image, (2) morphologically filtering the gradient image to remove local minima with small depthand extracting regional minima of flat areas in the resulting filtered image as seeds, (3) segmenting the multispectral image using the RG approach with reference to the seeds, and (4) merge the resulting initial segments to yield asegmentation map. In a test with a WorldView-2 multispectral image, the proposed method offered segmentation maps with nearly the same accuracy as several current methods. Xiuxia Li, Linhai Jing, Qizhong Lin, Hui Li 0008, Ru Xu, Yunwei Tang, Haifeng Ding, Qingjie Liu 0001 |
IGARSS | 8 |
| 2015 | Assessment of pan-sharpening methods applied to WorldView-2 image fusionabstractVarious multispectral (MS) and panchromatic (PAN) fusion (or pan-sharpening) algorithms were developed to produce an enhanced MS image of high spatial resolution. Regarding the novelty in both the PAN and MS bands of the WV-2 imagery, the objective of this study is to assess the performance of nine state-of-the-art pan-sharpening methods for the WV-2 imagery, using both image quality indices and information indices that used for urban information extraction. The comparison of the four quality indices (RASE, ERGAS, SAM, and Q4) demonstrated that the HR method performed the best for the WV-2 MS and PAN images. However, the comparison of the four information indices showed that a higher quality at data level does not signify better information preservation for object recognition. Hui Li 0008, Linhai Jing, Yunwei Tang, Qingjie Liu 0001, Haifeng Ding, Zhongchang Sun, Yu Chen 0057 |
IGARSS | 4 |
| 2014 | A novel multi-resolution segmentation algorithm for highresolution remote sensing imagery based on minimum spanning tree and minimum heterogeneity criterionabstractImage segmentation is the basis of object-based information extraction from remote sensing imagery. Image segmentation based on multiple features, multi-resolution, and spatial context is one current research focus. Combining graph theory based optimization with the multi-scale image segmentation framework of the eCognition software, a multi-scale image segmentation method is proposed in this paper. In this method, a coherent enhancement anisotropic diffusion filtering approach and a minimum spanning tree segmentation algorithm are employed to initially segment the image. After that, the resulting segments are merged regarding minimum heterogeneity criteria, which are based on both the spectral characteristics and the shape parameters of segments. Two test images were used for visual and quantitative comparisons of the proposed method with the multi-scale segmentation method FNEA employed in the eCognition software. The results show that the proposed method is effective, and is more sensitive to subtle spectral differences than the FNEA. Hui Li 0008, Yunwei Tang, Qingjie Liu 0001, Haifeng Ding, Linhai Jing, Qizhong Lin |
IGARSS | 3 |
| 2014 | A multiple-point geostatistical method for digital elevation models conflationabstractA data conflation method was developed based on a multiple-point geostatistical method. Geostatistics can quantify cross-correlation from different sources of data when integrating geospatial information. Multiple-point geostatistics (MPG) is a development of geostatistics. Pattern-based MPG can capture similar patterns using spatial correlation in the form of training image, and then reproduces the area of interest in a local window using sequential simulation. In the proposed method, MPG simulation was applied to the traditional geostatistical prediction. This method was tested on digital elevation models (DEMs). The aim was to simulate an image at a fine spatial resolution by conflating sparsely sampled elevation data and digital raster elevation at a coarser spatial resolution. The MPG simulation approach was compared to traditional geostatistical conflation using four different methods. The results show that the proposed method can achieve a more precise prediction than the benchmarks. Yunwei Tang, Jingxiong Zhang, Hui Li 0008, Haifeng Ding, Qingjie Liu 0001, Linhai Jing |
IGARSS | 5 |
| 2014 | Pan-sharpening based on weighted red black waveletsabstractPan‐sharpening is a technique which provides an efficient and economical solution to generate multi‐spectral (MS) images with high‐spatial resolution by fusing spectral information in MS images and spatial information in panchromatic (PAN) image. In this study, the authors propose a new pan‐sharpening method based on weighted red‐black (WRB) wavelets and adaptive principal component analysis (PCA), where the usage of WRB wavelet decomposition is to extract the spatial details in PAN image and the adaptive PCA is used to select the adequate principal component for injecting spatial details. WRB wavelets are data‐dependent second generation wavelets. Multi‐resolution analysis (MRA) based on WRB wavelet transform shows a better de‐correlation of the data compared with common linear translation‐invariant MRA, which makes it suitable for applications requiring manipulating image details. A local processing strategy is introduced to reduce the artefact effects and spectral distortions in the pan‐sharpened images. The proposed method is evaluated on the datasets acquired by QuickBird, IKONOS and Landsat‐7 ETM + satellites and compared with existing methods. Experimental results demonstrate that the authors method can provide promising fused MS images with high‐spatial resolution. Qingjie Liu 0001, Yunhong Wang 0001, Zhaoxiang Zhang 0001 |
IET Image Process. | 1 |
| 2012 | Locally linear embedding based example learning for pan-sharpening
Qingjie Liu 0001, Yunhong Wang 0001, Zhaoxiang Zhang 0001 |
ICPR | 1 |
| 2012 | Pan-sharpening using weighted red-black wavelet
Qingjie Liu 0001, Yunhong Wang 0001, Zhaoxiang Zhang 0001 |
ICPR | 1 |