EDBT 2026 Demo / reviewers in the wild / expert
Fan Liu 0003
dblp:56/2849-3
· DBLP profile ↗
65ranked-venue papers
18as first author
50since 2021 · last 2026
0000-0001-8746-9845ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 7 first-author · 18 since 2021Artificial intelligence and machine learning · 26 · 7 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 3 first-author · 13 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AirNavigation: Let UAV Navigation Tell Its Own StoryabstractTesting autonomous navigation algorithms of Unmanned Aerial Vehicles (UAVs) in real-world scenarios often entails significant safety risks. In this paper, we aim to build a flexible yet user-friendly UAV autonomous navigation simulator. Ideally, it should closely emulate real-world environments, support diverse UAV models and algorithms, and provide a flexible evaluation framework. Existing frameworks fail to satisfy all three requirements simultaneously. To this end, we present AirNavigation, an integrated simulation platform designed to support the end-to-end workflow of UAV navigation research. Specifically, our system leverages Unreal Engine to simulate highly realistic environments and diverse UAV models. It further facilitates semi-automated scene generation and multi-modal synthetic training data production. To lower the barrier of adoption, we develop a suite of user-friendly interfaces to enable seamless integration of diverse navigation algorithms. Moreover, we introduce a novel evaluation system powered by large language models to deliver personalized and fine-grained performance analysis. Jianyu Jiang, Zequan Wang, Liang Yao 0001, Shengxiang Xu, Fan Liu 0003 |
AAAI | 5 |
| 2026 | RemoteReasoner: Towards Unifying Geospatial Reasoning WorkflowabstractRemote sensing imagery presents vast, inherently unstructured spatial data, necessitating sophisticated reasoning to interpret complex user intents and contextual relationships beyond simple recognition tasks. In this paper, we aim to construct an Earth observation workflow to handle complex queries by reasoning about spatial context and user intent. As a reasoning workflow, it should autonomously explore and construct its own inference paths, rather than being confined to predefined ground‑truth sequences. Ideally, its architecture ought to be unified yet generalized, possessing capabilities to perform diverse reasoning tasks through one model without requiring additional fine-tuning. Existing remote sensing approaches rely on supervised fine-tuning paradigms and task‑specific heads, limiting both autonomous reasoning and unified generalization. To this end, we propose RemoteReasoner, a unified workflow for geospatial reasoning. The design of RemoteReasoner integrates a multi-modal large language model (MLLM) for interpreting user instructions and localizing targets, together with task transformation strategies that enable multi-granularity tasks, including object-, region-, and pixel-level. In contrast to existing methods, our framework is trained with reinforcement learning (RL) to endow the MLLM sufficient reasoning autonomy. At the inference stage, our transformation strategies enable diverse task output formats without requiring task-specific decoders or further fine-tuning. Experiments demonstrated that RemoteReasoner achieves state-of-the-art performance across multi-granularity reasoning tasks. Furthermore, it retains the MLLM's inherent generalization capability, demonstrating robust performance on unseen tasks and categories. Liang Yao 0001, Fan Liu 0003, Hongbo Lu, Chuanyi Zhang, Shengxiang Xu, Shimin Di |
AAAI | 2 |
| 2026 | SGPVT: Self-Generated Proximal Visual Tokens for Mitigating Proximal Collateral Damage in MLLM UnlearningabstractJiaqi Li, Zhijing Zhang, Jiahui Geng, Sheng Bi, Chuanyi Zhang, Fan Liu, Guilin Qi. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiaqi Li 0031, Zhijing Zhang, Jiahui Geng, Chuanyi Zhang, Fan Liu 0003, Guilin Qi |
ACL (1) | 6 |
| 2026 | E2SGNN: Reconciling Expression and Efficiency in Spiking Graph Neural NetworkabstractBy mimicking the brain's efficient spiking encoding paradigm, spiking graph neural networks exhibit significant potential for efficient graph data analysis. Due to the inherent expressive limitations of binary spiking signals adopted in spiking encoding, existing models typically enhance their expression by integrating numerous real-valued multiplication-additions or high-latency encoding. However, such integrations compromise the core efficiency superiority of spiking models, limiting their scalability in real-world applications. To simultaneously reconcile considerable expression and efficiency, we propose E2SGNN, a novel network comprising a dual-scale modulated spiking backbone and a latency-dynamic optimization module. The former backbone integrates global and local real-valued graph modulations into spiking graph convolution, enabling discriminative dual-scale neighbor embedding in the encoding process. It both breaks through binary spiking signals' expressive limitations and improves the content expressiveness of spiking graph representations, while retaining low-latency and addition-only efficient advantages. Moreover, to further reduce the latency redundancy for higher efficiency, the latter module adaptively customizes the latency for each graph data based on data complexity. In this way, our network can finally generate graph representations expressively and efficiently. Experiments on various datasets demonstrate the superiority of our network in expression and efficiency. Xu Yang 0019, Cheng Deng 0002, Fan Liu 0003 |
WWW | 4 |
| 2026 | Two-stream attentive spatial-temporal graph convolutional network for P300 detection in brain-computer interface
Jincen Wang, Yan Zhao 0037, Cunhang Fan, Yong Li 0032, Fan Liu 0003, Hailun Lian, Cheng Lu 0005 |
Expert Syst. Appl. | 5 |
| 2026 | Frequency-spatial decoupled co-modeling transformer for fine-grained remote sensing image segmentation
Xin Li 0090, Shangtuo Qian, Xin Lyu 0001, Yongze Song, Fan Liu 0003, Yiwei Fang, Zhennan Xu, André Kaup |
Inf. Sci. | 5 |
| 2026 | Image restoration model compression via mamba-oriented heterogeneous knowledge distillation
Sai Yang, Bin Hu 0023, Xiaoxin Wu 0004, Fan Liu 0003, Wanzhi Wen |
Neural Networks | 4 |
| 2026 | Heterogeneous Knowledge Distillation Fostered Pre-training for remote sensing object detection
Fan Liu 0003, Chuanyi Zhang, Liang Yao 0001 |
Pattern Recognit. | 1 |
| 2025 | Making Large Vision Language Models to Be Good Few-Shot LearnersabstractFew-shot classification (FSC) is a fundamental yet challenging task in computer vision that involves recognizing novel classes from limited data. While previous methods have focused on enhancing visual features or incorporating additional modalities, Large Vision Language Models (LVLMs) offer a promising alternative due to their rich knowledge and strong visual perception. However, LVLMs risk learning specific response formats rather than effectively extracting useful information from support data in FSC. In this paper, we investigate LVLMs' performance in FSC and identify key issues such as insufficient learning and the presence of severe position biases. To tackle above challenges, we adopt the meta-learning strategy to teach models ``learn to learn". By constructing a rich set of meta-tasks for instruction fine-tuning, LVLMs enhance the ability to extract information from few-shot support data for classification. Additionally, we further boost LVLM's few-shot learning capabilities through label augmentation (LA) and candidate selection (CS) in the fine-tuning and inference stages, respectively. LA is implemented via a character perturbation strategy to ensure the model focuses on support information. CS leverages attribute descriptions to filter out unreliable candidates and simplify the task. Extensive experiments demonstrate that our approach achieves superior performance on both general and fine-grained datasets. Furthermore, our candidate selection strategy has been proven beneficial for training-free LVLMs. Fan Liu 0003, Wenwen Cai, Jian Huo, Chuanyi Zhang, Delong Chen |
AAAI | 1 |
| 2025 | Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image CaptionsabstractWhile densely annotated image captions significantly facilitate the learning of robust visionlanguage alignment, methodologies for systematically optimizing human annotation efforts remain underexplored.We introduce CHAIN-OF-TALKERS (COTALK), an AI-in-the-loop methodology designed to maximize the number of annotated samples and improve their comprehensiveness under fixed budget constraints (e.g., total human annotation time).The framework is built upon two key insights.First, sequential annotation reduces redundant workload compared to conventional parallel annotation, as subsequent annotators only need to annotate the "residual"-the missing visual information that previous annotations have not covered.Second, humans process textual input faster by reading while outputting annotations with much higher throughput via talking; thus a multimodal interface enables optimized efficiency.We evaluate our framework from two aspects: intrinsic evaluations that assess the comprehensiveness of semantic units, obtained by parsing detailed captions into object-attribute trees and analyzing their effective connections; extrinsic evaluation measures the practical usage of the annotated captions in facilitating vision-language alignment.Experiments with eight participants show our CHAIN-OF-TALKERS (CoTalk) improves annotation speed (0.42 vs. 0.30 units/sec) and retrieval performance (41.13% vs. 40.52%)over the parallel method.per minute? a review and meta-analysis of reading rate. Delong Chen, Fan Liu 0003, Chuanyi Zhang, Liang Yao 0001, Yuhui Zheng |
EMNLP | 3 |
| 2025 | Prompting DirectSAM for Semantic Contour Extraction in Remote Sensing ImagesabstractThe Direct Segment Anything Model (DirectSAM) excels in class-agnostic contour extraction. In this paper, we explore its use by applying it to optical remote sensing imagery, where semantic contour extraction—such as identifying buildings, road networks, and coastlines-holds significant practical value. Those applications are currently handled via training specialized small models separately on small datasets in each domain. We introduce a foundation model derived from DirectSAM, termed DirectSAM-RS, which not only inherits the strong segmentation capability acquired from natural images, but also benefits from a large-scale dataset we created for remote sensing semantic contour extraction. This dataset comprises over 34k image-text-contour triplets, making it at least 30 times larger than individual dataset. DirectSAM-RS integrates a prompter module: a text encoder and cross-attention layers attached to the DirectSAM architecture, which allows flexible conditioning on target class labels or referring expressions. We evaluate the DirectSAM-RS in both zero-shot and fine-tuning setting, and demonstrate that it achieves state-of-the-art performance across several downstream benchmarks. Shiyu Miao, Delong Chen, Fan Liu 0003, Chuanyi Zhang, Yanhui Gu, Shengjie Guo, Jun Zhou 0011 |
ICASSP | 3 |
| 2025 | RemoteTrimmer: Adaptive Structural Pruning for Remote Sensing Image ClassificationabstractSince high resolution remote sensing image classifi-cation often requires a relatively high computation complexity, lightweight models tend to be practical and efficient. Model pruning is an effective method for model compression. However, existing methods rarely take into account the specificity of remote sensing images, resulting in significant accuracy loss after pruning. To this end, we propose an effective structural pruning approach for remote sensing image classification. Specifically, a pruning strategy that amplifies the differences in channel importance of the model is introduced. Then an adaptive mining loss function is designed for the fine-tuning process of the pruned model. Finally, we conducted experiments on two remote sensing classification datasets. The experimental results demonstrate that our method achieves minimal accuracy loss after compressing remote sensing classification models, achieving state-of-the-art (SoTA) performance. Guangwenjie Zou, Liang Yao 0001, Fan Liu 0003, Chuanyi Zhang, Xin Li 0090, Shengxiang Xu, Jun Zhou 0001 |
ICASSP | 3 |
| 2025 | RemoteSAM: Towards Segment Anything for Earth ObservationabstractWe aim to develop a robust yet flexible visual foundation model for Earth observation. It should possess strong capabilities in recognizing and localizing diverse visual targets while providing compatibility with various input-output interfaces required across different task scenarios. Current systems cannot meet these requirements, as they typically utilize task-specific architecture trained on narrow data domains with limited semantic coverage. Our study addresses these limitations from two aspects: data and modeling. We first introduce an automatic data engine that enjoys significantly better scalability compared to previous human annotation or rule-based approaches. It has enabled us to create the largest dataset of its kind to date, comprising 270K image-text-mask triplets covering an unprecedented range of diverse semantic categories and attribute specifications. Based on this data foundation, we further propose a task unification paradigm that centers around referring expression segmentation. It effectively handles a wide range of vision-centric perception tasks, including classification, detection, segmentation, grounding, etc, using a single model without any task-specific heads. Combining these innovations on data and modeling, we present RemoteSAM, a foundation model that establishes new SoTA on several earth observation perception benchmarks, outperforming other foundation models such as Falcon, GeoChat, and LHRS-Bot with significantly higher efficiency. Models and data are publicly available at https://github.com/1e12Leon/RemoteSAM. Liang Yao 0001, Fan Liu 0003, Delong Chen, Chuanyi Zhang, Ziyun Chen 0004, Shimin Di, Yuhui Zheng |
ACM Multimedia | 2 |
| 2025 | UEMM-Air: Enable UAVs to Undertake More Multi-modal TasksabstractThe development of multi-modal Unmanned Aerial Vehicles (UAVs) environment perception systems is hindered by three critical gaps in existing datasets: (1) insufficient modalities and pixel misalignment, (2) noisy labels, and (3) limited task types. To address these gaps, we propose an automatic data construction approach and construct a multi-modal UAV-based environment perception dataset, UEMM-Air. Its synthetic nature ensures scalability, reproducibility, and rare-event coverage, making it suitable for large-scale model pre-training. Benefiting from our automated data collection and annotation pipeline, UEMM-Air encompasses 120k data pairs across 6 aligned modalities and supports 4 perception tasks, significantly exceeding existing datasets (max 60k data, 3 modalities, 2 tasks). Compared to existing synthetic datasets like SynDrone, UEMM-Air provides more accurate annotations by avoiding noisy labels from direct coordinate computation. Notably, models pre-trained on UEMM-Air achieve a 5.8% accuracy improvement compared to those utilizing other synthetic datasets, while requiring less than half the data. This benchmark establishes performance evaluation of UAV multi-modal environmental perception models, and hopefully encourages more research efforts towards enabling UAVs to undertake more multi-modal tasks. The dataset and its generation engine are openly accessible under a permissive license at https://github.com/1e12Leon/UEMM-Air. Liang Yao 0001, Fan Liu 0003, Shengxiang Xu, Chuanyi Zhang, Shimin Di, Jianyu Jiang, Zequan Wang, Jun Zhou 0001 |
ACM Multimedia | 2 |
| 2025 | Unifying Foundation Model and Segment Anything Model for Remote Sensing Weakly Supervised Semantic SegmentationabstractDue to its reliance on fewer precise annotations, weakly supervised semantic segmentation (WSSS) techniques are in high demand in the field of remote sensing (RS) image processing. Despite mainstream WSSS approaches achieve remarkable dense prediction accuracies, they still face challenges such as insufficient pre-trained and ambiguous segment predictions. To this end, we propose to improve the accuracy of weakly supervised semantic segmentation by unifying the vision language foundation model and the Segment Anything Model (SAM). Specifically, we leverage a remote sensing vision language foundational model, RemoteCLIP, to provide sufficient pre-trained knowledge. Subsequently, we employ a decoder to transform the high-level feature representations extracted by RemoteCLIP into the segmentation predictions. Then, we introduce a multi-prompt fusion (MPF) approach via the Segment Anything Model (SAM) to obtain high-quality segment results with well-defined boundaries. To the best of our knowledge, this is the first study to apply a unified framework of foundation model and Segment Anything Model for RS WSSS. Experimental results demonstrate that our method achieves remarkable performance across three remote sensing datasets. Jinfeng Cui, Liang Yao 0001, Guoyan Xu, Fan Liu 0003 |
SMC | 7 |
| 2025 | Collaborative Semantic Contrastive for All-in-one Image Restoration
Bin Hu 0023, Sai Yang, Fan Liu 0003, Weiping Ding 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Graph guided local structure propagation for tensorial multi-view subspace clusteringabstractMulti-view subspace clustering (MSC) is widely studied owing to the ability to capture the diverse and complementary information hidden in multiple views. As a representative model, tensorial MSC can capture global information by leveraging high-order correlations across various perspectives, leading to promising results. However, this approach fails to reveal the local structure in the specific view and ignores the prior information of the self-representation tensor. To address the problems, we propose the novel graph-guided local structure propagation (GGLSP) for tensorial MSC. First, we improve the adaptive graph model to acquire a fused graph similarity matrix for extracting the relationships between samples and propagating the local structure information to the self-representation tensor. Subsequently, we introduce the weighted tensor Schatten p-norm to approximate the tensor rank function by exploiting the contributions of different singular values so that the self-representation tensor can better reveal the global low-rank structure information. Finally, we develop two efficient algorithms to solve the optimization problems. Large numbers of experiments on seven popular datasets confirm the superiority of our proposed GGLSP. Tao Zhang 0015, Yizhang Wang, Xiaobo Shen 0001, Fan Liu 0003 |
Intell. Data Anal. | 5 |
| 2025 | Integrating Global and Local Information for Remote Sensing Image-Text RetrievalabstractPre-trained Vision-Language Models (VLMs) have demonstrated promising performance in remote sensing image-text retrieval tasks. However, the scarcity of high-quality image-text datasets remains a challenge in fine-tuning VLMs for remote sensing. The captions in existing datasets tend to be uniform and lack details. To fully utilize rich detailed information from remote sensing images, we propose a method to fine-tune VLMs. We first construct a new visual-language dataset that balances both Global and Local information for Remote Sensing image-text retrieval (GLRS). Specifically, a Multi-modal Large Language Model (MLLM) is utilized to generate captions for local patches and global captions for the entire image. To effectively utilize local information, we propose a Global and Local image Captioning method (GLCap). With a Large Language Model (LLM), we further obtain higher-quality captions by merging both global and local captions. Finally, we fine-tune the weights of RS-M-CLIP with a progressive global-local fine-tuning strategy on GLRS. Experimental results demonstrate that our method outperforms state-of-the-art approaches on two common remote sensing image-text retrieval downstream tasks. The dataset will be publicly available once the paper is accepted. Ziyun Chen 0004, Fan Liu 0003, Zhangqingyun Guan, Xiaocong Zhou, Chuanyi Zhang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2025 | Domain-Invariant Progressive Knowledge Distillation for UAV-Based Object DetectionabstractKnowledge distillation (KD) is an effective method for compressing models in object detection tasks. Due to limited computational capability, unmanned aerial vehicle-based object detection (UAV-OD) widely adopt the KD technique to obtain lightweight detectors. Existing methods often overlook the significant differences in feature space caused by the large gap in scale between the teacher and student models. This limitation hampers the efficiency of knowledge transfer during the distillation process. Furthermore, the complex backgrounds in aerial images make it challenging for the student model to efficiently learn the object features. In this letter, we propose a novel KD framework for UAV-OD. Specifically, a progressive distillation approach is designed to alleviate the feature gap between teacher and student models. Then, a new feature alignment method is provided to extract object-related features for enhancing the student model’s knowledge reception efficiency. Finally, extensive experiments are conducted to validate the effectiveness of our proposed approach. The results demonstrate that our proposed method achieves state-of-the-art performance on two datasets. Liang Yao 0001, Fan Liu 0003, Chuanyi Zhang, Zhiquan Ou |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2025 | Multi-stage Bayesian Prototype Refinement with feature weighting for few-shot classification
Xiaocong Zhou, Shengxiang Xu, Fan Liu 0003, Chuanyi Zhang, Wenwen Cai, Jun Zhou 0011 |
Pattern Anal. Appl. | 4 |
| 2025 | IPT-ILR: Image Pyramid Transformer Coupled With Information Loss Regularization for All-in-One Image RestorationabstractAll-in-one image restoration has recently developed to be a new research trend in the low-level computer vision field, aiming to tackle multiple image degradation types simultaneously in a unified model. As a typical multi-task learning, existing approaches focus on modeling either the specificity or commonality among different image restoration tasks. To exploit the unique strengths of both worlds, we propose a method of Image Pyramid Transformer coupled with Information Loss Regularization (IPT-ILR), in which the multi-scale architecture structure can excavate more information for multiple restoration tasks concurrently, while the learning strategy can identify the difference among multiple restoration tasks depending on the degree of information loss in each restoration task. Specifically, it first establishes a new Image Pyramid Transformer Network (IPT-Network) to accommodate multiple image restoration tasks. Given original degraded images, the IPT-Network exploits the image pyramid technique to establish a series of images with different scales, which are then restored by transformer-like auto-encoders. Moreover, the restored image on a low-level scale is referenced to assist restoring the degraded image on a high-level scale. Next, Information Loss Regularization (ILR) is presented to optimize the IPT-Network. ILR calculates the average distance between degraded images and their clean counterparts as the weights, which automatically implement different penalties for different image restoration tasks, thus avoiding the short-cut phenomenon for the easy task while encouraging the hard task. Extensive experiments have been conducted with 6 image restoration tasks in the all-in-one setting. The results show our method performs favorably against numerous state-of-the-art methods across most tasks, including image denoising, image deblurring, image dehazing, image deraining, image desnowing, as well as low-light enhancement. Sai Yang, Bin Hu 0023, Fan Liu 0003, Xiaoxin Wu 0004, Weiping Ding 0001, Jun Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | A Euclidean Affinity-Augmented Hyperbolic Neural Network for Semantic Segmentation of Remote Sensing ImagesabstractSemantic segmentation of remote sensing images (RSIs) plays a pivotal role in advancing geospatial analyses and applications across diverse fields, such as urban planning and environmental monitoring. Traditional learning paradigms predominantly utilize Euclidean spaces for feature extraction. This approach can introduce spatial distortions when representing objects, as Euclidean architectures typically focus on locality and are optimized for grid data, not always yielding optimal geometrical representations for data structured in non-Euclidean spaces. To address these problems, we propose EAAHNet, the first fully hyperbolic neural network designed for semantic segmentation of RSIs. EAAHNet employs the Lorentz model to reformalize conventional Euclidean-based neural network operations, ensuring the preservation of hyperbolic properties. Furthermore, to account for the inherently Euclidean nature of ground objects, we propose a Euclidean affinity-augmented hyperbolic attention module (EAAHAM) that enriches contextual dependencies through an attention fusion manner. This enhancement significantly improves the network’s capacity to discern pixel-wise semantics. Extensive experiments conducted on the ISPRS Vaihingen, ISPRS Potsdam, and LoveDA datasets demonstrate EAAHNet’s superior performance over several state-of-the-art methods. Additionally, the ablation study verifies the impacts of EAAHAM. Xin Li 0090, Feng Xu 0008, Fan Liu 0003, Xin Lyu 0001, Hongmin Gao 0001, Jun Zhou 0011, André Kaup |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Boost UAV-Based Object Detection via Scale-Invariant Feature Disentanglement and Adversarial LearningabstractDetecting objects from Unmanned Aerial Vehicles (UAV) is often hindered by a large number of small objects, resulting in low detection accuracy. To address this issue, mainstream approaches typically utilize multi-stage inferences. Despite their remarkable detecting accuracies, real-time efficiency is sacrificed, making them less practical to handle real applications. To this end, we propose to improve the single-stage inference accuracy through learning scale-invariant features. Specifically, a Scale-Invariant Feature Disentangling module is designed to disentangle scale-related and scale-invariant features. Then an Adversarial Feature Learning scheme is employed to enhance disentanglement. Finally, scale-invariant features are leveraged for robust UAV-based object detection. Furthermore, we construct a multi-modal UAV object detection dataset, State-Air, which incorporates annotated UAV state parameters. We apply our approach to three lightweight detection frameworks on two benchmark datasets. Extensive experiments demonstrate that our approach can effectively improve model accuracy and achieve state-of-the-art (SoTA) performance on three datasets. Our code and dataset are publicly available at https://github.com/1e12Leon/SIFDAL. Fan Liu 0003, Liang Yao 0001, Chuanyi Zhang, Xiruo Jiang, Jun Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Robust Reversible Watermarking With Invisible Distortion Against VAE Watermark RemovalabstractOrthogonal Moment-based Robust Reversible Watermarking (OM-RRW) is crucial for intellectual property protection, providing the dual benefits of robustness and reversibility. However, OM-RRW embeds watermarks into visually sensitive global low-frequency features, which easily leads to ring-like distortions that expose watermark locations, making them vulnerable to removal through image inpainting. To address this issue, this paper makes the first attempt to introduce an innovative strategy to eliminate these visible distortions, thereby overcoming OM-RRW's inherent limitations. The strategy innovates on two fronts: first, it customizes varying embedding step sizes based on the stability differences of moment values to minimize distortion; second, it designs a texture-aware adaptive basis function fine-tuning strategy. This strategy adjusts the representation capability of the basis functions in different regions based on the human eye's sensitivity to various texture areas, helping to avoid visible ring-like distortions. The performance of the proposed method is evaluated using Polar Harmonic Transform (PHT) moments, comprising three moments that exhibit remarkable performance in existing OM-RRW methods. Extensive experiments show that the proposed method can embed 128-bit watermarks with no visible distortions while minimizing the loss of robustness. In addition, this paper finds that OM-RRW demonstrates satisfactory robustness against VAE watermark removal attacks. Bobiao Guo, Ping Ping, Fan Liu 0003, Feng Xu 0008 |
IEEE Trans. Image Process. | 3 |
| 2025 | ProtoCLIP: Prototypical Contrastive Language Image PretrainingabstractContrastive language image pretraining (CLIP) has received widespread attention since its learned representations can be transferred well to various downstream tasks. During the training process of the CLIP model, the InfoNCE objective aligns positive image-text pairs and separates negative ones. We show an underlying representation grouping effect during this process: the InfoNCE objective indirectly groups semantically similar representations together via randomly emerged within-modal anchors. Based on this understanding, in this article, prototypical contrastive language image pretraining (ProtoCLIP) is introduced to enhance such grouping by boosting its efficiency and increasing its robustness against the modality gap. Specifically, ProtoCLIP sets up prototype-level discrimination between image and text spaces, which efficiently transfers higher level structural knowledge. Furthermore, prototypical back translation (PBT) is proposed to decouple representation grouping from representation alignment, resulting in effective learning of meaningful representations under a large modality gap. The PBT also enables us to introduce additional external teachers with richer prior language knowledge. ProtoCLIP is trained with an online episodic training strategy, which means it can be scaled up to unlimited amounts of data. We trained our ProtoCLIP on conceptual captions (CCs) and achieved an +5.81% ImageNet linear probing improvement and an +2.01% ImageNet zero-shot classification improvement. On the larger YFCC-15M dataset, ProtoCLIP matches the performance of CLIP with 33% of training time. Delong Chen, Fan Liu 0003, Zaiquan Yang, Shaoqiu Zheng, Ying Tan 0002, Erjin Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Global-Local Multiple Granularity Learning for Cross-Modality Visible-Infrared Person ReidentificationabstractCross-modality visible-infrared person reidentification (VI-ReID), which aims to retrieve pedestrian images captured by both visible and infrared cameras, is a challenging but essential task for smart surveillance systems. The huge barrier between visible and infrared images has led to the large cross-modality discrepancy and intraclass variations. Most existing VI-ReID methods tend to learn discriminative modality-sharable features based on either global or part-based representations, lacking effective optimization objectives. In this article, we propose a novel global-local multichannel (GLMC) network for VI-ReID, which can learn multigranularity representations based on both global and local features. The coarse- and fine-grained information can complement each other to form a more discriminative feature descriptor. Besides, we also propose a novel center loss function that aims to simultaneously improve the intraclass cross-modality similarity and enlarge the interclass discrepancy to explicitly handle the cross-modality discrepancy issue and avoid the model fluctuating problem. Experimental results on two public datasets have demonstrated the superiority of the proposed method compared with state-of-the-art approaches in terms of effectiveness. Liyan Zhang 0001, Guodong Du 0005, Fan Liu 0003, Huawei Tu, Xiangbo Shu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | AerialFace: A Light Weight Framework for Unmanned Aerial Vehicle Face RecognitionabstractUnmanned Aerial Vehicle (UAV) are widely applied in multiple fields due to their simple structure and high flexibility. Applying facial recognition technology to UAV can improve their intelligence and diversity of application scenarios. However, UAV face recognition is often hindered by low resolution face, resulting in low accuracy. To alleviate these issues, we propose an efficient face recognition framework named AerialFace. Firstly, we utilize the Residual SRGAN (ResSR-GAN) model to enhance image quality and generate high-resolution face images. Then, we propose Semantic-improved MobileFaceNet (SeMFNet) to relieve the impact of complex backgrounds. Finally, we leverage two pruning algorithms for face detection and recognition models, respectively. It can reduce their parameters to meet the deployment requirements of the algorithm on UAVs. Furthermore, we apply our AerialFace on a UAV face dataset and employ it on an edge computing device. Extensive experiments demonstrate that our approach can effectively improve UAV face recognition accuracy and have real-time performance in embedded UAV devices. Zhiquan Ou, Liang Yao 0001, Fan Liu 0003 |
FG | 4 |
| 2024 | Locality-Enhanced Transformer for Semantic Segmentation of High-Resolution Remote Sensing ImagesabstractTransformers have emerged as a transformative tool in various computer vision tasks, excelling at capturing long-range dependencies. Their potential applicability and scalability in the interpretation of high-resolution remote sensing images (HRRSIs) have thus garnered substantial interest. However, unlike natural images, HRRSIs present intricate scenes characterized by scale variations and diverse appearances. These challenges underscore the importance of enabling networks to effectively assimilate both local intricacies and global context. In this letter, we introduce LETFormer, a semantic segmentation transformer. LETFormer balances capturing longrange dependencies with preserving local details through its unique LETFormer block, featuring an anchor token. This token aggregates localized contextual information within a designated window and promotes meaningful interactions among anchor tokens. With a mask transformer decoder, LETFormer gains ample contextual cues for precise semantic mask prediction. Empirical findings based on evaluations using the ISPRS Potsdam and LoveDA benchmarks unequivocally establish LETFormer’s superiority over state-of-the-art models. Additionally, we analyze the parameter size and floating-point operations per second (FLOPs) of LETFormer. Xin Li 0090, Feng Xu 0008, Runliang Xia, Nan Xu 0008, Fan Liu 0003, Chi Yuan, Qian Huang 0008, Xin Lyu 0001 |
ICASSP | 5 |
| 2024 | FreqFormer: A Frequency Transformer for Semantic Segmentation of Remote Sensing ImagesabstractSemantic segmentation of remote sensing images (RSIs) is vital for geospatial intelligence. However, traditional methods face challenges with mixed pixels and complex land cover types. Convolutional neural networks and transformers have led the field of RSI semantic segmentation by learning visual features in the spatial domain, but they often overlook the rich spectral features which can be well-described in the frequency domain, resulting in inadequate context modeling. In this paper, we present FreqFormer, a frequency transformer that enhances semantic segmentation by incorporating both spectral and spatial information through a devised frequency attention (FA) module. FA refines representations in the frequency domain through two parallel branches. Specifically, the high-frequency branch (HFB) utilizes a convolution layer with a Canny kernel to preserve high-frequency details, followed by multi-head self-attention to model high-frequency context. Followed by an element summation, high-frequency and low-frequency contexts are aggregated. Then, the formed FreqFormer block is sequentially deployed in the encoder stage with patch merging for spatial contraction. As for the decoder, the mask transformer decoder applies a scalar product to predict patch-wise semantics before upsampling. In experiments, FreqFormer outperforms state-of-the-art models on the ISPRS Potsdam and LoveDA datasets, demonstrating significant improvements in numerical evaluations. The integration of HFB significantly boosts the model’s ability to capture fine details, highlighting its potential for geospatial analysis. Xin Li 0090, Feng Xu 0008, Fan Liu 0003, Yiwei Fang, Xin Lyu 0001, Jun Zhou 0001 |
MMAsia | 4 |
| 2024 | Feature-weighted Multi-stage Bayesian Prototype for Few-shot ClassificationabstractFew-shot classification aims to recognize the query sample through a limited amount of support data, where a prototype classifier is commonly applied. However, although the prototype classifier is simple and non-parametric, it does not fully utilize the prior information of samples, leading to prototype bias. To this end, we propose a Feature-weighted Multi-stage Bayesian Prototype Classifier (FMBPC). Specifically, we utilize a feature weighting module to balance the effect of each support sample. Then, features of balanced support samples are utilized as prior information to construct the Bayesian prototype classifier, which can focus more on the important information. Ultimately, a multi-stage inferring strategy is adopted, where the support sample with the greatest distance is filtered in each stage. Prototypes and the corresponding classification score are updated after sample filtering. By integrating the multi-stage classification results, we successfully utilize multi-stage Bayesian inference to enhance the prototype classifier for more accurate few-shot classification results. Experimental results show the efficacy of our method, demonstrating notable advancements in few-shot classification accuracy. Xiaocong Zhou, Fan Liu 0003, Chuanyi Zhang, Wenwen Cai, Jun Zhou 0001 |
MMAsia | 2 |
| 2024 | Single Image Unlearning: Efficient Machine Unlearning in Multimodal Large Language ModelsabstractMachine unlearning (MU) empowers individuals with the `right to be forgotten' by removing their private or sensitive information encoded in machine learning models. However, it remains uncertain whether MU can be effectively applied to Multimodal Large Language Models (MLLMs), particularly in scenarios of forgetting the leaked visual data of concepts. To overcome the challenge, we propose an efficient method, Single Image Unlearning (SIU), to unlearn the visual recognition of a concept by fine-tuning a single associated image for few steps. SIU consists of two key aspects: (i) Constructing Multifaceted fine-tuning data. We introduce four targets, based on which we construct fine-tuning data for the concepts to be forgotten; (ii) Joint training loss. To synchronously forget the visual recognition of concepts and preserve the utility of MLLMs, we fine-tune MLLMs through a novel Dual Masked KL-divergence Loss combined with Cross Entropy loss. Alongside our method, we establish MMUBench, a new benchmark for MU in MLLMs and introduce a collection of metrics for its evaluation. Experimental results on MMUBench show that SIU completely surpasses the performance of existing methods. Furthermore, we surprisingly find that SIU can avoid invasive membership inference attacks and jailbreak attacks. To the best of our knowledge, we are the first to explore MU in MLLMs. We will release the code and benchmark in the near future. Jiaqi Li 0031, Qianshan Wei, Chuanyi Zhang, Guilin Qi, Miaozeng Du, Yongrui Chen 0002, Fan Liu 0003 |
NeurIPS | 8 |
| 2024 | Two-step affinity matrix learning for multi-view subspace clustering
Tao Zhang 0015, Yun-Hao Yuan 0001, Xiaobo Shen 0001, Fan Liu 0003 |
Expert Syst. Appl. | 4 |
| 2024 | AAFormer: Attention-Attended Transformer for Semantic Segmentation of Remote Sensing ImagesabstractThe rapid advancements in remote sensing technology have enabled the widespread availability of fine-resolution remote sensing images (RSIs), offering rich spatial details and semantics. Despite the applicability and scalability of transformers in semantic segmentation of RSIs by learning pairwise contextual affinity, they inevitably introduce irrelevant context, hindering accurate inference of patch semantics. To address this, we propose a novel multi-head attention-attended module (AAM) that refines the multi-head self-attention mechanism. The AAM filters out irrelevant context while highlighting informative ones by considering the relevance between self-attention maps and the query vector. The AAM generates an attention gate to complement contextual affinity and emphasize the useful ones with a higher weight simultaneously. Leveraging multi-head AAM as the core unit, we construct a lightweight attention-attended transformer block (ATB). Subsequently, we devise AAFormer, a pure transformer with a mask transformer decoder, for achieving semantic segmentation of RSIs. We extensively evaluate our approach on the ISPRS Potsdam and LoveDA datasets, demonstrating compelling performance compared to mainstream methods. Additionally, we conduct evaluations to analyze the effects of AAM. Xin Li 0090, Feng Xu 0008, Linyang Li, Nan Xu 0008, Fan Liu 0003, Chi Yuan, Xin Lyu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | A Cross-Domain Coupling Network for Semantic Segmentation of Remote Sensing ImagesabstractSemantic segmentation of remote sensing images (RSIs) is critical for various applications, including urban planning, agriculture, and disaster management. Existing methods often fail to capture fine-grained textures and periodic patterns in RSIs, leading to suboptimal results in complex terrains. To address these challenges, we propose a cross-domain coupling network (CDCNet) that leverages both domain-specific extraction and cross-domain coupling (CDC) to enrich contextual cues for semantic inference. Our CDCNet integrates a CDC layer within the encoder-decoder architecture to simultaneously refine representations in the frequency and spatial domains. This approach effectively models fine-grained textures and periodic patterns in the frequency domain, as well as edges, shapes, and broad structural elements in the spatial domain. Extensive experiments on the ISPRS Potsdam and LoveDA datasets demonstrate the superiority of CDCNet over several state-of-the-art methods. Ablation studies confirm the significant impact of the CDC layer, validating the effectiveness of our approach in handling RSIs. Xin Li 0090, Feng Xu 0008, Feifei Tao, Hongmin Gao 0001, Fan Liu 0003, Xin Lyu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2024 | Few-shot classification guided by generalization error bound
Fan Liu 0003, Sai Yang, Delong Chen, Huaxi Huang, Jun Zhou 0001 |
Pattern Recognit. | 1 |
| 2024 | A Frequency Domain Feature-Guided Network for Semantic Segmentation of Remote Sensing ImagesabstractSemantic segmentation of Remote Sensing Images (RSIs) entails assigning semantic labels to each pixel accurately. RSIs are rich in spatial and spectral data, revealing diverse material and object characteristics. Yet, current RSI-focused computer vision models struggle with significant intra-class variation and inter-class resemblance due to limited spectral data usage. We propose the Frequency Domain Feature-Guided Network (FFGNet) for RSI semantic segmentation, influenced by digital signal processing theories. FFGNet initially generates frequency domain features via patch partitioning and 2D discrete cosine transformation. Our Frequency Enhancement Attention module (FEA) then distinguishes and intensifies frequency components to retain detailed information. These enhanced features are integrated with the Spatial-Spectral Attention (SSA) for enriched spectral signals. In the inference phase, these features are upsampled and combined with decoded features, emphasizing spectral details. Additionally, our novel loss function combines frequency and cross-entropy losses. Experiments on LoveDA and ISPRS Potsdam datasets demonstrate FFGNet's effectiveness, surpassing other mainstream models. An ablation study further validates our dual-guidance design. Xin Li 0090, Feng Xu 0008, Hongmin Gao 0001, Fan Liu 0003, Xin Lyu 0001 |
IEEE Signal Process. Lett. | 4 |
| 2024 | Few-Shot Classification Model Compression via School LearningabstractFew-shot classification (FSC) is a challenging task due to limitation in accessing training data. Recent methods often employ highly complex networks to obtain high-quality features, but this may not be suitable for resource-limited applications. To tackle this challenge, we introduce Few-Shot Classification Model Compression (FSC-MC), a new task aimed at enhancing the FSC performance of lightweight and low-capacity models by learning from more complex models. We also propose a novel two-level learning strategy called School Learning to accomplish the FSC-MC task by mimicking the real learning process in the social school life. In this new learning paradigm, the first level performs preview learning, in which each student is equipped with a preparer to perform self-learning on the base set. The second level is the team learning, consisting of a complex teacher network and several lightweight student networks organized into a team. One student network is randomly chosen as the leader network, while the remaining student networks serve as member networks. The leader network simultaneously learns knowledge from the teacher network and all member networks. Conversely, each member network receives knowledge from both the teacher network and the leader network. Ultimately, the leader network is deployed for FSC evaluation, resulting in effective model compression. Extensive experiments in the FSC-MC setting demonstrate that School Learning outperforms 17 state-of-the-art knowledge distillation methods including both offline methods and online methods, enabling lightweight models to achieve outstanding FSC performance. Sai Yang, Fan Liu 0003, Delong Chen, Huaxi Huang, Jun Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Semantic Segmentation of Remote Sensing Images by Interactive Representation Refinement and Geometric Prior-Guided InferenceabstractHigh spatial resolution remote sensing images (HRRSIs) contain intricate details and varied spectral distributions, making their semantic segmentation a challenging task. To address this problem, it is crucial to adequately capture both local and global contexts to reduce semantic ambiguity. While self-attention modules in vision transformers capture long-range context, they tend to sacrifice local details. In this article, we propose a geometric prior-guided interactive network (GPINet), a hybrid network that refines features across encoder and decoder stages. First of all, a dual branch structure encoder with local-global interaction modules (LGIMs) is designed to fully exploit local and global contexts for feature refinement. Unlike commonly used skip connections or concatenations, the LGIMs bilaterally couple and exchange CNN features with transformer features by lossless transformation and elaborating cross-attention. Moreover, we introduce a geometric prior generation module (GPGM) that iteratively updates the randomly initialized geometric prior. Subsequently, the geometric priors are stored and used to guide feature recovery. Finally, a weighted summation is applied to the upsampled decoded features and geometric priors. By comprehensively capturing contexts and enabling lossless decoding and deterministic inference, GPINet allows the network to learn discriminative representations for accurately specifying pixel-level semantics. Experiments on three benchmark datasets demonstrate the superiority of the proposed GPINet over state-of-the-art methods. Furthermore, we validate the effectiveness of geometric priors and compare the model sizes. Xin Li 0090, Feng Xu 0008, Fan Liu 0003, Xin Lyu 0001, Jun Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | RemoteCLIP: A Vision Language Foundation Model for Remote SensingabstractGeneral-purpose foundation models have led to recent breakthroughs in artificial intelligence. In remote sensing, self-supervised learning (SSL) and Masked Image Modeling (MIM) have been adopted to build foundation models. However, these models primarily learn low-level features and require annotated data for fine-tuning. Moreover, they are inapplicable for retrieval and zero-shot applications due to the lack of language understanding. To address these limitations, we propose RemoteCLIP, the first vision-language foundation model for remote sensing that aims to learn robust visual features with rich semantics and aligned text embeddings for seamless downstream application. To address the scarcity of pre-training data, we leverage data scaling which converts heterogeneous annotations into a unified image-caption data format based on Box-to-Caption (B2C) and Mask-to-Box (M2B) conversion. By further incorporating UAV imagery, we produce a 12 × larger pretraining dataset than the combination of all available datasets. RemoteCLIP can be applied to a variety of downstream tasks, including zero-shot image classification, linear probing,k-NN classification, few-shot classification, image-text retrieval, and object counting in remote sensing images. Evaluation on 16 datasets, including a newly introduced RemoteCount benchmark to test the object counting ability, shows that RemoteCLIP consistently outperforms baseline foundation models across different model scales. Impressively, RemoteCLIP beats the state-of-the-art method by 9.14% mean recall on the RSITMD dataset and 8.92% on the RSICD dataset. For zero-shot classification, our RemoteCLIP outperforms the CLIP baseline by up to 6.39% average accuracy on 12 downstream datasets. Fan Liu 0003, Delong Chen, Zhangqingyun Guan, Xiaocong Zhou, Qiaolin Ye, Liyong Fu, Jun Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Few-shot Classification via Ensemble Learning with Multi-Order StatisticsabstractTransfer learning has been widely adopted for few-shot classification. Recent studies reveal that obtaining good generalization representation of images on novel classes is the key to improving the few-shot classification accuracy. To address this need, we prove theoretically that leveraging ensemble learning on the base classes can correspondingly reduce the true error in the novel classes. Following this principle, a novel method named Ensemble Learning with Multi-Order Statistics (ELMOS) is proposed in this paper. In this method, after the backbone network, we use multiple branches to create the individual learners in the ensemble learning, with the goal to reduce the storage cost. We then introduce different order statistics pooling in each branch to increase the diversity of the individual learners. The learners are optimized with supervised losses during the pre-training phase. After pre-training, features from different branches are concatenated for classifier evaluation. Extensive experiments demonstrate that each branch can complement the others and our method can produce a state-of-the-art performance on multiple few-shot classification benchmark datasets. Sai Yang, Fan Liu 0003, Delong Chen, Jun Zhou 0001 |
IJCAI | 2 |
| 2023 | Few-shot classification using Gaussianisation prototypical classifierabstractAbstract Few‐shot classification (FSC) aims at classifying query samples into correct classes given only a few labelled samples. Prototypical Classifier (PC) can be chosen to be an ideal classifier for settling this problem, as it has good properties of low‐capacity and parameter‐free. However, the mean‐based prototypes suffer from the issue of deviating from its ground‐truth centre. In order to solve such problem of prototype bias, Gaussianisation Prototypical Classifier (GPC) is proposed, which is a kind of one‐step prototype rectification method. Specifically, the authors first perform Gaussianisation operation over the feature extracted from the backbone network so that the features fit the particular Gaussian distribution. Second, the authors use prototype feature of the base class as prior information and employs Maximum a Posteriori estimation method to obtain the reliable prototype for each novel class. Finally, the query sample of novel class is classified to be its nearest prototype with non‐parametric classifiers. Extensive experiments have been conducted on multiple FSC benchmarks. Comparative results also demonstrate that the authors’ method is superior to existing state‐of‐the‐art FSC methods. Fan Liu 0003, Sai Yang |
IET Comput. Vis. | 1 |
| 2023 | Asymmetric exponential loss function for crack segmentation
Fan Liu 0003, Delong Chen, Chunmei Shen, Feng Xu 0008 |
Multim. Syst. | 1 |
| 2023 | MEP-3M: A large-scale multi-modal E-commerce product dataset
Fan Liu 0003, Delong Chen, Xiaoyu Du 0002, Ruizhuo Gao, Feng Xu 0008 |
Pattern Recognit. | 1 |
| 2023 | A Synergistical Attention Model for Semantic Segmentation of Remote Sensing ImagesabstractIn remotely sensed images, high intraclass variance and interclass similarity are ubiquitous due to complex scenes and objects with multivariate features, making semantic segmentation a challenging task. Deep convolutional neural networks can solve this problem by modeling the context of features and improving their discriminability. However, current learning paradigms model the feature affinity in spatial dimension and channel dimension separately and then fuse them in a sequential or parallel manner, leading to suboptimal performance. In this study, we first analyze this problem practically and summarize it as attention bias that reduces the capability of network in distinguishing weak and discretely distributed objects from wide-range objects with internal connectivity, when modeled only in spatial or channel domain. To jointly model both spatial and channel affinity, we design a synergistic attention module (SAM), which allows for channelwise affinity extraction while preserving spatial details. In addition, we propose a synergistic attention perception neural network (SAPNet) for the semantic segmentation of remote sensing images. The hierarchical-embedded synergistic attention perception module aggregates SAM-refined features and decoded features. As a result, SAPNet enriches inference clues with desired spatial and channel details. Experiments on three benchmark datasets show that SAPNet is competitive in accuracy and adaptability compared with state-of-the-art methods. The experiments also validate the hypothesis of attention bias and the efficiency of SAM. Xin Li 0090, Feng Xu 0008, Fan Liu 0003, Xin Lyu 0001, Zhennan Xu, Jun Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | A review of driver fatigue detection and its advances on the use of RGB-D camera and deep learning
Fan Liu 0003, Delong Chen, Jun Zhou 0001, Feng Xu 0008 |
Eng. Appl. Artif. Intell. | 1 |
| 2022 | Feature hallucination in hypersphere space for few-shot classificationabstractAbstract Few‐shot classification (FSC) targeting at classifying unseen classes with few labelled samples is still a challenging task. Recent works show that transfer‐learning based approaches are competitive with meta‐learning ones, which usually pre‐train a convolutional neural networks (CNN)‐based network using cross‐entropy (CE) loss and throw away the last layer to post‐process the novel classes. Hereby, they still suffer the issue of getting a more transferable extractor and lacking enough labelled novel samples. Thus, the authors propose the algorithm of feature hallucination in hypersphere space (FHHS) for FSC. On the first stage, the authors pre‐train a more transferable feature extractor using a hypersphere loss (HL), which supplies CE with supervised contrastive (SC) loss and self‐supervised loss (SSL), in which SC can map the base and novel images onto the hypersphere space densely. On the second stage, the authors generate new samples for unseen classes using their novel algorithm of synthetic novel sampling with the base (SNSB), which linearly interpolate between each novel class prototype and its K nearest neighbour base class prototypes. Comprehensive experiments on multiple popular FSC demonstrate that HL loss can enhance the performance of backbone network and the authors’ feature hallucination method is superior to the existing hallucination‐based methods. Sai Yang, Fan Liu 0003 |
IET Image Process. | 2 |
| 2022 | Self-Supervised Music Motion Synchronization Learning for Music-Driven Conducting Motion Generation
Fan Liu 0003, Delong Chen, Ruizhi Zhou, Sai Yang, Feng Xu 0008 |
J. Comput. Sci. Technol. | 1 |
| 2022 | Hybridizing Euclidean and Hyperbolic Similarities for Attentively Refining Representations in Semantic Segmentation of Remote Sensing ImagesabstractAttention mechanisms have revolutionized the semantic segmentation network in interpreting remotely sensed images (RSIs) due to their amazing ability in establishing contextual dependencies. Nevertheless, due to the complex scenes and diverse objects in RSIs, a variety of details and correlations are not available in Euclidean space. Therefore, a similarity-hybrid attention module (SHAM) is devised to attentively learn the hyperbolic and Euclidean attention maps between any two positions, followed by a weighted element-wise summation. The hybrid attention maps posses latent geometric properties of both Euclidean and hyperboloid. Taking commonly-used fully convolutional network (FCN) as baseline, HAENet that embeds SHAM, is presented. Experiments on ISPRS Potsdam and DeepGlobe benchmarks reveal its superiority to comparative methods. In addition, the ablation study validates the effectiveness of SHAM compared to other attention modules. Xin Li 0090, Feng Xu 0008, Fan Liu 0003, Runliang Xia, Linyang Li, Zhennan Xu, Xin Lyu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | A Survey of Convolutional Neural Networks: Analysis, Applications, and ProspectsabstractA convolutional neural network (CNN) is one of the most significant networks in the deep learning field. Since CNN made impressive achievements in many areas, including but not limited to computer vision and natural language processing, it attracted much attention from both industry and academia in the past few years. The existing reviews mainly focus on CNN's applications in different scenarios without considering CNN from a general perspective, and some novel ideas proposed recently are not covered. In this review, we aim to provide some novel ideas and prospects in this fast-growing field. Besides, not only 2-D convolution but also 1-D and multidimensional ones are involved. First, this review introduces the history of CNN. Second, we provide an overview of various convolutions. Third, some classic and advanced CNN models are introduced; especially those key points making them reach state-of-the-art results. Fourth, through experimental analysis, we draw some conclusions and provide several rules of thumb for functions and hyperparameter selection. Fifth, the applications of 1-D, 2-D, and multidimensional convolution are covered. Finally, some open issues and promising directions for CNN are discussed as guidelines for future work. Fan Liu 0003, Shouheng Peng, Jun Zhou 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Feature hallucination via Maximum A Posteriori for few-shot learning
Ning Dong 0001, Fan Liu 0003, Sai Yang, Jinglu Hu |
Knowl. Based Syst. | 3 |
| 2020 | Pavement Crack Detection Using Attention U-Net with Multiple Sources
Fan Liu 0003, Guoyan Xu, Tao Zhang 0015 |
PRCV (2) | 2 |
| 2020 | Multi-view generalized support vector machine via mining the inherent relationship between views with applications to face and fire smoke recognition
Yawen Cheng, Liyong Fu, Qiaolin Ye, Fan Liu 0003 |
Knowl. Based Syst. | 5 |
| 2020 | Robust discriminant feature selection via joint L2, 1-norm distance minimization and maximization
Zhangjing Yang, Qiaolin Ye, Qiao Chen 0004, Xu Ma 0005, Liyong Fu, Guowei Yang 0002, Fan Liu 0003 |
Knowl. Based Syst. | 8 |
| 2019 | Single sample face recognition via BoF using multistage KNN collaborative codingabstractIn this paper, we propose a multistage KNN collaborative coding based Bag-of-Feature (MKCC-BoF) method to address SSPP problem, which tries to weaken the semantic gap between facial features and facial identification. First, local descriptors are extracted from the single training face images and a visual dictionary is obtained offline by clustering a large set of descriptors with K-means. Then, we design a multistage KNN collaborative coding scheme to project local features into the semantic space, which is much more efficient than the most commonly used non-negative sparse coding algorithm in face recognition. To describe the spatial information as well as reduce the feature dimension, the encoded features are then pooled on spatial pyramid cells by max-pooling, which generates a histogram of visual words to represent a face image. Finally, a SVM classifier based on linear kernel is trained with the concatenated features from pooling results. Experimental results on three public face databases show that the proposed MKCC-BoF is much superior to those specially designed methods for SSPP problem. Moreover, it also has great robustness to expression, illumination, occlusion and, time variation. Fan Liu 0003, Sai Yang, Yuhua Ding, Feng Xu 0008 |
Multim. Tools Appl. | 1 |
| 2018 | A Text-Based CAPTCHA Cracking System with Generative Adversarial NetworksabstractAs a multimedia security mechanism, CAPTCHAs are completely automated public turing test to tell computers and humans apart. Although cracking CAPTCHA has been explored for many years, it is still a challenging problem for real practice. In this demo, we present a text based CAPTCHA cracking system by using convolutional neural networks(CNN). To solve small sample problem, we propose to combine conditional deep convolutional generative adversarial networks(cDCGAN) and CNN, which makes a tremendous progress in accuracy. In addition, we also select multiple models with low pearson correlation coefficients for majority voting ensemble, which further improves the accuracy. The experimental results show that the system has great advantages and provides a new mean for cracking CAPTCHAs. Fan Liu 0003, Xueyi Li 0008, Tanyue Lv |
ISM | 1 |
| 2018 | L1-Norm Distance Linear Discriminant Analysis Based on an Effective Iterative AlgorithmabstractRecent works have proposed two L1-norm distance measure-based linear discriminant analysis (LDA) methods, L1-LD and LDA-L1, which aim to promote the robustness of the conventional LDA against outliers. In LDA-L1, a gradient ascending iterative algorithm is applied, which, however, suffers from the choice of stepwise. In L1-LDA, an alternating optimization strategy is proposed to overcome this problem. In this paper, however, we show that due to the use of this strategy, L1-LDA is accompanied with some serious problems that hinder the derivation of the optimal discrimination for data. Then, we propose an effective iterative framework to solve a general L1-norm minimization-maximization (minmax) problem. Based on the framework, we further develop a effective L1-norm distance-based LDA (called L1-ELDA) method. Theoretical insights into the convergence and effectiveness of our algorithm are provided and further verified by extensive experimental results on image databases. Qiaolin Ye, Jian Yang 0003, Fan Liu 0003, Chunxia Zhao, Ning Ye 0001, Tongming Yin |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | A multi-phase sparse probability framework via entropy minimization for single sample face recognitionabstractIn this paper, we propose a robust probability based sparse method to solve single sample face recognition, which harvests the advantages of both local and global representation. Different from previous sparse representation methods that generate sparse coefficients by l1, we produce sparse class probability distribution by proposing a multi-phase sparse probability (MSP) framework. To create class probability distribution, we divide each face image into many local blocks and vote based on the classification results of all blocks. For classifying each block, we propose local similarity assumption that makes many conventional methods feasible to SSPP problem. Moreover, we also propose a heuristic multiphase class selection scheme to solve the entropy minimization problem, which finally provides a higher classification confidence from the global perspective. Experimental results on three popular databases show that our approach not only generalizes well to SSPP problem but also has strong robustness to expression, illumination, occlusion and time variation. Fan Liu 0003, Jinhui Tang 0001, Yan Song 0005, Qian Huang 0008, Feng Xu 0008 |
ICIP | 1 |
| 2016 | Nonnegative matrix factorization with endmember sparse graph learning for hyperspectral unmixingabstractNonnegative matrix factorization (NMF) based hyperspectral unmixing aims at estimating pure spectral signatures and their fractional abundances at each pixel. During the past several years, manifold structures have been introduced as regularization constraints into NMF. However, most methods only consider the constraints on abundance matrix while ignoring the geometric relationship of endmembers. Although such relationship can be described by traditional graph construction approaches based on k-nearest neighbors, its accuracy is questionable. In this paper, we propose a novel hyperspectral unmixing method, namely NMF with endmember sparse graph learning, to tackle the above drawbacks. This method first integrates endmember sparse graph structure into NMF, then simultaneously performs unmixing and graph learning. It is further extended by incorporating abundance smoothness constraint to improve the unmixing performance. Experimental results on both synthetic and real datasets have validated the effectiveness of the proposed method. Bin Qian 0006, Jun Zhou 0001, Xiaobo Shen 0001, Fan Liu 0003 |
ICIP | 5 |
| 2016 | Local structure based multi-phase collaborative representation for face recognition with single sample per person
Fan Liu 0003, Jinhui Tang 0001, Yan Song 0005, Ye Bi, Sai Yang |
Inf. Sci. | 1 |
| 2015 | Local Structure-Based Sparse Representation for Face RecognitionabstractThis article presents a simple yet effective face recognition method, called local structure-based sparse representation classification (LS_SRC). Motivated by the “divide-and-conquer” strategy, we first divide the face into local blocks and classify each local block, then integrate all the classification results to make the final decision. To classify each local block, we further divide each block into several overlapped local patches and assume that these local patches lie in a linear subspace. This subspace assumption reflects the local structure relationship of the overlapped patches, making sparse representation-based classification (SRC) feasible even when encountering the single-sample-per-person (SSPP) problem. To lighten the computing burden of LS_SRC, we further propose the local structure-based collaborative representation classification (LS_CRC). Moreover, the performance of LS_SRC and LS_CRC can be further improved by using the confusion matrix of the classifier. Experimental results on four public face databases show that our methods not only generalize well to SSPP problem but also have strong robustness to occlusion; little pose variation; and the variations of expression, illumination, and time. Fan Liu 0003, Jinhui Tang 0001, Yan Song 0005, Liyan Zhang 0001, Zhenmin Tang |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2015 | Real-Time System for Driver Fatigue Detection by RGB-D CameraabstractDrowsy driving is one of the major causes of fatal traffic accidents. In this article, we propose a real-time system that utilizes RGB-D cameras to automatically detect driver fatigue and generate alerts to drivers. By introducing RGB-D cameras, the depth data can be obtained, which provides extra evidence to benefit the task of head detection and head pose estimation. In this system, two important visual cues (head pose and eye state) for driver fatigue detection are extracted and leveraged simultaneously. We first present a real-time 3D head pose estimation method by leveraging RGB and depth data. Then we introduce a novel method to predict eye states employing the WLBP feature, which is a powerful local image descriptor that is robust to noise and illumination variations. Finally, we integrate the results from both head pose and eye states to generate the overall conclusion. The combination and collaboration of the two types of visual cues can reduce the uncertainties and resolve the ambiguity that a single cue may induce. The experiments were performed using an inside-car environment during the day and night, and theyfully demonstrate the effectiveness and robustness of our system as well as the proposed methods of predicting head pose and eye states. Liyan Zhang 0001, Fan Liu 0003, Jinhui Tang 0001 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2014 | Local structure based sparse representation for face recognition with single sample per personabstractIn this paper, we propose local structure based sparse representation classification (LS SRC) to solve single sample per person (SSPP) problem. By adopting the “divide-conquer-aggregate” strategy, we successfully alleviate the dilemma of high data dimensionality and small samples, where we first divide the face into local blocks, and classify each local block, and then integrate all the classification results by voting. For each block, we further divide it into overlapped patches and assume that these patches lie in a linear subspace. This subspace assumption reflects local structure relationship of the overlapped patches and makes SRC feasible for SSPP problem. To lighten the computing burden, we further propose local structure based collaborative representation classification (LS CRC). Experimental results on three public face databases show that our methods not only generalize well to SSPP problem but also have strong robustness to expression, illumination, little pose variation, occlusion and time variation. Fan Liu 0003, Jinhui Tang 0001, Yan Song 0005, Xinguang Xiang, Zhenmin Tang |
ICIP | 1 |
| 2014 | Body Surface Context: A New Robust Feature for Action Recognition From Depth VideosabstractHuman action recognition in videos is useful for many applications. However, there still exist huge challenges in real applications due to the variations in the appearance, lighting condition and viewing angle, of the subjects. In this consideration, depth data have advantages over red, green, blue (RGB) data because of their spatial information about the distance between object and viewpoint. Unlike existing works, we utilize the 3-D point cloud, which contains points in the 3-D real-world coordinate system to represent the external surface of human body. Specifically, we propose a new robust feature, the body surface context (BSC), by describing the distribution of relative locations of the neighbors for a reference point in the point cloud in a compact and descriptive way. The BSC encodes the cylindrical angular of the difference vector based on the characteristics of human body, which increases the descriptiveness and discriminability of the feature. As the BSC is an approximate object-centered feature, it is robust to transformations including translations and rotations, which are very common in real applications. Furthermore, we propose three schemes to represent human actions based on the new feature, including the skeleton-based scheme, the random-reference-point scheme, and the spatial-temporal scheme. In addition, to evaluate the proposed feature, we construct a human action dataset by a depth camera. Experiments on three datasets demonstrate that the proposed feature outperforms RGB-based features and other existing depth-based features, which validates that the BSC feature is promising in the field of human action recognition. Yan Song 0005, Jinhui Tang 0001, Fan Liu 0003, Shuicheng Yan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2013 | WLBP: Weber local binary pattern for local image description
Fan Liu 0003, Zhenmin Tang, Jinhui Tang 0001 |
Neurocomputing | 1 |
| 2012 | Abnormal behavior recognition system for ATM monitoring by RGB-D cameraabstractIn this demo, we present an effective real-time system for ATM intelligent monitoring by using Kinect of Microsoft. With Kinect, we can easily detect people in ATM room and get their position information. By analyzing position information and video content, the system detects abnormal behaviors such as face-hiding, peeping and wandering, while records the time of these abnormal videos. Therefore, it not only prevents crimes but also helps to find suspects quickly after crimes have happened. The experimental results show that the system has the advantages of robustness, and provides a new mean for preventing financial crimes. Fan Liu 0003, Jinhui Tang 0001, Ruizhen Zhao, Zhenmin Tang |
ACM Multimedia | 1 |