VLDB 2026 Research / reviewers in the wild / expert
Li Liu 0002
dblp:33/4528-2
· DBLP profile ↗
179ranked-venue papers
23as first author
134since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 90 · 13 first-author · 66 since 2021Graphics, computer vision, multimedia, augmented reality and games · 71 · 11 first-author · 46 since 2021Applied, interdisciplinary, general and emerging computing · 30 · 28 since 2021Computer networks · 4 · 4 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deep Lookup NetworkabstractConvolutional neural networks are constructed with massive operations with different types and are highly computationally intensive. Among these operations, multiplication operation is higher in computational complexity and usually requires more energy consumption with longer inference time than other operations, which hinders the deployment of convolutional neural networks on mobile devices. In many resource-limited edge devices, complicated operations can be calculated via lookup tables to reduce computational cost. Motivated by this, in this paper, we introduce a generic and efficient lookup operation which can be used as a basic operation for the construction of neural networks. Instead of calculating the multiplication of weights and activation values, simple yet efficient lookup operations are adopted to compute their responses. To enable end-to-end optimization of the lookup operation, we construct the lookup tables in a differentiable manner and propose several training strategies to promote their convergence. By replacing computationally expensive multiplication operations with our lookup operations, we develop lookup networks for the image classification, image super-resolution, and point cloud classification tasks. It is demonstrated that our lookup networks can benefit from the lookup operations to achieve higher efficiency in terms of energy consumption and inference speed while maintaining competitive performance to vanilla convolutional networks. Extensive experiments show that our lookup networks produce state-of-the-art performance on different tasks (both classification and regression tasks) and different data types (both images and point clouds). Yulan Guo, Longguang Wang, Wendong Mao, Yingqian Wang 0002, Li Liu 0002, Wei An 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Probing Deep Into Temporal Profile Makes the Infrared Small Target Detector Much BetterabstractInfrared small target (IRST) detection is challenging in simultaneously achieving precise, robust, and efficient performance due to extremely dim targets and strong interference. Current learning-based methods attempt to leverage "more" information from both the spatial and the short-term temporal domains, but suffer from unreliable performance under complex conditions while incurring computational redundancy. In this paper, we explore the "more essential" information from a more crucial domain for the detection. Through theoretical analysis, we reveal that the global temporal saliency and correlation information in the temporal profile demonstrate significant superiority in distinguishing target signals from other signals. To investigate whether such superiority is preferentially leveraged by well-trained networks, we built the first prediction attribution tool in this field and verified the importance of the temporal profile information. Inspired by the above conclusions, we remodel the IRST detection task as a one-dimensional signal anomaly detection task, and propose an efficient deep temporal probe network (DeepPro) that only performs calculations in the time dimension for IRST detection. We conducted extensive experiments to fully validate the effectiveness of our method. The experimental results are exciting, as our DeepPro outperforms existing state-of-the-art IRST detection methods on widely-used benchmarks with extremely high efficiency, and achieves a significant improvement on dim targets and in complex scenarios. We provide a new modeling domain, a new insight, a new method, and a new performance, which can promote the development of IRST detection. Ruojing Li, Wei An 0003, Yingqian Wang 0002, Xinyi Ying, Yimian Dai, Longguang Wang, Yulan Guo, Li Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2026 | Diving Into Epipolar Transformers for Light Field Super-Resolution and Disparity EstimationabstractLight field (LF) cameras capture the light rays of a 3D scene from multiple views simultaneously, and thus provide a more immersive experience of the real world as compared to traditional cameras. Although significant progress has been made in various LF image processing tasks, it remains challenging to effectively model the non-local spatial-angular correlations inherent in LF images, particularly when dealing with complex disparity variations. In this paper, we focus on orthogonal epipolar geometry of LF images and propose a generic Epipolar Transformer mechanism that incorporates geometrically meaningful correlations along the epipolar lines. Our Epipolar Transformer mechanism enjoys the following benefits: learning effective and diverse LF feature representations, delivering satisfactory results without redundant architectural designs, and enabling flexible extension to various LF-related tasks with simple adaptations. For LF spatial and angular super-resolution, our methods not only achieve state-of-the-art performance on benchmark datasets, but also demonstrate superior and robust performance on large disparity variations. For disparity estimation, we explore the use of geometry information encoded in our Epipolar Transformer to directly regress the disparity results, effectively avoiding the limitation of a fixed maximum disparity. Zhengyu Liang, Yingqian Wang 0002, Longguang Wang, Jun-Gang Yang, Yulan Guo, Li Liu 0002, Shilin Zhou 0001, Wei An 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | ATRNet-STAR: A Large Dataset and Benchmark Toward Remote Sensing Object Recognition in the WildabstractThe absence of publicly available, large-scale, high-quality datasets for Synthetic Aperture Radar Automatic Target Recognition (SAR ATR) has significantly hindered the application of rapidly advancing deep learning techniques, which hold huge potential to unlock new capabilities in this field. This is primarily because collecting large volumes of diverse target samples from SAR images is prohibitively expensive, largely due to privacy concerns, the characteristics of microwave radar imagery perception, and the need for specialized expertise in data annotation. Throughout the history of SAR ATR research, there have been only a number of small datasets, mainly including targets like ships, airplanes, buildings, etc. There is only one vehicle dataset MSTAR collected in the 1990 s, which has been a valuable source for SAR ATR. To fill this gap, this paper introduces a large-scale, new dataset named ATRNet-STAR with 40 different vehicle categories collected under various realistic imaging conditions and scenes. It marks a substantial advancement in dataset scale and diversity, comprising over 190,000 well-annotated samples-$10\times$ larger than its predecessor, the famous MSTAR. Building such a large dataset is a challenging task, and the data collection scheme will be detailed. Secondly, we illustrate the value of ATRNet-STAR via extensively evaluating the performance of 15 representative methods with 7 different experimental settings on challenging classification and detection benchmarks derived from the dataset. Finally, based on our extensive experiments, we identify valuable insights for SAR ATR and discuss potential future research directions in this field. We hope that the scale, diversity, and benchmark of ATRNet-STAR can significantly facilitate the advancement of SAR ATR. Yongxiang Liu, Li Liu 0002, Jie Zhou 0031, Bowen Peng, Xuying Xiong, Wei Yang 0046, Tianpeng Liu, Zhen Liu 0004, Xiang Li 0014 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | S4ST: A Strong, Self-Transferable, faSt, and Simple Scale Transformation for Data-Free Transferable Targeted Attack
Yongxiang Liu, Bowen Peng, Li Liu 0002, Xiang Li 0014 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Prompt is All You Need: Prompting Foundation Models for Large-Scale Self-Supervised Semantic SegmentationabstractThis paper addresses the important and challenging task of large-scale unsupervised semantic segmentation (LUSS). We present the first attempt to unleash the power of foundation models (FMs) for the challenging, dense prediction task LUSS, and our main objective is to present simple, effective yet efficient solutions for LUSS, namely Prompting foundation models for LUSS (PLUSS). Firstly, we proposed a cascade framework PLUSS$_\alpha$α by effectively marrying CLIPS, Grounding DINO, and SAM in a zero-shot manner. This cascade architecture automatically generates semantic and spatial prompts for SAM, establishing a strong baseline that significantly outperforms previous state-of-the-art methods. Building upon this foundation, we propose PLUSS$_\beta$β, which addresses the critical bottleneck of prompt quality through two novel tuner modules: a semantic tuner that enhances fine-grained category discrimination via visual prompt tuning, and a box tuner that improves object localization through cross-modal feature fusion. Both tuners are optimized by capitalizing on the knowledge already present within the foundation models themselves, deriving self-supervised signals from internal model consistency. This approach requires no external supervision or updates to the foundation models' parameters. Extensive experiments on ImageNet-S benchmarks demonstrate that PLUSS$_\beta$β achieves remarkable performance improvements, surpassing the previous best method by 39.6%, 27.3%, and 22.6% in mIoU for 50, 300, and 919 categories respectively. Our approach exhibits robust category-shape representation across varying object sizes and dataset scales, while maintaining strong generalization capabilities for open-vocabulary tasks. The proposed framework provides a solid baseline for adapting foundation models to downstream vision tasks. Jiaojiao Su, Qiwu Luo, Shuzhou Sun, Yuenan Hou, Xinyu Zhang 0010, Janne Heikkilä, Chunhua Yang 0001, Li Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2026 | Step-Wise Distribution-Aligned Style Prompt Tuning for Source-Free Cross-Domain Few-Shot LearningabstractExisting cross-domain few-shot learning (CDFSL) methods, which develop training strategies in the source domain to enhance model transferability, face challenges when applied to large-scale pre-trained models (LMs), as their source domains and training strategies are not accessible. Besides, fine-tuning LMs specifically for CDFSL requires substantial computational resources, which limits their practicality. Therefore, this paper investigates the source-free CDFSL (SF-CDFSL) problem to solve the few-shot learning (FSL) task in target domain using only a pre-trained model and a few target samples, without requiring source data or training strategies. However, the inaccessibility of source data prevents explicitly reducing the domain gaps between the source and target. To tackle this challenge, this paper proposes a novel approach, Step-wise Distribution-aligned Style Prompt Tuning (StepSPT), to implicitly narrow the domain gaps from the perspective of prediction distribution optimization. StepSPT initially proposes a style prompt that adjusts the target samples to mirror the expected distribution. Furthermore, StepSPT tunes the style prompt and classifier by exploring a dual-phase optimization process (external and internal processes). In the external process, a step-wise distribution alignment strategy is introduced to tune the proposed style prompt by factorizing the prediction distribution optimization problem into the multi-step distribution alignment problem. In the internal process, the classifier is updated via standard cross-entropy loss. Evaluation on 5 datasets illustrates the superiority of StepSPT over existing prompt tuning-based methods and state-of-the-art methods (SOTAs). Furthermore, ablation studies and performance analyzes highlight the efficacy of StepSPT. Huali Xu, Li Liu 0002, Tianpeng Liu, Shuaifeng Zhi, Shuzhou Sun, Ming-Ming Cheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Band-Kernel Stochastic Learning for Unsupervised Blind Hyperspectral Image Super-ResolutionabstractHyperspectral image super-resolution (HSI-SR) is fundamentally more difficult than RGB image SR, since its ultrahigh spectral dimensionality. Existing supervised methods rely on labeled training data to obtain data prior, which incurs prohibitive collection costs and limits generalization. Unsupervised methods individually preset the band and kernel with handcrafted priors, whereas this decoupling modeling artificially creates a complexity-performance trade-off in the selected band number. To address these issues, we propose BKX-HMM, a unified statistical framework for blind HSI-SR, which uniformly models the band selection, kernel estimation, and HSI restoration through the state transition of a hidden Markov model (HMM). BKX-HMM redefines the trade-off as a distributional fitting problem: each Markov transition progressively learns optimal parameters of full-band distribution via limited spectral observations. Based on BKX-HMM, we propose BKSR, the first unsupervised blind HSI-SR method, which consists of three synergistic modules: Gibbs sampling-based band selection (GBS), test-time-training kernel estimation (TKE), and robust HSI restoration (RHR). These modules form a closed-loop optimization cycle: i) In GBS, the dynamic ergodicity of Gibbs sampling provides a global spectral view for kernel estimation and HSI restoration while maintaining local spectral computations; ii) In TKE, the GBS-sampled bands guide the kernel estimator update, achieving a learnable sampling-based mechanism, which refines kernel estimation to regularize RHR's diffusion trajectory; iii) In RHR, a spectral hyper-Laplacian prior is integrated into the reverse process of an off-the-shelf diffusion model, which achieves non-i.i.d. noise robust HSI restoration, feedback reweights band and kernel importance for subsequent GBS and TKE iterations. Extensive experiments on both synthetic and real HSI datasets demonstrate our BKSR's superiority over baseline methods across diverse scenarios (e.g., unknown Gaussian/motion kernel, non-i.i.d. noise) while maintaining comparable computational costs to the classic band selection methods. Zhixiong Yang 0001, Jingyuan Xia, Shengxi Li, Lingyu Zheng, Shuanghui Zhang, Li Liu 0002, Yaowen Fu, Yongxiang Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | SRFormerV2: Taking a Closer Look at Permuted Self-Attention for Image Super-ResolutionabstractPrevious works have shown that increasing the window size for Transformer-based image super-resolution models (e.g., SwinIR) can significantly improve the model performance. Still, the computation overhead is also considerable when the window size gradually increases. In this paper, we present SRFormer, a simple but novel method that can enjoy the benefit of large window self-attention but introduces even less computational burden. The core of our SRFormer is the permuted self-attention (PSA), which strikes an appropriate balance between the channel and spatial information for self-attention. Without any bells and whistles, we show that our SRFormer achieves a 33.86 dB PSNR score on the Urban100 dataset, which is 0.46 dB higher than that of SwinIR but uses fewer parameters and computations. In addition, we also attempt to scale up the model by further enlarging the window size and channel numbers to explore the potential of Transformer-based models. Experiments show that our scaled model, named SRFormerV2, can further improve the results and achieves state-of-the-art. We hope our simple and effective approach could be useful for future research in super-resolution model design. Yupeng Zhou, Zhen Li 0031, Chunle Guo, Li Liu 0002, Ming-Ming Cheng, Qibin Hou |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Reinforcement Learning-Based Multi-Target Detection Method for MIMO Radar Assisted by Strong Target LimitationabstractUnder the background of co-located MIMO radar, the existing reinforcement learning (RL)-based multi-target detection methods generally perform poorly on weak targets. In our previous work, we have proposed a beam optimization scheme with strong target limitation and gave a solution approach based on multi-rank beamformer, to achieve focusing more radar transmit power on weak targets. In this letter, we further propose a solution approach based on inner convex approximation, which can achieve a higher power gain due to its improved freedom. In addition, we also design an approach for choosing the focused angle cells of radar by fusing the statistical prior information from previous time. Summarizing the above improvements, we propose a RL-based multi-target detection method for MIMO radar assisted by strong target limitation. The experiments show that our method owns better performance on weak targets than its competitors while maintaining the excellent performance on strong targets. Xijie Wu, Tianpeng Liu, Yongxiang Liu, Li Liu 0002 |
IEEE Signal Process. Lett. | 4 |
| 2026 | Policy Generalization Enhancement for UAV Active Object Detection via Divide-and-Conquer Sharpness-Aware Gradient MatchingabstractTarget detection in aerial images captured by unmanned aerial vehicles has long been hampered by occlusion. Active Object Detection (AOD) aims to fundamentally address this issue from the active vision perspective, typically realized through the Deep Reinforcement Learning (DRL) paradigm. However, the active observation policy often suffers from low generalization ability, thus limiting its practical application. In this paper, we propose Divide-and-Conquer Sharpness-Aware Gradient Matching (DC-SAGM), a novel sharpness-based Domain Generalization (DG) method, to effectively enhance the generalization capacity of the agent’s policy. Specifically, we train the agent to learn the active observation policy using the conventional DRL approach. Sharpness-Aware Gradient Matching (SAGM) is employed during training, improving the model’s generalization performance by minimizing the sharpness metric of the loss landscape. Nevertheless, the imperfect state representation and classifier preference in the AOD problem lead to fierce gradient conflicts, deteriorating the effectiveness of SAGM. We address this incompatibility by using a divide-and-conquer strategy and exclude gradient conflicts via the majority-rule gradient surgery operation. Extensive experimental results on the UEVAVD dataset validate DC-SAGM’s superiority in helping the agent’s policy achieve better generalization compared to extensive policy learning approaches. Xinhua Jiang, Tianpeng Liu, Li Liu 0002, Zhenghui Gong, Yongxiang Liu, Xiang Li 0014 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | A Policy-Driven Black-Box Adversarial Example With Location Optimization Against 3D Object DetectionabstractAdversarial attack strategies for 3D object detection have highlighted the critical importance of addressing security concerns in this domain. However, white-box methods require full access to the victim model in large-scale point cloud applications. To this end, we propose a novel Policy-Driven Black-box Attack (BAT) that is designed to optimize attack locations without necessitating detailed knowledge of the victim models. First, we introduce a density-aware pattern generator that creates scene-adaptive attack clusters. Second, we leverage the deep deterministic policy gradient in deep reinforcement learning to train an attack agent capable of targeting the victim model. Ultimately, the attack agent is iteratively directed towards optimal attack locations through the joint application of critic loss and actor loss. To the best of our knowledge, this represents the first reinforcement learning-based black-box attack applied to practical 3D object detection. Experimental results on the KITTI, nuScenes, and Waymo datasets demonstrate that BAT effectively diminishes the accuracy of notable models. Importantly, BAT significantly enhances the attack success rate (surpassing state-of-the-art both white-box and black-box methods) and increases transferability (by 20 times) through simple deep deterministic policy gradient, thus establishing a new baseline for adversarial attacks in 3D object detection. Ting Han 0001, Xiaobin Wu, Chaolei Wang, Huan Luo 0001, Xiaochun Cao, Li Liu 0002, Yiping Chen 0002 |
IEEE Trans. Image Process. | 7 |
| 2026 | Fast Diffusion-Based Camouflaged Object Detection via Asynchronous Denoising and Linear AttentionabstractCamouflaged object detection (COD) aims to precisely locate and identify objects concealed within their surrounding backgrounds, a task that has been significantly advanced by recent powerful diffusion models. Nevertheless, existing diffusion-based COD methods demonstrate suboptimal performance in inference speed, attributable to two key factors: time-consuming iterative sampling processes and the quadratic complexity introduced by self-attention mechanisms. To address these limitations, we propose Fast CamoDiff, a novel and efficient diffusion-based model for COD. To tackle the computational overhead, we incorporate an asynchronous denoising paradigm leveraging dynamic encoding and early termination, which significantly reduces computational costs during the sampling process. Additionally, we introduce the linear attention mechanism from state space models (SSM) into the diffusion process to achieve linear computational complexity. Meanwhile, to address the side effect of compromising fine-grained feature preservation through information compression, we design a novel bidirectional pooling layer that enhances the model's capability of preserving detailed features without sacrificing computational benefits. Extensive experimental results on four public datasets demonstrate that Fast-CamoDiff achieves superior detection accuracy with only 36.5M parameters (67.4% reduction) and 28 FPS (2× speedup) compared to previous state-of-the-art diffusion-based models. The source code will be available at https://github.com/wty-team/diff-ssm. Xinghua Xu, Changchong Sheng, Shaohua Qiu, Li Liu 0002, Denghua Guo |
IEEE Trans. Multim. | 5 |
| 2025 | FaDeN: Fast Depth-Supervised NeRFs with RGB-D Cameras
Janne Mustaniemi, Li Liu 0002, Janne Heikkilä |
CAIP (1) | 2 |
| 2025 | StyleSRN: Scene Text Image Super-Resolution with Text Style Embedding
Shengrong Yuan, Ke Hao, Xuqi Ma, Changxin Gao, Li Liu 0002, Nong Sang |
ICCV | 6 |
| 2025 | When Pixel Difference Patterns Meet ViT: PiDiViT for Few-Shot Object Detection
Yongxiang Liu, Canyu Mo, Bowen Peng, Li Liu 0002 |
ICCV | 6 |
| 2025 | Efficient Binarized Neural Network Intellectual Property ProtectionabstractBinary Neural Networks (BNNs) quantize weights and activations to −1 and +1 to achieve significant memory reduction and computational acceleration, which have been extensively explored in image and video tasks. The development for training high-accuracy BNNs holds substantial commercial value, underscoring the critical importance of emphasizing risks related to intellectual property (IP) infringement. However, existing IP protection methods focus on float-point models, neglecting protection for low-bit models. To this end, we provide a study tailored for IP protection of BNNs. We adopt a passport-based watermarking method as our baseline, known for resisting both removal and ambiguity attacks. We observe that the discretization of weights and activations introduces instability in the joint training stage, leading to a significant accuracy decrease in the target model. To overcome this challenge, we present a novel passport-aware module for BNNs to improve gradient optimization stability and reduce sensitivity to binary weights flipping. Furthermore, we conduct a comprehensive evaluation of the robustness of BNNs against various attacks. Extensive evaluations show our method improves model performance and robustness against attacks. We hope that our research will contribute to the further advancement of IP protection for low-bit networks, especially BNNs. Yuchen Sun 0001, Li Liu 0002 |
ICME | 4 |
| 2025 | ETA: Learning Optical Flow with Efficient Temporal AttentionabstractConsidering the potential of using multi-frame information to solve the occlusion problem, we introduce a novel idea of multi-frame information integration, which uses the attention mechanism to fuse the temporal information from the previous frame. The idea can effectively improve the estimation accuracy in occluded regions and optimize the inference speed under multi-frame settings. Meanwhile, we suggest the concept of attention confidence to provide an explicit value criterion for the model to utilize useful attention information more efficiently. Furthermore, we propose an Efficient Temporal Attention network (ETA), which achieves promising results on Sintel and KITTI benchmarks, especially with a 9.4% error reduction compared to the baseline method GMA on Sintel (test) Clean. Bo Wang 0144, Zhenping Sun, Yang Yu 0014, Li Liu 0002, Jian Li 0003, Dewen Hu |
IROS | 4 |
| 2025 | Luminance-Aware Statistical Quantization: Unsupervised Hierarchical Learning for Illumination EnhancementabstractLow-light image enhancement (LLIE) faces persistent challenges in balancing reconstruction fidelity with cross-scenario generalization. While existing methods predominantly focus on deterministic pixel-level mappings between paired low/normal-light images, they often neglect the continuous physical process of luminance transitions in real-world environments, leading to performance drop when normal-light references are unavailable. Inspired by empirical analysis of natural luminance dynamics revealing power-law distributed intensity transitions, this paper introduces Luminance-Aware Statistical Quantification (LASQ), a novel framework that reformulates LLIE as a statistical sampling process over hierarchical luminance distributions. Our LASQ re-conceptualizes luminance transition as a power-law distribution in intensity coordinate space that can be approximated by stratified power functions, therefore, replacing deterministic mappings with probabilistic sampling over continuous luminance layers. A diffusion forward process is designed to autonomously discover optimal transition paths between luminance layers, achieving unsupervised distribution emulation without normal-light references.
In this way, it considerably improves the performance in practical situations, enabling more adaptable and versatile light restoration. This framework is also readily applicable to cases with normal-light references, where it achieves superior performance on domain-specific datasets alongside better generalization-ability across non-reference datasets. The code is available at: https://github.com/XYLGroup/LASQ. Derong Kong, Zhixiong Yang 0001, Shengxi Li, Shuaifeng Zhi, Li Liu 0002, Zhen Liu 0004, Jingyuan Xia |
NeurIPS | 5 |
| 2025 | Camouflaged Object Detection with Adaptive Partition and Background Retrieval
Bowen Yin, Xuying Zhang, Li Liu 0002, Ming-Ming Cheng, Yongxiang Liu, Qibin Hou |
Int. J. Comput. Vis. | 3 |
| 2025 | TS-BiT: Two-Stage Binary Transformer for ORSI Salient Object DetectionabstractVision transformers (ViTs) have demonstrated superior performance in various remote sensing tasks, such as optical remote sensing image salient object detection (ORSI-SOD). However, the high resolution of remote sensing images and the substantial computational costs pose significant challenges for deploying existing methods on resource-constrained devices. Model binarization significantly reduces computational costs and storage requirements by constraining weights and activations to 1-bit representations, which has been widely explored in convolutional neural networks (CNNs). However, directly applying binary methods to ViTs poses challenges since quantization errors hinder the ability to capture the similarity between tokens, resulting in significant performance degradation in detecting salient objects in complex ORSI scenarios. To address this issue, we propose two-stage binary transformer (TS-BiT) for the ORSI-SOD task to preserve information on salient objects under 1-bit representation. Specifically, we design a two-stage central-aware softmax binarization (TCSB) strategy to reduce quantization errors arising from substantial discrepancies in the long-tail distribution of multihead attention. Furthermore, we develop a scalable hyperbolic tangent function to approximate the gradients of the Sign function within each binarization group, substantially mitigating quantization errors during the binarization of softmax attention. Extensive experiments demonstrate that our method outperforms existing binary ViT approaches on ORSSD, EORSSD, and ORSI-4199 datasets. Tianpeng Liu, Li Liu 0002 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | Generalized Semantic Contrastive Learning via Embedding Side Information for Few-Shot Object DetectionabstractThe objective of few-shot object detection (FSOD) is to detect novel objects with few training samples. The core challenge of this task is how to construct a generalized feature space for novel categories with limited data on the basis of the base category space, which could adapt the learned detection model to unknown scenarios. Most existing fine-tuning-based approaches tackle the challenge via pre-training a feature extractor based on the base categories and then fine-tuning the detector through the novel categories. However, limited by insufficient samples for novel categories, two issues still exist: (1) the features of the novel category are easily implicitly represented by the features of the base category, leading to inseparable classifier boundaries, (2) novel categories with fewer data are not enough to fully represent the distribution, where the model fine-tuning is prone to overfitting. To address these issues, we introduce the side information to alleviate the negative influences derived from the feature space and sample viewpoints and formulate a novel generalized feature representation learning method for FSOD. Specifically, we first utilize embedding side information to construct a knowledge matrix to quantify the semantic relationship between the base and novel categories. Then, to strengthen the discrimination between semantically similar categories, we further develop contextual semantic supervised contrastive learning which embeds side information. Furthermore, to prevent overfitting problems caused by sparse samples, a side-information guided region-aware masked module is introduced to augment the diversity of samples, which finds and abandons biased information that discriminates between similar categories via counterfactual explanation, and refines the discriminative representation space further. Finally, we theoretically analyze the generalization bound for introducing our proposed module and demonstrate that our proposed model can effectively reduce the upper bound of the generalization error. Extensive experiments using ResNet and ViT backbones on PASCAL VOC, MS COCO, LVIS V1, FSOD-1 K, and FSVOD-500 benchmarks demonstrate that our model outperforms the previous state-of-the-art methods, significantly improving the ability of FSOD in most shots/splits. The code is released athttps://github.com/RuoyuChen10/CCL-FSOD. Ruoyu Chen 0001, Hua Zhang 0008, Jingzhi Li 0002, Li Liu 0002, Zhen Huang 0006, Xiaochun Cao |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Enhancing Representations Through Heterogeneous Self-Supervised LearningabstractIncorporating heterogeneous representations from different architectures has facilitated various vision tasks, e.g., some hybrid networks combine transformers and convolutions. However, complementarity between such heterogeneous architectures has not been well exploited in self-supervised learning. Thus, we propose Heterogeneous Self-Supervised Learning (HSSL), which enforces a base model to learn from an auxiliary head whose architecture is heterogeneous from the base model. In this process, HSSL endows the base model with new characteristics in a representation learning way without structural changes. To comprehensively understand the HSSL, we conduct experiments on various heterogeneous pairs containing a base model and an auxiliary head. We discover that the representation quality of the base model moves up as their architecture discrepancy grows. This observation motivates us to propose a search strategy that quickly determines the most suitable auxiliary head for a specific base model to learn and several simple but effective methods to enlarge the model discrepancy. The HSSL is compatible with various self-supervised methods, achieving superior performances on various downstream tasks, including image classification, semantic segmentation, instance segmentation, and object detection. Zhongyu Li 0006, Bowen Yin, Yongxiang Liu, Li Liu 0002, Ming-Ming Cheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | A Causal Adjustment Module for Debiasing Scene Graph GenerationabstractWhile recent debiasing methods for Scene Graph Generation (SGG) have shown impressive performance, these efforts often attribute model bias solely to the long-tail distribution of relationships, overlooking the more profound causes stemming from skewed object and object pair distributions. In this paper, we employ causal inference techniques to model the causality among these observed skewed distributions. Our insight lies in the ability of causal inference to capture the unobservable causal effects between complex distributions, which is crucial for tracing the roots of model bias. Specifically, we introduce the Mediator-based Causal Chain Model (MCCM), which, in addition to modeling causality among objects, object pairs, and relationships, incorporates mediator variables, i.e., cooccurrence distribution, for complementing the causality. Following this, we propose the Causal Adjustment Module (CAModule) to estimate the modeled causal structure, using variables from MCCM as inputs to produce a set of adjustment factors aimed at correcting biased model predictions. Moreover, our method enables the composition of zero-shot relationships, thereby enhancing the model's ability to recognize such relationships. Experiments conducted across various SGG backbones and popular benchmarks demonstrate that CAModule achieves state-of-the-art mean recall rates, with significant improvements also observed on the challenging zero-shot recall rate metric. Li Liu 0002, Shuzhou Sun, Shuaifeng Zhi, Fan Shi 0003, Zhen Liu 0004, Janne Heikkilä, Yongxiang Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Sample Adaptive Localized Simple Multiple Kernel K-Means and its Application in Parcellation of Human Cerebral CortexabstractSimple multiple kernel k-means (SMKKM) introduces a new minimization-maximization learning paradigm for multi-view clustering and makes remarkable achievements in some applications. As one of its variants, localized SMKKM (LSMKKM) is recently proposed to capture the variation among samples, focusing on reliable pairwise samples, which should keep together and cut off unreliable, farther pairwise ones. Though demonstrating effectiveness, we observe that LSMKKM indiscriminately utilizes the variation of each sample, resulting in unsatisfying clustering performance. To overcome this limitation, we propose a sample adaptive localized SMKKM (SAL-SMKKM) algorithm where the weight of the local alignment for each sample can be adaptively adjusted, resulting in a more challenging tri-level minimization-minimization-maximization. To deal with it, we reformulate it into a minimization problem of an optimal function characterized by minimization-maximization dynamics, prove its differentiability, and develop a reduced gradient descent method to optimize it. We then theoretically analyze the clustering performance of the proposed SAL-SMKKM by deriving its generalization error bound. In addition, we empirically evaluate the clustering performance of the proposed SAL-SMKKM on several benchmark datasets. Experiment results clearly indicate that proposed algorithms consistently outperform state-of-the-art ones. Finally, we apply the proposed SAL-SMKKM to the multi-modal parcellation of the human cerebral cortex, which is essential and helpful to understanding brain organization and function. As seen, SAL-SMKKM achieves accurate parcellation in an automatic and objective manner without any manual intervention, which once again demonstrates its validity and effectiveness in practical applications. Xinwang Liu 0002, Yi Zhang 0104, Li Liu 0002, Chang Tang, Long Lan, Dewen Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Rapid Salient Object Detection With Difference Convolutional Neural NetworksabstractThis paper addresses the challenge of deploying salient object detection (SOD) on resource-constrained devices with real-time performance. While recent advances in deep neural networks have improved SOD, existing top-leading models are computationally expensive. We propose an efficient network design that combines traditional wisdom on SOD and the representation power of modern CNNs. Like biologically-inspired classical SOD methods relying on computing contrast cues to determine saliency of image regions, our model leverages Pixel Difference Convolutions (PDCs) to encode the feature contrasts. Differently, PDCs are incorporated in a CNN architecture so that the valuable contrast cues are extracted from rich feature maps. For efficiency, we introduce a difference convolution reparameterization (DCR) strategy that embeds PDCs into standard convolutions, eliminating computation and parameters at inference. Additionally, we introduce SpatioTemporal Difference Convolution (STDC) for video SOD, enhancing the standard 3D convolution with spatiotemporal contrast capture. Our models, SDNet for image SOD and STDNet for video SOD, achieve significant improvements in efficiency-accuracy trade-offs. On a Jetson Orin device, our models with $< $< 1M parameters operate at 46 FPS and 150 FPS on streamed images and videos, surpassing the second-best lightweight models in our experiments by more than $2\times$2× and $3\times$3× in speed with superior accuracy. Zhuo Su 0002, Li Liu 0002, Matthias Müller 0011, Diana Wofk, Ming-Ming Cheng, Matti Pietikäinen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | A Reverse Causal Framework to Mitigate Spurious Correlations for Debiasing Scene Graph GenerationabstractExisting two-stage Scene Graph Generation (SGG) frameworks typically incorporate a detector to extract relationship features and a classifier to categorize these relationships; therefore, the training paradigm follows a causal chain structure, where the detector's inputs determine the classifier's inputs, which in turn influence the final predictions. However, such a causal chain structure can yield spurious correlations between the detector's inputs and the final predictions, i.e., the prediction of a certain relationship may be influenced by other relationships. This influence can induce at least two observable biases: tail relationships are predicted as head ones, and foreground relationships are predicted as background ones; notably, the latter bias is seldom discussed in the literature. To address this issue, we propose reconstructing the causal chain structure into a reverse causal structure, wherein the classifier's inputs are treated as the confounder, and both the detector's inputs and the final predictions are viewed as causal variables. Specifically, we term the reconstructed causal paradigm as the Reverse causal Framework for SGG (RcSGG). RcSGG initially employs the proposed Active Reverse Estimation (ARE) to intervene on the confounder to estimate the reverse causality, i.e., the causality from final predictions to the classifier's inputs. Then, the Maximum Information Sampling (MIS) is suggested to enhance the reverse causality estimation further by considering the relationship information. Theoretically, RcSGG can mitigate the spurious correlations inherent in the SGG framework, subsequently eliminating the induced biases. Comprehensive experiments on popular benchmarks and diverse SGG frameworks show the state-of-the-art mean recall rate. Shuzhou Sun, Li Liu 0002, Tianpeng Liu, Shuaifeng Zhi, Ming-Ming Cheng, Janne Heikkilä, Yongxiang Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | SceneTracker: Long-Term Scene Flow Estimation NetworkabstractConsidering that scene flow estimation has the capability of the spatial domain to focus but lacks the coherence of the temporal domain, this study proposes long-term scene flow estimation (LSFE), a comprehensive task that can simultaneously capture the fine-grained and long-term 3D motion in an online manner. We introduce SceneTracker, the first LSFE network that adopts an iterative approach to approximate the optimal 3D trajectory. The network dynamically and simultaneously indexes and constructs appearance correlation and depth residual features. Transformers are then employed to explore and utilize long-range connections within and between trajectories. With detailed experiments, SceneTracker shows superior capabilities in addressing 3D spatial occlusion and depth noise interference, highly tailored to the needs of the LSFE task. We build a real-world evaluation dataset, LSFDriving, for the LSFE field and use it in experiments to further demonstrate the advantage of SceneTracker in generalization abilities. Bo Wang 0144, Jian Li 0003, Yang Yu 0014, Li Liu 0002, Zhenping Sun, Dewen Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Visible-Thermal Tiny Object Detection: A Benchmark Dataset and BaselinesabstractVisible-thermal small object detection (RGBT SOD) is a significant yet challenging task with a wide range of applications, including video surveillance, traffic monitoring, search and rescue. However, existing studies mainly focus on either visible or thermal modality, while RGBT SOD is rarely explored. Although some RGBT datasets have been developed, the insufficient quantity, limited diversity, unitary application, misaligned images and large target size cannot provide an impartial benchmark to evaluate RGBT SOD algorithms. In this paper, we build the first large-scale benchmark with high diversity for RGBT SOD (namely RGBT-Tiny), including 115 paired sequences, 93 K frames and 1.2 M manual annotations. RGBT-Tiny contains abundant objects (7 categories) and high-diversity scenes (8 types that cover different illumination and density variations). Note that, over 81% of objects are smaller than 16×16, and we provide paired bounding box annotations with tracking ID to offer an extremely challenging benchmark with wide-range applications, such as RGBT image fusion, object detection and tracking. In addition, we propose a scale adaptive fitness (SAFit) measure that exhibits high robustness on both small and large objects. The proposed SAFit can provide reasonable performance evaluation and promote detection performance. Based on the proposed RGBT-Tiny dataset, extensive evaluations have been conducted with IoU and SAFit metrics, including 30 recent state-of-the-art algorithms that cover four different types (i.e., visible generic object detection, visible SOD, thermal SOD and RGBT object detection). Xinyi Ying, Wei An 0003, Ruojing Li, Boyang Li 0007, Zhaoxu Li, Yingqian Wang 0002, Mingyuan Hu, Zaiping Lin, Shilin Zhou 0001, Li Liu 0002, Weidong Sheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 15 |
| 2025 | Few-Shot Class-Incremental Learning for Classification and Object Detection: A SurveyabstractFew-shot Class-Incremental Learning (FSCIL) presents a unique challenge in Machine Learning (ML), as it necessitates the Incremental Learning (IL) of new classes from sparsely labeled training samples without forgetting previous knowledge. While this field has seen recent progress, it remains an active exploration area. This paper aims to provide a comprehensive and systematic review of FSCIL. In our in-depth examination, we delve into various facets of FSCIL, encompassing the problem definition, the discussion of the primary challenges of unreliable empirical risk minimization and the stability-plasticity dilemma, general schemes, and relevant problems of IL and Few-shot Learning (FSL). Besides, we offer an overview of benchmark datasets and evaluation metrics. Furthermore, we introduce the Few-shot Class-incremental Classification (FSCIC) methods from data-based, structure-based, and optimization-based approaches and the Few-shot Class-incremental Object Detection (FSCIOD) methods from anchor-free and anchor-based approaches. Beyond these, we present several promising research directions within FSCIL that merit further investigation. Li Liu 0002, Olli Silvén, Matti Pietikäinen, Dewen Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | MaDiNet: Mamba Diffusion Network for SAR Target DetectionabstractThe fundamental challenge in SAR target detection lies in developing discriminative, efficient, and robust representations of target characteristics within intricate non-cooperative environments. However, accurate target detection is impeded by factors including the sparse distribution and discrete features of the targets, as well as complex background interference. In this study, we propose a Gamma Diffusion Model Network with MambaSAR module (MaDiNet) for SAR target detection. Specifically, MaDiNet leverages the Gamma distribution to model the statistical characteristics of SAR images, and conceptulizes SAR target detection as the task of generating target bounding boxes in the image space. Furthermore, we design a MambaSAR module to capture intricate spatial structural information of targets and enhance the capability of the model to differentiate between targets and complex backgrounds. The experimental results on multi-class target detection datasets have all achieved SOTA, with a particularly notable improvement of 6.7% in mAP50 on the ODSOG-1.0 dataset, proving the effectiveness of the proposed network. Code is available at https://github.com/JoyeZLearning/MaDiNet. Jie Zhou 0031, Yongxiang Liu, Bowen Peng, Li Liu 0002, Xiang Li 0014 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Observations Temporal Permutation-Based Self-Supervised Reinforcement Learning for UAV Active Object DetectionabstractIn passive ground target detection using Unmanned Aerial Vehicles (UAVs), some detrimental factors like occlusion significantly impact target detection performance. Active Object Detection offers an effective way to address it, which usually uses Deep Reinforcement Learning (DRL) to plan UAV’s viewpoint for favorable observations. However, existing DRL-based AOD methods often suffer from low sample efficiency and poor generalization due to inadequate state representation learned by the policy network. Inspired by human scene understanding where their spatial representation of the scene remains consistent despite different observation orders, we design a self-supervised state representation learning method based on Observations Temporal Permutation (OTP) to improve the state representation of the agent’s policy network. We require the policy network to output consistent action value estimates for observation sequences with the same content but different temporal orders. Besides, we use the state representation to predict the target orientation variations in the observation sequence, which further regularizes and facilitates the state representation learning process. Finally, we design multiple experiments based on the UEVAVD dataset to compare the proposed method with existing self-supervised state representation learning methods for the AOD task. The experimental results demonstrate that the OTP method can help the agent’s policy network learn a better state representation, thus achieving higher policy learning sample efficiency and stronger policy generalization. Xinhua Jiang, Tianpeng Liu, Li Liu 0002, Zhen Liu 0004, Yongxiang Liu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Infrared Small Target Detection in Satellite Videos: A New Dataset and a Novel Recurrent Feature Refinement FrameworkabstractMultiframe infrared small target (MIRST) detection in satellite videos has been a long-standing, fundamental yet challenging task for decades, and the challenges can be summarized as follows. First, the extremely small target size, highly complex clutter & noise and various satellite motions result in limited feature representation, high false alarms and difficult motion analyses. In addition, existing methods are primarily designed for static or slightly adjusted perspectives captured by short-distance platforms, which cannot generalize well to complex background motion in satellite videos. Second, the lack of a large-scale publicly available MIRST dataset in satellite videos greatly hinders the algorithm development. To address the aforementioned challenges, in this article, we first build a large-scale dataset for MIRST detection in satellite videos (namely IRSatVideo-LEO), and then develop a recurrent feature refinement (RFR) framework as the baseline method for satellite motion estimation and compensation. Specifically, IRSatVideo-LEO is a semi-simulated dataset with synthesized satellite motion, target appearance, trajectory, and intensity, which can provide a standard toolbox for satellite video generation and a reliable evaluation platform to facilitate algorithm development. For the baseline method, RFR is proposed to be equipped with existing powerful CNN-based methods for long-term temporal dependency exploitation and integrated motion compensation and MIRST detection. Specifically, a pyramid deformable alignment (PDA) module is proposed to achieve effective feature alignment, and a temporal-spatial–frequent modulation (TSFM) module is proposed to achieve efficient feature aggregation and enhancement. Extensive experiments have been conducted to demonstrate the effectiveness and superiority of our scheme. The comparative results show that ResUNet equipped with RFR outperforms the state-of-the-art MIRST detection methods. The dataset and code are available athttps://github.com/XinyiYing/RFR. Xinyi Ying, Li Liu 0002, Zaiping Lin, Yangsi Shi, Yingqian Wang 0002, Ruojing Li, Boyang Li 0007, Shilin Zhou 0001, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Heterogeneous Binary Pixel Difference Networks for Remote Sensing Object DetectionabstractRecent research in remote sensing object detection (RSOD) has significantly advanced the development of vision foundation models. However, deploying these models on resource-constrained edge devices is challenging due to their high computational demands. Binarized detectors utilize binary neural networks (BNNs) to achieve extreme compression by quantizing weights and activations to +1 or −1, which have been extensively studied for generic object detection tasks. In remote sensing images, the objects of interest typically exhibit weak responses, and the images often contain numerous unique local areas. Feature binarization in these images can lead to substantial loss of object contrast and scale prior information, which exacerbates performance issues, particularly for small objects, resulting in significant performance degradation. To address these challenges, we propose a novel binarized detector for RSOD named the heterogeneous binary pixel difference network (HBiPiDiNet). Initially, we developed a binary pixel difference convolution (BiPDC) that integrates local binary patterns (LBPs) to capture local contrast information with traditional binary convolution, thereby enhancing the representation of small objects. Subsequently, we constructed heterogeneous kernel fusion convolution blocks (HKFCB) based on BiPDC and standard binary convolution. The HKFCB comprises multiple BiPDCs at different scales, effectively representing BiPDC under multiscale LBP and multiscale binary convolutions. Extensive experiments demonstrate that our proposed method significantly enhances the performance of state-of-the-art binary detection methods across three remote sensing datasets: AI-TOD, VisDrone2019, and DIOR. We have released our code and models athttps://github.com/yuhua666/HBiPiDiNet/tree/main. Jialei Zhan, Liang Bai 0003, Tianpeng Liu, Fan Shi 0003, Yongxiang Liu, Li Liu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Advancing Segment Anything Model for Efficient Salient Object Detection in Remote Sensing ImagesabstractSalient object detection in optical remote sensing images (ORSI-SOD) often relies on leveraging pre-trained knowledge from natural images to achieve high accuracy with limited training data. Traditional methods typically employ vision backbones (e.g., Convolutional Neural Networks (CNNs) or Vision Transformers (ViTs)) pre-trained on ImageNet to extract features from ORSI scenes. However, these backbones exhibit limited generalization across diverse scenarios compared to recent vision foundation models. To this end, we propose ORSI-SAM, a novel ORSI-SOD framework based on the Segment Anything Model (SAM), leveraging its superior generalization capabilities to achieve an exceptional efficiency-accuracy trade-off. Specifically, ORSI-SAM adopts lightweight SAM as the backbone, effectively reducing parameter size and computational overhead to enable efficient deployment on satellite devices while retaining the rich knowledge learned from large-scale natural image datasets. To mitigate the impact of unavailable prompts in ORSI-SOD on the prediction capability of the SAM decoder, we introduce a Hierarchical Interaction Prompt Generator (HIPG), which aggregates hierarchical features and generates mask prompts tailored for salient objects to guide the decoder in producing high-quality saliency maps. Furthermore, to address the recognition challenges caused by the inherent characteristics of ORSIs, we propose a Semantic-Aware Refinement Decoder (SARD). SARD integrates structural details from low-level features to enrich fine-grained object information while leveraging high-level features to suppress redundant interference in shallow layers, thereby improving the detailed information in the predicted saliency map. ORSI-SAM is the first work to explore the accuracy-efficiency trade-offs for ORSI-SOD based on SAM architecture. Extensive experiments on benchmark datasets show that ORSI-SAM achieves superior performance compared to recent state-of-the-art methods with 12.2M parameters and 8.9G FLOPs. Li Liu 0002, Zhuo Su 0002, Tianpeng Liu, Zhen Liu 0004, Matti Pietikäinen |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | ARBiBench: Benchmarking and Analyzing Adversarial Robustness of Binarized Convolutional Neural NetworksabstractBinarized convolutional neural networks (BCNNs), which restrict the weights and activations of the model to +1 or −1, provide notable reductions in memory requirements and enhanced model inference speed during deployment. Current research on BCNNs primarily revolves around addressing the performance degradation resulting from binarization. However, the investigation of the effects of extreme discretization on the robustness of BCNNs has been largely overlooked, despite its critical relevance to real-world applications. To this end, we propose ARBiBench, a comprehensive benchmark for evaluating the adversarial robustness of BCNNs in the image classification task. The key contributions of ARBiBench include: 1) systematically evaluating the robustness of seven influential BCNN methods across various architectures; 2) rigorous validation of diverse adversarial attack methods; and 3) novel empirical findings showing that BCNNs exhibit weaker robustness than full-precision networks on small datasets but surprisingly stronger robustness on large-scale datasets. Leveraging Information Bottleneck theory, we further demonstrate how data scale and model capacity collectively determine BCNNs’ adversarial robustness. These findings not only challenge conventional assumptions about BCNN security, but also provide new insights for developing robust yet efficient neural network architectures. Li Liu 0002, Bowen Peng, Zhen Liu 0004, Longguang Wang, Yingmei Wei |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Refining Pseudo Labeling via Multi-Granularity Confidence Alignment for Unsupervised Cross Domain Object DetectionabstractMost state-of-the-art object detection methods suffer from poor generalization due to the domain shift between training and testing datasets. To resolve this challenge, unsupervised cross domain object detection is proposed to learn an object detector for an unlabeled target domain by transferring knowledge from an annotated source domain. Promising results have been achieved via Mean Teacher, however, pseudo labeling which is the bottleneck of mutual learning remains to be further explored. In this study, we find that confidence misalignment of the predictions, including category-level overconfidence, instance-level task confidence inconsistency, and image-level confidence misfocusing, leading to the injection of noisy pseudo labels in the training process, will bring suboptimal performance. Considering the above issue, we present a novel general framework termed Multi-Granularity Confidence Alignment Mean Teacher (MGCAMT) for unsupervised cross domain object detection, which alleviates confidence misalignment across category-, instance-, and image-levels simultaneously to refine pseudo labeling for better teacher-student learning. Specifically, to align confidence with accuracy at category level, we propose Classification Confidence Alignment (CCA) to model category uncertainty based on Evidential Deep Learning (EDL) and filter out the category incorrect labels via an uncertainty-aware selection strategy. Furthermore, we design Task Confidence Alignment (TCA) to mitigate the instance-level misalignment between classification and localization by enabling each classification feature to adaptively identify the optimal feature for regression. Finally, we develop imagery Focusing Confidence Alignment (FCA) adopting another way of pseudo label learning, i.e., we use the original outputs from the Mean Teacher network for supervised learning without label assignment to achieve a balanced perception of the image's spatial layout. When these three procedures are integrated into a single framework, they mutually benefit to improve the final performance from a cooperative learning perspective. Extensive experiments across multiple scenarios demonstrate that our method outperforms large foundational models, and surpasses other state-of-the-art approaches by a large margin. Jiangming Chen, Li Liu 0002, Wanxia Deng, Zhen Liu 0004, Yu Liu 0012, Yingmei Wei, Yongxiang Liu |
IEEE Trans. Image Process. | 2 |
| 2025 | SARATR-X: Toward Building a Foundation Model for SAR Target RecognitionabstractDespite the remarkable progress in synthetic aperture radar automatic target recognition (SAR ATR), recent efforts have concentrated on detecting and classifying a specific category, e.g., vehicles, ships, airplanes, or buildings. One of the fundamental limitations of the top-performing SAR ATR methods is that the learning paradigm is supervised, task-specific, limited-category, closed-world learning, which depends on massive amounts of accurately annotated samples that are expensively labeled by expert SAR analysts and have limited generalization capability and scalability. In this work, we make the first attempt towards building a foundation model for SAR ATR, termed SARATR-X. SARATR-X learns generalizable representations via self-supervised learning (SSL) and provides a cornerstone for label-efficient model adaptation to generic SAR target detection and classification tasks. Specifically, SARATR-X is trained on 0.18 M unlabelled SAR target samples, which are curated by combining contemporary benchmarks and constitute the largest publicly available dataset till now. Considering the characteristics of SAR images, a backbone tailored for SAR ATR is carefully designed, and a two-step SSL method endowed with multi-scale gradient features was applied to ensure the feature diversity and model scalability of SARATR-X. The capabilities of SARATR-X are evaluated on classification under few-shot and robustness settings and detection across various categories and scenes, and impressive performance is achieved, often competitive with or even superior to prior fully supervised, semi-supervised, or self-supervised algorithms. Our SARATR-X and the curated dataset are released at https://github.com/waterdisappear/SARATR-X to foster research into foundation models for SAR image interpretation. Wei Yang 0046, Yuenan Hou, Li Liu 0002, Yongxiang Liu, Xiang Li 0014 |
IEEE Trans. Image Process. | 4 |
| 2025 | CMoA: Contrastive Mixture of Adapters for Generalized Few-Shot Continual LearningabstractThe goal of Few-Shot Continual Learning (FSCL) is to incrementally learn novel tasks with limited labeled samples and preserve previous capabilities simultaneously. However, current FSCL works lack research on domain increment and domain generalization ability, which cannot cope with changes in the visual perception environment. In this paper, we set up a Generalized FSCL (GFSCL) protocol involving both class- and domain-incremental scenarios together with domain generalization assessment. Firstly, two benchmark datasets and protocols are newly arranged, and detailed baselines are provided for this unexplored configuration. Furthermore, we find that common continual learning methods have poor generalization ability on unseen domains and cannot better tackle catastrophic forgetting issue in cross-incremental tasks. Hence, we propose a rehearsal-free framework based on Vision Transformer (ViT) named Contrastive Mixture of Adapters (CMoA). It contains two non-conflicting parts: (1) By applying the fast-adaptation characteristic of adapter-embedded ViT, the mixture of Adapters (MoA) module is incorporated into ViT. For stability purpose, cosine similarity regularization and dynamic weighting are designed to make each adapter learn specific knowledge and concentrate on particular classes. (2) To further enhance domain generalization ability, we alleviate the intra-class variation by prototype-calibrated contrastive learning to improve domain-invariant representation learning. Finally, six evaluation indicators showing the overall performance and forgetting are compared by comprehensive experiments on two benchmark datasets to validate the efficacy of CMoA, and the results illustrate that CMoA can achieve comparative performance with rehearsal-based continual learning methods. Yawen Cui, Jian Zhao 0006, Zitong Yu, Rizhao Cai, Lei Jin 0003, Alex Chichung Kot, Li Liu 0002, Xuelong Li 0001 |
IEEE Trans. Multim. | 8 |
| 2025 | Boosting Convolutional Neural Networks With Middle Spectrum Grouped ConvolutionabstractThis article proposes a novel module called middle spectrum grouped convolution (MSGC) for efficient deep convolutional neural networks (DCNNs) with the mechanism of grouped convolution. It explores the broad "middle spectrum" area between channel pruning and conventional grouped convolution. Compared with channel pruning, MSGC can retain most of the information from the input feature maps due to the group mechanism; compared with grouped convolution, MSGC benefits from the learnability, the core of channel pruning, for constructing its group topology, leading to better channel division. The middle spectrum area is unfolded along four dimensions: groupwise, layerwise, samplewise, and attentionwise, making it possible to reveal more powerful and interpretable structures. As a result, the proposed module acts as a booster that can reduce the computational cost of the host backbones for general image recognition with even improved predictive accuracy. For example, in the experiments on the ImageNet dataset for image classification, MSGC can reduce the multiply-accumulates (MACs) of ResNet-18 and ResNet-50 by half but still increase the Top-1 accuracy by more than 1%. With a 35% reduction of MACs, MSGC can also increase the Top-1 accuracy of the MobileNetV2 backbone. Results on the MS COCO dataset for object detection show similar observations. Our code and trained models are available at https://github.com/hellozhuo/msgc. Zhuo Su 0002, Tianpeng Liu, Zhen Liu 0004, Shuanghui Zhang, Matti Pietikäinen, Li Liu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2025 | Locality Preservation for Unsupervised Multimodal Change Detection in Remote Sensing ImageryabstractMultimodal change detection (MCD) is a topic of increasing interest in remote sensing. Due to different imaging mechanisms, the multimodal images cannot be directly compared to detect the changes. In this article, we explore the topological structure of multimodal images and construct the links between class relationships (same/different) and change labels (changed/unchanged) of pairwise superpixels, which are imaging modality-invariant. With these links, we formulate the MCD problem within a mathematical framework termed the locality-preserving energy model (LPEM), which is used to maintain the local consistency constraints embedded in the links: the structure consistency based on feature similarity and the label consistency based on spatial continuity. Because the foundation of LPEM, i.e., the links, is intuitively explainable and universal, the proposed method is very robust across different MCD situations. Noteworthy, LPEM is built directly on the label of each superpixel, so it is a paradigm that outputs the change map (CM) directly without the need to generate intermediate difference image (DI) as most previous algorithms have done. Experiments on different real datasets demonstrate the effectiveness of the proposed method. Source code of the proposed method is made available at https://github.com/yulisun/LPEM. Yuli Sun, Lin Lei, Dongdong Guan, Gangyao Kuang, Li Liu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | S3INet: Semantic-Information Space Sharing Interaction Network for Arbitrary Shape Text DetectionabstractThe detecting arbitrary shape text is a challenging task due to the significant variation in text shape, size, and aspect ratio, as well as the complexity of scene backgrounds. The enhancing feature extraction capabilities is essential for the boosting text detection accuracy. However, traditional text feature extraction methods face several issues, including insufficient multiscale feature fusion, limited information transfer between different feature levels, and constrained receptive field expansion when using asymmetric convolutional kernels for long text detection. To address these challenges, this article introduces an arbitrarily shaped scene text detector called the semantic-information space sharing interaction network (S3INet). The proposed network leverages the semantic-information space sharing module (S3M) to generate a single-level feature map capable of capturing multiscale features with rich semantic information and prominent foreground elements. In addition, we propose the multibranch parallel asymmetric convolutional module (MPACM) group to enhance the representation of text features, thereby further enhancing text detection performance. Extensive experimental evaluations on five publicly available natural scene text datasets (CTW-1500, Total-Text, MSRA-TD500, ICDAR2015, and ICDAR2017-MLT) and two traffic text datasets (CTST-1600 and TPD) demonstrate the superiority of our method. The results indicate that S3INet significantly outperforms most existing state-of-the-art methods in both accuracy and robustness. The code will be released at: https://github.com/runminwang/S3INet. Yanbin Zhu, Xiaofei Cao, Zhenlin Zhu, Shengyou Qian, Changxin Gao, Li Liu 0002, Nong Sang |
IEEE Trans. Neural Networks Learn. Syst. | 9 |
| 2025 | A Forward and Backward Compatible Framework for Few-Shot Class-Incremental Pill RecognitionabstractAutomatic pill recognition (APR) systems are crucial for enhancing hospital efficiency, assisting visually impaired individuals, and preventing cross-infection. However, most existing deep learning-based pill recognition systems can only perform classification on classes with sufficient training data. In practice, the high cost of data annotation and the continuous increase in new pill classes necessitate the development of a few-shot class-incremental pill recognition (FSCIPR) system. This article introduces the first FSCIPR framework, discriminative and bidirectional compatible few-shot class-incremental learning (DBC-FSCIL). It encompasses forward-compatible and backward-compatible learning components. In forward-compatible learning, we propose an innovative virtual class generation strategy and a center-triplet (CT) loss to enhance discriminative feature learning. These virtual classes serve as placeholders in the feature space for future class updates, providing diverse semantic knowledge for model training. For backward-compatible learning, we develop a strategy to synthesize reliable pseudo-features of old classes using uncertainty quantification, facilitating data replay (DR) and knowledge distillation (KD). This approach allows for the flexible synthesis of features and effectively reduces additional storage requirements for samples and models. Additionally, we construct a new pill image dataset for FSCIL and assess various mainstream FSCIL methods, establishing new benchmarks. Our experimental results demonstrate that our framework surpasses existing state-of-the-art (SOTA) methods. Li Liu 0002, Kai Gao 0011, Dewen Hu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Optimal Indexing: An Efficient Feature-Based Indexing Framework for Similarity Data Sharing at the Network EdgeabstractEdge storage systems have drawn many efforts to extend the storage and service capabilities of cloud data centers. A pivotal aspect lies in the data-sharing mechanism, which integrates geographically dispersed weak edge servers into an efficient storage system. It enables users to launch data operations at any server and retrieve the desired data across the distributed system. However, it remains open to meeting the increasing demand for similarity retrieval across edge servers. The intrinsic reason is that the existing solutions can only return an exact data match for a query while more general edge applications require the data similar to a query input from any server. To fill this gap, this paper pioneers the similarity edge data sharing mechanism, a new paradigm to support high-dimensional similarity search at network edges. First, through deeply thinking about the nature of similarity data sharing, we propose the problem of Optimal Indexing and formulate it as the optimal transport problem from the data space to the network space. On this basis, we propose Prophet, the first known architecture for similarity data indexing at the edge. We first divide the feature space of data into plenty of subareas, then project both subareas and edge servers into a virtual space where the distance between any two points can reflect not only data similarity but also network latency. When any edge server submits a request for data insert, delete, or query, it computes the data feature and the virtual coordinate; and then iteratively forwards the request via greedy routing based on the forwarding tables and the virtual coordinates. By Prophet, similar high-dimensional features would be stored by a common server or several nearby servers. Compared with distributed hash tables in P2P networks, Prophet requires to visit logarithmic servers for a data request and reduces the network latency from the logarithmic to the constant level of the server number. Evaluation results indicate that Prophet achieves the comparable retrieval accuracy and significantly shortens the query latency compared with centralized schemes, while the load balancing performance is nearly optimal. Yuchen Sun 0001, Lailong Luo, Deke Guo, Li Liu 0002, Bangbang Ren |
IEEE Trans. Netw. | 4 |
| 2024 | Hide in Thicket: Generating Imperceptible and Rational Adversarial Perturbations on 3D Point CloudsabstractAdversarial attack methods based on point manipulation for 3D point cloud classification have revealed the fragility of 3D models, yet the adversarial examples they produce are easily perceived or defended against. The tradeoff between the imperceptibility and adversarial strength leads most point attack methods to inevitably introduce easily detectable outlier points upon a successful attack. An-other promising strategy, shape-based attack, can effectively eliminate outliers, but existing methods often suffer significant reductions in imperceptibility due to irrational deformations. We find that concealing deformation perturbations in areas insensitive to human eyes can achieve a better tradeoff between imperceptibility and adversarial strength, specifically in parts of the object surface that are complex and exhibit drastic curvature changes. Therefore, we propose a novel shape-based adversarial attack method, HiT-ADV, which initially conducts a two-stage search for attack regions based on saliency and imperceptibility scores, and then adds deformation perturbations in each attack region using Gaussian kernel functions. Additionally, HiT-ADV is extendable to physical attack. We propose that by employing benign resampling and benign rigid transformations, we can further enhance physical adversarial strength with little sacrifice to imperceptibility. Extensive experiments have validated the superiority of our method in terms of adversarial and imperceptible properties in both digital and physical spaces. Our code is avaliable at: https://github.com/TRLou/HiT-ADV. Tianrui Lou, Xiaojun Jia, Jindong Gu, Li Liu 0002, Siyuan Liang 0004, Bangyan He, Xiaochun Cao |
CVPR | 4 |
| 2024 | DFormer: Rethinking RGBD Representation Learning for Semantic SegmentationabstractWe present DFormer, a novel RGB-D pretraining framework to learn transferable representations for RGB-D segmentation tasks. DFormer has two new key innovations: 1) Unlike previous works that encode RGB-D information with RGB pretrained backbone, we pretrain the backbone using image-depth pairs from ImageNet-1K, and thus the DFormer is endowed with the capacity to encode RGB-D representations; 2) DFormer comprises a sequence of RGB-D blocks, which are tailored for encoding both RGB and depth information through a novel building block design. DFormer avoids the mismatched encoding of the 3D geometry relationships in depth maps by RGB pretrained backbones, which widely lies in existing methods but has not been resolved. We finetune the pretrained DFormer on two popular RGB-D tasks, i.e., RGB-D semantic segmentation and RGB-D salient object detection, with a lightweight decoder head. Experimental results show that our DFormer achieves new state-of-the-art performance on these two tasks with less than half of the computational cost of the current best methods on two RGB-D semantic segmentation datasets and five RGB-D salient object detection datasets. Code will be made publicly available. Bowen Yin, Xuying Zhang, Zhongyu Li 0006, Li Liu 0002, Ming-Ming Cheng, Qibin Hou |
ICLR | 4 |
| 2024 | Improved (Related-Key) Differential-Based Neural Distinguishers for SIMON and SIMECK Block CiphersabstractAbstract In CRYPTO 2019, Gohr made a pioneering attempt and successfully applied deep learning to the differential cryptanalysis against NSA block cipher Speck 32/64, achieving higher accuracy than the pure differential distinguishers. By its very nature, mining effective features in data plays a crucial role in data-driven deep learning. In this paper, in addition to considering the integrity of the information from the training data of the ciphertext pair, domain knowledge about the structure of differential cryptanalysis is also considered into the training process of deep learning to improve the performance. Meanwhile, taking the performance of the differential-neural distinguisher of Simon 32/64 as an entry point, we investigate the impact of input difference on the performance of the hybrid distinguishers to choose the proper input difference. Eventually, we improve the accuracy of the neural distinguishers of Simon 32/64, Simon 64/128, Simeck 32/64 and Simeck 64/128. We also obtain related-key differential-based neural distinguishers on round-reduced versions of Simon 32/64, Simon 64/128, Simeck 32/64 and Simeck 64/128 for the first time. Jinyu Lu, Bing Sun 0001, Chao Li 0002, Li Liu 0002 |
Comput. J. | 5 |
| 2024 | SplatFlow: Learning Multi-frame Optical Flow via Splatting
Bo Wang 0144, Jian Li 0003, Yang Yu 0014, Zhenping Sun, Li Liu 0002, Dewen Hu |
Int. J. Comput. Vis. | 6 |
| 2024 | An Integrated Network for SA-ISAR Image Processing With Adaptive Denoising and Super-Resolution ModulesabstractThis letter focuses on developing an effective and generalizable deep learning approach for inverse synthetic aperture radar (ISAR) image super-resolution (SR). Since the ISAR imaging process is typically carried out under sparse aperture (SA) conditions, imaging results may exhibit striped noise caused by echoes missing, making it challenging to apply conventional SR methods directly. In view of this, we present a blind SR (BSR) method specifically designed for ISAR images with striped noise. The proposed method employs an integrated network that includes an adaptive denoising module and a SR module (AD-SRNet). Experimental results on both synthetic and real ISAR samples demonstrate the superior performance and strong generalization capability of our approach. Mingyao Chen, Jingyuan Xia, Tianpeng Liu, Li Liu 0002 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | An Autonomous Feature Detection Method of Slow Small Targets on Sea Surface Based on Contextual BanditabstractUnder the background of slow small target detection on sea surface, it is a serious problem that the detection performance of existing feature detection methods decreases when the number of coherent pulses is less. In this letter, we first model the slow small target detection on sea surface as a contextual bandit problem. On this basis, we propose an autonomous feature detection method by modifying the classical feature detection process. The method can autonomously choose the detectors with better performance under current sea scene from the constructed 18 feature detectors, and obtain fine detection performance by fusing their detection results. The performance superiority and the real-time capability of proposed method are verified by the experiments on 7 CSIR datasets. Xijie Wu, Tianpeng Liu, Yongxiang Liu, Li Liu 0002 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Data Distribution Loss for Imbalanced SAR Vehicle Target RecognitionabstractThe data distribution of synthetic aperture radar (SAR) vehicle targets in the actual missions is often imbalanced. However, the recent algorithms for SAR target recognition are designed either under abundant samples, or the situation of few labeled samples among all the categories. These cases all avoid facing the difficulties of imbalanced data distribution,i.e. the difference between the number of labeled samples among categories is huge. The samples in the majority classes will get more chances to be learnt by the deep neural network, which impedes the regular algorithms from achieving a high recognition rate. In this letter, a design guideline for imbalance loss and an example of data distribution (DD) loss based on the guideline is proposed, which provides an extremely effective way of handling the problem of imbalanced SAR target recognition. The DD loss takes the sample distribution and the data quantity of SAR vehicle targets into consideration. It can cause images with fewer samples in their categories to decrease more gradients proportionally. Moreover, the proposed DD loss adds no more burden to the networks and compared to other imbalanced algorithms with complex processes, the DD loss can be conducted easily. Plenty of experiments, which involve two various kinds of imbalanced datasets, are implemented and the proposed DD loss shows excellent performance among these imbalanced datasets. When there are only 40 labeled samples in minority categories, the DD loss can achieve over 95% in nine different cases, which exceeds other methods and losses of at least 7%. Linbin Zhang, Xiangguang Leng, Xiaojie Ma, Kefeng Ji, Gangyao Kuang, Li Liu 0002 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2024 | Conditional Random Field-Based Adversarial Attack Against SAR Target DetectionabstractThe existence of adversarial examples causes serious security risks when deep neural networks are applied to synthetic aperture radar (SAR) target detection. In SAR image processing, the added small disturbances can cause the model to output incorrect predictions. Due to the multipath effect in the propagation of detection signals, there are complex interactions between targets and their surroundings serving as supportive clues for target detection. The interactions are manifested as tight correlations between pixels and contextual information in the SAR image (where context refers to various relationships, e.g., target-to-target co-occurrence relationships). In this letter, we proposed a novel conditional random field-based adversarial attack (CRFA) method, which disturbs the intrinsic interactions between the target and its surroundings. To the best of our knowledge, we are the first to exploit the contextual information for attacking the SAR target detector. We formulate the attack as an optimization problem and design the context information loss to calculate the energy differences in local feature patterns before and after perturbation. By maximizing the energy differences, the context area information around the target is destroyed, and the detector outputs the candidate box with a slight shift, even ignoring the ground truth and missing targets. Extensive experimental results on the SAR Ship Detection dataset (SSDD) demonstrate that our proposed algorithm reduces mAP by 4.29% on existing object detection models, validating the effectiveness of the method. Jie Zhou 0031, Jianyue Xie, Bowen Peng, Li Liu 0002, Xiang Li 0014 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | Mitigating SAR Out-of-Distribution Overconfidence Based on Evidential UncertaintyabstractSynthetic aperture radar (SAR) automatic target recognition (ATR) is extensively applied in both military and civilian sectors. Nevertheless, test and training data distribution may differ in the open world. Therefore, SAR out-of-distribution (OOD) detection is important because it enhances the reliability and adaptability of SAR systems. However, most OOD detection models are based on maximum likelihood estimation (MLE) and overlook the impact of data uncertainty, leading to overconfidence output for both in-distribution (ID) and OOD data. To address this issue, we consider the effect of data uncertainty on prediction probabilities, treating these probabilities as random variables and modeling them using Dirichlet distribution. Building on this, we propose an evidential uncertainty aware mean squared error (UMSE) loss function to guide the model in learning highly distinguishable output between ID and OOD data. Furthermore, to comprehensively evaluate OOD detection performance, we have compiled and organized some publicly available data and constructed a new SAR OOD detection dataset named SAR-OOD. Experimental results on SAR-OOD demonstrate that the UMSE approach achieves state-of-the-art (SOTA) performance. The code and data are available at:https://github.com/Xiaoyan-Zhou/UMSE-SAR-OOD-Detection. Tao Tang 0006, Zhongzhen Sun, Gangyao Kuang, Janne Heikkilä, Li Liu 0002 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2024 | DiffDet4SAR: Diffusion-Based Aircraft Target Detection Network for SAR ImagesabstractAircraft target detection in SAR images is a challenging task due to the discrete scattering points and severe background clutter interference. Currently, methods with convolution-based or transformer-based paradigms cannot adequately address these issues. In this letter, we explore diffusion models for SAR image aircraft target detection for the first time and propose a novel Diffusion-based aircraft target Detection network for SAR images (DiffDet4SAR). Specifically, the proposed DiffDet4SAR yields two main advantages for SAR aircraft target detection: 1) DiffDet4SAR maps the SAR aircraft target detection task to a denoising diffusion process of bounding boxes without heuristic anchor size selection, effectively enabling large variations in aircraft sizes to be accommodated; and 2) the dedicatedly designed Scattering Feature Enhancement (SFE) module further reduces the clutter intensity and enhances the target saliency during inference. Extensive experimental results on the SAR-AIRcraft-1.0 dataset show that the proposed DiffDet4SAR achieves 88.4% mAP50, outperforming the state-of-the-art methods by 6%. Code is availabel at https://github.com/JoyeZLearning/DiffDet4SAR. Jie Zhou 0031, Zhen Liu 0004, Li Liu 0002, Yongxiang Liu, Xiang Li 0014 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | Meta-learning based blind image super-resolution approach to different degradations
Zhixiong Yang 0001, Jingyuan Xia, Shengxi Li, Wende Liu, Shuaifeng Zhi, Shuanghui Zhang, Li Liu 0002, Yaowen Fu, Deniz Gündüz |
Neural Networks | 7 |
| 2024 | Editorial: Learning With Fewer Labels in Computer VisionabstractUndoubtedly, Deep Neural Networks (DNNs), from AlexNet to ResNet to Transformer, have sparked revolutionary advancements in diverse computer vision tasks. The scale of DNNs has grown exponentially due to the rapid development of computational resources. Despite the tremendous success, DNNs typically depend on massive amounts of training data (especially the recent various foundation models) to achieve high performance and are brittle in that their performance can degrade severely with small changes in their operating environment. Generally, collecting massive-scale training datasets is costly or even infeasible, as for certain fields, only very limited or no examples at all can be gathered. Nevertheless, collecting, labeling, and vetting massive amounts of practical training data is certainly difficult and expensive, as it requires the painstaking efforts of experienced human annotators or experts, and in many cases, prohibitively costly or impossible due to some reason, such as privacy, safety or ethic issues. Li Liu 0002, Timothy M. Hospedales, Yann LeCun, Mingsheng Long, Jiebo Luo 0001, Wanli Ouyang, Matti Pietikäinen, Tinne Tuytelaars |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Revisiting Computer-Aided Tuberculosis DiagnosisabstractTuberculosis (TB) is a major global health threat, causing millions of deaths annually. Although early diagnosis and treatment can greatly improve the chances of survival, it remains a major challenge, especially in developing countries. Recently, computer-aided tuberculosis diagnosis (CTD) using deep learning has shown promise, but progress is hindered by limited training data. To address this, we establish a large-scale dataset, namely the Tuberculosis X-ray (TBX11 K) dataset, which contains 11 200 chest X-ray (CXR) images with corresponding bounding box annotations for TB areas. This dataset enables the training of sophisticated detectors for high-quality CTD. Furthermore, we propose a strong baseline, SymFormer, for simultaneous CXR image classification and TB infection area detection. SymFormer incorporates Symmetric Search Attention (SymAttention) to tackle the bilateral symmetry property of CXR images for learning discriminative features. Since CXR images may not strictly adhere to the bilateral symmetry property, we also propose Symmetric Positional Encoding (SPE) to facilitate SymAttention through feature recalibration. To promote future research on CTD, we build a benchmark by introducing evaluation metrics, evaluating baseline models reformed from existing detectors, and running an online challenge. Experiments show that SymFormer achieves state-of-the-art performance on the TBX11 K dataset. Yun Liu 0011, Yu-Huan Wu, Shi-Chen Zhang, Li Liu 0002, Min Wu 0008, Ming-Ming Cheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Deep Learning for Visual Speech Analysis: A SurveyabstractVisual speech, referring to the visual domain of speech, has attracted increasing attention due to its wide applications, such as public security, medical treatment, military defense, and film entertainment. As a powerful AI strategy, deep learning techniques have extensively promoted the development of visual speech learning. Over the past five years, numerous deep learning based methods have been proposed to address various problems in this area, especially automatic visual speech recognition and generation. To push forward future research on visual speech, this paper will present a comprehensive review of recent progress in deep learning methods on visual speech analysis. We cover different aspects of visual speech, including fundamental problems, challenges, benchmark datasets, a taxonomy of existing methods, and state-of-the-art performance. Besides, we also identify gaps in current research and discuss inspiring future research directions. Changchong Sheng, Gangyao Kuang, Liang Bai 0003, Chenping Hou, Yulan Guo, Xin Xu 0001, Matti Pietikäinen, Li Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2024 | Highly Efficient and Unsupervised Framework for Moving Object Detection in Satellite VideosabstractMoving object detection in satellite videos (SVMOD) is a challenging task due to the extremely dim and small target characteristics. Current learning-based methods extract spatio-temporal information from multi-frame dense representation with labor-intensive manual labels to tackle SVMOD, which needs high annotation costs and contains tremendous computational redundancy due to the severe imbalance between foreground and background regions. In this paper, we propose a highly efficient unsupervised framework for SVMOD. Specifically, we propose a generic unsupervised framework for SVMOD, in which pseudo labels generated by a traditional method can evolve with the training process to promote detection performance. Furthermore, we propose a highly efficient and effective sparse convolutional anchor-free detection network by sampling the dense multi-frame image form into a sparse spatio-temporal point cloud representation and skipping the redundant computation on background regions. Coping these two designs, we can achieve both high efficiency (label and computation efficiency) and effectiveness. Extensive experiments demonstrate that our method can not only process 98.8 frames per second on 1024 ×1024 images but also achieve state-of-the-art performance. Wei An 0003, Yifan Zhang 0030, Zhuo Su 0002, Weidong Sheng, Matti Pietikäinen, Li Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2024 | An Adaptive Post-Processing Network With the Global-Local Aggregation for Semantic SegmentationabstractCurrent semantic segmentation methods mainly focus on modeling the context of the global image to obtain high-quality segmentation results. However, they ignore the role of local image patches, which contain complementary and effective context information. In this paper, we propose an adaptive post-processing network (APPNet) for semantic segmentation based on the predictions of current methods in the global image and local image patches. The key point of APPNet is the global-local aggregation module, which models the context between global predictions and local predictions to generate accurate pixel-wise representation. Furthermore, we develop an adaptive points replacement module to compensate for the lack of fine detail in global prediction and the overconfidence in local predictions. Our method can be readily integrated into existing segmentation methods (i.e., ConvNeXt, HRNet, ViT-Adapter) with little memory and without extra modification in current models. We empirically demonstrate our method brings performance improvements across diverse datasets (i.e., Cityscapes, ADE20K, PASCAL-Context, COCO-Stuff). The code and models will be publicly available athttps://github.com/zhu-gl-ux/APPN. Guilin Zhu, Zhenlin Zhu, Changxin Gao, Li Liu 0002, Nong Sang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Audio-Visual Kinship Verification: A New Dataset and a Unified Adaptive Adversarial Multimodal Learning ApproachabstractFacial kinship verification refers to automatically determining whether two people have a kin relation from their faces. It has become a popular research topic due to potential practical applications. Over the past decade, many efforts have been devoted to improving the verification performance from human faces only while lacking other biometric information, for example, speaking voice. In this article, to interpret and benefit from multiple modalities, we propose for the first time to combine human faces and voices to verify kinship, which we refer it as the audio-visual kinship verification study. We first establish a comprehensive audio-visual kinship dataset that consists of familial talking facial videos under various scenarios, called TALKIN-Family. Based on the dataset, we present the extensive evaluation of kinship verification from faces and voices. In particular, we propose a deep-learning-based fusion method, called unified adaptive adversarial multimodal learning (UAAML). It consists of the adversarial network and the attention module on the basis of unified multimodal features. Experiments show that audio (voice) information is complementary to facial features and useful for the kinship verification problem. Furthermore, the proposed fusion method outperforms baseline methods. In addition, we also evaluate the human verification ability on a subset of TALKIN-Family. It indicates that humans have higher accuracy when they have access to both faces and voices. The machine-learning methods could effectively and efficiently outperform the human ability. Finally, we include the future work and research opportunities with the TALKIN-Family dataset. Xiaoting Wu, Xueyi Zhang 0001, Xiaoyi Feng, Miguel Bordallo López, Li Liu 0002 |
IEEE Trans. Cybern. | 5 |
| 2024 | From Coarse to Fine: ISAR Object View Interpolation via Flow Estimation and GANabstractThis article focuses on the multiazimuth angle interpolation task of inverse synthetic aperture radar (ISAR) images for aircraft targets and complements incomplete ISAR image datasets. ISAR image automatic target recognition (ATR) has been widely applied in remote sensing and many fields. However, the imaging process is more challenging when compared to capturing optical and SAR image data, which reduces the accuracy and generalization performance of the ATR system. Therefore, in this article, we leverage existing limited ISAR data to achieve autonomous data expansion. This approach helps mitigate the impact of low sample quantity and unbalanced distribution, ultimately improving the accuracy of the ATR system for target recognition. Most existing methods use generative networks for ISAR image expansion, but few focus on generating ISAR images with specific azimuth angles. This article proposes a novel two-stage coarse-to-fine framework for ISAR object view interpolation (C2FIPNet) that combines flow estimation and GAN to interpolate ISAR images with intermediate azimuth angles using a set of ISAR image pairs. Flow estimation is employed for coarse-grained generation, determining the position and intensity of strong scattering points in the ISAR image. The GAN, on the other hand, is used for fine-grained completion to correct image distortion caused by flow estimation and enhance image details. In addition, a suitable loss function is designed, incorporating both global and local features, allowing for priority generation in the region of strong scattering points. In conclusion, extensive simulation and comparative experiments have demonstrated that the interpolated ISAR images generated by the proposed C2FIPNet exhibit greater pixel-level authenticity. Zhen Liu 0004, Weidong Jiang, Yongxiang Liu, Shuowei Liu, Li Liu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Azimuth-Aware Subspace Classifier for Few-Shot Class-Incremental SAR ATRabstractWith the rapid acquisition of high-resolution Synthetic Aperture Radar(SAR) images, new categories are continually observed with few-shot instances in openly non-cooperative scenarios. Powering a SAR Automatic Target Recognition (SAR ATR) system with an ability of few-shot class-incremental learning (FSCIL) is nontrivial. Observing the pronounced azimuth-dependence and part-sparsity of targets in SAR images, an Azimuth-aware Subspace Classifier (AASC) on the Grassmannian manifold is proposed to tackle the FSCIL of SAR ATR stably and accurately. In the AASC, losses covering both semantic and manifold facets, which include Semantic Margin Separation (SMS), Deep Subspace Separation (DSS), and Structure Less Forgetting (SLF), are designed to strike both the intrinsic model’s stability and plasticity dilemma and domain-specific challenges. For plasticity, the novel-to-old semantic margins are enlarged by the SMS loss for knowledge transferring while avoiding inappropriate adaptions. The DSS loss derived from the Grassmannian geometry aims to regularize class subspaces orthogonality. For stability, semantic drifts of target spatial and global structures are punished by the SLF loss. As the periodicity and volatility of target azimuth-aware patterns, an Azimuth-aware Exemplar Selection (AES) strategy is designed to select representative and complementary exemplars. In experiments, the advantages of the subspace classifier and the designed losses and strategies are deeply verified. Comprehensive experiments on three FSCIL scenarios derived from both airborne and spaceborne datasets, including the MSTAR, the SAR-AIRcraft-1.0, and self-collected data sets, show that our method significantly outperforms various task-specific benchmarks, verifying its effectiveness for the FSCIL in real SAR ATR scenarios. Yan Zhao 0026, Lingjun Zhao, Siqian Zhang, Kefeng Ji, Gangyao Kuang, Li Liu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | PMSA-DyTr: Prior-Modulated and Semantic-Aligned Dynamic Transformer for Strip Steel Defect DetectionabstractIn-process hot-rolled strip steel is suffering from some complicated yet unavoidable surface defects due to its harsh production environment. The automated visual inspection on defects consistently faces challenges of interclass similarity, intraclass difference, low contrast, and overlapping issue, which tend to trigger false or missed detections. This article proposes a prior-modulated and semantic-aligned dynamic transformer, called PMSA-DyTr. In this framework, a long short-term self-attention embedded with local convolution is designed for assisting an encoder to eliminate noise ambiguity between defects and backgrounds. Then, a semantic aligner is cleverly bridged between the encoder and the decoder to align the sematic for speeding up the convergence, and prior-modulated cross attention is proposed to alleviate the deficiency of samples for a data-driven transformer. Furthermore, a gate controller is innovatively constructed to dynamically select the minimal number of encoder blocks while preserving detection accuracy. The proposed PMSA-DyTr outperforms 19 state-of-the-art models on mean average precision with an inference time of 54.67 ms and visually performs best in detecting low-contrast and multiple small defects. Jiaojiao Su, Qiwu Luo, Chunhua Yang 0001, Weihua Gui 0001, Olli Silvén, Li Liu 0002 |
IEEE Trans. Ind. Informatics | 6 |
| 2024 | Enhancing Information Maximization With Distance-Aware Contrastive Learning for Source-Free Cross-Domain Few-Shot LearningabstractExisting Cross-Domain Few-Shot Learning (CDFSL) methods require access to source domain data to train a model in the pre-training phase. However, due to increasing concerns about data privacy and the desire to reduce data transmission and training costs, it is necessary to develop a CDFSL solution without accessing source data. For this reason, this paper explores a Source-Free CDFSL (SF-CDFSL) problem, in which CDFSL is addressed through the use of existing pretrained models instead of training a model with source data, avoiding accessing source data. However, due to the lack of source data, we face two key challenges: effectively tackling CDFSL with limited labeled target samples, and the impossibility of addressing domain disparities by aligning source and target domain distributions. This paper proposes an Enhanced Information Maximization with Distance-Aware Contrastive Learning (IM-DCL) method to address these challenges. Firstly, we introduce the transductive mechanism for learning the query set. Secondly, information maximization (IM) is explored to map target samples into both individual certainty and global diversity predictions, helping the source model better fit the target data distribution. However, IM fails to learn the decision boundary of the target task. This motivates us to introduce a novel approach called Distance-Aware Contrastive Learning (DCL), in which we consider the entire feature set as both positive and negative sets, akin to Schrödinger's concept of a dual state. Instead of a rigid separation between positive and negative sets, we employ a weighted distance calculation among features to establish a soft classification of the positive and negative sets for the entire feature set. We explore three types of negative weights to enhance the performance of CDFSL. Furthermore, we address issues related to IM by incorporating contrastive constraints between object features and their corresponding positive and negative sets. Evaluations of the 4 datasets in the BSCD-FSL benchmark indicate that the proposed IM-DCL, without accessing the source domain, demonstrates superiority over existing methods, especially in the distant domain task. Additionally, the ablation study and performance analysis confirmed the ability of IM-DCL to handle SF-CDFSL. The code will be made public at https://github.com/xuhuali-mxj/IM-DCL. Huali Xu, Li Liu 0002, Shuaifeng Zhi, Shaojing Fu, Zhuo Su 0002, Ming-Ming Cheng, Yongxiang Liu |
IEEE Trans. Image Process. | 2 |
| 2024 | TTDNet: An End-to-End Traffic Text Detection Framework for Open Driving EnvironmentsabstractTraffic text detection is crucial for traffic scene understanding in intelligent transportation systems (ITS). Although natural scene text detection has been extensively studied, yielding noteworthy results, little research has focused on traffic text detection. Traffic text, as a special type of natural scene text, faces not only the general challenges of natural scene text detection but also the significant impact of false alarms from non-traffic texts on system performance. In light of these challenges, we propose an end-to-end traffic text detection framework that can effectively detect traffic text captured by in-vehicle cameras in various driving scenarios. The key contributions of our proposed approach are: (1) an Image Enhancement Module designed to remove fog and enhance low-quality images; (2) a plug-and-play Text Feature Enhancement Module; (3) a Joint Loss Function; and (4) the creation of an open driving environment traffic text dataset (named ODETT-3000) containing various traffic environments. Comprehensive experimental studies have verified that our method achieves state-of-the-art performance on traffic text datasets such as CTST-1600, TPD, and ODETT-3000, as well as promising results on public natural scene text datasets such as MSRA-TD500, ICDAR 2015, and CTW 1500, thereby showcasing the superior performance and adaptability of our method. The code and our dataset (ODETT-3000) will be available athttps://github.com/yanbin-zhu/TTDNet. Yanbin Zhu, Zhenlin Zhu, Yajun Ding, Shengyou Qian, Changxin Gao, Li Liu 0002, Nong Sang |
IEEE Trans. Intell. Transp. Syst. | 9 |
| 2024 | Rethinking Few-Shot Class-Incremental Learning With Open-Set Hypothesis in Hyperbolic GeometryabstractBy training first with a large base dataset, Few-Shot Class-Incremental Learning (FSCIL) aims at continually learning a sequence of few-shot learning tasks with novel classes. There are mainly two challenges in FSCIL: the overfitting issue of novel classes with limited labeled samples and the catastrophic forgetting of previously seen classes. The current protocol of FSCIL is built by mimicking the general class-incremental learning setting by building a unified framework, while the existing frameworks for FSCIL on this protocol always bias to the classes in the base dataset because the dominant performance of the deep model is decided by the size of the training dataset. Moreover, it is difficult to handle the stability-plasticity constraint in a unified FSCIL framework. To solve these issues, we rethink the configuration of FSCIL with the open-set hypothesis by reserving the possibility in the first session for incoming categories. To find a better decision boundary of close space and open space, Hyperbolic Reciprocal Point Learning module (Hyper-RPL) is built on Reciprocal Point Learning with hyperbolic neural networks. Besides, when learning novel categories from limited labeled data, we incorporate a hyperbolic metric learning (Hyper-Metric) module into the distillation-based framework to alleviate the overfitting issue and better handle the trade-off issue between the preservation of old knowledge and the acquisition of new knowledge. Finally, the comprehensive assessments of the proposed configuration and modules on three benchmark datasets are executed to validate the effectiveness, and state-of-the-art results are achieved. Yawen Cui, Zitong Yu, Wei Peng 0009, Qi Tian 0001, Li Liu 0002 |
IEEE Trans. Multim. | 5 |
| 2024 | Uncertainty-Aware Distillation for Semi-Supervised Few-Shot Class-Incremental LearningabstractGiven a model well-trained with a large-scale base dataset, few-shot class-incremental learning (FSCIL) aims at incrementally learning novel classes from a few labeled samples by avoiding overfitting, without catastrophically forgetting all encountered classes previously. Currently, semi-supervised learning technique that harnesses freely available unlabeled data to compensate for limited labeled data can boost the performance in numerous vision tasks, which heuristically can be applied to tackle issues in FSCIL, i.e., the semi-supervised FSCIL (Semi-FSCIL). So far, very limited work focuses on the Semi-FSCIL task, leaving the adaptability issue of semi-supervised learning to the FSCIL task unresolved. In this article, we focus on this adaptability issue and present a simple yet efficient Semi-FSCIL framework named uncertainty-aware distillation with class-equilibrium (UaD-ClE), encompassing two modules: uncertainty-aware distillation (UaD) and class equilibrium (ClE). Specifically, when incorporating unlabeled data into each incremental session, we introduce the ClE module that employs a class-balanced self-training (CB_ST) to avoid the gradual dominance of easy-to-classified classes on pseudo-label generation. To distill reliable knowledge from the reference model, we further implement the UaD module that combines uncertainty-guided knowledge refinement with adaptive distillation. Comprehensive experiments on three benchmark datasets demonstrate that our method can boost the adaptability of unlabeled data with the semi-supervised learning technique in FSCIL tasks. The code is available at https://github.com/yawencui/UaD-ClE. Yawen Cui, Wanxia Deng, Haoyu Chen 0001, Li Liu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Image Regression With Structure Cycle Consistency for Heterogeneous Change DetectionabstractChange detection (CD) between heterogeneous images is an increasingly interesting topic in remote sensing. The different imaging mechanisms lead to the failure of homogeneous CD methods on heterogeneous images. To address this challenge, we propose a structure cycle consistency-based image regression method, which consists of two components: the exploration of structure representation and the structure-based regression. We first construct a similarity relationship-based graph to capture the structure information of image; here, a k -selection strategy and an adaptive-weighted distance metric are employed to connect each node with its truly similar neighbors. Then, we conduct the structure-based regression with this adaptively learned graph. More specifically, we transform one image to the domain of the other image via the structure cycle consistency, which yields three types of constraints: forward transformation term, cycle transformation term, and sparse regularization term. Noteworthy, it is not a traditional pixel value-based image regression, but an image structure regression, i.e., it requires the transformed image to have the same structure as the original image. Finally, change extraction can be achieved accurately by directly comparing the transformed and original images. Experiments conducted on different real datasets show the excellent performance of the proposed method. The source code of the proposed method will be made available at https://github.com/yulisun/AGSCC. Yuli Sun, Lin Lei, Dongdong Guan, Junzheng Wu, Gangyao Kuang, Li Liu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | KMSharing: The Framework and Space Abstraction for Efficient Data Sharing at the Network EdgeabstractEdge storage promises to be crucial for edge computing infrastructure, which enables users to access data within a low delay from widespread storage nodes at the network edge. The key challenge is how to integrate massive geographically distributed weak edge nodes to form an efficient storage system, enabling users to launch data operations from any node or retrieve the desired data across the entire distributed system. To address this data-sharing problem, researchers from both the traditional peer-to-peer (P2P) overlay networking and emerging edge computing fields have proposed some decentralized indexing mechanisms. However, existing studies lack insightful descriptions and analyses about the nature of the data-sharing problem at the network edge. It motivates us to rethink the edge data-sharing framework and provide the problem reformulation for analyzing the limitations of existing schemes. We reveal that the existing data-sharing schemes fail in complex network topologies which can be regarded as high-dimensional network spaces beyond the representation of low-dimensional Euclidean spaces or other existing hash spaces. A better space abstraction is an urgent need to alleviate the performance degradation due to the dimensional mismatch between network spaces and virtual spaces. To fill this gap, this paper proposes the Kautz metric space, a novel space abstraction extended from Kautz graphs, where the coordinates and the metric are defined as Kautz strings and Kautz distances (i.e., the shortest distances in undirected Kautz graphs), respectively. We design a dynamic programming algorithm to directly compute the Kautz distances. Then, we propose KMSharing, an efficient edge data-sharing scheme: both nodes and data are represented in a Kautz metric space, where the Kautz distance of any two Kautz strings reflects the network delay of the corresponding nodes. The workflow of KMSharing consists of three core components: the virtual address allocation represents edge nodes in the Kautz metric space; the data-to-node mapping ensures the uniqueness of target nodes; and forwarding table construction ensures the data delivery. Theoretical analyses confirm that KMSharing ideally achieves$\mathcal {O}\left ({{ \tau }}\right)$network delays,$\mathcal {O}\left ({{ \log N }}\right)$overlay hops, and$\mathcal {O}\left ({{ 1 }}\right)$forwarding entries in an N-node edge system with the network radius$\tau $, while the successive ensuring data delivery. Its worst-case network delay$\mathcal {O}\left ({{ \tau \log N }}\right)$is also much better than${\mathcal {O}\left ({{ \tau N^{\alpha } }}\right)},\alpha \mathrm {\in }(0,1)$, the worst case of the baselines using Euclidean spaces. Evaluation on various network topologies also shows that our KMSharing effectively reduces network delays and indexing costs than existing data-sharing schemes. Yuchen Sun 0001, Lailong Luo, Deke Guo, Li Liu 0002 |
IEEE/ACM Trans. Netw. | 4 |
| 2023 | Mapping Degeneration Meets Label Evolution: Learning Infrared Small Target Detection with Single Point SupervisionabstractTraining a convolutional neural network (CNN) to detect infrared small targets in a fully supervised manner has gained remarkable research interests in recent years, but is highly labor expensive since a large number of per-pixel annotations are required. To handle this problem, in this paper, we make the first attempt to achieve infrared small target detection with point-level supervision. Interestingly, during the training phase supervised by point labels, we discover that CNNs first learn to segment a cluster of pixels near the targets, and then gradually converge to predict groundtruth point labels. Motivated by this “mapping degeneration” phenomenon, we propose a label evolution framework named label evolution with single point supervision (LESPS) to progressively expand the point label by leveraging the intermediate predictions of CNNs. In this way, the network predictions can finally approximate the updated pseudo labels, and a pixel-level target mask can be obtained to train CNNs in an end-to-end manner. We conduct extensive experiments with insightful visualizations to validate the effectiveness of our method. Experimental results show that CNNs equipped with LESPS can well recover the target masks from corresponding point labels, and can achieve over 70% and 95% of their fully supervised performance in terms of pixel-level intersection over union (IoU) and object-level probability of detection (Pd), respectively. Code is available at https://github.com/XinyiYing/LESPS. Xinyi Ying, Li Liu 0002, Yingqian Wang 0002, Ruojing Li, Zaiping Lin, Weidong Sheng, Shilin Zhou 0001 |
CVPR | 2 |
| 2023 | ROFusion: Efficient Object Detection Using Hybrid Point-Wise Radar-Optical Fusion
Shuaifeng Zhi, Zhenhua Du, Li Liu 0002, Xinyu Zhang 0010, Kai Huo, Weidong Jiang |
ICANN (7) | 4 |
| 2023 | Cross-Domain Few-Shot Classification Via Inter-Source StylizationabstractThe goal of Cross-Domain Few-Shot Classification (CDFSC) is to accurately classify a target dataset with limited labelled data by exploiting the knowledge of a richly labelled auxiliary dataset, despite the differences between the domains of the two datasets. Some existing approaches require labelled samples from multiple domains for model training. However, these methods fail when the sample labels are scarce. To overcome this challenge, this paper proposes a solution that makes use of multiple source domains without the need for additional labeling costs. Specifically, one of the source domains is completely tagged, while the others are untagged. An Inter-Source Stylization Network (ISSNet) is then introduced to enhance stylisation across multiple source domains, enriching data distribution and model’s generalization capabilities. Experiments on 8 target datasets show that ISSNet leverages unlabelled data from multiple source data and significantly reduces the negative impact of domain gaps on classification performance compared to several baseline methods. Huali Xu, Shuaifeng Zhi, Li Liu 0002 |
ICIP | 3 |
| 2023 | Evidential Uncertainty and Diversity Guided Active Learning for Scene Graph Generation
Shuzhou Sun, Shuaifeng Zhi, Janne Heikkilä, Li Liu 0002 |
ICLR | 4 |
| 2023 | Variational Information Bottleneck for Cross Domain Object DetectionabstractCross domain object detection leverages a labeled source domain to learn an object detector which performs well in a novel unlabeled target domain. Most existing works mainly align the distribution utilizing the entire image knowledge ignoring the obstacles of task-uncorrelated information to alleviate the domain discrepancy. To tackle this issue, we propose a novel module called Variational Instance Disentanglement (VID) based on information theory which aims to decouple the information of task-correlated while filtering out the task-uncorrelated factors at the instance level. Notably, the proposed VID can be used as a plug-and-play module without bringing extra network parameter cost. We equip it with adversarial network and self-training network forming Variational Instance Disentanglement Adversarial Network (VIDAN) and Variational Instance Disentanglement Self-training Network (VIDSN), respectively. Extensive experiments on multiple widely-used scenarios show that the proposed method improves the performance of the popular frameworks and outperforms state-of-the-art methods. Jiangming Chen, Wanxia Deng, Tianpeng Liu, Yingmei Wei, Li Liu 0002 |
ICME | 6 |
| 2023 | Score-based causal feature selection for cancer risk predictionabstractThe primary goal of cancer risk prediction is to find dominant features, with each being responsible for cancer diagnosis. As a result, selected feature sets often converge to inexplicable and implausible results. Existing score-based causal models adopt a score function to learn a local causal structure to explicitly characterize a unique causal configuration as a variable number of nodes and links, which however suffer from nonconvex optimization and global incompleteness. This leads us to present a score-based approach to construct a causal network by optimizing a score function with a convex solution under the constraint of causal Markov property. It can be analytically shown that the resulting causal network satisfies the causal Markov property, and as a result, all cause-effect dependencies can be retained and are globally consistent. An additional node selector is introduced to choose the most dominant causal features. Empirical evaluations on three benchmarks and one in-house cancer risk datasets suggest our approach significantly outperforms the state-of-the-arts. Shanshan Huang 0004, Lei Wang 0197, Yuanhao Wang 0008, Li Liu 0002 |
ICME | 5 |
| 2023 | Prophet: An Efficient Feature Indexing Mechanism for Similarity Data Sharing at Network Edge
Yuchen Sun 0001, Deke Guo, Lailong Luo, Li Liu 0002, Xinyi Li 0001 |
INFOCOM | 4 |
| 2023 | Discovering and Explaining the Noncausality of Deep Learning in SAR ATRabstractIn recent years, deep learning has been widely used in SAR ATR and achieved excellent performance on the MSTAR dataset. However, due to constrained imaging conditions, MSTAR has data biases such as background correlation,i.e., background clutter properties have a spurious correlation with target classes. Deep learning can overfit clutter to reduce training errors. Therefore, the degree of overfitting for clutter reflects the non-causality of deep learning in SAR ATR. Existing methods only qualitatively analyze this phenomenon. In this paper, we quantify the contributions of different regions to target recognition based on the Shapley value. The Shapley value of clutter measures the degree of overfitting. Moreover, we explain how data bias and model bias contribute to non-causality. Concisely, data bias leads to comparable signal-to-clutter ratios and clutter textures in training and test sets. And various model structures have different degrees of overfitting for these biases. The experimental results of various models under standard operating conditions on the MSTAR dataset support our conclusions. Our code is available at https://github.com/waterdisappear/Data-Bias-in-MSTAR. Wei Yang 0046, Li Liu 0002, Wenpeng Zhang 0002, Yongxiang Liu |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | Learning Invariant Representation Via Contrastive Feature Alignment for Clutter Robust SAR ATRabstractThe deep neural networks (DNNs) have freed the synthetic aperture radar automatic target recognition (SAR ATR) from expertise-based feature designing and demonstrated superiority over conventional solutions. There has been shown the unique deficiency of ground vehicle benchmarks in shapes of strong background correlation results in DNNs overfitting the clutter and being non-robust to unfamiliar surroundings. However, the gap between fixed background model training and varying background application remains underexplored. This letter proposes a solution called Contrastive Feature Alignment (CFA) aiming to learn invariant representation for robust recognition. The proposed method contributes a mixed clutter variants generation strategy and a new inference branch equipped with channel-weighted mean square error (CWMSE) loss for invariant representation learning. In specific, the generation strategy is delicately designed to better attract clutter-sensitive deviation in feature space. The CWMSE loss is further devised to bettercontrastthis deviation andalignthe deep features activated by the original images and corresponding clutter variants. The proposed CFA combines both classification and CWMSE losses to train the model jointly, which allows for the progressive learning of invariant target representation. Extensive evaluations conducted on the MSTAR dataset and six DNN models prove the effectiveness of our proposal. The results demonstrate that the CFA-trained models are capable of recognizing targets among unfamiliar surroundings that are not included in the dataset, and are robust to varying signal-to-clutter ratios. Bowen Peng, Jianyue Xie, Li Liu 0002 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | Deep Learning for Instance Retrieval: A SurveyabstractIn recent years a vast amount of visual content has been generated and shared from many fields, such as social media platforms, medical imaging, and robotics. This abundance of content creation and sharing has introduced new challenges, particularly that of searching databases for similar content - Content Based Image Retrieval (CBIR) - a long-established research area in which improved efficiency and accuracy are needed for real-time retrieval. Artificial intelligence has made progress in CBIR and has significantly facilitated the process of instance search. In this survey we review recent instance retrieval works that are developed based on deep learning algorithms and techniques, with the survey organized by deep feature extraction, feature embedding and aggregation methods, and network fine-tuning strategies. Our survey considers a wide variety of recent methods, whereby we identify milestone work, reveal connections among various methods and present the commonly used benchmarks, evaluation results, common challenges, and propose promising future directions. Wei Chen 0072, Yu Liu 0012, Weiping Wang 0002, Erwin M. Bakker, Theodoros Georgiou 0001, Paul W. Fieguth, Li Liu 0002, Michael S. Lew |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | Lightweight Pixel Difference Networks for Efficient Visual Representation LearningabstractRecently, there have been tremendous efforts in developing lightweight Deep Neural Networks (DNNs) with satisfactory accuracy, which can enable the ubiquitous deployment of DNNs in edge devices. The core challenge of developing compact and efficient DNNs lies in how to balance the competing goals of achieving high accuracy and high efficiency. In this paper we propose two novel types of convolutions, dubbed Pixel Difference Convolution (PDC) and Binary PDC (Bi-PDC) which enjoy the following benefits: capturing higher-order local differential information, computationally efficient, and able to be integrated with existing DNNs. With PDC and Bi-PDC, we further present two lightweight deep networks named Pixel Difference Networks (PiDiNet) and Binary PiDiNet (Bi-PiDiNet) respectively to learn highly efficient yet more accurate representations for visual tasks including edge detection and object recognition. Extensive experiments on popular datasets (BSDS500, ImageNet, LFW, YTF, etc.) show that PiDiNet and Bi-PiDiNet achieve the best accuracy-efficiency trade-off. For edge detection, PiDiNet is the first network that can be trained without ImageNet, and can achieve the human-level performance on BSDS500 at 100 FPS and with 1 M parameters. For object recognition, among existing Binary DNNs, Bi-PiDiNet achieves the best accuracy and a nearly 2× reduction of computational cost on ResNet18. Zhuo Su 0002, Longguang Wang, Hua Zhang 0008, Zhen Liu 0004, Matti Pietikäinen, Li Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | Unbiased Scene Graph Generation via Two-Stage Causal ModelingabstractDespite the impressive performance of recent unbiased Scene Graph Generation (SGG) methods, the current debiasing literature mainly focuses on the long-tailed distribution problem, whereas it overlooks another source of bias, i.e., semantic confusion, which makes the SGG model prone to yield false predictions for similar relationships. In this paper, we explore a debiasing procedure for the SGG task leveraging causal inference. Our central insight is that the Sparse Mechanism Shift (SMS) in causality allows independent intervention on multiple biases, thereby potentially preserving head category performance while pursuing the prediction of high-informative tail relationships. However, the noisy datasets lead to unobserved confounders for the SGG task, and thus the constructed causal models are always causal-insufficient to benefit from SMS. To remedy this, we propose Two-stage Causal Modeling (TsCM) for the SGG task, which takes the long-tailed distribution and semantic confusion as confounders to the Structural Causal Model (SCM) and then decouples the causal intervention into two stages. The first stage is causal representation learning, where we use a novel Population Loss (P-Loss) to intervene in the semantic confusion confounder. The second stage introduces the Adaptive Logit Adjustment (AL-Adjustment) to eliminate the long-tailed distribution confounder to complete causal calibration learning. These two stages are model agnostic and thus can be used in any SGG model that seeks unbiased predictions. Comprehensive experiments conducted on the popular SGG backbones and benchmarks show that our TsCM can achieve state-of-the-art performance in terms of mean recall rate. Furthermore, TsCM can maintain a higher recall rate than other debiasing methods, which indicates that our method can achieve a better tradeoff between head and tail relationships. Shuzhou Sun, Shuaifeng Zhi, Qing Liao 0001, Janne Heikkilä, Li Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Redundancy-Reduced Sparsity-Based Adaptive Beamforming for Polarization-Sensitive ArraysabstractA sparse reconstruction approach for adaptive beamforming (ABF) with polarization-sensitive arrays (PSA) is introduced in this letter. It first represents the spatial sparsity of incoming signals as the row sparsity of a power-scaled polarization matrix, which arises from the matrization of the redundancy-reduced covariance vector. Then the row sparsity issue is relaxed to an$\ell _{2,1}$norm minimization form and solved in a gridless way via a compact formulation, where a dimension reduction method is introduced to reduce the problem size. Compared to existing techniques, the proposed method processes the polarization information holistically and derives each signal parameter in the continuous domain. Simulation results substantiate the advantages of the proposed method over competing methods. Tianpeng Liu, Junpeng Shi, Li Liu 0002, Yongxiang Liu |
IEEE Signal Process. Lett. | 4 |
| 2023 | A Model-Agnostic Approach to Mitigate Gradient Interference for Multi-Task LearningabstractMultitask learning (MTL) is a powerful technique for jointly learning multiple tasks. However, it is difficult to achieve a tradeoff between tasks during iterative training, as some tasks may compete with each other. Existing methods manually design specific network models to mitigate task conflicts, but they require considerable manual effort and prior knowledge about task relationships to tune the model so as to obtain the best performance for each task. Moreover, few works have offered formal descriptions of task conflicts and theoretical explanations for the cause of task conflict problems. In this article, we provide a formal description of task conflicts that are caused by the gradient interference problem of tasks. To alleviate this issue, we propose a novel model-agnostic approach to mitigate gradient interference (MAMG) by designing a gradient clipping rule that directly modifies the interfering components on the gradient interfering direction. Specifically, MAMG is model-agnostic and thus it can be applied to a large number of multitask models. We also theoretically prove the convergence of MAMG and its superiority to existing MTL methods. We evaluate our method on a variety of real-world large datasets, and extensive experimental results confirm that MAMG can outperform some state-of-the-art algorithms on different types of tasks and can be easily applied to various methods. Heyan Chai 0001, Ye Ding 0002, Li Liu 0002, Binxing Fang, Qing Liao 0001 |
IEEE Trans. Cybern. | 4 |
| 2023 | Structural Regression Fusion for Unsupervised Multimodal Change DetectionabstractMultimodal change detection (MCD) is an increasingly interesting but very challenging topic in remote sensing, which is due to the unavailability of detecting changes by directly comparing multimodal images from different domains. In this paper, we first analyze the structural asymmetry between multitemporal images and show their negative impact on the previous MCD methods using image structures. Specifically, when there is a structural asymmetry, previous structure based methods can only complete a structure comparison or image regression in one direction and fails in the other direction, that is, they cannot transform or convert from complex structural images (with more categories) to simple structural images (with fewer categories). To reduce the influence of structural asymmetry, we propose a structural regression fusion based method (SRF) that simultaneously transforms the pre-event and post-event images into the image domain of each other, calculating the forward and backward changed images, respectively. Noteworthy, different from previous late fusion methods that fuse the forward and backward changed images in the post-processing stage, SRF incorporates fusion into the regression process, which can fully explore the connection between changed images, and thus improve image transformation performance and obtain better changed images. Specifically, SRF yields three types of constraints to perform the fused image transformation: structure consistency based regression term, change smoothness and alignment based fusion term, and prior sparsity based penalty term. Finally, the changes can be extracted by comparing the transformed and original images. The proposed SRF is verified on six real data sets by comparing with some state-of-the-art methods. Source code of the proposed method will be made available at https://github.com/yulisun/SRF. Yuli Sun, Lin Lei, Li Liu 0002, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Incorporating Deep Background Prior Into Model-Based Method for Unsupervised Moving Vehicle Detection in Satellite VideosabstractBackground reconstruction is a key step of moving object detection in satellite videos. Most existing model-based methods exploit low-rank prior to recover background, which have achieved good performance but suffered degradation under complex and dynamic scenes. In this paper, we introduce a deep background prior into model-based methods for moving vehicle detection in satellite videos. Our deep background prior is obtained by a background reconstruction network, which can learn to reconstruct background from consecutive frames. By applying our deep background prior into model-based methods, a closed-form solution can be obtained via alternating direction method of multipliers (ADMM) and then detection results can be acquired through iterative optimization. More importantly, our background reconstruction network can be trained in an unsupervised way by introducing specifically designed loss, thus relieving the dependence on large-scale labeled dataset. Extensive experimental results demonstrate the efficiency and effectiveness of the proposed method. Ting Liu 0017, Xinyi Ying, Yingqian Wang 0002, Li Liu 0002, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Few-Shot Class-Incremental SAR Target Recognition via Cosine Prototype LearningabstractRecent years have witnessed a remarkable breakthrough in Synthetic Aperture Radar Automatic Target Recognition (SAR ATR) with the development of deep learning (DL). Nonetheless, once deployed, the DL-based methods’ ability to incrementally learn new knowledge from few-shot samples without forgetting the old is fragile, hindering them from discriminating unseen targets in real-world situations. In this paper, we propose a Cosine Prototype Learning (CPL) framework to first unlock few-shot class-incremental learning (FSCIL) in the SAR ATR field inspired by the intrinsic relationships between target azimuth-aware knowledge and semantic features under the cosine criterion. By condensing class-specific characteristics into individual prototypes, stable profiles of targets are depicted without losing generalization. For the model’s plasticity, a pairwise structure separation (PSS) loss is introduced to separate old and new classes and compact intra-class features. Meanwhile, the model’s transferability on new classes is guaranteed by a prototype consistency (PC) loss. For the model’s stability, we propose a prototype-exemplar distillation (PED) loss and a prototype re-calibration (PR) strategy to penalize semantic drifts of old-class feature spaces and alleviate the misalignment of the learned prototypes successively. At inference, a nearest-class-mean (NCM) classifier is adopted for evaluation by comparing cosine similarity scores between testing samples and class-specific prototypes. In experiments, the proposed components of our method are explored by ablation studies. Strong baselines are established, and extensive experiments conducted on the MSTAR dataset show that our method outperforms state-of-the-art methods under various FSCIL conditions, verifying its effectiveness for the FSCIL of SAR ATR. Yan Zhao 0026, Lingjun Zhao, Dewen Hu, Gangyao Kuang, Li Liu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | New Wine Old Bottles: Feistel Structure RevisedabstractThis paper mainly investigates the iterative structures whose decryption is similar to the encryption. Firstly, we unify many well-known structures which share similar procedures between the decryption and the encryption, and give a sufficient and necessary condition for this structure to be bijective, which reveals many new insights into the Feistel structure as well as the Lai-Massey structure. Secondly, we analyze the security of the unified structure against the known cryptanalysis. By extending the dual structure from a Feistel structure to the unified structure, we prove that a differential of the unified structure is impossible if and only if it is a zero-correlation linear hull of its dual structure, which presents a generalized link between the impossible differential and zero-correlation linear cryptanalysis shown in CRYPTO 2015. Significantly, several constraints on the linear components of the cipher and the permutation on the branches of the cipher are specified to make the structure resilient to differential and linear cryptanalysis. Furthermore, in the case that the order of the permutation equals the number of the branches$n$, we prove that there always exist a$(3n-1)$-round impossible differential and a$(3n-1)$-round zero-correlation linear hull of the structure, and also present an algorithm to construct these distinguishers. Finally, we propose some novel structures which might be used in future block cipher designs. Bing Sun 0001, Li Liu 0002, Hua Zhang 0008, Chao Li 0002 |
IEEE Trans. Inf. Theory | 5 |
| 2023 | Lifelong Fine-Grained Image RetrievalabstractFine-grained image retrieval has been extensively explored in a zero-shot manner. A deep model is trained on the seen part and then evaluated the generalization performance on the unseen part. However, this setting is infeasible for many real-world applications since (1) the retrieval dataset can be non-fixed so that new data are added constantly, and (2) data samples of the seen categories are also common in practice and are important for evaluation. In this paper, we explore lifelong fine-grained image retrieval (LFGIR), which learns continuously on a sequence of new tasks with data from different datasets. We first use knowledge distillation to minimize catastrophic forgetting on old tasks. Training continuously on different datasets causes large domain shifts between the old and new tasks while image retrieval is sensitive to even small shifts in the features. This tends to weaken the effectiveness of knowledge distillation by the frozen teacher. To mitigate the impact of domain shifts, we use the network inversion method to generate images of the old tasks. In addition, we design an on-the-fly teacher which transfers knowledge captured on a new task to the student to improve better generalization performance, thereby achieving a better balance between old and new tasks in the end. We name the whole framework as Dual Knowledge Distillation (DKD), whose efficacy is demonstrated by extensive experimental results on sequential tasks including seven datasets. Wei Chen 0072, Haoyang Xu, Nan Pu, Yu Liu 0012, Mingrui Lao, Weiping Wang 0002, Li Liu 0002, Michael S. Lew |
IEEE Trans. Multim. | 7 |
| 2023 | Uncertainty-Guided Semi-Supervised Few-Shot Class-Incremental Learning With Knowledge DistillationabstractClass-Incremental Learning (CIL) aims at incrementally learning novel classes without forgetting old ones. This capability becomes more challenging when novel tasks contain one or a few labeled training samples, which leads to a more practical learning scenario,i.e., Few-Shot Class- Incremental Learning (FSCIL). The dilemma on FSCIL lies in serious overfitting and exacerbated catastrophic forgetting caused by the limited training data from novel classes. In this paper, excited by the easy accessibility of unlabeled data, we conduct a pioneering work and focus on a Semi-Supervised Few-Shot Class-Incremental Learning (Semi-FSCIL) problem, which requires the model incrementally to learn new classes from extremely limited labeled samples and a large number of unlabeled samples. To address this problem, a simple but efficient framework is first constructed based on the knowledge distillation technique to alleviate catastrophic forgetting. To efficiently mitigate the overfitting problem on novel categories with unlabeled data, uncertainty-guided semi-supervised learning is incorporated into this framework to select unlabeled samples into incremental learning sessions considering the model uncertainty. This process provides extra reliable supervision for the distillation process and contributes to better formulating the class means. Our extensive experiments on CIFAR100, miniImageNet and CUB200 datasets demonstrate the promising performance of our proposed method, and define baselines in this new research direction. Yawen Cui, Wanxia Deng, Xin Xu 0001, Zhen Liu 0004, Zhong Liu 0002, Matti Pietikäinen, Li Liu 0002 |
IEEE Trans. Multim. | 7 |
| 2023 | Importance-Aware Information Bottleneck Learning Paradigm for Lip ReadingabstractLip reading is the task of decoding text from speakers' mouth movements. Numerous deep learning-based methods have been proposed to address this task. However, these existing deep lip reading models suffer from poor generalization due to overfitting the training data. To resolve this issue, we present a novel learning paradigm that aims to improve the interpretability and generalization of lip reading models. In specific, a Variational Temporal Mask (VTM) module is customized to automatically analyze the importance of frame-level features. Furthermore, the prediction consistency constraints of global information and local temporal important features are introduced to strengthen the model generalization. We evaluate the novel learning paradigm with multiple lip reading baseline models on the LRW and LRW-1000 datasets. Experiments show that the proposed framework significantly improves the generalization performance and interpretability of lip reading models. Changchong Sheng, Li Liu 0002, Wanxia Deng, Liang Bai 0003, Zhong Liu 0002, Songyang Lao, Gangyao Kuang, Matti Pietikäinen |
IEEE Trans. Multim. | 2 |
| 2023 | Late Fusion Multiple Kernel Clustering With Proxy Graph RefinementabstractMultiple kernel clustering (MKC) optimally utilizes a group of pre-specified base kernels to improve clustering performance. Among existing MKC algorithms, the recently proposed late fusion MKC methods demonstrate promising clustering performance in various applications and enjoy considerable computational acceleration. However, we observe that the kernel partition learning and late fusion processes are separated from each other in the existing mechanism, which may lead to suboptimal solutions and adversely affect the clustering performance. In this article, we propose a novel late fusion multiple kernel clustering with proxy graph refinement (LFMKC-PGR) framework to address these issues. First, we theoretically revisit the connection between late fusion kernel base partition and traditional spectral embedding. Based on this observation, we construct a proxy self-expressive graph from kernel base partitions. The proxy graph in return refines the individual kernel partitions and also captures partition relations in graph structure rather than simple linear transformation. We also provide theoretical connections and considerations between the proposed framework and the multiple kernel subspace clustering. An alternate algorithm with proved convergence is then developed to solve the resultant optimization problem. After that, extensive experiments are conducted on 12 multi-kernel benchmark datasets, and the results demonstrate the effectiveness of our proposed algorithm. The code of the proposed algorithm is publicly available at https://github.com/wangsiwei2010/graphlatefusion_MKC. Siwei Wang 0001, Xinwang Liu 0002, Li Liu 0002, Sihang Zhou 0001, En Zhu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Sim2Word: Explaining Similarity with Representative Attribute Words via Counterfactual ExplanationsabstractRecently, we have witnessed substantial success using the deep neural network in many tasks. Although there still exist concerns about the explainability of decision making, it is beneficial for users to discern the defects in the deployed deep models. Existing explainable models either provide the image-level visualization of attention weights or generate textual descriptions as post hoc justifications. Different from existing models, in this article we propose a new interpretation method that explains the image similarity models by salience maps and attribute words. Our interpretation model contains visual salience maps generation and the counterfactual explanation generation. The former has two branches: global identity relevant region discovery and multi-attribute semantic region discovery. The first branch aims to capture the visual evidence supporting the similarity score, which is achieved by computing counterfactual feature maps. The second branch aims to discover semantic regions supporting different attributes, which helps to understand which attributes in an image might change the similarity score. Then, by fusing visual evidence from two branches, we can obtain the salience maps indicating important response evidence. The latter will generate the attribute words that best explain the similarity using the proposed erasing model. The effectiveness of our model is evaluated on the classical face verification task. Experiments conducted on two benchmarks—VGGFace2 and Celeb-A—demonstrate that our model can provide convincing interpretable explanations for the similarity. Moreover, our algorithm can be applied to evidential learning cases, such as finding the most characteristic attributes in a set of face images, and we verify its effectiveness on the VGGFace2 dataset. Ruoyu Chen 0001, Jingzhi Li 0002, Hua Zhang 0008, Changchong Sheng, Li Liu 0002, Xiaochun Cao |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2022 | SVNet: Where SO(3) Equivariance Meets Binarization on Point Cloud RepresentationabstractEfficiency and robustness are increasingly needed for applications on 3D point clouds, with the ubiquitous use of edge devices in scenarios like autonomous driving and robotics, which often demand real-time and reliable responses. The paper tackles the challenge by designing a general framework to construct 3D learning architectures with SO(3) equivariance and network binarization. However, a naive combination of equivariant networks and binarization either causes sub-optimal computational efficiency or geometric ambiguity. We propose to locate both scalar and vector features in our networks to avoid both cases. Precisely, the presence of scalar features makes the major part of the network binarizable, while vector features serve to retain rich structural information and ensure SO(3) equivariance. The proposed approach can be applied to general backbones like PointNet and DGCNN. Meanwhile, experiments on ModelNet40, ShapeNet, and the real-world dataset ScanObjectNN, demonstrated that the method achieves a great trade-off between efficiency, rotation robustness, and accuracy. The codes are available at https://github.com/zhuoinoulu/svnet. Zhuo Su 0002, Max Welling, Matti Pietikäinen, Li Liu 0002 |
3DV | 4 |
| 2022 | Decoupling Makes Weakly Supervised Local Feature BetterabstractWeakly supervised learning can help local feature methods to overcome the obstacle of acquiring a large-scale dataset with densely labeled correspondences. However, since weak supervision cannot distinguish the losses caused by the detection and description steps, directly conducting weakly supervised learning within a joint training describe-then-detect pipeline suffers limited performance. In this paper, we propose a decoupled training describe-then-detect pipeline tailored for weakly supervised local feature learning. Within our pipeline, the detection step is decoupled from the description step and postponed until discriminative and robust descriptors are learned. In addition, we introduce a line-to-window search strategy to explicitly use the camera pose information for better descriptor learning. Extensive experiments show that our method, namely PoSFeat (Camera Pose Supervised Feature), outperforms previous fully and weakly supervised methods and achieves state-of-the-art performance on a wide range of downstream task. Kunhong Li 0001, Longguang Wang, Li Liu 0002, Qing Ran, Kai Xu 0004, Yulan Guo |
CVPR | 3 |
| 2022 | Learnable Lookup Table for Neural Network QuantizationabstractNeural network quantization aims at reducing bit-widths of weights and activations for memory and computational efficiency. Since a linear quantizer (i.e., round(·) function) cannot well fit the bell-shaped distributions of weights and activations, many existing methods use predefined functions (e.g., exponential function) with learnable parameters to build the quantizer for joint optimization. However, these complicated quantizers introduce considerable computational overhead during inference since activation quantization should be conducted online. In this paper, we formulate the quantization process as a simple lookup operation and propose to learn lookup tables as quantizers. Specifically, we develop differentiable lookup tables and introduce several training strategies for optimization. Our lookup tables can be trained with the network in an end-to-end manner to fit the distributions in different layers and have very small additional computational cost. Comparison with previous methods show that quantized networks using our lookup tables achieve state-of-the-art performance on image classification, image super-resolution, and point cloud classification tasks. Longguang Wang, Yingqian Wang 0002, Li Liu 0002, Wei An 0003, Yulan Guo |
CVPR | 4 |
| 2022 | Highly-efficient Incomplete Largescale Multiview Clustering with Consensus Bipartite GraphabstractMultiview clustering has received increasing attention due to its effectiveness in fusing complementary information without manual annotations. Most previous methods hold the assumption that each instance appears in all views. However, it is not uncommon to see that some views may contain some missing instances, which gives rise to incomplete multi-view clustering (IMVC) in literature. Although many IMVC methods have been recently proposed, they always encounter high complexity and expensive time expenditure from being applied into large-scale tasks. In this paper, we present a flexible highly-efficient incomplete large-scale multi-view clustering approach based on bipartite graph framework to solve these issues. Specifically, we formalize multi-view anchor learning and incomplete bipartite graph into a unified framework, which coordinates with each other to boost cluster performance. By introducing the flexible bipartite graph framework to handle IMVC for the first practice, our proposed method enjoys linear complexity respecting to instance numbers, which is more applicable for large-scale IMVC tasks. Comprehensive experimental results on various benchmark datasets demonstrate the effectiveness and efficiency of our proposed algorithm against other IMVC competitors. The code is available at11https://github.com/wangsiwei2010/CVPR22-IMVC-CBG. Siwei Wang 0001, Xinwang Liu 0002, Li Liu 0002, Wenxuan Tu, Xinzhong Zhu, Jiyuan Liu 0003, Sihang Zhou 0001, En Zhu |
CVPR | 3 |
| 2022 | Dynamic Binary Neural Network by Learning Channel-Wise ThresholdsabstractBinary neural networks (BNNs) constrain weights and activations to +1 or -1 with limited storage and computational cost, which is hardware-friendly for portable devices. Recently, BNNs have achieved remarkable progress and been adopted into various fields. However, the performance of BNNs is sensitive to activation distribution. The existing BNNs utilized the Sign function with predefined or learned static thresholds to binarize activations. This process limits representation capacity of BNNs since different samples may adapt to unequal thresholds. To address this problem, we propose a dynamic BNN (DyBNN) incorporating dynamic learnable channel-wise thresholds of Sign function and shift parameters of PReLU. The method aggregates the global information into the hyper function and effectively increases the feature expression ability. The experimental results prove that our method is an effective and straightforward way to reduce information loss and enhance performance of BNNs. The DyBNN based on two backbones of ReActNet (MobileNetV1 and ResNet18) achieve 71.2% and 67.4% top1-accuracy on ImageNet dataset, outperforming baselines by a large margin (i.e., 1.8% and 1.5% respectively). Zhuo Su 0002, Yang-He Feng, Xin Lu 0002, Matti Pietikäinen, Li Liu 0002 |
ICASSP | 6 |
| 2022 | Boosting Lip Reading with a Multi-View Fusion NetworkabstractLip reading aims to decode speech information by analyzing lip movement without involving audio. Numerous deep learning based methods are proposed to address this task. Generally, most existing methods extract visual features only based on the lip appearance, while ignoring the shape dynamic information of the lip region. Motivated by this, we propose a Multi-View Fusion Network (MVFN), which can extract more discriminative visual representations by incorporating appearance and shape information. Besides, a novel adaptive graph convolutional network model called Adaptive Spatial Graph Model(ASGM) is proposed to learn lip spatial topology and lip shape dynamics automatically. Experiments on LRW (word-level) and OuluVS2 (phrase-level) clearly show that the proposed method significantly outperforms the baseline methods by a large margin and achieves state-of-the-art performance. Xueyi Zhang 0001, Jinping Sui, Changchong Sheng, Wanxia Deng, Li Liu 0002 |
ICME | 6 |
| 2022 | Facial Kinship Verification: A Comprehensive Review and OutlookabstractThe goal of Facial Kinship Verification (FKV) is to automatically determine whether two individuals have a kin relationship or not from their given facial images or videos. It is an emerging and challenging problem that has attracted increasing attention due to its practical applications. Over the past decade, significant progress has been achieved in this new field. Handcrafted features and deep learning techniques have been widely studied in FKV. The goal of this paper is to conduct a comprehensive review of the problem of FKV. We cover different aspects of the research, including problem definition, challenges, applications, benchmark datasets, a taxonomy of existing methods, and state-of-the-art performance. In retrospect of what has been achieved so far, we identify gaps in current research and discuss potential future research directions. Xiaoting Wu, Xiaoyi Feng, Xiaochun Cao, Xin Xu 0001, Dewen Hu, Miguel Bordallo López, Li Liu 0002 |
Int. J. Comput. Vis. | 7 |
| 2022 | Speckle-Variant Attack: Toward Transferable Adversarial Attack to SAR Target RecognitionabstractRecent advances of deep neural networks (DNNs) highlight the success on synthetic aperture radar automatic target recognition (SAR ATR) with superiority effectiveness and efficiency. However, the DNNs are known to be vulnerable to the adversarial examples, whose performance will be dramatically reduced when the imperceptible perturbation exists. In optical image processing, invisible perturbations are typically embedded in the way of a full-scaled distribution in purely digital setting. Whereas, it is not feasible to achieve this in SAR ATR tasks due to the inaccessibility of SAR system and unique imaging mechanism. In practical, the subtle perturbations could be produced by physical approaches that change the scattering property of the target. Therefore, the adversarial perturbations for SAR ATR should be of good transferability to achieve effective attack on major DNNs classifiers, as well as accessible additive region in SAR images with respect to the realistic target locations. In this letter, we present a novel approach, namely speckle variant attack (SVA). The proposed SVA is composed of two major modules: an iterative gradient based perturbation generator and a target region extractor. The perturbation generator implements a speckle variant transformation that continuously reconstruct the speckle noise pattern during each of the iterations for strong transferability. The target region extractor ensures the feasibility of the additive adversarial perturbations in practical scenarios through restricting the region of the perturbation. Therefore, the proposed SVA is capable of producing adversarial examples that are more transferable and physically feasible. Extensive evaluations on the MSTAR dataset show that the SVA has achieved the superior transferability and competitive time consumption compared with the SOTA transformation-based techniques, including the diverse inputs method and the scale-invariant method. Bowen Peng, Jie Zhou 0031, Jingyuan Xia, Li Liu 0002 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | DSFNet: Dynamic and Static Fusion Network for Moving Object Detection in Satellite VideosabstractMoving object detection (MOD) in satellite videos remains challenging due to the extremely small size of the interested targets and the highly complex background. Both the intra-frame (static) and inter-frame (dynamic) information are of great importance to MOD. In this letter, we propose a two-stream detection network named dynamic and static fusion network (DSFNet) to tackle the MOD problem in satellite videos. Specifically, the DSFNet is composed of a 2-D backbone to extract static context information from a single frame and a lightweight 3-D backbone to extract dynamic motion cues from consecutive frames. Then the extracted static and dynamic features are fused and fed into the detection head to detect the moving targets in satellite videos. We conduct extensive experiments on videos collected from Jilin-1 satellite and the results have demonstrated the effectiveness and robustness of the proposed DSFNet. Experimental results show that our DSFNet achieves the-state-of-the-art performance. Xinyi Ying, Ruojing Li, Shuanglin Wu, Li Liu 0002, Wei An 0003 |
IEEE Geosci. Remote. Sens. Lett. | 7 |
| 2022 | Coarse-to-fine pseudo supervision guided meta-task optimization for few-shot object classification
Yawen Cui, Qing Liao 0001, Dewen Hu, Wei An 0003, Li Liu 0002 |
Pattern Recognit. | 5 |
| 2022 | Scale-selective and noise-robust extended local binary pattern for texture classification
Qiwu Luo, Jiaojiao Su, Chunhua Yang 0001, Olli Silvén, Li Liu 0002 |
Pattern Recognit. | 5 |
| 2022 | Deep Ladder-Suppression Network for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) aims at learning a classifier for an unlabeled target domain by transferring knowledge from a labeled source domain with a related but different distribution. Most existing approaches learn domain-invariant features by adapting the entire information of the images. However, forcing adaptation of domain-specific variations undermines the effectiveness of the learned features. To address this problem, we propose a novel, yet elegant module, called the deep ladder-suppression network (DLSN), which is designed to better learn the cross-domain shared content by suppressing domain-specific variations. Our proposed DLSN is an autoencoder with lateral connections from the encoder to the decoder. By this design, the domain-specific details, which are only necessary for reconstructing the unlabeled target data, are directly fed to the decoder to complete the reconstruction task, relieving the pressure of learning domain-specific variations at the later layers of the shared encoder. As a result, DLSN allows the shared encoder to focus on learning cross-domain shared content and ignores the domain-specific variations. Notably, the proposed DLSN can be used as a standard module to be integrated with various existing UDA frameworks to further boost performance. Without whistles and bells, extensive experimental results on four gold-standard domain adaptation datasets, for example: 1) Digits; 2) Office31; 3) Office-Home; and 4) VisDA-C, demonstrate that the proposed DLSN can consistently and significantly improve the performance of various popular UDA frameworks. Wanxia Deng, Lingjun Zhao, Gangyao Kuang, Dewen Hu, Matti Pietikäinen, Li Liu 0002 |
IEEE Trans. Cybern. | 6 |
| 2022 | Dynamic Sparse Subspace Clustering for Evolving High-Dimensional Data StreamsabstractIn an era of ubiquitous large-scale evolving data streams, data stream clustering (DSC) has received lots of attention because the scale of the data streams far exceeds the ability of expert human analysts. It has been observed that high-dimensional data are usually distributed in a union of low-dimensional subspaces. In this article, we propose a novel sparse representation-based DSC algorithm, called evolutionary dynamic sparse subspace clustering (EDSSC). It can cope with the time-varying nature of subspaces underlying the evolving data streams, such as subspace emergence, disappearance, and recurrence. The proposed EDSSC consists of two phases: 1) static learning and 2) online clustering. During the first phase, a data structure for storing the statistic summary of data streams, called EDSSC summary, is proposed which can better address the dilemma between the two conflicting goals: 1) saving more points for accuracy of subspace clustering (SC) and 2) discarding more points for the efficiency of DSC. By further proposing an algorithm to estimate the subspace number, the proposed EDSSC does not need to know the number of subspaces. In the second phase, a more suitable index, called the average sparsity concentration index (ASCI), is proposed, which dramatically promotes the clustering accuracy compared to the conventionally utilized SCI index. In addition, the subspace evolution detection model based on the Page-Hinkley test is proposed where the appearing, disappearing, and recurring subspaces can be detected and adapted. Extinct experiments on real-world data streams show that the EDSSC outperforms the state-of-the-art online SC approaches. Jinping Sui, Zhen Liu 0004, Li Liu 0002, Alexander Jung 0001, Xiang Li 0014 |
IEEE Trans. Cybern. | 3 |
| 2022 | Scattering Model Guided Adversarial Examples for SAR Target Recognition: Attack and DefenseabstractDeep Neural Networks (DNNs) based Synthetic Aperture Radar (SAR) Automatic Target Recognition (ATR) systems have shown to be highly vulnerable to adversarial perturbations that are deliberately designed yet almost imperceptible but can bias DNN inference when added to targeted objects. This leads to serious safety concerns when applying DNNs to high-stakes SAR ATR applications. Therefore, enhancing the adversarial robustness of DNNs is essential for applying DNNs to modern real-world SAR ATR systems. Toward building more robust DNN-based SAR ATR models, this article explores the domain knowledge of SAR imaging process and proposes a novel Scattering Model Guided Adversarial Attack (SMGAA) algorithm which can generate adversarial perturbations in the form of electromagnetic scattering response (called adversarial scatterers). The proposed SMGAA consists of two parts: 1) a parametric scattering model and corresponding imaging method and 2) a customized gradient-based optimization algorithm. First, we introduce the effective Attributed Scattering Center Model (ASCM) and a general imaging method to describe the scattering behavior of typical geometric structures in the SAR imaging process. By further devising several strategies to take the domain knowledge of SAR target images into account and relax the greedy search procedure, the proposed method does not need to be prudentially finetuned, and can efficiently find the effective ASCM parameters to fool the SAR classifiers and facilitate the robust model training. Comprehensive evaluations on the MSTAR dataset show that the adversarial scatterers generated by SMGAA are more robust to perturbations and transformations in the SAR processing chain than the currently studied attacks, and are effective to construct a defensive model against the malicious scatterers. Bowen Peng, Jie Zhou 0031, Jianyue Xie, Li Liu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Graph Signal Processing for Heterogeneous Change DetectionabstractThis paper provides a new strategy for the heterogeneous change detection (HCD) problem: solving HCD from the perspective of graph signal processing (GSP). We construct a graph to represent the structure of each image, and treat each image as a graph signal defined on the graph. In this way, we convert the HCD into a GSP problem: a comparison of the responses of signals on systems defined on the graphs, which attempts to find structural differences and signal differences due to the changes between heterogeneous images. Firstly, we analyze the GSP for HCD from the vertex domain. We show that once a region has changed, the local structure of image changes,i.e. the connectivity of the vertex containing this region changes. Therefore, we can compare the output signals of the same input graph signal passing through filters defined on the two graphs to detect changes. We analyze the negative effects of changing regions on the change detection results from the viewpoint of signal propagation, and we also design different filters from the vertex domain to explore the high-order neighborhood information hidden in original graphs. Secondly, we analyze the GSP for HCD from the spectral domain. We explore the spectral properties of different images on the same graph, and show that their spectra exhibit commonalities and dissimilarities. Specifically, it is the change that leads to the dissimilarities of their spectra. With the help of graph spectral analysis, we propose a regression model for the HCD, which decomposes the source signal into the regressed signal and changed signal, and constrains the spectral property of the regressed signal. Experiments conducted on seven real data sets show the effectiveness of the vertex domain filtering based and spectral domain analysis based HCD methods. Source code will be made available at https://github.com/yulisun/HCD-GSP. Yuli Sun, Lin Lei, Dongdong Guan, Gangyao Kuang, Li Liu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Hyperspectral Estimation of Soil Copper Concentration Based on Improved TabNet Model in the Eastern Junggar CoalfieldabstractChina is the largest coal consumer in the world. The massive exploitation and utilization of coal resources has resulted in serious problems of heavy metal pollution and environmental contamination, such as soil degradation, water pollution, crop damage, and even threatening human lives. Therefore, monitoring soil heavy metal pollution quickly and in real time is an urgent task at present. This research not only formulated a new preprocessing method enlightened by few-shot learning for soil hyperspectral data, but also combined it with other soil-related auxiliary information to extract effective information from the soil hyperspectrum, at the end of which different regression methods were adopted to predict soil heavy metal contamination. This test used 168 actual soil samples from the Eastern Junggar coalfield in Xinjiang for verification. Since copper in the soil is a trace element and the corresponding spectral characteristics are affected by other impurities, improper use of hyperspectral preprocessing methods may introduce interference information or may delete useful information, which makes the model effect unsatisfied. To effectively address the above problems, the preprocessing method of this experiment combined the second-order differential derivation, data enhancement method together with the addition of auxiliary information to allow more effective features to be entered into the model. Next, the Attentive Interpretable Tabular Learning (TabNet) model was improved in three different ways using the original TabNet model and three improved TabNet models to create regression models. One of the improved TabNet models had the best effect, with a list of the top 30 features according to the degree of importance. Meanwhile, the regression prediction of Cu content using four different convolutional neural networks (CNN) revealed that the model with the residual block was the strongest and slightly outperformed the improved TabNet model, but lacked interpretation of the input data. Besides, this experiment also employed different pre-processing methods for regression prediction on various models, and found that the traditional pre-processing methods performed best in traditional regression models (e.g., PLSR) and underperformed in deep learning models. The selected optimal model was compared with partial least square regression (PLSR), and convolutional neural network (CNN) models. The results indicated that both the improved TabNet model and improved CNN model had better performance using the new preprocessing approach proposed in this paper, with improved TabNet yielding a coefficient of determination (R2), root mean square error (RMSE) and ratio of performance to interquartile range (RPIQ) of 0.94, 1.341 and 4.474, respectively. The improved CNN model had a coefficient of determination of 0.942, a root mean square error of 1.324 and an interquartile range of 4.531 in the test dataset. Yuan Wang 0034, Abdugheni Abliz, Hongbing Ma, Li Liu 0002, Alishir Kurban, Ümüt Halik, Matti Pietikäinen |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Domain Knowledge Powered Two-Stream Deep Network for Few-Shot SAR Vehicle RecognitionabstractSynthetic aperture radar (SAR) target recognition faces the challenge that there are very little labeled data. Although few-shot learning methods are developed to extract more information from a small amount of labeled data to avoid overfitting problems, recent few-shot or limited-data SAR target recognition algorithms overlook the unique SAR imaging mechanism. Domain knowledge-powered two-stream deep network (DKTS-N) is proposed in this study, which incorporates SAR domain knowledge related to the azimuth angle, the amplitude, and the phase data of vehicles, making it a pioneering work in few-shot SAR vehicle recognition. The two-stream deep network, extracting the features of the entire image and image patches, is proposed for more effective use of the SAR domain knowledge. To measure the structural information distance between the global and local features of vehicles, the deep Earth mover’s distance is improved to cope with the features from a two-stream deep network. Considering the sensitivity of the azimuth angle in SAR vehicle recognition, the nearest neighbor classifier replaces the structured fully connected layer for$K$-shot classification. All experiments are conducted under the configuration that the SARSIM and the Moving and Stationary Target Acquisition and Recognition (MSTAR) dataset work as a source and target task, respectively. Our proposed DKTS-N achieved 49.26% and 96.15% under ten-way one-shot and ten-way 25-shot, whose labeled samples are randomly selected from the training set. In standard operating condition (SOC) as well as three extended operating conditions (EOCs), DKTS-N demonstrated overwhelming advantages in accuracy and time consumption compared with other few-shot learning methods in$K$-shot recognition tasks. Linbin Zhang, Xiangguang Leng, Sijia Feng, Xiaojie Ma, Kefeng Ji, Gangyao Kuang, Li Liu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2022 | Attentional Feature Refinement and Alignment Network for Aircraft Detection in SAR ImageryabstractAircraft detection in synthetic aperture radar (SAR) imagery is a challenging task in SAR automatic target recognition (SAR ATR) areas due to aircraft’s extremely discrete appearance, obvious intraclass variation, small size, and serious background’s interference. In this article, a single shot detector (SSD), namely, attentional feature refinement and alignment network (AFRAN), is proposed for detecting aircraft in SAR images with competitive accuracy and speed. Specifically, three significant components, including attention feature fusion module (AFFM), deformable lateral connection module (DLCM), and anchor-guided detection module (ADM), are carefully designed in our method for refining and aligning informative characteristics of aircraft. To represent the characteristics of aircraft with less interference, low-level textural and high-level semantic features of aircraft are fused and refined in AFFM thoroughly. The alignment between aircraft’s discrete backscatting points and convolutional sampling spots is promoted in DLCM. Eventually, the locations of aircraft are predicted precisely in ADM based on aligned features revised by refined anchors. To evaluate the performance of our method, a self-built SAR aircraft sliced dataset and a large scene SAR image are collected. Extensive quantitative and qualitative experiments with detailed analysis illustrate the effectiveness of the three proposed components. Furthermore, the topmost detection accuracy and competitive speed are achieved by our method compared with other domain-specific methods, e.g., dense attention pyramid network (DAPN) and pyramid attention dilated network (PADN), and general convolutional neural network (CNN)-based methods, e.g., Feature Pyramid Network (FPN), Cascade R-CNN, SSD, RefineDet, and RepPoints Detector (RPDet). Yan Zhao 0026, Lingjun Zhao, Zhong Liu 0002, Dewen Hu, Gangyao Kuang, Li Liu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Temporal Self-Ensembling Teacher for Semi-Supervised Object DetectionabstractThis paper focuses on the semi-supervised object detection (SSOD) which makes good use of unlabeled data to boost performance. We face the following obstacles when adapting the knowledge distillation (KD) framework in SSOD. (1) The teacher model serves a dual role as a teacher and a student, such that the teacher predictions on unlabeled images may limit the upper bound of the student. (2) The data imbalance issue caused by the large quantity of consistent predictions between the teacher and student hinders an efficient knowledge transfer between them. To mitigate these issues, we propose a novel SSOD model called Temporal Self-Ensembling Teacher (TSET). Our teacher model ensembles its temporal predictions for unlabeled images under stochastic perturbations. Then, our teacher model ensembles its model weights with those of the student model by an exponential moving average. These ensembling strategies ensure data and model diversity, and lead to better teacher predictions for unlabeled images. In addition, we adapt the focal loss to formulate the consistency loss for handling the data imbalance issue. Together with a thresholding method, the focal loss automatically reweights the inconsistent predictions, which preserves the knowledge for difficult objects to detect in the unlabeled images. The mAP of our model reaches 80.73% and 40.52% on the VOC2007 test set and the COCO2014minival5kset, respectively, and outperforms a strong fully supervised detector by 2.37% and 1.49%, respectively. Furthermore, the mAP of our model (80.73%) sets a new state-of-the-art performance in SSOD on the VOC2007 test set. Shouyang Dong, Kunlin Cao, Li Liu 0002, Yuanhao Guo |
IEEE Trans. Multim. | 5 |
| 2022 | Feature Estimations Based Correlation Distillation for Incremental Image RetrievalabstractDeep learning for fine-grained image retrieval in an incremental context is less investigated. In this paper, we explore this task to realize the model’s continuous retrieval ability. That means, the model enables to perform well on new incoming data and reduce forgetting of the knowledge learned on preceding old tasks. For this purpose, we distill semantic correlations knowledge among the representations extracted from the new data only so as to regularize the parameters updates using the teacher-student framework. In particular, for the case of learning multiple tasks sequentially, aside from the correlations distilled from the penultimate model, we estimate the representations for all prior models and further their semantic correlations by using the representations extracted from the new data. To this end, the estimated correlations are used as an additional regularization and further prevent catastrophic forgetting over all previous tasks, and it is unnecessary to save the stream of models trained on these tasks. Extensive experiments demonstrate that the proposed method performs favorably for retaining performance on the already-trained old tasks and achieving good accuracy on the current task when new data are added at once or sequentially. Wei Chen 0072, Yu Liu 0012, Nan Pu, Weiping Wang 0002, Li Liu 0002, Michael S. Lew |
IEEE Trans. Multim. | 5 |
| 2022 | Informative Feature Disentanglement for Unsupervised Domain AdaptationabstractUnsupervised Domain Adaptation (UDA) aims at learning a classifier for an unlabeled target domain by transferring knowledge from a labeled source domain with a related but different distribution. The strategy of aligning the two domains in latent feature space via metric discrepancy or adversarial learning has achieved considerable progress. However, these existing approaches mainly focus on adapting the entire image and ignore the bottleneck that occurs when forced adaptation of uninformative domain-specific variations undermines the effectiveness of learned features. To address this problem, we propose a novel component called Informative Feature Disentanglement (IFD), which is equipped with the adversarial network or the metric discrepancy model, respectively. Accordingly, the new network architectures, named IFDAN and IFDMN, enable informative feature refinement before the adaptation. The proposed IFD is designed to disentangle informative features from the uninformative domain-specific variations, which are produced by a Variational Autoencoder (VAE) with lateral connections from the encoder to the decoder. We cooperatively apply the IFD to conduct supervised disentanglement for the source domain and unsupervised disentanglement for the target domain. In this way, informative features are disentangled from the domain-specific details before the adaptation. Extensive experimental results on three gold-standard domain adaptation datasets, e.g., Office31, Office-Home and VisDA-C, demonstrate the effectiveness of the proposed IFDAN and IFDMN models for UDA. Wanxia Deng, Lingjun Zhao, Qing Liao 0001, Deke Guo, Gangyao Kuang, Dewen Hu, Matti Pietikäinen, Li Liu 0002 |
IEEE Trans. Multim. | 8 |
| 2022 | Adaptive Semantic-Spatio-Temporal Graph Convolutional Network for Lip ReadingabstractThe goal of this work is to recognize words, phrases, and sentences being spoken by a talking face without given the audio. Current deep learning approaches for lip reading focus on exploring the appearance and optical flow information of videos. However, these methods do not fully exploit the characteristics of lip motion. In addition to appearance and optical flow, the mouth contour deformation usually conveys significant information that is complementary to others. However, the modeling of dynamic mouth contour has received little attention than that of appearance and optical flow. In this work, we propose a novel model of dynamic mouth contours called Adaptive Semantic-Spatio-Temporal Graph Convolution Network (ASST-GCN), to go beyond previous methods by automatically learning both the spatial and temporal information from videos. To combine the complementary information from appearance and mouth contour, a two-stream visual front-end network is proposed. Experimental results demonstrate that the proposed method significantly outperforms the state-of-the-art lip reading methods on several large-scale lip reading benchmarks. Changchong Sheng, Xinzhong Zhu, Matti Pietikäinen, Li Liu 0002 |
IEEE Trans. Multim. | 5 |
| 2021 | Semi-Supervised Few-Shot Object Detection with a Teacher-Student Network
Wuti Xiong, Yawen Cui, Li Liu 0002 |
BMVC | 3 |
| 2021 | Median Pixel Difference Convolutional Network for Robust Face Recognition
Zhuo Su 0002, Li Liu 0002 |
BMVC | 3 |
| 2021 | Pixel Difference Networks for Efficient Edge DetectionabstractRecently, deep Convolutional Neural Networks (CNNs) can achieve human-level performance in edge detection with the rich and abstract edge representation capacities. However, the high performance of CNN based edge detection is achieved with a large pretrained CNN backbone, which is memory and energy consuming. In addition, it is surprising that the previous wisdom from the traditional edge detectors, such as Canny, Sobel, and LBP are rarely investigated in the rapid-developing deep learning era. To address these issues, we propose a simple, lightweight yet effective architecture named Pixel Difference Network (PiDiNet) for efficient edge detection. PiDiNet adopts novel pixel difference convolutions that integrate the traditional edge detection operators into the popular convolutional operations in modern CNNs for enhanced performance on the task, which enjoys the best of both worlds. Extensive experiments on BSDS500, NYUD, and Multicue are provided to demonstrate its effectiveness, and its high training and inference efficiency. Surprisingly, when training from scratch with only the BSDS500 and VOC datasets, PiDiNet can surpass the recorded result of human perception (0.807 vs. 0.803 in ODS F-measure) on the BSDS500 dataset with 100 FPS and less than 1M parameters. A faster version of PiDiNet with less than 0.1M parameters can still achieve comparable performance among state of the arts with 200 FPS. Results on the NYUD and Multicue datasets show similar observations. The codes are available at https://github.com/zhuoinoulu/pidinet. Zhuo Su 0002, Zitong Yu, Dewen Hu, Qing Liao 0001, Qi Tian 0001, Matti Pietikäinen, Li Liu 0002 |
ICCV | 8 |
| 2021 | One-pass Multi-view Clustering for Large-scale DataabstractExisting non-negative matrix factorization based multi-view clustering algorithms compute multiple coefficient matrices respect to different data views, and learn a common consensus concurrently. The final partition is always obtained from the consensus with classical clustering techniques, such as k-means. However, the non-negativity constraint prevents from obtaining a more discriminative embedding. Meanwhile, this two-step procedure fails to unify multi-view matrix factorization with partition generation closely, resulting in unpromising performance. Therefore, we propose an one-pass multi-view clustering algorithm by removing the non-negativity constraint and jointly optimize the aforementioned two steps. In this way, the generated partition can guide multi-view matrix factorization to produce more purposive coefficient matrix which, as a feedback, improves the quality of partition. To solve the resultant optimization problem, we design an alternate strategy which is guaranteed to be convergent theoretically. Moreover, the proposed algorithm is free of parameter and of linear complexity, making it practical in applications. In addition, the proposed algorithm is compared with recent advances in literature on benchmarks, demonstrating its effectiveness, superiority and efficiency. Jiyuan Liu 0003, Xinwang Liu 0002, Yuexiang Yang, Li Liu 0002, Siqi Wang 0001, Weixuan Liang, Jiangyong Shi |
ICCV | 4 |
| 2021 | Localized Simple Multiple Kernel K-meansabstractAs a representative of multiple kernel clustering (MKC), simple multiple kernel k-means (SimpleMKKM) is recently put forward to boosting the clustering performance by optimally fusing a group of pre-specified kernel matrices. Despite achieving significant improvement in a variety of applications, we find out that SimpleMKKM could indiscriminately force all sample pairs to be equally aligned with the same ideal similarity. As a result, it does not sufficiently take the variation of samples into consideration, leading to unsatisfying clustering performance. To address these issues, this paper proposes a novel MKC algorithm with a "local" kernel alignment, which only requires that the similarity of a sample to its k-nearest neighbours be aligned with the ideal similarity matrix. Such an alignment helps the clustering algorithm to focus on closer sample pairs that shall stay together and avoids involving unreliable similarity evaluation for farther sample pairs. After that, we theoretically show that the objective of SimpleMKKM is a special case of this local kernel alignment criterion with normalizing each base kernel matrix. Based on this observation, the proposed localized SimpleMKKM can be readily implemented by existing SimpleMKKM package. Moreover, we conduct extensive experiments on several widely used benchmark datasets to evaluate the clustering performance of localized SimpleMKKM. The experimental results have demonstrated that our algorithm consistently outperforms the state-of-the-art ones, verifying the effectiveness of the proposed local kernel alignment criterion. The code of Localized SimpleMKKM is publicly available at: https://github.com/xinwangliu/LocalizedSMKKM. Xinwang Liu 0002, Sihang Zhou 0001, Li Liu 0002, Chang Tang, Siwei Wang 0001, Jiyuan Liu 0003, Yi Zhang 0104 |
ICCV | 3 |
| 2021 | Semi-Supervised Few-Shot Class-Incremental LearningabstractThe capability of incrementally learning new classes and learning from a few examples is one of the hallmarks of human intelligence. It is crucial to endow a practical recognition system with such ability. Therefore, in this paper, we conduct pioneering work and focus on a challenging yet practical Semi-Supervised Few-Shot Class-Incremental Learning (SSFSCIL) problem, which requires CNN models incrementally learn new classes from very few labeled samples and a large number of unlabeled samples, without forgetting the previously learned ones. To address this problem, a simple and efficient solution for SSFSCIL is proposed to learn novel categories using a self-training strategy in a semi-supervised manner and avoid catastrophic forgetting by distillation-based methods. Our extensive experiments on CIFAR100, mini ImageNet and CUB200 datasets demonstrate the promising performance of our proposed method, and define baselines in this new research direction. Yawen Cui, Wuti Xiong, Mohammad Tavakolian, Li Liu 0002 |
ICIP | 4 |
| 2021 | Transferable Discriminative Feature Mining For Unsupervised Domain AdaptationabstractUnsupervised Domain Adaptation (UDA) aims to seek an effective model for unlabeled target domain by leveraging knowledge from a labeled source domain with a related but different distribution. Many existing approaches ignore the underlying discriminative features of the target data and the discrepancy of conditional distributions. To address these two issues simultaneously, the paper presents a Transferable Discriminative Feature Mining (TDFM) approach for UDA, which can naturally unify the mining of domain-invariant discriminative features and the alignment of class-wise features into one single framework. To be specific, to achieve the domain-invariant discriminative features, TDFM jointly learns a shared encoding representation for two tasks: supervised classification of labeled source data, and discriminative clustering of unlabeled target data. It then conducts the class-wise alignment by decreasing intra-class variations and increasing inter-class differences across domains, encouraging the emergence of transferable discriminative features. When combined, these two procedures are mutually beneficial. Comprehensive experiments verify that TDFM can obtain remarkable margins over state-of-the-art domain adaptation methods. Lingjun Zhao, Wanxia Deng, Gangyao Kuang, Dewen Hu, Li Liu 0002 |
ICIP | 5 |
| 2021 | One Pass Late Fusion Multi-view ClusteringabstractExisting late fusion multi-view clustering (LFMVC) optimally integrates a group of pre-specified base partition matrices to learn a consensus one. It is then taken as the input of the widely used k-means to generate the cluster labels. As observed, the learning of the consensus partition matrix and the generation of cluster labels are separately done. These two procedures lack necessary negotiation and can not best serve for each other, which may adversely affect the clustering performance. To address this issue, we propose to unify the aforementioned two learning procedures into a single optimization, in which the consensus partition matrix can better serve for the generation of cluster labels, and the latter is able to guide the learning of the former. To optimize the resultant optimization problem, we develop a four-step alternate algorithm with proved convergence. We theoretically analyze the clustering generalization error of the proposed algorithm on unseen data. Comprehensive experiments on multiple benchmark datasets demonstrate the superiority of our algorithm in terms of both clustering accuracy and computational efficiency. It is expected that the simplicity and effectiveness of our algorithm will make it a good option to be considered for practical multi-view clustering applications. Xinwang Liu 0002, Li Liu 0002, Qing Liao 0001, Siwei Wang 0001, Yi Zhang 0104, Wenxuan Tu, Chang Tang, Jiyuan Liu 0003, En Zhu |
ICML | 2 |
| 2021 | Informative Class-Conditioned Feature Alignment for Unsupervised Domain AdaptationabstractThe goal of unsupervised domain adaptation is to learn a task classifier that performs well for the unlabeled target domain by borrowing rich knowledge from a well-labeled source domain. Although remarkable breakthroughs have been achieved in learning transferable representation across domains, two bottlenecks remain to be further explored. First, many existing approaches focus primarily on the adaptation of the entire image, ignoring the limitation that not all features are transferable and informative for the object classification task. Second, the features of the two domains are typically aligned without considering the class labels; this can lead the resulting representations to be domain-invariant but non-discriminative to the category. To overcome the two issues, we present a novel Informative Class-Conditioned Feature Alignment (IC2FA) approach for UDA, which utilizes a twofold method: informative feature disentanglement and class-conditioned feature alignment, designed to address the above two challenges, respectively. More specifically, to surmount the first drawback, we cooperatively disentangle the two domains to obtain informative transferable features; here, Variational Information Bottleneck (VIB) is employed to encourage the learning of task-related semantic representations and suppress task-unrelated information. With regard to the second bottleneck, we optimize a new metric, termed Conditional Sliced Wasserstein Distance (CSWD), which explicitly estimates the intra-class discrepancy and the inter-class margin. The intra-class and inter-class CSWDs are minimized and maximized, respectively, to yield the domain-invariant discriminative features. IC2FA equips class-conditioned feature alignment with informative feature disentanglement and causes the two procedures to work cooperatively, which facilitates informative discriminative features adaptation. Extensive experimental results on three domain adaptation datasets confirm the superiority of IC2FA. Wanxia Deng, Yawen Cui, Zhen Liu 0004, Gangyao Kuang, Dewen Hu, Matti Pietikäinen, Li Liu 0002 |
ACM Multimedia | 7 |
| 2021 | Cross-modal Self-Supervised Learning for Lip Reading: When Contrastive Learning meets Adversarial TrainingabstractThe goal of this work is to learn discriminative visual representations for lip reading without access to manual text annotation. Recent advances in cross-modal self-supervised learning have shown that the corresponding audio can serve as a supervisory signal to learn effective visual representations for lip reading. However, existing methods only exploit the natural synchronization of the video and the corresponding audio. We find that both video and audio are actually composed of speech-related information, identity-related information, and modal information. To make the visual representations (i) more discriminative for lip reading and (ii) indiscriminate with respect to the identities and modals, we propose a novel self-supervised learning framework called Adversarial Dual-Contrast Self-Supervised Learning (ADC-SSL), to go beyond previous methods by explicitly forcing the visual representations disentangled from speech-unrelated information. Experimental results clearly show that the proposed method outperforms state-of-the-art cross-modal self-supervised baselines by a large margin. Besides, ADC-SSL can outperform its supervised counterpart without any finetune. Changchong Sheng, Matti Pietikäinen, Qi Tian 0001, Li Liu 0002 |
ACM Multimedia | 4 |
| 2021 | New Ideas and Trends in Deep Multimodal Content Understanding: A ReviewabstractThe focus of this survey is on the analysis of two modalities of multimodal deep learning: image and text. Unlike classic reviews of deep learning where monomodal image classifiers such as VGG, ResNet and Inception module are central topics, this paper will examine recent multimodal deep models and structures, including auto-encoders, generative adversarial nets and their variants. These models go beyond the simple image classifiers in which they can do uni-directional (e.g. image captioning, image generation) and bi-directional (e.g. cross-modal retrieval, visual question answering) multimodal tasks. Besides, we analyze two aspects of the challenge in terms of better content understanding in deep multimodal applications. We then introduce current ideas and trends in deep multimodal feature learning, such as feature embedding approaches and objective function design, which are crucial in overcoming the aforementioned challenges. Finally, we include several promising directions for future research. Wei Chen 0072, Weiping Wang 0002, Li Liu 0002, Michael S. Lew |
Neurocomputing | 3 |
| 2021 | Robust semi-supervised classification based on data augmented online ELMs with deep features
Xiaochang Hu, Yujun Zeng, Xin Xu 0001, Sihang Zhou 0001, Li Liu 0002 |
Knowl. Based Syst. | 5 |
| 2021 | Deep Learning for 3D Point Clouds: A SurveyabstractPoint cloud learning has lately attracted increasing attention due to its wide applications in many areas, such as computer vision, autonomous driving, and robotics. As a dominating technique in AI, deep learning has been successfully used to solve various 2D vision problems. However, deep learning on point clouds is still in its infancy due to the unique challenges faced by the processing of point clouds with deep neural networks. Recently, deep learning on point clouds has become even thriving, with numerous methods being proposed to address different problems in this area. To stimulate future research, this paper presents a comprehensive review of recent progress in deep learning methods for point clouds. It covers three major tasks, including 3D shape classification, 3D object detection and tracking, and 3D point cloud segmentation. It also presents comparative results on several publicly available datasets, together with insightful observations and inspiring future research directions. Yulan Guo, Hanyun Wang, Qingyong Hu, Hao Liu 0061, Li Liu 0002, Mohammed Bennamoun |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | Efficient and Effective Regularized Incomplete Multi-View ClusteringabstractIncomplete multi-view clustering (IMVC) optimally combines multiple pre-specified incomplete views to improve clustering performance. Among various excellent solutions, the recently proposed multiple kernel k-means with incomplete kernels (MKKM-IK) forms a benchmark, which redefines IMVC as a joint optimization problem where the clustering and kernel matrix imputation tasks are alternately performed until convergence. Though demonstrating promising performance in various applications, we observe that the manner of kernel matrix imputation in MKKM-IK would incur intensive computational and storage complexities, over-complicated optimization and limitedly improved clustering performance. In this paper, we first propose an Efficient and Effective Incomplete Multi-view Clustering (EE-IMVC) algorithm to address these issues. Instead of completing the incomplete kernel matrices, EE-IMVC proposes to impute each incomplete base matrix generated by incomplete views with a learned consensus clustering matrix. Moreover, we further improve this algorithm by incorporating prior knowledge to regularize the learned consensus clustering matrix. Two three-step iterative algorithms are carefully developed to solve the resultant optimization problems with linear computational complexity, and their convergence is theoretically proven. After that, we theoretically study the generalization bound of the proposed algorithms. Furthermore, we conduct comprehensive experiments to study the proposed algorithms in terms of clustering accuracy, evolution of the learned consensus clustering matrix and the convergence. As indicated, our algorithms deliver their effectiveness by significantly and consistently outperforming some state-of-the-art ones. Xinwang Liu 0002, Miaomiao Li 0001, Chang Tang, Jingyuan Xia, Jian Xiong 0002, Li Liu 0002, Marius Kloft, En Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2021 | Deep ladder reconstruction-classification network for unsupervised domain adaptation
Wanxia Deng, Zhuo Su 0002, Qiang Qiu 0001, Lingjun Zhao, Gangyao Kuang, Matti Pietikäinen, Huaxin Xiao, Li Liu 0002 |
Pattern Recognit. Lett. | 8 |
| 2021 | A Graphical Social Topology Model for RGB-D Multi-Person TrackingabstractTracking multiple persons is a challenging task especially when persons move in groups and occlude one another. Existing research have investigated the problems of group division and segmentation; however, lacking overall person-group topology modeling limits the ability to handle complex person and group dynamics. We propose a Graphical Social Topology (GST) model in the RGB-D data domain, and estimate object group dynamics by jointly modeling the group structure and states of persons using RGB-D topological representation. With our topology representation, moving persons are not only assigned to groups, but also dynamically connected with each other, which enables in-group individuals to be correctively associated and the cohesion of each group to be precisely modeled. Using the learned typical topology pattern and group online update modules, we infer the birth/death and merging/splitting of dynamic groups. With the GST model, the proposed multi-person tracker can naturally facilitate the occlusion problem by treating the occluded object and other in-group members as a whole, while leveraging overall state transition. Experiments on different RGB-D and RGB datasets confirm that the proposed multi-person tracker improves the state-of-the-arts. Shan Gao 0003, Qixiang Ye, Li Liu 0002, Arjan Kuijper, Xiangyang Ji |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Fragmentary Multi-Instance ClassificationabstractMulti-instance learning (MIL) has been extensively applied to various real tasks involving objects with bags of instances, such as in drugs and images. Previous studies on MIL assume that data are entirely complete. However, in many real tasks, the instance is fragmentary. In this article, we present probably the first study on multi-instance classification with fragmentary data. In our proposed framework, called fragmentary multi-instance classification (FIC), the fragmentary data are completed and the multi-instance classifier is learned jointly. To facilitate the integration between the completion and classifier learning, FIC establishes the weighting mechanism to measure the importance levels of different instances. To validate the compatibility of our framework, four typical MIL methods, including multi-instance support vector machine (MI-SVM), expectation maximization diverse density (EM-DD), citation- K nearest neighbors (Citation-KNNs), and MIL with discriminative bag mapping (MILDM), are embedded into the framework to obtain the corresponding FIC versions. As an illustration, an efficient solving algorithm is developed to address the problem for MI-SVM, together with the proof of convergence behavior. The experimental results on various types of real-world datasets demonstrate the effectiveness. Wenzhang Zhuge, Xinwang Liu 0002, Li Liu 0002, Chenping Hou |
IEEE Trans. Cybern. | 4 |
| 2021 | Semi-Supervised Natural Face De-OcclusionabstractOcclusions are often present in face images in the wild, e.g., under video surveillance and forensic scenarios. Existing face de-occlusion methods are limited as they require the knowledge of an occlusion mask. To overcome this limitation, we propose in this paper a new generative adversarial network (named OA-GAN) for natural face de-occlusion without an occlusion mask, enabled by learning in a semi-supervised fashion using (i) paired images with known masks of artificial occlusions and (ii) natural images without occlusion masks. The generator of our approach first predicts an occlusion mask, which is used for filtering the feature maps of the input image as a semantic cue for de-occlusion. The filtered feature maps are then used for face completion to recover a non-occluded face image. The initial occlusion mask prediction might not be accurate enough, but it gradually converges to the accurate one because of the adversarial loss we use to perceive which regions in a face image need to be recovered. The discriminator of our approach consists of an adversarial loss, distinguishing the recovered face images from natural face images, and an attribute preserving loss, ensuring that the face image after de-occlusion can retain the attributes of the input face image. Experimental evaluations on the widely used CelebA dataset and a dataset with natural occlusions we collected show that the proposed approach can outperform the state of the art methods in natural face de-occlusion. Jiancheng Cai, Hu Han 0001, Jiyun Cui, Jie Chen 0001, Li Liu 0002, Shaohua Kevin Zhou |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2021 | Joint Clustering and Discriminative Feature Alignment for Unsupervised Domain AdaptationabstractUnsupervised Domain Adaptation (UDA) aims to learn a classifier for the unlabeled target domain by leveraging knowledge from a labeled source domain with a different but related distribution. Many existing approaches typically learn a domain-invariant representation space by directly matching the marginal distributions of the two domains. However, they ignore exploring the underlying discriminative features of the target data and align the cross-domain discriminative features, which may lead to suboptimal performance. To tackle these two issues simultaneously, this paper presents a Joint Clustering and Discriminative Feature Alignment (JCDFA) approach for UDA, which is capable of naturally unifying the mining of discriminative features and the alignment of class-discriminative features into one single framework. Specifically, in order to mine the intrinsic discriminative information of the unlabeled target data, JCDFA jointly learns a shared encoding representation for two tasks: supervised classification of labeled source data, and discriminative clustering of unlabeled target data, where the classification of the source domain can guide the clustering learning of the target domain to locate the object category. We then conduct the cross-domain discriminative feature alignment by separately optimizing two new metrics: 1) an extended supervised contrastive learning, i.e., semi-supervised contrastive learning 2) an extended Maximum Mean Discrepancy (MMD), i.e., conditional MMD, explicitly minimizing the intra-class dispersion and maximizing the inter-class compactness. When these two procedures, i.e., discriminative features mining and alignment are integrated into one framework, they tend to benefit from each other to enhance the final performance from a cooperative learning perspective. Experiments are conducted on four real-world benchmarks (e.g., Office-31, ImageCLEF-DA, Office-Home and VisDA-C). All the results demonstrate that our JCDFA can obtain remarkable margins over state-of-the-art domain adaptation methods. Comprehensive ablation studies also verify the importance of each key component of our proposed algorithm and the effectiveness of combining two learning strategies into a framework. Wanxia Deng, Qing Liao 0001, Lingjun Zhao, Deke Guo, Gangyao Kuang, Dewen Hu, Li Liu 0002 |
IEEE Trans. Image Process. | 7 |
| 2020 | Dynamic Group Convolution for Accelerating Convolutional Neural Networks
Zhuo Su 0002, Linpu Fang, Wenxiong Kang, Dewen Hu, Matti Pietikäinen, Li Liu 0002 |
ECCV (6) | 6 |
| 2020 | JGR-P2O: Joint Graph Reasoning Based Pixel-to-Offset Prediction Network for 3D Hand Pose Estimation from a Single Depth Image
Linpu Fang, Xingyan Liu, Li Liu 0002, Wenxiong Kang |
ECCV (6) | 3 |
| 2020 | Deep Learning for Generic Object Detection: A SurveyabstractAbstract Object detection, one of the most fundamental and challenging problems in computer vision, seeks to locate object instances from a large number of predefined categories in natural images. Deep learning techniques have emerged as a powerful strategy for learning feature representations directly from data and have led to remarkable breakthroughs in the field of generic object detection. Given this period of rapid evolution, the goal of this paper is to provide a comprehensive survey of the recent achievements in this field brought about by deep learning techniques. More than 300 research contributions are included in this survey, covering many aspects of generic object detection: detection frameworks, object feature representation, object proposal generation, context modeling, training strategies, and evaluation metrics. We finish the survey by identifying promising directions for future research. Li Liu 0002, Wanli Ouyang, Xiaogang Wang 0001, Paul W. Fieguth, Jie Chen 0001, Xinwang Liu 0002, Matti Pietikäinen |
Int. J. Comput. Vis. | 1 |
| 2020 | Efficient Visual Recognition
Li Liu 0002, Matti Pietikäinen, Jie Qin 0004, Wanli Ouyang, Luc Van Gool |
Int. J. Comput. Vis. | 1 |
| 2020 | Absent Multiple Kernel Learning AlgorithmsabstractMultiple kernel learning (MKL) has been intensively studied during the past decade. It optimally combines the multiple channels of each sample to improve classification performance. However, existing MKL algorithms cannot effectively handle the situation where some channels of the samples are missing, which is not uncommon in practical applications. This paper proposes three absent MKL (AMKL) algorithms to address this issue. Different from existing approaches where missing channels are first imputed and then a standard MKL algorithm is deployed on the imputed data, our algorithms directly classify each sample based on its observed channels, without performing imputation. Specifically, we define a margin for each sample in its own relevant space, a space corresponding to the observed channels of that sample. The proposed AMKL algorithms then maximize the minimum of all sample-based margins, and this leads to a difficult optimization problem. We first provide two two-step iterative algorithms to approximately solve this problem. After that, we show that this problem can be reformulated as a convex one by applying the representer theorem. This makes it readily be solved via existing convex optimization packages. In addition, we provide a generalization error bound to justify the proposed AMKL algorithms from a theoretical perspective. Extensive experiments are conducted on nine UCI and six MKL benchmark datasets to compare the proposed algorithms with existing imputation-based methods. As demonstrated, our algorithms achieve superior performance and the improvement is more significant with the increase of missing ratio. Xinwang Liu 0002, Lei Wang 0001, Xinzhong Zhu, Miaomiao Li 0001, En Zhu, Tongliang Liu, Li Liu 0002, Yong Dou, Jianping Yin |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2020 | Self-supervised pain intensity estimation from facial videos via statistical spatiotemporal distillation
Mohammad Tavakolian, Miguel Bordallo López, Li Liu 0002 |
Pattern Recognit. Lett. | 3 |
| 2020 | Deep Video Super-Resolution Using HR Optical Flow EstimationabstractVideo super-resolution (SR) aims at generating a sequence of high-resolution (HR) frames with plausible and temporally consistent details from their low-resolution (LR) counterparts. The key challenge for video SR lies in the effective exploitation of temporal dependency between consecutive frames. Existing deep learning based methods commonly estimate optical flows between LR frames to provide temporal dependency. However, the resolution conflict between LR optical flows and HR outputs hinders the recovery of fine details. In this paper, we propose an end-to-end video SR network to super-resolve both optical flows and images. Optical flow SR from LR frames provides accurate temporal dependency and ultimately improves video SR performance. Specifically, we first propose an optical flow reconstruction network (OFRnet) to infer HR optical flows in a coarse-to-fine manner. Then, motion compensation is performed using HR optical flows to encode temporal dependency. Finally, compensated LR inputs are fed to a super-resolution network (SRnet) to generate SR results. Extensive experiments have been conducted to demonstrate the effectiveness of HR optical flows for SR performance improvement. Comparative results on the Vid4 and DAVIS-10 datasets show that our network achieves the state-of-the-art performance. Longguang Wang, Yulan Guo, Li Liu 0002, Zaiping Lin, Xinpu Deng, Wei An 0003 |
IEEE Trans. Image Process. | 3 |
| 2020 | Multiple Kernel Clustering With Neighbor-Kernel Subspace SegmentationabstractMultiple kernel clustering (MKC) has been intensively studied during the last few decades. Even though they demonstrate promising clustering performance in various applications, existing MKC algorithms do not sufficiently consider the intrinsic neighborhood structure among base kernels, which could adversely affect the clustering performance. In this paper, we propose a simple yet effective neighbor-kernel-based MKC algorithm to address this issue. Specifically, we first define a neighbor kernel, which can be utilized to preserve the block diagonal structure and strengthen the robustness against noise and outliers among base kernels. After that, we linearly combine these base neighbor kernels to extract a consensus affinity matrix through an exact-rank-constrained subspace segmentation. The naturally possessed block diagonal structure of neighbor kernels better serves the subsequent subspace segmentation, and in turn, the extracted shared structure is further refined through subspace segmentation based on the combined neighbor kernels. In this manner, the above two learning processes can be seamlessly coupled and negotiate with each other to achieve better clustering. Furthermore, we carefully design an efficient iterative optimization algorithm with proven convergence to address the resultant optimization problem. As a by-product, we reveal an interesting insight into the exact-rank constraint in ridge regression by careful theoretical analysis: it back-projects the solution of the unconstrained counterpart to its principal components. Comprehensive experiments have been conducted on several benchmark data sets, and the results demonstrate the effectiveness of the proposed algorithm. Sihang Zhou 0001, Xinwang Liu 0002, Miaomiao Li 0001, En Zhu, Li Liu 0002, Changwang Zhang, Jianping Yin |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2019 | BIRD: Learning Binary and Illumination Robust Descriptor for Face Recognition
Zhuo Su 0002, Matti Pietikäinen, Li Liu 0002 |
BMVC | 3 |
| 2019 | Sparse Subspace Clustering for Evolving Data StreamsabstractThe data streams arising in many applications can be modeled as a union of low-dimensional subspaces known as multi-subspace data streams (MSDSs). Clustering MSDSs according to their underlying low-dimensional subspaces is a challenging problem which has not been resolved satisfactorily by existing data stream clustering (DSC) algorithms. In this paper, we propose a sparse-based DSC algorithm, which we refer to as dynamic sparse subspace clustering (D-SSC). This algorithm recovers the low-dimensional subspaces (structures) of high-dimensional data streams and finds an explicit assignment of points to subspaces in an online manner. Moreover, as an online algorithm, D-SSC is able to cope with the time-varying structure of MSDSs. The effectiveness of D-SSC is evaluated using numerical experiments. Jinping Sui, Zhen Liu 0004, Li Liu 0002, Alexander Jung 0001, Tianpeng Liu, Xiang Li 0014 |
ICASSP | 3 |
| 2019 | Dynamic Texture Recognition Using 3D Random FeaturesabstractIn this paper, we present a novel, simple but effective approach for dynamic texture recognition using 3D random features. Compared with the existing dynamic texture recognition approaches using carefully designed features for high performance, our method use only a few 3D random filters to extract spatio-temporal features from local dynamic texture blocks, which are further encoded into a low-dimensional feature vector. To explore the representative power of the 3D random features, we use two different encoding schemes, the learning-based Fisher vector encoding and the learning-free binary encoding. The proposed method is tested on the UCLA and DynTex databases with various evaluation protocols. Experimental results demonstrate the high performance of our method for dynamic texture recognition. Xiaochao Zhao, Yaping Lin, Li Liu 0002 |
ICASSP | 3 |
| 2019 | From BoW to CNN: Two Decades of Texture Representation for Texture ClassificationabstractTexture is a fundamental characteristic of many types of images, and texture representation is one of the essential and challenging problems in computer vision and pattern recognition which has attracted extensive research attention over several decades. Since 2000, texture representations based on Bag of Words and on Convolutional Neural Networks have been extensively studied with impressive performance. Given this period of remarkable evolution, this paper aims to present a comprehensive survey of advances in texture representation over the last two decades. More than 250 major publications are cited in this survey covering different aspects of the research, including benchmark datasets and state of the art results. In retrospect of what has been achieved so far, the survey discusses open challenges and directions for future research. Li Liu 0002, Jie Chen 0001, Paul W. Fieguth, Guoying Zhao 0001, Rama Chellappa, Matti Pietikäinen |
Int. J. Comput. Vis. | 1 |
| 2019 | Guest Editors' Introduction to the Special Section on Compact and Efficient Feature Representation and Learning in Computer VisionabstractThe papers in this special section examine compact and efficient feature representation and learning in computer vision. Li Liu 0002, Matti Pietikäinen, Jie Chen 0001, Guoying Zhao 0001, Xiaogang Wang 0001, Rama Chellappa |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | CycleMatch: A cycle-consistent embedding network for image-text matching
Yu Liu 0012, Yanming Guo, Li Liu 0002, Erwin M. Bakker, Michael S. Lew |
Pattern Recognit. | 3 |
| 2019 | Texture Classification in Extreme Scale Variations Using GANetabstractResearch in texture recognition often concentrates on recognizing textures with intraclass variations, such as illumination, rotation, viewpoint, and small-scale changes. In contrast, in real-world applications, a change in scale can have a dramatic impact on texture appearance to the point of changing completely from one texture category to another. As a result, texture variations due to changes in scale are among the hardest to handle. In this paper, we conduct the first study of classifying textures with extreme variations in scale. To address this issue, we first propose and then reduce scale proposals on the basis of dominant texture patterns. Motivated by the challenges posed by this problem, we propose a new GANet network where we use a genetic algorithm to change the filters in the hidden layers during network training in order to promote the learning of more informative semantic texture patterns. Finally, we adopt a Fisher vector pooling of a convolutional neural network filter bank feature encoder for global texture representation. Because extreme scale variations are not necessarily present in most standard texture databases, to support the proposed extreme-scale aspects of texture understanding, we are developing a new dataset, the extreme scale variation textures (ESVaT), to test the performance of our framework. It is demonstrated that the proposed framework significantly outperforms the gold-standard texture features by more than 10% on ESVaT. We also test the performance of our proposed approach on the KTHTIPS2b and OS datasets and a further dataset synthetically derived from Forrest, showing the superior performance compared with the state-of-the-art. Li Liu 0002, Jie Chen 0001, Guoying Zhao 0001, Paul W. Fieguth, Xilin Chen 0001, Matti Pietikäinen |
IEEE Trans. Image Process. | 1 |
| 2019 | SwapGAN: A Multistage Generative Approach for Person-to-Person Fashion Style TransferabstractFashion style transfer has attracted significant attention because it both has interesting scientific challenges and it is also important to the fashion industry. This paper focuses on addressing a practical problem in fashion style transfer, person-to-person clothing swapping, which aims to visualize what the person would look like with the target clothes worn on another person instead of dressing them physically. This problem remains challenging due to varying pose deformations between different person images. In contrast to traditional nonparametric methods that blend or warp the target clothes for the reference person, in this paper we propose a multistage deep generative approach named SwapGAN that exploits three generators and one discriminator in a unified framework to fulfill the task end-to-end. The first and second generators are conditioned on a human pose map and a segmentation map, respectively, so that we can simultaneously transfer the pose style and the clothes style. In addition, the third generator is used to preserve the human body shape during the image synthesis process. The discriminator needs to distinguish two fake image pairs from the real image pair. The entire SwapGAN is trained by integrating the adversarial loss and the mask-consistency loss. The experimental results on the DeepFashion dataset demonstrate the improvements of SwapGAN over other existing approaches through both quantitative and qualitative evaluations. Moreover, we conduct ablation studies on SwapGAN and provide a detailed analysis about its effectiveness. Yu Liu 0012, Wei Chen 0072, Li Liu 0002, Michael S. Lew |
IEEE Trans. Multim. | 3 |
| 2019 | Dynamic Texture Classification Using Unsupervised 3D Filter Learning and Local Binary EncodingabstractLocal binary descriptors, such as local binary pattern (LBP) and its various variants, have been studied extensively in texture and dynamic texture analysis due to their outstanding characteristics, such as grayscale invariance, low computational complexity and good discriminability. Most existing local binary feature extraction methods extract spatio-temporal features from three orthogonal planes of a spatio-temporal volume by viewing a dynamic texture in 3D space. For a given pixel in a video, only a proportion of its surrounding pixels is incorporated in the local binary feature extraction process. We argue that the ignored pixels contain discriminative information that should be explored. To fully utilize the information conveyed by all the pixels in a local neighborhood, we propose extracting local binary features from the spatio-temporal domain with 3D filters that are learned in an unsupervised manner so that the discriminative features along both the spatial and temporal dimensions are captured simultaneously. The proposed approach consists of three components: 1) 3D filtering; 2) binary hashing; and 3) joint histogramming. Densely sampled 3D blocks of a dynamic texture are first normalized to have zero mean and are then filtered by 3D filters that are learned in advance. To preserve more of the structure information, the filter response vectors are decomposed into two complementary components, namely, the signs and the magnitudes, which are further encoded separately into binary codes. The local mean pixels of the 3D blocks are also converted into binary codes. Finally, three types of binary codes are combined via joint or hybrid histograms for the final feature representation. Extensive experiments are conducted on three commonly used dynamic texture databases: 1) UCLA; 2) DynTex; and 3) YUVL. The proposed method provides comparable results to, and even outperforms, many state-of-the-art methods. Xiaochao Zhao, Yaping Lin, Li Liu 0002, Janne Heikkilä, Wenming Zheng |
IEEE Trans. Multim. | 3 |
| 2018 | Super Wide Regression Network for Unsupervised Cross-Database Facial Expression RecognitionabstractUnsupervised cross-database facial expression recognition (FER) is a challenging problem, in which the training and testing samples belong to different facial expression databases. For this reason, the training (source) and testing (target) facial expression samples would have different feature distributions and hence the performance of lots of existing FER methods may decrease. To solve this problem, in this paper we propose a novel super wide regression network (SWiRN) model, which serves as the regression parameter to bridge the original feature space and the label space and herein in each layer the maximum mean discrepancy (MMD) criterion is used to enforce the source and target facial expression samples to share the same or similar feature distributions. Consequently, the learned SWiRN is able to predict the expression categories of the target samples although we have no access to any label information of target samples. We conduct extensive cross-database FER experiments on CK+, eNTERFACE, and Oulu-CASIA VIS facial expression databases to evaluate the proposed SWiRN. Experimental results show that our SWiRN model achieves more promising performance than recent proposed cross-database emotion recognition methods. Baofeng Zhang, Yuan Zong, Li Liu 0002, Jie Chen 0001, Guoying Zhao 0001, Junchao Zhu |
ICASSP | 4 |
| 2018 | Unsupervised Cross-Corpus Speech Emotion Recognition Using Domain-Adaptive Subspace LearningabstractIn this paper, we investigate an interesting problem, i.e., unsupervised cross-corpus speech emotion recognition (SER), in which the training and testing speech signals come from two different speech emotion corpora. Meanwhile, the training speech signals are labeled, while the label information of the testing speech signals is entirely unknown. Due to this setting, the training (source) and testing (target) speech signals may have different feature distributions and therefore lots of existing SER methods would not work. To deal with this problem, we propose a domain-adaptive subspace learning (DoSL) method for learning a projection matrix with which we can transform the source and target speech signals from the original feature space to the label space. The transformed source and target speech signals in the label space would have similar feature distributions. Consequently, the classifier learned on the labeled source speech signals can effectively predict the emotional states of the unlabeled target speech signals. To evaluate the performance of the proposed DoSL method, we carry out extensive cross-corpus SER experiments on three speech emotion corpora including EmoDB, eNTERFACE, and AFEW 4.0. Compared with recent state-of-the-art cross-corpus SER methods, the proposed DoSL can achieve more satisfactory overall results. Yuan Zong, Baofeng Zhang, Li Liu 0002, Jie Chen 0001, Guoying Zhao 0001, Junchao Zhu |
ICASSP | 4 |
| 2018 | A Dual Prediction Network for Image CaptioningabstractGeneral captioning practice involves a single forward prediction, with the aim of predicting the word in the next timestep given the word in the current timestep. In this paper, we present a novel captioning framework, namely Dual Prediction Network (DPN), which is end-to-end trainable and addresses the captioning problem with dual predictions. Specifically, the dual predictions consist of a forward prediction to generate the next word from the current input word, as well as a backward prediction to reconstruct the input word using the predicted word. DPN has two appealing properties: 1) By introducing an extra supervision signal on the prediction, DPN can better capture the interplay between the input and the target; 2) Utilizing the reconstructed input, DPN can make another new prediction. During the test phase, we average both predictions to formulate the final target sentence. Experimental results on the MS COCO dataset demonstrate that, benefiting from the reconstruction step, both generated predictions in DPN outperform the predictions of methods based on the general captioning practice (single forward prediction), and averaging them can bring a further accuracy boost. Overall, DPN achieves competitive results with state-of-the-art approaches, across multiple evaluation metrics. Yanming Guo, Yu Liu 0012, Maaike de Boer, Li Liu 0002, Michael S. Lew |
ICME | 4 |
| 2018 | Localized Incomplete Multiple Kernel k-meansabstractThe recently proposed multiple kernel k-means with incomplete kernels (MKKM-IK) optimally integrates a group of pre-specified incomplete kernel matrices to improve clustering performance. Though it demonstrates promising performance in various applications, we observe that it does not \emph{sufficiently consider the local structure among data and indiscriminately forces all pairwise sample similarity to equally align with their ideal similarity values}. This could make the incomplete kernels less effectively imputed, and in turn adversely affect the clustering performance. In this paper, we propose a novel localized incomplete multiple kernel k-means (LI-MKKM) algorithm to address this issue. Different from existing MKKM-IK, LI-MKKM only requires the similarity of a sample to its k-nearest neighbors to align with their ideal similarity values. This helps the clustering algorithm to focus on closer sample pairs that shall stay together and avoids involving unreliable similarity evaluation for farther sample pairs. We carefully design a three-step iterative algorithm to solve the resultant optimization problem and theoretically prove its convergence. Comprehensive experiments on eight benchmark datasets demonstrate that our algorithm significantly outperforms the state-of-the-art comparable algorithms proposed in the recent literature, verifying the advantage of considering local structure. Xinzhong Zhu, Xinwang Liu 0002, Miaomiao Li 0001, En Zhu, Li Liu 0002, Zhiping Cai, Jianping Yin, Wen Gao 0001 |
IJCAI | 5 |
| 2018 | Learning visual and textual representations for multimodal matching and classification
Yu Liu 0012, Li Liu 0002, Yanming Guo, Michael S. Lew |
Pattern Recognit. | 2 |
| 2017 | Robust local features for remote face recognition
Jie Chen 0001, Vishal M. Patel, Li Liu 0002, Vili Kellokumpu, Guoying Zhao 0001, Matti Pietikäinen, Rama Chellappa |
Image Vis. Comput. | 3 |
| 2017 | Local binary features for texture classification: Taxonomy and experimental study
Li Liu 0002, Paul W. Fieguth, Yulan Guo, Xiaogang Wang 0001, Matti Pietikäinen |
Pattern Recognit. | 1 |
| 2016 | Evaluation of LBP and Deep Texture Descriptors with a New Robustness Benchmark
Li Liu 0002, Paul W. Fieguth, Xiaogang Wang 0001, Matti Pietikäinen, Dewen Hu |
ECCV (3) | 1 |
| 2016 | An attention model based on spatial transformers for scene recognitionabstractScene recognition is an important and challenging task in computer vision. We propose an end-to-end pipeline by combing convolutional neural networks (CNNs) with explicit attention model to determine several meaningful regions of original images for scene recognition. In the proposed pipeline, the spatial transformer network is leveraged as the attention module, which can automatically learn the scales and movements of centers of attention windows. As for feature extraction, the basic CNN architecture is utilized. Furthermore, the stronger descriptors of scenes are constructed by feature fusion. The highlight of our proposed network is that it is capable to localize discriminative regions from an image in a data-driven manner without any additional supervision. We conduct experiments on a subset of the Places205 database to evaluate the performance of the proposed basic network and the involved parameters. Our model achieves state-of-the-art top-1 accuracy 82.10% on the evaluation dataset comparing with fine-tuned PlacesCNN (80.98%). We find that our model is able to learn informative attention regions for discriminating scene categories. Shuxuan Guo, Li Liu 0002, Wei Wang 0115, Songyang Lao, Liang Wang 0001 |
ICPR | 2 |
| 2016 | RoLoD: Robust local descriptors for computer vision
Jie Chen 0001, Zhen Lei 0001, Li Liu 0002, Guoying Zhao 0001, Matti Pietikäinen |
Neurocomputing | 3 |
| 2016 | Extended local binary patterns for face recognition
Li Liu 0002, Paul W. Fieguth, Guoying Zhao 0001, Matti Pietikäinen, Dewen Hu |
Inf. Sci. | 1 |
| 2016 | Random projections and Single BoW for fast and Robust texture segmentation
Li Liu 0002, Liansheng Wang 0002, Lingjun Zhao, Paul W. Fieguth |
Inf. Sci. | 1 |
| 2016 | EI3D: Expression-invariant 3D face recognition based on feature and shape matching
Yulan Guo, Yinjie Lei, Li Liu 0002, Yan Wang 0059, Mohammed Bennamoun, Ferdous Sohel |
Pattern Recognit. Lett. | 3 |
| 2016 | Median Robust Extended Local Binary Pattern for Texture ClassificationabstractLocal binary patterns (LBP) are considered among the most computationally efficient high-performance texture features. However, the LBP method is very sensitive to image noise and is unable to capture macrostructure information. To best address these disadvantages, in this paper, we introduce a novel descriptor for texture classification, the median robust extended LBP (MRELBP). Different from the traditional LBP and many LBP variants, MRELBP compares regional image medians rather than raw image intensities. A multiscale LBP type descriptor is computed by efficiently comparing image medians over a novel sampling scheme, which can capture both microstructure and macrostructure texture information. A comprehensive evaluation on benchmark data sets reveals MRELBP's high performance-robust to gray scale variations, rotation changes and noise-but at a low computational cost. MRELBP produces the best classification scores of 99.82%, 99.38%, and 99.77% on three popular Outex test suites. More importantly, MRELBP is shown to be highly robust to image noise, including Gaussian noise, Gaussian blur, salt-and-pepper noise, and random pixel corruption. Li Liu 0002, Songyang Lao, Paul W. Fieguth, Yulan Guo, Xiaogang Wang 0001, Matti Pietikäinen |
IEEE Trans. Image Process. | 1 |
| 2015 | Median robust extended local binary pattern for texture classificationabstractLocal Binary Patterns (LBP) are among the most computationally efficient amongst high-performance texture features. However, LBP is very sensitive to image noise and is unable to capture macrostructure information. To best address these disadvantages, in this paper we introduce a novel descriptor for texture classification, the Median Robust Extended Local Binary Pattern (MRELBP). In contrast to traditional LBP and many LBP variants, MRELBP compares local image medians instead of raw image intensities. We develop a multiscale LBP-type descriptor by efficiently comparing image medians over a novel sampling scheme, which can capture both microstructure and macrostructure. A comprehensive evaluation on benchmark datasets reveals MRELBP's remarkable performance (robust to gray scale variations, rotation changes and noise) relative to state-of-the-art algorithms, but nevertheless at a low computational cost, producing the best classification scores of 99.82%, 99.38% and 99.77% on three popular Outex test suites. Furthermore, MRELBP is also shown to be highly robust to image noise including Gaussian noise, Gaussian blur, Salt-and-Pepper noise and random pixel corruption. Li Liu 0002, Paul W. Fieguth, Matti Pietikäinen, Songyang Lao |
ICIP | 1 |
| 2015 | A novel specific image scenes detection method
Yuxiang Xie, Xiao-Ping Zhang 0002, Xidao Luan, Li Liu 0002, Xin Zhang 0029 |
Multim. Tools Appl. | 4 |
| 2015 | Fusing Sorted Random Projections for Robust Texture and Material ClassificationabstractThis paper presents a conceptually simple, and robust, yet highly effective, approach to both texture classification and material categorization. The proposed system is composed of three components: 1) local, highly discriminative, and robust features based on sorted random projections (RPs), built on the universal and information-preserving properties of RPs; 2) an effective bag-of-words global model; and 3) a novel approach for combining multiple features in a support vector machine classifier. The proposed approach encompasses the simplicity, broad applicability, and efficiency of the three methods. We have tested the proposed approach on eight popular texture databases, including Flickr Materials Database, a highly challenging materials database. We compare our method with 13 recent state-of-the-art methods, and the experimental results show that our texture classification system yields the best classification rates of which we are aware of 99.37% for Columbia-Utrecht, 97.16% for Brodatz, 99.30% for University of Maryland Database, and 99.29% for Kungliga Tekniska högskolan-textures under varying illumination, pose, and scale. Moreover, the proposed approach significantly outperforms the current state-of-the-art approach in materials categorization, with an improvement to classification accuracy of 67%. Li Liu 0002, Paul W. Fieguth, Dewen Hu, Yingmei Wei, Gangyao Kuang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | Extended local binary pattern fusion for face recognitionabstractThis paper presents a simple, novel, yet highly effective approach for robust face recognition. Given LBP-like descriptors based on local accumulated pixel differences, Angular Differences (AD) and Radial Differences (RD), the local differences are decomposed into complementary components of signs and magnitudes. The proposed descriptors have desirable features: (1) robustness to lighting, pose, and expression; (2) computation efficiency; (3) encoding of both microstructures and macrostructures; (4) consistent in form with traditional LBP, thus inheriting the merits of LBP; and (5) no required training, improving generalizability. From a given face image, we obtain six histogram features, each of which is obtained by concatenating spatial histograms extracted from nonoverlapping subregions. The Whitened PCA technique is used for dimensionality reduction, followed by Nearest Neighbor classification. We have evaluated the effectiveness of the proposed method on the Extended Yale B and CAS-PEAL-R1 databases. The proposed method impressively outperforms other well known systems, including what we believe to be the best reported performance for the the CAS-PEAL-R1 lighting probe set with a recognition rate of 72.3%. Li Liu 0002, Paul W. Fieguth, Guoying Zhao 0001, Matti Pietikäinen |
ICIP | 1 |
| 2014 | BRINT: Binary Rotation Invariant and Noise Tolerant Texture ClassificationabstractIn this paper, we propose a simple, efficient, yet robust multiresolution approach to texture classification-binary rotation invariant and noise tolerant (BRINT). The proposed approach is very fast to build, very compact while remaining robust to illumination variations, rotation changes, and noise. We develop a novel and simple strategy to compute a local binary descriptor based on the conventional local binary pattern (LBP) approach, preserving the advantageous characteristics of uniform LBP. Points are sampled in a circular neighborhood, but keeping the number of bins in a single-scale LBP histogram constant and small, such that arbitrarily large circular neighborhoods can be sampled and compactly encoded over a number of scales. There is no necessity to learn a texton dictionary, as in methods based on clustering, and no tuning of parameters is required to deal with different data sets. Extensive experimental results on representative texture databases show that the proposed BRINT not only demonstrates superior performance to a number of recent state-of-the-art LBP variants under normal conditions, but also performs significantly and consistently better in presence of noise due to its high distinctiveness and robustness. This noise robustness characteristic of the proposed BRINT is evaluated quantitatively with different artificially generated types and levels of noise (including Gaussian, salt and pepper, and speckle noise) in natural texture images. Li Liu 0002, Yunli Long, Paul W. Fieguth, Songyang Lao, Guoying Zhao 0001 |
IEEE Trans. Image Process. | 1 |
| 2013 | BRINT: A binary rotation invariant and noise tolerant texture descriptorabstractLocal Binary Pattern (LBP) and its variants are effective and popular descriptors for texture classification. Most LBP like descriptors have disadvantages including sensitiveness to noise and inability to capture long distance texture information. In this paper we propose a simple, efficient, yet robust multi-resolution descriptor to texture classification - Binary Rotation Invariant and Noise Tolerant (BRINT). The proposed descriptor is very fast to build, very compact while remaining robust to illumination variations, rotation changes and noise. We develop a novel and simple strategy - averaging before binarization - to compute a local binary descriptor based on the conventional LBP approach. Points are sampled in a circular neighborhood, but keeping the number of bins in a single-scale LBP histogram constant and small by averaging over several contiguous pixels in the circle. There is no need for pre-training, no texton dictionary, and no tuning of parameters to deal with different datasets. Experiments on the Outex test suite demonstrate that the proposed approach is very robust to noise and significantly outperforms the state-of-the-art in terms of classifying noise corrupted textures. Li Liu 0002, Paul W. Fieguth, Yingmei Wei |
ICIP | 1 |
| 2012 | Extended local binary patterns for texture classification
Li Liu 0002, Lingjun Zhao, Yunli Long, Gangyao Kuang, Paul W. Fieguth |
Image Vis. Comput. | 1 |
| 2012 | Polarimetric SAR Target Detection Using the Reflection SymmetryabstractThis letter addresses the polarimetric synthetic aperture radar target detection using the magnitude of the (2, 3) term in the sample averaged coherency matrix. The theoretical analysis demonstrates that such term reveals the difference between the nonreflection symmetric targets and natural clutters. The statistical models for such term are derived within different degrees of homogeneity. Based on the statistical models, an automatic constant-false-alarm-rate detection scheme is completed. The parameter estimation and the solution for the detection threshold are given in detail. Experimental results demonstrate the capability of the proposed approach for detecting ships, oil stores, buildings, etc., in homogeneous and heterogeneous areas. Na Wang 0002, Gongtao Shi, Li Liu 0002, Lingjun Zhao, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2012 | Texture Classification from Random FeaturesabstractInspired by theories of sparse representation and compressed sensing, this paper presents a simple, novel, yet very powerful approach for texture classification based on random projection, suitable for large texture database applications. At the feature extraction stage, a small set of random features is extracted from local image patches. The random features are embedded into a bag-of-words model to perform texture classification; thus, learning and classification are carried out in a compressed domain. The proposed unconventional random feature extraction is simple, yet by leveraging the sparse nature of texture images, our approach outperforms traditional feature extraction methods which involve careful design and complex steps. We have conducted extensive experiments on each of the CUReT, the Brodatz, and the MSRC databases, comparing the proposed approach to four state-of-the-art texture classification methods: Patch, Patch-MRF, MR8, and LBP. We show that our approach leads to significant improvements in classification accuracy and reductions in feature dimensionality. Li Liu 0002, Paul W. Fieguth |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | Sorted random projections for robust rotation-invariant texture classification
Li Liu 0002, Paul W. Fieguth, David A. Clausi, Gangyao Kuang |
Pattern Recognit. | 1 |
| 2011 | Sorted Random Projections for robust texture classificationabstractThis paper presents a simple and highly effective system for robust texture classification, based on (1) random local features, (2) a simple global Bag-of-Words (BoW) representation, and (3) Support Vector Machines (SVMs) based classification. The key contribution in this work is to apply a sorting strategy to a universal yet information-preserving random projection (RP) technique, then comparing two different texture image representations (histograms and signatures) with various kernels in the SVMs. We have tested our texture classification system on six popular and challenging texture databases for exemplar based texture classification, comparing with 12 recent state-of-the-art methods. Experimental results show that our texture classification system yields the best classification rates of which we are aware of 99.37% for CUReT, 97.16% for Brodatz, 99.30% for UMD and 99.29% for KTH-TIPS. Moreover, combining random features significantly outperforms the state-of-the-art descriptors in material categorization. Li Liu 0002, Paul W. Fieguth, Gangyao Kuang, Hongbin Zha |
ICCV | 1 |
| 2011 | Combining sorted random features for texture classificationabstractThis paper explores the combining of powerful local texture descriptors and the advantages over single descriptors for texture classification. The proposed system is composed of three components: (i) highly discriminative and robust sorted random projections (SRP) features; (ii) a global Bag-of-Words (BoW) model; and (iii) the use of multiple kernel Support Vector Machines (SVMs) combining multiple features. The proposed system is also very simple, stemming from (1) the effortless extraction of the SRP features, (2) the simple orderless histogramming in the BoW model, (3) a strategy with low computational complexity for multiple kernel SVMs. We have tested our texture classification system on three popular and challenging texture databases and find that the SVMs combining of SRP features produces outstanding classification results, out-performing the state-of-the-art for CUReT (99.37%) and KTH-TIPS (99.29%), and with highly competitive results for UIUC (98.56%). Li Liu 0002, Paul W. Fieguth, Gangyao Kuang |
ICIP | 1 |
| 2010 | Compressed Sensing for Robust Texture Classification
Li Liu 0002, Paul W. Fieguth, Gangyao Kuang |
ACCV (1) | 1 |
| 2009 | An Adaptive and Fast CFAR Algorithm Based on Automatic Censoring for Target Detection in High-Resolution SAR ImagesabstractAn adaptive and fast constant false alarm rate (CFAR) algorithm based on automatic censoring (AC) is proposed for target detection in high-resolution synthetic aperture radar (SAR) images. First, an adaptive global threshold is selected to obtain an index matrix which labels whether each pixel of the image is a potential target pixel or not. Second, by using the index matrix, the clutter environment can be determined adaptively to prescreen the clutter pixels in the sliding window used for detecting. The$G^{0}$distribution, which can model multilook SAR images within an extensive range of degree of homogeneity, is adopted as the statistical model of clutter in this paper. With the introduction of AC, the proposed algorithm gains good CFAR detection performance for homogeneous regions, clutter edge, and multitarget situations. Meanwhile, the corresponding fast algorithm greatly reduces the computational load. Finally, target clustering is implemented to obtain more accurate target regions. According to the theoretical performance analysis and the experiment results of typical real SAR images, the proposed algorithm is shown to be of good performance and strong practicability. Gui Gao, Li Liu 0002, Lingjun Zhao, Gongtao Shi, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 2 |