Zhen Liu 0004

dblp:77/35-4 · DBLP profile ↗
← Back
33ranked-venue papers
4as first author
24since 2021 · last 2026
0000-0002-1233-1494ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 8 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ATRNet-STAR: A Large Dataset and Benchmark Toward Remote Sensing Object Recognition in the Wild
abstract
The absence of publicly available, large-scale, high-quality datasets for Synthetic Aperture Radar Automatic Target Recognition (SAR ATR) has significantly hindered the application of rapidly advancing deep learning techniques, which hold huge potential to unlock new capabilities in this field. This is primarily because collecting large volumes of diverse target samples from SAR images is prohibitively expensive, largely due to privacy concerns, the characteristics of microwave radar imagery perception, and the need for specialized expertise in data annotation. Throughout the history of SAR ATR research, there have been only a number of small datasets, mainly including targets like ships, airplanes, buildings, etc. There is only one vehicle dataset MSTAR collected in the 1990 s, which has been a valuable source for SAR ATR. To fill this gap, this paper introduces a large-scale, new dataset named ATRNet-STAR with 40 different vehicle categories collected under various realistic imaging conditions and scenes. It marks a substantial advancement in dataset scale and diversity, comprising over 190,000 well-annotated samples-$10\times$ larger than its predecessor, the famous MSTAR. Building such a large dataset is a challenging task, and the data collection scheme will be detailed. Secondly, we illustrate the value of ATRNet-STAR via extensively evaluating the performance of 15 representative methods with 7 different experimental settings on challenging classification and detection benchmarks derived from the dataset. Finally, based on our extensive experiments, we identify valuable insights for SAR ATR and discuss potential future research directions in this field. We hope that the scale, diversity, and benchmark of ATRNet-STAR can significantly facilitate the advancement of SAR ATR.
Yongxiang Liu, Li Liu 0002, Jie Zhou 0031, Bowen Peng, Xuying Xiong, Wei Yang 0046, Tianpeng Liu, Zhen Liu 0004, Xiang Li 0014
IEEE Trans. Pattern Anal. Mach. Intell.10
2025 Fusion Meets Diverse Conditions: A High-Diversity Benchmark and Baseline for UAV-Based Multimodal Object Detection with Condition Cues
abstract
Unmanned aerial vehicles (UAV)-based object detection with visible (RGB) and infrared (IR) images facilitates robust around-the-clock detection, driven by advancements in deep learning techniques and the availability of high-quality dataset. However, the existing dataset struggles to fully capture real-world complexity for limited imaging conditions. To this end, we introduce a high-diversity dataset ATR-UMOD covering varying scenarios, spanning altitudes from 80m to 300m, angles from 0° to 75°, and all-day, all-year time variations in rich weather and illumination conditions. Moreover, each RGB-IR image pair is annotated with 6 condition attributes, offering valuable high-level contextual information. To meet the challenge raised by such diverse conditions, we propose a novel prompt-guided condition-aware dynamic fusion (PCDF) to adaptively reassign multimodal contributions by leveraging annotated condition cues. By encoding imaging conditions as text prompts, PCDF effectively models the relationship between conditions and multimodal contributions through a task-specific soft-gating transformation. A prompt-guided condition-decoupling module further ensures the availability in practice without condition annotations. Experiments on ATR-UMOD dataset reveal the effectiveness of PCDF.
Chen Chen 0152, Kangcheng Bin, Jiahao Qi, Tianpeng Liu, Zhen Liu 0004, Yongxiang Liu, Ping Zhong 0001
ICCV7
2025 Luminance-Aware Statistical Quantization: Unsupervised Hierarchical Learning for Illumination Enhancement
abstract
Low-light image enhancement (LLIE) faces persistent challenges in balancing reconstruction fidelity with cross-scenario generalization. While existing methods predominantly focus on deterministic pixel-level mappings between paired low/normal-light images, they often neglect the continuous physical process of luminance transitions in real-world environments, leading to performance drop when normal-light references are unavailable. Inspired by empirical analysis of natural luminance dynamics revealing power-law distributed intensity transitions, this paper introduces Luminance-Aware Statistical Quantification (LASQ), a novel framework that reformulates LLIE as a statistical sampling process over hierarchical luminance distributions. Our LASQ re-conceptualizes luminance transition as a power-law distribution in intensity coordinate space that can be approximated by stratified power functions, therefore, replacing deterministic mappings with probabilistic sampling over continuous luminance layers. A diffusion forward process is designed to autonomously discover optimal transition paths between luminance layers, achieving unsupervised distribution emulation without normal-light references. In this way, it considerably improves the performance in practical situations, enabling more adaptable and versatile light restoration. This framework is also readily applicable to cases with normal-light references, where it achieves superior performance on domain-specific datasets alongside better generalization-ability across non-reference datasets. The code is available at: https://github.com/XYLGroup/LASQ.
Derong Kong, Zhixiong Yang 0001, Shengxi Li, Shuaifeng Zhi, Li Liu 0002, Zhen Liu 0004, Jingyuan Xia
NeurIPS6
2025 A Causal Adjustment Module for Debiasing Scene Graph Generation
abstract
While recent debiasing methods for Scene Graph Generation (SGG) have shown impressive performance, these efforts often attribute model bias solely to the long-tail distribution of relationships, overlooking the more profound causes stemming from skewed object and object pair distributions. In this paper, we employ causal inference techniques to model the causality among these observed skewed distributions. Our insight lies in the ability of causal inference to capture the unobservable causal effects between complex distributions, which is crucial for tracing the roots of model bias. Specifically, we introduce the Mediator-based Causal Chain Model (MCCM), which, in addition to modeling causality among objects, object pairs, and relationships, incorporates mediator variables, i.e., cooccurrence distribution, for complementing the causality. Following this, we propose the Causal Adjustment Module (CAModule) to estimate the modeled causal structure, using variables from MCCM as inputs to produce a set of adjustment factors aimed at correcting biased model predictions. Moreover, our method enables the composition of zero-shot relationships, thereby enhancing the model's ability to recognize such relationships. Experiments conducted across various SGG backbones and popular benchmarks demonstrate that CAModule achieves state-of-the-art mean recall rates, with significant improvements also observed on the challenging zero-shot recall rate metric.
Li Liu 0002, Shuzhou Sun, Shuaifeng Zhi, Fan Shi 0003, Zhen Liu 0004, Janne Heikkilä, Yongxiang Liu
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Large Language Models Can Achieve Explainable and Training-Free One-Shot HRRP ATR
abstract
This letter introduces a pioneering, training-free and explainable framework for High-Resolution Range Profile (HRRP) automatic target recognition (ATR) utilizing large-scale pre-trained Large Language Models (LLMs). Diverging from conventional methods requiring extensive task-specific training or fine-tuning, our approach converts one-dimensional HRRP signals into textual scattering center representations. Prompts are designed to align LLMs’ semantic space for ATR via few-shot in-context learning, effectively leveraging its vast pre-existing knowledge without any parameter update.
Lingfeng Chen, Panhe Hu, Zhiliang Pan, Qi Liu 0058, Shuanghui Zhang, Zhen Liu 0004
IEEE Signal Process. Lett.6
2025 Observations Temporal Permutation-Based Self-Supervised Reinforcement Learning for UAV Active Object Detection
abstract
In passive ground target detection using Unmanned Aerial Vehicles (UAVs), some detrimental factors like occlusion significantly impact target detection performance. Active Object Detection offers an effective way to address it, which usually uses Deep Reinforcement Learning (DRL) to plan UAV’s viewpoint for favorable observations. However, existing DRL-based AOD methods often suffer from low sample efficiency and poor generalization due to inadequate state representation learned by the policy network. Inspired by human scene understanding where their spatial representation of the scene remains consistent despite different observation orders, we design a self-supervised state representation learning method based on Observations Temporal Permutation (OTP) to improve the state representation of the agent’s policy network. We require the policy network to output consistent action value estimates for observation sequences with the same content but different temporal orders. Besides, we use the state representation to predict the target orientation variations in the observation sequence, which further regularizes and facilitates the state representation learning process. Finally, we design multiple experiments based on the UEVAVD dataset to compare the proposed method with existing self-supervised state representation learning methods for the AOD task. The experimental results demonstrate that the OTP method can help the agent’s policy network learn a better state representation, thus achieving higher policy learning sample efficiency and stronger policy generalization.
Xinhua Jiang, Tianpeng Liu, Li Liu 0002, Zhen Liu 0004, Yongxiang Liu
IEEE Trans. Geosci. Remote. Sens.4
2025 Advancing Segment Anything Model for Efficient Salient Object Detection in Remote Sensing Images
abstract
Salient object detection in optical remote sensing images (ORSI-SOD) often relies on leveraging pre-trained knowledge from natural images to achieve high accuracy with limited training data. Traditional methods typically employ vision backbones (e.g., Convolutional Neural Networks (CNNs) or Vision Transformers (ViTs)) pre-trained on ImageNet to extract features from ORSI scenes. However, these backbones exhibit limited generalization across diverse scenarios compared to recent vision foundation models. To this end, we propose ORSI-SAM, a novel ORSI-SOD framework based on the Segment Anything Model (SAM), leveraging its superior generalization capabilities to achieve an exceptional efficiency-accuracy trade-off. Specifically, ORSI-SAM adopts lightweight SAM as the backbone, effectively reducing parameter size and computational overhead to enable efficient deployment on satellite devices while retaining the rich knowledge learned from large-scale natural image datasets. To mitigate the impact of unavailable prompts in ORSI-SOD on the prediction capability of the SAM decoder, we introduce a Hierarchical Interaction Prompt Generator (HIPG), which aggregates hierarchical features and generates mask prompts tailored for salient objects to guide the decoder in producing high-quality saliency maps. Furthermore, to address the recognition challenges caused by the inherent characteristics of ORSIs, we propose a Semantic-Aware Refinement Decoder (SARD). SARD integrates structural details from low-level features to enrich fine-grained object information while leveraging high-level features to suppress redundant interference in shallow layers, thereby improving the detailed information in the predicted saliency map. ORSI-SAM is the first work to explore the accuracy-efficiency trade-offs for ORSI-SOD based on SAM architecture. Extensive experiments on benchmark datasets show that ORSI-SAM achieves superior performance compared to recent state-of-the-art methods with 12.2M parameters and 8.9G FLOPs.
Li Liu 0002, Zhuo Su 0002, Tianpeng Liu, Zhen Liu 0004, Matti Pietikäinen
IEEE Trans. Geosci. Remote. Sens.5
2025 Sparse Aperture ISAR Autofocusing and Imaging Algorithm Based on Log-Sum Regularization
abstract
Mathematically, the autofocusing and imaging model for sparse aperture inverse synthetic aperture radar (ISAR) has an infinite number of solutions, even with the addition of some sparsity constraints. What relationship exists between these infinite solutions should be answered. In addition, some existing approaches suffer from troublesome manual fine-tuning parameters, or high algorithmic complexity, or low reconstruction accuracy. To deal with the above problems, a sparse aperture ISAR autofocusing and imaging method combining phase estimation and log-sum minimization is proposed in this paper, which has high estimation accuracy and computational efficiency, and avoids complicated manual parameter adjustment. Focusing on the model with a phase diagonal matrix and log-sum function, we reveal why there are countless solutions and what relationships exist between them. These features are shared by other models including different regularization functions. Without sparsity prior, the regularization parameter of the algorithm is changed automatically at each iteration accroding to the ratio of the selected element to the maximum absolute value element in an auxiliary matrix, which controls that only a few elements participate in the computation at each iteration. Coupled with the situation that computationally heavy matrix inversion operations are not required, the calculation speed of the approach is greatly boosted. To show the stability of the algorithm, its convergence is proven without any assumptions. Experiments based on both simulation and measurement data demonstrate that compared with the recently proposed algorithms by others, the proposed algorithm can achieve well-focused ISAR images in less than one second and is very efficient to implement.
Jianjun Shen, Zhen Liu 0004, Yongxiang Liu, Bin Xue 0003
IEEE Trans. Geosci. Remote. Sens.4
2025 ARBiBench: Benchmarking and Analyzing Adversarial Robustness of Binarized Convolutional Neural Networks
abstract
Binarized convolutional neural networks (BCNNs), which restrict the weights and activations of the model to +1 or −1, provide notable reductions in memory requirements and enhanced model inference speed during deployment. Current research on BCNNs primarily revolves around addressing the performance degradation resulting from binarization. However, the investigation of the effects of extreme discretization on the robustness of BCNNs has been largely overlooked, despite its critical relevance to real-world applications. To this end, we propose ARBiBench, a comprehensive benchmark for evaluating the adversarial robustness of BCNNs in the image classification task. The key contributions of ARBiBench include: 1) systematically evaluating the robustness of seven influential BCNN methods across various architectures; 2) rigorous validation of diverse adversarial attack methods; and 3) novel empirical findings showing that BCNNs exhibit weaker robustness than full-precision networks on small datasets but surprisingly stronger robustness on large-scale datasets. Leveraging Information Bottleneck theory, we further demonstrate how data scale and model capacity collectively determine BCNNs’ adversarial robustness. These findings not only challenge conventional assumptions about BCNN security, but also provide new insights for developing robust yet efficient neural network architectures.
Li Liu 0002, Bowen Peng, Zhen Liu 0004, Longguang Wang, Yingmei Wei
IEEE Trans. Inf. Forensics Secur.5
2025 Refining Pseudo Labeling via Multi-Granularity Confidence Alignment for Unsupervised Cross Domain Object Detection
abstract
Most state-of-the-art object detection methods suffer from poor generalization due to the domain shift between training and testing datasets. To resolve this challenge, unsupervised cross domain object detection is proposed to learn an object detector for an unlabeled target domain by transferring knowledge from an annotated source domain. Promising results have been achieved via Mean Teacher, however, pseudo labeling which is the bottleneck of mutual learning remains to be further explored. In this study, we find that confidence misalignment of the predictions, including category-level overconfidence, instance-level task confidence inconsistency, and image-level confidence misfocusing, leading to the injection of noisy pseudo labels in the training process, will bring suboptimal performance. Considering the above issue, we present a novel general framework termed Multi-Granularity Confidence Alignment Mean Teacher (MGCAMT) for unsupervised cross domain object detection, which alleviates confidence misalignment across category-, instance-, and image-levels simultaneously to refine pseudo labeling for better teacher-student learning. Specifically, to align confidence with accuracy at category level, we propose Classification Confidence Alignment (CCA) to model category uncertainty based on Evidential Deep Learning (EDL) and filter out the category incorrect labels via an uncertainty-aware selection strategy. Furthermore, we design Task Confidence Alignment (TCA) to mitigate the instance-level misalignment between classification and localization by enabling each classification feature to adaptively identify the optimal feature for regression. Finally, we develop imagery Focusing Confidence Alignment (FCA) adopting another way of pseudo label learning, i.e., we use the original outputs from the Mean Teacher network for supervised learning without label assignment to achieve a balanced perception of the image's spatial layout. When these three procedures are integrated into a single framework, they mutually benefit to improve the final performance from a cooperative learning perspective. Extensive experiments across multiple scenarios demonstrate that our method outperforms large foundational models, and surpasses other state-of-the-art approaches by a large margin.
Jiangming Chen, Li Liu 0002, Wanxia Deng, Zhen Liu 0004, Yu Liu 0012, Yingmei Wei, Yongxiang Liu
IEEE Trans. Image Process.4
2025 Boosting Convolutional Neural Networks With Middle Spectrum Grouped Convolution
abstract
This article proposes a novel module called middle spectrum grouped convolution (MSGC) for efficient deep convolutional neural networks (DCNNs) with the mechanism of grouped convolution. It explores the broad "middle spectrum" area between channel pruning and conventional grouped convolution. Compared with channel pruning, MSGC can retain most of the information from the input feature maps due to the group mechanism; compared with grouped convolution, MSGC benefits from the learnability, the core of channel pruning, for constructing its group topology, leading to better channel division. The middle spectrum area is unfolded along four dimensions: groupwise, layerwise, samplewise, and attentionwise, making it possible to reveal more powerful and interpretable structures. As a result, the proposed module acts as a booster that can reduce the computational cost of the host backbones for general image recognition with even improved predictive accuracy. For example, in the experiments on the ImageNet dataset for image classification, MSGC can reduce the multiply-accumulates (MACs) of ResNet-18 and ResNet-50 by half but still increase the Top-1 accuracy by more than 1%. With a 35% reduction of MACs, MSGC can also increase the Top-1 accuracy of the MobileNetV2 backbone. Results on the MS COCO dataset for object detection show similar observations. Our code and trained models are available at https://github.com/hellozhuo/msgc.
Zhuo Su 0002, Tianpeng Liu, Zhen Liu 0004, Shuanghui Zhang, Matti Pietikäinen, Li Liu 0002
IEEE Trans. Neural Networks Learn. Syst.4
2024 A Dynamic Kernel Prior Model for Unsupervised Blind Image Super-Resolution
abstract
Deep learning-based methods have achieved significant successes on solving the blind super-resolution (BSR) problem. However, most of them request supervised pretraining on labelled datasets. This paper proposes an unsupervised kernel estimation model, named dynamic kernel prior (DKP), to realize an unsupervised and pretraining-free learning-based algorithm for solving the BSR problem. DKP can adaptively learn dynamic kernel priors to realize real-time kernel estimation, and thereby enables superior HR image restoration performances. This is achieved by a Markov chain Monte Carlo sampling process on random kernel distributions. The learned kernel prior is then assigned to optimize a blur kernel estimation network, which entails a network-based Langevin dynamic optimization strategy. These two techniques ensure the accuracy of the kernel estimation. DKP can be easily used to replace the kernel estimation models in the existing methods, such as Double-DIP and FKP-DIP, or be added to the off-the-shelf image restoration model, such as diffusion model. In this paper, we incorporate our DKP model with DIP and diffusion model, referring to DIP-DKP and Diff-DKP, for validations. Extensive simulations on Gaussian and motion kernel scenarios demonstrate that the proposed DKP model can significantly improve the kernel estimation with comparable runtime and memory usage, leading to state-of-the-art BSR results. The code is available at https://github.com/XYLGroup/DKP.
Zhixiong Yang 0001, Jingyuan Xia, Shengxi Li, Xinghua Huang, Shuanghui Zhang, Zhen Liu 0004, Yaowen Fu, Yongxiang Liu
CVPR6
2024 DiffDet4SAR: Diffusion-Based Aircraft Target Detection Network for SAR Images
abstract
Aircraft target detection in SAR images is a challenging task due to the discrete scattering points and severe background clutter interference. Currently, methods with convolution-based or transformer-based paradigms cannot adequately address these issues. In this letter, we explore diffusion models for SAR image aircraft target detection for the first time and propose a novel Diffusion-based aircraft target Detection network for SAR images (DiffDet4SAR). Specifically, the proposed DiffDet4SAR yields two main advantages for SAR aircraft target detection: 1) DiffDet4SAR maps the SAR aircraft target detection task to a denoising diffusion process of bounding boxes without heuristic anchor size selection, effectively enabling large variations in aircraft sizes to be accommodated; and 2) the dedicatedly designed Scattering Feature Enhancement (SFE) module further reduces the clutter intensity and enhances the target saliency during inference. Extensive experimental results on the SAR-AIRcraft-1.0 dataset show that the proposed DiffDet4SAR achieves 88.4% mAP50, outperforming the state-of-the-art methods by 6%. Code is availabel at https://github.com/JoyeZLearning/DiffDet4SAR.
Jie Zhou 0031, Zhen Liu 0004, Li Liu 0002, Yongxiang Liu, Xiang Li 0014
IEEE Geosci. Remote. Sens. Lett.4
2024 Sparsity-Based Adaptive Beamforming for Coherent Signals With Polarized Sensor Arrays
abstract
A sparsity-based adaptive beamforming (ABF) method is introduced to effectively process coherent signals with polarized sensor arrays (PSA). This method exploits the spatial sparsity of observed signals by transforming it into row-sparsity within a waveform-polarization composite matrix through data reorganization. This row-sparsity is subsequently cast as an$\ell _{2,1}$norm minimization problem, characterized by a gridless and compact mathematical expression with a Hermitian Toeplitz matrix. Then, a matrix factorization-based gradient descent (GD) algorithm is introduced to effectively resolve this optimization problem. The experimental evaluations demonstrate that the GD algorithm significantly outperforms the MOSEK solver in terms of computational efficiency. Further comparative analysis demonstrates that the proposed method outperforms the existing techniques, especially in contexts of low signal-to-noise ratio (SNR), with a moderate increase in computational runtime.
Tianpeng Liu, Junpeng Shi, Zhen Liu 0004, Yongxiang Liu
IEEE Signal Process. Lett.4
2024 Meta-Learning Based Domain Prior With Application to Optical-ISAR Image Translation
abstract
This paper focuses on generating Inverse Synthetic Aperture Radar (ISAR) images from optical images, in particular, for orbit space targets. ISAR images are widely applied in space target observation and classification tasks, whereas, limited to the expensive cost of ISAR sample collection, training deep learning-based ISAR image classifiers with insufficient samples and generating ISAR samples from emulation optical images via image translation techniques have attracted increasing attention. Image translation has highlighted significant success and popularity in computer vision, remote sensing and data generation societies. However, most of the existing methods are implemented under the discipline of extracting the explicit pixel-level features and do not perform effectively while entailing translation to domains with specific implicit features, such as ISAR image does. We propose a meta-learning based domain prior to implicit feature modelling and apply it to CycleGAN and UNIT models to realize effective translations between the ISAR and optical domains. Two representative implicit features, ISAR scattering distribution feature from the physical domain and the classification identifying feature from the task domain, are elaborately formulated with explicit modelling in statistic form. A meta-learning based training scheme is introduced to leverage the mutual knowledge of domain priors across different samples, and thus allows few-shot learning capacity with dramatically reduced training samples. Extensive simulations validate that the obtained ISAR images have better visible-authenticity and training-effectiveness than the existing image translation approaches on various synthetic datasets. Source codes are available at.
Huaizhang Liao, Jingyuan Xia, Zhixiong Yang 0001, Fulin Pan, Zhen Liu 0004, Yongxiang Liu
IEEE Trans. Circuits Syst. Video Technol.5
2024 From Coarse to Fine: ISAR Object View Interpolation via Flow Estimation and GAN
abstract
This article focuses on the multiazimuth angle interpolation task of inverse synthetic aperture radar (ISAR) images for aircraft targets and complements incomplete ISAR image datasets. ISAR image automatic target recognition (ATR) has been widely applied in remote sensing and many fields. However, the imaging process is more challenging when compared to capturing optical and SAR image data, which reduces the accuracy and generalization performance of the ATR system. Therefore, in this article, we leverage existing limited ISAR data to achieve autonomous data expansion. This approach helps mitigate the impact of low sample quantity and unbalanced distribution, ultimately improving the accuracy of the ATR system for target recognition. Most existing methods use generative networks for ISAR image expansion, but few focus on generating ISAR images with specific azimuth angles. This article proposes a novel two-stage coarse-to-fine framework for ISAR object view interpolation (C2FIPNet) that combines flow estimation and GAN to interpolate ISAR images with intermediate azimuth angles using a set of ISAR image pairs. Flow estimation is employed for coarse-grained generation, determining the position and intensity of strong scattering points in the ISAR image. The GAN, on the other hand, is used for fine-grained completion to correct image distortion caused by flow estimation and enhance image details. In addition, a suitable loss function is designed, incorporating both global and local features, allowing for priority generation in the region of strong scattering points. In conclusion, extensive simulation and comparative experiments have demonstrated that the interpolated ISAR images generated by the proposed C2FIPNet exhibit greater pixel-level authenticity.
Zhen Liu 0004, Weidong Jiang, Yongxiang Liu, Shuowei Liu, Li Liu 0002
IEEE Trans. Geosci. Remote. Sens.2
2023 Lightweight Pixel Difference Networks for Efficient Visual Representation Learning
abstract
Recently, there have been tremendous efforts in developing lightweight Deep Neural Networks (DNNs) with satisfactory accuracy, which can enable the ubiquitous deployment of DNNs in edge devices. The core challenge of developing compact and efficient DNNs lies in how to balance the competing goals of achieving high accuracy and high efficiency. In this paper we propose two novel types of convolutions, dubbed Pixel Difference Convolution (PDC) and Binary PDC (Bi-PDC) which enjoy the following benefits: capturing higher-order local differential information, computationally efficient, and able to be integrated with existing DNNs. With PDC and Bi-PDC, we further present two lightweight deep networks named Pixel Difference Networks (PiDiNet) and Binary PiDiNet (Bi-PiDiNet) respectively to learn highly efficient yet more accurate representations for visual tasks including edge detection and object recognition. Extensive experiments on popular datasets (BSDS500, ImageNet, LFW, YTF, etc.) show that PiDiNet and Bi-PiDiNet achieve the best accuracy-efficiency trade-off. For edge detection, PiDiNet is the first network that can be trained without ImageNet, and can achieve the human-level performance on BSDS500 at 100 FPS and with 1 M parameters. For object recognition, among existing Binary DNNs, Bi-PiDiNet achieves the best accuracy and a nearly 2× reduction of computational cost on ResNet18.
Zhuo Su 0002, Longguang Wang, Hua Zhang 0008, Zhen Liu 0004, Matti Pietikäinen, Li Liu 0002
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Uncertainty-Guided Semi-Supervised Few-Shot Class-Incremental Learning With Knowledge Distillation
abstract
Class-Incremental Learning (CIL) aims at incrementally learning novel classes without forgetting old ones. This capability becomes more challenging when novel tasks contain one or a few labeled training samples, which leads to a more practical learning scenario,i.e., Few-Shot Class- Incremental Learning (FSCIL). The dilemma on FSCIL lies in serious overfitting and exacerbated catastrophic forgetting caused by the limited training data from novel classes. In this paper, excited by the easy accessibility of unlabeled data, we conduct a pioneering work and focus on a Semi-Supervised Few-Shot Class-Incremental Learning (Semi-FSCIL) problem, which requires the model incrementally to learn new classes from extremely limited labeled samples and a large number of unlabeled samples. To address this problem, a simple but efficient framework is first constructed based on the knowledge distillation technique to alleviate catastrophic forgetting. To efficiently mitigate the overfitting problem on novel categories with unlabeled data, uncertainty-guided semi-supervised learning is incorporated into this framework to select unlabeled samples into incremental learning sessions considering the model uncertainty. This process provides extra reliable supervision for the distillation process and contributes to better formulating the class means. Our extensive experiments on CIFAR100, miniImageNet and CUB200 datasets demonstrate the promising performance of our proposed method, and define baselines in this new research direction.
Yawen Cui, Wanxia Deng, Xin Xu 0001, Zhen Liu 0004, Zhong Liu 0002, Matti Pietikäinen, Li Liu 0002
IEEE Trans. Multim.4
2022 Deep Alternating Projection Networks for Gridless DOA Estimation With Nested Array
abstract
Recently, deep unfolding networks with interpretable parameters have been widely utilized in direction of arrival (DOA) estimation due to the faster convergence speed and better generalization ability. However, few consider the nested array for gridless DOA estimation. In this letter, we propose a deep alternating projection network to address the problem. We first convert the covariance matrix into a measurement vector in the form of atomic norm, which can reduce the matrix dimension during projection. We then train the proposed network to alternately obtain the positive semi-definite matrix and the corresponding irregular Hermitian Toeplitz matrix, where the loss function is derived by employing the trace of network output. Finally, we apply the irregular root Multiple Signal Classification (MUSIC) method to obtain gridless DOA via nested array. We demonstrate that the proposed networks can accelerate the convergence rate and reduce computational cost. Simulations verify the performance of proposed networks in comparison with the existing methods.
Xiaolong Su, Panhe Hu, Zhen Liu 0004, Junpeng Shi, Xiang Li 0014
IEEE Signal Process. Lett.3
2022 Dynamic Sparse Subspace Clustering for Evolving High-Dimensional Data Streams
abstract
In an era of ubiquitous large-scale evolving data streams, data stream clustering (DSC) has received lots of attention because the scale of the data streams far exceeds the ability of expert human analysts. It has been observed that high-dimensional data are usually distributed in a union of low-dimensional subspaces. In this article, we propose a novel sparse representation-based DSC algorithm, called evolutionary dynamic sparse subspace clustering (EDSSC). It can cope with the time-varying nature of subspaces underlying the evolving data streams, such as subspace emergence, disappearance, and recurrence. The proposed EDSSC consists of two phases: 1) static learning and 2) online clustering. During the first phase, a data structure for storing the statistic summary of data streams, called EDSSC summary, is proposed which can better address the dilemma between the two conflicting goals: 1) saving more points for accuracy of subspace clustering (SC) and 2) discarding more points for the efficiency of DSC. By further proposing an algorithm to estimate the subspace number, the proposed EDSSC does not need to know the number of subspaces. In the second phase, a more suitable index, called the average sparsity concentration index (ASCI), is proposed, which dramatically promotes the clustering accuracy compared to the conventionally utilized SCI index. In addition, the subspace evolution detection model based on the Page-Hinkley test is proposed where the appearing, disappearing, and recurring subspaces can be detected and adapted. Extinct experiments on real-world data streams show that the EDSSC outperforms the state-of-the-art online SC approaches.
Jinping Sui, Zhen Liu 0004, Li Liu 0002, Alexander Jung 0001, Xiang Li 0014
IEEE Trans. Cybern.2
2021 Generalized Thinned Coprime Array for DOA Estimation
abstract
Owing to the large degrees of freedom and reduced mutual coupling by producing difference coarrays, nonuniform linear arrays have aroused great interest in direction of arrival (DOA) estimation. Previous works have presented some new sparse arrays, such as the thinned coprime array. In this paper, we propose a generalized thinned coprime array by introducing the flexible inter-element spacings, where the conventional one can be seen as a special case. We derive closedform expression for the range of consecutive lags, written as the functions of the antenna numbers and inter-element spacings. We show that, after optimization, the proposed array can achieve more consecutive lags than the other coprime arrays. In particular, the optimized results also provide the minimum number of antenna pairs with small separation. Simulation results demonstrate the superiority of the proposed GTCA using the subspace-based method.
Junpeng Shi, Yongxiang Liu, Fangqing Wen, Zhen Liu 0004, Panhe Hu, Zhenghui Gong
ICASSP4
2021 Parameter Identifiability Of Spatial-Smoothing-Based Bistatic Mimo Radar
abstract
Diversity smoothing has been widely developed for angle estimation with bistatic multiple input multiple output (MIMO) radar in the presence of coherent targets, the parameter identifiability of which is an important issue. In this paper, we are devoted to establishing more accurate conditions by studying the positive definiteness of smoothed target covariance matrix. The antenna numbers of transmit and receive arrays are derived as functions of the target number and target structure. We show that the new results improve upon previous ones and recover them in special cases. Simulation results are presented that corroborate our theoretical findings.
Junpeng Shi, Fangqing Wen, Yongxiang Liu, Qinmu Shen, Zhihui Li 0002, Zhen Liu 0004
ICASSP6
2021 Informative Class-Conditioned Feature Alignment for Unsupervised Domain Adaptation
abstract
The goal of unsupervised domain adaptation is to learn a task classifier that performs well for the unlabeled target domain by borrowing rich knowledge from a well-labeled source domain. Although remarkable breakthroughs have been achieved in learning transferable representation across domains, two bottlenecks remain to be further explored. First, many existing approaches focus primarily on the adaptation of the entire image, ignoring the limitation that not all features are transferable and informative for the object classification task. Second, the features of the two domains are typically aligned without considering the class labels; this can lead the resulting representations to be domain-invariant but non-discriminative to the category. To overcome the two issues, we present a novel Informative Class-Conditioned Feature Alignment (IC2FA) approach for UDA, which utilizes a twofold method: informative feature disentanglement and class-conditioned feature alignment, designed to address the above two challenges, respectively. More specifically, to surmount the first drawback, we cooperatively disentangle the two domains to obtain informative transferable features; here, Variational Information Bottleneck (VIB) is employed to encourage the learning of task-related semantic representations and suppress task-unrelated information. With regard to the second bottleneck, we optimize a new metric, termed Conditional Sliced Wasserstein Distance (CSWD), which explicitly estimates the intra-class discrepancy and the inter-class margin. The intra-class and inter-class CSWDs are minimized and maximized, respectively, to yield the domain-invariant discriminative features. IC2FA equips class-conditioned feature alignment with informative feature disentanglement and causes the two procedures to work cooperatively, which facilitates informative discriminative features adaptation. Extensive experimental results on three domain adaptation datasets confirm the superiority of IC2FA.
Wanxia Deng, Yawen Cui, Zhen Liu 0004, Gangyao Kuang, Dewen Hu, Matti Pietikäinen, Li Liu 0002
ACM Multimedia3
2021 Convolution Neural Networks for Localization of Near-Field Sources via Symmetric Double-Nested Array
abstract
We present the convolution neural networks (CNNs) to achieve the localization of near‐field sources via the symmetric double‐nested array (SDNA). Considering that the incoherent near‐field sources can be separated in the frequency spectrum, we first calculate the phase difference matrices and consider the typical elements as the inputs of the networks. In order to guarantee the precision of the angle‐of‐arrival (AOA) estimation, we implement the autoencoders to divide the AOA subregions and construct the corresponding classification CNNs to obtain the AOAs of near‐field sources. Then, we construct a particular range vector without the estimated AOAs and utilize the regression CNN to obtain the range parameters of near‐field sources. The proposed algorithm is robust to the off‐grid parameters and suitable for the scenarios with the different number of near‐field sources. Moreover, the proposed method outperforms the existing method for near‐field source localization.
Xiaolong Su, Panhe Hu, Zhenghui Gong, Zhen Liu 0004, Junpeng Shi, Xiang Li 0014
Wirel. Commun. Mob. Comput.4
2020 PGNet: A Part-based Generative Network for 3D object reconstruction
Yang Zhang 0036, Kai Huo, Zhen Liu 0004, Yongxiang Liu, Xiang Li 0014, Cheng Wang 0003
Knowl. Based Syst.3
2020 Large-Scale Point Cloud Contour Extraction via 3-D-Guided Multiconditional Residual Generative Adversarial Network
abstract
As one of the most important features for human perception, contours are widely applied in graphics and mapping applications. However, it is considerably challenging to extract contours from large-scale point clouds due to the irregular distribution of point clouds. In this letter, we propose a 3-D-guided multiconditional residual generative adversarial network (3-D-GMRGAN), the first deep-learning framework to generate contours for large-scale outdoor point clouds. To make the network handle huge amounts of points, we operate contours in the parametric space rather than raw point space, associated with a parametric chamfer distance. Then, to gather contour features from potential positions and avoid the huge solution space, we propose a guided residual generative adversarial framework, by utilizing a simple feature-based method to get the “over extraction” potential contour distribution. Experiments demonstrate that the proposed method is able to generate contours efficiently for large-scale point clouds, with fewer outliers and pseudo contours compared with state-of-the-art approaches.
Yang Zhang 0036, Zhen Liu 0004, Tianpeng Liu, Xiang Li 0014
IEEE Geosci. Remote. Sens. Lett.2
2019 Sparse Subspace Clustering for Evolving Data Streams
abstract
The data streams arising in many applications can be modeled as a union of low-dimensional subspaces known as multi-subspace data streams (MSDSs). Clustering MSDSs according to their underlying low-dimensional subspaces is a challenging problem which has not been resolved satisfactorily by existing data stream clustering (DSC) algorithms. In this paper, we propose a sparse-based DSC algorithm, which we refer to as dynamic sparse subspace clustering (D-SSC). This algorithm recovers the low-dimensional subspaces (structures) of high-dimensional data streams and finds an explicit assignment of points to subspaces in an online manner. Moreover, as an online algorithm, D-SSC is able to cope with the time-varying structure of MSDSs. The effectiveness of D-SSC is evaluated using numerical experiments.
Jinping Sui, Zhen Liu 0004, Li Liu 0002, Alexander Jung 0001, Tianpeng Liu, Xiang Li 0014
ICASSP2
2014 Aliasing-free micro-Doppler analysis based on short-time compressed sensing
abstract
Time–frequency distribution (TFD) has been widely used for micro‐Doppler analysis in radar signal processing. However, the spectrogram will suffer from aliasing if the maximum Doppler frequency exceeds half of the pulse repetition frequency, which may lead to false estimation of the targets' kinematic properties. In this study, by transmitting a series of random pulse repetition interval (RPRI) pulses, a concise TFD approach named short‐time compressed sensing (STCS) is proposed for aliasing‐free micro‐Doppler analysis. In STCS, precise analysis and synthesis of the random sampling time series can be achieved by exploiting the signal's sparsity in the frequency domain. Furthermore, adaptive to the data, the widths of the particular rectangle windows are determined by sequential processing with a proper optimisation rule. To speed up the STCS procedure, the smoothed L0 algorithm is chosen for sparse recovery, where the pseudoinverse of the dictionaries can be calculated iteratively. The simulation results indicate that the proposed STCS approach can achieve both preferable TFD and acceptable computational cost. The effectiveness of the STCS is finally verified by the application for micro‐Doppler estimating in RPRI radar.
Zhen Liu 0004, Xizhang Wei, Xiang Li 0014
IET Signal Process.1
2014 Super-Resolution Reconstruction of Radar Tomographic Image Based on Image Decomposition
abstract
In this letter, the relationship between target scattering function and point spread function of radar tomographic imaging is discussed, and an inherent super-resolution reconstruction algorithm based on image decomposition (ID) is proposed. By removing the cross disturbance due to the resolution limitation, the exact radar images can be obtained. Furthermore, the total least-square regulation is applied to the ID-based algorithm in cases of additive white Gaussian noise and small parameter estimation error. Finally, the effectiveness of the proposed algorithm is demonstrated via numerical simulations.
Xizhang Wei, Zhen Liu 0004, Xiaofeng Ding 0005, Meimei Fan
IEEE Geosci. Remote. Sens. Lett.2
2013 Dynamic ISAR Imaging of Maneuvering Targets Based on Sequential SL0
abstract
For maneuvering targets, the time-varying Doppler shifts will produce blurred inverse synthetic aperture radar (ISAR) images for a long coherent processing interval (CPI). By exploiting sparsity of the target scene, sparse recovery (SR) algorithms have been applied to achieve high cross-range resolution within a short CPI, during which the Doppler shifts nearly remain constant. For practical applications, however, the required pulse number for attaining an acceptable image is difficult to designate in various scenarios, and the common recovery procedure suffers from low efficiency because of having to solve a new SR problem from scratch when the new echo pulses are sequentially available. In this letter, we present a dynamic ISAR imaging algorithm based on sequential smoothed L0, which is proposed as an efficient recursive implementation of the SR approach. Furthermore, by defining the proper stopping rules, we can seek the optimal pulse number required in each CPI. Simulation results show that the proposed dynamic algorithm is more suitable for ISAR imaging of uncooperative targets.
Zhen Liu 0004, Peng You, Xizhang Wei, Xiang Li 0014
IEEE Geosci. Remote. Sens. Lett.1
2013 Correction to "Dynamic ISAR Imaging of Maneuvering Targets Based on Sequential SL0"
abstract
There is an error in Fig. 3 in the above paper (ibid., vol. 10, no. 5, pp. 1041-1045, Sep. 2013). Bottom panels 3(g) and 3(h) are missing. The corrected figure is published here. We are sorry for the error.
Zhen Liu 0004, Peng You, Xizhang Wei, Xiang Li 0014
IEEE Geosci. Remote. Sens. Lett.1
2013 Resolution Analyses of Radar Tomographic Imaging for Arbitrary Angle Aperture
abstract
The technique of tomography, which collects a set of electromagnetic waves from a target over various angles, is a basic method to form 2-D images of a tridimensional object. However, the resolution ability of the radar tomographic imaging (RTI) has not been fully addressed. Therefore, based on the typical RTI algorithm called inverse Radon transform, this letter deduces its resolution performance in the cases of whole-angle aperture, small-angle aperture, and arbitrary-angle aperture, respectively. It has been proved that the resolution performance of inverse-synthetic-aperture-radar imaging is actually equivalent to that of the RTI at a small-angle aperture. Furthermore, aside from bandwidth, the carrier frequency and the relative position between the target and the radar also play crucial roles on the resolution ability of the RTI.
Xizhang Wei, Xiaofeng Ding 0005, Zhen Liu 0004, Meimei Fan
IEEE Geosci. Remote. Sens. Lett.3
2012 CS-based moving target detection in random PRI radar
abstract
Based on the compressed sensing (CS), we present a novel framework of moving target detection in random PRI radar. Firstly, the statistical characteristics of correlation output are analyzed to reflect the sidelobe pedestal. Then the equivalent sensing matrix is verified to approximately accord with the restricted isometry property by comparing it to a typical random CS matrix in a statistical sense. In order to cover the concerned range and velocity multi-channel processing is used. The simulation results demonstrate that this scheme has high performance of detection and large unambiguous scope, which can also shorten the coherent processing interval compared to traditional staggered PRI mode.
Zhen Liu 0004, Xizhang Wei, Xiang Li 0014
IGARSS1