Mou Wang

dblp:213/6786 · DBLP profile ↗
← Back
37ranked-venue papers
12as first author
32since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 25 · 9 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Pitch-Assistant Harmonic Recovery for Efficient Speech Enhancement
abstract
With the rapid development of low-resource online speech enhancement models, noise suppression can now be achieved with significantly fewer model parameters. However, these models often suffer from limited effectiveness in preserving speech quality. In particular, most existing online speech enhancement methods tend to distort the harmonic structure of speech while performing noise reduction, leading to noticeable degradation in perceptual quality. In this paper, we propose a novel model architecture called Pitch-Assistant Harmonic Recovery for Efficient Speech Enhancement (PHRSE). The model operates on low-dimensional Bark-scale spectral features to perform noise suppression, while leveraging estimated fundamental frequency information to guide the reconstruction of harmonic components. This pitch-guided strategy enables the model to preserve the speech’s natural harmonic structure more effectively. Experimental results demonstrate that PHRSE not only achieves higher perceptual speech quality compared to existing benchmarks, but also maintains real-time performance with significantly lower computational overhead, making it suitable for online and resource-constrained scenarios.
Zengqiang Shang, Haoyuan Xie, Mou Wang, Pengyuan Zhang
ASRU4
2025 Restoring Harmonics: Enhancing Speech Quality with Deep Mask and Harmonic Restoration Network
Zengqiang Shang, Mou Wang, Pengyuan Zhang
INTERSPEECH3
2025 Multi-granularity acoustic information fusion for sound event detection
Han Yin, Jisheng Bai, Mou Wang, Susanto Rahardja, Dongyuan Shi, Woon-Seng Gan
Signal Process.4
2025 SA-ISAR Imaging via Detail Enhancement Operator and Adaptive Threshold Sensing
abstract
Sparse aperture inverse synthetic aperture radar (SA-ISAR) aims to reconstruct target images by undersampled data. Traditional algorithms are limited in their application scope and exhibit weak capabilities in reconstructing target details. To address these issues, an adaptive threshold sensing (ATS) sparse reconstruction algorithm based on alternating direction method of multiplier (ADMM), named ATS-ADMM, is proposed. Within our framework, a detail-enhancing operator (DEO) is designed and combined with the$l_{1}$-norm to form a joint constrained optimization function to facilitate the recovery of weak scatterers. The matrix inversion operation within the ADMM framework is optimized to efficiently solve the multiconstrained problem. To enhance clutter suppression, an adaptive threshold network is designed based on deep convolutional networks. Inspired by deep learning, the DEO is set as a learnable operator, and the parameters are trained using an unsupervised network. Finally, the performance of ATS-ADMM is validated by comparing it with advanced algorithms using both simulated and real data. The results demonstrate that ATS-ADMM effectively focuses images, is suitable for diverse imaging scenarios, and is the fastest among ADMM-based algorithms.
Mou Wang, Yanbo Wen, Shunjun Wei, Jiangbo Hu, Wei Yi 0002, Jun Shi 0002
IEEE Trans. Geosci. Remote. Sens.1
2025 Exploring Spatial Feature Regularization in Deep-Learning-Based TomoSAR Reconstruction: A Preliminary Study and Performance Analysis
abstract
Tomographic synthetic aperture radar (TomoSAR) shows great potential for high-quality 3-D mapping, especially in urban areas. As TomoSAR reconstruction methods advance into the deep learning (DL) era, current studies have demonstrated DL’s strengths in both precision and efficiency. However, for reconstructing urban areas with prominent spatial features from building structures, current studies focus on pixel-by-pixel reconstruction without leveraging the potential benefits of these features. In this context, an exploratory study to introduce spatial feature regularization in DL reconstruction is proposed for the first time, focusing on feature description, modeling, and regularization. Spatial features are analyzed and summarized by sharp edges and regular geometric shapes within the scene. To model these features, 2-D slices are used as the basic reconstruction units, and a general intraslice and interslice strategy is proposed to harness features within and between slices. Two-dimensional slices are fused into the entire 3-D scene. Two methods of fusion are designed: parallel and serial. To regularize these features, a new computational framework called light reconstruction and enhancement is designed, which includes two stages: light reconstruction with sparsity feature regularization and enhancement with spatial feature regularization. Finally, to evaluate performance, we design an extensive evaluation framework. A newly self-constructed compound urban building simulation dataset, combined with two public measured data, forms six different tests ranging from a classical close point resolution test to a diverse urban landscape challenge test. Evaluation results reveal the effectiveness of the designs and the boost provided by spatial feature regularization, resulting in higher reconstruction precision, more complete building spatial structure retrieval, and fewer outliers.
Tianjiao Zeng, Xu Zhan, Xiangdong Ma, Jun Shi 0002, Shunjun Wei, Mou Wang, Xiaoling Zhang 0002
IEEE Trans. Geosci. Remote. Sens.8
2025 Unified Learning and Reconstruction for Robust Tomographic SAR Reconstruction: A Model-Driven Framework
abstract
Tomographic synthetic aperture radar (tomoSAR) imaging is a powerful tool for urban 3D reconstruction. While recent deep learning methods have improved reconstruction quality, their reliance on simulated measurement-scene-image pairs for supervised training raises concerns over robustness in real-world scenarios due to the distribution shifts, such as varying observation geometry, observed scene distributions, and signal/noise levels. These concerns have received limited attention until now, motivating us to explore an alternative approach that leverages the strength of deep learning for tomoSAR reconstruction without requiring paired measurement and scene-image training data. Therefore, we propose the Unified Learning and Reconstruction (ULAR), a model-driven method for robust tomoSAR reconstruction, trained without such paired data. ULAR integrates physical-model consistency with scene feature regularization (capturing spatial structures) in a unified optimization process. Specifically, it jointly performs image reconstruction and spatial structure refinement by alternating between physics-guided updates and mainly self-supervised spatial-structure learning. The approach incorporates two complementary components for spatial-structure learning: a one-step self-supervised generative model for local spatial structures and a pretrained denoiser for nonlocal ones. And the denoiser is further enhanced with equivariance properties to improve its robustness. Experimental results on both simulated and real measured datasets demonstrate that ULAR achieves reconstruction accuracy comparable to supervised methods, and even surpasses them when distribution shifts exist, revealing strong robustness while not relying on simulated measurement-ground truth paired data. These results demonstrate the robustness of self-supervised learning for tomoSAR reconstruction, while highlighting its potential for better practicality and reliability in real-world applications.
Xu Zhan, Tianjiao Zeng, Xiangdong Ma, Mou Wang, Jun Shi 0002, Shunjun Wei, Xiaoling Zhang 0002
IEEE Trans. Geosci. Remote. Sens.6
2024 Ensemble of Deep Variational Mixture Models for Unsupervised Clustering
abstract
Deep variational mixture models (DVMMs) have demonstrated promising performance in unsupervised clustering for complicated high-dimensional data such as images. However, their prediction accuracy is often unstable and significantly influenced by randomness, particularly during the initialization of parameters. To reduce this uncertainty, we propose an ensemble approach that combines the predictions of multiple base models. Specifically, we introduce two individual ensemble strategies: voting and merging. In the voting strategy, the final label is determined by selecting the predicted class label with the most votes and lowest Shannon entropy. In the merging strategy, the class probability vectors (scaled by the temperature parameter) from different models are combined to predict the final class label. Experimental results on two image datasets demonstrate that these proposed methods yield reliable and superior clustering performance.
Xu Tan 0004, Junqi Chen 0001, Jiawei Yang 0001, Sylwan Rahardja, Mou Wang, Susanto Rahardja
ICIP5
2024 Audiolog: LLMs-Powered Long Audio Logging with Hybrid Token-Semantic Contrastive Learning
abstract
Previous studies in automated audio captioning have faced difficulties in accurately capturing the complete temporal details of acoustic scenes and events within long audio sequences. This paper presents AudioLog, a large language models (LLMs)-powered audio logging system with hybrid token-semantic contrastive learning. Specifically, we propose to fine-tune the pre-trained hierarchical token-semantic audio Transformer by incorporating contrastive learning between hybrid acoustic representations. We then leverage LLMs to generate audio logs that summarize textual descriptions of the acoustic environment. Finally, we evaluate the AudioLog system on two datasets with both scene and event annotations. Experiments show that the proposed system achieves exceptional performance in acoustic scene classification and sound event detection, surpassing existing methods in the field. Further analysis of the prompts to LLMs demonstrates that AudioLog can effectively summarize long audio sequences1. To the best of our knowledge, this approach is the first attempt to leverage LLMs for summarizing long audio sequences.
Jisheng Bai, Han Yin, Mou Wang, Dongyuan Shi, Woon-Seng Gan, Susanto Rahardja
ICME3
2024 IAM-ACGAN: A High-Accuracy Approach for SAR Image Augmentation
abstract
Limited by the scarcity of synthetic aperture radar (SAR) systems, image augmentation is of great significance to SAR image detection, target recognition, and other application fields. However, traditional image augmentation methods rarely consider the SAR imaging mechanism, resulting in the inability to accurately reflect the anisotropic characteristics of target scattering. This paper introduces a novel SAR image augmentation method based on rebooting auxiliary classifier generative adversarial networks (Re-ACGAN), named IAM-ACGAN (Integrating Attention Mechanism with ACGAN). In this scheme, IAM-ACGAN integrates two attention mechanisms, channel attention (CA) and spatial attention (SA), into the discriminator of the GAN backbone to enhance classification accuracy. These two mechanisms can enhance the channel and spatial features of the input SAR images respectively. A self-constructed simulation ship dataset and a MSTAR real dataset both demonstrate the effectiveness of IAM-ACGAN. Compared with ACGAN and Re-ACGAN augmentation methods, IAM-ACGAN can provide higher image generation accuracy.
Shunjun Wei, Yifei Hu, Mou Wang, Xiaoling Zhang 0002, Yuanyuan Zhou 0007
IGARSS4
2024 Contrastive learning for deep tone mapping operator
Di Li 0006, Mou Wang, Susanto Rahardja
Signal Process. Image Commun.2
2024 Smoothed Frame-Level SINR and Its Estimation for Sensor Selection in Distributed Acoustic Sensor Networks
abstract
Distributed acoustic sensor network (DASN) refers to a sound acquisition system that consists of a collection of microphones randomly distributed across a wide acoustic area. Theory and methods for DASN are gaining increasing attention as the associated technologies can be used in a broad range of applications to solve challenging problems. However, unlike traditional microphone arrays or centralized systems, properly exploiting the redundancy among different channels in DASN is facing many challenges including but not limited to variations in pre-amplification gains, clocks, sensors' response, and signal-to-interference-plus-noise ratios (SINRs). Selecting appropriate sensors relevant to the task at hand is therefore crucial in DASN. In this work, we propose a speaker-dependent smoothed frame-level SINR estimation method for sensor selection in multi-speaker scenarios, specifically addressing source movement within DASN. Additionally, we devise an approach for similarity measurement to generate dynamic speaker embeddings resilient to variations in reference speech levels. Furthermore, we introduce a novel loss function that integrates classification and ordinal regression within a unified framework. Extensive simulations are performed and the results demonstrate the efficacy of the proposed method in accurately estimating smoothed frame-level SINR dynamically, yielding state-of-the-art performance.
Shanzheng Guan, Mou Wang, Zhongxin Bai, Jianyu Wang 0007, Jingdong Chen, Jacob Benesty
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 Non-Line-of-Sight Sparse Aperture ISAR Imaging via a Novel Detail-Aware Regularization
abstract
Non-line-of-sight (NLOS) moving target imaging is an emerging and challenging technology with potential applications in autonomous driving, security detection, disaster response, and more. In this article, a novel algorithm dubbed NLOS detail recovery via alternating direction method of multipliers (NDR-ADMMs) is proposed for NLOS moving target imaging. In our scheme, the static clutter filter (SCF) we proposed is utilized for NLOS clutter suppression, which facilitates hidden motion target echo extraction. To address the sparsity of echoes caused by scene complexity and target motion, we introduce a regularization constraint termed detail-aware regularization (DAR), which enhances details and suppresses noise in NLOS scenes by incorporating information from neighboring cells and expanding the receptive field of the image. Then, we propose the NDR-ADMM that combines DAR,$\ell _{1}$-norm, and ADMM to reconstruct high-resolution NLOS moving target images. Further, the corresponding fast version, NDR-ADMM+, is derived by mapping the NDR-ADMM to the adaptive parameter learning network for improving robustness and convergence. Finally, the proposed NDR-ADMM and NDR-ADMM+ are verified by simulated data and measured data we collected via millimeter-wave (MMW) radar in various NLOS scenarios. Compared to other state-of-the-art methods, NDR-ADMM+ demonstrates superior performance, robustness, and noise immunity, with NDR-ADMM following closely behind. This is attributed to DAR’s ability to capture details and suppress interference. Additionally, NDR-ADMM+ and AF-AMPnet offer the fastest processing speeds.
Yanbo Wen, Shunjun Wei, Xiang Cai, Yifei Hu, Mou Wang, Guolong Cui, Xiuhe Li, Jinhe Ran
IEEE Trans. Geosci. Remote. Sens.5
2024 CTV-Net: Complex-Valued TV-Driven Network With Nested Topology for 3-D SAR Imaging
abstract
regularization model is hindered by their hypothesis of inherent sparsity, causing unreal estimations of surface-like targets. Inspired by the edge-preserving property of total variation (TV), we propose a new complex-valued TV (CTV)-driven interpretable neural network with nested topology, i.e., CTV-Net, for 3-D SAR imaging. In our scheme, based on the 2-D holography imaging operator, the CTV-driven optimization model is constructed to pursue precise estimations in weakly sparse scenarios. Subsequently, a nested algorithmic framework, i.e., complex-valued TV-driven fast iterative shrinkage thresholding (CTV-FIST), is derived from the theory of proximal gradient descent (PGD) and FIST algorithm, theoretically supporting the design of CTV-Net. In CTV-Net, the trainable weights are layer-varied and functionally relevant to the hyperparameters of CTV-FIST, which aims to constrain the algorithmic parameters to update in a well-conditioned tendency. All weights are learned by end-to-end training based on a two-term cost function, which bounds the measurement fidelity and TV norm simultaneously. Under the guidance of the SAR signal model, a reasonably sized training set is generated, by randomly selecting reference images from the MNIST set and consequently synthesizing complex-valued label signals. Finally, the methodology is validated, numerically and visually, by extensive SAR simulations and real-measured experiments, and the results demonstrate the viability and efficiency of the proposed CTV-Net in the cases of recovering 3-D SAR images from incomplete echoes.
Mou Wang, Shunjun Wei, Zichen Zhou, Jun Shi 0002, Xiaoling Zhang 0002, Yongxin Guo 0002
IEEE Trans. Neural Networks Learn. Syst.1
2023 3D Audio Signal Processing Systems for Speech Enhancement and Sound Localization and Detection
abstract
The L3DAS23 of ICASSP Signal Processing Grand Challenge encourages research on 3D audio signal processing, such as 3D speech enhancement (SE) and 3D sound localization and detection (SELD). In this paper, we propose a two-stage system based on DPRNN and UNet for the SE task and a Conformer-based system for the SELD task. The proposed SE and SELD systems are evaluated on the L3DAS23 blind test sets. Results show that the proposed methods achieve state-of-the-art performance for 3D SE and SELD.
Jisheng Bai, Siwei Huang, Han Yin, Yafei Jia, Mou Wang
ICASSP5
2023 Correction: A lightweight classification of adaptor proteins using transformer networks
Sylwan Rahardja, Mou Wang, Binh P. Nguyen, Pasi Fränti, Susanto Rahardja
BMC Bioinform.2
2023 End-to-End Multi-Modal Speech Recognition on an Air and Bone Conducted Speech Corpus
abstract
Automatic speech recognition (ASR) has been significantly improved in the past years. However, most robust ASR systems are based on air-conducted (AC) speech, and their performances in low signal-to-noise-ratio (SNR) conditions are not satisfactory. Bone-conducted (BC) speech is intrinsically insensitive to environmental noise, and therefore can be used as an auxiliary source for improving the performance of an ASR at low SNR. In this paper, we first develop a multi-modal Mandarin corpus, which contains air- and bone-conducted synchronized speech (ABCS). The multi-modal speeches are recorded with a headset equipped with both AC and BC microphones. To our knowledge, it is by far the largest corpus for conducting bone conduction ASR research. Then, we propose a multi-modal conformer ASR system based on a novel multi-modal transducer (MMT). The proposed system extracts semantic embeddings from the AC and BC speech signals by a conformer-based encoder and a transformer-based truncated decoder. The semantic embeddings of the two speech sources are fused dynamically with adaptive weights by the MMT module. Experimental results demonstrate the proposed multi-modal system outperforms single-modal systems with either AC or BC modality and multi-modal baseline system by a large margin at various SNR levels. It also shows the two modalities complement with each other, and our method can effectively utilize the complementary information of different sources.
Mou Wang, Junqi Chen 0001, Xiao-Lei Zhang 0001, Susanto Rahardja
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 3-D SAR Imaging via Perceptual Learning Framework With Adaptive Sparse Prior
abstract
Mathematically, 3-D synthetic aperture radar (SAR) imaging is a typical inverse problem, which, by nature, can be solved by applying the theory of sparse signal recovery. However, many reconstruction algorithms are constructed by exploring the inherent sparsity of imaging space, which may cause unsatisfactory estimations in weakly sparse cases. To address this issue, we propose a new perceptual learning framework, dubbed as PeFIST-Net, for 3-D SAR imaging, by unfolding the fast iterative shrinkage-thresholding algorithm (FISTA) and exploring the sparse prior offered by the convolutional neural network (CNN). We first introduce a pair of approximated sensing operators in lieu of the conventional sensing matrices, by which the computational efficiency is highly improved. Then, to improve the reconstruction accuracy in inherently nonsparse cases, a mirror-symmetric CNN structure is designed to explore an optimal sparse representation of roughly estimated SAR images. The network weights control the hyperparameters of FISTA by elaborated regularization functions, ensuring a well-behaved updating tendency. Unlike directly using pixelwise loss function in existing unfolded networks, we introduce the perceptual loss by defining loss term based on high-level features extracted from the pretrained VGG-16 model, which brings higher reconstruction quality in terms of visual perception. Finally, the methodology is validated on simulations and measured SAR experiments. The experimental results indicate that the proposed method can obtain well-focused SAR images from highly incomplete echoes while maintaining fast computational speed.
Mou Wang, Shunjun Wei, Jun Shi 0002, Xiaoling Zhang 0002, Yongxin Guo 0002
IEEE Trans. Geosci. Remote. Sens.1
2022 End-To-End Multi-Modal Speech Recognition with Air and Bone Conducted Speech
abstract
Improving the performance of automatic speech recognition (ASR) in adverse acoustic environments is a long-term tough task. Although many robust ASR systems based on conventional microphones have been developed, their performance with air-conducted (AC) speech is still far from satisfactory in low signal-to-noise-ratio (SNR) environments. Bone-conducted (BC) speech is relatively insensitive to ambient noise, and has a potential of promoting the ASR performance at such low SNR environments as an auxiliary source. In this paper, we propose a conformer-based multi-modal speech recognition system. It uses a conformer encoder and a transformer-based truncated decoder to extract the semantic information from AC and BC channels respectively. The semantic information of the two channels are re-weighted and integrated by a novel multi-modal transducer. Experimental results show the effectiveness of the proposed method. For example, given a 0 dB SNR environment, it yields a character error rate of over 59.0% lower than a noise-robust baseline conducted on AC channel only, and over 12.7% lower than a multi-modal baseline that takes the concatenated features of AC and BC speech as the input.
Junqi Chen 0001, Mou Wang, Xiao-Lei Zhang 0001, Zhiyong Huang 0001, Susanto Rahardja
ICASSP2
2022 A lightweight classification of adaptor proteins using transformer networks
abstract
BACKGROUND: Adaptor proteins play a key role in intercellular signal transduction, and dysfunctional adaptor proteins result in diseases. Understanding its structure is the first step to tackling the associated conditions, spurring ongoing interest in research into adaptor proteins with bioinformatics and computational biology. Our study aims to introduce a small, new, and superior model for protein classification, pushing the boundaries with new machine learning algorithms. RESULTS: We propose a novel transformer based model which includes convolutional block and fully connected layer. We input protein sequences from a database, extract PSSM features, then process it via our deep learning model. The proposed model is efficient and highly compact, achieving state-of-the-art performance in terms of area under the receiver operating characteristic curve, Matthew's Correlation Coefficient and Receiver Operating Characteristics curve. Despite merely 20 hidden nodes translating to approximately 1% of the complexity of previous best known methods, the proposed model is still superior in results and computational efficiency. CONCLUSIONS: The proposed model is the first transformer model used for recognizing adaptor protein, and outperforms all existing methods, having PSSM profiles as inputs that comprises convolutional blocks, transformer and fully connected layers for the use of classifying adaptor proteins.
Sylwan Rahardja, Mou Wang, Binh P. Nguyen, Pasi Fränti, Susanto Rahardja
BMC Bioinform.2
2022 Perceptual Loss-Constrained Adversarial Autoencoder Networks for Hyperspectral Unmixing
abstract
Recently, the use of a deep autoencoder-based method in blind spectral unmixing has attracted great attention as the method can achieve superior performance. However, most autoencoder-based unmixing methods use non-structured reconstruction loss to train networks, leading to the ignorance of band-to-band-dependent characteristics and fine-grained information. To cope with this issue, we propose a general perceptual loss-constrained adversarial autoencoder network for hyperspectral unmixing. Specifically, the adversarial training process is used to update our framework. The discriminate network is found to be efficient in discovering the discrepancy between the reconstructed pixels and their corresponding ground truth. Moreover, the general perceptual loss is combined with the adversarial loss to further improve the consistency of high-level representations. Ablation studies verify the effectiveness of the proposed components of our framework, and experiments with both synthetic and real data illustrate the superiority of our framework when compared with other competing methods.
Min Zhao 0014, Mou Wang, Jie Chen 0022, Susanto Rahardja
IEEE Geosci. Remote. Sens. Lett.2
2022 Lightweight FISTA-Inspired Sparse Reconstruction Network for mmW 3-D Holography
abstract
Integrating compressed sensing (CS) with millimeter-wave (mmW) holography has shown great potential to achieve lightweight onboard hardware, low sampling ratio, and high-speed sensing. However, conventional CS-driven algorithms are always limited by nontrivial adjusting of parameters and excessive computational cost caused by plenty of iterations. To address this problem, we propose a lightweight model-based deep learning framework (LFIST-Net) for mmW 3-D holography, by combining the interpretability of fast iterative shrinkage-thresholding algorithm (FISTA) and tuning-free merit of data-driven deep neural network. First, the single-frequency (SF) holographic imaging technique is integrated into FISTA, which serves as the sensing kernels, to avoid large-scale matrix multiplications. Subsequently, the kernel-based FISTA (KFISTA) is mapped into layer-fixed and parameter-learnable LFIST-Net, whose weights are relaxed to be layer-varied. The updating of key parameters in LFIST-Net, including step sizes, thresholds, and momentum coefficients, are regularized by soft-plus function to ensure the non-negativity and monotonicity. As for 3-D holography implementation, the “1-D + 2-D” scheme is adopted, where the matched filtering (MF) and well-trained LFIST-Net are used for range focusing and reconstructions of azimuth slices. Without losing efficiency, the range-focused subechoes are processed parallelly in 3-D cube form. Experiments, including both simulated and measured tests based on a commercial mmW radar, prove that LFIST-Net is capable of reconstructing the imaging scene precisely. In particular, in near-field mmW 3-D holography tests, both numerical and visual results demonstrate LFIST-Net yields compelling reconstruction performance while maintaining high computational speed compared with MF-based, conventional CS-driven, and network-based methods.
Mou Wang, Shunjun Wei, Jiadian Liang, Jun Shi 0002, Xiaoling Zhang 0002
IEEE Trans. Geosci. Remote. Sens.1
2022 RMIST-Net: Joint Range Migration and Sparse Reconstruction Network for 3-D mmW Imaging
abstract
Compressed sensing (CS) demonstrates significant potential to improve image quality in 3-D millimeter-wave imaging compared with conventional matched filtering (MF). However, existing sparsity-driven 3-D imaging algorithms always suffer from large-scale storage, excessive computational cost, and nontrivial tuning of parameters due to the huge-dimensional matrix–vector multiplication in complicated iterative optimization steps. In this article, we present a novel range migration (RM) kernel-based iterative-shrinkage thresholding network, dubbed as RMIST-Net, by combining the traditional model-based CS method and data-driven deep learning method for near-field 3-D millimeter-wave (mmW) sparse imaging. First, the measurement matrices in ISTA optimization steps are replaced by RM kernels, by which matrix–vector multiplication is converted to the Hadamard product. Then, the modified ISTA optimization is unrolled into a deep hierarchical architecture, in which all parameters are learned automatically instead of manually tuned. Subsequently, 1000 pairs of oracle images with randomly distributed targets and their corresponding echoes are simulated to train the network. A well-trained RMIST-Net produces high-quality 3-D images from range-focused echoes. Finally, we experimentally prove that RMIST-Net is capable process$512 \times 512$large-scale imaging tasks within 1 s. Besides, we compare RMIST-Net with other state-of-the-art methods in near-field 3-D imaging applications. Both simulations and real-measured experiments demonstrate that RMIST-Net produces impressive reconstruction performance while maintaining high computational speed compared with conventional and sparse imaging algorithms.
Mou Wang, Shunjun Wei, Jiadian Liang, Xiangfeng Zeng, Chen Wang 0041, Jun Shi 0002, Xiaoling Zhang 0002
IEEE Trans. Geosci. Remote. Sens.1
2022 Efficient ADMM Framework Based on Functional Measurement Model for mmW 3-D SAR Imaging
abstract
Compressed sensing (CS) shows significant potential in the field of active millimeter-wave (mmW) synthetic aperture radar (SAR) imaging due to the merits of reducing system complexity and achieving high-speed sensing. However, most CS-driven imaging methods suffer from the excessive computational burden, since the calculative steps always rely on vectorization and consequently lead to extremely large-scale matrix operations. To address this issue, we propose an efficient alternating direction method of multipliers (ADMMs) framework for mmW 3-D SAR imaging. In our scheme, we utilize the single-frequency holographic (SFH) technique and construct SFH-based forward/inverse sensing operators rather than converting the imaging process into a special case of “linear inverse problems,” by which the large-scale matrix inversions are avoided and consequently the computational complexity is reduced. Based on the SFH functional measurement model, the SFH-ADMM is derived to reconstruct the 3-D image from sparsely sampled measurement echo while suppressing noisy clutters and ambiguities. Besides, the SFH-ADMM iteration steps undergird a neural network design, yielding a tailored SFH-ADMM-Net with trainable parameters and layer-fixed structures, which further shorten the execution time and improve reconstruction performance. The network is trained by simulated data, which are generated according to the radar signal model. Extensive experiments, including simulations and laboratory tests, demonstrate the superiority of the proposed algorithms in terms of both reconstruction accuracy and computational speed.
Mou Wang, Shunjun Wei, Zichen Zhou, Jun Shi 0002, Xiaoling Zhang 0002
IEEE Trans. Geosci. Remote. Sens.1
2022 3-D SAR Data-Driven Imaging via Learned Low-Rank and Sparse Priors
abstract
In the research topic of three-dimensional (3D) SAR imaging, the sparsity-enforcing techniques offer promise in shortening sensing time and improving reconstruction accuracy. However, many of them only explore the sparse prior of 3D SAR images, which leads to biased estimations in cases of non-sparse scenarios. To remedy this problem, we propose a new network with learned low-rank and sparse priors, i.e., LLRS-Net, to obtain improved reconstructions from sparsely sampled 3D SAR echoes. In our scheme, a two-stage reconstruction algorithmic framework (LSRA) is derived based on sparse and low-rank priors. Wherein, the first stage recovers the measurements from their limited observations by exploring the low-rank prior, while the second estimates the final 3D SAR images with a fast-iterative optimization. Theoretically inspired by LRSA, the LLRS-Net is designed into a cascaded network structure. In LLRS-Net, the trainable weights serve as independent variables and control the algorithmic hyper-parameters via regularizing functions, ensuring a well-conditioned updating tendency. By end-to-end training, the network weights are updated automatically under the guidance of a compound loss function constraining both the outputs of two stages. Finally, the methodology is validated on simulations and measured experiments. These results show that the proposed framework outperforms many state-of-the-art imaging algorithms in recovering 3D SAR images from incomplete echo data.
Mou Wang, Shunjun Wei, Zichen Zhou, Jun Shi 0002, Xiaoling Zhang 0002, Yongxin Guo 0002
IEEE Trans. Geosci. Remote. Sens.1
2022 3-D SAR Autofocusing With Learned Sparsity
abstract
Inevitable inaccuracies of 3-D synthetic aperture radar (3-D SAR) imaging geometry may cause undesired blurs in reconstructed images. Recent advances show impressive results in integrating error estimation into sparse imaging. However, the concept is still challenging in 3-D SAR due to the cumbersome high-dimensional processing. To address this problem, we propose a model-driven 3-D SAR autofocusing network with learned sparsity (AFLS-Net) by applying the recent emerging deep unfolding technique. In our scheme, we first construct a kernel-based observation model with consideration of motion-induced phase errors, which avoids the memory-consuming matrix calculations in the conventional matrix–vector form. Then, a joint sparse imaging and autofocusing algorithm is derived based on the framework of block coordinate descent. In addition, by mapping the computational steps, the AFLS-Net is designed to further improve the autofocusing accuracy and efficiency in which a shallow two-path convolutional neural network (CNN) is embedded to explore the implicit sparse prior, by which the reconstruction accuracy can be improved. Meanwhile, the batchwise autofocusing module is designed to obtain a robust estimation by jointly optimizing subcost functions associated with a batch of independent measurements. Finally, the methodology is validated in both simulations and laboratory 3-D SAR experiments. The experimental results suggest that the proposed method obtains better autofocusing quality compared to other comparison baselines in reconstructing 3-D SAR images from incomplete and error-polluted echoes.
Mou Wang, Shunjun Wei, Zichen Zhou, Jun Shi 0002, Xiaoling Zhang 0002, Yongxin Guo 0002
IEEE Trans. Geosci. Remote. Sens.1
2022 AF-AMPNet: A Deep Learning Approach for Sparse Aperture ISAR Imaging and Autofocusing
abstract
Inverse synthetic aperture radar (ISAR) imaging and autofocusing are challenging under sparse aperture (SA) conditions. Traditional imaging or autofocusing methods fail to obtain satisfying results due to the nonuniform and incomplete data caused by SA. To address this problem, a novel compressive sensing (CS)-based imaging and autofocusing framework is proposed to obtain high cross-range resolution for SA ISAR. To achieve well-focused imaging results of better performance and higher efficiency simultaneously, we merge the phase error estimation into the CS framework, then iteratively solve the compound CS problem in matrix form with approximate message-passing (AMP), dubbed as AF-AMP. Moreover, a deep learning approach is also proposed by mapping AF-AMP into a deep network, dubbed as AF-AMPNet, with extensive modifications to further improve the efficiency. The adaptively and layer-wisely optimal parameters learned by the training process are also promising to enhance the performance and robustness against noise. Besides, the loss function for training is subjoined with regularized$\ell _{1} $and$\ell _{2} $constraints to ensure the sparsity and quality of imaging results. Furthermore, the proposed AF-AMP and corresponding network-based AF-AMPNet are verified by simulated and measured experiments, both of which show superior performance, robustness, and higher efficiency than other state-of-the-art methods. AF-AMPNet can achieve the best performance in much less computational time.
Shunjun Wei, Jiadian Liang, Mou Wang, Jun Shi 0002, Xiaoling Zhang 0002, Jinhe Ran
IEEE Trans. Geosci. Remote. Sens.3
2022 Nonline-of-Sight 3-D Imaging Using Millimeter-Wave Radar
abstract
Nonline-of-sight (NLOS) radar imaging is a novel technique that can inverse the scattering characteristics of targets in the NLOS area, which has been one of the hot pots of radar imaging field. However, the existing NLOS radar mainly focuses on 1-D or 2-D imaging, which inevitably suffers from the geometric loss of real 3-D scenes, and its applications are restricted in the urban environment. In this article, we propose an NLOS radar 3-D imaging model and method for looking around corner (LAC) situation by multi-input–multioutput (MIMO) millimeter-wave (mmW) array antennas. In this scheme, first, the model of NLOS radar 3-D imaging with mmW MIMO antennas is established and the multipath scattering of targets with this model is analyzed. Then, the theoretical resolution of LAC 3-D imaging is derived and discussed. Second, exploiting the three bounces of LAC and extraction of linear structure, an effective imaging algorithm with mirror projection theory and Radon transform, dubbed as mirror symmetry backprojection (MSBP), is proposed for 3-D image focusing. Moreover, to suppress the uncertainties of phase caused by both LAC and system error, the minimum entropy principle is introduced to MSBP. Finally, an NLOS 3-D imaging system with 77-GHz mmW MIMO radio frequency module and 2-D rails is developed. Different types of targets, such as metal balls and ornaments, are tested in LAC. The results demonstrate that our NLOS technique can not only provide a high-quality 3-D focusing of the hidden targets but also extract positions of targets without prior knowledge of the NLOS area.
Shunjun Wei, Jinshan Wei, Xinyuan Liu 0002, Mou Wang, Xiaoling Zhang 0002, Jun Shi 0002, Guolong Cui
IEEE Trans. Geosci. Remote. Sens.4
2022 Learning-Based Split Unfolding Framework for 3-D mmW Radar Sparse Imaging
abstract
The application of the compressed sensing (CS) method in the radar field enables the radar imaging system to satisfy both low data cost and high reconstruction quality, however, it is accompanied by enormous iterative operations and difficult adjustments of parameters. In this paper, we propose a learning-based split unfolding framework, dubbed as split iterative sparse reconstruction network (SISR-Net), for near-field 3-D millimeter-wave (mmW) radar sparse imaging. Firstly, a sparse reconstruction algorithm, i.e., SISRA, is proposed to theoretically guide the structure of the imaging framework. Subsequently, by combining the model-based CS method and data-driven deep learning method, SISR-Net is constructed by SISRA to produce 3-D mmW radar images efficiently with excellent explainability and generalization ability. Joint the radar-imaging kernel, echo-generation kernel, and the split Bregman method, the efficiency and stability of SISR-Net are guaranteed, all parameters are layer-varied and learned steadily by end-to-end training to improve the convergence and robustness of the imaging network. Simulated data and the echo from a high-resolution mmW radar dataset 3DRIED, are used to train and test the SISR-Net based on the Adam optimizer. For both simulation and extensive 3-D mmW radar measured experiments, the proposed SISR-Net outperforms other state-of-the-art imaging methods in terms of imaging accuracy and generalization ability.
Shunjun Wei, Zichen Zhou, Mou Wang, Hao Zhang 0103, Jun Shi 0002, Xiaoling Zhang 0002, Ling Fan
IEEE Trans. Geosci. Remote. Sens.3
2022 Hyperspectral Unmixing for Additive Nonlinear Models With a 3-D-CNN Autoencoder Network
abstract
Spectral unmixing is an important task in hyperspectral image processing for separating the mixed spectral data pertaining to various materials observed aiming at analyzing the material components in observed pixels. Recently, nonlinear spectral unmixing has received particular attention in hyperspectral image processing, as there are many situations in which the linear mixture model may not be appropriate and could be advantageously replaced by a nonlinear one. Existing nonlinear unmixing approaches are often based on specific assumptions on the nonlinearity and can be less effective when used for scenes with unknown nonlinearity. This article presents an unsupervised nonlinear spectral unmixing method that addresses a general model that consists of a linear mixture part and an additive nonlinear mixture part. The structure of a deep autoencoder network, which has a clear physical interpretation, is specifically designed to achieve this purpose. Moreover, a convolutional neural network (CNN) is used to capture the spectral-spatial priors from hyperspectral data. Extensive experiments with synthetic and real data illustrate the generality and effectiveness of this scheme compared with state-of-the-art methods.
Min Zhao 0014, Mou Wang, Jie Chen 0022, Susanto Rahardja
IEEE Trans. Geosci. Remote. Sens.2
2022 SAF-3DNet: Unsupervised AMP-Inspired Network for 3-D MMW SAR Imaging and Autofocusing
abstract
The sparse imaging method based on compressed sensing (CS) is widely used in the field of millimeter-wave (MMW) synthetic aperture radar (SAR) imaging. However, 3D sparse imaging is limited by the difficult parameter tuning, the huge computational load, and the low processing efficiency. In addition, due to the motion errors and model mismatch, it is difficult to obtain well-focused results without error correction techniques. To address these issues, we propose a deep learning framework that integrates 3D sparse imaging and autofocusing, named 3D Sparse Autofocusing Network (SAF-3DNet) for MMW SAR data processing. The network is constructed based on an auto-encoder, which can optimize parameters without effective ground truth. The backbone structure of the encoder is expanded by approximate message-passing (AMP), and the operators in the frequency domain are used to replace the traditional matrix-vector CS model, which avoids large-scale matrix multiplication and other operations, and greatly improves the operation efficiency. In addition, the 2D phase error estimation in the cross-range plane is embedded into the sparse imaging models, enabling simultaneous 3D imaging and autofocusing. The decoder is designed as a mapping from the autofocusing results to the echo data. Experimental results based on both simulated and measured data demonstrate the proposed SAF-3DNet can achieve well-focused 3D reconstruction within an ephemeral time, which expresses the potential of 3D MMW SAR real-time and high-quality imaging.
Zichen Zhou, Shunjun Wei, Hao Zhang 0103, Rong Shen, Mou Wang, Jun Shi 0002, Xiaoling Zhang 0002
IEEE Trans. Geosci. Remote. Sens.5
2021 Non-Line-Of-Sight Imaging by Millimeter Wave Radar
abstract
Non-line-of-sight (NLOS) radar imaging technique aims to reconstruct hidden targets that illuminated wave cannot reach directly, which can greatly expand the range of radar detection. In this paper, inspired by synthetic aperture radar (SAR), an effective two-dimensional (2-D) NLOS imaging technique via multiple input multiple output (MIMO) millimeter-wave (MMW) radar is proposed. In the scheme, a 2-D virtual antenna array is synthesized by MIMO antenna scanning, and the multi-bounces echoes is used to obtain 2-D NLOS imaging. Then an algorithm via mirror symmetry back-projection (MSBP) is presented for 2-D high-precision focusing of these NLOS echoes. Moreover, a cost-effective 79GHz MMW NLOS experiment system is developed for technical validation. The effectiveness of MMW NLOS radar imaging is verified by near-field multi-targets experiment, and high-precision 2-D imaging results of hidden knives are obtained by MSBP method.
Jinshan Wei, Shunjun Wei, Xinyuan Liu 0002, Mou Wang, Jun Shi 0002, Xiaoling Zhang 0002
IGARSS4
2021 TPSSI-Net: Fast and Enhanced Two-Path Iterative Network for 3D SAR Sparse Imaging
abstract
The emerging field of combining compressed sensing (CS) and three-dimensional synthetic aperture radar (3D SAR) imaging has shown significant potential to reduce sampling rate and improve image quality. However, the conventional CS-driven algorithms are always limited by huge computational costs and non-trivial tuning of parameters. In this article, to address this problem, we propose a two-path iterative framework dubbed TPSSI-Net for 3D SAR sparse imaging. By mapping the AMP into a layer-fixed deep neural network, each layer of TPSSI-Net consists of four modules in cascade corresponding to four steps of the AMP optimization. Differently, the Onsager terms in TPSSI-Net are modified to be differentiable and scaled by learnable coefficients. Rather than manually choosing a sparsifying basis, a two-path convolutional neural network (CNN) is developed and embedded in TPSSI-Net for nonlinear sparse representation in the complex-valued domain. All parameters are layer-varied and optimized by end-to-end training based on a channel-wise loss function, bounding both symmetry constraint and measurement fidelity. Finally, extensive SAR imaging experiments, including simulations and real-measured tests, demonstrate the effectiveness and high efficiency of the proposed TPSSI-Net.
Mou Wang, Shunjun Wei, Jiadian Liang, Zichen Zhou, Qizhe Qu, Jun Shi 0002, Xiaoling Zhang 0002
IEEE Trans. Image Process.1
2020 Kernel Rotational Network for Synthetic Aperture Radar Target Recognition
abstract
Convolutional Neural Networks (CNNs) have excellent ability in image recognition, however, the requirement of a large amount of labeled dataset limits its application in the field of synthetic aperture radar (SAR) image processing. In this paper, a kernel rotational network (KR-Net) for SAR target recognition is constructed. When the labeled dataset is small, the KR-net can achieve higher classification rate than standard CNNs benefit from its inherent rotational convolution units. Also, weights sharing strategy is introduced to increase network capacity without multiplying the number of weights parameters. Meanwhile, a simple and feasible multi-branch feature converging method for the KR-Net is proposed to fuse features of rotational convolution units. Experimental results show that our network can achieve state-of-art result in the MSTAR dataset, especially when the training set is small.
Yuanyuan Zhou 0007, Yao Hu 0006, Chen Wang 0041, Mou Wang, Jun Shi 0002, Shunjun Wei
IGARSS4
2020 ISAR Compressive Sensing Imaging Using Convolution Neural Network with Interpretable Optimization
abstract
Compressive Sensing(CS) has been widely utilized in Inverse synthetic aperture radar(ISAR) imaging since real ISAR data is easier to be non-completed, and CS-based methods can obtain high-quality imaging results using under-sampled data. However, traditional CS-based methods need pre-defined parameters, sparse transforms and iterative reconstruction processes. Optimal parameters as well as transforms are tough to be hand-crafted, and iterative reconstruction consumes plenty of time, which limit practical applications in ISAR imaging. Given that Convolution Neural Network(CNN) has great power to learn rapidly, we compose CNN with traditional Iterative Shrinkage-Thresholding Algorithm(ISTA) to propose CNN-ISTA(CIST)-based ISAR imaging method. CIST is capable of learning optimal parameters and transforms throughout the training (i.e. the optimization process is interpretable) instead of manually defined. Compared with traditional state-of-the-art CS imaging methods, the experimental results demonstrate that our proposed CIST-based imaging method is superior in both imaging quality and computational efficiency.
Jiadian Liang, Shunjun Wei, Mou Wang, Jun Shi 0002, Xiaoling Zhang 0002
IGARSS3
2020 Linear Array 3-D SAR Sparse Imaging via Convolutional Neural Network
abstract
Compressed sensing theory has attracted extensive attention in the field of linear array 3-D Synthetic Aperture Radar (SAR) sparse imaging. However, conventional CS-based algorithms always suffer from quite huge computational cost. In this paper, we propose a new method for 3-D SAR sparse imaging based on convolutional neural network (CNN). Inspired by the work of ISTA-NET, a complex-valued version for imaging tasks is modified. Furthermore, we introduce a approximate phase correction scheme for 3-D imaging, it makes the proposed method works with only a constant measurement matrix corresponding to any slice. Moreover, Using a random training strategy, ISTA-NET networks for 3-D SAR imaging are effectively trained. Experimental results demonstrate that the proposed method outperforms conventional ISTA large margins in both accuracy and speed.
Mou Wang, Shunjun Wei, Jun Shi 0002, Yue Wu 0028, Jiadian Liang, Qizhe Qu
IGARSS1
2020 Efficient Insar Imaging Based on Frequency-Domain Back Projection Algorithm
abstract
High resolution imaging of interferometric synthetic aperture radar (InSAR) usually requires fine focusing and phase-preserving. Time-domain back projection (TDBP) method outperforms other conventional methods at focusing and phase-preserving, but suffer from huge computational complexity when the underlying scene is large. In this article, an efficient method exploiting by frequency-domain back projection (FDBP) is presented for high-resolution InSAR imaging. In the scheme, the coherent integration of focusing is efficient achieved by frequency-domain Fourier transform, and a delayed-distance is compensated to phase-preserving of InSAR. Simulation and experiment results demonstrates that FDBP algorithm improves the computational efficiency by three times while maintaining the similar focusing accuracy compared with the conventional TDBP method.
Yue Wu 0028, Shunjun Wei, Mou Wang, Jiadian Liang, Xiaoling Zhang 0002
IGARSS3
2019 Nonlinear Unmixing of Hyperspectral Data via Deep Autoencoder Networks
abstract
Nonlinear spectral unmixing is an important and challenging problem in hyperspectral image processing. Classical nonlinear algorithms are usually derived based on specific assumptions on the nonlinearity. In recent years, deep learning shows its advantage in addressing general nonlinear problems. However, existing ways of using deep neural networks for unmixing are limited and restrictive. In this letter, we develop a novel blind hyperspectral unmixing scheme based on a deep autoencoder network. Both encoder and decoder of the network are carefully designed so that we can conveniently extract estimated endmembers and abundances simultaneously from the nonlinearly mixed data. Because an autoencoder is essentially an unsupervised algorithm, this scheme only relies on the current data and, therefore, does not require additional training. Experimental results validate the proposed scheme and show its superior performance over several existing algorithms.
Mou Wang, Min Zhao 0014, Jie Chen 0022, Susanto Rahardja
IEEE Geosci. Remote. Sens. Lett.1