Yipeng Liu 0001

dblp:26/6297-1 · DBLP profile ↗
← Back
76ranked-venue papers
7as first author
48since 2021 · last 2026
0000-0003-2084-8781ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 52 · 3 first-author · 29 since 2021Artificial intelligence and machine learning · 16 · 2 first-author · 12 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021
YearPublicationVenuePosition
2026 T-MLA: A targeted multiscale log-exponential attack framework for neural image compression
Nikolay I. Kalmykov, Razan Dibo, Kaiyu Shen, Zhonghan Xu, Anh Huy Phan 0001, Yipeng Liu 0001, Ivan V. Oseledets
Inf. Sci.6
2026 Coupled tensor train decomposition in federated learning
Xiangtao Zhang, Eleftherios Kofidis, Ruituo Wu, Ce Zhu, Le Zhang 0001, Yipeng Liu 0001
Pattern Recognit.6
2026 TERM Model: Tensor Ring Mixture Model for Density Estimation
abstract
Probabilistic modeling is a core challenge in statistical machine learning. Tensor-based probabilistic graph methods address interpretability and stability concerns encountered in neural network approaches and allow tractable inference (e.g., marginal inference and conditional inference). In this paper, we introduce tensor ring decomposition for density estimation, which reduces the number of permutation candidates compared to existing methods, while simultaneously enhancing expressive power and maintaining tractable inference. Different non-negative strategies for density function results in two variants: Born TRDE offers simpler inference and sampling but with slightly lower accuracy, while Energy TRDE, though more complex, achieves superior performance. Furthermore, a mixture model that incorporates multiple permutation candidates with adaptive weights is designed, resulting in increased expressive flexibility and comprehensiveness. Unlike existing methods that focus on finding a single optimal permutation, our approach, inspired by ensemble learning, demonstrates that combining multiple suboptimal permutations can yield superior results. Experiments demonstrate that the proposed approach excels in estimating probability density functions and sampling, capturing intricate details with competitive or superior performance compared to existing state-of-the-art (SOTA) tractable density methods.
Ruituo Wu, Jiani Liu 0002, Bing Li 0002, Anh Huy Phan 0001, Ivan V. Oseledets, Ce Zhu, Yipeng Liu 0001
IEEE Trans. Big Data7
2026 Content-Adaptive Unfolding Wavelet Transformer for Hyperspectral Image Super-Resolution
abstract
In recent years, fusing high-resolution multispectral images (HR-MSIs) and low-resolution hyperspectral images (LR-HSIs) has become a widely used approach for hyperspectral image super-resolution (HSI-SR). The deep unfolding framework has attracted significant attention thanks to its ability to formulate the problem into a data module and a prior module. However, there are still two critical issues that hinder the performance enhancement of the existing methods: 1) Parameters in the data module are fixed (though learnable) at each iteration, i.e., lacking the adaptivity to comprehensive data; 2) The Transformer in the prior module cannot effectively capture high-frequency information. To resolve these issues, we propose a Content-Adaptive Unfolding Wavelet Transformer (CAUWT) for HSI-SR, where the parameters are adaptively learned based on the reconstructed HSI at each iteration. Moreover, we propose a novel Wavelet-Assisted Transformer (WAT), by integrating the Discrete Wavelet Transform (DWT) and the Hybrid Spectral-Spatial Attention Block (HSSAB) to further upgrade the high-frequency information quality of HSI at no cost of extra branch structures, where the former is for multi-scale and multi-frequency details and the latter is for correlations between and within sub-band components. Extensive experiments performed on both simulated and real datasets well demonstrate the effectiveness of the proposed method. In comparison with mainstream HSI-SR methods, our method exhibits superior performance and lower computational overhead.
Yipeng Liu 0001, Zhen Long, Chong-Yung Chi, Ce Zhu
IEEE Trans. Image Process.2
2026 FunOTTA: On-the-Fly Adaptation on Cross-Domain Fundus Image via Stable Test-Time Training
abstract
Fundus images are essential for the early screening and detection of eye diseases. While deep learning models using fundus images have significantly advanced the diagnosis of multiple eye diseases, variations in images from different imaging devices and locations (known as domain shifts) pose challenges for deploying pre-trained models in real-world applications. To address this, we propose a novel Fundus On-the-fly Test-Time Adaptation (FunOTTA) framework that effectively generalizes a fundus image diagnosis model to unseen environments, even under strong domain shifts. FunOTTA stands out for its stable adaptation process by performing dynamic disambiguation in the memory bank while minimizing harmful prior knowledge bias. We also introduce a new training objective during adaptation that enables the classifier to incrementally adapt to target patterns with reliable class conditional estimation and consistency regularization. We compare our method with several state-of-the-art test-time adaptation (TTA) pipelines. Experiments on cross-domain fundus image benchmarks across two diseases demonstrate the superiority of the overall framework and individual components under different backbone networks. Code is available at https://github.com/Casperqian/FunOTTA.
Le Zhang 0001, Yipeng Liu 0001, Ce Zhu, Fan Zhang 0013
IEEE Trans. Medical Imaging3
2025 SLR-MVTC: Smooth Low-Rank Multi-View Tensor Clustering
abstract
Multi-view tensor clustering (MVTC) has gained much attention for its effectiveness in capturing global high-order correlations across views. However, current MVTC methods suffer from two limitations: 1) adopting a two-stage process to learn the latent features for clustering, and 2) either ignoring local similarities within views or treating local similarities and global high-order correlations equally. In this paper, we propose a smooth low-rank MVTC (SLR-MVTC) method, which aims to extract latent features that are smooth within each view and low-rank across views, enhancing clustering performance. Specifically, we first learn latent features from each view using orthogonal projection and then construct the latent feature tensor by concatenation and rotation. Then, we introduce a new smooth tensor nuclear norm to depict the low-rank components of the low-frequency parts in the feature tensor. Benefiting from the fast Fourier transform along the sample dimension, the obtained low-frequency components effectively capture local smoothness within views, while their low-rank parts further explore global correlations across views. Experimental results on six multi-view datasets demonstrate that SLR-MVTC outperforms state-of-the-art algorithms in terms of clustering performance and CPU time.
Zhen Long, Yipeng Liu 0001, Yazhou Ren 0001, Ce Zhu
AAAI2
2025 Subspace Constraint and Contribution Estimation for Heterogeneous Federated Learning
abstract
Heterogeneous Federated Learning (HFL) has received widespread attention due to its adaptability to different models and data. The HFL approach utilizing auxiliary models for knowledge transfer can further enhance flexibility. However, existing frameworks face the challenges of local overfitting and aggregation bias. To address these issues, we propose FedSCE. By restricting specific layers of the local model updates to a subspace, FedSCE reduces the degrees of freedom of the update, enhances generalization, and mitigates the risk of overfitting. The subspace is dynamically updated to ensure coverage of the latest model update trajectory. Additionally, FedSCE evaluates client contributions based on the update distance of the auxiliary model in feature space and parameter space, achieving adaptive weighted aggregation. We validate our approach in both feature-skewed and label-skewed scenarios, demonstrating that on Office10, our method exceeds the best baseline by 3.87%. The code will be available at https://github.com/AVC2-UESTC/FedSCE.git.
Xiangtao Zhang, Ao Li 0007, Yipeng Liu 0001, Fan Zhang 0013, Ce Zhu, Le Zhang 0001
CVPR4
2025 Codar: Complex-valued Neural Network for Crossing-Floor Intrusion Detection via WiFi
abstract
WiFi systems offer enormous potential for device-free human intrusion detection. Current methods often require routers to be deployed in multiple adjacent rooms on the same floor, which is redundant and costly. To solve this, we introduce the first work on intrusion detection in the crossing-floor scenario via WiFi. Routers on different floors are utilized without major modifications to the existing router layout. Many previous works require a high sample rate and ignore the phase information. In this paper, we propose Codar, a complex-valued LSTM-CNN neural network. The LSTM effectively captures temporal dependencies at a low sample rate in harsh propagation environments. Moreover, amplitude and phase features are explored jointly by complex-valued operations. Experimental results demonstrate Codar achieves 95%, 94.5%, and 99% accuracy for intrusion detection, user identification, and intruded floor identification, surpassing competitive methods. The code and dataset are available at https://github.com/ouweiting/Codar.
Weiting Ou, Yipeng Liu 0001, Bing Li 0002, Le Zhang 0001, Ce Zhu
ICASSP2
2025 Incrementally Constrained Tucker Decomposition for Feature Extraction of Structural Diffusion Tensor Imaging Data
abstract
Diffusion Tensor Imaging (DTI) is the only in vivo technique capable of characterizing microstructural changes in the brain. The resulting feature maps, such as fractional anisotropy (FA), are three-dimensional and contain spatial details. Processing these feature maps without disrupting their structure is essential for accurate analysis. Tucker decomposition is a widely used feature extraction method for high-order data. However, it has been rarely adopted for characterizing structural DTI data. In addition, few work systematically studies the influence of its constraints on data characterization. In this study, we design the Incrementally Constrained Tucker Decomposition (ICTD) framework that progressively applies orthogonality and non-negativity constraints on decomposed factors to characterize DTI data and determine suitable constraints. The entanglement entropy is introduced to evaluate the entanglement of extracted features. Two public DTI datasets are adopted in the classification experiments. Our results demonstrate that Tucker decomposition is suitable for characterizing DTI data and the constraints are important for effective data characterization.
Houji Du, Fan Zhang 0013, Yipeng Liu 0001, Ce Zhu
ICME4
2025 Unified Line Segment Detection and Description
abstract
Line segments are fundamental elements in computer vision. However, aside from a few computationally expensive deep learning-based methods, most existing approaches treat their detection and description as independent tasks, leading to redundant computations and suboptimal performance. This paper introduces a Unified approach for Line Segment Detection and Description (ULSD2), designed for real-time vision tasks with minimal computational overhead. The core insight is to unify line segment detection and description by analyzing dedicated level lines and their differences, derived from gradients, which effectively capture the intrinsic characteristics of line segments. Furthermore, instead of the traditional scalar-based description, the use of level lines and their differences in local patches across multiple granularities enables a vectorized representation that encodes line segments from coarse to fine. Experiments demonstrate that ULSD2 outperforms other non-deep learning-based methods and competes with state-of-the-art deep learning-based methods while significantly improving efficiency. The code is available at https://github.com/roylin1229/ULSD2.
Yingjie Zhou 0001, Zhen Long, Yipeng Liu 0001, Lu Yang 0002, Ce Zhu
ICME4
2025 TRR-LGF: a Simple yet Efficient Classification Network
abstract
Hybrid models that combine convolution and self attention are popular for efficient local feature extraction and capturing long-range dependencies. However, these models often:1) only explore local and global features; 2) flatten high-order features at the output layer, which limit feature hierarchy exploration and the feature utility in the output layer. To address these issues, this paper introduces Tensor Ring Regression with Local-to-Global Features (TRR-LGF), a simple and effective classification network. It uses a local-to-global learning framework to capture diverse features at multiple scales. Additionally, a tensor ring regression layer replaces the linear output layer, preserving high-order feature structure and reducing parameters. Experimental results show that TRR-LGF outperforms existing state-of-the-art methods on various datasets, especially in noisy and sample-imbalanced settings. Furthermore, the model utilizes 7.6M parameters, reducing computational requirements by about 50% compared to the multilayer perceptron output layer. The code is available at https://github.com/Calcium-Oxide/TRR-LGF.
Zhen Long, Hu Yao, Yipeng Liu 0001, Le Zhang 0001, Ce Zhu
ICME4
2025 Flexibly Constrained Tucker Decomposition for High-Order Spectral Analysis
abstract
Spectral analysis is widely used for sequence signals, such as electroencephalogram (EEG), speech, and radar signals, by calculating spectra using Fourier or other transforms for feature extraction. Recently, high-order spectral analysis, which models spectra into a high-order tensor and characterizes it through tensor decomposition, has gained popularity. Existing decomposition models often impose prior constraints on data. The most popular one is the orthogonality. However, enforcing strict orthogonality might fail to capture the intrinsic structure of data with complex structures. To achieve flexible data analysis, we propose flexibly constrained Tucker decomposition (FCTD) model. FCTD imposes a relaxed orthogonality constraint for each data dimension, allowing the orthogonal strength of the decomposed factors to be controlled. In addition, FCTD can be easily transformed into models with mixed constraints. Two publicly available datasets for epilepsy detection and emotional speech classification are adopted to demonstrate the effectiveness of the proposed FCTD. FCTD achieves the best performance among all models with fixed constraints.
Houji Du, Nipon Theera-Umpon, Yipeng Liu 0001, Ce Zhu
MMSP4
2025 TLRLF4MVC: Tensor Low-Rank and Low-Frequency for Scalable Multi-View Clustering
abstract
Anchor-based multi-view clustering has garnered much attention for its effectiveness in handling massive datasets. However, current methods either fail to consider intra-view similarity or require ($\mathcal {O}(N^{3})$O(N3)) for exploring intra-view similarity, making efficient large-scale multi-view clustering difficult. This paper introduces a novel tensor low-frequency component (TLFC) operator, which achieves smooth representation among samples. Furthermore, this TLFC operator, which explores intra-view similarity, incorporates tensor nuclear norm (TNN) operator and consensus regularization that explore inter-view correlations, resulting in the development of tensor low-rank and low-frequency for scalable multi-view clustering (TLRLF4MVC). Iteratively, as intra-view sample similarity and complementary information across views achieve balance, the learned embedding features are mapped into a smooth and compact subspace, ultimately leading to outstanding clustering performance. Extensive experiments on six large-scale multi-view datasets demonstrate that TLRLF4MVC not only significantly outperforms state-of-the-art methods in terms of clustering accuracy but also achieves remarkable computational efficiency, particularly when handling massive data.
Zhen Long, Yazhou Ren 0001, Yipeng Liu 0001, Ce Zhu
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 HOPE: Enhanced Position Image Priors via High-Order Implicit Representations
abstract
Deep Image Prior (DIP) has shown that networks with stochastic initialization and custom architectures can effectively address inverse imaging challenges. Despite its potential, DIP requires significant computational resources, whereas the lighter Implicit Neural Positional Image Prior (PIP) often yields overly smooth solutions due to exacerbated spectral bias. Research on lightweight, high-performance solutions for inverse imaging remains limited. This paper proposes a novel framework, Enhanced Positional Image Priors through High-Order Implicit Representations (HOPE), incorporating high-order interactions between layers within a conventional cascade structure. This approach reduces the spectral bias commonly seen in PIP, enhancing the model's ability to capture both low- and high-frequency components for optimal inverse problem performance. We theoretically demonstrate that HOPE's expanded representational space, narrower convergence range, and improved Neural Tangent Kernel (NTK) diagonal properties enable more precise frequency representations than PIP. Comprehensive experiments across tasks such as signal representation (audio, image, volume) and inverse image processing (denoising, super-resolution, CT reconstruction, inpainting) confirm that HOPE establishes new benchmarks for recovery quality and training efficiency.
Ruituo Wu, Junhui Hou, Ce Zhu, Yipeng Liu 0001
IEEE Trans. Image Process.5
2025 DA-Flow: Dual Attention Normalizing Flow for Skeleton-Based Video Anomaly Detection
abstract
Cooperation between temporal convolutional networks (TCN) and graph convolutional networks (GCN) as a processing module has shown promising results in skeleton-based video anomaly detection (SVAD). However, to maintain a lightweight model with low computational and storage complexity, shallow GCN and TCN blocks are constrained by small receptive fields and a lack of cross-dimension interaction capture. To tackle this limitation, we propose a lightweight module called the Dual Attention Module (DAM) for capturing cross-dimension interaction relationships in spatio-temporal skeletal data. It employs the frame attention mechanism to identify the most significant frames and the skeleton attention mechanism to capture broader relationships across fixed partitions with minimal parameters and total Floating Point Operations (FLOPs). Furthermore, the proposed Dual Attention Normalizing Flow (DA-Flow) integrates the DAM as a post-processing unit after GCN within the normalizing flow framework. Simulations show that the proposed model is robust against noise and negative samples. Experimental results show that DA-Flow reaches competitive or better performance than the existing state-of-the-art (SOTA) methods in terms of the micro AUC metric with the fewest parameters and FLOPs. Moreover, we found that even without training, simply using random projection without dimensionality reduction on skeleton data enables substantial anomaly detection capabilities.
Ruituo Wu, Bing Li 0002, Jicong Fan 0001, Frédéric Dufaux, Ce Zhu, Yipeng Liu 0001
IEEE Trans. Multim.8
2025 Online Nonconvex Robust Tensor Principal Component Analysis
abstract
Robust tensor principal component analysis (RTPCA) based on tensor singular value decomposition (t-SVD) separates the low-rank component and the sparse component from the multiway data. For streaming data, online RTPCA (ORTPCA) processes tensor data sequentially, where the low-rank component is updated based on the latest estimation and the newly arrived sample. It enhances both computation and storage efficiency. However, in most of the existing ORTPCA methods, the relaxation from tensor multirank to the convex tensor nuclear norm (TNN) may have a certain modeling error, which leads to unavoidable tracking accuracy loss. In this article, a tensor Schatten-p norm ( $0\lt p\lt 1$ ) is applied to provide a tighter approximation of the tensor rank. A Lemma is deduced to divide the Schatten-p norm into terms to be updated in an online way. Based on it, the corresponding online nonconvex RTPCA (ONRTPCA) method is proposed for efficient tensor subspace tracking. Moreover, we incorporate the dynamic forgetting window into ONRTPCA to adaptively track varying subspaces. In addition, this article also provides convergence analysis and complexity analysis. Experimental results on synthetic data and real-world video data show that our proposed method achieves superior subspace tracking accuracy in comparison with a series of state-of-the-art methods while maintaining a high convergence speed and low memory requirement.
Lanlan Feng, Yipeng Liu 0001, Ce Zhu
IEEE Trans. Neural Networks Learn. Syst.2
2024 S2MVTC: A Simple Yet Efficient Scalable Multi-View Tensor Clustering
abstract
Anchor-based large-scale multi-view clustering has attracted considerable attention for its effectiveness in handling massive datasets. However, current methods mainly seek the consensus embedding feature for clustering by exploring global correlations between anchor graphs or projection matrices. In this paper, we propose a simple yet efficient scalable multi-view tensor clustering (S2MVTC) approach, where our focus is on learning correlations of embedding features within and across views. Specifically, we first construct the embedding feature tensor by stacking the embedding features of different views into a tensor and rotating it. Additionally, we build a novel tensor low-frequency approximation (TLFA) operator, which incorporates graph similarity into embedding feature learning, efficiently achieving smooth representation of embedding features within different views. Furthermore, consensus constraints are applied to embedding features to ensure inter-view semantic consistency. Experimental results on six large-scale multi-view datasets demonstrate that S2MVTC significantly outperforms state-of-the-art algorithms in terms of clustering performance and CPU execution time, especially when handling massive data. The code of S2MVTC is publicly available at https://github.com/longzhen520/S2MVTC.
Zhen Long, Yazhou Ren 0001, Yipeng Liu 0001, Ce Zhu
CVPR4
2024 Phase Retrieval by Tensor Total Least Squares
abstract
Phase retrieval seeks to reconstruct a series of image sequences from measurements that only capture their magnitudes. Current approaches either flatten and stack the image sequences, disregarding their multidimensional structural information, or fail to account for errors within the sensing vectors/tensors. To address these two issues simultaneously, we propose a unified framework for the phase retrieval problem, namely tensor total least squares (TTLS). Specifically, we set up a tensor representation for image sequences and the corresponding measurement model, and for the first time employ the advanced tensor ring network to effectively explore the inherent multidimensional structure for more accurate estimation. Moreover, in addition to the additive noise, the multiplicative errors within the sensing tensor can be also well-corrected, leading to a more robust estimation. Experimental results on both simulated data and real videos demonstrate the superiority of the proposed method.
Jiani Liu 0002, Ce Zhu, Xiaolin Huang, Yipeng Liu 0001
ICASSP5
2024 Multi-Band Speech Tensor Decomposition for Interactive Feature Extraction in Early Dysphagia Screening
abstract
Dysphagia is a prevalent symptom in numerous neurological disorders among older adults. Current dysphagia diagnostic systems either involve invasive procedures or necessitate the ingestion of liquids. Some researchers have devised automatic dysphagia detection methods based on vowels that are easy to collect and sensitive to vocal cord states. These methods extract features from each vowel separately and fuse them to train models. Nonetheless, they neglect potential interrelations among different vowels. Vowels collected from the same speaker could share subspaces since they are produced from the same vocal system. In this study, we introduce a tensor-based method that can simultaneously extract interactive information from all vowels across different modes. This method designs multi-band speech tensors and core-pruned tensor networks to investigate crucial frequency bands and connections for dysphagia screening. Experimental results show our model exceeds previous methods by approximately 10 percentage points in the classification accuracy.
Yipeng Liu 0001, Da Shen, Yangyang Jiang, Ce Zhu
ICASSP2
2024 Efficient Black-Box Adversarial Attack on Deep Clustering Models
abstract
Despite the significant progress made by deep clustering models in high-dimensional data processing, they remain vulnerable to adversarial examples. However, research on adversarial attacks against deep clustering algorithms appears to be relatively underexplored. To fill this gap, we propose a query-efficient black-box attack on deep clustering models, which leverages the transferability between different deep clustering models. Initially, we train a generator using a substitute deep clustering model, reducing the number of queries to the target model. Subsequently, when targeting an unknown deep clustering model, we employ the target query information to update both the substitute deep clustering model and the generator. Experimental evaluations on four state-of-the-art deep clustering models across three datasets demonstrate the efficacy of our method in disrupting clustering performance. The results indicate that our approach surpasses the performance of existing methods.
Zhen Long, Xiaolin Huang, Ce Zhu, Yipeng Liu 0001
ICIP6
2024 Epilepsy Detection with Personal Identification Based on Regularized O-minus Decomposition
abstract
Epilepsy is a common neurological disease that seriously affects the patient’s life quality. Electroencephalogram (EEG) is an important modality for epilepsy diagnosis and treatment. Here, we propose a tensor-based method using EEG, which can detect patients with seizures and identify them. In this method, we transform EEG signals into high-order tensors and design a regularized O-minus tensor network to calculate representative signals of channels. The interactive information among all tensor modes can be extracted by the regularized O-minus. Finally, the brain network is calculated based on the extracted representative signals and used as a shared feature set for epilepsy detection and patient identity. Leave one subject out cross-validation is used in seizure detection to verify the generalization of the model for first-time patients. The accuracies of seizure detection and patient identification are 88.14% and 98.59%, respectively. To our knowledge, this is the first time that epilepsy detection and patient identification are implemented within a unified framework, which could provide timely and personalized treatment for patients.
Da Shen, Zhongrong Wang, Ce Zhu, Yipeng Liu 0001
ISCAS6
2024 A Comprehensive Review of Image Line Segment Detection and Description: Taxonomies, Comparisons, and Challenges
abstract
An image line segment is a fundamental low-level visual feature that delineates straight, slender, and uninterrupted portions of objects and scenarios within images. Detection and description of line segments lay the basis for numerous vision tasks. Although many studies have aimed to detect and describe line segments, a comprehensive review is lacking, obstructing their progress. This study fills the gap by comprehensively reviewing related studies on detecting and describing two-dimensional image line segments to provide researchers with an overall picture and deep understanding. Based on their mechanisms, two taxonomies for line segment detection and description are presented to introduce, analyze, and summarize these studies, facilitating researchers to learn about them quickly and extensively. The key issues, core ideas, advantages and disadvantages of existing methods, and their potential applications for each category are analyzed and summarized, including previously unknown findings. The challenges in existing methods and corresponding insights for potentially solving them are also provided to inspire researchers. In addition, some state-of-the-art line segment detection and description algorithms are evaluated without bias, and the evaluation code will be publicly available. The theoretical analysis, coupled with the experimental results, can guide researchers in selecting the best method for their intended vision applications. Finally, this study provides insights for potentially interesting future research directions to attract more attention from researchers to this field.
Yingjie Zhou 0001, Yipeng Liu 0001, Ce Zhu
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 TS-RTPM-Net: Data-Driven Tensor Sketching for Efficient CP Decomposition
abstract
Tensor decomposition is widely used in feature extraction, data analysis, and other fields. As a means of tensor decomposition, the robust tensor power method based on tensor sketch (TS-RTPM) can quickly mine the potential features of tensor, but in some cases, its approximation performance is limited. In this paper, we propose a data-driven framework called TS-RTPM-Net, which improves the estimation accuracy of TS-RTPM by jointly training the TS value matrices with the RTPM initial matrices. It also uses two greedy initialization algorithms to optimize the TS location matrices. In addition, TS-RTPM-Net accelerates TS-RTPM by using fast power iteration modules. Comparative experiments on real-world datasets verify that TS-RTPM-Net outperforms TS-RTPM in terms of estimation accuracy, running speed, and memory consumption.
Xingyu Cao, Xiangtao Zhang, Ce Zhu, Jiani Liu 0002, Yipeng Liu 0001
IEEE Trans. Big Data5
2024 CS2DIPs: Unsupervised HSI Super-Resolution Using Coupled Spatial and Spectral DIPs
abstract
In recent years, fusing high spatial resolution multispectral images (HR-MSIs) and low spatial resolution hyperspectral images (LR-HSIs) has become a widely used approach for hyperspectral image super-resolution (HSI-SR). Various unsupervised HSI-SR methods based on deep image prior (DIP) have gained wide popularity thanks to no pre-training requirement. However, DIP-based methods often demonstrate mediocre performance in extracting latent information from the data. To resolve this performance deficiency, we propose a coupled spatial and spectral deep image priors (CS2DIPs) method for the fusion of an HR-MSI and an LR-HSI into an HR-HSI. Specifically, we integrate the nonnegative matrix-vector tensor factorization (NMVTF) into the DIP framework to jointly learn the abundance tensor and spectral feature matrix. The two coupled DIPs are designed to capture essential spatial and spectral features in parallel from the observed HR-MSI and LR-HSI, respectively, which are then used to guide the generation of the abundance tensor and spectral signature matrix for the fusion of the HSI-SR by mode-3 tensor product, meanwhile taking some inherent physical constraints into account. Free from any training data, the proposed CS2DIPs can effectively capture rich spatial and spectral information. As a result, it exhibits much superior performance and convergence speed over most existing DIP-based methods. Extensive experiments are provided to demonstrate its state-of-the-art overall performance including comparison with benchmark peer methods.
Yipeng Liu 0001, Chong-Yung Chi, Zhen Long, Ce Zhu
IEEE Trans. Image Process.2
2024 Adaptively Topological Tensor Network for Multi-View Subspace Clustering
abstract
Multi-view subspace clustering employs learned self-representation from multiple tensor decompositions to exploit the low-rank information. However, the data structures embedded with self-representation tensors may vary in different multi-view datasets. Therefore, a pre-defined decomposition may not fully exploit low-rank information from various data, resulting in sub-optimal multi-view clustering performance. To alleviate this, we proposed the adaptively topological tensor network (ATTN). ATTN can learn a suitable decomposition structure that can represent the low-rank structure and high-order correlation of the self-representation tensors better in a data-driven way, which can capture the intra-view and inter-view information better. Firstly, instead of connecting the tensor network blindly, ATTN utilizes the correlation between adjacent factors to prune redundant connections from the fully connected tensor networks, making the tensor network more expressive. Furthermore, a greedy adaptive rank-increasing strategy is applied to optimize the pruned tensor network structure, which improves the capacity of capturing low-rank structure. We apply ATTN on a multi-view subspace clustering task and utilize the alternating direction method of multipliers(ADMM) method to optimize it. Experiments show that multi-view subspace clustering based on ATTN has better performance on nine multi-view datasets.
Yipeng Liu 0001, Jie Chen 0086, Yingcong Lu, Weiting Ou, Zhen Long, Ce Zhu
IEEE Trans. Knowl. Data Eng.1
2024 Feature Space Recovery for Efficient Incomplete Multi-View Clustering
abstract
T-SVD based incomplete multi-view clustering (IMVC) has received wide attention due to its ability to capture high-order correlations. However, t-SVD suffers from rotation sensitivity, failing to fully explore both inter- and intra-view consistencies. Besides, current methods mainly consider inter- or intra-view correlations, ignoring the low-rank information of sample features within views. To address these weaknesses, we first propose a feature space recovery based IMVC (FSR-IMVC) method, where low-rank feature space recovery and low-rank tensor ring based consistency learning are considered into a unified framework. Furthermore, we extend FSR-IMVC by incorporating anchor learning on the latent feature space, resulting in a scalable FSR-IMVC (sFSR-IMVC) approach that is well-suited to large-scale data. In an iterative way, the learned inter- and intra-view correlations will guide the recovery of missing features, while the explored low-rank information from feature spaces will in turn facilitate consistency exploration, eventually achieving outstanding clustering performance. Experimental results show that FSR-IMVC provides a significant improvement over known state-of-the-art algorithms in terms of ACC, NMI and Purity. Compared with FSR-IMVC, sFSR-IMVC performs slightly worse in clustering accuracy, but offers a notable advantage in computational efficiency, particularly for large-scale datasets. The codes of FSR-IMVC and sFSR-IMVC are publicly available athttps://github.com/longzhen520/sFSR-IMVC.
Zhen Long, Ce Zhu, Pierre Comon, Yazhou Ren 0001, Yipeng Liu 0001
IEEE Trans. Knowl. Data Eng.5
2024 Multi-View MERA Subspace Clustering
abstract
Tensor-based multi-view subspace clustering (MSC) can capture high-order correlation in the self-representation tensor. Current tensor decompositions for MSC suffer from highly unbalanced unfolding matrices or rotation sensitivity, failing to fully explore inter/intra-view information. Using the advanced tensor network, namely, multi-scale entanglement renormalization ansatz (MERA), we propose a low-rank MERA based MSC (MERA-MSC) algorithm, where MERA factorizes a tensor into contractions of one top core factor and the rest orthogonal/semi-orthogonal factors. Benefiting from multiple interactions among orthogonal/semi-orthogonal (low-rank) factors, the low-rank MERA has a strong representation power to capture the complex inter/intra-view information in the self-representation tensor. The alternating direction method of multipliers is adopted to solve the optimization model. Experimental results on five multi-view datasets demonstrate MERA-MSC has superiority against the compared algorithms on six evaluation metrics. Furthermore, we extend MERA-MSC by incorporating anchor learning and develop a scalable low-rank MERA based multi-view clustering method (sMREA-MVC). To our knowledge, this is the first work to introduce MERA to the multi-view clustering topic. The effectiveness and efficiency of sMERA-MVC have been validated on three large-scale multi-view datasets.
Zhen Long, Ce Zhu, Jie Chen 0086, Yazhou Ren 0001, Yipeng Liu 0001
IEEE Trans. Multim.6
2023 Level-Line Guided Edge Drawing for Robust Line Segment Detection
abstract
Line segment detection plays a cornerstone role in computer vision tasks. Among numerous detection methods that have been recently proposed, the ones based on edge drawing attract increasing attention owing to their excellent detection efficiency. However, the existing methods are not robust enough due to the inadequate usage of image gradients for edge drawing and line segment fitting. Based on the observation that the line segments should locate on the edge points with both consistent coordinates and level-line information, i.e., the unit vector perpendicular to the gradient orientation, this paper proposes a level-line guided edge drawing for robust line segment detection (GEDRLSD). The level-line information provides potential directions for edge tracking, which could be served as a guideline for accurate edge drawing. Additionally, the level-line information is fused in line segment fitting to improve the robustness. Numerical experiments show the superiority of the proposed GEDRLSD1algorithm compared with state-of-the-art methods.
Yingjie Zhou 0001, Yipeng Liu 0001, Ce Zhu
ICASSP3
2023 Efficient and Effective Multi-Camera Pose Estimation with Weighted M-Estimate Sample Consensus
abstract
Camera pose estimation is a fundamental module for many vision tasks. It is usually based on feature correspondences, i.e., feature matches across different images. However, correspondences always contain non-negligible outliers, which may negatively affect pose estimation efficiency and accuracy. This paper proposes a multi-camera pose estimation method by leveraging point and line correspondences with non-negligible outliers, in which a weighted M-Estimate Sample Consensus (w-MSAC) based on the customized weights and the coarse pose prior is introduced to improve the efficiency and accuracy of pose estimation. The customized weights could decrease the iterations of the pose hypothesis and improve the pose estimation accuracy. The coarse pose prior is used to perform the pre-validation of the pose hypothesis, eliminating many unnecessary validations. Experiments demonstrate the superiority of the proposed w-MSAC1over existing state-of-the-art methods, e.g., improving 22% positioning and 24% orientation accuracy meanwhile decreasing 15% iterations and 92% validations than the MSAC.
Yingjie Zhou 0001, Xun Zhang 0002, Yipeng Liu 0001, Ce Zhu
ICASSP4
2023 Tensorized LSSVMS For Multitask Regression
abstract
Multitask learning (MTL) can utilize the relatedness between multiple tasks for performance improvement. The advent of multimodal data allows tasks to be referenced by multiple indices. High-order tensors are capable of providing efficient representations for such tasks, while preserving structural task-relations. In this paper, a new MTL method is proposed by leveraging low-rank tensor analysis and constructing tensorized Least Squares Support Vector Machines, namely the tLSSVM-MTL, where multilinear modelling and its nonlinear extensions can be flexibly exerted. We employ a high-order tensor for all the weights with each mode relating to an index and factorize it with CP decomposition, assigning a shared factor for all tasks and retaining task-specific latent factors along each index. Then an alternating algorithm is derived for the nonconvex optimization, where each resulting subproblem is solved by a linear system. Experimental results demonstrate promising performances of our tLSSVM-MTL.
Jiani Liu 0002, Qinghua Tao, Ce Zhu, Yipeng Liu 0001, Johan A. K. Suykens
ICASSP4
2023 Feature Space Recovery for Incomplete Multi-View Clustering
abstract
Incomplete multi-view clustering (IMVC), based on imputation and clustering unification, has received wide attention due to its ability to exploit hidden information from missing views. However, current methods mainly consider inter/intra-view correlations, ignoring the structural information of sample features within views. In this paper, we propose a feature space recovery based IMVC method, where low-rank feature space recovery and consensus representation learning of inter/intra-views are considered into a unified framework. Moreover, low-rank tensor ring approximation is used to capture the correlations of self-representation tensor. In an iterative way, the learned inter/intra-view correlations will guide the recovery of missing features, while the explored low-rank information from feature spaces will in turn facilitate self-representation learning, eventually achieving out-standing clustering performance. Experimental results show our method has a very significant improvement over known state-of-the-art algorithms in terms of ACC, NMI and Purity.
Zhen Long, Ce Zhu, Pierre Comon, Yipeng Liu 0001
ICASSP4
2023 Optimal Low-Rank Tensor Tree Completion
abstract
Tensor completion is a powerful technique for recovering missing entries from partial observations. Tensor tree network, with a hierarchical structure, has gained widespread attention for its ability to balance and effectively explore the correlations in high-order data. However, the performance of tensor trees is influenced by the order of modes, leading to variations in their effectiveness. To address this, we propose to minimize the loss of entanglement entropy to determine the optimal mode order within the tensor tree network, thereby optimizing its representation performance. We correspondingly construct an optimal low-rank tensor tree completion model, where the optimal low-rank tensor tree network captures global structures and total variation investigates local structures. The alternating direction method of multipliers is employed to solve the optimization problem. Experimental results on color images and light field images demonstrate that our method outperforms state-of-the-art algorithms in terms of recovery performance.
Ce Zhu, Zhen Long, Yipeng Liu 0001
MMSP4
2023 A dedicated benchmark for contour-based corner detection evaluation
Yingjie Zhou 0001, Yipeng Liu 0001, Ce Zhu
Image Vis. Comput.3
2023 Low Dimensional Trajectory Hypothesis is True: DNNs Can Be Trained in Tiny Subspaces
abstract
Deep neural networks (DNNs) usually contain massive parameters, but there is redundancy such that it is guessed that they could be trained in low-dimensional subspaces. In this paper, we propose a Dynamic Linear Dimensionality Reduction (DLDR) based on the low-dimensional properties of the training trajectory. The reduction method is efficient, supported by comprehensive experiments: optimizing DNNs in 40-dimensional spaces can achieve comparable performance as regular training over thousands or even millions of parameters. Since there are only a few variables to optimize, we develop an efficient quasi-Newton-based algorithm, obtain robustness to label noise, and improve the performance of well-trained models, which are three follow-up experiments that can show the advantages of finding such low-dimensional subspaces. The code is released (Pytorch: https://github.com/nblt/DLDR and Mindspore: https://gitee.com/mindspore/docs/tree/r1.6/docs/sample_code/dimension_reduce_training).
Tao Li 0054, Zhehao Huang, Qinghua Tao, Yipeng Liu 0001, Xiaolin Huang
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Level Line Guided Interest Point Detection
abstract
Detection of interest points,e.g., corners and blobs, lays the foundation for many vision tasks. Numerous methods have been proposed to improve the detection performance, and the gradient-based ones are the most investigated. However, existing gradient-based methods lack an adequate utilization of gradient orientations. In this letter, we show that the level line,i.e., the unit vector orthogonal to the gradient orientation of a specific point, is particularly important to interest point detection. The support level lines of an interest point,i.e., the level lines used to identify corners/blobs, exhibit a significantly different pattern from those of other points. Based on this observation, this letter proposes two robust interest point detectors for finding corners and blobs, respectively. For each detector, a specific type of level line difference is defined, and the corresponding differences are leveraged with different weights. Numerical experiments show the superior performance of the proposed detectors. The code will be publicly available athttps://github.com/roylin1229/LLD-IP.
Yingjie Zhou 0001, Yipeng Liu 0001, Ce Zhu
IEEE Signal Process. Lett.3
2023 Multiplex Transformed Tensor Decomposition for Multidimensional Image Recovery
abstract
Low-rank tensor completion aims to recover the missing entries of multi-way data, which has become popular and vital in many fields such as signal processing and computer vision. It varies with different tensor decomposition frameworks. Compared with matrix SVD, recently emerging transform t-SVD can better characterize the low-rank structure of order-3 data. However, it suffers from rotation sensitivity, and dimensional limitation (i.e., only effective for order-3 tensors). To alleviate these deficiencies, we develop a novel multiplex transformed tensor decomposition (MTTD) framework, which can characterize the global low-rank structure along all modes for any order- N tensor. Based on MTTD, we propose a related multi-dimensional square model for low-rank tensor completion. Besides, a total variation term is also introduced to utilize the local piecewise smoothness of the tensor data. The classic alternating direction method of multipliers is used to solve the convex optimization problems. For performance testing, we choose three linear invertible transforms including FFT, DCT, and a group of unitary transform matrices for our proposed methods. The simulated and real-data experiments demonstrate the superior recovery accuracy and computational efficiency of our method compared with state-of-the-art ones.
Lanlan Feng, Ce Zhu, Zhen Long, Jiani Liu 0002, Yipeng Liu 0001
IEEE Trans. Image Process.5
2022 Long-term Visual Localization Using Illumination Insensitive Descriptors
abstract
This demo shows a long-term visual localization system based on illumination insensitive descriptors (IID) of points and lines in multiple cameras. The system can robustly match the features in captured images for localization against those in localization database (DB). The developed localization system achieves remarkable performance and seasonal-time-spanned localization results in complex and changing environments.
Yingjie Zhou 0001, Yipeng Liu 0001, Ce Zhu
MMSP3
2022 Dysphagia diagnosis system with integrated speech analysis from throat vibration
Hengling Zhao, Yangyang Jiang, Shenghan Wang, Fangzhou Ren, Ce Zhu, Jirong Yue, Yipeng Liu 0001
Expert Syst. Appl.11
2022 Multi-Scale Spatial and Temporal Speech Associations to Swallowing for Dysphagia Screening
abstract
Dysphagia is a common symptom of many neurological diseases. It often occurs in older adults and increases the risk of aspiration pneumonia. Existing diagnosis systems of dysphagia are invasive or require patients to swallow liquids, which are costly and harmful to the patients. In this work, we propose an early screening system of dysphagia based on two kinds of throat signals, i.e., vowels and sentences. Based on the vowels, two new categories of speech features are developed: PET (pitch/energy trajectory) and FS-Conts (full spectrogram contours). The PET feature set focuses on the prominent resonance energy of speech to track the pitch and energy fluctuations. It can reflect the stability of vocal cords in the speech generation process. The FS-Conts feature set is proposed to emphasize the spatial details of formants based on three-dimensional contours. Concerning the sentences, three categories of speech features are proposed, called LSSDL (log symmetric spectral difference level), C-coes (crucial energy coefficients), and LDF (local dynamic features). The three features explore the speech representations of dysphagia from global variations to local associations. The LSSDL feature set is designed to highlight the global spectral differences in the interested frequency region. The C-coes and LDF feature sets locate local speech differences in specific frequency regions and time duration. In addition, a new feature selection algorithm is developed based on a newly designed precise matching analysis technique to search for distinguishing features. In the classification experiments, the SVM classifier is adopted and the dysphagia detection accuracy reaches 95.07%. The comparative experiments are conducted. The results indicate that our system performs better than the existing methods.
Ce Zhu, Yipeng Liu 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2022 Smooth Compact Tensor Ring Regression
abstract
In learning tasks with high order correlations, the low-rank approximation of the regression coefficient tensor has become increasingly important. Tensor ring can capture more correlation information among tensor networks. However, its optimal rank is generally unknown and needs to be tuned from multiple combinations. To address the issue, we propose a novel tensor regression framework with a group sparsity constraint on latent factors for tensor ring rank estimation. Specifically, the proposed group sparsity term constrained matrix factorization problem is first shown to be equivalent to a better approximation of matrix rank, namely Schatten-$1/2$quasi-norm. Extending it into tensor, the tensor ring rank can be inferred during the learning process to balance the prediction error and the model complexity. Besides, a total variation term is introduced to enhance the local consistency of the predicted response, which is useful for reducing the adverse effects of random noise. Experiments on the simulation dataset show that the proposed method can exactly obtain the tensor ring rank, and the effectiveness and robustness of the proposed algorithm is further verified on a real dataset for human motion capture tasks.
Jiani Liu 0002, Ce Zhu, Yipeng Liu 0001
IEEE Trans. Knowl. Data Eng.3
2021 Multi-mode Tensor Singular Value Decomposition for Low-Rank Image Recovery
Lanlan Feng, Ce Zhu, Yipeng Liu 0001
ICIG (2)3
2021 Deep Learning-Based Regional Sub-models Integration for Parkinson's Disease Diagnosis Using Diffusion Tensor Imaging
Hengling Zhao, Chih-Chien Tsai, Ce Zhu, Mingyi Zhou, Jiun-Jie Wang, Yipeng Liu 0001
ICIG (2)6
2021 Smart Dysphagia Detection System with Adaptive Boosting Analysis of Throat Signals
abstract
Dysphagia is a symptom of many neurological disorders. Existing diagnosis systems are either invasive or require swallowing liquids, which are costly and harmful to humans. In this work, we design a smart dysphagia detection system based on speech signals. Rather than the voice data acquired by traditional microphones, we apply a bone conduction headset for vibration signal acquisition from the throat to get cleaner speech signals. After speech feature extraction, under-sampling is performed to deal with the imbalanced data problem, and principal component analysis is used for dimensionality reduction. In this paper, we construct an ensemble adaptive boosting classifier to detect the dysphagia patient. Experimental results show that the testing classification accuracy of the proposed system reaches 71.2%. Sensitivity and specificity can reach 66.6 % and 76 %, respectively.
Shenghan Wang, Yangyang Jiang, Hengling Zhao, Ce Zhu, Yipeng Liu 0001
ISCAS8
2021 Deep Learning Based Gait Analysis for Contactless Dementia Detection System from Video Camera
abstract
Dementia is a neurodegenerative disease with a high incidence in the elderly. However, there is no effective treatment for this disease, and early intervention has a great effect to slow the deterioration. Currently, the detection of dementia is mainly achieved using questionnaire-like neuropsychological tests. Such ways usually cost a lot of time. To this end, we design a contactless dementia detection system based on gait analysis from surveillance video, and it can serve as a home-based healthcare system. This system applies a Kinect 2.0 camera to capture the human video and extract the skeleton joints at a rate of 15 frames per second. Two different gaits are collected for detection, namely single-task gait and dual-task gait. In this paper, we design a convolutional neural network based classifier to extract features in a data-driven way from these two groups of videos, but not take hand-crafted features. Experimental results show that we achieve a sensitivity of 74.10% on the test set using this system, and the processing only takes several minutes for early dementia detection.
Yangyang Jiang, Xingyu Cao, Ce Zhu, Yipeng Liu 0001
ISCAS7
2021 Low-rank tensor ring learning for multi-linear regression
Jiani Liu 0002, Ce Zhu, Zhen Long, Huyan Huang, Yipeng Liu 0001
Pattern Recognit.5
2021 A Simple Local Minimal Intensity Prior and an Improved Algorithm for Blind Image Deblurring
abstract
Blind image deblurring is a long standing challenging problem in image processing and low-level vision. Recently, sophisticated priors such as dark channel prior, extreme channel prior, and local maximum gradient prior, have shown promising effectiveness. However, these methods are computationally expensive. Meanwhile, since these priors involved subproblems cannot be solved explicitly, approximate solution is commonly used, which limits the best exploitation of their capability. To address these problems, this work firstly proposes a simplified sparsity prior of local minimal pixels, namely patch-wise minimal pixels (PMP). The PMP of clear images is much more sparse than that of blurred ones, and hence is very effective in discriminating between clear and blurred images. Then, a novel algorithm is designed to efficiently exploit the sparsity of PMP in deblurring. The new algorithm flexibly imposes sparsity inducing on the PMP under the maximum a posterior (MAP) framework rather than directly uses the half quadratic splitting algorithm. By this, it avoids non-rigorous approximation solution in existing algorithms, while being much more computationally efficient. Extensive experiments demonstrate that the proposed algorithm can achieve better practical stability compared with state-of-the-arts. In terms of deblurring quality, robustness and computational efficiency, the new algorithm is superior to state-of-the-arts. Code for reproducing the results of the new method is available at https://github.com/FWen/deblur-pmp.git.
Fei Wen 0005, Rendong Ying, Yipeng Liu 0001, Trieu-Kien Truong
IEEE Trans. Circuits Syst. Video Technol.3
2021 Bayesian Low Rank Tensor Ring for Image Recovery
abstract
Low rank tensor ring based data recovery can recover missing image entries in signal acquisition and transformation. The recently proposed tensor ring (TR) based completion algorithms generally solve the low rank optimization problem by alternating least squares method with predefined ranks, which may easily lead to overfitting when the unknown ranks are set too large and only a few measurements are available. In this article, we present a Bayesian low rank tensor ring completion method for image recovery by automatically learning the low-rank structure of data. A multiplicative interaction model is developed for low rank tensor ring approximation, where sparsity-inducing hierarchical prior is placed over horizontal and frontal slices of core factors. Compared with most of the existing methods, the proposed one is free of parameter-tuning, and the TR ranks can be obtained by Bayesian inference. Numerical experiments, including synthetic data, real-world color images and YaleFace dataset, show that the proposed method outperforms state-of-the-art ones, especially in terms of recovery accuracy.
Zhen Long, Ce Zhu, Jiani Liu 0002, Yipeng Liu 0001
IEEE Trans. Image Process.4
2021 AMP-Net: Denoising-Based Deep Unfolding for Compressive Image Sensing
abstract
Most compressive sensing (CS) reconstruction methods can be divided into two categories, i.e. model-based methods and classical deep network methods. By unfolding the iterative optimization algorithm for model-based methods onto networks, deep unfolding methods have the good interpretation of model-based methods and the high speed of classical deep network methods. In this article, to solve the visual image CS problem, we propose a deep unfolding model dubbed AMP-Net. Rather than learning regularization terms, it is established by unfolding the iterative denoising process of the well-known approximate message passing algorithm. Furthermore, AMP-Net integrates deblocking modules in order to eliminate the blocking artifacts that usually appear in CS of visual images. In addition, the sampling matrix is jointly trained with other network parameters to enhance the reconstruction performance. Experimental results show that the proposed AMP-Net has better reconstruction accuracy than other state-of-the-art methods with high reconstruction speed and a small number of network parameters.
Yipeng Liu 0001, Jiani Liu 0002, Fei Wen 0005, Ce Zhu
IEEE Trans. Image Process.2
2020 DaST: Data-Free Substitute Training for Adversarial Attacks
abstract
Machine learning models are vulnerable to adversarial examples. For the black-box setting, current substitute attacks need pre-trained models to generate adversarial examples. However, pre-trained models are hard to obtain in real-world tasks. In this paper, we propose a data-free substitute training method (DaST) to obtain substitute models for adversarial black-box attacks without the requirement of any real data. To achieve this, DaST utilizes specially designed generative adversarial networks (GANs) to train the substitute models. In particular, we design a multi-branch architecture and label-control loss for the generative model to deal with the uneven distribution of synthetic samples. The substitute model is then trained by the synthetic samples generated by the generative model, which are labeled by the attacked model subsequently. The experiments demonstrate the substitute models produced by DaST can achieve competitive performance compared with the baseline models which are trained by the same train set with attacked models. Additionally, to evaluate the practicability of the proposed method on the real-world task, we attack an online machine learning model on the Microsoft Azure platform. The remote model misclassifies 98.35% of the adversarial examples crafted by our method. To the best of our knowledge, we are the first to train a substitute model for adversarial attacks without any real data.
Mingyi Zhou, Jing Wu 0021, Yipeng Liu 0001, Shuaicheng Liu, Ce Zhu
CVPR3
2020 Smooth robust tensor principal component analysis for compressed sensing of dynamic MRI
Yipeng Liu 0001, Tengteng Liu, Jiani Liu 0002, Ce Zhu
Pattern Recognit.1
2020 Robust block tensor principal component analysis
Lanlan Feng, Yipeng Liu 0001, Longxi Chen, Xiang Zhang 0006, Ce Zhu
Signal Process.2
2020 Provable tensor ring completion
Huyan Huang, Yipeng Liu 0001, Jiani Liu 0002, Ce Zhu
Signal Process.2
2020 Low CP Rank and Tucker Rank Tensor Completion for Estimating Missing Components in Image Data
abstract
Tensor completion recovers missing components of multi-way data. The existing methods use either the Tucker rank or the CANDECOMP/PARAFAC (CP) rank in low-rank tensor optimization for data completion. In fact, these two kinds of tensor ranks represent different high-dimensional data structures. In this paper, we propose to exploit the two kinds of data structures simultaneously for image recovery through jointly minimizing the CP rank and Tucker rank in the low-rank tensor approximation. We use the alternating direction method of multipliers (ADMM) to reformulate the optimization model with two tensor ranks into its two sub-problems, and each has only one tensor rank optimization. For the two main sub-problems in the ADMM, we apply rank-one tensor updating and weighted sum of matrix nuclear norms minimization methods to solve them, respectively. The numerical experiments on some image and video completion applications demonstrate that the proposed method is superior to the state-of-the-art methods.
Yipeng Liu 0001, Zhen Long, Huyan Huang, Ce Zhu
IEEE Trans. Circuits Syst. Video Technol.1
2020 Low-Rank Tensor Train Coefficient Array Estimation for Tensor-on-Tensor Regression
abstract
The tensor-on-tensor regression can predict a tensor from a tensor, which generalizes most previous multilinear regression approaches, including methods to predict a scalar from a tensor, and a tensor from a scalar. However, the coefficient array could be much higher dimensional due to both high-order predictors and responses in this generalized way. Compared with the current low CANDECOMP/PARAFAC (CP) rank approximation-based method, the low tensor train (TT) approximation can further improve the stability and efficiency of the high or even ultrahigh-dimensional coefficient array estimation. In the proposed low TT rank coefficient array estimation for tensor-on-tensor regression, we adopt a TT rounding procedure to obtain adaptive ranks, instead of selecting ranks by experience. Besides, an l2constraint is imposed to avoid overfitting. The hierarchical alternating least square is used to solve the optimization problem. Numerical experiments on a synthetic data set and two real-life data sets demonstrate that the proposed method outperforms the state-of-the-art methods in terms of prediction accuracy with comparable computational complexity, and the proposed method is more computationally efficient when the data are high dimensional with small size in each mode.
Yipeng Liu 0001, Jiani Liu 0002, Ce Zhu
IEEE Trans. Neural Networks Learn. Syst.1
2019 Early diagnosis of Parkinson's disease from multiple voice recordings by simultaneous sample and feature selection
Ce Zhu, Mingyi Zhou, Yipeng Liu 0001
Expert Syst. Appl.4
2019 Robust corner detection using altitude to chord ratio accumulation
Ce Zhu, Yipeng Liu 0001, Qian Zhang 0047
Multim. Tools Appl.3
2019 Low rank tensor completion for multiway visual data
Zhen Long, Yipeng Liu 0001, Longxi Chen, Ce Zhu
Signal Process.2
2019 Editorial to The Special Issue on Tensor Image Processing
Yipeng Liu 0001, Qibin Zhao, Shuchin Aeron
Signal Process. Image Commun.1
2019 Tensor rank learning in CP decomposition via convolutional neural network
Mingyi Zhou, Yipeng Liu 0001, Zhen Long, Longxi Chen, Ce Zhu
Signal Process. Image Commun.2
2019 Image Completion Using Low Tensor Tree Rank and Total Variation Minimization
abstract
Tensor completion recovers missing entries of multiway data. Most of the current methods exploit the low-rank tensor structure for image completion applications. In this paper, we simultaneously exploit the globally multidimensional structure and locally piecewise smoothness to further enhance the performance. In the proposed optimization model, the low tensor tree rank minimization is used for the global data structure, and the total variation minimization is used for the local structure. Two kinds of total variation functions are discussed. The optimization problem is transformed into several subproblems by alternating direction method of multipliers. The subproblem on low tensor tree rank minimization is solved by singular value thresholding, and the subproblem on total variation minimization can be solved by soft thresholding. Numerical experiments on color images and light field images demonstrate that the proposed method outperforms most of the state-of-the-art methods in terms of recovery accuracy and computational complexity.
Yipeng Liu 0001, Zhen Long, Ce Zhu
IEEE Trans. Multim.1
2018 Robust Tensor Principal Component Analysis in All Modes
abstract
Robust tensor principal component analysis extracts the low rank and sparse component of multi-dimensional data by tensor singular value decomposition (t-SVD), which can be used for many data analysis problems. However, the current t-SVD based methods cannot fully extract the low rank component in tensor data, and low rank structure still exists in the core tensor, because t-SVD does not decompose data in the third mode. To fully exploit the low rank structure, we further extract the low rank component using low rank plus sparsity for the core matrix whose entries are from the diagonal elements of the frontal slices in the core tensor. The proposed method is applied to three groups of numerical experiments on image denoising, illumination normalization for face images and motion separation for surveillance videos, respectively, and the results show that the proposed method outperforms state-of-the-art methods in terms of both accuracy and computational complexity.
Longxi Chen, Yipeng Liu 0001, Ce Zhu
ICME2
2018 Image Ordinal Classification and Understanding: Grid Dropout with Masking Label
abstract
Image ordinal classification refers to predicting a discrete target value which carries ordering correlation among image categories. The limited size of labeled ordinal data renders modern deep learning approaches easy to overfit. To tackle this issue, neuron dropout and data augmentation were proposed which, however, still suffer from over-parameterization and breaking spatial structure, respectively. To address the issues, we first propose a grid dropout method that randomly dropout/blackout some areas of the training image. Then we combine the objective of predicting the blackout patches with classification to take advantage of the spatial information. Finally we demonstrate the effectiveness of both approaches by visualizing the Class Activation Map (CAM) and discover that grid dropout is more aware of the whole facial areas and more robust than neuron dropout for small training dataset. Experiments are conducted on a challenging age estimation dataset-Adience dataset with very competitive results compared with state-of-the-art methods.
Chao Zhang 0072, Ce Zhu, Jimin Xiao, Xun Xu 0002, Yipeng Liu 0001
ICME5
2018 Extended smoothlets: An efficient multi-resolution adaptive transform
Qian Zhang 0047, Yipeng Liu 0001, Ce Zhu, Chang Duan
J. Vis. Commun. Image Represent.4
2018 Visual aesthetic understanding: Sample-specific aesthetic classification and deep activation map visualization
Chao Zhang 0072, Ce Zhu, Xun Xu 0002, Yipeng Liu 0001, Jimin Xiao, Tammam Tillo
Signal Process. Image Commun.4
2018 Fast Signal Recovery From Saturated Measurements by Linear Loss and Nonconvex Penalties
abstract
Sign information is the key for overcoming the inevitable saturation error in compressive sensing system, which causes loss of information and may result in great bias. For sparse signal recovery from saturation, we propose to use linear loss to improve the effectiveness from the existing methods that utilize hard constraints/hinge loss for sign consistency. Due to the use of linear loss, analytical solution in the update progress is obtained and some nonconvex penalties are applicable, e.g., minimax concave penalty, ℓ0norm, and sorted ℓ1norm. Theoretical analysis reveals that the estimation error can still be bounded. Generally, with linear loss and nonconvex penalties, the recovery performance can be significantly improved and the computational time is also largely saved, which is verified by the numerical experiments.
Xiaolin Huang, Yipeng Liu 0001, Ming Yan 0006
IEEE Signal Process. Lett.3
2017 Iterative block tensor singular value thresholding for extraction of lowrank component of image data
abstract
Tensor principal component analysis (TPCA) is a multi-linear extension of principal component analysis which converts a set of correlated measurements into several principal components. In this paper, we propose a new robust TPCA method to extract the principal components of the multi-way data based on tensor singular value decomposition. The tensor is split into a number of blocks of the same size. The low rank component of each block tensor is extracted using iterative tensor singular value thresholding method. The principal components of the multi-way data are the concatenation of all the low rank components of all the block tensors. We give the block tensor incoherence conditions to guarantee the successful decomposition. This factorization has similar optimality properties to that of low rank matrix derived from singular value decomposition. Experimentally, we demonstrate its effectiveness in two applications, including motion separation for surveillance videos and illumination normalization for face images.
Longxi Chen, Yipeng Liu 0001, Ce Zhu
ICASSP2
2017 Attribute-controlled face photo synthesis from simple line drawing
abstract
Face photo synthesis from simple line drawing is a one-to-many task as simple line drawing merely contains the contour of human face. Previous exemplar-based methods are over-dependent on the datasets and are hard to generalize to complicated natural scenes. Recently, several works utilize deep neural networks to increase the generalization, but they are still limited in the controllability of the users. In this paper, we propose a deep generative model to synthesize face photo from simple line drawing controlled by face attributes such as hair color and complexion. In order to maximize the controllability of face attributes, an attribute-disentangled variational auto-encoder (AD-VAE) is firstly introduced to learn latent representations disentangled with respect to specified attributes. Then we conduct photo synthesis from simple line drawing based on AD-VAE. Experiments show that our model can well disentangle the variations of attributes from other variations of face photos and synthesize detailed photorealistic face images with desired attributes. Regarding background and illumination as the style and human face as the content, we can also synthesize face photos with the target style of a style photo.
Ce Zhu, Zhiqiang Xia, Yipeng Liu 0001
ICIP5
2017 Towards thinner convolutional neural networks through gradually global pruning
abstract
Deep network pruning is an effective method to reduce the storage and computation cost of deep neural networks when applying them to resource-limited devices. Among many pruning granularities, neuron level pruning will remove redundant neurons and filters in the model and result in thinner networks. In this paper, we propose a gradually global pruning scheme for neuron level pruning. In each pruning step, a small percent of neurons were selected and dropped across all layers in the model. We also propose a simple method to eliminate the biases in evaluating the importance of neurons to make the scheme feasible. Compared with layer-wise pruning scheme, our scheme avoid the difficulty in determining the redundancy in each layer and is more effective for deep networks. Our scheme would automatically find a thinner sub-network in original network under a given performance.
Ce Zhu, Zhiqiang Xia, Yipeng Liu 0001
ICIP5
2017 Learning based 3D keypoint detection with local and global attributes in multi-scale space
abstract
Over the last few decades various methods have been proposed by researchers to extract 3D keypoints from the surface of 3D mesh models, but most of them are geometric ones, which are not flexible enough for various applications. In this paper, we propose a new 3D keypoint detection method based on multi-scale neural network (MSNN), which is a tiny neural network and can effectively merge multi-scale information to detect 3D keypoints. Traditional end-to-end learning systems usually require large-scale dataset to do training. However, there are not enough 3D data with ground truth of 3D keypoints. To solve this problem, we perform delicate preprocessing, which effectively enhance the performance of the MSNN based approach. Numerical experiments show that the proposed MSNN 3D keypoint detector not only outperforms other six state-of-the-art geometric based methods, but also achieves better performance than a learning-based method using random forest.
Ce Zhu, Qian Zhang 0047, Mengxue Wang, Yipeng Liu 0001
MMSP5
2017 Efficient and Robust Corner Detectors Based on Second-Order Difference of Contour
abstract
As one of the most significant local features of image, corner is widely used in many computer vision tasks. Corner detection aims to achieve the highest possible detection accuracy while minimizing the computational complexity. In this letter, we first introduce a new measurement termed as second-order difference of contour (SODC), and then examine its regular distribution, which is found to provide useful information to distinguish corners from noncorners. Based on the SODC distribution characteristics, we propose two novel corner detectors to measure the response of contour points using Manhattan distance and Euclidean distance, respectively. Numerical experiments demonstrate that the Manhattan detector greatly decreases the computational complexity, while the Euclidean detector outperforms the state-of-the-art corner detectors in terms of repeatability and localization error.
Ce Zhu, Qian Zhang 0047, Xiaolin Huang, Yipeng Liu 0001
IEEE Signal Process. Lett.5
2017 A Bayesian Approach to Camouflaged Moving Object Detection
abstract
Moving object detection is about foreground and background separation based on motion detection. Detecting moving objects from similarly colored background (known as camouflage problem) has been a long-standing open question in this field. Discriminative modeling (DM), which focuses on enhancing the performance to distinguish foreground from background with discriminative features and well-designed classifiers, has been widely used for moving object detection. However, DM may tend to fail when encountering the camouflage problem, as the class separability in camouflaged areas is generally poor. In this paper, we propose a new strategy, camouflage modeling (CM), to identify camouflaged foreground pixels. In view of the fact that camouflage involves both foreground and background, we need to model both the background and the foreground, and compare them in a well-designed way in camouflage detection. Specifically, we develop a global model for the background, and an integration of global and local models for the foreground, respectively. Based on both background and foreground models, we introduce a factor to measure the degree of camouflage, and further identify truly camouflaged areas. In view of the fact that a moving object is usually composed of both camouflaged and noncamouflaged areas, CM and DM are fused in a Bayesian framework to perform complete object detection. Experiments are conducted on testing sequences to demonstrate the effectiveness of the proposed algorithm.
Xiang Zhang 0006, Ce Zhu, Yipeng Liu 0001, Mao Ye 0001
IEEE Trans. Circuits Syst. Video Technol.4
2017 Hybrid CS-DMRI: Periodic Time-Variant Subsampling and Omnidirectional Total Variation Based Reconstruction
abstract
Compressive sensing (CS) has been used to accelerate dynamic magnetic resonance imaging (DMRI). Currently, the online CS-DMRI is faster, whereas the offline CS-DMRI provides higher accuracy for image reconstruction. To achieve good image reconstruction performance in terms of both speed and accuracy, we propose a hybrid CS-DMRI method using periodic time-variant subsampling for different frames. In each period, there is one reference frame that is sampled at a higher subsampling ratio. The two nearby reference frames with good reconstruction quality can be used to provide rough predictions of the other frames between them. To finely recover the current frame, one structural regularization in the optimization model for reconstruction is a 2-D omnidirectional total variation (OTV) for exploiting the sparsity of the difference between the predicted and estimated frames, and the other is a 3-D OTV as a regularization term for exploiting the bilateral spatio-temporal coherence between the forward reference frame, current frame, and backward reference frame. Compared with classical total variation, the proposed OTV fully utilizes the correlations of all the possible directions of the data. The formulated optimization model can be solved using iterative reweighted least squares with the pre-conditioned conjugate gradient method. Numerical experiments demonstrate that the proposed method has better reconstruction accuracy than all the existing methods and low computational complexity that is comparable to the existing online methods.
Yipeng Liu 0001, Xiaolin Huang, Ce Zhu
IEEE Trans. Medical Imaging1
2016 Robust sparse recovery for compressive sensing in impulsive noise using ℓp-norm model fitting
abstract
This work considers the robust sparse recovery problem in compressive sensing (CS) in the presence of impulsive measurement noise. We propose a robust formulation for sparse recovery using the generalized lp-norm with 0 < p < 2 as the metric for the residual error under l1-norm regularization. An alternative direction method (ADM) has been proposed to solve this formulation efficiently. Moreover, a smoothing strategy has been used to derive a convergent method for the nonconvex case of p < 1. The convergence conditions of the proposed algorithm for both the convex and nonconvex cases have been provided. Numerical simulations demonstrated that the new algorithm can achieve state-of-the-art robust performance in highly impulsive noise.
Fei Wen 0005, Yipeng Liu 0001, Robert C. Qiu, Wenxian Yu
ICASSP3
2016 3D interest point detection based on geometric measures and sparse refinement
abstract
Three dimensional (3D) interest point detection plays a fundamental role in computer vision. In this paper, we introduce a new method for detecting 3D interest points of 3D mesh models based on geometric measures and sparse refinement (GMSR). The key point of our approach is to calculate the 3D saliency measure using two novel geometric measures, which are defined in multi-scale space to effectively distinguish 3D interest points from edges and flat areas. Those points with local maxima of 3D saliency measure are selected as the candidates of 3D interest points. Finally, we utilize an l0norm based optimization method to refine the candidates of 3D interest points by constraining the number of 3D interest points. Numerical experiments show that the proposed GMSR based 3D interest point detector outperforms current six state-of-the-art methods for different kinds of 3D mesh models.
Ce Zhu, Qian Zhang 0047, Yipeng Liu 0001
MMSP4
2015 Two-level ℓ1 minimization for compressed sensing
Xiaolin Huang, Yipeng Liu 0001, Lei Shi 0010, Sabine Van Huffel, Johan A. K. Suykens
Signal Process.2
2010 Adaptive Inter-Atom Interference Mitigation Approach to Sparse Multi-Path Channel Estimation
abstract
Impulse response of multi-path channel can be estimated by using a short training sequence when the channel is sparse. Though the ordinary orthogonal matching pursuit (OMP) provides fast sparse multi-path channel (SMPC) estimation, it suffers from inter-atom interference (IAI), especially in the case of SMPC with a large delay spread and short training sequence. Herein, an adaptive IAI mitigation method is proposed to improve OMP algorithm based on a sensing dictionary, which is utilized to prevent false atoms from being selected due to serious IAI. Numeral experiments illustrate that the improved OMP algorithm based on adaptive IAI mitigation outperforms both the ordinary OMP algorithm.
Ruiming Yang, Qun Wan, Yipeng Liu 0001, Wan-Lin Yang
VTC Spring3