VLDB 2026 Research / reviewers in the wild / expert
Ce Zhu
dblp:69/2369
· DBLP profile ↗
273ranked-venue papers
14as first author
133since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 204 · 14 first-author · 80 since 2021Artificial intelligence and machine learning · 58 · 48 since 2021Systems, architecture and hardware · 11 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 6 since 2021Databases, data management, data science and information retrieval · 6 · 5 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | ViS2T: A vision and scene graph to text model for captive giant panda video captioning
Chenyu Ma, Chang Duan, Ke Zhang 0022, Mengnan He, Zhongrong Wang, Ping Zhang 0023, Ce Zhu |
Expert Syst. Appl. | 8 |
| 2026 | DMGNet: Discriminative multi-view geometry learning with hybrid-domain enhancement for light field occlusion removal
Jieyu Chen, Ping An 0001, Xinpeng Huang, Chao Yang 0021, Ce Zhu, Sanghoon Lee 0001 |
Expert Syst. Appl. | 5 |
| 2026 | Relative depth knowledge distillation for generalizable monocular depth estimation
Mankun Li, Meng Yang 0002, Xuguang Lan, Ce Zhu |
Neurocomputing | 5 |
| 2026 | PG2CTDL: Pixel and gradient decorrelation guided coupled tensor dictionary learning for multimodal image fusion
Zhen Long, Lanlan Feng, Ce Zhu |
Knowl. Based Syst. | 4 |
| 2026 | GeoRL: Adaptive tokenization via reinforcement learning for remote sensing foundation modelsabstractVision Foundation Models (VFMs) have had success in transferring learned visual representations from discrete tasks to other tasks. However, VFMs for interpreting remote sensing images are limited due to heterogeneous properties found in Earth observation data: spatial scale, semantic complexity, and geographical context. To address this issue, we introduce GeoRL, an Adaptive Visual Tokenization Framework that treats token allocation as an MDP and allows VFMs dynamic token allocation decisions based on information density of a region. We developed a lightweight Policy Network trained with Proximal Policy Optimization (PPO) that learns to maximize the reward function composition of tokenization performance and computational efficiency. We also propose Hierarchical Semantic Anchoring (HSA) for interpreting tokenization policies learned. By performing extensive experiments on four established benchmark datasets (DOTA, iSAID, LoveDA, and xView), we have established that GeoRL outperforms all other competitors across three tasks of scene classification, semantic segmentation, and object detection, while also achieving a reduction in computational cost of 47%–63% when compared with uniform tokenization techniques. The policy learned by GeoRL can be used directly with other sources of imagery (e.g., SAR, multispectral) without requiring retraining, indicating that GeoRL leverages the spatial reasoning patterns inherent to all remote sensing imagery regardless of the modality. Theoretically, GeoRL has proof of convergence guarantees with mild assumptions and provides insight into the sample complexity associated with learning a policy in the remote sensing domain. Hengyang He, Ce Zhu |
Pattern Recognit. | 3 |
| 2026 | Coupled tensor train decomposition in federated learning
Xiangtao Zhang, Eleftherios Kofidis, Ruituo Wu, Ce Zhu, Le Zhang 0001, Yipeng Liu 0001 |
Pattern Recognit. | 4 |
| 2026 | Implicit neural functional tensor train for multivariate function approximationabstractThe accurate approximation of multivariate functions is fundamental to high-dimensional data analysis. Among various strategies, tensor-based methods have been widely explored to address the curse of dimensionality for scalable and efficient approximation. While traditional techniques employ either structured tensor-product grid discretization or coefficient tensor factorization, they face two critical limitations: (1) limited adaptability to irregular data structures common in real-world applications and (2) dependence on manually designed basis functions that restrict flexibility and require substantial domain expertise. This paper introduces the Implicit Neural Functional Tensor Train (INFTT), which couples functional tensor–train decomposition with implicit neural representations to obtain continuous, low-rank approximations of multivariate functions. Specifically, instead of manual basis design, we parameterize each univariate core with a SIREN-based implicit network, preserving the chain structure while enabling adaptive, high-fidelity modeling. Our theoretical analysis demonstrates that INFTT inherently embeds low-rank structural priors and Lipschitz continuity, ensuring compact and smooth approximations. Extensive experiments on synthetic and real-world scientific datasets validate the superior accuracy and robustness of the proposed method compared to state-of-the-art approaches for both discretized and continuous representations. Jiani Liu, André Lima Férrer de Almeida, Ce Zhu |
Signal Process. | 4 |
| 2026 | TERM Model: Tensor Ring Mixture Model for Density EstimationabstractProbabilistic modeling is a core challenge in statistical machine learning. Tensor-based probabilistic graph methods address interpretability and stability concerns encountered in neural network approaches and allow tractable inference (e.g., marginal inference and conditional inference). In this paper, we introduce tensor ring decomposition for density estimation, which reduces the number of permutation candidates compared to existing methods, while simultaneously enhancing expressive power and maintaining tractable inference. Different non-negative strategies for density function results in two variants: Born TRDE offers simpler inference and sampling but with slightly lower accuracy, while Energy TRDE, though more complex, achieves superior performance. Furthermore, a mixture model that incorporates multiple permutation candidates with adaptive weights is designed, resulting in increased expressive flexibility and comprehensiveness. Unlike existing methods that focus on finding a single optimal permutation, our approach, inspired by ensemble learning, demonstrates that combining multiple suboptimal permutations can yield superior results. Experiments demonstrate that the proposed approach excels in estimating probability density functions and sampling, capturing intricate details with competitive or superior performance compared to existing state-of-the-art (SOTA) tractable density methods. Ruituo Wu, Jiani Liu 0002, Bing Li 0002, Anh Huy Phan 0001, Ivan V. Oseledets, Ce Zhu, Yipeng Liu 0001 |
IEEE Trans. Big Data | 6 |
| 2026 | Progressive Reasoning-Based Group Activity RecognitionabstractGroup activity recognition (GAR) plays a crucial role in computer vision, enabling the exploration and comprehension of human behavior patterns. Existing methods mainly focus on dyad-level interactions within a group, but sociological studies have highlighted the importance of individual features, subgroup-level interactions, and overall group structure for understanding group activities. Therefore, we propose a new framework, the progressive group activity reasoning model (PGAR), which models these four aspects for GAR. Initially, we construct a person-person graph (PPG) using individual features to capture dyadic interactions. Subsequently, the PPG is fed into a novel ingredient graph model (Ingredient-GNN) for capturing subgroup-level interactions. Finally, we fuse the dyad-level and subgroup-level interactions with global information of group structure, obtained through an F-Formation modeling module, to form comprehensive representations for GAR. The F-Formation modeling module decouples the group structure into position, orientation, and skeleton graphs, and subsequently performs attribute recoupling at the individual level using the designed Tri-Coupling Transformer to form a global representation of the group structure. Extensive experiments on four public datasets demonstrate that our final model effectively integrates multi-level representations for group activity understanding, with our F-Formation modeling module outperforming comparable methods that rely solely on non-visual data. Lindong Li, Linbo Qing, Wang Tang, Pingyu Wang, Haosong Gou, Ce Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Content-Adaptive Unfolding Wavelet Transformer for Hyperspectral Image Super-ResolutionabstractIn recent years, fusing high-resolution multispectral images (HR-MSIs) and low-resolution hyperspectral images (LR-HSIs) has become a widely used approach for hyperspectral image super-resolution (HSI-SR). The deep unfolding framework has attracted significant attention thanks to its ability to formulate the problem into a data module and a prior module. However, there are still two critical issues that hinder the performance enhancement of the existing methods: 1) Parameters in the data module are fixed (though learnable) at each iteration, i.e., lacking the adaptivity to comprehensive data; 2) The Transformer in the prior module cannot effectively capture high-frequency information. To resolve these issues, we propose a Content-Adaptive Unfolding Wavelet Transformer (CAUWT) for HSI-SR, where the parameters are adaptively learned based on the reconstructed HSI at each iteration. Moreover, we propose a novel Wavelet-Assisted Transformer (WAT), by integrating the Discrete Wavelet Transform (DWT) and the Hybrid Spectral-Spatial Attention Block (HSSAB) to further upgrade the high-frequency information quality of HSI at no cost of extra branch structures, where the former is for multi-scale and multi-frequency details and the latter is for correlations between and within sub-band components. Extensive experiments performed on both simulated and real datasets well demonstrate the effectiveness of the proposed method. In comparison with mainstream HSI-SR methods, our method exhibits superior performance and lower computational overhead. Yipeng Liu 0001, Zhen Long, Chong-Yung Chi, Ce Zhu |
IEEE Trans. Image Process. | 5 |
| 2026 | Learning Feature Encoder With Synthetic Anomalies for Weakly Supervised Graph Anomaly DetectionabstractWeakly supervised graph anomaly detection aims to unveil unusual graph instances, e.g., nodes, whose behaviors significantly differ from normal ones, given only a limited number of annotated anomalies and abundant unlabeled samples. A major challenge is to learn a meaningful latent feature representation that reduces intra-class variance among normal data while remaining highly sensitive to anomalies. Although recent works have applied self-supervised feature learning for graph anomaly detection, their strategies are not specifically tailored to its unique requirements, motivating our exploration of a more domain-specific approach. In this paper, we introduce a weakly supervised graph anomaly detection method that leverages a feature learning strategy tailored for graph anomalies. Our approach is built upon a multi-task learning scheme that extracts robust feature representations through synthesized anomalies. We generate synthetic anomalies by perturbing the normal graph in various ways and assign a dedicated detection head to each anomaly type, ensuring that learned features are sensitive to potential deviations from normal patterns. Although synthetic anomalies may not perfectly replicate real-world patterns, they provide valuable auxiliary data for effective feature learnin, much like features learned from ImageNet classification transfer to downstream vision tasks. Additionally, we adopt a two-phase learning strategy: an initial warm-up phase using only synthetic samples, followed by a full-training phase integrating both tasks, to balance the influence of synthetic and real data. Extensive experiments on public datasets demonstrate the superior performance of our method over its competitors. Code is available at https://github.com/yj-zhou/SAWGAD. Yingjie Zhou 0001, Yuqin Xie, Fanxing Liu, Dongjin Song, Ce Zhu, Lingqiao Liu |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2026 | FunOTTA: On-the-Fly Adaptation on Cross-Domain Fundus Image via Stable Test-Time TrainingabstractFundus images are essential for the early screening and detection of eye diseases. While deep learning models using fundus images have significantly advanced the diagnosis of multiple eye diseases, variations in images from different imaging devices and locations (known as domain shifts) pose challenges for deploying pre-trained models in real-world applications. To address this, we propose a novel Fundus On-the-fly Test-Time Adaptation (FunOTTA) framework that effectively generalizes a fundus image diagnosis model to unseen environments, even under strong domain shifts. FunOTTA stands out for its stable adaptation process by performing dynamic disambiguation in the memory bank while minimizing harmful prior knowledge bias. We also introduce a new training objective during adaptation that enables the classifier to incrementally adapt to target patterns with reliable class conditional estimation and consistency regularization. We compare our method with several state-of-the-art test-time adaptation (TTA) pipelines. Experiments on cross-domain fundus image benchmarks across two diseases demonstrate the superiority of the overall framework and individual components under different backbone networks. Code is available at https://github.com/Casperqian/FunOTTA. Le Zhang 0001, Yipeng Liu 0001, Ce Zhu, Fan Zhang 0013 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | WiFi CSI Based Temporal Activity Detection via Dual Pyramid NetworkabstractWe address the challenge of WiFi-based temporal activity detection and propose an efficient Dual Pyramid Network that integrates Temporal Signal Semantic Encoders and Local Sensitive Response Encoders. The Temporal Signal Semantic Encoder splits feature learning into high and low-frequency components, using a novel Signed Mask-Attention mechanism to emphasize important areas and downplay unimportant ones, with the features fused using ContraNorm. The Local Sensitive Response Encoder captures fluctuations without learning. These feature pyramids are then combined using a new cross-attention fusion mechanism. We also introduce a dataset with over 2,114 activity segments across 553 WiFi CSI samples, each lasting around 85 seconds. Extensive experiments show our method outperforms challenging baselines. Le Zhang 0001, Bing Li 0002, Yingjie Zhou 0001, Zhenghua Chen, Ce Zhu |
AAAI | 6 |
| 2025 | SLR-MVTC: Smooth Low-Rank Multi-View Tensor ClusteringabstractMulti-view tensor clustering (MVTC) has gained much attention for its effectiveness in capturing global high-order correlations across views. However, current MVTC methods suffer from two limitations: 1) adopting a two-stage process to learn the latent features for clustering, and 2) either ignoring local similarities within views or treating local similarities and global high-order correlations equally. In this paper, we propose a smooth low-rank MVTC (SLR-MVTC) method, which aims to extract latent features that are smooth within each view and low-rank across views, enhancing clustering performance. Specifically, we first learn latent features from each view using orthogonal projection and then construct the latent feature tensor by concatenation and rotation. Then, we introduce a new smooth tensor nuclear norm to depict the low-rank components of the low-frequency parts in the feature tensor. Benefiting from the fast Fourier transform along the sample dimension, the obtained low-frequency components effectively capture local smoothness within views, while their low-rank parts further explore global correlations across views. Experimental results on six multi-view datasets demonstrate that SLR-MVTC outperforms state-of-the-art algorithms in terms of clustering performance and CPU time. Zhen Long, Yipeng Liu 0001, Yazhou Ren 0001, Ce Zhu |
AAAI | 4 |
| 2025 | Rethinking Token Reduction with Parameter-Efficient Fine-Tuning in ViT for Pixel-Level TasksabstractParameter-Efficient fine-tuning (PEFT) adapts pre-trained models to new tasks by updating only a small subset of parameters, achieving efficiency but still facing significant inference costs driven by input token length. This challenge is even more pronounced in pixel-level tasks, which require longer input sequences compared to image-level tasks. Although token reduction (TR) techniques can help reduce computational demands, they often lead to homogeneous attention patterns that compromise performance in pixel-level scenarios. This study underscores the importance of maintaining attention diversity for these tasks and proposes to enhance attention diversity while ensuring the completeness of token sequences. Our approach effectively reduces the number of tokens processed within transformer blocks, improving computational efficiency without sacrificing performance on several pixel-level tasks. We also demonstrate the superior generalization capability of our proposed method compared to challenging baseline models. The source code will be made available at https://github.com/AVC2-UESTC/DAR-TR-PEFT. Ao Li 0007, Hu Yao, Ce Zhu, Le Zhang 0001 |
CVPR | 4 |
| 2025 | Subspace Constraint and Contribution Estimation for Heterogeneous Federated LearningabstractHeterogeneous Federated Learning (HFL) has received widespread attention due to its adaptability to different models and data. The HFL approach utilizing auxiliary models for knowledge transfer can further enhance flexibility. However, existing frameworks face the challenges of local overfitting and aggregation bias. To address these issues, we propose FedSCE. By restricting specific layers of the local model updates to a subspace, FedSCE reduces the degrees of freedom of the update, enhances generalization, and mitigates the risk of overfitting. The subspace is dynamically updated to ensure coverage of the latest model update trajectory. Additionally, FedSCE evaluates client contributions based on the update distance of the auxiliary model in feature space and parameter space, achieving adaptive weighted aggregation. We validate our approach in both feature-skewed and label-skewed scenarios, demonstrating that on Office10, our method exceeds the best baseline by 3.87%. The code will be available at https://github.com/AVC2-UESTC/FedSCE.git. Xiangtao Zhang, Ao Li 0007, Yipeng Liu 0001, Fan Zhang 0013, Ce Zhu, Le Zhang 0001 |
CVPR | 6 |
| 2025 | Fully Connected Tensor Network based Brain Structural Feature Extraction for Early Alzheimer's Disease DetectionabstractAlzheimer’s disease (AD) is an incurable neurodegenerative disease that involves structural changes in the brain. Early diagnosis of AD helps provide timely treatment and delay its progressive process. Many studies have been conducted based on brain images to detect AD. However, these works are mostly designed based on multi-modal brain data. The existing methods on uni-modal data, e.g., diffusion magnetic resonance imaging (dMRI) data, are not very effective or inexplainable. In this work, we propose the first tensor-based method for detecting early-stage AD. First, the brain data are segmented into blocks and concatenated as high-order tensors. The structural information in each block is preserved. Second, the fully connected tensor network (FCTN) is adopted. FCTN has strong expression capability for high-order data. This is the first work that exploits the advantage of FCTN for structural feature extraction. In the classification experiments, our method surpasses the existing methods 5% to 15% in accuracy. Ce Zhu |
ICASSP | 3 |
| 2025 | Distribution Alignment Informed Thresholding for Semi-Supervised Curvilinear Structure SegmentationabstractCurvilinear structure segmentation using deep neural networks is often limited by the high cost of annotation. Semi-supervised learning (SSL) helps mitigate this dependency on extensive annotated data. State-of-the-art SSL approaches generate pseudo-labels for unlabeled data, which are then used for further model training. These methods primarily focus on calibrating thresholds to binarize the predictions. In this work, we assume that when labeled and unlabeled data are similar, the foreground-to-background ratio should be consistent between them. To leverage this assumption, we calibrate the threshold by minimizing the distribution gap between labeled ground truth and pseudo-labels on unlabeled data. Our proposed threshold calibration can be integrated with existing SSL methods. We evaluate its effectiveness on four datasets, demonstrating that our method outperforms current state-of-the-art SSL techniques, especially in scenarios with very low labeled data. Yuhao Mo, Bihan Wen, Xulei Yang, Ce Zhu, Xun Xu 0002 |
ICASSP | 5 |
| 2025 | Codar: Complex-valued Neural Network for Crossing-Floor Intrusion Detection via WiFiabstractWiFi systems offer enormous potential for device-free human intrusion detection. Current methods often require routers to be deployed in multiple adjacent rooms on the same floor, which is redundant and costly. To solve this, we introduce the first work on intrusion detection in the crossing-floor scenario via WiFi. Routers on different floors are utilized without major modifications to the existing router layout. Many previous works require a high sample rate and ignore the phase information. In this paper, we propose Codar, a complex-valued LSTM-CNN neural network. The LSTM effectively captures temporal dependencies at a low sample rate in harsh propagation environments. Moreover, amplitude and phase features are explored jointly by complex-valued operations. Experimental results demonstrate Codar achieves 95%, 94.5%, and 99% accuracy for intrusion detection, user identification, and intruded floor identification, surpassing competitive methods. The code and dataset are available at https://github.com/ouweiting/Codar. Weiting Ou, Yipeng Liu 0001, Bing Li 0002, Le Zhang 0001, Ce Zhu |
ICASSP | 6 |
| 2025 | Channel and space-based joint rate allocation algorithmabstractRate control is a critical component for image and video compression Particularly under limited network bandwidth conditions, bitrate control is essential to ensure efficient image transmission by effectively allocation channel resources. In this research, since both Channel and Spatial have relationship with rate allocation, we first propose a joint Channel-wise and Spatial-wise Quantization scheme to determine optimal quantization parameters. Subsequently, we develop a quantization step estimation network to obtain parameters to efficiently allocate rate according to target rate. Experiments demonstrate that our algorithm significantly improve compressed image quality with minimal bitrate distortion and achieve accurate rate control with nearly 3% average bitrate error. Yu Sun 0003, Xin Lu 0001, Frédéric Dufaux, Ce Zhu |
ICASSP | 7 |
| 2025 | MobileIE: An Extremely Lightweight and Effective ConvNet for Real-Time Image Enhancement on Mobile DevicesabstractRecent advancements in deep neural networks have driven significant progress in image enhancement (IE). However, deploying deep learning models on resource-constrained platforms, such as mobile devices, remains challenging due to high computation and memory demands. To address these challenges and facilitate real-time IE on mobile, we introduce an extremely lightweight Convolutional Neural Network (CNN) framework with around 4K parameters. Our approach integrates reparameterization with an Incremental Weight Optimization strategy to ensure efficiency. Additionally, we enhance performance with a Feature Self-Transform module and a Hierarchical Dual-Path Attention mechanism, optimized with a Local Variance-Weighted loss. With this efficient framework, we are the first to achieve real-time IE inference at up to 1,100 frames per second (FPS) while delivering competitive image quality, achieving the best trade-off between speed and performance across multiple IE tasks. The code will be available at https://github.com/AVC2-UESTC/MobileIE.git. Hailong Yan, Ao Li 0007, Xiangtao Zhang, Zhe Liu 0019, Zenglin Shi, Ce Zhu, Le Zhang 0001 |
ICCV | 6 |
| 2025 | A Multi-Layer End-to-End 360 Image Compressionabstract360° images have attracted increasing attention due to their wide field of view. However, spherical 360° images need to be converted into 2D equi-rectangular projection (ERP) images for compression. This conversion often leads to pixel overstretching in the ERP image, which results in a lot of redundancy in the texture. Performing a direct prediction without considering the stretching will inevitably make it difficult to achieve optimal results. To tackle this problem, we propose a multi-layer adaptive scale-block scheme for ERP image compression. In particular, we introduce a multi-layer structure based on the overstretching rate and use multi-scale convolution kernels to better match each layer and extract features more effectively. Subsequently, we employ an adaptive scale-block method to effectively reduce bitrate redundancy in overstretched and less important regions. Finally, we propose a new end-to-end model that is efficient for 360° image compression. Experimental results demonstrate that our scheme outperforms other image compression methods and reduces bitrate by nearly 16% compared to the latest learned 360° image compression model. Yubiao Zhou, Yu Sun 0003, Frédéric Dufaux, Weian Li, Ce Zhu |
ICIP | 7 |
| 2025 | Incrementally Constrained Tucker Decomposition for Feature Extraction of Structural Diffusion Tensor Imaging DataabstractDiffusion Tensor Imaging (DTI) is the only in vivo technique capable of characterizing microstructural changes in the brain. The resulting feature maps, such as fractional anisotropy (FA), are three-dimensional and contain spatial details. Processing these feature maps without disrupting their structure is essential for accurate analysis. Tucker decomposition is a widely used feature extraction method for high-order data. However, it has been rarely adopted for characterizing structural DTI data. In addition, few work systematically studies the influence of its constraints on data characterization. In this study, we design the Incrementally Constrained Tucker Decomposition (ICTD) framework that progressively applies orthogonality and non-negativity constraints on decomposed factors to characterize DTI data and determine suitable constraints. The entanglement entropy is introduced to evaluate the entanglement of extracted features. Two public DTI datasets are adopted in the classification experiments. Our results demonstrate that Tucker decomposition is suitable for characterizing DTI data and the constraints are important for effective data characterization. Houji Du, Fan Zhang 0013, Yipeng Liu 0001, Ce Zhu |
ICME | 5 |
| 2025 | Unified Line Segment Detection and DescriptionabstractLine segments are fundamental elements in computer vision. However, aside from a few computationally expensive deep learning-based methods, most existing approaches treat their detection and description as independent tasks, leading to redundant computations and suboptimal performance. This paper introduces a Unified approach for Line Segment Detection and Description (ULSD2), designed for real-time vision tasks with minimal computational overhead. The core insight is to unify line segment detection and description by analyzing dedicated level lines and their differences, derived from gradients, which effectively capture the intrinsic characteristics of line segments. Furthermore, instead of the traditional scalar-based description, the use of level lines and their differences in local patches across multiple granularities enables a vectorized representation that encodes line segments from coarse to fine. Experiments demonstrate that ULSD2 outperforms other non-deep learning-based methods and competes with state-of-the-art deep learning-based methods while significantly improving efficiency. The code is available at https://github.com/roylin1229/ULSD2. Yingjie Zhou 0001, Zhen Long, Yipeng Liu 0001, Lu Yang 0002, Ce Zhu |
ICME | 6 |
| 2025 | TRR-LGF: a Simple yet Efficient Classification NetworkabstractHybrid models that combine convolution and self attention are popular for efficient local feature extraction and capturing long-range dependencies. However, these models often:1) only explore local and global features; 2) flatten high-order features at the output layer, which limit feature hierarchy exploration and the feature utility in the output layer. To address these issues, this paper introduces Tensor Ring Regression with Local-to-Global Features (TRR-LGF), a simple and effective classification network. It uses a local-to-global learning framework to capture diverse features at multiple scales. Additionally, a tensor ring regression layer replaces the linear output layer, preserving high-order feature structure and reducing parameters. Experimental results show that TRR-LGF outperforms existing state-of-the-art methods on various datasets, especially in noisy and sample-imbalanced settings. Furthermore, the model utilizes 7.6M parameters, reducing computational requirements by about 50% compared to the multilayer perceptron output layer. The code is available at https://github.com/Calcium-Oxide/TRR-LGF. Zhen Long, Hu Yao, Yipeng Liu 0001, Le Zhang 0001, Ce Zhu |
ICME | 6 |
| 2025 | Fast CU Partition Algorithm For 360-Degree Videos on VVCabstract360-degree videos (abbreviated as 360 videos) have gained widespread popularity due to their immersive experience. The massive amount of video data resulting from ultra-high resolution of 360 videos makes the coding process extremely complex and seriously limits their widespread applications. In this paper, we propose a new fast-Coding Unit (CU) partition algorithm for 360 videos based on Versatile Video Coding (VVC). The major novelty is that we establish databases and develop models based on stretching and split mode (SM) distribution for 360 videos on VVC. Specifically, first, we establish stretching-based CU partition databases for 360 videos. Second, we propose a new stretching-based multi-scale convolution kernel network structure and a corresponding synthetical loss function, which further considers both Rate-Distortion cost and imbalance issue. We then develop a unique multi-threshold selection scheme to select candidate SMs. Experimental results demonstrate that the proposed algorithm improves encoding speed by 69.63%, with only a 2.28% increase in Bjøntegaard Delta Bit Rate (BDBR), outperforming the state-of-the-art methods. Shijie Du, Yu Sun 0003, Shuyin Xia, Frédéric Dufaux, Hongwei Guo 0001, Guoyin Wang 0001, Ce Zhu |
ICME | 8 |
| 2025 | Flexibly Constrained Tucker Decomposition for High-Order Spectral AnalysisabstractSpectral analysis is widely used for sequence signals, such as electroencephalogram (EEG), speech, and radar signals, by calculating spectra using Fourier or other transforms for feature extraction. Recently, high-order spectral analysis, which models spectra into a high-order tensor and characterizes it through tensor decomposition, has gained popularity. Existing decomposition models often impose prior constraints on data. The most popular one is the orthogonality. However, enforcing strict orthogonality might fail to capture the intrinsic structure of data with complex structures. To achieve flexible data analysis, we propose flexibly constrained Tucker decomposition (FCTD) model. FCTD imposes a relaxed orthogonality constraint for each data dimension, allowing the orthogonal strength of the decomposed factors to be controlled. In addition, FCTD can be easily transformed into models with mixed constraints. Two publicly available datasets for epilepsy detection and emotional speech classification are adopted to demonstrate the effectiveness of the proposed FCTD. FCTD achieves the best performance among all models with fixed constraints. Houji Du, Nipon Theera-Umpon, Yipeng Liu 0001, Ce Zhu |
MMSP | 5 |
| 2025 | IdCo: Joint Identification and Contrastive Learning for Masked Face RecognitionabstractDeep face recognition models suffer significant accuracy declines when encountering masked faces, limiting their real-world applicability in the wild. This decline stems from misalignment between masked and regular faces in the feature space. However, improper adjustments to correct this misalignment would degrade regular face feature extracting, which is crucial for practical use. To obtain better performance in both masked face recognition (MFR) and regular face recognition (FR), we propose a joint Identification and Contrastive learning framework (IdCo), which combines identification learning with instance discrimination-based contrastive learning. In this framework, a modified loss function—Supervised Contrastive loss with Hard negative samples (SupConH)—is introduced in the contrastive learning module. This loss incorporates a hard negative sampling technique into the Supervised Contrastive loss (SupCon) to enhance learning effectiveness. Our framework promotes alignment between masked and regular faces without compromising discriminative power of learned features. To reduce the high training costs of contrastive learning, we adopt a stage-wise training strategy. In the first stage, the model is trained using identification learning alone, and IdCo is incorporated in the second stage. Extensive experiments on popular masked and regular face datasets show that IdCo outperforms state-of-the-art methods in both MFR and FR tasks. Qingtong Xu, Chao Zhang 0072, Ce Zhu |
MMSP | 5 |
| 2025 | KaRF: Weakly-Supervised Kolmogorov-Arnold Networks-based Radiance Fields for Local Color EditingabstractRecent advancements have suggested that neural radiance fields (NeRFs) show great potential in color editing within the 3D domain. However, most existing NeRF-based editing methods continue to face significant challenges in local region editing, which usually lead to imprecise local object boundaries, difficulties in maintaining multi-view consistency, and over-reliance on annotated data. To address these limitations, in this paper, we propose a novel weakly-supervised method called KaRF for local color editing, which facilitates high-fidelity and realistic appearance edits in arbitrary regions of 3D scenes. At the core of the proposed KaRF approach is a unified two-stage Kolmogorov-Arnold Networks (KANs)-based radiance fields framework, comprising a segmentation stage followed by a local recoloring stage. This architecture seamlessly integrates geometric priors from NeRF to achieve weakly-supervised learning, leading to superior performance. More specifically, we propose a residual adaptive gating KAN structure, which integrates KAN with residual connections, adaptive parameters, and gating mechanisms to effectively enhance segmentation accuracy and refine specific editing effects. Additionally, we propose a palette-adaptive reconstruction loss, which can enhance the accuracy of additive mixing results. Extensive experiments demonstrate that the proposed KaRF algorithm significantly outperforms many state-of-the-art methods both qualitatively and quantitatively. Our code and more results are available at: https://github.com/PaiDii/KARF.git. Wudi Chen, Zhiyuan Zha, Shigang Wang 0003, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Zipei Fan, Ce Zhu |
NeurIPS | 9 |
| 2025 | Multimodal Causal Reasoning for UAV Object DetectionabstractUnmanned Aerial Vehicle (UAV) object detection faces significant challenges due to complex environmental conditions and different imaging conditions. These factors introduce significant changes in scale and appearance, particularly for small objects that occupy limited pixels and exhibit limited information, complicating detection tasks. To address these challenges, we propose a Multimodel Causal Reasoning framework based on YOLO backbone for UAV Object Detection (MCR-UOD). The key idea is to use the backdoor adjustment to discover the condition-invariant object representation for easy detection. Specifically, the YOLO backbone is first adjusted to incorporate the pre-trained vision-language model. The original category labels are replaced with semantic text prompts, and the detection head is replaced with text-image contrastive learning. Based on this backbone, our method consists of two parts. The first part, named language guided region exploration, discovers the regions with high probability of object existence using text embeddings based on vision-language model such as CLIP. Another part is the backdoor adjustment casual reasoning module, which constructs a confounder dictionary tailored to different imaging conditions to capture global image semantics and derives a prior probability distribution of shooting conditions. During causal inference, we use the confounder dictionary and the prior to intervene on local instance features, disentangling condition variations, and obtaining condition-invariant representations. Experimental results on several public datasets confirm the state-of-the-art performance of our approach. The code, data and models will be released upon publication of this paper. Nianxin Li, Mao Ye 0001, Lihua Zhou, Shuaifeng Li, Song Tang 0001, Luping Ji, Ce Zhu |
NeurIPS | 7 |
| 2025 | Multimodal Sensing for Intelligent V2X: A Review of Recent Advances Toward DeploymentabstractThe integration of multi-modal sensing and communication systems advances the integrated sensing and communication (ISAC) concept by combining camera, LiDAR, radar, and radio-frequency (RF) data. This approach provides a promising avenue to enhance wireless link reliability and spectral efficiency while reducing end-to-end latency in 6G vehicle-to-everything (V2X) networks. Surveys on sensing and communication integration primarily concentrate on RF technologies. While initial studies have explored the combination of various sensing methods, there is a paucity of comprehensive reviews regarding how such data is utilized to enhance V2X communication performance. Therefore, this review comprehensively analyzes recent research on multi-modal sensing-assisted vehicular networks to enhance communication performance, centered around channel prediction, beam management, mobility management, and multi-task solutions. Since data-centric techniques are the main approach in this field, relevant datasets and simulators supporting data-driven designs are also thoroughly examined. Finally, by critically analyzing current research, the survey highlights unresolved issues and future trends in research cornerstones, method implementation, and system design. It also addresses key deployment challenges—spectrum availability, roadside hardware costs, and privacy-aware learning—thereby charting the course for developing real-world intelligent V2X networks. Kang Tan, Ce Zhu |
IEEE Internet Things J. | 2 |
| 2025 | Correction to: Deep negative correlation classification
Le Zhang 0001, Qibin Hou, Yun Liu 0011, Jiawang Bian, Xun Xu 0002, Joey Tianyi Zhou, Ce Zhu |
Mach. Learn. | 7 |
| 2025 | Towards Real Zero-Shot Camouflaged Object Segmentation Without Camouflaged AnnotationsabstractCamouflaged Object Segmentation (COS) faces significant challenges due to the scarcity of annotated data, where meticulous pixel-level annotation is both labor-intensive and costly, primarily due to the intricate object-background boundaries. Addressing the core question, "Can COS be effectively achieved in a zero-shot manner without manual annotations for any camouflaged object?", we propose an affirmative solution. We examine the learned attention patterns for camouflaged objects and introduce a robust zero-shot COS framework. Our findings reveal that while transformer models for salient object segmentation (SOS) prioritize global features in their attention mechanisms, camouflaged object segmentation exhibits both global and local attention biases. Based on these findings, we design a framework that adapts with the inherent local pattern bias of COS while incorporating global attention patterns and a broad semantic feature space derived from SOS. This enables efficient zero-shot transfer for COS. Specifically, We incorporate a Masked Image Modeling (MIM) based image encoder optimized for Parameter-Efficient Fine-Tuning (PEFT), a Multimodal Large Language Model (M-LLM), and a Multi-scale Fine-grained Alignment (MFA) mechanism. The MIM encoder captures essential local features, while the PEFT module learns global and semantic representations from SOS datasets. To further enhance semantic granularity, we leverage the M-LLM to generate caption embeddings conditioned on visual cues, which are meticulously aligned with multi-scale visual features via MFA. This alignment enables precise interpretation of complex semantic contexts. Moreover, we introduce a learnable codebook to represent the M-LLM during inference, significantly reducing computational demands while maintaining performance. Our framework demonstrates its versatility and efficacy through rigorous experimentation, achieving state-of-the-art performance in zero-shot COS with $F_{\beta }^{w}$Fβw scores of 72.9% on CAMO and 71.7% on COD10K. By removing the M-LLM during inference, we achieve an inference speed comparable to that of traditional end-to-end models, reaching 18.1 FPS. Additionally, our method excels in polyp segmentation, and underwater scene segmentation, outperforming challenging baselines in both zero-shot and supervised settings, thereby implying its potentiality in various segmentation tasks. Tian-Zhu Xiang, Ao Li 0007, Ce Zhu, Le Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Exploring Frequency-Inspired Optimization in Transformer for Efficient Single Image Super-ResolutionabstractTransformer-based methods have exhibited remarkable potential in single image super-resolution (SISR) by effectively extracting long-range dependencies. However, most of the current research in this area has prioritized the design of transformer blocks to capture global information, while overlooking the importance of incorporating high-frequency priors, which we believe could be beneficial. In our study, we conducted a series of experiments and found that transformer structures are more adept at capturing low-frequency information, but have limited capacity in constructing high-frequency representations when compared to their convolutional counterparts. Our proposed solution, the cross-refinement adaptive feature modulation transformer (CRAFT), integrates the strengths of both convolutional and transformer structures. It comprises three key components: the high-frequency enhancement residual block (HFERB) for extracting high-frequency information, the shift rectangle window attention block (SRWAB) for capturing global information, and the hybrid fusion block (HFB) for refining the global representation. To tackle the inherent intricacies of transformer structures, we introduce a frequency-guided post-training quantization (PTQ) method aimed at enhancing CRAFT's efficiency. These strategies incorporate adaptive dual clipping and boundary refinement. To further amplify the versatility of our proposed approach, we extend our PTQ strategy to function as a general quantization method for transformer-based SISR techniques. Our experimental findings showcase CRAFT's superiority over current state-of-the-art methods, both in full-precision and quantization scenarios. These results underscore the efficacy and universality of our PTQ strategy. Ao Li 0007, Le Zhang 0001, Yun Liu 0011, Ce Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | TLRLF4MVC: Tensor Low-Rank and Low-Frequency for Scalable Multi-View ClusteringabstractAnchor-based multi-view clustering has garnered much attention for its effectiveness in handling massive datasets. However, current methods either fail to consider intra-view similarity or require ($\mathcal {O}(N^{3})$O(N3)) for exploring intra-view similarity, making efficient large-scale multi-view clustering difficult. This paper introduces a novel tensor low-frequency component (TLFC) operator, which achieves smooth representation among samples. Furthermore, this TLFC operator, which explores intra-view similarity, incorporates tensor nuclear norm (TNN) operator and consensus regularization that explore inter-view correlations, resulting in the development of tensor low-rank and low-frequency for scalable multi-view clustering (TLRLF4MVC). Iteratively, as intra-view sample similarity and complementary information across views achieve balance, the learned embedding features are mapped into a smooth and compact subspace, ultimately leading to outstanding clustering performance. Extensive experiments on six large-scale multi-view datasets demonstrate that TLRLF4MVC not only significantly outperforms state-of-the-art methods in terms of clustering accuracy but also achieves remarkable computational efficiency, particularly when handling massive data. Zhen Long, Yazhou Ren 0001, Yipeng Liu 0001, Ce Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Deep-Learning-Empowered Super Resolution: A Comprehensive Survey and Future ProspectsabstractSuper-resolution (SR) has garnered significant attention within the computer vision community, driven by advances in deep learning (DL) techniques and the growing demand for high-quality visual applications. With the expansion of this field, numerous surveys have emerged. Most existing surveys focus on specific domains, lacking a comprehensive overview of this field. Here, we present an in-depth review of diverse SR methods, encompassing single-image SR (SISR), video SR (VSR), stereo SR (SSR), and light field SR (LFSR). We extensively cover over 150 SISR methods, nearly 70 VSR approaches, and approximately 30 techniques for SSR and LFSR. We analyze methodologies, datasets, evaluation protocols, empirical results, and complexity. In addition, we conducted a taxonomy based on each backbone structure according to the diverse purposes. We also explore valuable yet understudied open issues in the field. We believe that this work will serve as a valuable resource and offer guidance to researchers in this domain. To facilitate access to related work, we created a dedicated repository available at https://github.com/AVC2-UESTC/Holistic-Super-Resolution-Review Le Zhang 0001, Ao Li 0007, Qibin Hou, Ce Zhu, Yonina C. Eldar |
Proc. IEEE | 4 |
| 2025 | Continuous Patch Stitching for Block-Wise Image CompressionabstractMost recently, learned image compression methods have outpaced traditional hand-crafted standard codecs. However, their inference typically requires to input the whole image at the cost of heavy computing resources, especially for high-resolution image compression; otherwise, the block artefact can exist when compressed by blocks within existing learned image compression methods. To address this issue, we propose a novel continuous patch stitching (CPS) framework for block-wise image compression that is able to achieve seamlessly patch stitching and mathematically eliminate block artefact, thus capable of significantly reducing the required computing resources when compressing images. More specifically, the proposed CPS framework is achieved by padding-free operations throughout, with a newly established parallel overlapping stitching strategy to provide a general upper bound for ensuring the continuity. Upon this, we further propose functional residual blocks with even-sized kernels to achieve down-sampling and up-sampling, together with bottleneck residual blocks retaining feature size to increase network depth. Experimental results demonstrate that our CPS framework achieves the state-of-the-art performance against existing baselines, whilst requiring less than half of computing resources of existing models. The source code and trained models are available athttps://github.com/bblgbr/SPL-CPS. Zifu Zhang, Shengxi Li, Henan Liu, Mai Xu, Ce Zhu |
IEEE Signal Process. Lett. | 5 |
| 2025 | Joint Resources Optimization for Soft Video Transmission Over IRS-Assisted SR NetworkabstractIntelligent reflective surface (IRS) assisted symbiotic radio (SR) network has been proposed as a promising solution for the sixth generation (6G) mobile wireless system, which achieves mutualistic spectrum sharing and highly reliable backscattering communication with extremely low energy cost. On the other hand, exponential growth in video traffic makes wireless video transmission more challenging in the 6G era. With the assistance of IRS based secondary link in SR network, an efficient soft video transmission scheme (IRSCast) is proposed to achieve linear quality transition under the drastically varying wireless channel. To minimize the transmission distortion of the video signal, a multivariable optimization problem is formulated to jointly optimize the wireless resources, including transmission power, active beamforming of the primary transmitter (PTx), and passive beamforming of the secondary transmitter (STx). Then, an alternating optimization method is utilized to decouple the multivariate optimization problem into multiple univariate sub-problems that are finally solved by semi-positive definite relaxation and Lagrange multiplier methods. The simulation results demonstrated that the proposed IRSCast method significantly improves the objective and subjective quality of the received video. Lei Luo 0003, Zhi Jin 0002, Hongwei Guo 0001, Ce Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Triply Laplacian Scale Mixture Modeling for Seismic Data Noise Suppression
Sirui Pan, Zhiyuan Zha, Shigang Wang 0003, Yue Li 0003, Zipei Fan, Bihan Wen, Ce Zhu |
IEEE Trans. Geosci. Remote. Sens. | 9 |
| 2025 | HOPE: Enhanced Position Image Priors via High-Order Implicit RepresentationsabstractDeep Image Prior (DIP) has shown that networks with stochastic initialization and custom architectures can effectively address inverse imaging challenges. Despite its potential, DIP requires significant computational resources, whereas the lighter Implicit Neural Positional Image Prior (PIP) often yields overly smooth solutions due to exacerbated spectral bias. Research on lightweight, high-performance solutions for inverse imaging remains limited. This paper proposes a novel framework, Enhanced Positional Image Priors through High-Order Implicit Representations (HOPE), incorporating high-order interactions between layers within a conventional cascade structure. This approach reduces the spectral bias commonly seen in PIP, enhancing the model's ability to capture both low- and high-frequency components for optimal inverse problem performance. We theoretically demonstrate that HOPE's expanded representational space, narrower convergence range, and improved Neural Tangent Kernel (NTK) diagonal properties enable more precise frequency representations than PIP. Comprehensive experiments across tasks such as signal representation (audio, image, volume) and inverse image processing (denoising, super-resolution, CT reconstruction, inpainting) confirm that HOPE establishes new benchmarks for recovery quality and training efficiency. Ruituo Wu, Junhui Hou, Ce Zhu, Yipeng Liu 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Texture-Consistent 3D Scene Style Transfer via Transformer-Guided Neural Radiance FieldsabstractRecent advancements have suggested that neural radiance fields (NeRFs) show great potential in 3D style transfer. However, most existing NeRF-based style transfer methods still face considerable challenges in generating stylized images that simultaneously preserve clear scene textures and maintain strong cross-view consistency. To address these limitations, in this paper, we propose a novel transformer-guided approach for 3D scene style transfer. Specifically, we first design a transformer-based style transfer network to capture long-range dependencies and generate 2D stylized images with initial consistency, which serve as supervision for the 3D stylized generation. To enable fine-grained control over style, we propose a latent style vector as a conditional feature and design a style network that projects this style information into the 3D space. We further develop a merge network that integrates style features with scene geometry to render 3D stylized images that are both visually coherent and stylistically consistent. In addition, we propose a texture consistency loss to preserve scene structure and enhance texture fidelity across views. Extensive quantitative and qualitative experimental results demonstrate that our proposed approach outperforms many state-of-the-art methods in terms of visual perception, image quality and multi-view consistency. Our code and more results are available at: https://github.com/PaiDii/TGTC-Style.git. Wudi Chen, Zhiyuan Zha, Shigang Wang 0003, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu |
IEEE Trans. Image Process. | 8 |
| 2025 | Hierarchical Semantic Compression for Consistent Image Semantic RestorationabstractThe emerging semantic compression has been receiving increasing research efforts most recently, capable of achieving high fidelity restoration during compression, even at extremely low bitrates. However, existing semantic compression methods typically combine standard pipelines with either pre-defined or high-dimensional semantics, thus suffering from deficiency in compression. To address this issue, we propose a novel hierarchical semantic compression (HSC) framework that purely operates within intrinsic semantic spaces from generative models, which is able to achieve efficient compression for consistent semantic restoration. More specifically, we first analyse the entropy models for the semantic compression, which motivates us to employ a hierarchical architecture based on a newly developed general inversion encoder. Then, we propose the feature compression network (FCN) and semantic compression network (SCN), such that the middle-level semantic feature and core semantics are hierarchically compressed to restore both accuracy and consistency of image semantics, via an entropy model progressively shared by channel-wise context. Experimental results demonstrate that the proposed HSC framework achieves the state-of-the-art performance on subjective quality and consistency for human vision, together with superior performances on machine vision tasks given compressed bitstreams. This essentially coincides with human visual system in understanding images, thus providing a new framework for future image/video compression paradigms. The source code and trained models are available at https://github.com/bblgbr/HSC-TIP2025. Shengxi Li, Zifu Zhang, Mai Xu, Lai Jiang 0004, Yufan Liu 0001, Ce Zhu |
IEEE Trans. Image Process. | 6 |
| 2025 | DSMT: Dual-Stage Multiscale Transformer for Hyperspectral Snapshot Compressive ImagingabstractSnapshot compressive imaging (SCI) compresses a 3D hyperspectral image (HSI) into a 2D measurement, significantly improving imaging efficiency while preserving the spatial and spectral information inherent in HSI. However, reconstructing high-quality HSIs from compressed measurements remains a core challenge due to the complexity of the inverse problem. Transformer-based methods have recently shown promising performance in HSI reconstruction. Nonetheless, effectively capturing local information, long-range dependencies, and multi-scale features within a reasonable computational cost remains a significant challenge. In this paper, we propose a dual-stage multiscale Transformer (DSMT) tailored for HSI reconstruction, which adopts a coarse-to-fine framework to enhance reconstruction accuracy and network generalization. Specifically, we design a novel U-Net architecture with a dual-branch encoder, where two separate branches process distinct features and are fused to achieve more refined reconstruction results. Full-scale skip connections are introduced to strengthen feature fusion across different stages. To further improve performance, we develop a novel self-attention mechanism called dual-window multiscale multi-head self-attention (DWM-MSA). By utilizing two differently sized windows, DWM-MSA captures long-range dependencies and local information at multiple scales, significantly boosting reconstruction quality. Additionally, we introduce a hybrid positional embedding method, conditional/relative positional embedding (CRPE), which dynamically models both spatial and spectral dependencies, effectively enhancing the Transformer's capacity for HSI reconstruction. Extensive quantitative and qualitative experiments on both the simulated and the real data are conducted to demonstrate the superior performance, stability, and generalization ability of our DSMT. Code of this project is at https://github.com/chenx2000/DSMT. Fulin Luo, Xi Chen 0087, Tan Guo, Xiuwen Gong, Lefei Zhang, Ce Zhu |
IEEE Trans. Image Process. | 6 |
| 2025 | Hypergraph Mamba Reasoning-Based Social Relation RecognitionabstractRecognizing social relations from images is crucial for improving machine perception of social interactions. Current studies mainly focus on exploring single-type relation reasoning frameworks, such as the relation between father, mother and son in a family. However, real-world scenarios often involve complex hybrid relations, such as friendships and professional relations, which pose a challenge for current methods due to the difficulty of establishing robust logical connections between these relations. In fact, in this hybrid social relation recognition setting, the interactions extend beyond dyadic to multipartite structures. To effectively explore these multipartite interactions, we propose a novel Hypergraph Mamba (HGM) framework. Specifically, we construct two hypergraphs, i.e., Person-Person Hypergraphs (PPH) and Person-Object Hypergraphs (POH), to model these high-order multipartite interactions. The HGM module performs social relation reasoning within these hypergraph structures, which includes a Vertex Selection Algorithm to mitigate inference confusion by filtering out confounders, and a Vertex Interaction Operator to find optimal global vertex neighborhoods by capturing long-range vertex dependencies. In addition, a Multilevel Transformer is proposed to adaptively align the PPH and POH inferred knowledge and visual signals to facilitate information fusion. We validate the effectiveness of our proposed HGM model on several public datasets and perform extensive ablation studies to elucidate the reasons contributing to its superior performance. Experimental results indicate that our HGM model achieves superior accuracy in predicting social relations compared to the state-of-the-art methods. Codes and datasets are available at: https://github.com/tw-repository/HGM-SRR. Wang Tang, Linbo Qing, Pingyu Wang, Lindong Li, Ce Zhu |
IEEE Trans. Image Process. | 5 |
| 2025 | Towards Open-Vocabulary Video Semantic SegmentationabstractSemantic segmentation in videos has been a focal point of recent research. However, existing models encounter challenges when faced with unfamiliar categories. To address this, we introduce the Open Vocabulary Video Semantic Segmentation (OV-VSS) task, designed to accurately segment every pixel across a wide range of open-vocabulary categories, including those that are novel or previously unexplored. To enhance OV-VSS performance, we propose a robust baseline, OV2VSS, which integrates a spatial-temporal fusion module, allowing the model to utilize temporal relationships across consecutive frames. Additionally, we incorporate a random frame enhancement module, broadening the model's understanding of semantic context throughout the entire video sequence. Our approach also includes video text encoding, which strengthens the model's capability to interpret textual information within the video context. Comprehensive evaluations on benchmark datasets such as VSPW and Cityscapes highlight OV-VSS's zero-shot generalization capabilities, especially in handling novel categories. The results validate OV2VSS's effectiveness, demonstrating improved performance in semantic segmentation tasks across diverse video datasets. Yun Liu 0011, Guolei Sun, Min Wu 0008, Le Zhang 0001, Ce Zhu |
IEEE Trans. Multim. | 6 |
| 2025 | Depth Map Super-Resolution via Deep Cross-Modality and Cross-Scale GuidanceabstractGuided depth super-resolution is essential in many applications, which enhances low-resolution (LR) depth maps using high-resolution (HR) RGB images from the same scene. However, the challenge lies in avoiding the texture-copy artifacts issue caused by structural inconsistencies between two modalities. To mitigate, we propose a cross-modality and cross-scale guided depth super-resolution network (D2CNet). We first design a novel two-stage feature integration module to effectively fuse multi-modal RGB and depth while minimizing texture-copy artifacts. That is, a cross-modality fusion stage transfers consistent structures from RGB to depth in a multi-scale manner, and a cross-scale refinement stage mitigates inconsistent structures across modalities. In addition, we design a convolution group as the basic module to well extract high-frequency features and an LR and HR domain projection strategy to enrich features between the fusion and refinement stages. We then develop a new network architecture by progressively repeating the feature integration module and the convolution group, which is flexibly controllable to strike a balance between accuracy and cost for easy implementation in real world. Extensive experiments on multiple benchmarks demonstrate that our D2CNet consistently achieves superior accuracy and generalization ability across sampling scales in both qualitative and quantitative evaluations, when compared to state-of-the-art baselines. Shuzhe Liu, Delong Suzhang, Meng Yang 0002, Xinhu Zheng, Ce Zhu |
IEEE Trans. Multim. | 5 |
| 2025 | DA-Flow: Dual Attention Normalizing Flow for Skeleton-Based Video Anomaly DetectionabstractCooperation between temporal convolutional networks (TCN) and graph convolutional networks (GCN) as a processing module has shown promising results in skeleton-based video anomaly detection (SVAD). However, to maintain a lightweight model with low computational and storage complexity, shallow GCN and TCN blocks are constrained by small receptive fields and a lack of cross-dimension interaction capture. To tackle this limitation, we propose a lightweight module called the Dual Attention Module (DAM) for capturing cross-dimension interaction relationships in spatio-temporal skeletal data. It employs the frame attention mechanism to identify the most significant frames and the skeleton attention mechanism to capture broader relationships across fixed partitions with minimal parameters and total Floating Point Operations (FLOPs). Furthermore, the proposed Dual Attention Normalizing Flow (DA-Flow) integrates the DAM as a post-processing unit after GCN within the normalizing flow framework. Simulations show that the proposed model is robust against noise and negative samples. Experimental results show that DA-Flow reaches competitive or better performance than the existing state-of-the-art (SOTA) methods in terms of the micro AUC metric with the fewest parameters and FLOPs. Moreover, we found that even without training, simply using random projection without dimensionality reduction on skeleton data enables substantial anomaly detection capabilities. Ruituo Wu, Bing Li 0002, Jicong Fan 0001, Frédéric Dufaux, Ce Zhu, Yipeng Liu 0001 |
IEEE Trans. Multim. | 7 |
| 2025 | Online Nonconvex Robust Tensor Principal Component AnalysisabstractRobust tensor principal component analysis (RTPCA) based on tensor singular value decomposition (t-SVD) separates the low-rank component and the sparse component from the multiway data. For streaming data, online RTPCA (ORTPCA) processes tensor data sequentially, where the low-rank component is updated based on the latest estimation and the newly arrived sample. It enhances both computation and storage efficiency. However, in most of the existing ORTPCA methods, the relaxation from tensor multirank to the convex tensor nuclear norm (TNN) may have a certain modeling error, which leads to unavoidable tracking accuracy loss. In this article, a tensor Schatten-p norm ( $0\lt p\lt 1$ ) is applied to provide a tighter approximation of the tensor rank. A Lemma is deduced to divide the Schatten-p norm into terms to be updated in an online way. Based on it, the corresponding online nonconvex RTPCA (ONRTPCA) method is proposed for efficient tensor subspace tracking. Moreover, we incorporate the dynamic forgetting window into ONRTPCA to adaptively track varying subspaces. In addition, this article also provides convergence analysis and complexity analysis. Experimental results on synthetic data and real-world video data show that our proposed method achieves superior subspace tracking accuracy in comparison with a series of state-of-the-art methods while maintaining a high convergence speed and low memory requirement. Lanlan Feng, Yipeng Liu 0001, Ce Zhu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | RCBEVDet: Radar-Camera Fusion in Bird's Eye View for 3D Object DetectionabstractThree-dimensional object detection is one of the key tasks in autonomous driving. To reduce costs in practice, low-cost multi-view cameras for 3D object detection are proposed to replace the expansive LiDAR sensors. However, relying solely on cameras is difficult to achieve highly accurate and robust 3D object detection. An effective solution to this issue is combining multi-view cameras with the economical millimeter-wave radar sensor to achieve more reliable multi-modal 3D object detection. In this paper, we introduce RCBEVDet, a radar-camera fusion 3D object detection method in the bird's eye view (BEV). Specifically, we first design RadarBEVNet for radar BEV feature extraction. RadarBEVNet consists of a dual-stream radar backbone and a Radar Cross-Section (RCS) aware BEV encoder. In the dual-stream radar backbone, a point-based encoder and a transformer-based encoder are proposed to extract radar features, with an injection and extraction module to facilitate communication between the two encoders. The RCS-aware BEV encoder takes RCS as the object size prior to scattering the point feature in BEV. Besides, we present the Cross-Attention Multi-layer Fusion module to automatically align the multi-modal BEV feature from radar and camera with the deformable attention mechanism, and then fuse the feature with channel and spatial fusion layers. Experimental results show that RCBEVDet achieves new state-of-the-art radar-camera fusion results on nuScenes and view-of-delft (VoD) 3D object detection benchmarks. Furthermore, RCBEVDet achieves better 3D detection results than all real-time camera-only and radar-camera 3D object detectors with a faster inference speed at 21∼28 FPS. The source code will be released at https://github.com/VDIGPKU/RCBEVDet. Zhongyu Xia, Yongtao Wang, Shengxiang Qi, Nan Dong, Ce Zhu |
CVPR | 10 |
| 2024 | S2MVTC: A Simple Yet Efficient Scalable Multi-View Tensor ClusteringabstractAnchor-based large-scale multi-view clustering has attracted considerable attention for its effectiveness in handling massive datasets. However, current methods mainly seek the consensus embedding feature for clustering by exploring global correlations between anchor graphs or projection matrices. In this paper, we propose a simple yet efficient scalable multi-view tensor clustering (S2MVTC) approach, where our focus is on learning correlations of embedding features within and across views. Specifically, we first construct the embedding feature tensor by stacking the embedding features of different views into a tensor and rotating it. Additionally, we build a novel tensor low-frequency approximation (TLFA) operator, which incorporates graph similarity into embedding feature learning, efficiently achieving smooth representation of embedding features within different views. Furthermore, consensus constraints are applied to embedding features to ensure inter-view semantic consistency. Experimental results on six large-scale multi-view datasets demonstrate that S2MVTC significantly outperforms state-of-the-art algorithms in terms of clustering performance and CPU execution time, especially when handling massive data. The code of S2MVTC is publicly available at https://github.com/longzhen520/S2MVTC. Zhen Long, Yazhou Ren 0001, Yipeng Liu 0001, Ce Zhu |
CVPR | 5 |
| 2024 | Binomial Self-Compensation for Motion Error in Dynamic 3D Scanning
Geyou Zhang, Ce Zhu, Kai Liu 0012 |
ECCV (20) | 2 |
| 2024 | Phase Retrieval by Tensor Total Least SquaresabstractPhase retrieval seeks to reconstruct a series of image sequences from measurements that only capture their magnitudes. Current approaches either flatten and stack the image sequences, disregarding their multidimensional structural information, or fail to account for errors within the sensing vectors/tensors. To address these two issues simultaneously, we propose a unified framework for the phase retrieval problem, namely tensor total least squares (TTLS). Specifically, we set up a tensor representation for image sequences and the corresponding measurement model, and for the first time employ the advanced tensor ring network to effectively explore the inherent multidimensional structure for more accurate estimation. Moreover, in addition to the additive noise, the multiplicative errors within the sensing tensor can be also well-corrected, leading to a more robust estimation. Experimental results on both simulated data and real videos demonstrate the superiority of the proposed method. Jiani Liu 0002, Ce Zhu, Xiaolin Huang, Yipeng Liu 0001 |
ICASSP | 2 |
| 2024 | Multi-Band Speech Tensor Decomposition for Interactive Feature Extraction in Early Dysphagia ScreeningabstractDysphagia is a prevalent symptom in numerous neurological disorders among older adults. Current dysphagia diagnostic systems either involve invasive procedures or necessitate the ingestion of liquids. Some researchers have devised automatic dysphagia detection methods based on vowels that are easy to collect and sensitive to vocal cord states. These methods extract features from each vowel separately and fuse them to train models. Nonetheless, they neglect potential interrelations among different vowels. Vowels collected from the same speaker could share subspaces since they are produced from the same vocal system. In this study, we introduce a tensor-based method that can simultaneously extract interactive information from all vowels across different modes. This method designs multi-band speech tensors and core-pruned tensor networks to investigate crucial frequency bands and connections for dysphagia screening. Experimental results show our model exceeds previous methods by approximately 10 percentage points in the classification accuracy. Yipeng Liu 0001, Da Shen, Yangyang Jiang, Ce Zhu |
ICASSP | 6 |
| 2024 | Fast Intra Mode Prediction Algorithms for SCBS in VVC SCCabstractVersatile Video Coding (VVC) now supports Screen Content Coding (SCC) by integrating two efficient coding modes: Intra Block Copy (IBC) and Palette (PLT). However, the numerous modes and the Quad-Tree Plus Multi-Type Tree (QTMT) structure inherent to VVC contribute to a very high coding complexity. To effectively reduce the computational complexity of VVC SCC, we propose a fast Intra mode prediction algorithm for VVC SCC. More specifically, we first use the difference of minimum Sum of Absolute Transformed Differences (SATD) value of four Directional Modes (DMs) of Intra and the SATD value of the IBC-merge mode to determine whether to early skip Intra checking. Subsequently, we use a decision tree to determine whether to early terminate the checking after block differential pulse coded modulation (BDPCM). Finally, we employ a decision tree to determine whether to early skip multiple transform selection (MTS) and low frequency non-separable transform (LFNST) checking. The results demonstrate that our algorithm achieves an average encoding time reduction of 34.34% with a negligible Bjøntegaard delta bitrate increase of 0.46%. Yishen Deng, Weisheng Li 0001, Xin Lu 0001, Frédéric Dufaux, Bo Hang, Ce Zhu |
ICASSP | 7 |
| 2024 | Fast Coding Mode Prediction for Intra Prediction in VVC SCCabstractCurrently, screen content video applications are increasingly widespread in our daily lives. The latest Screen Content Coding (SCC) standard, known as Versatile Video Coding (VVC) SCC, employs screen content Coding Modes (CMs) selection. While VVC SCC achieves high coding efficiency, its coding complexity poses a significant obstacle to the further widespread adoption of screen content video. Hence, it is crucial to enhance the coding speed of VVC SCC. In this paper, we propose a fast mode and splitting decision for Intra prediction in VVC SCC. Specifically, we initially exploit deep learning techniques to predict content types for all CUs. Subsequently, we examine CM distributions of different content types to predict candidate CMs for CUs. We then introduce early skip and early terminate CM decisions for different content types of CUs to further eliminate unlikely CMs. Finally, we develop Block-based Differential Pulse-Code Modulation (BDPCM) early termination to improve coding speed. Experimental results demonstrate that the proposed algorithm can improve coding speed by $34.95 \%$ on average while maintaining almost the same coding efficiency. Junyi Yu, Xin Lu 0001, Frédéric Dufaux, Hongwei Guo 0001, Ce Zhu |
ICIP | 7 |
| 2024 | Efficient Black-Box Adversarial Attack on Deep Clustering ModelsabstractDespite the significant progress made by deep clustering models in high-dimensional data processing, they remain vulnerable to adversarial examples. However, research on adversarial attacks against deep clustering algorithms appears to be relatively underexplored. To fill this gap, we propose a query-efficient black-box attack on deep clustering models, which leverages the transferability between different deep clustering models. Initially, we train a generator using a substitute deep clustering model, reducing the number of queries to the target model. Subsequently, when targeting an unknown deep clustering model, we employ the target query information to update both the substitute deep clustering model and the generator. Experimental evaluations on four state-of-the-art deep clustering models across three datasets demonstrate the efficacy of our method in disrupting clustering performance. The results indicate that our approach surpasses the performance of existing methods. Zhen Long, Xiaolin Huang, Ce Zhu, Yipeng Liu 0001 |
ICIP | 5 |
| 2024 | Epilepsy Detection with Personal Identification Based on Regularized O-minus DecompositionabstractEpilepsy is a common neurological disease that seriously affects the patient’s life quality. Electroencephalogram (EEG) is an important modality for epilepsy diagnosis and treatment. Here, we propose a tensor-based method using EEG, which can detect patients with seizures and identify them. In this method, we transform EEG signals into high-order tensors and design a regularized O-minus tensor network to calculate representative signals of channels. The interactive information among all tensor modes can be extracted by the regularized O-minus. Finally, the brain network is calculated based on the extracted representative signals and used as a shared feature set for epilepsy detection and patient identity. Leave one subject out cross-validation is used in seizure detection to verify the generalization of the model for first-time patients. The accuracies of seizure detection and patient identification are 88.14% and 98.59%, respectively. To our knowledge, this is the first time that epilepsy detection and patient identification are implemented within a unified framework, which could provide timely and personalized treatment for patients. Da Shen, Zhongrong Wang, Ce Zhu, Yipeng Liu 0001 |
ISCAS | 5 |
| 2024 | Causal Context Adjustment Loss for Learned Image CompressionabstractIn recent years, learned image compression (LIC) technologies have surpassed conventional methods notably in terms of rate-distortion (RD) performance. Most present learned techniques are VAE-based with an autoregressive entropy model, which obviously promotes the RD performance by utilizing the decoded causal context. However, extant methods are highly dependent on the fixed hand-crafted causal context. The question of how to guide the auto-encoder to generate a more effective causal context benefit for the autoregressive entropy models is worth exploring. In this paper, we make the first attempt in investigating the way to explicitly adjust the causal context with our proposed Causal Context Adjustment loss (CCA-loss). By imposing the CCA-loss, we enable the neural network to spontaneously adjust important information into the early stage of the autoregressive entropy model. Furthermore, as transformer technology develops remarkably, variants of which have been adopted by many state-of-the-art (SOTA) LIC techniques. The existing computing devices have not adapted the calculation of the attention mechanism well, which leads to a burden on computation quantity and inference latency. To overcome it, we establish a convolutional neural network (CNN) image compression model and adopt the unevenly channel-wise grouped strategy for high efficiency. Ultimately, the proposed CNN-based LIC network trained with our Causal Context Adjustment loss attains a great trade-off between inference latency and rate-distortion performance. Minghao Han, Shiyin Jiang, Shengxi Li, Xin Deng 0002, Mai Xu, Ce Zhu, Shuhang Gu |
NeurIPS | 6 |
| 2024 | Image compressive sensing reconstruction via nonlocal low-rank residual-based ADMM framework
Junhao Zhang 0005, Kim-Hui Yap, Lap-Pui Chau, Ce Zhu |
Comput. Vis. Image Underst. | 4 |
| 2024 | Towards real-time practical image compression with lightweight attention
Minfeng Huang, Lei Luo 0003, Xu Yang 0030, Ce Zhu |
Expert Syst. Appl. | 5 |
| 2024 | Analyzing the pregnancy status of giant pandas with hierarchical behavioral information
Xianggang Li, Jing Wu 0021, Rong Hou, Zhangyu Zhou, Chang Duan, Mengnan He, Yingjie Zhou 0001, Ce Zhu |
Expert Syst. Appl. | 10 |
| 2024 | Disentanglement then reconstruction: Unsupervised domain adaptation by twice distribution alignments
Lihua Zhou, Mao Ye 0001, Xinpeng Li 0005, Ce Zhu, Yiguang Liu, Xue Li 0001 |
Expert Syst. Appl. | 4 |
| 2024 | Deep negative correlation classification
Le Zhang 0001, Qibin Hou, Yun Liu 0011, Jiawang Bian, Xun Xu 0002, Joey Tianyi Zhou, Ce Zhu |
Mach. Learn. | 7 |
| 2024 | A Comprehensive Review of Image Line Segment Detection and Description: Taxonomies, Comparisons, and ChallengesabstractAn image line segment is a fundamental low-level visual feature that delineates straight, slender, and uninterrupted portions of objects and scenarios within images. Detection and description of line segments lay the basis for numerous vision tasks. Although many studies have aimed to detect and describe line segments, a comprehensive review is lacking, obstructing their progress. This study fills the gap by comprehensively reviewing related studies on detecting and describing two-dimensional image line segments to provide researchers with an overall picture and deep understanding. Based on their mechanisms, two taxonomies for line segment detection and description are presented to introduce, analyze, and summarize these studies, facilitating researchers to learn about them quickly and extensively. The key issues, core ideas, advantages and disadvantages of existing methods, and their potential applications for each category are analyzed and summarized, including previously unknown findings. The challenges in existing methods and corresponding insights for potentially solving them are also provided to inspire researchers. In addition, some state-of-the-art line segment detection and description algorithms are evaluated without bias, and the evaluation code will be publicly available. The theoretical analysis, coupled with the experimental results, can guide researchers in selecting the best method for their intended vision applications. Finally, this study provides insights for potentially interesting future research directions to attract more attention from researchers to this field. Yingjie Zhou 0001, Yipeng Liu 0001, Ce Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Compressed-SDR to HDR Video ReconstructionabstractThe new generation of organic light emitting diode display is designed to enable the high dynamic range (HDR), going beyond the standard dynamic range (SDR) supported by the traditional display devices. However, a large quantity of videos are still of SDR format. Further, most pre-existing videos are compressed at varying degrees for minimizing the storage and traffic flow demands. To enable movie-going experience on new generation devices, converting the compressed SDR videos to the HDR format (i.e., compressed-SDR to HDR conversion) is in great demands. The key challenge with this new problem is how to solve the intrinsic many-to-many mapping issue. However, without constraining the solution space or simply imitating the inverse camera imaging pipeline in stages, existing SDR-to-HDR methods can not formulate the HDR video generation process explicitly. Besides, they ignore the fact that videos are often compressed. To address these challenges, in this work we propose a novel imaging knowledge-inspired parallel networks (termed as KPNet) for compressed-SDR to HDR (CSDR-to-HDR) video reconstruction. KPNet has two key designs: Knowledge-Inspired Block (KIB) and Information Fusion Module (IFM). Concretely, mathematically formulated using some priors with compressed videos, our conversion from a CSDR-to-HDR video reconstruction is conceptually divided into four synergistic parts: reducing compression artifacts, recovering missing details, adjusting imaging parameters, and reducing image noise. We approximate this process by a compact KIB. To capture richer details, we learn HDR representations with a set of KIBs connected in parallel and fused with the IFM. Extensive evaluations show that our KPNet achieves superior performance over the state-of-the-art methods. Mao Ye 0001, Xiatian Zhu, Shuai Li 0005, Xue Li 0001, Ce Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Structured residual sparsity for video compressive sensing reconstruction
Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiachao Zhang, Jiantao Zhou 0001, Ce Zhu |
Signal Process. | 6 |
| 2024 | TS-RTPM-Net: Data-Driven Tensor Sketching for Efficient CP DecompositionabstractTensor decomposition is widely used in feature extraction, data analysis, and other fields. As a means of tensor decomposition, the robust tensor power method based on tensor sketch (TS-RTPM) can quickly mine the potential features of tensor, but in some cases, its approximation performance is limited. In this paper, we propose a data-driven framework called TS-RTPM-Net, which improves the estimation accuracy of TS-RTPM by jointly training the TS value matrices with the RTPM initial matrices. It also uses two greedy initialization algorithms to optimize the TS location matrices. In addition, TS-RTPM-Net accelerates TS-RTPM by using fast power iteration modules. Comparative experiments on real-world datasets verify that TS-RTPM-Net outperforms TS-RTPM in terms of estimation accuracy, running speed, and memory consumption. Xingyu Cao, Xiangtao Zhang, Ce Zhu, Jiani Liu 0002, Yipeng Liu 0001 |
IEEE Trans. Big Data | 3 |
| 2024 | Enhanced Pseudo-Label Generation With Self-Supervised Training for Weakly- Supervised Semantic SegmentationabstractDue to the high cost of pixel-level labels required for fully-supervised semantic segmentation, weakly-supervised segmentation has emerged as a more viable option recently. Existing weakly-supervised methods tried to generate pseudo-labels without pixel-level labels for semantic segmentation, but a common problem is that the generated pseudo-labels contain insufficient semantic information, resulting in poor accuracy. To address this challenge, a novel method is proposed, which generates class activation/attention maps (CAMs) containing sufficient semantic information as pseudo-labels for the semantic segmentation training without pixel-level labels. In this method, the attention-transfer module is designed to preserve salient regions on CAMs while avoiding the suppression of inconspicuous regions of the targets, which results in the generation of pseudo-labels with sufficient semantic information. A pixel relevance focused-unfocused module has also been developed for better integrating contextual information, with both attention mechanisms employed to extract focused relevant pixels and multi-scale atrous convolution employed to expand receptive field for establishing distant pixel connections. The proposed method has been experimentally demonstrated to achieve competitive performance in weakly-supervised segmentation, and even outperforms many saliency-joined methods. Zhen Qin 0002, Guosong Zhu, Erqiang Zhou, Yingjie Zhou 0001, Yicong Zhou, Ce Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | Multiple Complementary Priors for Multispectral Image Compressive Sensing ReconstructionabstractCompressive sensing (CS) techniques using a few compressed measurements have drawn considerable interest in reconstructing multispectral imagery (MSI). Nonlocal-based tensor methods have been widely used for MSI-CS reconstruction, which employ the nonlocal self-similarity (NSS) property of MSI to obtain satisfactory results. However, such methods only consider the internal priors of MSI while ignoring important external image information, for example deep-driven priors learned from a corpus of natural image datasets. Meanwhile, they usually suffer from annoying ringing artifacts due to the aggregation of overlapping patches. In this article, we propose a novel approach for highly effective MSI-CS reconstruction using multiple complementary priors (MCPs). The proposed MCP jointly exploits nonlocal low-rank and deep image priors under a hybrid plug-and-play framework, which contains multiple pairs of complementary priors, namely, internal and external, shallow and deep, and NSS and local spatial priors. To make the optimization tractable, a well-known alternating direction method of multiplier (ADMM) algorithm based on the alternating minimization framework is developed to solve the proposed MCP-based MSI-CS reconstruction problem. Extensive experimental results demonstrate that the proposed MCP algorithm outperforms many state-of-the-art CS techniques in MSI reconstruction. The source code of the proposed MCP-based MSI-CS reconstruction algorithm is available at: https://github.com/zhazhiyuan/MCP_MSI_CS_Demo.git. Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiachao Zhang, Jiantao Zhou 0001, Xudong Jiang 0001, Ce Zhu |
IEEE Trans. Cybern. | 7 |
| 2024 | CS2DIPs: Unsupervised HSI Super-Resolution Using Coupled Spatial and Spectral DIPsabstractIn recent years, fusing high spatial resolution multispectral images (HR-MSIs) and low spatial resolution hyperspectral images (LR-HSIs) has become a widely used approach for hyperspectral image super-resolution (HSI-SR). Various unsupervised HSI-SR methods based on deep image prior (DIP) have gained wide popularity thanks to no pre-training requirement. However, DIP-based methods often demonstrate mediocre performance in extracting latent information from the data. To resolve this performance deficiency, we propose a coupled spatial and spectral deep image priors (CS2DIPs) method for the fusion of an HR-MSI and an LR-HSI into an HR-HSI. Specifically, we integrate the nonnegative matrix-vector tensor factorization (NMVTF) into the DIP framework to jointly learn the abundance tensor and spectral feature matrix. The two coupled DIPs are designed to capture essential spatial and spectral features in parallel from the observed HR-MSI and LR-HSI, respectively, which are then used to guide the generation of the abundance tensor and spectral signature matrix for the fusion of the HSI-SR by mode-3 tensor product, meanwhile taking some inherent physical constraints into account. Free from any training data, the proposed CS2DIPs can effectively capture rich spatial and spectral information. As a result, it exhibits much superior performance and convergence speed over most existing DIP-based methods. Extensive experiments are provided to demonstrate its state-of-the-art overall performance including comparison with benchmark peer methods. Yipeng Liu 0001, Chong-Yung Chi, Zhen Long, Ce Zhu |
IEEE Trans. Image Process. | 5 |
| 2024 | Adaptively Topological Tensor Network for Multi-View Subspace ClusteringabstractMulti-view subspace clustering employs learned self-representation from multiple tensor decompositions to exploit the low-rank information. However, the data structures embedded with self-representation tensors may vary in different multi-view datasets. Therefore, a pre-defined decomposition may not fully exploit low-rank information from various data, resulting in sub-optimal multi-view clustering performance. To alleviate this, we proposed the adaptively topological tensor network (ATTN). ATTN can learn a suitable decomposition structure that can represent the low-rank structure and high-order correlation of the self-representation tensors better in a data-driven way, which can capture the intra-view and inter-view information better. Firstly, instead of connecting the tensor network blindly, ATTN utilizes the correlation between adjacent factors to prune redundant connections from the fully connected tensor networks, making the tensor network more expressive. Furthermore, a greedy adaptive rank-increasing strategy is applied to optimize the pruned tensor network structure, which improves the capacity of capturing low-rank structure. We apply ATTN on a multi-view subspace clustering task and utilize the alternating direction method of multipliers(ADMM) method to optimize it. Experiments show that multi-view subspace clustering based on ATTN has better performance on nine multi-view datasets. Yipeng Liu 0001, Jie Chen 0086, Yingcong Lu, Weiting Ou, Zhen Long, Ce Zhu |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2024 | Feature Space Recovery for Efficient Incomplete Multi-View ClusteringabstractT-SVD based incomplete multi-view clustering (IMVC) has received wide attention due to its ability to capture high-order correlations. However, t-SVD suffers from rotation sensitivity, failing to fully explore both inter- and intra-view consistencies. Besides, current methods mainly consider inter- or intra-view correlations, ignoring the low-rank information of sample features within views. To address these weaknesses, we first propose a feature space recovery based IMVC (FSR-IMVC) method, where low-rank feature space recovery and low-rank tensor ring based consistency learning are considered into a unified framework. Furthermore, we extend FSR-IMVC by incorporating anchor learning on the latent feature space, resulting in a scalable FSR-IMVC (sFSR-IMVC) approach that is well-suited to large-scale data. In an iterative way, the learned inter- and intra-view correlations will guide the recovery of missing features, while the explored low-rank information from feature spaces will in turn facilitate consistency exploration, eventually achieving outstanding clustering performance. Experimental results show that FSR-IMVC provides a significant improvement over known state-of-the-art algorithms in terms of ACC, NMI and Purity. Compared with FSR-IMVC, sFSR-IMVC performs slightly worse in clustering accuracy, but offers a notable advantage in computational efficiency, particularly for large-scale datasets. The codes of FSR-IMVC and sFSR-IMVC are publicly available athttps://github.com/longzhen520/sFSR-IMVC. Zhen Long, Ce Zhu, Pierre Comon, Yazhou Ren 0001, Yipeng Liu 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2024 | Cross-Domain Low-Dose CT Image Denoising With Semantic Preservation and Noise AlignmentabstractDeep learning (DL)-based Low-dose CT (LDCT) image denoising methods may face domain shift problem, where data from different domains (i.e., hospitals) may have similar anatomical regions but exhibit different intrinsic noise characteristics. Therefore, we propose a plug-and-play model called Lowand High-frequency Alignment (LHFA) to address this issue by leveraging semantic features and aligning noise distributions of different CT datasets, while maintaining diagnostic image quality and suppressing noise. Specifically, the LHFA model consists of a Low-frequency Alignment (LFA) module that preserves semantic features (i.e., low-frequency components) with fewer perturbations from both domains for reconstruction. Notably, a Highfrequency Alignment (HFA) module is proposed to quantify the discrepancy between noise representations (i.e., high-frequency components) in a latent space mapped by an auto-encoder. Experimental results demonstrate that the LHFA model effectively alleviates the domain shift problem and significantly improves the performance of DL-based methods on cross-domain LDCT image denoising task, outperforming other domain adaptationbased methods. Jiaxin Huang 0006, Kecheng Chen, Yazhou Ren 0001, Xiaorong Pu, Ce Zhu |
IEEE Trans. Multim. | 6 |
| 2024 | Multi-View MERA Subspace ClusteringabstractTensor-based multi-view subspace clustering (MSC) can capture high-order correlation in the self-representation tensor. Current tensor decompositions for MSC suffer from highly unbalanced unfolding matrices or rotation sensitivity, failing to fully explore inter/intra-view information. Using the advanced tensor network, namely, multi-scale entanglement renormalization ansatz (MERA), we propose a low-rank MERA based MSC (MERA-MSC) algorithm, where MERA factorizes a tensor into contractions of one top core factor and the rest orthogonal/semi-orthogonal factors. Benefiting from multiple interactions among orthogonal/semi-orthogonal (low-rank) factors, the low-rank MERA has a strong representation power to capture the complex inter/intra-view information in the self-representation tensor. The alternating direction method of multipliers is adopted to solve the optimization model. Experimental results on five multi-view datasets demonstrate MERA-MSC has superiority against the compared algorithms on six evaluation metrics. Furthermore, we extend MERA-MSC by incorporating anchor learning and develop a scalable low-rank MERA based multi-view clustering method (sMREA-MVC). To our knowledge, this is the first work to introduce MERA to the multi-view clustering topic. The effectiveness and efficiency of sMERA-MVC have been validated on three large-scale multi-view datasets. Zhen Long, Ce Zhu, Jie Chen 0086, Yazhou Ren 0001, Yipeng Liu 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Progressive Graph Reasoning-Based Social Relation RecognitionabstractIdentifying relationships between people from images is essential for studying social activities and interactions, and this has significant potential to further the understanding of human social behaviors. Existing image-based research mainly explores social relationships at the dyadic level, i.e., recognizing pairwise relationships based on visual features of persons, objects, and scenes and their logical constraints. Notably, social relational structures are hierarchically nested, i.e., individuals and dyads are nested within group structures, as indicated in the social relations model (SRM) of social psychology. However, existing computer vision-based studies fail to consider hierarchical nested structures, thus overlooking some of the most important interactions, which leads to poor relation reasoning. To improve the performance of reasoning neural networks, we propose a novel SRM framework for progressive graph reasoning (PGR) to explore social interactions. Specifically, we construct individual–dyad and dyad–group graphs to progressively explore the impact of individuals and groups on recognition of dyadic relationships. A transformer is utilized to fuse visual features and graph reasoning knowledge into a comprehensive representation of social relationships. We demonstrate the effectiveness of the proposed model based on PGR using several public datasets and perform extensive ablation studies to explore the reasons behind its superior performance. Experimental results demonstrate that our proposed model successfully predicts social relationships with higher accuracy than state-of-the-art methods. Codes and datasets are available at:https://github.com/tw-repository/PGRSRR. Wang Tang, Linbo Qing, Lindong Li, Ce Zhu |
IEEE Trans. Multim. | 5 |
| 2023 | Level-Line Guided Edge Drawing for Robust Line Segment DetectionabstractLine segment detection plays a cornerstone role in computer vision tasks. Among numerous detection methods that have been recently proposed, the ones based on edge drawing attract increasing attention owing to their excellent detection efficiency. However, the existing methods are not robust enough due to the inadequate usage of image gradients for edge drawing and line segment fitting. Based on the observation that the line segments should locate on the edge points with both consistent coordinates and level-line information, i.e., the unit vector perpendicular to the gradient orientation, this paper proposes a level-line guided edge drawing for robust line segment detection (GEDRLSD). The level-line information provides potential directions for edge tracking, which could be served as a guideline for accurate edge drawing. Additionally, the level-line information is fused in line segment fitting to improve the robustness. Numerical experiments show the superiority of the proposed GEDRLSD1algorithm compared with state-of-the-art methods. Yingjie Zhou 0001, Yipeng Liu 0001, Ce Zhu |
ICASSP | 4 |
| 2023 | Efficient and Effective Multi-Camera Pose Estimation with Weighted M-Estimate Sample ConsensusabstractCamera pose estimation is a fundamental module for many vision tasks. It is usually based on feature correspondences, i.e., feature matches across different images. However, correspondences always contain non-negligible outliers, which may negatively affect pose estimation efficiency and accuracy. This paper proposes a multi-camera pose estimation method by leveraging point and line correspondences with non-negligible outliers, in which a weighted M-Estimate Sample Consensus (w-MSAC) based on the customized weights and the coarse pose prior is introduced to improve the efficiency and accuracy of pose estimation. The customized weights could decrease the iterations of the pose hypothesis and improve the pose estimation accuracy. The coarse pose prior is used to perform the pre-validation of the pose hypothesis, eliminating many unnecessary validations. Experiments demonstrate the superiority of the proposed w-MSAC1over existing state-of-the-art methods, e.g., improving 22% positioning and 24% orientation accuracy meanwhile decreasing 15% iterations and 92% validations than the MSAC. Yingjie Zhou 0001, Xun Zhang 0002, Yipeng Liu 0001, Ce Zhu |
ICASSP | 5 |
| 2023 | Tensorized LSSVMS For Multitask RegressionabstractMultitask learning (MTL) can utilize the relatedness between multiple tasks for performance improvement. The advent of multimodal data allows tasks to be referenced by multiple indices. High-order tensors are capable of providing efficient representations for such tasks, while preserving structural task-relations. In this paper, a new MTL method is proposed by leveraging low-rank tensor analysis and constructing tensorized Least Squares Support Vector Machines, namely the tLSSVM-MTL, where multilinear modelling and its nonlinear extensions can be flexibly exerted. We employ a high-order tensor for all the weights with each mode relating to an index and factorize it with CP decomposition, assigning a shared factor for all tasks and retaining task-specific latent factors along each index. Then an alternating algorithm is derived for the nonconvex optimization, where each resulting subproblem is solved by a linear system. Experimental results demonstrate promising performances of our tLSSVM-MTL. Jiani Liu 0002, Qinghua Tao, Ce Zhu, Yipeng Liu 0001, Johan A. K. Suykens |
ICASSP | 3 |
| 2023 | Feature Space Recovery for Incomplete Multi-View ClusteringabstractIncomplete multi-view clustering (IMVC), based on imputation and clustering unification, has received wide attention due to its ability to exploit hidden information from missing views. However, current methods mainly consider inter/intra-view correlations, ignoring the structural information of sample features within views. In this paper, we propose a feature space recovery based IMVC method, where low-rank feature space recovery and consensus representation learning of inter/intra-views are considered into a unified framework. Moreover, low-rank tensor ring approximation is used to capture the correlations of self-representation tensor. In an iterative way, the learned inter/intra-view correlations will guide the recovery of missing features, while the explored low-rank information from feature spaces will in turn facilitate self-representation learning, eventually achieving out-standing clustering performance. Experimental results show our method has a very significant improvement over known state-of-the-art algorithms in terms of ACC, NMI and Purity. Zhen Long, Ce Zhu, Pierre Comon, Yipeng Liu 0001 |
ICASSP | 2 |
| 2023 | A Novel Mode Selection-Based Fast Intra Prediction Algorithm for Spatial SHVCabstractDue to multi-layer encoding and Inter-layer prediction, Spatial Scalable High-Efficiency Video Coding (SSHVC) has extremely high coding complexity. It is very crucial to improve its coding speed so as to promote widespread and cost-effective SSHVC applications. In this paper, we have proposed a novel Mode Selection-Based Fast Intra Prediction algorithm for SSHVC. We reveal the RD costs of Inter-layer Reference (ILR) mode and Intra mode have a significant difference, and the RD costs of these two modes follow Gaussian distribution. Based on this observation, we propose to apply the classic Gaussian Mixture Model and Expectation Maximization in machine learning to determine whether ILR is the best mode so as to skip the Intra mode. Experimental results demonstrate that the proposed algorithm can significantly improve the coding speed with negligible coding efficiency loss. Yu Sun 0003, Weisheng Li 0001, Lele Xie, Xin Lu 0001, Frédéric Dufaux, Ce Zhu |
ICASSP | 7 |
| 2023 | Hyperspectral Image Denoising Via Nonlocal Rank Residual ModelingabstractNonlocal low-rank (LR) tensor modeling has shown great potential in hyperspectral image (HSI) denoising, which first uses the nonlocal self-similarity (NSS) prior to search for many similar full-band patches to form three-dimensional nonlocal full-band groups (tensors), and then usually enforces an LR penalty on each nonlocal full-band group. However, in most existing methods, the LR tensor is only approximated directly from the degraded nonlocal full-band tensor, which is subject to certain issues (e.g., in heavy noise environments) in obtaining a suboptimal tensor approximation, and thus leading to unsatisfactory denoising results. In this paper, we propose a novel nonlocal rank residual (NRR) approach for highly effective HSI denoising, which progressively approximates the underlying L-R tensor via minimizing the rank residual. Towards this end, we first obtain a good estimate of the original nonlocal full-band group by using the NSS prior, and then the rank residual between the de-graded nonlocal full-band group with the corresponding estimated nonlocal full-band group is minimized to achieve a more accurate LR tensor. Moreover, the global spectral LR prior is employed to reduce the spectral redundancy of HSI in the proposed denoising framework. Finally, we develop a simple yet effective alternating minimization algorithm to jointly refine global spectral information and nonlocal full-band groups. Experimental results clearly show that the proposed NRR algorithm outperforms many state-of-the-art HSI denoising methods. The source code of the proposed NRR algorithm for HSI denoising is available at: https://github.com/zhazhiyuan/NRR_HSI_Denoising_Demo.git. Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu |
ICASSP | 5 |
| 2023 | Feature Modulation Transformer: Cross-Refinement of Global Representation via High-Frequency Prior for Image Super-ResolutionabstractTransformer-based methods have exhibited remarkable potential in single image super-resolution (SISR) by effectively extracting long-range dependencies. However, most of the current research in this area has prioritized the design of transformer blocks to capture global information, while overlooking the importance of incorporating high-frequency priors, which we believe could be beneficial. In our study, we conducted a series of experiments and found that transformer structures are more adept at capturing low-frequency information, but have limited capacity in constructing high-frequency representations when compared to their convolutional counterparts. Our proposed solution, the cross-refinement adaptive feature modulation transformer (CRAFT), integrates the strengths of both convolutional and transformer structures. It comprises three key components: the high-frequency enhancement residual block (HFERB) for extracting high-frequency information, the shift rectangle window attention block (SRWAB) for capturing global information, and the hybrid fusion block (HFB) for refining the global representation. Our experiments on multiple datasets demonstrate that CRAFT outperforms state-of-the-art methods by up to 0.29dB while using fewer parameters. The source code will be made available at: https://github.com/AVC2-UESTC/CRAFT-SR.git. Ao Li 0007, Le Zhang 0001, Yun Liu 0011, Ce Zhu |
ICCV | 4 |
| 2023 | Fast Learning-Based Split Type Prediction Algorithm for VVCabstractAs the latest video coding standard, Versatile Video Coding (VVC) is highly efficient at the cost of very high coding complexity, which seriously hinders its widespread application. Therefore, it is very crucial to improve its coding speed. In this paper, we propose a learning-based fast split type (ST) prediction algorithm for VVC using a deep learning approach. We first construct a large-scale database containing sufficient STs with diverse video resolution and content. Next, since the ST distributions of coding units (CUs) of different sizes are significantly distinct, so we separately design neural networks for all different CU sizes. Then, we merge ambiguous STs into four merged classes (MCs) to train models to obtain probabilities of MCs and skip unlikely ones. Experimental results demonstrate that the proposed algorithm can reduce the encoding time of VVC by 67.53% with 1.89% increase in Bjøntegaard delta bit-rate (BDBR) on average. Liulin Chen, Xin Lu 0001, Frédéric Dufaux, Weisheng Li 0001, Ce Zhu |
ICIP | 6 |
| 2023 | A Probability-Based All-Zero Block Early Termination Algorithm for QSHVCabstractTo seamlessly adapt to time-varying network bandwidths, Quality Scalable High-Efficiency Video Coding (QSHVC) is developed. However, its coding process is overwhelmingly complex, and this seriously limits its wide applications in realtime environments. Therefore, it is of great significance to study fast coding algorithms for QSHVC. In this paper, we propose a novel probability-based All-Zero Block (AZB) early termination algorithm for QSHVC. We observe that the generated residual coefficients follow the Laplace distribution if a CU is accurately predicted. Based on this observation, we derive the sum of squared differences-based AZB decision condition. Second, the probability of each coding mode and coding depth being chosen as the best ones are combined with AZBs to derive the probability-based early termination condition. The experimental results show that the proposed algorithm can improve the average coding speed by 74.95% with a 0.26% increase in BDBR. Xin Lu 0001, Frédéric Dufaux, Qianmin Wang, Weisheng Li 0001, Bo Hang, Ce Zhu |
ICIP | 7 |
| 2023 | Nonlocal Low-Rank Residual Modeling for Image Compressive Sensing ReconstructionabstractThe nonlocal low-rank (LR) modeling has proven to be an effective approach in image compressive sensing (CS) reconstruction, which starts by clustering similar patches using the nonlocal self-similarity (NSS) prior into nonlocal image groups and then imposes an L-R penalty on each nonlocal image group. However, most existing methods only approximate the LR matrix directly from the degraded nonlocal image group, which may lead to suboptimal LR matrix approximation and thus obtain unsatisfactory reconstruction results. This paper proposes a novel nonlocal low-rank residual (NLRR) approach for image CS reconstruction, which progressively approximates the underlying LR matrix by minimizing the LR residual. To do this, we first use the NSS prior to obtain a good estimate of the original nonlocal image group, and then the LR residual between the degraded nonlocal image group and the estimated nonlocal image group is minimized to derive a more accurate LR matrix. To ensure the optimization is both feasible and reliable, we employ an alternative direction multiplier method (ADMM) to solve the NLRR-based image CS reconstruction problem. Our experimental results show that the proposed NLRR algorithm achieves superior performance against many popular or state-of-the-art image CS reconstruction methods, both in objective metrics and subjective perceptual quality. Junhao Zhang 0005, Kim-Hui Yap, Lap-Pui Chau, Ce Zhu |
ICIP | 4 |
| 2023 | Federated Deep Multi-View Clustering with Global Self-SupervisionabstractFederated multi-view clustering has the potential to learn a global clustering model from data distributed across multiple devices. In this setting, label information is unknown and data privacy must be preserved, leading to two major challenges. First, views on different clients often have feature heterogeneity, and mining their complementary cluster information is not trivial. Second, the storage and usage of data from multiple clients in a distributed environment can lead to incompleteness of multi-view data. To address these challenges, we propose a novel federated deep multi-view clustering method that can mine complementary cluster structures from multiple clients, while dealing with data incompleteness and privacy concerns. Specifically, in the server environment, we propose sample alignment and data extension techniques to explore the complementary cluster structures of multiple views. The server then distributes global prototypes and global pseudo-labels to each client as global self-supervised information. In the client environment, multiple clients use the global self-supervised information and deep autoencoders to learn view-specific cluster assignments and embedded features, which are then uploaded to the server for refining the global self-supervised information. Finally, the results of our extensive experiments demonstrate that our proposed method exhibits superior performance in addressing the challenges of incomplete multi-view data in distributed environments. Xinyue Chen 0004, Jie Xu 0044, Yazhou Ren 0001, Xiaorong Pu, Ce Zhu, Xiaofeng Zhu 0001, Zhifeng Hao 0005, Lifang He 0001 |
ACM Multimedia | 5 |
| 2023 | Optimal Low-Rank Tensor Tree CompletionabstractTensor completion is a powerful technique for recovering missing entries from partial observations. Tensor tree network, with a hierarchical structure, has gained widespread attention for its ability to balance and effectively explore the correlations in high-order data. However, the performance of tensor trees is influenced by the order of modes, leading to variations in their effectiveness. To address this, we propose to minimize the loss of entanglement entropy to determine the optimal mode order within the tensor tree network, thereby optimizing its representation performance. We correspondingly construct an optimal low-rank tensor tree completion model, where the optimal low-rank tensor tree network captures global structures and total variation investigates local structures. The alternating direction method of multipliers is employed to solve the optimization problem. Experimental results on color images and light field images demonstrate that our method outperforms state-of-the-art algorithms in terms of recovery performance. Ce Zhu, Zhen Long, Yipeng Liu 0001 |
MMSP | 2 |
| 2023 | Predicting the Invariance Behind Residuals: A Novel GAN Inversion Method for Image Editing and Detail RetainingabstractGenerative adversarial network (GAN) inversion has been serving as the vehicle to enable the restoration of real-world images by GANs, rather than the realistic generation from random noise. Existing GAN inversion methods, however, suffer from the fidelity-editability trade-off, which mainly invert the semantics within images and fail to reconstruct the details. To address this issue, we propose a novel adaptive detail compensation method for GAN inversion (ADC-GInv), which automatically locates and restores the non-semantic details, whilst maintaining the editability on the semantics of real-world images. More specifically, we first develop an adversarial reciprocal learning GAN (ARL-GAN) so as to seamlessly optimise the reconstruction during the training of GANs, followed by a sophisticated fine-tuning technique for ARL-GAN inversion. This ensures superior restoration and editability on the semantic cues of images. Then, regarding the non-semantic details, our ADC-GInv method adaptively locates the details by predicting the invariance given edited and non-edited residuals of restoration, which are then compensated at the pixel-level for high fidelity. As a consequence, the experimental results have verified the superior performance of our ADC-GInv, on both fidelity and editability during inversion. Zhimo Yan, Hengyang He, Shengxi Li, Mai Xu, Ce Zhu |
MMSP | 7 |
| 2023 | Efficient Lightweight Attention Based Learned Image CompressionabstractThe CNN-based end-to-end learned image compression methods have already achieved a significant improvement in terms of coding efficiency. Moreover, with the capability of modeling long-range global correlation, the transformer-based image compression has further elevated the coding efficiency to outperform the latest Versatile Video Coding (VVC) standard. Nonetheless, the high computational burden of the self-attention mechanism in the transformer design present a significant obstacle for practical applications. To address this concern, we propose a relatively low-complexity end-to-end learned image compression approach by integrating an efficient lightweight attention module, which effectively mitigates the computational overhead associated with self-attention in the transformer design. The experimental results demonstrate that the proposed method achieves better RD performance than VVC. Furthermore, as compared to the prevailing state-of-the-art transformer-based approach, our method accelerates the coding speed by over 6 times while maintaining comparable coding efficiency. Lei Luo 0003, Le Zhang 0001, Hongwei Guo 0001, Ce Zhu |
VCIP | 5 |
| 2023 | Single-image HDR reconstruction by dual learning the camera imaging process
Lei She, Mao Ye 0001, Shuai Li 0005, Ce Zhu |
Eng. Appl. Artif. Intell. | 5 |
| 2023 | SCN: Self-Calibration Network for fast and accurate image super-resolution
Haoran Yang 0008, Xiaomin Yang, Kai Liu 0012, Gwanggil Jeon, Ce Zhu |
Expert Syst. Appl. | 5 |
| 2023 | A dedicated benchmark for contour-based corner detection evaluation
Yingjie Zhou 0001, Yipeng Liu 0001, Ce Zhu |
Image Vis. Comput. | 4 |
| 2023 | DifFormer: Multi-Resolutional Differencing Transformer With Dynamic Ranging for Time Series AnalysisabstractTime series analysis is essential to many far-reaching applications of data science and statistics including economic and financial forecasting, surveillance, and automated business processing. Though being greatly successful of Transformer in computer vision and natural language processing, the potential of employing it as the general backbone in analyzing the ubiquitous times series data has not been fully released yet. Prior Transformer variants on time series highly rely on task-dependent designs and pre-assumed "pattern biases", revealing its insufficiency in representing nuanced seasonal, cyclic, and outlier patterns which are highly prevalent in time series. As a consequence, they can not generalize well to different time series analysis tasks. To tackle the challenges, we propose DifFormer, an effective and efficient Transformer architecture that can serve as a workhorse for a variety of time-series analysis tasks. DifFormer incorporates a novel multi-resolutional differencing mechanism, which is able to progressively and adaptively make nuanced yet meaningful changes prominent, meanwhile, the periodic or cyclic patterns can be dynamically captured with flexible lagging and dynamic ranging operations. Extensive experiments demonstrate DifFormer significantly outperforms state-of-the-art models on three essential time-series analysis tasks, including classification, regression, and forecasting. In addition to its superior performances, DifFormer also excels in efficiency - a linear time/memory complexity with empirically lower time consumption. Bing Li 0002, Wei Cui 0002, Le Zhang 0001, Ce Zhu, Wei Wang 0011, Ivor W. Tsang, Joey Tianyi Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Pre-encoding based temporal dependent rate-distortion optimization for HEVC
Hongwei Guo 0001, Ce Zhu, Mao Ye 0001, Lei Luo 0003, Xu Yang 0030 |
Signal Process. Image Commun. | 2 |
| 2023 | Level Line Guided Interest Point DetectionabstractDetection of interest points,e.g., corners and blobs, lays the foundation for many vision tasks. Numerous methods have been proposed to improve the detection performance, and the gradient-based ones are the most investigated. However, existing gradient-based methods lack an adequate utilization of gradient orientations. In this letter, we show that the level line,i.e., the unit vector orthogonal to the gradient orientation of a specific point, is particularly important to interest point detection. The support level lines of an interest point,i.e., the level lines used to identify corners/blobs, exhibit a significantly different pattern from those of other points. Based on this observation, this letter proposes two robust interest point detectors for finding corners and blobs, respectively. For each detector, a specific type of level line difference is defined, and the corresponding differences are leveraged with different weights. Numerical experiments show the superior performance of the proposed detectors. The code will be publicly available athttps://github.com/roylin1229/LLD-IP. Yingjie Zhou 0001, Yipeng Liu 0001, Ce Zhu |
IEEE Signal Process. Lett. | 4 |
| 2023 | Nonlocal Structured Sparsity Regularization Modeling for Hyperspectral Image DenoisingabstractThe non-local-based model for hyperspectral image (HSI) denoising first uses non-local self-similarity (NSS) prior to group similar full-band patches into three-dimensional non-local full-band groups (tensors) using a block matching (BM) operation, and then a low-rank (LR) penalty is typically applied to each non-local full-band group to reduce noise. While non-local-based methods have shown promising performance in HSI denoising, most existing methods have only considered the LR property of the non-local full-band group while ignoring the strong correlation between sparse coefficients. Moreover, such methods often result in unsatisfactory visual artifacts due to the noise sensitivity of BM operations, while requiring expensive computations. To address these limitations, this paper proposes a novel non-local structured sparsity regularization (NLSSR) approach for HSI denoising. First, to mitigate the noise sensitivity of the BM operation, we propose a graph-based domain distance scheme to index similar full-band patches to form the non-local full-band group. Second, we design an adaptive unidirectional low-rank (LR) dictionary with low complexity that takes into account the differences in intrinsic structure correlation among different modes of the non-local full-band tensor. Third, we utilize a global spectral LR prior to reduce spectral redundancy. Fourth, we develop a generalized soft-thresholding (GST) algorithm based on the alternating minimization framework to solve the NLSSR-based HSI denoising problem. We perform extensive experiments on both simulated and real data to show that the proposed NLSSR algorithm outperforms many popular or state-of-the-art HSI denoising methods in both quantitative and visual evaluations. Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiachao Zhang, Jiantao Zhou 0001, Yilong Lu, Ce Zhu |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2023 | Multiplex Transformed Tensor Decomposition for Multidimensional Image RecoveryabstractLow-rank tensor completion aims to recover the missing entries of multi-way data, which has become popular and vital in many fields such as signal processing and computer vision. It varies with different tensor decomposition frameworks. Compared with matrix SVD, recently emerging transform t-SVD can better characterize the low-rank structure of order-3 data. However, it suffers from rotation sensitivity, and dimensional limitation (i.e., only effective for order-3 tensors). To alleviate these deficiencies, we develop a novel multiplex transformed tensor decomposition (MTTD) framework, which can characterize the global low-rank structure along all modes for any order- N tensor. Based on MTTD, we propose a related multi-dimensional square model for low-rank tensor completion. Besides, a total variation term is also introduced to utilize the local piecewise smoothness of the tensor data. The classic alternating direction method of multipliers is used to solve the convex optimization problems. For performance testing, we choose three linear invertible transforms including FFT, DCT, and a group of unitary transform matrices for our proposed methods. The simulated and real-data experiments demonstrate the superior recovery accuracy and computational efficiency of our method compared with state-of-the-art ones. Lanlan Feng, Ce Zhu, Zhen Long, Jiani Liu 0002, Yipeng Liu 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | RGB-Guided Depth Map Recovery by Two-Stage Coarse-to-Fine Dense CRF ModelsabstractDepth maps generally suffer from large erroneous areas even in public RGB-Depth datasets. Existing learning-based depth recovery methods are limited by insufficient high-quality datasets and optimization-based methods generally depend on local contexts not to effectively correct large erroneous areas. This paper develops an RGB-guided depth map recovery method based on the fully connected conditional random field (dense CRF) model to jointly utilize local and global contexts of depth maps and RGB images. A high-quality depth map is inferred by maximizing its probability conditioned upon a low-quality depth map and a reference RGB image based on the dense CRF model. The optimization function is composed of redesigned unary and pairwise components, which constraint local structure and global structure of depth map, respectively, with the guidance of RGB image. In addition, the texture-copy artifacts problem is handled by two-stage dense CRF models in a coarse-to-fine way. A coarse depth map is first recovered by embedding RGB image in a dense CRF model in unit of $3\times 3$ blocks. It is refined afterward by embedding RGB image in another model in unit of individual pixels and restricting the model mainly work in discontinued regions. Extensive experiments on six datasets verify that the proposed method considerably outperforms a dozen of baseline methods in correcting erroneous areas and diminishing texture-copy artifacts of depth maps. Haotian Wang 0009, Meng Yang 0002, Ce Zhu, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 3 |
| 2023 | Spatio-Temporal Detail Information Retrieval for Compressed Video Quality EnhancementabstractThe past few years have witnessed the great success of multi-frame quality enhancement for compressed video. Although the existing methods based on deformable alignment have achieved the state-of-the-art performance, they do not pay enough attention to the recovery of detail information. In this work, we propose a Spatio-Temporal Detail Retrieval (STDR) method to promote the recovery of detail information. To alleviate the problem of inaccurate deformable offsets caused by the fixed receptive field, motivated by multi-task learning, we design a plug-and-play Multi-path Deformable Alignment (MDA) module to generate more accurate offsets by integrating the alignment features of different receptive fields, so that the temporal detail information can be better recovered. For the spatial detail information restoration, several residual dense blocks with channel attention layer are utilized in the reconstruction module to explore valuable high-frequency spatial information from the fused multi-path alignment features. Meanwhile, a complementary loss function based on the Pearson correlation coefficient is developed to ameliorate the over-smoothing shortcoming caused by pixel-wise mean square or absolute value loss. Experimental results demonstrate that the proposed STDR network achieves superior performance compared with the state-of-the-art methods in both quantitative and qualitative evaluations. Dengyan Luo, Mao Ye 0001, Shuai Li 0005, Ce Zhu, Xue Li 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Low-Rankness Guided Group Sparse Representation for Image RestorationabstractAs a spotlighted nonlocal image representation model, group sparse representation (GSR) has demonstrated a great potential in diverse image restoration tasks. Most of the existing GSR-based image restoration approaches exploit the nonlocal self-similarity (NSS) prior by clustering similar patches into groups and imposing sparsity to each group coefficient, which can effectively preserve image texture information. However, these methods have imposed only plain sparsity over each individual patch of the group, while neglecting other beneficial image properties, e.g., low-rankness (LR), leads to degraded image restoration results. In this article, we propose a novel low-rankness guided group sparse representation (LGSR) model for highly effective image restoration applications. The proposed LGSR jointly utilizes the sparsity and LR priors of each group of similar patches under a unified framework. The two priors serve as the complementary priors in LGSR for effectively preserving the texture and structure information of natural images. Moreover, we apply an alternating minimization algorithm with an adaptively adjusted parameter scheme to solve the proposed LGSR-based image restoration problem. Extensive experiments are conducted to demonstrate that the proposed LGSR achieves superior results compared with many popular or state-of-the-art algorithms in various image restoration tasks, including denoising, inpainting, and compressive sensing (CS). Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu, Alex Chichung Kot |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Simultaneous Nonlocal Low-Rank And Deep Priors For Poisson DenoisingabstractPoisson noise is a common electronic noise, which has widely occurred in various photo-limited imaging systems. However, due to signal-dependent and multiplicative characteristics for Poisson noise, Poisson denoising is still an open problem. In this paper, we propose a novel approach using simultaneous nonlocal low-rank and deep priors (SNLDP) for Poisson denoising. The proposed SNLD-P simultaneously employs nonlocal self-similarity and deep image priors under the hybrid plug and play framework, which comprises multiple pairs of complementary priors, namely, nonlocal and local, shallow and deep, and internal and external. To make the optimization tractable, an effective alternating direction method of multiplier (ADMM) algorithm under the alternative minimization framework is provided to solve the proposed SNLDP-based Poisson denoising problem. Experimental results demonstrate the superiority of the proposed SNLDP over many popular or state-of-the-art Poisson denoising algorithms in terms of quantitative and visual perception. Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu |
ICASSP | 5 |
| 2022 | Light Field Integral Image Coding Optimization under 2D Hierarchical Coding StructureabstractIn integral photography, one important acquisition method of light field, a 3D scene is captured from different perspectives and the data with repetitive views is contained in one integral image. To encode such integral images, 2D hierarchical coding structure (2D-HCS) is used to exploit the dependency among images of different views. However, the fixed 2D-HCS cannot fully explore the inter-view dependency, making the coding not optimal. Therefore, in this paper, an inter-view dependent rate-distortion optimization (RDO) method is proposed for light field integral image coding under 2D-HCS. Experimental results demonstrate that, compared to 2D-HCS, the proposed method can achieve averaged {Y, U, V} BD-rate savings of {13.2%, 9.2%, 12.0%} only by adapting the Lagrange Multiplier and up to {21.6%, 18.4%, 21.8%} together with QP adaptation. Yanbo Gao, Ce Zhu |
ICIP | 3 |
| 2022 | KUNet: Imaging Knowledge-Inspired Single HDR Image ReconstructionabstractRecently, with the rise of high dynamic range (HDR) display devices, there is a great demand to transfer traditional low dynamic range (LDR) images into HDR versions. The key to success is how to solve the many-to-many mapping problem. However, the existing approaches either do not consider constraining solution space or just simply imitate the inverse camera imaging pipeline in stages, without directly formulating the HDR image generation process. In this work, we address this problem by integrating LDR-to-HDR imaging knowledge into an UNet architecture, dubbed as Knowledge-inspired UNet (KUNet). The conversion from LDR-to-HDR image is mathematically formulated, and can be conceptually divided into recovering missing details, adjusting imaging parameters and reducing imaging noise. Accordingly, we develop a basic knowledge-inspired block (KIB) including three subnetworks corresponding to the three procedures in this HDR imaging process. The KIB blocks are cascaded in the similar way to the UNet to construct HDR image with rich global information. In addition, we also propose a knowledge inspired jump-connect structure to fit a dynamic range gap between HDR and LDR images. Experimental results demonstrate that the proposed KUNet achieves superior performance compared with the state-of-the-art methods. The code, dataset and appendix materials are available at https://github.com/wanghu178/KUNet.git. Mao Ye 0001, Xiatian Zhu, Shuai Li 0005, Ce Zhu, Xue Li 0001 |
IJCAI | 5 |
| 2022 | Long-term Visual Localization Using Illumination Insensitive DescriptorsabstractThis demo shows a long-term visual localization system based on illumination insensitive descriptors (IID) of points and lines in multiple cameras. The system can robustly match the features in captured images for localization against those in localization database (DB). The developed localization system achieves remarkable performance and seasonal-time-spanned localization results in complex and changing environments. Yingjie Zhou 0001, Yipeng Liu 0001, Ce Zhu |
MMSP | 4 |
| 2022 | Dysphagia diagnosis system with integrated speech analysis from throat vibration
Hengling Zhao, Yangyang Jiang, Shenghan Wang, Fangzhou Ren, Ce Zhu, Jirong Yue, Yipeng Liu 0001 |
Expert Syst. Appl. | 8 |
| 2022 | Deep anomaly detection in packet payload
Xucheng Song, Yingjie Zhou 0001, Yanru Zhang, Dapeng Oliver Wu, Ce Zhu |
Neurocomputing | 8 |
| 2022 | Siamese Network for RGB-D Salient Object Detection and BeyondabstractExisting RGB-D salient object detection (SOD) models usually treat RGB and depth as independent information and design separate networks for feature extraction from each. Such schemes can easily be constrained by a limited amount of training data or over-reliance on an elaborately designed training process. Inspired by the observation that RGB and depth modalities actually present certain commonality in distinguishing salient objects, a novel joint learning and densely cooperative fusion (JL-DCF) architecture is designed to learn from both RGB and depth inputs through a shared network backbone, known as the Siamese architecture. In this paper, we propose two effective components: joint learning (JL), and densely cooperative fusion (DCF). The JL module provides robust saliency feature learning by exploiting cross-modal commonality via a Siamese network, while the DCF module is introduced for complementary feature discovery. Comprehensive experiments using 5 popular metrics show that the designed framework yields a robust RGB-D saliency detector with good generalization. As a result, JL-DCF significantly advances the SOTAs by an average of ~2.0% (F-measure) across 7 challenging datasets. In addition, we show that JL-DCF is readily applicable to other related multi-modal detection tasks, including RGB-T SOD and video SOD, achieving comparable or better performance. Keren Fu, Deng-Ping Fan, Ge-Peng Ji, Qijun Zhao, Jianbing Shen, Ce Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Unsupervised Real-World Image Super-Resolution via Dual Synthetic-to-Realistic and Realistic-to-Synthetic TranslationsabstractDue to the challenges of collecting paired low-resolution (LR) and high-resolution (HR) images in real-world scenarios, most existing deep convolutional neural network (CNN)-based single image super-resolution (SR) models are trained with artificially synthesized LR-HR image pairs. However, the domain gap between the synthetic data for model training and the realistic data for testing degrades SR performance significantly, which discourages the application of SR models in practice. One possible solution is to learn from unpaired real-world LR and HR images for their accessibility. Predominant strategies are mainly based on unsupervised domain translation. Despite great advances, there are still noticeable domain gaps between the realistic-like/synthetic-like images generated by unpaired translation and the true realistic/synthetic ones. To address this problem, this letter proposes an effective unsupervised SR framework based on dual synthetic-to-realistic and realistic-to-synthetic translations, namely DTSR. Specifically, to bridge the domain gap between testing and training data, the SR model is optimized using HR images and their realistic-like LR counterparts produced by the synthetic-to-realistic translation. In turn, we propose to narrow the domain gap further via applying the realistic-to-synthetic translation to realistic LR images prior to super-resolving, which also makes the SR model super-resolve simpler examples in testing relative to model training. Moreover, focal frequency and bilateral filtering losses are particularly introduced into DTSR for better details restoration and artifacts suppression. Extensive experiments show that our DTSR outperforms several state-of-the-art models in terms of both quantitative and qualitative comparisons. Honggang Chen, Ling Dong, Xiaohai He, Ce Zhu |
IEEE Signal Process. Lett. | 5 |
| 2022 | Symmetrical Epipolar Features Over Normalized Camera/Projector Calibration Matrices for Real-Time Structured Light IlluminationabstractStructured light illumination is a 3D scanning technique based on projecting a series of striped patterns and reconstructing depth based on the observed warping of the pattern across the target surface. Extensively studied, real-time performance is of paramount importance in many applications. In this Letter, we build on prior researches with epipolar geometry to further simplify the computations of 3D point clouds, with the symmetrical epipolar features derived by extending epipolar geometry over normalized calibration matrices of the camera and projector. Experiments show that the proposed two new processes are of the same accuracy over the normalized calibration matrices with substantially fewer calculations. Kai Liu 0012, Songlin Ying, Daniel L. Lau, Ce Zhu |
IEEE Signal Process. Lett. | 4 |
| 2022 | Multi-Scale Spatial and Temporal Speech Associations to Swallowing for Dysphagia ScreeningabstractDysphagia is a common symptom of many neurological diseases. It often occurs in older adults and increases the risk of aspiration pneumonia. Existing diagnosis systems of dysphagia are invasive or require patients to swallow liquids, which are costly and harmful to the patients. In this work, we propose an early screening system of dysphagia based on two kinds of throat signals, i.e., vowels and sentences. Based on the vowels, two new categories of speech features are developed: PET (pitch/energy trajectory) and FS-Conts (full spectrogram contours). The PET feature set focuses on the prominent resonance energy of speech to track the pitch and energy fluctuations. It can reflect the stability of vocal cords in the speech generation process. The FS-Conts feature set is proposed to emphasize the spatial details of formants based on three-dimensional contours. Concerning the sentences, three categories of speech features are proposed, called LSSDL (log symmetric spectral difference level), C-coes (crucial energy coefficients), and LDF (local dynamic features). The three features explore the speech representations of dysphagia from global variations to local associations. The LSSDL feature set is designed to highlight the global spectral differences in the interested frequency region. The C-coes and LDF feature sets locate local speech differences in specific frequency regions and time duration. In addition, a new feature selection algorithm is developed based on a newly designed precise matching analysis technique to search for distinguishing features. In the classification experiments, the SVM classifier is adopted and the dysphagia detection accuracy reaches 95.07%. The comparative experiments are conducted. The results indicate that our system performs better than the existing methods. Ce Zhu, Yipeng Liu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Learning Clustering for Motion SegmentationabstractSubspace clustering has been extensively studied from the hypothesis-and-test, algebraic, and spectral clustering-based perspectives. Most assume that only a single type/class of subspace is present. Generalizations to multiple types are non-trivial, plagued by challenges such as choice of types and numbers of models, sampling imbalance and parameter tuning. In many real world problems, data may not lie perfectly on a linear subspace and hand designed linear subspace models may not fit into these situations. In this work, we formulate the multi-type subspace clustering problem as one of learning non-linear subspace filters via deep multi-layer perceptrons (mlps). The response to the learnt subspace filters serve as the feature embedding that is clustering-friendly, i.e., points of the same clusters will be embedded closer together through the network. For inference, we apply K-means to the network output to cluster the data. Experiments are carried out on synthetic data and real world motion segmentation problems, producing state-of-the-art results. Xun Xu 0002, Le Zhang 0001, Loong Fah Cheong, Zhuwen Li, Ce Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Nonconvex Structural Sparsity Residual Constraint for Image RestorationabstractThis article proposes a novel nonconvex structural sparsity residual constraint (NSSRC) model for image restoration, which integrates structural sparse representation (SSR) with nonconvex sparsity residual constraint (NC-SRC). Although SSR itself is powerful for image restoration by combining the local sparsity and nonlocal self-similarity in natural images, in this work, we explicitly incorporate the novel NC-SRC prior into SSR. Our proposed approach provides more effective sparse modeling for natural images by applying a more flexible sparse representation scheme, leading to high-quality restored images. Moreover, an alternating minimizing framework is developed to solve the proposed NSSRC-based image restoration problems. Extensive experimental results on image denoising and image deblocking validate that the proposed NSSRC achieves better results than many popular or state-of-the-art methods over several publicly available datasets. Zhiyuan Zha, Xin Yuan 0002, Bihan Wen, Jiachao Zhang, Ce Zhu |
IEEE Trans. Cybern. | 5 |
| 2022 | MA-GANet: A Multi-Attention Generative Adversarial Network for Defocus Blur DetectionabstractBackground clutters pose challenges to defocus blur detection. Existing approaches often produce artifact predictions in background areas with clutter and relatively low confident predictions in boundary areas. In this work, we tackle the above issues from two perspectives. Firstly, inspired by the recent success of self-attention mechanism, we introduce channel-wise and spatial-wise attention modules to attentively aggregate features at different channels and spatial locations to obtain more discriminative features. Secondly, we propose a generative adversarial training strategy to suppress spurious and low reliable predictions. This is achieved by utilizing a discriminator to identify predicted defocus map from ground-truth ones. As such, the defocus network (generator) needs to produce ‘realistic’ defocus map to minimize discriminator loss. We further demonstrate that the generative adversarial training allows exploiting additional unlabeled data to improve performance, a.k.a. semi-supervised learning, and we provide the first benchmark on semi-supervised defocus detection. Finally, we demonstrate that the existing evaluation metrics for defocus detection generally fail to quantify the robustness with respect to thresholding. For a fair and practical evaluation, we introduce an effective yet efficient$AUF_\beta $metric. Extensive experiments on three public datasets verify the superiority of the proposed methods compared against state-of-the-art approaches. Xun Xu 0002, Le Zhang 0001, Chao Zhang 0072, Chuan-Sheng Foo, Ce Zhu |
IEEE Trans. Image Process. | 6 |
| 2022 | Depth Map Recovery Based on a Unified Depth Boundary Distortion ModelabstractDepth maps acquired by either physical sensors or learning methods are often seriously distorted due to boundary distortion problems, including missing, fake, and misaligned boundaries (compared with RGB images). An RGB-guided depth map recovery method is proposed in this paper to recover true boundaries in seriously distorted depth maps. Therefore, a unified model is first developed to observe all these kinds of distorted boundaries in depth maps. Observing distorted boundaries is equivalent to identifying erroneous regions in distorted depth maps, because depth boundaries are essentially formed by contiguous regions with different intensities. Then, erroneous regions are identified by separately extracting local structures of RGB image and depth map with Gaussian kernels and comparing their similarity on the basis of the SSIM index. A depth map recovery method is then proposed on the basis of the unified model. This method recovers true depth boundaries by iteratively identifying and correcting erroneous regions in recovered depth map based on the unified model and a weighted median filter. Because RGB image generally includes additional textural contents compared with depth maps, texture-copy artifacts problem is further addressed in the proposed method by restricting the model works around depth boundaries in each iteration. Extensive experiments are conducted on five RGB-depth datasets including depth map recovery, depth super-resolution, depth estimation enhancement, and depth completion enhancement. The results demonstrate that the proposed method considerably improves both the quantitative and visual qualities of recovered depth maps in comparison with fifteen competitive methods. Most object boundaries in recovered depth maps are corrected accurately, and kept sharply and well aligned with the ones in RGB images. Haotian Wang 0009, Meng Yang 0002, Xuguang Lan, Ce Zhu, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | Smooth Compact Tensor Ring RegressionabstractIn learning tasks with high order correlations, the low-rank approximation of the regression coefficient tensor has become increasingly important. Tensor ring can capture more correlation information among tensor networks. However, its optimal rank is generally unknown and needs to be tuned from multiple combinations. To address the issue, we propose a novel tensor regression framework with a group sparsity constraint on latent factors for tensor ring rank estimation. Specifically, the proposed group sparsity term constrained matrix factorization problem is first shown to be equivalent to a better approximation of matrix rank, namely Schatten-$1/2$quasi-norm. Extending it into tensor, the tensor ring rank can be inferred during the learning process to balance the prediction error and the model complexity. Besides, a total variation term is introduced to enhance the local consistency of the predicted response, which is useful for reducing the adverse effects of random noise. Experiments on the simulation dataset show that the proposed method can exactly obtain the tensor ring rank, and the effectiveness and robustness of the proposed algorithm is further verified on a real dataset for human motion capture tasks. Jiani Liu 0002, Ce Zhu, Yipeng Liu 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | Joint Contrast Enhancement and Exposure Fusion for Real-World Image DehazingabstractDue to the complexity of real environment and potential defects of current simulation datasets, either prior-based or deep learning-based single image dehazing methods may not work well in certain scenarios. In this work, we propose an efficient joint contrast enhancement and exposure fusion (CEEF) framework to formulate image dehazing task as a problem of enhancing local visibility and global contrast. In the contrast enhancement stage, several intermediate images are generated through two pre-processing steps. Specifically, gamma correction (GC) is used to adjust local visibility of an input hazy image. To address the issue of applying adaptive histogram equalization (AHE) to each color channel independently, we introduce color-preserving AHE (CP-AHE) to improve global contrast of the input hazy image. In the fusion stage, we develop a fast structural patch decomposition-based fusion strategy with an adaptive kernel size to fuse the inputs obtained by GC and CP-AHE. Extensive experiments on the real-world datasets demonstrate superiority of the proposed method to state-of-the-art methods in terms of visual and quantitative evaluation. Particularly for nighttime hazy scenes, our approach is shown to retain fine details and reduce color artifacts against three latest nighttime defogging methods. Moreover, we discuss potential applications of our CP-AHE in low-light enhancement and image editing. Xiaoning Liu 0003, Hui Li 0029, Ce Zhu |
IEEE Trans. Multim. | 3 |
| 2022 | A Hybrid Structural Sparsification Error Model for Image RestorationabstractRecent works on structural sparse representation (SSR), which exploit image nonlocal self-similarity (NSS) prior by grouping similar patches for processing, have demonstrated promising performance in various image restoration applications. However, conventional SSR-based image restoration methods directly fit the dictionaries or transforms to the internal (corrupted) image data. The trained internal models inevitably suffer from overfitting to data corruption, thus generating the degraded restoration results. In this article, we propose a novel hybrid structural sparsification error (HSSE) model for image restoration, which jointly exploits image NSS prior using both the internal and external image data that provide complementary information. Furthermore, we propose a general image restoration scheme based on the HSSE model, and an alternating minimization algorithm for a range of image restoration applications, including image inpainting, image compressive sensing and image deblocking. Extensive experiments are conducted to demonstrate that the proposed HSSE-based scheme outperforms many popular or state-of-the-art image restoration methods in terms of both objective metrics and visual perception. Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu, Alex Chichung Kot |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Feature Encoding With Autoencoders for Weakly Supervised Anomaly DetectionabstractWeakly supervised anomaly detection aims at learning an anomaly detector from a limited amount of labeled data and abundant unlabeled data. Recent works build deep neural networks for anomaly detection by discriminatively mapping the normal samples and abnormal samples to different regions in the feature space or fitting different distributions. However, due to the limited number of annotated anomaly samples, directly training networks with the discriminative loss may not be sufficient. To overcome this issue, this article proposes a novel strategy to transform the input data into a more meaningful representation that could be used for anomaly detection. Specifically, we leverage an autoencoder to encode the input data and utilize three factors, hidden representation, reconstruction residual vector, and reconstruction error, as the new representation for the input data. This representation amounts to encode a test sample with its projection on the training data manifold, its direction to its projection, and its distance to its projection. In addition to this encoding, we also propose a novel network architecture to seamlessly incorporate those three factors. From our extensive experiments, the benefits of the proposed strategy are clearly demonstrated by its superior performance over the competitive methods. Code is available at: https://github.com/yj-zhou/Feature_Encoding_with_AutoEncoders_for_Weakly-supervised_Anomaly_Detection. Yingjie Zhou 0001, Xucheng Song, Yanru Zhang, Fanxing Liu, Ce Zhu, Lingqiao Liu |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Prototype-Based Multisource Domain AdaptationabstractUnsupervised domain adaptation aims to transfer knowledge from labeled source domain to unlabeled target domain. Recently, multisource domain adaptation (MDA) has begun to attract attention. Its performance should go beyond simply mixing all source domains together for knowledge transfer. In this article, we propose a novel prototype-based method for MDA. Specifically, for solving the problem that the target domain has no label, we use the prototype to transfer the semantic category information from source domains to target domain. First, a feature extraction network is applied to both source and target domains to obtain the extracted features from which the domain-invariant features and domain-specific features will be disentangled. Then, based on these two kinds of features, the named inherent class prototypes and domain prototypes are estimated, respectively. Then a prototype mapping to the extracted feature space is learned in the feature reconstruction process. Thus, the class prototypes for all source and target domains can be constructed in the extracted feature space based on the previous domain prototypes and inherent class prototypes. By forcing the extracted features are close to the corresponding class prototypes for all domains, the feature extraction network is progressively adjusted. In the end, the inherent class prototypes are used as a classifier in the target domain. Our contribution is that through the inherent class prototypes and domain prototypes, the semantic category information from source domains is transformed into the target domain by constructing the corresponding class prototypes. In our method, all source and target domains are aligned twice at the feature level for better domain-invariant features and more closer features to the class prototypes, respectively. Several experiments on public data sets also prove the effectiveness of our method. Lihua Zhou, Mao Ye 0001, Ce Zhu, Luping Ji |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Multi-mode Tensor Singular Value Decomposition for Low-Rank Image Recovery
Lanlan Feng, Ce Zhu, Yipeng Liu 0001 |
ICIG (2) | 2 |
| 2021 | Deep Learning-Based Regional Sub-models Integration for Parkinson's Disease Diagnosis Using Diffusion Tensor Imaging
Hengling Zhao, Chih-Chien Tsai, Ce Zhu, Mingyi Zhou, Jiun-Jie Wang, Yipeng Liu 0001 |
ICIG (2) | 3 |
| 2021 | Low-Rank Regularized Joint Sparsity for Image DenoisingabstractNonlocal sparse representation models such as group sparse representation (GSR), low-rankness and joint sparsity (JS) have shown great potentials in image denoising studies, by effectively exploiting image nonlocal self-similarity (NSS) property. Popular dictionary-based JS algorithms apply convex JS penalties in their objective functions, which avoid NP-hard sparse coding step, but lead to only approximately sparse representation. Such approximated JS models fail to impose low-rankness of the underlying image data, resulting in degraded quality in image restoration. To simultaneously exploit the low-rank and JS priors, we propose a novel low-rank regularized joint sparsity model, dubbed LRJS, to enhance the dependency (i. e., low-rankness) of similar patches, thus better suppress independent noise. Moreover, to make the optimization tractable and robust, an alternating minimization algorithm with an adaptive parameter adjustment strategy is developed to solve the proposed LRJS-based image denoising problem. Experimental results demonstrate that the proposed LRJS outperforms many popular or state-of-the-art denoising algorithms in terms of both objective and visual perception met- Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu |
ICIP | 5 |
| 2021 | Smart Dysphagia Detection System with Adaptive Boosting Analysis of Throat SignalsabstractDysphagia is a symptom of many neurological disorders. Existing diagnosis systems are either invasive or require swallowing liquids, which are costly and harmful to humans. In this work, we design a smart dysphagia detection system based on speech signals. Rather than the voice data acquired by traditional microphones, we apply a bone conduction headset for vibration signal acquisition from the throat to get cleaner speech signals. After speech feature extraction, under-sampling is performed to deal with the imbalanced data problem, and principal component analysis is used for dimensionality reduction. In this paper, we construct an ensemble adaptive boosting classifier to detect the dysphagia patient. Experimental results show that the testing classification accuracy of the proposed system reaches 71.2%. Sensitivity and specificity can reach 66.6 % and 76 %, respectively. Shenghan Wang, Yangyang Jiang, Hengling Zhao, Ce Zhu, Yipeng Liu 0001 |
ISCAS | 6 |
| 2021 | Deep Learning Based Gait Analysis for Contactless Dementia Detection System from Video CameraabstractDementia is a neurodegenerative disease with a high incidence in the elderly. However, there is no effective treatment for this disease, and early intervention has a great effect to slow the deterioration. Currently, the detection of dementia is mainly achieved using questionnaire-like neuropsychological tests. Such ways usually cost a lot of time. To this end, we design a contactless dementia detection system based on gait analysis from surveillance video, and it can serve as a home-based healthcare system. This system applies a Kinect 2.0 camera to capture the human video and extract the skeleton joints at a rate of 15 frames per second. Two different gaits are collected for detection, namely single-task gait and dual-task gait. In this paper, we design a convolutional neural network based classifier to extract features in a data-driven way from these two groups of videos, but not take hand-crafted features. Experimental results show that we achieve a sensitivity of 74.10% on the test set using this system, and the processing only takes several minutes for early dementia detection. Yangyang Jiang, Xingyu Cao, Ce Zhu, Yipeng Liu 0001 |
ISCAS | 5 |
| 2021 | Unsupervised feature selection via multi-step markov probability relationship
Yan Min, Mao Ye 0001, Yulin Jian, Ce Zhu, Shangming Yang |
Neurocomputing | 5 |
| 2021 | Learning various length dependence by dual recurrent neural networks
Chenpeng Zhang, Shuai Li 0005, Mao Ye 0001, Ce Zhu, Xue Li 0001 |
Neurocomputing | 4 |
| 2021 | Deep embedded multi-view clustering with collaborative training
Jie Xu 0044, Yazhou Ren 0001, Guofeng Li, Lili Pan 0001, Ce Zhu, Zenglin Xu |
Inf. Sci. | 5 |
| 2021 | Low-rank tensor ring learning for multi-linear regression
Jiani Liu 0002, Ce Zhu, Zhen Long, Huyan Huang, Yipeng Liu 0001 |
Pattern Recognit. | 2 |
| 2021 | Single depth map super-resolution via joint non-local self-similarity modeling and local multi-directional gradient-guided regularization
Chao Ren 0002, Honggang Chen, Ce Zhu, Kai Liu 0012 |
Signal Process. Image Commun. | 4 |
| 2021 | Bayesian Low Rank Tensor Ring for Image RecoveryabstractLow rank tensor ring based data recovery can recover missing image entries in signal acquisition and transformation. The recently proposed tensor ring (TR) based completion algorithms generally solve the low rank optimization problem by alternating least squares method with predefined ranks, which may easily lead to overfitting when the unknown ranks are set too large and only a few measurements are available. In this article, we present a Bayesian low rank tensor ring completion method for image recovery by automatically learning the low-rank structure of data. A multiplicative interaction model is developed for low rank tensor ring approximation, where sparsity-inducing hierarchical prior is placed over horizontal and frontal slices of core factors. Compared with most of the existing methods, the proposed one is free of parameter-tuning, and the TR ranks can be obtained by Bayesian inference. Numerical experiments, including synthetic data, real-world color images and YaleFace dataset, show that the proposed method outperforms state-of-the-art ones, especially in terms of recovery accuracy. Zhen Long, Ce Zhu, Jiani Liu 0002, Yipeng Liu 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Image Restoration via Reconciliation of Group Sparsity and Low-Rank ModelsabstractImage nonlocal self-similarity (NSS) property has been widely exploited via various sparsity models such as joint sparsity (JS) and group sparse coding (GSC). However, the existing NSS-based sparsity models are either too restrictive, e.g., JS enforces the sparse codes to share the same support, or too general, e.g., GSC imposes only plain sparsity on the group coefficients, which limit their effectiveness for modeling real images. In this paper, we propose a novel NSS-based sparsity model, namely, low-rank regularized group sparse coding (LR-GSC), to bridge the gap between the popular GSC and JS. The proposed LR-GSC model simultaneously exploits the sparsity and low-rankness of the dictionary-domain coefficients for each group of similar patches. An alternating minimization with an adaptive adjusted parameter strategy is developed to solve the proposed optimization problem for different image restoration tasks, including image denoising, image deblocking, image inpainting, and image compressive sensing. Extensive experimental results demonstrate that the proposed LR-GSC algorithm outperforms many popular or state-of-the-art methods in terms of objective and perceptual metrics. Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu |
IEEE Trans. Image Process. | 5 |
| 2021 | Triply Complementary Priors for Image RestorationabstractRecent works that utilized deep models have achieved superior results in various image restoration (IR) applications. Such approach is typically supervised, which requires a corpus of training images with distributions similar to the images to be recovered. On the other hand, the shallow methods, which are usually unsupervised remain promising performance in many inverse problems, e.g., image deblurring and image compressive sensing (CS), as they can effectively leverage nonlocal self-similarity priors of natural images. However, most of such methods are patch-based leading to the restored images with various artifacts due to naive patch aggregation in addition to the slow speed. Using either approach alone usually limits performance and generalizability in IR tasks. In this paper, we propose a joint low-rank and deep (LRD) image model, which contains a pair of triply complementary priors, namely, internal and external, shallow and deep, and non-local and local priors. We then propose a novel hybrid plug-and-play (H-PnP) framework based on the LRD model for IR. Following this, a simple yet effective algorithm is developed to solve the proposed H-PnP based IR problems. Extensive experimental results on several representative IR tasks, including image deblurring, image CS and image deblocking, demonstrate that the proposed H-PnP algorithm achieves favorable performance compared to many popular or state-of-the-art IR methods in terms of both objective and visual perception. Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Joey Tianyi Zhou, Jiantao Zhou 0001, Ce Zhu |
IEEE Trans. Image Process. | 6 |
| 2021 | AMP-Net: Denoising-Based Deep Unfolding for Compressive Image SensingabstractMost compressive sensing (CS) reconstruction methods can be divided into two categories, i.e. model-based methods and classical deep network methods. By unfolding the iterative optimization algorithm for model-based methods onto networks, deep unfolding methods have the good interpretation of model-based methods and the high speed of classical deep network methods. In this article, to solve the visual image CS problem, we propose a deep unfolding model dubbed AMP-Net. Rather than learning regularization terms, it is established by unfolding the iterative denoising process of the well-known approximate message passing algorithm. Furthermore, AMP-Net integrates deblocking modules in order to eliminate the blocking artifacts that usually appear in CS of visual images. In addition, the sampling matrix is jointly trained with other network parameters to enhance the reconstruction performance. Experimental results show that the proposed AMP-Net has better reconstruction accuracy than other state-of-the-art methods with high reconstruction speed and a small number of network parameters. Yipeng Liu 0001, Jiani Liu 0002, Fei Wen 0005, Ce Zhu |
IEEE Trans. Image Process. | 5 |
| 2020 | Distribution-Aware Coordinate Representation for Human Pose EstimationabstractWhile being the de facto standard coordinate representation for human pose estimation, heatmap has not been investigated in-depth. This work fills this gap. For the first time, we find that the process of decoding the predicted heatmaps into the final joint coordinates in the original image space is surprisingly significant for the performance. We further probe the design limitations of the standard coordinate decoding method, and propose a more principled distributionaware decoding method. Also, we improve the standard coordinate encoding process (i.e. transforming ground-truth coordinates to heatmaps) by generating unbiased/accurate heatmaps. Taking the two together, we formulate a novel Distribution-Aware coordinate Representation of Keypoints (DARK) method. Serving as a model-agnostic plug-in, DARK brings about significant performance boost to existing human pose estimation models. Extensive experiments show that DARK yields the best results on two common benchmarks, MPII and COCO. Besides, DARK achieves the 2nd place entry in the ICCV 2019 COCO Keypoints Challenge. The code is available online. Feng Zhang 0052, Xiatian Zhu, Hanbin Dai, Mao Ye 0001, Ce Zhu |
CVPR | 5 |
| 2020 | DaST: Data-Free Substitute Training for Adversarial AttacksabstractMachine learning models are vulnerable to adversarial examples. For the black-box setting, current substitute attacks need pre-trained models to generate adversarial examples. However, pre-trained models are hard to obtain in real-world tasks. In this paper, we propose a data-free substitute training method (DaST) to obtain substitute models for adversarial black-box attacks without the requirement of any real data. To achieve this, DaST utilizes specially designed generative adversarial networks (GANs) to train the substitute models. In particular, we design a multi-branch architecture and label-control loss for the generative model to deal with the uneven distribution of synthetic samples. The substitute model is then trained by the synthetic samples generated by the generative model, which are labeled by the attacked model subsequently. The experiments demonstrate the substitute models produced by DaST can achieve competitive performance compared with the baseline models which are trained by the same train set with attacked models. Additionally, to evaluate the practicability of the proposed method on the real-world task, we attack an online machine learning model on the Microsoft Azure platform. The remote model misclassifies 98.35% of the adversarial examples crafted by our method. To the best of our knowledge, we are the first to train a substitute model for adversarial attacks without any real data. Mingyi Zhou, Jing Wu 0021, Yipeng Liu 0001, Shuaicheng Liu, Ce Zhu |
CVPR | 5 |
| 2020 | A Hybrid Structural Sparse Error Model for Image DeblockingabstractInspired by the image nonlocal self-similarity (NSS) prior, structural sparse representation (SSR) models exploit each group as the basic unit for sparse representation, which have achieved promising results in various image restoration applications. However, conventional SSR models only exploited the group within the input degraded (internal) image for image restoration, which can be limited by over-fitting to data corruption. In this paper, we propose a novel hybrid structural sparse error (HSSE) model for image deblocking. The proposed HSSE model exploits image NSS prior over both the internal image and external image corpus, which can be complementary in both feature space and image plane. Moreover, we develop an alternating minimization with an adaptive parameter setting strategy to solve the proposed HSSE model. Experimental results demonstrate that the proposed HSSE-based image deblocking algorithm outperforms many state-of-the-art image deblocking methods in terms of objective and visual perception. Zhiyuan Zha, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu, Bihan Wen |
ICASSP | 4 |
| 2020 | The Power Of Triply Complementary Priors For Image Compressive SensingabstractRecent works that utilized deep models have achieved superior results in various image restoration applications. Such approach is typically supervised which requires a corpus of training images with distribution similar to the images to be recovered. On the other hand, the shallow methods which are usually unsupervised remain promising performance in many inverse problems, e.g., image compressive sensing (CS), as they can effectively leverage non-local self-similarity priors of natural images. However, most of such methods are patch-based leading to the restored images with various ringing artifacts due to naive patch aggregation. Using either approach alone usually limits performance and generalizability in image restoration tasks. In this paper, we propose a joint low-rank and deep (LRD) image model, which contains a pair of triply complementary priors, namely external and internal, deep and shallow, and local and nonlocal priors. We then propose a novel hybrid plug-and-play (H-PnP) framework based on the LRD model for image CS. To make the optimization tractable, a simple yet effective algorithm is proposed to solve the proposed H-PnP based image CS problem. Extensive experimental results demonstrate that the proposed H-PnP algorithm significantly outperforms the state-of-the-art techniques for image CS recovery such as SCSNet and WNNM. Zhiyuan Zha, Xin Yuan 0002, Joey Tianyi Zhou, Jiantao Zhou 0001, Bihan Wen, Ce Zhu |
ICIP | 6 |
| 2020 | Reconciliation Of Group Sparsity And Low-Rank Models For Image RestorationabstractImage nonlocal self-similarity (NSS) property has been widely exploited via various sparsity models such as joint sparsity (JS) and group sparse coding (GSC). However, the existing NSS-based sparsity models are either too restrictive, i.e., JS enforces the sparse codes to share the same support, or too general, i.e., GSC imposes only plain sparsity on the group coefficients, which limit their effectiveness for modeling real images. In this paper, we propose a novel NSS-based sparsity model, namely low-rank regularized group sparse coding (LR-GSC), to bridge the gap between the popular GSC and JS. The proposed LR-GSC model simultaneously exploits the sparsity and low-rankness of the dictionary-domain coefficients for each group of similar patches. To make the proposed scheme tractable and robust, an alternating minimization with an adaptive adjusted parameter strategy is developed to solve the proposed optimization problem. Experimental results on both image deblocking and denoising demonstrate that the proposed LR-GSC image restoration algorithms outperform many popular or state-of-the-art methods, in terms of both the objective and perceptual quality. Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu |
ICME | 5 |
| 2020 | Bidirectional Independently Recurrent Neural Network for Skeleton-Based Hand Gesture RecognitionabstractGestures are a common form of human communication and important for Human-Computer Interaction (HCI). In this paper, we propose a new approach for skeleton-based hand gesture recognition based on the Independently Recurrent Neural Network (IndRNN). First, a bidirectional IndRNN (Bi-IndRNN) is developed to extend the IndRNN with the capability of bidirectional processing. Then, a deep Bi-IndRNN network is constructed for gesture recognition, where, in addition to the joint coordinates, the temporal displacement of each joint is also used to enhance the input features. Experimental results demonstrate that the proposed method achieves the state-of-the-art performance on the widely used DHG dataset with an accuracy of 93.15% for the 14 gesture classes case and 91.13% for the 28 gesture classes case. Shuai Li 0005, Longfei Zheng, Ce Zhu, Yanbo Gao |
ISCAS | 3 |
| 2020 | MultiANet: a Multi-Attention Network for Defocus Blur DetectionabstractDefocus blur detection is a challenging task because of obscure homogenous regions and interferences of background clutter. Most existing deep learning-based methods mainly focus on building wider or deeper network to capture multi-level features, neglecting to extract the feature relationships of intermediate layers, thus hindering the discriminative ability of network. Moreover, fusing features at different levels have been demonstrated to be effective. However, direct integrating without distinction is not optimal because low-level features focus on fine details only and could be distracted by background clutters. To address these issues, we propose the Multi-Attention Network for stronger discriminative learning and spatial guided low-level feature learning. Specifically, a channel-wise attention module is applied to both high-level and low-level feature maps to capture channel-wise global dependencies. In addition, a spatial attention module is employed to low-level features maps to emphasize effective detailed information. Experimental results show the performance of our network is superior to the state-of-the-art algorithms. Xun Xu 0002, Chao Zhang 0072, Ce Zhu |
MMSP | 4 |
| 2020 | Single depth map super-resolution via joint non-local and local modelingabstractDepth maps are widely used in 3D imaging techniques because of the appearance of the consumer depth cameras. However, the practical application of the depth map is limited by the poor image quality. In this paper, we propose a novel framework for the single depth map super-resolution via joint the local and non-local constraints simultaneously in the depth map. For the non-local constraint, we use the group-based sparse representation to explore the non-local self-similarity of the depth map. For the local constraint, we first estimate gradient images in different directions of the desired high-resolution (HR) depth map, and then build a multi-directional gradient guided regularizer using these estimated gradient images to describe depth gradients with different orientations. Finally, the two complementary regularizers are cast into a unified optimization framework to obtain the desired HR image. The experimental results show that the proposed method can achieve better depth super-resolution performance than state-of-the-art methods. Chao Ren 0002, Honggang Chen, Ce Zhu |
MMSP | 4 |
| 2020 | Smooth robust tensor principal component analysis for compressed sensing of dynamic MRI
Yipeng Liu 0001, Tengteng Liu, Jiani Liu 0002, Ce Zhu |
Pattern Recognit. | 4 |
| 2020 | Robust block tensor principal component analysis
Lanlan Feng, Yipeng Liu 0001, Longxi Chen, Xiang Zhang 0006, Ce Zhu |
Signal Process. | 5 |
| 2020 | Provable tensor ring completion
Huyan Huang, Yipeng Liu 0001, Jiani Liu 0002, Ce Zhu |
Signal Process. | 4 |
| 2020 | Low CP Rank and Tucker Rank Tensor Completion for Estimating Missing Components in Image DataabstractTensor completion recovers missing components of multi-way data. The existing methods use either the Tucker rank or the CANDECOMP/PARAFAC (CP) rank in low-rank tensor optimization for data completion. In fact, these two kinds of tensor ranks represent different high-dimensional data structures. In this paper, we propose to exploit the two kinds of data structures simultaneously for image recovery through jointly minimizing the CP rank and Tucker rank in the low-rank tensor approximation. We use the alternating direction method of multipliers (ADMM) to reformulate the optimization model with two tensor ranks into its two sub-problems, and each has only one tensor rank optimization. For the two main sub-problems in the ADMM, we apply rank-one tensor updating and weighted sum of matrix nuclear norms minimization methods to solve them, respectively. The numerical experiments on some image and video completion applications demonstrate that the proposed method is superior to the state-of-the-art methods. Yipeng Liu 0001, Zhen Long, Huyan Huang, Ce Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Fast Depth and Mode Decision in Intra Prediction for Quality SHVCabstractScalable High Efficiency Video Coding (SHVC) is the extension of High Efficiency Video Coding (HEVC). In intra prediction for quality SHVC, a Coding Unit (CU) is recursively divided into a quadtree-based structure from the largest 64×64 CU to the smallest 8×8 CU, in which 35 intra prediction modes and Inter-Layer Reference (ILR) mode are checked to determine the best possible mode. This leads to very high coding efficiency but also results in an extremely high coding complexity. To improve coding speed while maintaining coding efficiency, in this paper, we propose a new efficient algorithm for fast intra prediction for enhancement layer in SHVC. First, temporal and spatial correlations, as well as their correlation degrees, are combined in a Naive Bayes classifier to predict depth probabilities and skip depths with low likelihood. Second, for a given depth candidate, we combine ILR mode probability with Partial Zero Blocks (PZBs) based on the Sum of Squared Differences (SSD) to determine whether the ILR mode is the best one. In that case, we can skip intra prediction, which requires very high complexity. Third, initial Intra Modes (IMs) are obtained through Sobel operator, and are combined with the relationship between IMs and their corresponding Hadamard Cost (HC) values to predict candidate IMs in Rough Mode Decision (RMD). Then, an analytical criterion of early termination is developed based on the HC values of two neighboring IMs in the Rate-Distortion Optimization (RDO) process. Finally, we combine depth probabilities and the distribution of residual coefficients at the current depth to early terminate depth selection. The proposed scheme can significantly decrease the complexity of depth determination while reducing the complexity of mode decision for a depth candidate. Our experimental results demonstrate that the proposed scheme can achieve a speed up gain of more than 80% in average, while maintaining coding efficiency. Yu Sun 0003, Ce Zhu, Weisheng Li 0001, Frédéric Dufaux, Jiangtao Luo |
IEEE Trans. Image Process. | 3 |
| 2020 | Group Sparsity Residual Constraint With Non-Local Priors for Image RestorationabstractGroup sparse representation (GSR) has made great strides in image restoration producing superior performance, realized through employing a powerful mechanism to integrate the local sparsity and nonlocal self-similarity of images. However, due to some form of degradation (e.g., noise, down-sampling or pixels missing), traditional GSR models may fail to faithfully estimate sparsity of each group in an image, thus resulting in a distorted reconstruction of the original image. This motivates us to design a simple yet effective model that aims to address the above mentioned problem. Specifically, we propose group sparsity residual constraint with nonlocal priors (GSRC-NLP) for image restoration. Through introducing the group sparsity residual constraint, the problem of image restoration is further defined and simplified through attempts at reducing the group sparsity residual. Towards this end, we first obtain a good estimation of the group sparse coefficient of each original image group by exploiting the image nonlocal self-similarity (NSS) prior along with self-supervised learning scheme, and then the group sparse coefficient of the corresponding degraded image group is enforced to approximate the estimation. To make the proposed scheme tractable and robust, two algorithms, i.e., iterative shrinkage/thresholding (IST) and alternating direction method of multipliers (ADMM), are employed to solve the proposed optimization problems for different image restoration tasks. Experimental results on image denoising, image inpainting and image compressive sensing (CS) recovery, demonstrate that the proposed GSRC-NLP based image restoration algorithm is comparable to state-of-the-art denoising methods and outperforms several state-of-the-art image inpainting and image CS recovery methods in terms of both objective and perceptual quality metrics. Zhiyuan Zha, Xin Yuan 0002, Bihan Wen, Jiantao Zhou 0001, Ce Zhu |
IEEE Trans. Image Process. | 5 |
| 2020 | From Rank Estimation to Rank Approximation: Rank Residual Constraint for Image RestorationabstractIn this paper, we propose a novel approach for the rank minimization problem, termed rank residual constraint (RRC). Different from existing low-rank based approaches, such as the well-known nuclear norm minimization (NNM) and the weighted nuclear norm minimization (WNNM), which estimate the underlying low-rank matrix directly from the corrupted observation, we progressively approximate (approach) the underlying low-rank matrix via minimizing the rank residual. Through integrating the image nonlocal self-similarity (NSS) prior with the proposed RRC model, we apply it to image restoration tasks, including image denoising and image compression artifacts reduction. Toward this end, we first obtain a good reference of the original image groups by using the image NSS prior, and then the rank residual of the image groups between this reference and the degraded image is minimized to achieve a better estimate to the desired image. In this manner, both the reference and the estimated image in each iteration are improved gradually and jointly. Based on the group-based sparse representation model, we further provide a theoretical analysis on the feasibility of the proposed RRC model. Experimental results demonstrate that the proposed RRC model outperforms many state-of-the-art schemes in both the objective and perceptual qualities. Zhiyuan Zha, Xin Yuan 0002, Bihan Wen, Jiantao Zhou 0001, Jiachao Zhang, Ce Zhu |
IEEE Trans. Image Process. | 6 |
| 2020 | A Benchmark for Sparse Coding: When Group Sparsity Meets Rank MinimizationabstractSparse coding has achieved a great success in various image processing tasks. However, a benchmark to measure the sparsity of image patch/group is missing since sparse coding is essentially an NP-hard problem. This work attempts to fill the gap from the perspective of rank minimization. We firstly design an adaptive dictionary to bridge the gap between group-based sparse coding (GSC) and rank minimization. Then, we show that under the designed dictionary, GSC and the rank minimization problems are equivalent, and therefore the sparse coefficients of each patch group can be measured by estimating the singular values of each patch group. We thus earn a benchmark to measure the sparsity of each patch group because the singular values of the original image patch groups can be easily computed by the singular value decomposition (SVD). This benchmark can be used to evaluate performance of any kind of norm minimization methods in sparse coding through analyzing their corresponding rank minimization counterparts. Towards this end, we exploit four well-known rank minimization methods to study the sparsity of each patch group and the weighted Schatten p-norm minimization (WSNM) is found to be the closest one to the real singular values of each patch group. Inspired by the aforementioned equivalence regime of rank minimization and GSC, WSNM can be translated into a non-convex weighted ℓp-norm minimization problem in GSC. By using the earned benchmark in sparse coding, the weighted ℓp-norm minimization is expected to obtain better performance than the three other norm minimization methods, i.e., ℓ1-norm, ℓp-norm and weighted ℓ1-norm. To verify the feasibility of the proposed benchmark, we compare the weighted ℓp-norm minimization against the three aforementioned norm minimization methods in sparse coding. Experimental results on image restoration applications, namely image inpainting and image compressive sensing recovery, demonstrate that the proposed scheme is feasible and outperforms many state-of-the-art methods. Zhiyuan Zha, Xin Yuan 0002, Bihan Wen, Jiantao Zhou 0001, Jiachao Zhang, Ce Zhu |
IEEE Trans. Image Process. | 6 |
| 2020 | Image Restoration Using Joint Patch-Group-Based Sparse RepresentationabstractSparse representation has achieved great success in various image processing and computer vision tasks. For image processing, typical patch-based sparse representation (PSR) models usually tend to generate undesirable visual artifacts, while group-based sparse representation (GSR) models lean to produce over-smooth effects. In this paper, we propose a new sparse representation model, termed joint patch-group based sparse representation (JPG-SR). Compared with existing sparse representation models, the proposed JPG-SR provides an effective mechanism to integrate the local sparsity and nonlocal self-similarity of images. We then apply the proposed JPG-SR to image restoration tasks, including image inpainting and image deblocking. An iterative algorithm based on the alternating direction method of multipliers (ADMM) framework is developed to solve the proposed JPG-SR based image restoration problems. Experimental results demonstrate that the proposed JPG-SR is effective and outperforms many state-of-the-art methods in both objective and perceptual quality. Zhiyuan Zha, Xin Yuan 0002, Bihan Wen, Jiachao Zhang, Jiantao Zhou 0001, Ce Zhu |
IEEE Trans. Image Process. | 6 |
| 2020 | Image Restoration via Simultaneous Nonlocal Self-Similarity PriorsabstractThrough exploiting the image nonlocal self-similarity (NSS) prior by clustering similar patches to construct patch groups, recent studies have revealed that structural sparse representation (SSR) models can achieve promising performance in various image restoration tasks. However, most existing SSR methods only exploit the NSS prior from the input degraded (internal) image, and few methods utilize the NSS prior from external clean image corpus; how to jointly exploit the NSS priors of internal image and external clean image corpus is still an open problem. In this paper, we propose a novel approach for image restoration by simultaneously considering internal and external nonlocal self-similarity (SNSS) priors that offer mutually complementary information. Specifically, we first group nonlocal similar patches from images of a training corpus. Then a group-based Gaussian mixture model (GMM) learning algorithm is applied to learn an external NSS prior. We exploit the SSR model by integrating the NSS priors of both internal and external image data. An alternating minimization with an adaptive parameter adjusting strategy is developed to solve the proposed SNSS-based image restoration problems, which makes the entire algorithm more stable and practical. Experimental results on three image restoration applications, namely image denoising, deblocking and deblurring, demonstrate that the proposed SNSS produces superior results compared to many popular or state-of-the-art methods in both objective and perceptual quality measurements. Zhiyuan Zha, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu, Bihan Wen |
IEEE Trans. Image Process. | 4 |
| 2020 | Deep Cascade Model-Based Face Recognition: When Deep-Layered Learning Meets Small DataabstractSparse representation based classification (SRC), nuclear-norm matrix regression (NMR), and deep learning (DL) have achieved a great success in face recognition (FR). However, there still exist some intrinsic limitations among them. SRC and NMR based coding methods belong to one-step model, such that the latent discriminative information of the coding error vector cannot be fully exploited. DL, as a multi-step model, can learn powerful representation, but relies on large-scale data and computation resources for numerous parameters training with complicated back-propagation. Straightforward training of deep neural networks from scratch on small-scale data is almost infeasible. Therefore, in order to develop efficient algorithms that are specifically adapted for small-scale data, we propose to derive the deep models of SRC and NMR. Specifically, in this paper, we propose an end-to-end deep cascade model (DCM) based on SRC and NMR with hierarchical learning, nonlinear transformation and multi-layer structure for corrupted face recognition. The contributions include four aspects. First, an end-to-end deep cascade model for small-scale data without back-propagation is proposed. Second, a multi-level pyramid structure is integrated for local feature representation. Third, for introducing nonlinear transformation in layer-wise learning, softmax vector coding of the errors with class discrimination is proposed. Fourth, the existing representation methods can be easily integrated into our DCM framework. Experiments on a number of small-scale benchmark FR datasets demonstrate the superiority of the proposed model over state-of-the-art counterparts. Additionally, a perspective that deep-layered learning does not have to be convolutional neural network with back-propagation optimization is consolidated. The demo code is available in https://github.com/liuji93/DCM. Lei Zhang 0038, Ji Liu 0002, Bob Zhang 0001, David Zhang 0001, Ce Zhu |
IEEE Trans. Image Process. | 5 |
| 2020 | Fast Depth and Inter Mode Prediction for Quality Scalable High Efficiency Video CodingabstractThe scalable high efficiency video coding (SHVC) is an extension of high efficiency video coding (HEVC). It introduces multiple layers and inter-layer prediction, thus significantly increases the coding complexity on top of the already complicated HEVC encoder. In inter prediction for quality SHVC, in order to determine the best possible mode at each depth level, a coding tree unit can be recursively split into four depth levels, including merge mode, inter2N×2N, inter2N×N, interN×2N, interN×N, inter2N×nU, inter2N×nD, internL×2N and internRx×2N, intra modes and inter-layer reference (ILR) mode. This can obtain the highest coding efficiency, but also result in very high coding complexity. Therefore, it is crucial to improve coding speed while maintaining coding efficiency. In this research, we have proposed a new depth level and inter mode prediction algorithm for quality SHVC. First, the depth level candidates are predicted based on inter-layer correlation, spatial correlation and its correlation degree. Second, for a given depth candidate, we divide mode prediction into square and non-square mode predictions respectively. Third, in the square mode prediction, ILR and merge modes are predicted according to depth correlation, and early terminated whether residual distribution follows a Gaussian distribution. Moreover, ILR mode, merge mode and inter2N×2N are early terminated based on significant differences in Rate Distortion (RD) costs. Fourth, if the early termination condition cannot be satisfied, non-square modes are further predicted based on significant differences in expected values of residual coefficients. Finally, inter-layer and spatial correlations are combined with residual distribution to examine whether to early terminate depth selection. Experimental results have demonstrated that, on average, the proposed algorithm can achieve a time saving of 71.14%, with a bit rate increase of 1.27%. Yu Sun 0003, Ce Zhu, Weisheng Li 0001, Frédéric Dufaux |
IEEE Trans. Multim. | 3 |
| 2020 | Low-Rank Tensor Train Coefficient Array Estimation for Tensor-on-Tensor RegressionabstractThe tensor-on-tensor regression can predict a tensor from a tensor, which generalizes most previous multilinear regression approaches, including methods to predict a scalar from a tensor, and a tensor from a scalar. However, the coefficient array could be much higher dimensional due to both high-order predictors and responses in this generalized way. Compared with the current low CANDECOMP/PARAFAC (CP) rank approximation-based method, the low tensor train (TT) approximation can further improve the stability and efficiency of the high or even ultrahigh-dimensional coefficient array estimation. In the proposed low TT rank coefficient array estimation for tensor-on-tensor regression, we adopt a TT rounding procedure to obtain adaptive ranks, instead of selecting ranks by experience. Besides, an l2constraint is imposed to avoid overfitting. The hierarchical alternating least square is used to solve the optimization problem. Numerical experiments on a synthetic data set and two real-life data sets demonstrate that the proposed method outperforms the state-of-the-art methods in terms of prediction accuracy with comparable computational complexity, and the proposed method is more computationally efficient when the data are high dimensional with small size in each mode. Yipeng Liu 0001, Jiani Liu 0002, Ce Zhu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | C3AE: Exploring the Limits of Compact Model for Age EstimationabstractAge estimation is a classic learning problem in computer vision. Many larger and deeper CNNs have been proposed with promising performance, such as AlexNet, VggNet, GoogLeNet and ResNet. However, these models are not practical for the embedded/mobile devices. Recently, MobileNets and ShuffleNets have been proposed to reduce the number of parameters, yielding lightweight models. However, their representation has been weakened because of the adoption of depth-wise separable convolution. In this work, we investigate the limits of compact model for small-scale image and propose an extremely Compact yet efficient Cascade Context-based Age Estimation model(C3AE). This model possesses only 1/9 and 1/2000 parameters compared with MobileNets/ShuffleNets and VggNet, while achieves competitive performance. In particular, we re-define age estimation problem by two-points representation, which is implemented by a cascade model. Moreover, to fully utilize the facial context information, multi-branch CNN network is proposed to aggregate multi-scale context. Experiments are carried out on three age estimation datasets. The state-of-the-art performance on compact model has been achieved with a relatively large margin. Chao Zhang 0072, Shuaicheng Liu, Xun Xu 0002, Ce Zhu |
CVPR | 4 |
| 2019 | A Pruning Method Based on Feature Abstraction Capability of Filters
Xiang Zhang 0006, Ce Zhu |
ICIG (2) | 3 |
| 2019 | A Comparative Study for the Nuclear Norms Minimization MethodsabstractThe nuclear norm minimization (NNM) is commonly used to approximate the matrix rank by shrinking all singular values equally. However, the singular values have clear physical meanings in many practical problems, and NNM may not be able to faithfully approximate the matrix rank. To alleviate the above-mentioned limitation of NNM, recent studies have suggested that the weighted nuclear norm minimization (WNNM) can achieve a better rank estimation than NNM, which heuristically set the weight being inverse to the singular values. However, it still lacks a rigorous explanation why WNNM is more effective than NMM in various applications. In this paper, we analyze NNM and WNNM from the perspective of group sparse representation (GSR). Concretely, an adaptive dictionary learning method is devised to connect the rank minimization and GSR models. Based on the proposed dictionary, we prove that NNM and WNNM are equivalent to ℓ1-norm minimization and the weighted ℓ1-norm minimization in GSR, respectively. Inspired by enhancing sparsity of the weighted ℓ1-norm minimization in comparison with ℓ1-norm minimization in sparse representation, we thus explain that WNNM is more effective than NMM. By integrating the image nonlocal self-similarity (NSS) prior with the WNNM model, we then apply it to solve the image denoising problem. Experimental results demonstrate that WNNM is more effective than NNM and outperforms several state-of-the-art methods in both objective and perceptual quality. Zhiyuan Zha, Bihan Wen, Jiachao Zhang, Jiantao Zhou 0001, Ce Zhu |
ICIP | 5 |
| 2019 | Simultaneous Nonlocal Self-Similarity Prior for Image DenoisingabstractNonlocal image representation has achieved great success in various image processing tasks such as image denoising, image deblurring and image deblocking. Particularly, by exploiting the image nonlo-cal self-similarity (NSS) prior, many nonlocal similar patches can be searched across the whole image for a given patch, which has significantly boosted the performance of image restoration. To the best of our knowledge, most existing methods only consider the NSS prior of the input degraded image, while few methods exploit the NSS prior from external clean image corpus. However, how to utilize the NSS priors of input degraded image and external clean image corpus simultaneously is still an open problem. In this paper, we propose a novel approach for image denoising, which exploits simultaneous nonlocal self-similarity (SNSS) by integrating the NSS priors of both the input degraded image and external clean image corpus. Firstly, we search and group nonlocal similar patches from a clean image corpus, and a group-based Gaussian Mixture Model (GMM) learning algorithm is developed to learn an external NSS prior. Then, an optimal group is selected from the best suitable Gaussian component for a group of the noisy image. By integrating the group of the noisy image and the corresponding group of the Gaussian component with a low-rank constraint, an iterative algorithm is developed to solve the proposed SNSS model. Experimental results demonstrate that the proposed SNSS-based denoising method produces superior results compared with many state-of-the-art denoising methods in both objective and perceptual quality. Zhiyuan Zha, Xin Yuan 0002, Bihan Wen, Jiachao Zhang, Jiantao Zhou 0001, Ce Zhu |
ICIP | 6 |
| 2019 | Fast Inter Mode Predictions for SHVCabstractThe Scalable High Efficiency Video Coding (SHVC) has very high coding efficiency, but its computational complexity is also very high. This definitely limits its wide applications, particularly for real-time video applications. Therefore, it is crucial to improve the coding speed. In this research, we have proposed a new inter mode prediction algorithm for quality SHVC, in order to improve the coding speed while maintaining coding efficiency. First, we divide mode prediction into square mode prediction and non-square mode prediction. Second, in the square mode prediction, Inter-Layer Reference (ILR) and merge modes are predicted based on depth correlation. Moreover, ILR mode, merge mode and inter 2N×2N are early terminated based on Rate Distortion (RD) cost. Third, if the early termination condition cannot be satisfied, nonsquare modes are further predicted based on the distribution of residual coef-ficients. Experimental results have demonstrated that the proposed algorithm can significantly improve the coding speed with negligible coding efficiency loss. Yu Sun 0003, Weisheng Li 0001, Ce Zhu, Frédéric Dufaux |
ICME | 4 |
| 2019 | Early diagnosis of Parkinson's disease from multiple voice recordings by simultaneous sample and feature selection
Ce Zhu, Mingyi Zhou, Yipeng Liu 0001 |
Expert Syst. Appl. | 2 |
| 2019 | A fully trainable network with RNN-based pooling
Shuai Li 0005, Wanqing Li 0001, Chris Cook, Ce Zhu, Yanbo Gao |
Neurocomputing | 4 |
| 2019 | Feature level MRI fusion based on 3D dual tree compactly supported Shearlet transform
Chang Duan, Qi Hong Huang, Ce Zhu, Yuanyuan Xu 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2019 | Robust corner detection using altitude to chord ratio accumulation
Ce Zhu, Yipeng Liu 0001, Qian Zhang 0047 |
Multim. Tools Appl. | 2 |
| 2019 | Low rank tensor completion for multiway visual data
Zhen Long, Yipeng Liu 0001, Longxi Chen, Ce Zhu |
Signal Process. | 4 |
| 2019 | Tensor rank learning in CP decomposition via convolutional neural network
Mingyi Zhou, Yipeng Liu 0001, Zhen Long, Longxi Chen, Ce Zhu |
Signal Process. Image Commun. | 5 |
| 2019 | Source Distortion Temporal Propagation Analysis for Random-Access Hierarchical Video Coding OptimizationabstractDue to the widely used inter prediction in the current video coding standards, encoding units in different frames is of temporal dependency in that the rate-distortion optimization (RDO) of one unit may affect the coding performance of the following units in the temporal domain. To achieve optimal coding solution for a given video sequence, temporal dependency among units needs to be considered in the RDO process, which is known as the temporally dependent RDO (TD-RDO). The hierarchical coding structure (HCS) employed in the High Efficiency Video Coding (HEVC) standard further complicates this problem by grouping frames into different layers of varying coding strategies, leading to a more complex temporal relationship. In our earlier work, we addressed TD-RDO for the low delay HCS (LD-HCS), where only uni-prediction is considered. This paper aims to address more complicated TD-RDO under random access HCS (RA-HCS), where both uni-prediction and bi-prediction are considered, making the temporal relationship even more intricate. The temporal dependency introduced in the RA-HCS is thoroughly examined and an RA-based TD-RDO scheme is formulated for each layer by modeling temporal propagation of distortion under different prediction types. Based on the formulation, the global Lagrange multiplier can be obtained analytically. Moreover, the effect of random access point pictures is considered in the RA-based TD-RDO scheme. The proposed method can be simply realized by updating the Lagrange multiplier as in the independent RDO formulation or combined with adjusting quantization parameter (QP) for better results in terms of BD-rate saving. Experimental results show that under RA-HCS, the proposed method, by adapting the Lagrange multiplier only, can achieve about 2.2% bitrate savings in average. With multi-QP optimization, an average BD-rate gain of 5.2% can be obtained. Yanbo Gao, Ce Zhu, Shuai Li 0005, Tianwu Yang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Adaptive Deep Convolutional Neural Networks for Scene-Specific Object DetectionabstractA deep convolutional neural network (CNN) becomes a widely used tool for object detection. Many previous works have achieved excellent performance on object detection benchmarks. However, these works present generic detectors whose performance will drop rapidly when they are applied to a surveillance scene. In this paper, we propose an efficient method to construct a scene-specific regression model based on a generic CNN-based classifier. Our regression model is an adaptive deep CNN (ADCNN), which can predict object locations in the surveillance scene. First, we transfer the generic CNN-based classifier to the surveillance scene by selecting useful kernels. Second, we learn the context information of the surveillance scene in our regression model for accurate location prediction. Our main contributions are: 1) a transfer learning method that selects useful kernels in the generic CNN-based classifier; 2) a special architecture that can effectively learn the local and global context information in the surveillance scene; and 3) a new objective function to effectively train parameters in ADCNN. Compared with some state-of-the-art models, ADCNN achieves the best performance on three surveillance data sets for pedestrian detection and one surveillance data set for vehicle detection. Xudong Li 0001, Mao Ye 0001, Yiguang Liu, Ce Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | A Deep Learning Approach for Multi-Frame In-Loop Filter of HEVCabstractAn extensive study on the in-loop filter has been proposed for a high efficiency video coding (HEVC) standard to reduce compression artifacts, thus improving coding efficiency. However, in the existing approaches, the in-loop filter is always applied to each single frame, without exploiting the content correlation among multiple frames. In this paper, we propose a multi-frame in-loop filter (MIF) for HEVC, which enhances the visual quality of each encoded frame by leveraging its adjacent frames. Specifically, we first construct a large-scale database containing encoded frames and their corresponding raw frames of a variety of content, which can be used to learn the in-loop filter in HEVC. Furthermore, we find that there usually exist a number of reference frames of higher quality and of similar content for an encoded frame. Accordingly, a reference frame selector (RFS) is designed to identify these frames. Then, a deep neural network for MIF (known as MIF-Net) is developed to enhance the quality of each encoded frame by utilizing the spatial information of this frame and the temporal information of its neighboring higher-quality frames. The MIF-Net is built on the recently developed DenseNet, benefiting from its improved generalization capacity and computational efficiency. In addition, a novel block-adaptive convolutional layer is designed and applied in the MIF-Net, for handling the artifacts influenced by coding tree unit (CTU) structure in HEVC. Extensive experiments show that our MIF approach achieves on average 11.621% saving of the Bjøntegaard delta bit-rate (BD-BR) on the standard test set, significantly outperforming the standard in-loop filter in HEVC and other state-of-the-art approaches. Tianyi Li 0004, Mai Xu, Ce Zhu, Zulin Wang, Zhenyu Guan 0002 |
IEEE Trans. Image Process. | 3 |
| 2019 | Efficient Multi-Strategy Intra Prediction for Quality Scalable High Efficiency Video CodingabstractAs an extension of High Efficiency Video Coding (HEVC), the Scalable High Efficiency Video Coding (SHVC) introduces multiple layers with inter-layer predictions, which greatly increases the complexity on top of the already complicated HEVC encoder. In Intra prediction for Quality SHVC, Coding Tree Unit (CTU) allows recursive splitting into four depth levels, which considers 35 Intra prediction modes and interlayer reference (ILR) mode to determine the best possible mode at each depth level. This achieves the highest coding efficiency but incurs a substantially high computational complexity. In this paper, we propose a novel Intra prediction scheme to effectively speed up the enhancement layer Intra-coding in Quality SHVC. The new features of the proposed framework include: First, spatial correlation and its correlation degree are combined to predict most probable depth level candidates. Second, for a given depth candidate, based on the probabilities of the ILR mode, we check the ILR mode by examining the residual distribution based on skewness and kurtosis to determine whether the residuals follow a Gaussian distribution. In that case, the Intra prediction comparisons, which require a high complexity, are skipped. Third, during Intra prediction selection from 35 Intra prediction modes, spatial and inter-layer correlations are combined with the local monotonicity of the Hadamard costs associated with the modes in a small neighborhood, to examine only a portion of Intra prediction modes. Finally, a hypothesis testing on the currently selected depth level is performed to examine whether the residuals present significant differences within their block to early terminate depth selection. The proposed multi-step multistrategy scheme aims to minimize the number of depth selections while greatly reducing the mode decision complexity for a depth candidate in a hierarchical fashion. Our experimental results demonstrate that the proposed scheme can achieve a speedup gain of more than 75% in average on the test video sequences, while maintaining almost the same coding efficiency. . Ce Zhu, Yu Sun 0003, Frédéric Dufaux, Yuanyuan Huang 0007 |
IEEE Trans. Image Process. | 2 |
| 2019 | Image Completion Using Low Tensor Tree Rank and Total Variation MinimizationabstractTensor completion recovers missing entries of multiway data. Most of the current methods exploit the low-rank tensor structure for image completion applications. In this paper, we simultaneously exploit the globally multidimensional structure and locally piecewise smoothness to further enhance the performance. In the proposed optimization model, the low tensor tree rank minimization is used for the global data structure, and the total variation minimization is used for the local structure. Two kinds of total variation functions are discussed. The optimization problem is transformed into several subproblems by alternating direction method of multipliers. The subproblem on low tensor tree rank minimization is solved by singular value thresholding, and the subproblem on total variation minimization can be solved by soft thresholding. Numerical experiments on color images and light field images demonstrate that the proposed method outperforms most of the state-of-the-art methods in terms of recovery accuracy and computational complexity. Yipeng Liu 0001, Zhen Long, Ce Zhu |
IEEE Trans. Multim. | 3 |
| 2019 | Joint Texture/Depth Power Allocation for 3-D Video SoftCastabstractRecently, a novel uncoded (pseudoanalog) scheme called SoftCast is proposed for wireless video transmission, which eliminates the cliff effect of the state-of-the-art source-channel coding based schemes and achieves linear quality transition within a wide range of channel signal-to-noise ratio. Therefore, SoftCast-like uncoded and hybrid transmission has become an attractive research issue for natural 2-D video. However, very few studies focus on the SoftCast-based wireless transmission of the 3-D video (3DV) currently. One critical issue of 3DV SoftCast is how to allocate the limited power budget of the transmitter to the texture videos and depth maps of the 3DV to achieve the optimal overall quality on the receiver side, including the transmission quality of the reference views and the synthesis quality of the virtual views. This paper attempts to solve the optimal joint power allocation problem in an efficient way. First, we formulate the target problem as a constrained power-distortion optimization (PDO) problem mathematically. Then, each part of the distortion is analyzed and formulated in a closed form. Finally, the PDO problem is mapped to an unconstrained convex optimization problem and solved by the Lagrangian multiplier method. Simulation results demonstrate that the performance of the proposed method is close to that of the full search method, which can provide the best performance theoretically. Nevertheless, the complexity of the proposed method is negligible compared with that of the full search method. In addition, as compared with the fixed ratio (e.g., 1:1) power allocation between texture and depth, the proposed method can achieve a PNSR gain up to 1.8 dB. Lei Luo 0003, Taihai Yang, Ce Zhu, Zhi Jin 0002, Shu Tang |
IEEE Trans. Multim. | 3 |
| 2019 | Efficient Estimation of View Synthesis Distortion for Depth Coding OptimizationabstractDepth coding in depth-based three-dimensional (3-D) video is unique in that its quality is measured by view synthesis distortion (VSD) rather than the depth distortion itself, which further complicates the coding optimization as the VSD is related to quality of both the associated depth and texture videos. In this paper, an efficient VSD estimation scheme is developed to measure the effect of depth errors on the VSD for a block given its depth distortion in mean-squared error. Unlike other relevant VSD models which involve computationally intensive parameter training or Fourier transform, the proposed scheme is free of parameter training, while taking the advantage of integer 4 × 4 discrete Cosine transform to replace Fourier transform, thus well-saving computational cost and diminishing sensitivity to training dataset of video. The proposed scheme is then incorporated on the coding unit basis into the rate-distortion optimization for depth coding optimization, coupled with adapting quantization parameter accordingly to accommodate local effect of the depth errors on the VSD. Experimental results show that our solution obtains better results in depth coding than three testing solutions, on the platform of H.264/AVC reference software JM16.0. Benefiting from the efficiency of the VSD estimation, low coding complexity is obtained as well. The proposed solution is further evaluated on the reference software HTM13.0 of the latest 3-D high-efficiency video coding standard, exhibiting better and comparable results compared against the HTM codec with the view synthesis optimization disabled and enabled, respectively. Meng Yang 0002, Ce Zhu, Xuguang Lan, Nanning Zheng 0001 |
IEEE Trans. Multim. | 2 |
| 2018 | Joint Patch-Group Based Sparse Representation for Image InpaintingabstractSparse representation has achieved great successes in various machine learning and image processing tasks. For image processing, typical patch-based sparse representation (PSR) models usually tend to generate undesirable visual artifacts, while group-based sparse representation (GSR) models produce over-smooth phenomena. In this paper, we propose a new sparse representation model, termed joint patch-group based sparse representation (JPG-SR). Compared with existing sparse representation models, the proposed JPG-SR provides a powerful mechanism to integrate the local sparsity and nonlocal self-similarity of images. We then apply the proposed JPG-SR model to a low-level vision problem, namely, image inpainting. To make the proposed scheme tractable and robust, an iterative algorithm based on the alternating direction method of multipliers (ADMM) framework is developed to solve the proposed JPG-SR model. Experimental results demonstrate that the proposed model is efficient and outperforms several state-of-the-art methods in both objective and perceptual quality. Zhiyuan Zha, Xin Yuan 0002, Bihan Wen, Jiantao Zhou 0001, Ce Zhu |
ACML | 5 |
| 2018 | Independently Recurrent Neural Network (IndRNN): Building a Longer and Deeper RNNabstractRecurrent neural networks (RNNs) have been widely used for processing sequential data. However, RNNs are commonly difficult to train due to the well-known gradient vanishing and exploding problems and hard to learn long-term patterns. Long short-term memory (LSTM) and gated recurrent unit (GRU) were developed to address these problems, but the use of hyperbolic tangent and the sigmoid action functions results in gradient decay over layers. Consequently, construction of an efficiently trainable deep network is challenging. In addition, all the neurons in an RNN layer are entangled together and their behaviour is hard to interpret. To address these problems, a new type of RNN, referred to as independently recurrent neural network (IndRNN), is proposed in this paper, where neurons in the same layer are independent of each other and they are connected across layers. We have shown that an IndRNN can be easily regulated to prevent the gradient exploding and vanishing problems while allowing the network to learn long-term dependencies. Moreover, an IndRNN can work with non-saturated activation functions such as relu (rectified linear unit) and be still trained robustly. Multiple IndRNNs can be stacked to construct a network that is deeper than the existing RNNs. Experimental results have shown that the proposed IndRNN is able to process very long sequences (over 5000 time steps), can be used to construct very deep networks (21 layers used in the experiment) and still be trained robustly. Better performances have been achieved on various tasks by using IndRNNs compared with the traditional RNN and LSTM. Shuai Li 0005, Wanqing Li 0001, Chris Cook, Ce Zhu, Yanbo Gao |
CVPR | 4 |
| 2018 | Robust Tensor Principal Component Analysis in All ModesabstractRobust tensor principal component analysis extracts the low rank and sparse component of multi-dimensional data by tensor singular value decomposition (t-SVD), which can be used for many data analysis problems. However, the current t-SVD based methods cannot fully extract the low rank component in tensor data, and low rank structure still exists in the core tensor, because t-SVD does not decompose data in the third mode. To fully exploit the low rank structure, we further extract the low rank component using low rank plus sparsity for the core matrix whose entries are from the diagonal elements of the frontal slices in the core tensor. The proposed method is applied to three groups of numerical experiments on image denoising, illumination normalization for face images and motion separation for surveillance videos, respectively, and the results show that the proposed method outperforms state-of-the-art methods in terms of both accuracy and computational complexity. Longxi Chen, Yipeng Liu 0001, Ce Zhu |
ICME | 3 |
| 2018 | Image Ordinal Classification and Understanding: Grid Dropout with Masking LabelabstractImage ordinal classification refers to predicting a discrete target value which carries ordering correlation among image categories. The limited size of labeled ordinal data renders modern deep learning approaches easy to overfit. To tackle this issue, neuron dropout and data augmentation were proposed which, however, still suffer from over-parameterization and breaking spatial structure, respectively. To address the issues, we first propose a grid dropout method that randomly dropout/blackout some areas of the training image. Then we combine the objective of predicting the blackout patches with classification to take advantage of the spatial information. Finally we demonstrate the effectiveness of both approaches by visualizing the Class Activation Map (CAM) and discover that grid dropout is more aware of the whole facial areas and more robust than neuron dropout for small training dataset. Experiments are conducted on a challenging age estimation dataset-Adience dataset with very competitive results compared with state-of-the-art methods. Chao Zhang 0072, Ce Zhu, Jimin Xiao, Xun Xu 0002, Yipeng Liu 0001 |
ICME | 2 |
| 2018 | Extended smoothlets: An efficient multi-resolution adaptive transform
Qian Zhang 0047, Yipeng Liu 0001, Ce Zhu, Chang Duan |
J. Vis. Commun. Image Represent. | 5 |
| 2018 | Video analytical coding: When video coding meets video analysis
Ce Zhu, Min Mao, Fangliang Song, Frédéric Dufaux, Xiang Zhang 0006 |
Signal Process. Image Commun. | 2 |
| 2018 | Visual aesthetic understanding: Sample-specific aesthetic classification and deep activation map visualization
Chao Zhang 0072, Ce Zhu, Xun Xu 0002, Yipeng Liu 0001, Jimin Xiao, Tammam Tillo |
Signal Process. Image Commun. | 2 |
| 2018 | Hole Filling With Multiple Reference Views in DIBR View SynthesisabstractDepth-image-based rendering (DIBR) oriented view synthesis has been widely employed in the current depth-based 3-D video systems by synthesizing a virtual view from an arbitrary viewpoint. However, holes may appear in the synthesized view due to disocclusion, thus significantly degrading the quality. Consequently, efforts have been made on developing effective and efficient hole-filling algorithms. Current hole-filling techniques generally extrapolate/interpolate the hole regions with the neighboring information based on an assumption that the texture pattern in the holes is similar to that of the neighboring background information. However, in many scenarios, especially of complex texture, the assumption may not hold. In other words, hole-filling techniques can only provide an estimation for a hole which may not be good enough or may even be erroneous considering a wide variety of complex scene of images. In this paper, we first examine the view interpolation with multiple reference views, demonstrating that the problem of emerging holes in a target virtual view can be greatly alleviated by making good use of other neighboring complementary views in addition to its two (commonly used) most neighboring primary views. The effects of using multiple views for view extrapolation in reducing holes are also investigated in this paper. In view of the 3D Video and ongoing free-viewpoint TV standardization, we propose a new view synthesis framework, which employs multiple views to synthesize output virtual views. Furthermore, a scheme of selective warping of complementary views is developed by efficiently locating a small number of useful pixels in the complementary views for hole reduction, to avoid full warping of additional complementary views thus lowering greatly the warping complexity. Experimental results show that the hole size based on two primary reference views may be reduced by up to about 70% with the help of two complementary reference views in the case of view interpolation, while the hole size based on one primary reference view may be reduced by about 27% with the help of one more complementary reference view in view extrapolation. Moreover, it is shown that by using one more pair of views in view interpolation and one more view in view extrapolation, 10% hole pixels may be reduced additionally. Shuai Li 0005, Ce Zhu, Ming-Ting Sun |
IEEE Trans. Multim. | 2 |
| 2017 | Iterative block tensor singular value thresholding for extraction of lowrank component of image dataabstractTensor principal component analysis (TPCA) is a multi-linear extension of principal component analysis which converts a set of correlated measurements into several principal components. In this paper, we propose a new robust TPCA method to extract the principal components of the multi-way data based on tensor singular value decomposition. The tensor is split into a number of blocks of the same size. The low rank component of each block tensor is extracted using iterative tensor singular value thresholding method. The principal components of the multi-way data are the concatenation of all the low rank components of all the block tensors. We give the block tensor incoherence conditions to guarantee the successful decomposition. This factorization has similar optimality properties to that of low rank matrix derived from singular value decomposition. Experimentally, we demonstrate its effectiveness in two applications, including motion separation for surveillance videos and illumination normalization for face images. Longxi Chen, Yipeng Liu 0001, Ce Zhu |
ICASSP | 3 |
| 2017 | Attribute-controlled face photo synthesis from simple line drawingabstractFace photo synthesis from simple line drawing is a one-to-many task as simple line drawing merely contains the contour of human face. Previous exemplar-based methods are over-dependent on the datasets and are hard to generalize to complicated natural scenes. Recently, several works utilize deep neural networks to increase the generalization, but they are still limited in the controllability of the users. In this paper, we propose a deep generative model to synthesize face photo from simple line drawing controlled by face attributes such as hair color and complexion. In order to maximize the controllability of face attributes, an attribute-disentangled variational auto-encoder (AD-VAE) is firstly introduced to learn latent representations disentangled with respect to specified attributes. Then we conduct photo synthesis from simple line drawing based on AD-VAE. Experiments show that our model can well disentangle the variations of attributes from other variations of face photos and synthesize detailed photorealistic face images with desired attributes. Regarding background and illumination as the style and human face as the content, we can also synthesize face photos with the target style of a style photo. Ce Zhu, Zhiqiang Xia, Yipeng Liu 0001 |
ICIP | 2 |
| 2017 | Towards thinner convolutional neural networks through gradually global pruningabstractDeep network pruning is an effective method to reduce the storage and computation cost of deep neural networks when applying them to resource-limited devices. Among many pruning granularities, neuron level pruning will remove redundant neurons and filters in the model and result in thinner networks. In this paper, we propose a gradually global pruning scheme for neuron level pruning. In each pruning step, a small percent of neurons were selected and dropped across all layers in the model. We also propose a simple method to eliminate the biases in evaluating the importance of neurons to make the scheme feasible. Compared with layer-wise pruning scheme, our scheme avoid the difficulty in determining the redundancy in each layer and is more effective for deep networks. Our scheme would automatically find a thinner sub-network in original network under a given performance. Ce Zhu, Zhiqiang Xia, Yipeng Liu 0001 |
ICIP | 2 |
| 2017 | Temporal corelation based hierarchical quantization parameter determination for HEVC video codingabstractThe quantization parameter (QP) value and Lagrangian multiplier (λ) are the key factors for an encoder to achieve the trade-off between visual quality and bit-rate in next generation multimedia communications. In this work, we propose a novel temporal redundancy ratio (TRR) model to determinate hierarchical QPs. Taking the temporal redundancy information, intensity and fluctuation into consideration, the TRR model is constructed as the ratio of mean and variance from the temporal redundancy, and then used to allocate a proper QP value for an individual picture to achieve bit-rate saving and improve the visual quality. We implement the TRR model based QP determination scheme into the reference software HM16.7. Under the common test condition (CTC), simulation results show that the proposed TRR model obtains 1.28% and 1.52% BD-Rate gain over HM16.7 for Low-Delay and Random-Access in 8-bit video coding respectively. Meanwhile, more BD-Rate gains are obtained in 10-bit coding and the screen coding class. Yimin Zhou 0002, Ling Tian, Ce Zhu |
ICIP | 4 |
| 2017 | Memory-based pedestrian detection through sequence learningabstractHuman recognize an object through eyes scanning in a certain order. We think that the proper order is helpful for capturing useful characteristics, which makes our recognition process rapidly and accurately. Therefore, we propose a memory-based sequence learning model to simulate the human recognition process. Firstly, we divide the image without overlapping to generate the sequence. Then, a convolutional neural network is used for feature extraction. Next, the sequence is re-sorted by order of importance. Finally, a long short-term memory successively receives the sequence to memorize the sequential patterns and predict the sequence label. In addition, we propose a joint learning method to make our model efficiently learn both of the sequence order and the sequence patterns. Our model is applied in the region-based detection framework for pedestrian detection. Compared with the state-of-the-art methods on two pedestrian datasets, our method achieves the comparable performance in term of accuracy and speed. Xudong Li 0001, Mao Ye 0001, Yiguang Liu, Ce Zhu |
ICME | 4 |
| 2017 | Status-aware projection metric learning for kinship verificationabstractThis paper develops a status-aware projection metric learning (SPML) method for facial image-based kinship verification, especially for the parent-child kinship. Kinship verification for parent-child is considered to be an asymmetrical metric process, in that parents and children are associated with different status where the parents are priority known to be significantly older than the children. Accordingly, an SPML is proposed to address the asymmetric metric learning. The proposed SPML takes advantage of two status-specific projections to capture the significant appearance commonality between parents and children, respectively, which generally outperforms the one Mahalanobis distance metric. Extensive experimental results and comparisons with state-of-the-art approaches and baseline methods demonstrate the effectiveness of the proposed SPML for kinship verification. Ce Zhu |
ICME | 2 |
| 2017 | A frame-level rate control scheme for low delay video coding in HEVCabstractR-λ rate control scheme is recommended in the High Efficiency Video Coding (HEVC) standard, which shows high accuracy of bit rate control but lower rate distortion performance. In order to minimize the distortion subject to a target bit rate, a λ domain frame-level rate control scheme for low delay coding of HEVC was proposed. Firstly, an improved parameter updating method is presented for the frame level rate distortion model, which full uses the information of encoded frames in the previous group of pictures (GOP). Then, an adaptive dynamic frame level bit allocation scheme is proposed by employing global rate distortion optimization theory. Finally, to further improve the coding efficiency, the bits allocation adjustment is made for the first several GOPs in video sequence according to the rate distortion dependency. The experimental results show that the proposed method can greatly improve the rate distortion performance under the condition of high accuracy of bit rate control. Hongwei Guo 0001, Ce Zhu, Yanbo Gao, Shichang Song |
MMSP | 2 |
| 2017 | Learning based 3D keypoint detection with local and global attributes in multi-scale spaceabstractOver the last few decades various methods have been proposed by researchers to extract 3D keypoints from the surface of 3D mesh models, but most of them are geometric ones, which are not flexible enough for various applications. In this paper, we propose a new 3D keypoint detection method based on multi-scale neural network (MSNN), which is a tiny neural network and can effectively merge multi-scale information to detect 3D keypoints. Traditional end-to-end learning systems usually require large-scale dataset to do training. However, there are not enough 3D data with ground truth of 3D keypoints. To solve this problem, we perform delicate preprocessing, which effectively enhance the performance of the MSNN based approach. Numerical experiments show that the proposed MSNN 3D keypoint detector not only outperforms other six state-of-the-art geometric based methods, but also achieves better performance than a learning-based method using random forest. Ce Zhu, Qian Zhang 0047, Mengxue Wang, Yipeng Liu 0001 |
MMSP | 2 |
| 2017 | Analytical distortion aware video coding for computer based video analysisabstractWith the development of artificial intelligence, more and more multimedia applications for various tasks have emerged in our daily life. Meanwhile, as one of the main information sources of the applications, a huge amount of video data has been being generated by portable or mounted cameras in daily basis for varying purposes including surveillance, in which case we may need computers to "watch" videos to save labor cost. However, most video coding standards are designed for the highest human perceptual quality given a bit rate by minimizing a fidelity cost function (e.g., mean squared error, MSE), assuming the content will be consumed by human beings. In view of the above considerations, this paper proposes a new rate-analytical-distortion optimization method (RADO) for video analysis. Specifically, we consider moving object detection as the analysis task. Accordingly, we develop a novel rate analytical distortion (RAD) model for video coding, where the analytical distortion is related to the object detection performance expressed in terms of F-measure. As shown in the experimental results, the performance of the video analysis task can be significantly improved (up to 40% reduction of analytical distortion) with a slight bit rate increase. Ce Zhu, Min Mao, Fangliang Song, Frédéric Dufaux, Xiang Zhang 0006 |
MMSP | 2 |
| 2017 | Efficient and Robust Corner Detectors Based on Second-Order Difference of ContourabstractAs one of the most significant local features of image, corner is widely used in many computer vision tasks. Corner detection aims to achieve the highest possible detection accuracy while minimizing the computational complexity. In this letter, we first introduce a new measurement termed as second-order difference of contour (SODC), and then examine its regular distribution, which is found to provide useful information to distinguish corners from noncorners. Based on the SODC distribution characteristics, we propose two novel corner detectors to measure the response of contour points using Manhattan distance and Euclidean distance, respectively. Numerical experiments demonstrate that the Manhattan detector greatly decreases the computational complexity, while the Euclidean detector outperforms the state-of-the-art corner detectors in terms of repeatability and localization error. Ce Zhu, Qian Zhang 0047, Xiaolin Huang, Yipeng Liu 0001 |
IEEE Signal Process. Lett. | 2 |
| 2017 | A Bayesian Approach to Camouflaged Moving Object DetectionabstractMoving object detection is about foreground and background separation based on motion detection. Detecting moving objects from similarly colored background (known as camouflage problem) has been a long-standing open question in this field. Discriminative modeling (DM), which focuses on enhancing the performance to distinguish foreground from background with discriminative features and well-designed classifiers, has been widely used for moving object detection. However, DM may tend to fail when encountering the camouflage problem, as the class separability in camouflaged areas is generally poor. In this paper, we propose a new strategy, camouflage modeling (CM), to identify camouflaged foreground pixels. In view of the fact that camouflage involves both foreground and background, we need to model both the background and the foreground, and compare them in a well-designed way in camouflage detection. Specifically, we develop a global model for the background, and an integration of global and local models for the foreground, respectively. Based on both background and foreground models, we introduce a factor to measure the degree of camouflage, and further identify truly camouflaged areas. In view of the fact that a moving object is usually composed of both camouflaged and noncamouflaged areas, CM and DM are fused in a Bayesian framework to perform complete object detection. Experiments are conducted on testing sequences to demonstrate the effectiveness of the proposed algorithm. Xiang Zhang 0006, Ce Zhu, Yipeng Liu 0001, Mao Ye 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Temporally Dependent Rate-Distortion Optimization for Low-Delay Hierarchical Video CodingabstractLow-delay hierarchical coding structure (LD-HCS), as one of the most important components in the latest High Efficiency Video Coding (HEVC) standard, greatly improves coding performance. It groups consecutive P/B frames into different layers and encodes them with different quantization parameters (QPs) and reference mechanisms in such a way that temporal dependency among frames can be exploited. However, due to varying characteristics of video contents, temporal dependency among coding units differs significantly from each other in the same or different layers, while a fixed LD-HCS scheme cannot take full advantage of the dependency, leading to a substantial loss in coding performance. This paper addresses the temporally dependent rate distortion optimization (RDO) problem by attempting to exploit varying temporal dependency of different units. First, the temporal relationship of different frames under the LD-HCS is examined, and hierarchical temporal propagation chains are constructed to represent the temporal dependency among coding units in different frames. Then, a hierarchical temporally dependent RDO scheme is developed specifically for the LD-HCS based on a source distortion propagation model. Experimental results show that our proposed scheme can achieve 2.5% and 2.3% BD-rate gain in average compared with the HEVC codec under the same configuration of P and B frames, respectively, with a negligible increase in encoding time. Furthermore, coupled with QP adaption, our proposed method can achieve higher coding gains, e.g., with multi-QP optimization, about 5.4% and 5.0% BD-rate saving in average over the HEVC codec under the same setting of P and B frames, respectively. Yanbo Gao, Ce Zhu, Shuai Li 0005, Tianwu Yang |
IEEE Trans. Image Process. | 2 |
| 2017 | A Novel Method of Minimizing View Synthesis Distortion Based on Its Non-Monotonicity in 3D VideoabstractIn depth-based 3D video, the view synthesis distortion (VSD), is generally measured by modeling the effect of texture and depth errors separately. With such a development, it has been referred that the VSD changes monotonically with respect to to both the texture and depth distortions. In this paper, we find that the VSD does not always change monotonically with them by both theoretical analysis and experimental test, when the effect of the texture and depth errors is considered together. Specifically, first, we prove that the VSD is non-monotonic with the texture distortion. That is, the VSD increases with the increasing texture distortion at higher distortion range but conversely decreases with it at lower range. It is different from the general scenario that only considering the effect of the texture errors. We also analytically depict their relationship with low computational cost and identify the turning point at which the change of the VSD is converted. Second, we confirm that the VSD is always monotonic with the depth distortion, which is consistent with the general scenario that only considering the effect of the depth errors. The non-monotonicity property of the VSD can be utilized to improve the viewing performance of 3D video in relevant applications, since a minimal value of the VSD exists at the turning point. We conduct two applications for this purpose. First, it is used to generate the synthesis view of minimal distortion, which achieves 0.51-dB gain of PSNR on average for the tested scenarios. Second, it is used for lossy compression of texture videos in 3D video, which reduces the coding rate by 24% on average for the tested scenarios, meanwhile, keeps the VSD not increased simultaneously. Meng Yang 0002, Nanning Zheng 0001, Ce Zhu, Fei Wang 0008 |
IEEE Trans. Image Process. | 3 |
| 2017 | Hybrid CS-DMRI: Periodic Time-Variant Subsampling and Omnidirectional Total Variation Based ReconstructionabstractCompressive sensing (CS) has been used to accelerate dynamic magnetic resonance imaging (DMRI). Currently, the online CS-DMRI is faster, whereas the offline CS-DMRI provides higher accuracy for image reconstruction. To achieve good image reconstruction performance in terms of both speed and accuracy, we propose a hybrid CS-DMRI method using periodic time-variant subsampling for different frames. In each period, there is one reference frame that is sampled at a higher subsampling ratio. The two nearby reference frames with good reconstruction quality can be used to provide rough predictions of the other frames between them. To finely recover the current frame, one structural regularization in the optimization model for reconstruction is a 2-D omnidirectional total variation (OTV) for exploiting the sparsity of the difference between the predicted and estimated frames, and the other is a 3-D OTV as a regularization term for exploiting the bilateral spatio-temporal coherence between the forward reference frame, current frame, and backward reference frame. Compared with classical total variation, the proposed OTV fully utilizes the correlations of all the possible directions of the data. The formulated optimization model can be solved using iterative reweighted least squares with the pre-conditioned conjugate gradient method. Numerical experiments demonstrate that the proposed method has better reconstruction accuracy than all the existing methods and low computational complexity that is comparable to the existing online methods. Yipeng Liu 0001, Xiaolin Huang, Ce Zhu |
IEEE Trans. Medical Imaging | 5 |
| 2017 | An Imbalance Compensation Framework for Background SubtractionabstractClass imbalance refers to the instance where the number of training samples for the majority classes is far more than that of the minority classes (relative imbalance), and the quality of training samples for the minority classes is inferior to that of the majority classes (absolute imbalance), which are further complicated by other imbalance factors, e.g., data overlapping. Video background subtraction aims to classify each pixel into two classes: foreground and background. This paper first reveals that background subtraction is a class imbalance problem, where the foreground and background are the minority and majority classes, respectively. By exploring spatial and temporal correlation inherent in video data, we present an imbalance compensation framework for background subtraction, which consists of two sequential modules, imbalance-compensated bilayer modeling, and imbalance-compensated Bayesian classification. In the first module, spatio-temporal oversampling (SOS) and selective downsampling (SDS) are proposed to compensate the imbalance at data level. SOS attempts to synthesize representative samples appended to the minority sample set, while SDS selectively deletes a number of majority samples in data overlapping areas. The rebalanced samples are then used to learn a bilayer model. In the second module, novel cost functions are proposed to compensate the effect of class imbalance at algorithm level. The cost functions are based on imbalance measurement, and used to construct the prior term in the Bayesian classification scheme. Experiments are conducted on public databases to demonstrate the effectiveness of the proposed method. Xiang Zhang 0006, Ce Zhu, Honggang Wu, Zhi Liu 0003, Yuanyuan Xu 0001 |
IEEE Trans. Multim. | 2 |
| 2016 | Hierarchical temporal dependent rate-distortion optimization for low-delay codingabstractHierarchical coding structure (HCS) is one of the most important components in High Efficiency Video Coding (HEVC) that improves the coding performance greatly, especially for Low-Delay (LD) coding. It groups frames into different layers and enc odes them with different quantization parameters (QP) and different reference mechanisms. Due to the extensively used inter-prediction, the coding of frames in different layers is highly dependent and an appropriate QP and reference selection scheme may significantly improve the performance by taking advantage of such temporal dependency. However in the current HEVC codec, a predefined HCS, such as the Low-Delay HCS (LD-HCS), is performed without considering the different characteristic of different video contents, thus leading to a suboptimal coding solution. In this paper, the hierarchical temporal relationship under LD-HCS is first investigated and a hierarchical temporal propagation chain is constructed to describe the temporal dependency among frames. Then a hierarchical temporal dependent rate-distortion optimization scheme is developed specifically for the LD-HCS in HEVC. Experiments results show that the proposed scheme achieves BD-rate saving of 2.9% and 2.8% in average against HEVC codec under LD-HCS of P and B frames, respectively, with a negligible increase in encoding time. Yanbo Gao, Ce Zhu, Shuai Li 0005 |
ISCAS | 2 |
| 2016 | Layer-based temporal dependent rate-distortion optimization in Random-Access hierarchical video codingabstractRate-distortion optimization (RDO) plays an important part in improving the coding efficiency of High Efficiency Video Coding (HEVC), especially for the hierarchical coding structure defined in the Random-Access (RA) configuration, noted as Random-Access Hierarchical Video Coding (RA-HVC), where different frames are assigned to different temporal layers and further coded with different coding parameters. Due to the inter-frame prediction, coding result of one unit may affect the coding performance of the following temporally related units. Therefore, the temporal dependency among units needs to be considered in the coding process. However, the RDO process in the current video codec is performed without considering the varying temporal dependency, thus compromising the rate-distortion performance significantly. To address this problem, a layer-based temporal dependent RDO method is proposed in this paper where the temporal dependency among different frames in the same or different layers is examined. By reformulating the temporal dependent RDO for the RA-HVC, we show that it can be implemented in a way of simply refining the Lagrange multiplier. Experimental results show that the proposed method achieves, in average, about 1.4% BD-rate savings with a negligible increase in encoding time for the random-access configuration. Yanbo Gao, Ce Zhu, Shuai Li 0005, Tianwu Yang |
MMSP | 2 |
| 2016 | 3D interest point detection based on geometric measures and sparse refinementabstractThree dimensional (3D) interest point detection plays a fundamental role in computer vision. In this paper, we introduce a new method for detecting 3D interest points of 3D mesh models based on geometric measures and sparse refinement (GMSR). The key point of our approach is to calculate the 3D saliency measure using two novel geometric measures, which are defined in multi-scale space to effectively distinguish 3D interest points from edges and flat areas. Those points with local maxima of 3D saliency measure are selected as the candidates of 3D interest points. Finally, we utilize an l0norm based optimization method to refine the candidates of 3D interest points by constraining the number of 3D interest points. Numerical experiments show that the proposed GMSR based 3D interest point detector outperforms current six state-of-the-art methods for different kinds of 3D mesh models. Ce Zhu, Qian Zhang 0047, Yipeng Liu 0001 |
MMSP | 2 |
| 2016 | Guest Editors' Introduction: Special issue on deep learning with applications to visual representation and analysis
Lei Wang 0001, Ce Zhu, Jieping Ye, Juergen Gall |
Signal Process. Image Commun. | 2 |
| 2016 | Lagrangian Multiplier Adaptation for Rate-Distortion Optimization With Inter-Frame DependencyabstractRate-distortion optimization (RDO) is widely used in video coding, which plays a critical role in enhancing the coding efficiency substantially. Currently, the RDO process is performed in a way that coding efficiency of each coding unit (CU) is maximized independently without considering the dependency among CUs. As we know, in the current hybrid video coding structure, spatial/temporal prediction techniques are extensively used, which introduce strong dependency among CUs. In this paper, we investigate RDO with inter-frame dependency, where the impact of coding performance of the current CU on that of the following frames is considered. Accordingly, an RDO scheme taking the inter-frame dependency into account is proposed by adapting the Lagrangian multiplier. The experimental results show that the proposed scheme can achieve about 3.22% and 3.19% BD-rate saving in average over the state-of-the-art High Efficiency Video Coding (HEVC) reference software HM15.0 in the low-delay $P$ (LDP) and low-delay $B$ (LDB) coding structures, respectively, with no extra encoding time. The proposed scheme can obtain a significantly higher coding gain than the multiple quantization parameter (MQP) (±3) optimization technique that would greatly increase the encoding time by a factor of about six. Coupled with MQP optimization, the proposed scheme can further achieve about 5.96% and 5.57% BD-rate savings in average over the HEVC and about 4.03% and 4.07% over the HEVC with MQP optimization, under the specified common test conditions for LDP and LDB coding structures, respectively. Shuai Li 0005, Ce Zhu, Yanbo Gao, Yimin Zhou 0002, Frédéric Dufaux, Ming-Ting Sun |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | A Reversible Data Hiding Scheme Using Ordered Cluster-Based VQ Index Tables for Complex Images
Junlan Bai, Chin-Chen Chang 0001, Ce Zhu |
ICIG (1) | 3 |
| 2015 | Statistical approach for motion estimation skipping (SAMEK)abstractHigh Efficient Video Coding (HEVC) standard has achieved significant rate-distortion improvement over the previous standard H.264/AVC. However, the complexity that comes from its flexible data structure representation is an obstacle for its wide application. To reduce the overall complexity and encoding time, this paper proposes a statistical approach for motion estimation skipping (SAMEK). The SAMEK method avoids some unnecessary motion estimations in units with less probability of being referenced. These units can be recognized by the two rules of SAMEK method, namely ZeroCase and DecreaseCase. These two rules are summarized by statistically analyzing the relationships between each PU and its references. The experimental results demonstrate that our method can save up to 9.5% encoding time (averagely 6.87%) with negligible rate-distortion losses (averagely 0.006 dB) when compared with HEVC encoder with Test Zone search (TZ search) enabled. Li Yu 0004, Jimin Xiao, Tammam Tillo, Ce Zhu |
ICIP | 4 |
| 2015 | Inter-frame dependent rate-distortion optimization using lagrangian multiplier adaptionabstractIt is known that, in the current hybrid video coding structure, spatial and temporal prediction techniques are extensively used which introduce strong dependency among coding units. Such dependency poses a great challenge to perform a global rate-distortion optimization (RDO) when encoding a video sequence. RDO is usually performed in a way that coding efficiency of each coding unit is optimized independently without considering dependeny among coding units, leading to a suboptimal coding result for the whole sequence. In this paper, we investigate the inter-frame dependent RDO, where the impact of coding performance of the current coding unit on that of the following frames is considered. Accordingly, an inter-frame dependent rate-distortion optimization scheme is proposed and implemented on the newest video coding standard High Efficiency Video Coding (HEVC) platform. Experimental results show that the proposed scheme can achieve about 3.19% BD-rate saving in average over the state-of-the-art HEVC codec (HM15.0) in the low-delay B coding structure, with no extra encoding time. It obtains a significantly higher coding gain than the multiple QP (±3) optimization technique which would greatly increase the encoding time by a factor of about 6. Coupled with the multiple QP optimization, the proposed scheme can further achieve a higher BD-rate saving of 5.57% and 4.07% in average than the HEVC codec and the multiple QP optimization enabled HEVC codec, respectively. Shuai Li 0005, Ce Zhu, Yanbo Gao, Yimin Zhou 0001, Frédéric Dufaux, Ming-Ting Sun |
ICME | 2 |
| 2015 | Parameter-free view synthesis distortion model with application to depth video codingabstractDepth coding in 3D video is unique in that its quality is measured by the view synthesis distortion (VSD) rather than its own depth distortion, which further complicates the coding optimization as VSD is related to both the depth and texture quality. We propose a parameter-free VSD model to directly estimate the impact of depth errors on VSD on the small block basis, given its depth distortion. The proposed model is incorporated into rate-distortion optimization of depth coding, by adapting the Lagrange multiplier and optimizing the selection of quantization parameters. Simulation results show that our proposed scheme improves the BDPSNR and BDBR by 0.6dB PSNR and 20% bits saving on average compared with H.264 coding standard in depth coding, while keeping the implementation easy. Meng Yang 0002, Ce Zhu, Xuguang Lan, Nanning Zheng 0001 |
ISCAS | 2 |
| 2015 | Reducing Wedgelet lookup table size with down-sampling for depth map coding in 3D-HEVCabstractIn 3D-HEVC, a new coding standard of 3D video, Explicit Wedgelet mode within Depth Modelling Modes (DMM) partitions a depth block into two non-rectangular regions using a Wedgelet pattern selected from a Wedgelet lookup table. The Wedgelet lookup table is constructed during both encoder and decoder initialization for each block size ranging from 4×4 to 16×16, while the Wedgelet lookup table consists of a large number of Wedgelet bi-patterns, which involves relatively high computational complexity and may cause cache storing problem. In this paper, we propose a down-sampling method to reduce the number of Wedgelet patterns in the Wedgelet lookup table. The experimental results show our proposed down-sampling method can achieve storage size reduction by 27.8% with negligible 0.03% and 0.05% coding loss in both configurations of Common Test Conditions (CTC) and All Intra (AI), respectively. Shuying Ma, Ce Zhu, Yongbing Lin, Jianhua Zheng |
MMSP | 3 |
| 2015 | Down-/up-sampling based depth coding by utilizing interview correlationabstractIn depth-based 3D video compression, a framework that down-samples a depth map before encoding and up-samples the depth map at the decoder has shown more high coding efficiency, where the down- and up-sampling components are key to the coding performance. In this paper, a down-/up-sampling scheme for 3D video coding is developed by making good use of interview correlation. At the encoder, to obtain a uniform distribution of sampled pixels across views, an interlaced down-sampling approach is applied among different views. Then, at the decoder, a three-step up-sampling approach is proposed accordingly. First, sampled pixels in one view are warped to the to-be-up-sampled view based on their disparity vectors. Second, in order not to introduce new values, the warped pixel is further refined or updated by one of its nearest four sampled pixels. Finally, the unknown pixels left are filled with a known pixel around it. Among all the three steps, a winner-takes-all approach is employed to select a pixel with the largest weight which is determined by considering the depth similarity, texture similarity or geometric closeness. Experimental results demonstrate that the proposed scheme can achieve better down-/up-sampling result and improved the coding efficiency, compared with the conventional algorithms, especially for a large sampling scale. Jianjun Song, Ce Zhu |
MMSP | 4 |
| 2015 | Fast crowd density estimation with convolutional neural networks
Pei Xu 0009, Xudong Li 0001, Qihe Liu, Mao Ye 0001, Ce Zhu |
Eng. Appl. Artif. Intell. | 6 |
| 2015 | Exploiting entropy masking in perceptual graphic rendering
Lu Dong 0001, Yuming Fang 0001, Weisi Lin, Chenwei Deng, Ce Zhu, Seah Hock Soon |
Signal Process. Image Commun. | 5 |
| 2015 | Depth Coding Based on Depth-Texture Motion and Structure SimilaritiesabstractThis paper addresses high performance depth coding in 3D video by making good use of its coded texture video counterpart. The relationship between the depth and its associated texture video in terms of coding mode and motion vector is carefully examined. Our statistical study suggests that the skip-coding mode and its associated motion vectors in the coded texture can be shared for depth coding by saving bit rate at the cost of little increase of distortion, which subsequently results in a nonsequential coding of the depth map. In this sense, coding/prediction of a block can be performed using the skip-coded blocks below and right, which are not available in the conventional sequential coding, thus producing the so-called omnidirectional blocks predicted in the intra-coding by making the best use of (at most) four neighboring blocks. Moreover, in view of the depth-texture structure similarity, a depth-texture cooperative clustering-based prediction method is proposed for cluster-based depth prediction in the intra-coding, which exploits the structure similarity for the current coding block and its neighboring pixels around the block. On the other hand, some large prediction errors may be present for the depth-texture misaligned pixels, which may greatly compromise the coding performance. To deal with these large residuals induced by the depth-texture misalignment, a simple yet effective detection and rectification approach is incorporated in the proposed depth coding scheme. Experimental results show that our proposed depth coding scheme achieves superior rate-distortion performance compared with other relevant coding methods. Jianjun Lei 0001, Shuai Li 0005, Ce Zhu, Ming-Ting Sun, Chunping Hou |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | Scalable Bit Allocation Between Texture and Depth Views for 3-D Video Streaming Over Heterogeneous NetworksabstractIn the multiview video plus depth (MVD) coding format, both texture and depth views are jointly compressed to represent the 3-D video content. The MVD format enables synthesis of virtual views through depth-image-based rendering; hence, distortion in the texture and depth views affects the quality of the synthesized virtual views. Bit allocation between texture and depth views has been studied with some promising results. However, to the best of our knowledge, most of the existing bit-allocation methods attempt to allocate a fixed amount of total bit rate between texture and depth views; that is, to select appropriate pair of quantization parameters for texture and depth views to maximize the synthesized view quality subject to a fixed total bit rate. In this paper we propose a scalable bit-allocation scheme, where a single ordering of texture and depth packets is derived and used to obtain optimal bit allocation between texture and depth views for any total target rates. In the proposed scheme, both texture and depth views are encoded using the quality scalable coding method; that is, medium grain scalable (MGS) coding of the Scalable Video Coding (SVC) extension of the Advanced Video Coding (H.264/AVC) standard. For varying target total bit rates, optimal bit truncation points for both texture and depth views can be obtained using the proposed scheme. Moreover, we propose to order the enhancement layer packets of the H.264/SVC MGS encoded depth view according to their contribution to the reduction of the synthesized view distortion. On one hand, this improves the depth view packet ordering when considered the rate-distortion performance of synthesized views, which is demonstrated by the experimental results. On the other hand, the information obtained in this step is used to facilitate optimal bit allocation between texture and depth views. Experimental results demonstrate the effectiveness of the proposed scalable bit-allocation scheme for texture and depth views. Jimin Xiao, Miska M. Hannuksela, Tammam Tillo, Moncef Gabbouj, Ce Zhu, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2015 | Silhouette Analysis for Human Action Recognition Based on Supervised Temporal t-SNE and Incremental LearningabstractThis paper develops a human action recognition method for human silhouette sequences based on supervised temporal t-stochastic neighbor embedding (ST-tSNE) and incremental learning. Inspired by the SNE and its variants, ST-tSNE is proposed to learn the underlying relationship between action frames in a manifold, where the class label information and temporal information are introduced to well represent those frames from the same action class. As to the incremental learning, an important step for action recognition, we introduce three methods to perform the low-dimensional embedding of new data. Two of them are motivated by local methods, locally linear embedding and locality preserving projection. Those two techniques are proposed to learn explicit linear representations following the local neighbor relationship, and their effectiveness is investigated for preserving the intrinsic action structure. The rest one is based on manifold-oriented stochastic neighbor projection to find a linear projection from high-dimensional to low-dimensional space capturing the underlying pattern manifold. Extensive experimental results and comparisons with the state-of-the-art methods demonstrate the effectiveness and robustness of the proposed ST-tSNE and incremental learning methods in the human action silhouette analysis. Jian Cheng 0003, Haijun Liu 0001, Feng Wang 0015, Hongsheng Li 0001, Ce Zhu |
IEEE Trans. Image Process. | 5 |
| 2014 | Multipath Routing of Multiple Description Coded Images in Wireless Networks
Yuanyuan Xu 0001, Ce Zhu, Lu Yu 0003 |
J. Comput. Sci. Technol. | 2 |
| 2014 | A robust elastic net approach for feature learning
Ling Wang 0013, Hong Cheng 0002, Zicheng Liu 0001, Ce Zhu |
J. Vis. Commun. Image Represent. | 4 |
| 2014 | Pixel-Based Inter Prediction in Coded Texture Assisted Depth CodingabstractThis letter presents a pixel-based motion estimation scheme assisted with the coded texture video for depth inter-prediction, in view of motion similarity between depth and texture video. The proposed scheme can achieve higher inter-prediction gain without transmitting any motion vector in the pixel-based motion estimation. Coupled with depth-texture structure similarity, the inter prediction method is further extended to an integrated prediction approach by making use of both intra and inter information. Experimental results show that our proposed method achieves superior rate-distortion performance. Shuai Li 0005, Jianjun Lei 0001, Ce Zhu, Lu Yu 0003, Chunping Hou |
IEEE Signal Process. Lett. | 3 |
| 2013 | To exploit uncertainty masking for adaptive image renderingabstractFor high-quality image rendering using Monte Carlo methods, a large number of samples are required to be computed for each pixel. Adaptive sampling aims to decrease the total number of samples by concentrating samples on difficult regions. However, existing adaptive sampling schemes haven't fully exploited the potential of image regions with complex structures to the reduction of sample numbers. To solve this problem, we propose to exploit uncertainty masking in adaptive sampling. Experimental results show that incorporation of uncertainty information leads to significant sample reduction and therefore time-savings. Lu Dong 0001, Weisi Lin, Chenwei Deng, Ce Zhu, Seah Hock Soon |
ISCAS | 4 |
| 2013 | Generalized Gradient Vector Flow for Snakes: New Observations, Analysis, and ImprovementabstractSnakes, or active contours, have been widely used in image processing applications. An external force for snakes called gradient vector flow (GVF) attempts to address traditional snake problems of initialization sensitivity and poor convergence to concavities, while generalized GVF (GGVF) aims to improve GVF snake convergence to long and thin indentations (LTIs). In this paper, we find and show that both GVF and GGVF snakes essentially yield the same performance in capturing LTIs of odd widths, and generally neither can converge to even-width LTIs. Based on a thorough investigation of the GVF and GGVF fields within the LTI during their iterative processes, we identify the crux of the convergence problem, and accordingly propose a novel external force termed as component-normalized GGVF (CN-GGVF) to eliminate the problem. CN-GGVF is obtained by normalizing each component of initial GGVF vectors with respect to its own magnitude. Experimental results and comparisons against GGVF snakes show that the proposed CN-GGVF snakes can capture LTIs regardless of odd or even widths with a remarkably faster convergence speed, while preserving other desirable properties of GGVF snakes with lower computational complexity in vector normalization. Lunming Qin, Ce Zhu, Yao Zhao 0001, Huihui Bai 0001, Huawei Tian |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | End-to-End Rate-Distortion Optimized Description Generation for H.264 Multiple Description Video CodingabstractIn this paper, H.264/AVC primary and redundant slices interlaced multiple description video coding (PRSI-MDVC) is studied due to its high coding efficiency and ease of constructing multiple descriptions by interleaving primary and redundant slices, where the problem of optimal description generation in the rate-distortion sense is addressed. Optimal description generation requires minimization of end-to-end distortion consisting of both source coding distortion determined by quality of primary slices and channel distortion associated with redundant slice coding, subject to a rate constraint. The relevant existing works on the description generation mainly focus on the estimation of channel distortion to determine the amount of inserted redundancy (quality of redundant slices) but ignore the source distortion estimation for the quality of primary slices, thus a comprehensive end-to-end rate-distortion optimization is still unavailable. In this paper, we attack the minimization of the end-to-end distortion by fully exploring temporal coding dependency. Specifically, on one hand, a most recently developed source distortion temporal propagation model is employed to determine coding options of primary slices in the PRSI-MDVC. On the other hand, channel distortion estimation is mainly concerned with the mismatch error estimation when primary slices are lost. Unlike the existing channel distortion estimation approach under the asymptotic fine quantization assumption which is not valid in most practical cases (e.g., at low or medium coding rates), we develop a novel and more feasible estimation scheme, based on which coding parameters of the redundant slices can be better determined. Simulation results show the effectiveness of the proposed frame-level rate-distortion optimized description generation scheme compared with the relevant approaches. Yuanyuan Xu 0001, Ce Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2012 | Source Distortion Temporal Propagation Model for Motion Compensated Video Coding OptimizationabstractRate-distortion optimization (RDO) is widely employed to maximize coding efficiency in hybrid video coding. Due to an extensive use of spatial-temporal predictions in video coding, a global RDO problem becomes so complex that the processing of each coding unit (e.g. a macro block) is dependent and entangled each other. Usually the RDO is simply performed for each coding unit individually and independently, thus compromising the RDO performance significantly. In this paper, we examine temporal dependent RDO by developing a novel source distortion temporal propagation (SDTP) model, which takes into account the influence of a current coding macro block to future macro blocks in its temporal propagation chain. Accordingly estimations of the influence and the corresponding Lagrange multiplier for the proposed SDTP-based RDO are obtained. Experimental results show consistently a significant coding gain of the proposed scheme over H.264/AVC. Tianwu Yang, Ce Zhu, Xiaojiu Fan, Qiang Peng |
ICME | 2 |
| 2012 | Video organization: Near-Duplicate Video clusteringabstractIt is not uncommon to see several videos of almost identical content on the internet. These near duplicates, coupled with the sheer number of videos, pose a big challenge to the effective organization of video clips online. We propose an adaptive classification approach to detect near-duplicate versions, and an integrated voting strategy to group clusters and to elect a representative for each cluster. Our proposed methods are based on our observation that near-duplicate videos usually span a small, albeit variable area in the feature space, while videos of different contents are scattered far apart. The classification method aims to select a suitable threshold by maximizing the margin for each video sequence in the similarity space, and the voting scheme focuses on merging subsets with mutual information based on neighbor information and inverted indices. Experimental results on an unconstrained web dataset including over 10000 videos demonstrate the efficacy of the proposed methods. Tzu-Yi Hung, Ce Zhu, Gao Yang 0001, Yap-Peng Tan |
ISCAS | 2 |
| 2012 | Reducing location map in prediction-based difference expansion for reversible image data embedding
Minglei Liu, Seah Hock Soon, Ce Zhu, Weisi Lin, Feng Tian 0006 |
Signal Process. | 3 |
| 2012 | Multi-description multipath video streaming in wireless ad hoc networks
Yuanyuan Xu 0001, Ce Zhu |
Signal Process. Image Commun. | 2 |
| 2012 | Multiple description coded video streaming in peer-to-peer networks
Yuanyuan Xu 0001, Ce Zhu, Wenjun Zeng 0001, Xue Jun Li |
Signal Process. Image Commun. | 2 |
| 2012 | Robust Image Hashing Based on Random Gabor Filtering and Dithered Lattice Vector QuantizationabstractIn this paper, we propose a robust-hash function based on random Gabor filtering and dithered lattice vector quantization (LVQ). In order to enhance the robustness against rotation manipulations, the conventional Gabor filter is adapted to be rotation invariant, and the rotation-invariant filter is randomized to facilitate secure feature extraction. Particularly, a novel dithered-LVQ-based quantization scheme is proposed for robust hashing. The dithered-LVQ-based quantization scheme is well suited for robust hashing with several desirable features, including better tradeoff between robustness and discrimination, higher randomness, and secrecy, which are validated by analytical and experimental results. The performance of the proposed hashing algorithm is evaluated over a test image database under various content-preserving manipulations. The proposed hashing algorithm shows superior robustness and discrimination performance compared with other state-of-the-art algorithms, particularly in the robustness against rotations (of large degrees). Yuenan Li 0001, Zheming Lu 0001, Ce Zhu, Xiamu Niu |
IEEE Trans. Image Process. | 3 |
| 2011 | Corrigendum to "Adaptive coset partition for distributed video coding" [Signal Processing 90 (2010) 2480-2486]
Ce Zhu |
Signal Process. | 2 |
| 2011 | Binocular Just-Noticeable-Difference Model for Stereoscopic ImagesabstractConventional 2-D Just-Noticeable-Difference (JND) models measure the perceptible distortion of visual signal based on monocular vision properties by presenting a single image for both eyes. However, they are not applicable for stereoscopic displays in which a pair of stereoscopic images is presented to a viewer's left and right eyes, respectively. Some unique binocular vision properties, e.g., binocular combination and rivalry, need to be considered in the development of a JND model for stereoscopic images. In this letter, we propose a binocular JND (BJND) model based on psychophysical experiments which are conducted to model the basic binocular vision properties in response to asymmetric noises in a pair of stereoscopic images. The first experiment exploits the joint visibility thresholds according to the luminance masking effect and the binocular combination of noises. The second experiment examines the reduction of visual sensitivity in binocular vision due to the contrast masking effect. Based on these experiments, the developed BJND model measures the perceptible distortion of binocular vision for stereoscopic images. Subjective evaluations on stereoscopic images validate of the proposed BJND model. Yin Zhao, Ce Zhu, Yap-Peng Tan, Lu Yu 0003 |
IEEE Signal Process. Lett. | 3 |
| 2011 | Frame Fusion for Video Copy DetectionabstractContent-based video copy detection is very important for copyright protection in view of the growing popularity of video sharing websites, which deals with not only whether a copy occurs in a query video stream but also where the copy is located and where the copy is originated from. While a lot of work has addressed the problem with good performance, less effort has been made to consider the copy detection problem in the case of a continuous query stream, for which precise temporal localization and some complex video transformations like frame insertion and video editing need to be handled. We attempt to attack the problem by presenting a frame fusion based copy detection approach, which converts video copy detection to frame similarity search and frame fusion under a temporal consistency assumption. Our work focuses mainly on the frame fusion stage due to its critical role in copy detection performance. The proposed frame fusion scheme is based on a Viterbi-like algorithm, comprising an online back-tracking strategy with three relaxed constraints. The experimental results show that the proposed approach achieves high localization accuracy in both the query stream and the reference database even when a query video stream undergoes some complex transformations, while achieving comparable performance compared with state-of-the-art copy detection methods. Shikui Wei, Yao Zhao 0001, Ce Zhu, Changsheng Xu, Zhenfeng Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2011 | Face Region Based Conversational Video CodingabstractFace regions are visual focuses in conversational video communications, thus better reconstruction quality of the regions of interest (ROI) is highly desired or necessary in the bandwidth-constrained conversational video coding. In this paper, we introduce an efficient motion based face detection method to identify face blocks in the first step, which can reduce computational complexity substantially without any loss in face detection results. Then an active contour model is applied to find face contours for more refined and compact face regions. Based on the well-located and compact face regions, facial feature priority based bit allocation is proposed for face ROI based conversational video coding. Experimental results demonstrate that the proposed face region based coding can considerably improve the coding results in the face regions, compared with two other relevant video coding schemes, in terms of objective rate-distortion performance as well as subjective visual quality. Bing Xiong 0005, Xiaojiu Fan, Ce Zhu, Xuan Jing, Qiang Peng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2011 | Video Quality Assessment Based on Measuring Perceptual Noise From Spatial and Temporal PerspectivesabstractVideo quality assessment (VQA) exploits important properties of the sophisticated human visual system (HVS). In this paper, we study a series of fundamental HVS characteristics for subjective video quality assessment, and incorporate them into a systematic framework to simulate subjective evaluation on impaired videos. Based on this framework, we develop a novel full-reference metric, namely, perceptual quality index (PQI). Specifically, the proposed PQI metric comprises four major modules: 1) visual performance equation for the foveal and extra-foveal vision based on the cortical magnification theory; 2) perceptible noise detection using a spatial-temporal just noticeable difference model, and its quantification in both spatial and temporal channels, considering the varying error sensitivity due to the contrast and motion masking effects; 3) instantaneous error summation with inhibition of weak local distortions, and quality degradation accumulation over time that models the visual persistence and recency effect; and 4) fusion of the spatial and temporal noise intensities into a perceptual quality index. Compared with some state-of-the-art VQA models, the PQI metric, which exploits multiple visual properties, measures video quality more accurately and reliably on two VQA databases. Yin Zhao, Lu Yu 0003, Ce Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2011 | Depth No-Synthesis-Error Model for View Synthesis in 3-D VideoabstractCurrently, 3-D Video targets at the application of disparity-adjustable stereoscopic video, where view synthesis based on depth-image-based rendering (DIBR) is employed to generate virtual views. Distortions in depth information may introduce geometry changes or occlusion variations in the synthesized views. In practice, depth information is stored in 8-bit grayscale format, whereas the disparity range for a visually comfortable stereo pair is usually much less than 256 levels. Thus, several depth levels may correspond to the same integer (or sub-pixel) disparity value in the DIBR-based view synthesis such that some depth distortions may not result in geometry changes in the synthesized view. From this observation, we develop a depth no-synthesis-error (D-NOSE) model to examine the allowable depth distortions in rendering a virtual view without introducing any geometry changes. We further show that the depth distortions prescribed by the proposed D-NOSE profile also do not compromise the occlusion order in view synthesis. Therefore, a virtual view can be synthesized losslessly if depth distortions follow the D-NOSE specified thresholds. Our simulations validate the proposed D-NOSE model in lossless view synthesis and demonstrate the gain with the model in depth coding. Yin Zhao, Ce Zhu, Lu Yu 0003 |
IEEE Trans. Image Process. | 2 |
| 2010 | Suppressing texture-depth misalignment for boundary noise removal in view synthesisabstractDuring view synthesis based on depth maps, also known as Depth-Image-Based Rendering (DIBR), annoying artifacts are often generated around foreground objects, yielding the visual effects that slim silhouettes of foreground objects are scattered into the background. The artifacts are referred as the boundary noises. We investigate the cause of boundary noises, and find out that they result from the misalignment between texture and depth information along object boundaries. Accordingly, we propose a novel solution to remove such boundary noises by applying restrictions during forward warping on the pixels within the texture-depth misalignment regions. Experiments show this algorithm can effectively eliminate most boundary noises and it is also robust for view synthesis with compressed depth and texture information. Yin Zhao, Dong Tian, Ce Zhu, Lu Yu 0003 |
PCS | 4 |
| 2010 | Adaptive coset partition for distributed video coding
Ce Zhu |
Signal Process. | 2 |
| 2009 | Joint Multiple Description Coding and Network Coding for Wireless Image MulticastabstractMultiple description coding (MDC) is an effective technique to combat transmission loss over unreliable lossy networks. Network coding which allows coding at the intermediate nodes in the network promises to increase throughput of the whole network and better utilizes network resources. To provide both robustness and efficiency for image multicast over wireless ad hoc networks, a scheme based on joint multiple description coding (MDC) and network coding is proposed in this paper. Multiple description lattice vector quantization (MDLVQ) is employed to encode an image at the source node, and different descriptions are transmitted on an individually or mixed base. Linear network coding is used to mix selected packets at the intermediate nodes provided that corresponding receivers can decode the mixed packet. At the destination nodes, received original packets and mixed packets can result in a reproduction with certain quality. Experimental results validate that the proposed scheme can benefit the image transmission with better reconstructed image quality, lower failure rate and less energy consumptions. Yuanyuan Xu 0001, Ce Zhu |
ICIG | 2 |
| 2009 | Enhancing Two-Stage Multiple Description Scalar QuantizationabstractIn this letter we consider enhancing coding performance of two-stage multiple description scalar quantization. An enhancement scheme is proposed to make the product of central and side distortions closer to the rate-distortion bound of multiple description coding under the high-resolution assumption. We show analytically that the second stage refinement information, i.e., the quantized residual errors, can be used to further reduce the side distortions apart from the central distortion, which is substantiated with a memoryless Gaussian source. Minglei Liu, Ce Zhu |
IEEE Signal Process. Lett. | 2 |
| 2009 | Forward Error Correction-Based 2-D Layered Multiple Description Coding for Error-Resilient H.264 SVC Video TransmissionabstractIn this paper, we propose a novel 2-D layered multiple description coding (2DL-MDC) for error-resilient video transmission over unreliable networks. The proposed 2DL-MDC scheme allocates multiple description sub-bitstreams of a 2-D scalable bitstream to two network paths with unequal loss rates. We formulate the 2-D scalable rate-distortion problem and derive the expected distortion for the proposed scheme. To minimize the end-to-end distortion given the total rate budget and packet loss probabilities, we need to optimally allocate source and channel rates for each hierarchical sublayer of the scalable bitstream. The conventional Lagrangian multiplier method can be utilized to solve this problem but with overwhelming computational complexity. Therefore, we consider the use of the genetic algorithm to solve the rate-distortion optimization problem. The simulation results verify that the proposed method is able to achieve significant performance gain as opposed to the conventional equal rate allocation method. Wei Xiang 0001, Ce Zhu, Chee Kheong Siew, Yuanyuan Xu 0001, Minglei Liu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | Efficient Block Matching Motion Estimation Using Multilevel Intra- and Inter-Subblock Features Subblock-Based SATDabstractBlock matching based motion estimation employs conventionally pixel-based sum of absolute differences (SAD) metric for distortion measure. Many fast matching mechanisms have been developed to lower the computational load for block matching, such as those exploiting block feature based or attribute-based SAD calculations. By examining a subblock-sum-based SAD measure, where the subblock-sums can be considered to be intra-subblock features, we propose both intra- and inter-subblock features based SAD measure for more effective block matching in terms of achieving better coding efficiency. Interestingly, the proposed feature based SAD measure can be interpreted as subblock-based sum of absolute Hadamard transformed differences (SATD). The new features can be constructed in a multilevel structure, which may provide a flexible and scalable means for a tradeoff between computation load and coding efficiency. Encoding results are shown to compare the proposed scheme against other relevant feature based SAD measures. Bing Xiong 0005, Ce Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | Multiple Description Video Coding Based on Hierarchical B PicturesabstractA general multiple description video coding (MDVC) framework based on hierarchical B pictures is proposed in this paper. Two or more descriptions are generated by employing the hierarchical B pictures of H.264/AVC scalable extension, where temporal-level-based key pictures are selected in a staggered way among different descriptions. Based on this hierarchical and staggered structure, inter-description redundancy control is studied to achieve a good central/side-distortion-rate tradeoff. Moreover, to better exploit multiple complementary descriptions, a linear combination of received descriptions is employed to optimize decoding results. This proposed MDVC framework is H.264/AVC-compliant for each temporal scalable description. Some existing temporal-splitting MDVC techniques can be considered as a degraded case in the proposed structure. Experimental results validate the effectiveness of the proposed design for MDVC. Ce Zhu, Minglei Liu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2009 | Transform-Exempted Calculation of Sum of Absolute Hadamard Transformed DifferencesabstractSum of absolute Hadamard transformed differences (SATD) is an important distortion metric applied in the latest video coding standard H.264/AVC, which is an alternative to the sum of absolute differences (SAD) to improve coding efficiency. However, the SATD requires more computation load due to the Hadamard transform involved. Inspired by a geometric interpretation of the simplest two-point SATD calculation, an efficient way to compute the SATD is proposed, which considers the joint effect of the Hadamard transform and the following SAD processing in the SATD calculation and thus enables the computation of SATD without performing the Hadamard transform separately. We further extend the two-point transform- exempted SATD (TE-SATD) computation scheme to four-point and 4 times 4 block SATD calculation. With the same coding performance, the proposed TE-SATD-based fast algorithms can save 38% and 17% operations in computing a 4 times 4 block SATD compared with the conventional SATD calculation and the fast Hadamard transform-based calculation, respectively. Ce Zhu, Bing Xiong 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2008 | Priority Encoding Transmission Based Multiple Description Video Coding over Packet Loss NetworkabstractIn this paper, we attempt to overcome the limitation of specific scalable video codec and apply FEC-MDC to a common video coder, such as the standard H.264. The proposed scheme is explained as follows. Firstly, according to motion vector changes, an original video sequence is divided into several sub-sequences as messages, so in each message better temporal correlation can be maintained for better estimation when information losses occur. Secondly, the standard H.264 encoder is used to encode the messages. Thirdly, based on priority encoding transmission, unequal protections are assigned in each message. Lastly, at the decoder, the segments whose priorities are not higher than the fraction of packets received can be recover totally. Huihui Bai 0001, Yao Zhao 0001, Ce Zhu |
DCC | 3 |
| 2008 | Index assignment for 3-description lattice vector quantization based on A2 lattice
Minglei Liu, Ce Zhu |
Signal Process. | 2 |
| 2008 | Two-Stage Diversity-Based Multiple Description Image CodingabstractIn this letter, a diversity-based two-description image coding scheme is firstly presented and analyzed, for which a two-stage coding is introduced to facilitate the tuning of central/side distortion tradeoff. By subsampling the central decoded errors from the first stage to constitute the second part for each description, respectively, we show that not only the central distortion but also the side distortion can be reduced with the second stage information. We further generalize the proposed scheme to 4-description coding. Experiment results demonstrate that the proposed scheme outperforms the state-of-the-art techniques in terms of both central and side coding performance. Chunyu Lin, Yao Zhao 0001, Ce Zhu |
IEEE Signal Process. Lett. | 3 |
| 2008 | Two-Description Image Coding With SteganographyabstractFor two-description image coding, a conventional scheme is to partition an image into two parts and then to produce each description by alternatively concatenating a finely coded bitstream of one part and a coarsely coded bitstream of the other part. This letter presents a new two-description image coding approach using steganography. Specifically, we propose forming each description by embedding (hiding) the coarsely coded part into the finely coded part based on a least-significant bit (LSB) steganographic method. In this way, the bit budget for the coarsely coded part in each description can be saved with little reconstruction degradation for the finely coded part if the embedding process is well designed. The experimental results substantiate the effectiveness of the proposed method. Ce Zhu |
IEEE Signal Process. Lett. | 2 |
| 2008 | A New Multiplication-Free Block Matching CriterionabstractIn block matching, the matching criterion plays a pivotal role in both matching accuracy and computational complexity. Mean squared error (MSE) is one most widely accepted benchmark for matching accuracy. To avoid multiplication operations for simpler implementation, the sum of absolute difference (SAD) is normally taken as a substitute to approximate MSE. In this paper, we first examine statistically the quantitative deviation of SAD from MSE in terms of maximum and average deviations. To minimize the average deviation, a new measure, namely, weighted sum of absolute difference (WSAD), is proposed in a two-element case first, for which an optimal weight is obtained theoretically. The weight is approximately adapted to 1/2 so that the weighting operation can be simplified to a shift operation, and the resulting multiplication-free WSAD also shows better matching accuracy with a smaller deviation from MSE than SAD. The two-element WSAD is further extended to a multi-element case in a multilevel pyramid structure. The proposed WSAD is experimentally validated by applying it in block motion estimation for video coding, showing better rate-distortion performance than SAD. Bing Xiong 0005, Ce Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | A Joint Source-Channel Video Coding Scheme Based on Distributed Source CodingabstractRecently, several error resilient schemes have been proposed to tackle the error propagation problem in the motion-compensated predictive video coding based on a promising technique - distributed source coding (DSC). However, these schemes mainly apply the distributed source codes for channel error correction, while under-utilizing their capability for data compression. A channel-aware joint source-channel video coding scheme based on DSC is proposed to eliminate the error propagation problem in predictive video coding in a more efficient way. It is known that near Slepian-Wolf bound DSC is achieved using powerful channel codes, assuming the source and its reference (also known as side-information) are connected by a virtual error-prone channel. In the proposed scheme, the virtual and real error-prone channels are fused so that a unified single channel code is applied to encode the current frame thus accomplishing a joint source-channel coding. Our analysis of the rate efficiency in recovering error propagation shows that the joint scheme can achieve a lower rate compared with performing source and channel coding separately. Simulation results show that the number of bits used for recovering from error propagation can be reduced by up to 10% using the proposed scheme compared to Sehgal-Jagmohan-Ahuja's DSC-based error resilient scheme. Ce Zhu, Kim-Hui Yap |
IEEE Trans. Multim. | 2 |
| 2007 | Multiple Description Video Coding using Adaptive Temporal Sub-SamplingabstractMultiple description coding (MDC) is a promising alternative for robust transmission of information over non-prioritized and unpredictable networks. Especially, MDC has emerged as an attractive approach for video applications where retransmission is unacceptable or infeasible. In this paper, an effective MD video codec is designed based on pre-and post-processing of video sequences, without any modification to the source or channel codec. Considering different motion information of inter-frame, adaptive temporal sub-sampling is employed to the original video data. As a result, adaptive redundancy can make a better tradeoff between the reconstruction quality and compression efficiency. The experimental results exhibit better performance of the proposed scheme than other schemes. Huihui Bai 0001, Yao Zhao 0001, Ce Zhu |
ICME | 3 |
| 2007 | Multiple Description Video Coding using Hierarchical B PicturesabstractA generalized multiple description video coding based on hierarchical B pictures is proposed in this paper. Two descriptions are generated by employing the H.264/MPEG4-AVC codec conforming coding with hierarchical B pictures on the original video sequence. The only difference between the two descriptions lies in the selection of key pictures. Temporal scalability for each description is achieved benefiting from the high-efficiency hierarchical structure. More importantly, good central-side-distortion-rate tradeoff is obtained by our proposed scheme. Experimental results show that it is an effective and efficient design. Minglei Liu, Ce Zhu |
ICME | 2 |
| 2007 | Optimized Multiple Description Lattice Vector Quantization for Wavelet Image CodingabstractMultiple description (MD) coding is a promising alternative for robust transmission of information over non-prioritized and unpredictable networks. In this paper, an effective MD image coding scheme is introduced based on the MD lattice vector quantization (MDLVQ) for the wavelet transformed images. In view of the characteristics of wavelet coefficients in different frequency subbands, MDLVQ is applied in an optimized way, including an appropriate construction of wavelet coefficient vectors, the optimization of MDLVQ encoding parameters such as the choice of sublattice index values and the quantization accuracy for different subbands. More importantly, optimized side decoding is employed to predict lost information based on inter-vector correlation and an alternative transmission way for further reducing side distortion. Experimental results validate the effectiveness of the proposed scheme with better performance than some other tested MD image codecs including that based on optimized MD scalar quantization. Huihui Bai 0001, Ce Zhu, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2006 | Multiple Description Shifted Lattice Vector Quantization for Progressive Wavelet Image CodingabstractMultiple description (MD) coding is a promising alternative for robust transmission of information over non-prioritized and unpredictable networks. Furthermore, practical variable-bandwidth channels also require fine grain scalability of the descriptions (bit streams). In this paper, according to the geometrical structure and the special relationship of lattice vector quantizers, a MD quantizer called shifted lattice vector quantization (SLVQ) is employed in MD image coding to realize progressive transmission over unreliable channels. In view of the characteristics of wavelet coefficients in different frequency subbands, besides an appropriate construction of wavelet coefficient vectors, the algorithm of modified zerotree coding is also applied to improve compression performance. Experimental results validate the effectiveness of the proposed scheme with better performance than the other schemes based on MD scalar quantization for progressive transmission. Huihui Bai 0001, Yao Zhao 0001, Ce Zhu |
ICIP | 3 |
| 2006 | Index assignment design for three-description lattice vector quantizationabstractIn this paper, we propose a new index assignment scheme for the 3-description case, which aims to find a 3-tuple sublattice points to represent a fine lattice point. The design is made such that each edge of the triangle formed by the 3-tuple points needs to be as short as possible while the gravity center of the triangle is as close as possible to the fine lattice point. With a delicate sublattice partition, a well designed construction and mapping is developed to minimize the expected distortion. Minglei Liu, Ce Zhu, Xiaolin Wu 0001 |
ISCAS | 2 |
| 2006 | Distributed video coding based on adaptive binningabstractDistributed video coding enables video coding systems to shift coding complexity from encoder to decoder. In this paper, based on the analysis of different binning methods, we propose a distributed video coding scheme by adaptively combining two binning methods according to the estimated motion activity in video to minimize decoding errors. Simulation results show the improvement of the proposed method over the other binning algorithms Ce Zhu |
ISCAS | 2 |
| 2005 | Predictive fine granularity successive elimination for fast optimal block-matching motion estimationabstractGiven the number of checking points, the speed of block motion estimation depends on how fast the block matching is. In this paper, a new framework, fine granularity successive elimination (FGSE), is proposed for fast optimal block matching in motion estimation. The FGSE features providing a sequence of nondecreasing fine-grained boundary levels to reject a checking point using as little computation as possible, where block complexity is utilized to determine the order of partitioning larger sub-blocks into smaller subblocks in the creation of the fine-grained boundary levels. It is shown that the well-known successive elimination algorithm (SEA) and multilevel successive elimination algorithm (MSEA) are just two special cases in the FGSE framework. Moreover, in view that two adjacent checking points (blocks) share most of the block pixels with just one pixel shifting horizontally or vertically, we develop a scheme to predict the rejection level for a candidate by exploiting the correlation of matching errors between two adjacent checking points. The resulting predictive FGSE algorithm can further reduce computation load by skipping some redundant boundary levels. Experimental results are presented to verify substantial computational savings of the proposed algorithm in comparison with the SEA/MSEA. Ce Zhu, Wei-Song Qi, Wee Ser |
IEEE Trans. Image Process. | 1 |
| 2005 | An unequal packet loss resilience scheme for video over the InternetabstractWe present an unequal packet loss resilience scheme for robust transmission of video over the Internet. By jointly exploiting the unequal importance existing in different levels of syntax hierarchy in video coding schemes, GOP-level and Resynchronization-packet-level Integrated Protection (GRIP) is designed for joint unequal loss protection (ULP) in these two levels using forward error correction (FEC) across packets. Two algorithms are developed to achieve efficient FEC assignment for the proposed GRIP framework: a model-based FEC assignment algorithm and a heuristic FEC assignment algorithm. The model-based FEC assignment algorithm is to achieve optimal allocation of FEC codes based on a simple but effective performance metric, namely distortion-weighted expected length of error propagation, which is adopted to quantify the temporal propagation effect of packet loss on video quality degradation. The heuristic FEC assignment algorithm aims at providing a much simpler yet effective FEC assignment with little computational complexity. The proposed GRIP together with any of the two developed FEC assignment algorithms demonstrates strong robustness against burst packet losses with adaptation to different channel status. Xiaokang Yang 0001, Ce Zhu, Zhengguo Li, Xiao Lin 0001, Nam Ling |
IEEE Trans. Multim. | 2 |
| 2004 | Reducing drift for FGS coding based on multiframe motion compensation [video coding]abstractFine granularity scalability (FGS) in MPEG-4 and some enhanced FGS coding techniques have been studied recently for video streaming. However, a problem of drifting may emerge due to the difference of reference frames used in the encoder and decoder for those enhanced FGS coding schemes. In this paper, we propose incorporating multiframe-based motion compensation into an FGS coding scheme to further improve the coding efficiency and error robustness. It has been shown that the multiple frames based FGS coding scheme has better coding efficiency than the conventional one-frame based FGS video coding, and more importantly, the new one can alleviate the drifting problem significantly. The proposed approach is implemented within a one-loop FGS coding scheme based on an H.26L framework. Ce Zhu, Lap-Pui Chau |
ICASSP (3) | 1 |
| 2004 | Efficient inner search for faster diamond search
Ce Zhu, Xiao Lin 0001, Lap-Pui Chau, Hock-Ann Ang, Choo-Yin Ong |
Signal Process. | 1 |
| 2004 | Enhanced hexagonal search for fast block motion estimationabstractFast block motion estimation normally consists of low-resolution coarse search and the following fine-resolution inner search. Most motion estimation algorithms developed attempt to speed up the coarse search without considering accelerating the focused inner search. On top of the hexagonal search method recently developed, an enhanced hexagonal search algorithm is proposed to further improve the performance in terms of reducing number of search points and distortion, where a novel fast inner search is employed by exploiting the distortion information of the evaluated points. Our experimental results substantially justify the merits of the proposed algorithm. Ce Zhu, Xiao Lin 0001, Lap-Pui Chau, Lai-Man Po |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2003 | Smooth constrained block matching criterion for motion estimationabstractIn this paper, a novel and efficient criterion for block matching motion estimation is presented. The proposed criterion is to enhance the conventional mean absolute difference (MAD) scheme with a new smoothness constraint on the residue block. The objective is to reduce the bit rate for encoding the residue image without any degradation of the reconstructed image quality. Simulation results show that by applying the new criterion in motion estimation, both the improvement in peak signal-to-noise ratio (PSNR) and the reduction in bit rate up to 4.3% can be achieved compared to MAD. Xuan Jing, Ce Zhu, Lap-Pui Chau |
ICASSP (3) | 2 |
| 2003 | A fast octagon-based search algorithm for motion estimation
Lap-Pui Chau, Ce Zhu |
Signal Process. | 2 |
| 2003 | Smooth constrained motion estimation for video coding
Xuan Jing, Ce Zhu, Lap-Pui Chau |
Signal Process. | 2 |
| 2003 | Unequal loss protection for robust transmission of motion compensated video over the internet
Xiaokang Yang 0001, Ce Zhu, Zhengguo Li, Xiao Lin 0001, Zhengguo Feng, Si Wu 0004, Nam Ling |
Signal Process. Image Commun. | 2 |
| 2003 | A unified architecture for real-time video-coding systemsabstractThis paper presents a unified architecture for a live video over the Internet with emphasis on solving some challenging problems such as network bandwidth adaptation for rate and congestion, loss packet recovery, joint source and channel coding, and packetization. In our architecture, a time-varying bit rate for the source coding and time-varying ratios for the channel coding are simultaneously computed by a new congestion-control protocol. An adaptive rate-control scheme is then proposed to calculate quantization parameters and to determine the number of skipping frames corresponding to the bit rate. An adaptive unequal error-control scheme is also provided to protect the bitstream. Furthermore, a simple and MPEG-4 standard compatible algorithm is designed to packetize generated bitstream at the SyncLayer by using the existing resynchronization marker approach. With the proposed architecture, the coding efficiency and the robustness of the whole system are improved greatly. Zhengguo Li, Ce Zhu, Nam Ling, Xiaokang Yang 0001, Genan Feng, Si Wu 0004, Feng Pan 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2002 | Unequal error protection for motion compensated video streaming over the InternetabstractThis paper presents an unequal error protection scheme for motion compensated video over the Internet. The forward error correction (FEC) codes are optimally assigned to different frames in group of picture (GOP) by exploiting the temporal dependency among frames. To achieve optimal allocation of FEC assigned in GOP, we propose a performance criterion for the measurement of allocation, namely the expected length of error propagation (ELEP), which makes sense intuitively, as fewer frames corrupted implies better quality of reconstruction. Experiment results show the proposed scheme is robust to burst packet loss in the Internet. More importantly, graceful degradation of video quality is achieved by the proposed scheme as the packet loss probability of an Internet connection increases. Xiaokang Yang 0001, Ce Zhu, Zhengguo Li, Genan Feng, Si Wu 0004, Nam Ling |
ICIP (2) | 2 |
| 2002 | Degressive error protection algorithm for MPEG-4 FGS video streamingabstractThis paper presents a novel degressive error protection (DEP) algorithm adaptive to time varying packet loss rate for robustly transporting MPEG-4 fine granularity scalability (FGS) bit-stream over the packet erasure networks. By exploiting the embedded nature of FGS enhancement-layer bit-stream, a rate-distortion (R-D) optimization framework is developed to optimally partition the FGS enhancement-layer bit-stream into blocks and then to apply DEP to blocks by forward error correction (FEC). Experimental results show that this algorithm achieves graceful degradation of video quality in a wide range of packet loss rates. Xiaokang Yang 0001, Ce Zhu, Zhengguo Li, Genan Feng, Si Wu 0004, Nam Ling |
ICIP (3) | 2 |
| 2002 | A router based unequal error control scheme for video over the InternetabstractThis paper presents a router based unequal error protection scheme for video over the Internet. The proposed scheme classifies the whole bitstream into two priorities according to their importance and packetizes them into packets with two priorities. To reduce the loss ratio of video packets with higher priority, an active queue management scheme is designed at each router to provide more chance for these packets to transmit over the Internet. Compared to an end-to-end based approach, in which some redundancy packets are generated to protect the packets with higher priority, our proposed scheme utilizes the whole bandwidth for the source coding. Therefore, the final picture quality is improved. Xiaokang Yang 0001, Ce Zhu, Zhengguo Li, Genan Feng, Si Wu 0004, Nam Ling, Feng Pan 0002 |
ICIP (2) | 2 |
| 2002 | Adaptive unequal error control for video over the InternetabstractWe propose an adaptive unequal error protection scheme to help protect a video stream over the Internet. The data partition approach and the resynchronization marker are applied to generate two types of packets with different importance. Our scheme is very adaptive to network traffic conditions and can be easily implemented via software. With the proposed scheme, more protection is provided for the more important packets. The final quality can therefore be improved at a higher packet loss ratio. Zhengguo Li, Nam Ling, Ce Zhu, Xiaokang Yang 0001, Genan Feng, Si Wu 0004, Feng Pan 0002 |
ICME (1) | 3 |
| 2002 | A new subsampling-based predictive vector quantization for image coding
Ce Zhu |
Signal Process. Image Commun. | 1 |
| 2002 | A congestion control strategy for multipoint videoconferencingabstractWe formulate a congestion control problem for multipoint videoconferencing by introducing the concept of generalized fairness. A numerical method is provided to solve such a congestion control problem. The proposed method is much simpler than an earlier proposed method. This is desirable because congestion control is executed in real time. Zhengguo Li, Xiao Lin 0001, Ce Zhu, Xiaokang Yang 0001, Genan Feng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2002 | Hexagon-based search pattern for fast block motion estimationabstractIn block motion estimation, a search pattern with a different shape or size has a very important impact on search speed and distortion performance. A square-shaped search pattern is adopted in many popular fast algorithms. Recently, a diamond-shaped search pattern was introduced in fast block motion estimation and has exhibited a faster search speed. Based on an in-depth examination of the influence of the search pattern on speed performance, we propose a novel algorithm using a hexagon-based search pattern to achieve further improvement. The hexagon-based search pattern is investigated in comparison with diamond search pattern and demonstrates significant speedup gain over the diamond-based search. Analysis shows that a speed improvement rate of the hexagon-based search (HEXBS) algorithm over the diamond search (DS) algorithm can be over 80% for locating some motion vectors in certain scenarios. In short, the proposed HEXBS algorithm can find the same motion vector with fewer search points than the DS algorithm. Generally speaking, the larger the motion vector, the more search points the. HEXBS algorithm can save, which is further justified by experimental results. Ce Zhu, Xiao Lin 0001, Lap-Pui Chau |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2001 | A novel hexagon-based search algorithm for fast block motion estimationabstractIn block motion estimation, search patterns with different shape or size have a very important impact on search speed and distortion performance. In this paper, we propose a novel algorithm using a hexagon-based search (HEXBS) pattern for fast block motion estimation. The proposed HEXBS algorithm may find any motion vector with fewer search points than the diamond search (DS) algorithm. The speedup gain of the HEXBS method over the DS algorithm is more striking for finding large motion vectors. Experimental results substantially justify the fastest performance of the HEXBS algorithm compared with several other popular fast algorithms. Ce Zhu, Xiao Lin 0001, Lap-Pui Chau, Keng-Pang Lim, Hock-Ann Ang, Choo-Yin Ong |
ICASSP | 1 |
| 1999 | Vector Quantization with Minimax L Distortion for Image CodingabstractA technique for vector quantization which minimizes a maximal L/sub /spl infin// distortion is introduced in this paper. A centroid rule is obtained and a codebook design method is developed. The problem of minimizing the number of occurrences of quantization errors larger than a preselected threshold is further investigated. Compared to the methods that minimize an averaged L/sub /spl infin// distortion recently proposed by Mathews et al., the new approaches that minimize the maximal L/sub /spl infin// distortion achieve more desirable results. Ce Zhu, Yingbo Hua |
ICIP (1) | 1 |
| 1999 | Image vector quantization with minimax L∞ distortionabstractA new framework of vector quantization that minimizes maximal L/sub /spl infin// distortion is introduced. Based on the concept of minimax L/sub /spl infin// distortion, a new centroid rule and a new codebook design method are developed. The problem of minimizing the number of occurrences of quantization errors larger than a preselected threshold is also investigated. Compared to the methods that minimize an averaged L/sub /spl infin// distortion recently proposed by Mathews et al. (1997), the new methods with minimax L/sub /spl infin// distortion achieve more desirable results. Ce Zhu, Yingbo Hua |
IEEE Signal Process. Lett. | 1 |
| 1998 | Minimax partial distortion competitive learning for optimal codebook designabstractThe design of the optimal codebook for a given codebook size and input source is a challenging puzzle that remains to be solved. The key problem in optimal codebook design is how to construct a set of codevectors efficiently to minimize the average distortion. A minimax criterion of minimizing the maximum partial distortion is introduced in this paper. Based on the partial distortion theorem, it is shown that minimizing the maximum partial distortion and minimizing the average distortion will asymptotically have the same optimal solution corresponding to equal and minimal partial distortion. Motivated by the result, we incorporate the alternative minimax criterion into the on-line learning mechanism, and develop a new algorithm called minimax partial distortion competitive learning (MMPDCL) for optimal codebook design. A computation acceleration scheme for the MMPDCL algorithm is implemented using the partial distance search technique, thus significantly increasing its computational efficiency. Extensive experiments have demonstrated that compared with some well-known codebook design algorithms, the MMPDCL algorithm consistently produces the best codebooks with the smallest average distortions. As the codebook size increases, the performance gain becomes more significant using the MMPDCL algorithm. The robustness and computational efficiency of this new algorithm further highlight its advantages. Ce Zhu, Lai-Man Po |
IEEE Trans. Image Process. | 1 |
| 1995 | Analysis of learning vector quantization algorithms for pattern classificationabstractAlthough the family of LVQ algorithms have been widely used for pattern classification and have achieved a great success, the rigorous theoretical studies on the classification performance of LVQ algorithms have seldom been made. In this paper, the asymptotical performance of LVQ1, LVQ2 and LVQ2.1 algorithms have been studied thoroughly, and three significant conclusions have been achieved respectively. Furthermore, a simple modification scheme to LVQ2 algorithm has been developed and analyzed on the asymptotical performance, which can produce the optimal or nearly-optimal classifier in the stable equilibrium state for the classification problems with classes overlapping. Ce Zhu, Jun Wang 0002, Taijun Wang |
ICASSP | 1 |
| 1995 | Neural Network Approaches to fast and Low Rate Vector QuantizationabstractIn this paper, two codebook search methods and a coding scheme are proposed for fast and low rate vector quantization using the self-organizing feature maps (SOFM). Based on the topology preservation property of the SOFM, the search methods use the distance between adjacent input vectors to guide the codebook search process and to determine searching sequence of codevectors. The novel coding scheme, which can be considered as a vector version of delta modulation, eliminates the correlation buried in the source sequence and hence reduces the rate. Simulation results demonstrate the effectiveness of proposed methods and better performances than those obtained previously. Jun Wang 0002, Ce Zhu, Chenwu Wu, Zhenya He |
ISCAS | 2 |
| 1994 | A new competitive learning algorithm for vector quantizationabstractIn this paper, a new competitive learning algorithm based on the partial distortion theorem is proposed for the on-line vector quantizer design. The novel algorithm is called partial-distortion-equivalent competitive learning (PDECL) algorithm, which aims at making the partial distortions for each neuron (code-vector) be uniform to overcome the neuron underuse problem as well as to minimize the average distortion for the designed vector quantizer. Compared with the Kohonen learning algorithm (KLA), the frequency-sensitive competitive learning (FSCL) algorithm and the soft competition scheme (SCS) algorithm, the PDECL consistently shows the better performance than all of them and the LBG algorithm for the design of vector quantizers with different codebook sizes especially when the codebook size is large enough.> Ce Zhu, Lihua Li 0002, Zhenya He, Jun Wang 0002 |
ICASSP (2) | 1 |