Zhihai He

dblp:23/4027 · DBLP profile ↗
← Back
162ranked-venue papers
32as first author
47since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 128 · 29 first-author · 36 since 2021Artificial intelligence and machine learning · 34 · 22 since 2021Computer networks · 9 · 2 first-author · 1 since 2021Systems, architecture and hardware · 8 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Deeply-conditioned image compression via self-generated priors
Zhineng Zhao, Zhihai He, Zikun Zhou, Siwei Ma 0001, Yaowei Wang 0001
Neurocomputing2
2026 STD2Vformer: A Free-Form Spatiotemporal Forecasting Model
abstract
Spatiotemporal forecasting plays a vital role in modeling and managing complex dynamic systems, such as traffic networks, power grids, and industrial diagnostic systems. However, most existing spatiotemporal models focus primarily on improving prediction accuracy within fixed scenarios, while overlooking the challenge of adapting to dynamically changing forecasting demands. As a result, when prediction requirements shift, these models often need to be retrained to remain accurate—leading to resource inefficiency, production delays, and heightened safety risks. To address this issue, we propose a novel spatiotemporal prediction framework that effectively captures dynamic spatiotemporal dependencies and can be directly applied to spatiotemporal prediction tasks with varying prediction lengths and arbitrary starting points, requiring only a single training phase. Specifically, we introduce the spatiotemporal Date2Vec embedding method, which generates past and future timestamp embeddings by explicitly modeling intrinsic spatiotemporal relationships. Furthermore, we design a fusion module to model the direct mapping relationship between past and future timestamp embeddings, thereby enabling rapid adaptation to dynamic prediction demands. Extensive experiments on seven real-world public datasets show that our model exhibits superior adaptability across four distinct domains and higher predictive accuracy—achieving an average 4.55% improvement over the best-performing baseline on the fixed-horizon prediction task and an average 9.10% improvement on the free-form prediction task—while also providing lower computational complexity and faster inference compared with state-of-the-art methods.
Liwei Deng 0004, Hao Wang 0075, Junhao Tan, Xinhe Niu, Shiyao Zhang 0001, Zhihai He
IEEE Trans. Ind. Informatics7
2026 Audio-Driven Multi-Modal Unobtrusive Health Monitoring and Inference for Smart Eldercare at Home
abstract
As the aging population grows and more elderly individuals live independently, the demand for reliable, unobtrusive home health monitoring becomes increasingly important. Existing in-home health monitoring systems often face limitations such as privacy concerns, dependence on unreliable wearable devices, degraded accuracy in complex environments, and lack of continuous monitoring capability. To address these challenges, we propose a long-term home health monitoring system that primarily relies on audio sensing, supplemented by other noninvasive modalities. Our approach is able to accurately detect and recognize overlapping acoustic events with fine-grained temporal resolution, surpassing conventional audio-based methods for activity recognition. The system incorporates a transformer-based time-frequency fusion module and a category dynamic threshold strategy to improve detection performance under semi supervised conditions. Experiments on real-world dataset demonstrate that our method outperforms existing baselines, achieving PSDS$_{1}$, PSDS$_{2}$, and EB-F1 scores of 0.581, 0.930 and 55.1%, with improvements of 0.054, 0.019, and 2.3%, respectively. In addition, a 30 day field deployment involving 10 elderly participants confirms the robustness and practicality of the system for real-world applications. By allowing continuous passive monitoring of daily activities and abnormal acoustic events, our system has significant potentials for early detection of health risks, behavioral anomalies, and long-term wellness tracking in aging in place scenarios.
Xinhua Fan, Youming Li, Zhongchao Huang, Zhihai He
IEEE J. Biomed. Health Informatics4
2026 Training-Free Dual Hyperbolic Adapters for Better Cross-Modal Reasoning
abstract
Recent research in Vision-Language Models (VLMs) has significantly advanced our capabilities in cross-modal reasoning. However, existing methods suffer from performance degradation with domain changes or require substantial computational resources for fine-tuning in new domains. To address this issue, we develop a new adaptation method for large vision-language models, calledTraining-free Dual Hyperbolic Adapters(T-DHA). We characterize vision-language relationship between semantic concepts, which typically has a hierarchical tree structure, in the hyperbolic space instead of the traditional Euclidean space. Hyperbolic spaces exhibit exponential volume growth with radius, unlike the polynomial growth in Euclidean space. We find that this unique property is particularly effective for embedding hierarchical data structures using the Poincaré ball model, achieving significantly improved representation and discrimination power. Coupled with negative learning, it provides more accurate and robust classifications with fewer feature dimensions. Our extensive experimental results on various datasets demonstrate that the T-DHA method significantly outperforms existing state-of-the-art methods in few-shot image recognition and domain generalization tasks.
Yi Zhang 0109, Chun-Wun Cheng, Ke Yu 0004, Yushun Tang, Carola-Bibiane Schönlieb, Zhihai He, Angelica I. Avilés-Rivero
IEEE Trans. Multim.7
2025 Conditional Latent Coding with Learnable Synthesized Reference for Deep Image Compression
abstract
In this paper, we study how to synthesize a dynamic reference from an external dictionary to perform conditional coding of the input image in the latent domain and how to learn the conditional latent synthesis and coding modules in an end-to-end manner. Our approach begins by constructing a universal image feature dictionary using a multi-stage approach involving modified spatial pyramid pooling, dimension reduction, and multi-scale feature clustering. For each input image, we learn to synthesize a conditioning latent by selecting and synthesizing relevant features from the dictionary, which significantly enhances the model's capability in capturing and exploring image source correlation. This conditional latent synthesis involves a correlation-based feature matching and alignment strategy, comprising a Conditional Latent Matching (CLM) module and a Conditional Latent Synthesis (CLS) module. The synthesized latent is then used to guide the encoding process, allowing for more efficient compression by exploiting the correlation between the input image and the reference dictionary. According to our theoretical analysis, the proposed conditional latent coding (CLC) method is robust to perturbations in the external dictionary samples and the selected conditioning latent, with an error bound that scales logarithmically with the dictionary size, ensuring stability even with large and diverse dictionaries. Experimental results on benchmark datasets show that our new method improves the coding performance by a large margin (up to 1.2 dB) with a very small overhead of approximately 0.5% bits per pixel.
Yinda Chen, Zhihai He
AAAI4
2025 Cross-Modal Few-Shot Learning with Second-Order Neural Ordinary Differential Equations
abstract
We introduce SONO, a novel method leveraging Second-Order Neural Ordinary Differential Equations (Second-Order NODEs) to enhance cross-modal few-shot learning. By employing a simple yet effective architecture consisting of a Second-Order NODEs model paired with a cross-modal classifier, SONO addresses the significant challenge of overfitting, which is common in few-shot scenarios due to limited training examples. Our second-order approach can approximate a broader class of functions, enhancing the model's expressive power and feature generalization capabilities. We initialize our cross-modal classifier with text embeddings derived from class-relevant prompts, streamlining training efficiency by avoiding the need for frequent text encoder processing. Additionally, we utilize text-based image augmentation, exploiting CLIP’s robust image-text correlation to enrich training data significantly. Extensive experiments across multiple datasets demonstrate that SONO outperforms existing state-of-the-art methods in few-shot learning performance.
Yi Zhang 0109, Chun-Wun Cheng, Zhihai He, Carola-Bibiane Schönlieb, Yuyan Chen, Angelica I. Avilés-Rivero
AAAI4
2025 LatentEdit: Adaptive Latent Control for Consistent Semantic Editing
Weiming Chen 0001, Yushun Tang, Zhihai He
PRCV (6)4
2025 Spore: Spatio-Temporal Collaborative Perception and representation space disentanglement for remote heart rate measurement
Zexing Zhang, Huimin Lu 0004, Zhihai He, Adil Al-Azzawi, Songzhe Ma, Chenglin Lin
Neurocomputing3
2024 Cross-Constrained Progressive Inference for 3D Hand Pose Estimation with Dynamic Observer-Decision-Adjuster Networks
abstract
Generalization is very important for pose estimation, especially for 3D pose estimation where small changes in the 2D images could trigger structural changes in the 3D space. To achieve generalization, the system needs to have the capability of detecting estimation errors by double-checking the projection coherence between the 3D and 2D spaces and adapting its network inference process based on this feedback. Current pose estimation is one-time feed-forward and lacks the capability to gather feedback and adapt the inference outcome. To address this problem, we propose to explore the concept of progressive inference where the network learns an observer to continuously detect the prediction error based on constraints matching, as well as an adjuster to refine its inference outcome based on these constraints errors. Within the context of 3D hand pose estimation, we find that this observer-adjuster design is relatively unstable since the observer is operating in the 2D image domain while the adjuster is operating in the 3D domain. To address this issue, we propose to construct two sets of observers-adjusters with complementary constraints from different perspectives. They operate in a dynamic sequential manner controlled by a decision network to progressively improve the 3D pose estimation. We refer to this method as Cross-Constrained Progressive Inference (CCPI). Our extensive experimental results on FreiHAND and HO-3D benchmark datasets demonstrate that the proposed CCPI method is able to significantly improve the generalization capability and performance of 3D hand pose estimation.
Zhehan Kan, Xueting Hu, Ke Yu 0004, Zhihai He
AAAI5
2024 Concept-Guided Prompt Learning for Generalization in Vision-Language Models
abstract
Contrastive Language-Image Pretraining (CLIP) model has exhibited remarkable efficacy in establishing cross-modal connections between texts and images, yielding impressive performance across a broad spectrum of downstream applications through fine-tuning. However, for generalization tasks, the current fine-tuning methods for CLIP, such as CoOp and CoCoOp, demonstrate relatively low performance on some fine-grained datasets. We recognize the underlying reason is that these previous methods only projected global features into the prompt, neglecting the various visual concepts, such as colors, shapes, and sizes, which are naturally transferable across domains and play a crucial role in generalization tasks. To address this issue, in this work, we propose Concept-Guided Prompt Learning (CPL) for vision-language models. Specifically, we leverage the well-learned knowledge of CLIP to create a visual concept cache to enable conceptguided prompting. In order to refine the text features, we further develop a projector that transforms multi-level visual features into text features. We observe that this concept-guided prompt learning approach is able to achieve enhanced consistency between visual and linguistic modalities. Extensive experimental results demonstrate that our CPL method significantly improves generalization capabilities compared to the current state-of-the-art methods.
Yi Zhang 0109, Ce Zhang 0009, Ke Yu 0004, Yushun Tang, Zhihai He
AAAI5
2024 Window-Based Channel Attention for Wavelet-Enhanced Learned Image Compression
Bowen Hai, Yushun Tang, Zhihai He
ACCV (7)4
2024 Dual-Path Adversarial Lifting for Domain Shift Correction in Online Test-Time Adaptation
Yushun Tang, Shuoshuo Chen, Zhihe Lu, Xinchao Wang, Zhihai He
ECCV (67)5
2024 Conceptual Codebook Learning for Vision-Language Models
Yi Zhang 0109, Ke Yu 0004, Zhihai He
ECCV (77)4
2024 Learning Inference-Time Drift Sensor-Actuator for Domain Generalization
abstract
In machine learning tasks, models trained in the source domain often suffer from performance degradation in the target domain due to domain drift or distribution shift. In this paper, we explore the concept of sensor-actuator design in adaptive control to address this domain drift problem and develop a new approach, called learning inference-time drift sensor-actuator (LIDSA) for domain generalization. The drift sensor network consists of a constraint network and a data converter. The constraint network is learned to extract a set of constraints in the source domain and sense the domain drift by detecting the deviation from these constraints, called constraint error, which is correlated with the classification error. The data converter network then maps this constraint error into an effective guidance signal, which can guide the actuator network to adjust the feature to achieve improved discrimination power and better generalization performance. Our extensive experimental results demonstrate that the proposed LIDSA approach improves the performance of domain generalization over the baseline method.
Shuoshuo Chen, Yushun Tang, Zhehan Kan, Zhihai He
ICASSP4
2024 Domain-Conditioned Transformer for Fully Test-time Adaptation
Yushun Tang, Shuoshuo Chen, Jiyuan Jia, Yi Zhang 0109, Zhihai He
ACM Multimedia5
2024 Training-Free Feature Reconstruction with Sparse Optimization for Vision-Language Models
abstract
In this paper, we address the challenge of adapting vision-language models (VLMs) to few-shot image recognition in a training-free manner. We observe that existing methods are not able to effectively characterize the semantic relationship between support and query samples in a training-free setting. We recognize that, in the semantic feature space, the feature of the query image is a linear and sparse combination of support image features since support-query pairs are from the class and share the same small set of distinctive visual attributes. Motivated by this interesting observation, we propose a novel method called Training-free Feature ReConstruction with Sparse optimization (TaCo), which formulates the few-shot image recognition task as a feature reconstruction and sparse optimization problem. Specifically, we exploit the VLM to encode the query and support images into features. We utilize sparse optimization to reconstruct the query feature from the corresponding support features. The feature reconstruction error is then used to define the reconstruction similarity. Coupled with the text-image similarity provided by the VLM, our reconstruction similarity analysis accurately characterizes the relationship between support and query images. This results in significantly improved performance in few-shot image recognition. Our extensive experimental results on few-shot recognition demonstrate that our method outperforms existing state-of-the-art approaches by substantial margins.
Yi Zhang 0109, Ke Yu 0004, Angelica I. Avilés-Rivero, Jiyuan Jia, Yushun Tang, Zhihai He
ACM Multimedia6
2024 Learning to Adapt CLIP for Few-Shot Monocular Depth Estimation
abstract
Pre-trained Vision-Language Models (VLMs), such as CLIP, have shown enhanced performance across a range of tasks that involve the integration of visual and linguistic modalities. When CLIP is used for depth estimation tasks, the patches, divided from the input images, can be combined with a series of semantic descriptions of the depth information to obtain similarity results. The coarse estimation of depth is then achieved by weighting and summing the depth values, called depth bins, corresponding to the predefined semantic descriptions. The zero-shot approach circumvents the computational and time-intensive nature of traditional fully-supervised depth estimation methods. However, this method, utilizing fixed depth bins, may not effectively generalize as images from different scenes may exhibit distinct depth distributions. To address this challenge, we propose a few-shot-based method which learns to adapt the VLMs for monocular depth estimation to balance training costs and generalization capabilities. Specifically, it assigns different depth bins for different scenes, which can be selected by the model during inference. Additionally, we incorporate learnable prompts to preprocess the input text to convert the easily human-understood text into easily model-understood vectors and further enhance the performance. With only one image per scene for training, our extensive experiment results on the NYU V2 and KITTI dataset demonstrate that our method outperforms the previous state-of-the-art method by up to 10.6% in terms of MARE1.
Xueting Hu, Ce Zhang 0009, Yi Zhang 0109, Bowen Hai, Ke Yu 0004, Zhihai He
WACV6
2024 Cross-Modal Concept Learning and Inference for Vision-Language Models
Yi Zhang 0109, Ce Zhang 0009, Yushun Tang, Zhihai He
Neurocomputing4
2023 BDC-Adapter: Brownian Distance Covariance for Better Vision-Language Reasoning
Yi Zhang 0109, Ce Zhang 0009, Yushun Tang, Zhihai He
BMVC5
2023 Self-Correctable and Adaptable Inference for Generalizable Human Pose Estimation
abstract
A central challenge in human pose estimation, as well as in many other machine learning and prediction tasks, is the generalization problem. The learned network does not have the capability to characterize the prediction error, generate feedback information from the test sample, and correct the prediction error on the fly for each individual test sample, which results in degraded performance in generalization. In this work, we introduce a self-correctable and adaptable inference (SCAI) method to address the generalization challenge of network prediction and use human pose estimation as an example to demonstrate its effectiveness and performance. We learn a correction network to correct the prediction result conditioned by a fitness feedback error. This feedback error is generated by a learned fitness feedback network which maps the prediction result to the original input domain and compares it against the original input. Interestingly, we find that this self-referential feedback error is highly correlated with the actual prediction error. This strong correlation suggests that we can use this error as feedback to guide the correction process. It can be also used as a loss function to quickly adapt and optimize the correction network during the inference process. Our extensive experimental results on human pose estimation demonstrate that the proposed SCAI method is able to significantly improve the generalization capability and performance of human pose estimation.
Zhehan Kan, Shuoshuo Chen, Ce Zhang 0009, Yushun Tang, Zhihai He
CVPR5
2023 Neuro-Modulated Hebbian Learning for Fully Test-Time Adaptation
abstract
Fully test-time adaptation aims to adapt the network model based on sequential analysis of input samples during the inference stage to address the cross-domain performance degradation problem of deep neural networks. We take inspiration from the biological plausibility learning where the neuron responses are tuned based on a local synapse-change procedure and activated by competitive lateral inhibition rules. Based on these feed-forward learning rules, we design a soft Hebbian learning process which provides an unsupervised and effective mechanism for online adaptation. We observe that the performance of this feed-forward Hebbian learning for fully test-time adaptation can be significantly improved by incorporating a feedback neuromodulation layer. It is able to fine-tune the neuron responses based on the external feedback generated by the error backpropagation from the top inference layers. This leads to our proposed neuro-modulated Hebbian learning (NHL) method for fully test-time adaptation. With the unsupervised feed-forward soft Hebbian learning being combined with a learned neuromodulator to capture feedback from external responses, the source model can be effectively adapted during the testing process. Experimental results on benchmark datasets demonstrate that our proposed method can significantly improve the adaptation performance of network models and outperforms existing state-of-the-art methods.
Yushun Tang, Ce Zhang 0009, Shuoshuo Chen, Luziwei Leng, Qinghai Guo, Zhihai He
CVPR8
2023 Cross-Inferential Networks for Source-Free Unsupervised Domain Adaptation
abstract
One central challenge in source-free unsupervised domain adaptation (UDA) is the lack of an effective approach to evaluate the prediction results of the adapted network model in the target domain. To address this challenge, we propose to explore a new method called cross-inferential networks (CIN). Our main idea is that, when we adapt the network model to predict the sample labels from encoded features, we use these prediction results to construct new training samples with derived labels to learn a new examiner network that performs a different but compatible task in the target domain. Specifically, in this work, the base network model is performing image classification while the examiner network is tasked to perform relative ordering of triplets of samples whose training labels are carefully constructed from the prediction results of the base network model. Two similarity measures, cross-network correlation matrix similarity and attention consistency, are then developed to provide important guidance for the UDA process. Our experimental results on benchmark datasets demonstrate that our proposed CIN approach can significantly improve the performance of source-free UDA.
Yushun Tang, Qinghai Guo, Zhihai He
ICIP3
2023 Unsupervised Prototype Adapter for Vision-Language Models
Yi Zhang 0109, Ce Zhang 0009, Xueting Hu, Zhihai He
PRCV (1)4
2023 Contrastive Bayesian Analysis for Deep Metric Learning
abstract
Recent methods for deep metric learning have been focusing on designing different contrastive loss functions between positive and negative pairs of samples so that the learned feature embedding is able to pull positive samples of the same class closer and push negative samples from different classes away from each other. In this work, we recognize that there is a significant semantic gap between features at the intermediate feature layer and class labels at the final output layer. To bridge this gap, we develop a contrastive Bayesian analysis to characterize and model the posterior probabilities of image labels conditioned by their features similarity in a contrastive learning setting. This contrastive Bayesian analysis leads to a new loss function for deep metric learning. To improve the generalization capability of the proposed method onto new classes, we further extend the contrastive Bayesian loss with a metric variance constraint. Our experimental results and ablation studies demonstrate that the proposed contrastive Bayesian metric learning method significantly improves the performance of deep metric learning in both supervised and pseudo-supervised scenarios, outperforming existing methods by a large margin.
Shichao Kan, Zhiquan He, Yi-Gang Cen, Yang Li 0091, Vladimir Mladenovic, Zhihai He
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Cycle optimization metric learning for few-shot classification
Qifan Liu, Wenming Cao 0001, Zhihai He
Pattern Recognit.3
2023 Jointly-Learnt Networks for Future Action Anticipation via Self-Knowledge Distillation and Cycle Consistency
abstract
Future action anticipation aims to infer future actions from the observation of a small set of past video frames. In this paper, we propose a novel Jointly-learnt Action Anticipation Network (J-AAN) via Self-Knowledge Distillation (Self-KD) and cycle consistency for future action anticipation. In contrast to the current state-of-the-art methods which anticipate the future actions either directly or recursively, our proposed J-AAN anticipates the future actions jointly in both direct and recursive ways. However, when dealing with future action anticipation, one important challenge to address is the future’s uncertainty since multiple action sequences may come from or be followed by the same action. Training an action anticipation model with one-hot-encoded hard labels that assign zero probabilities to incorrect yet semantically similar actions may not handle the uncertain future. To address this challenge, we design a Self-KD mechanism to train our J-AAN, where the J-AAN gradually distills its own knowledge during the training to soften the hard labels to model the uncertainty on future action anticipation. Furthermore, we design a forward and backward action anticipation framework with our proposed J-AAN based on a cyclic consistency constraint. The forward J-AAN anticipates the future actions from the observed past actions, and the backward J-AAN verifies the anticipation of the forward J-AAN by anticipating the past actions from the anticipated future actions. The proposed method outperforms all the latest state-of-the-art action anticipation methods on the Breakfast, 50Salads, and EPIC-Kitchens-55 datasets. This project will be publicly available onhttps://github.com/MoniruzzamanMd/J-AAN.
Md. Moniruzzaman 0002, Zhaozheng Yin, Zhihai He, Ming C. Leu, Ruwen Qin
IEEE Trans. Circuits Syst. Video Technol.3
2022 Self-Constrained Inference Optimization on Structural Groups for Human Pose Estimation
Zhehan Kan, Shuoshuo Chen, Zhihai He
ECCV (5)4
2022 Coded Residual Transform for Generalizable Deep Metric Learning
abstract
A fundamental challenge in deep metric learning is the generalization capability of the feature embedding network model since the embedding network learned on training classes need to be evaluated on new test classes. To address this challenge, in this paper, we introduce a new method called coded residual transform (CRT) for deep metric learning to significantly improve its generalization capability. Specifically, we learn a set of diversified prototype features, project the feature map onto each prototype, and then encode its features using their projection residuals weighted by their correlation coefficients with each prototype. The proposed CRT method has the following two unique characteristics. First, it represents and encodes the feature map from a set of complimentary perspectives based on projections onto diversified prototypes. Second, unlike existing transformer-based feature representation approaches which encode the original values of features based on global correlation analysis, the proposed coded residual transform encodes the relative differences between the original features and their projected prototypes. Embedding space density and spectral decay analysis show that this multi perspective projection onto diversified prototypes and coded residual representation are able to achieve significantly improved generalization capability in metric learning. Finally, to further enhance the generalization performance, we propose to enforce the consistency on their feature similarity matrices between coded residual transforms with different sizes of projection prototypes and embedding dimensions. Our extensive experimental results and ablation studies demonstrate that the proposed CRT method outperform the state-of-the-art deep metric learning methods by large margins and improving upon the current best method by up to 4.28% on the CUB dataset.
Shichao Kan, Yixiong Liang, Min Li 0007, Yi-Gang Cen, Jianxin Wang 0001, Zhihai He
NeurIPS6
2022 Geometric machine learning: research and applications
Wenming Cao 0001, Canta Zheng, Zhiyue Yan, Zhihai He, Weixin Xie
Multim. Tools Appl.4
2022 Self-guided information for few-shot classification
Zhineng Zhao, Qifan Liu, Wenming Cao 0001, Deliang Lian, Zhihai He
Pattern Recognit.5
2022 Reciprocal Twin Networks for Pedestrian Motion Learning and Future Path Prediction
abstract
Modeling the moving behaviors and predicting the future paths of pedestrians, especially for those in complex scenes, remain a challenging problem in machine learning. We recognize that human motion trajectories, governed by social norms and constrained by physical structures of the surrounding environment, are both forward predictable and backward predictable. Motivated by this observation, we develop a new approach, calledreciprocal twin networks, for human trajectory learning and prediction. We design two networks, a forward prediction network to predict future trajectory from past observations and a backward prediction that performs the trajectory prediction backward in time. The backward prediction network serves as the inverse operation of the forward prediction network, forming a reciprocal constraint. During the training stage, this reciprocal constraint allows them to be jointly learned for accurate and robust human trajectory prediction. During the inference stage, we borrow the concept of adversarial attack of deep neural networks, which iteratively modifies the input of the network to match the given or forced network output, and develop a new method, calledreciprocal attack for matched prediction, to achieve accurate human trajectory prediction. Our experimental results on benchmark datasets demonstrate that our new method outperforms the state-of-the-art methods for human trajectory prediction.
Hao Sun 0024, Zhiqun Zhao, Zhaozheng Yin, Zhihai He
IEEE Trans. Circuits Syst. Video Technol.4
2022 Local Semantic Correlation Modeling Over Graph Neural Networks for Deep Feature Embedding and Image Retrieval
abstract
Deep feature embedding aims to learn discriminative features or feature embeddings for image samples which can minimize their intra-class distance while maximizing their inter-class distance. Recent state-of-the-art methods have been focusing on learning deep neural networks with carefully designed loss functions. In this work, we propose to explore a new approach to deep feature embedding. We learn a graph neural network to characterize and predict the local correlation structure of images in the feature space. Based on this correlation structure, neighboring images collaborate with each other to generate and refine their embedded features based on local linear combination. Graph edges learn a correlation prediction network to predict the correlation scores between neighboring images. Graph nodes learn a feature embedding network to generate the embedded feature for a given image based on a weighted summation of neighboring image features with the correlation scores as weights. Our extensive experimental results under the image retrieval settings demonstrate that our proposed method outperforms the state-of-the-art methods by a large margin, especially for top-1 recalls.
Shichao Kan, Yi-Gang Cen, Yang Li 0091, Vladimir Mladenovic, Zhihai He
IEEE Trans. Image Process.5
2022 Human Action Recognition by Discriminative Feature Pooling and Video Segment Attention Model
abstract
We Introduce a simple yet effective network that embeds a novel Discriminative Feature Pooling (DFP) mechanism and a novel Video Segment Attention Model (VSAM), for video-based human action recognition from both trimmed and untrimmed videos. Our DFP module introduces an attentional pooling mechanism for 3D Convolutional Neural Networks that attentionally pools 3D convolutional feature maps to emphasize the most critical spatial, temporal, and channel-wise features related to the actions within a video segment, while our VSAM ensembles these most critical features from all video segments and learns (1) class-specific attention weights to classify the video segments into the corresponding action categories, and (2) class-agnostic attention weights to rank the video segments based on their relevance to the action class. Our action recognition network can be trained from both trimmed videos in a fully-supervised way and untrimmed videos in a weakly-supervised way. For untrimmed videos with weak labels, our network learns attention weights without the requirement of precise temporal annotations of action occurrences in videos. Evaluated on the untrimmed video datasets of THUMOS14 and ActivityNet1.2, and trimmed video datasets of HMDB51, UCF101, and HOLLYWOOD2, our network achieves promising performance, compared to the latest state-of-the-art method. The implementation code is available athttps://github.com/MoniruzzamanMd/DFP-VSAM-Networks.
Md. Moniruzzaman 0002, Zhaozheng Yin, Zhihai He, Ruwen Qin, Ming C. Leu
IEEE Trans. Multim.3
2021 Spatial Assembly Networks for Image Representation Learning
abstract
It has been long recognized that deep neural networks are sensitive to changes in spatial configurations or scene structures. Image augmentations, such as random translation, cropping, and resizing, can be used to improve the robustness of deep neural networks under spatial transforms. However, changes in object part configurations, spatial layout of object, and scene structures of the images may still result in major changes in the their feature representations generated by the network, creating significant challenges for various visual learning tasks, including representation or metric learning, image classification and retrieval. In this work, we introduce a new learnable module, called spatial assembly network (SAN), to address this important issue. This SAN module examines the input image and performs a learned re-organization and assembly of feature points from different spatial locations conditioned by feature maps from previous network layers so as to maximize the discriminative power of the final feature representation. This differentiable module can be flexibly incorporated into existing network architectures, improving their capabilities in handling spatial variations and structural changes of the image scene. We demonstrate that the proposed SAN module is able to significantly improve the performance of various metric / representation learning, image retrieval and classification tasks, in both supervised and unsupervised learning scenarios.
Yang Li 0091, Shichao Kan, Jianhe Yuan, Wenming Cao 0001, Zhihai He
CVPR5
2021 Relative Order Analysis and Optimization for Unsupervised Deep Metric Learning
abstract
In unsupervised learning of image features without labels, especially on datasets with fine-grained object classes, it is often very difficult to tell if a given image belongs to one specific object class or another, even for human eyes. However, we can reliably tell if image C is more similar to image A than image B. In this work, we propose to explore how this relative order can be used to learn discriminative features with an unsupervised metric learning method. Instead of resorting to clustering or self-supervision to create pseudo labels for an absolute decision, which often suffers from high label error rates, we construct reliable relative orders for groups of image samples and learn a deep neural network to predict these relative orders. During training, this relative order prediction network and the feature embedding network are tightly coupled, providing mutual constraints to each other to improve overall metric learning performance in a cooperative manner. During testing, the predicted relative orders are used as constraints to optimize the generated features and refine their feature distance-based image retrieval results using a constrained optimization procedure. Our experimental results demonstrate that the proposed relative orders for unsupervised learning (ROUL) method is able to significantly improve the performance ofunsupervised deep metric learning.
Shichao Kan, Yi-Gang Cen, Yang Li 0091, Vladimir Mladenovic, Zhihai He
CVPR5
2021 Consistency-Sensitivity Guided Ensemble Black-Box Adversarial Attacks in Low-Dimensional Spaces
abstract
Black-box attacks aim to generate adversarial noise to fail the victim deep neural network in the black box. The central task in black-box attack method design is to estimate and characterize the victim model in the high-dimensional model space based on feedback results of queries submitted to the victim network. The central performance goal is to minimize the number of queries needed for successful at-tack. Existing attack methods directly search and refine the adversarial noise in an extremely high-dimensional space, requiring hundreds or even thousands queries to the victim network. To address this challenge, we propose to explore a consistency and sensitivity guided ensemble attack (CSEA) method in a low-dimensional space. Specifically, we estimate the victim model in the black box using a learned linear composition of an ensemble of surrogate models with diversified network structures. Using random block masks on the input image, these surrogate models jointly construct and submit randomized and sparsified queries to the victim model. Based on these query results and guided by a consistency constraint, the surrogate models can be trained using a very small number of queries such that their learned composition is able to accurately approximate the victim model in the high-dimensional space. The randomized and sparsified queries also provide important information for us to construct an attack sensitivity map for the input image, with which the adversarial attack can be locally refined to further increase its success rate. Our extensive experimental results demonstrate that our proposed approach significantly reduces the number of queries to the victim network while maintaining very high success rates, outperforming existing black-box attack methods by large margins.
Jianhe Yuan, Zhihai He
ICCV2
2021 Structure-Oriented Progressive Low-Rank Image Restoration for Defending Adversarial Attacks
abstract
Deep neural networks recognize objects by analyzing local image details and summarizing their information along the inference layers to derive the final decision. Because of this, they are prone to adversarial attacks. On the other hand, human eyes recognize objects based on their global structures and semantic cues, instead of local image textures. In this work, we propose to develop a structure-oriented progressive low-rank image completion method to remove unneeded texture details from the input images and shift the bias of deep neural networks towards global object structures and semantic cues. We formulate the problem into a low-rank matrix completion problem with progressively smoothed rank functions to avoid local minimums. Our experimental results demonstrate the proposed method is able to successfully remove the insignificant local image details while preserving important global object structures.
Zhiqun Zhao, Hengyou Wang, Hao Sun 0024, Wenming Cao 0001, Zhihai He
ICME5
2021 A Networked Social Virtual Reality Learning Environment Platform for Special Education
abstract
Delivering curriculum using desktop-based virtual learning environment (VLE) technologies in a collaborative group setting has been shown to reduce the social skill limitations of students with learning disabilities. However, the lack of the immersiveness and effective generalization of acquiring knowledge and skills among students remains a critical challenge in the interactive tools used in current VLEs. In this paper, we present a networked social virtual reality learning environment (VRLE) system viz., vSocial that has been redesigned based on iterative user feedback and developed in order to leverage the latest advances in integration of smart devices such as VR headsets for virtual content delivery. We describe a comparative study to evaluate technology trade-offs in the development process of transitioning from a VLE to a VRLE, from both technological and user (e.g., student/instructor) perspectives. Lastly, we outline open issues in using VRLEs which include: system complexity, emotion recognition, cybersickness and system sustainability.
Roland Oruche, Vaibhav Akashe, Samaikya Valluripally, Aniket Gulhane, Prasad Calyam, Janine Stichter, Zhihai He
LCN7
2021 Context-Aware Question-Answer for Interactive Media Experiences
abstract
Media content has become a primary source of information, entertainment, and even education. The ability to provide video content querying as well as interactive experiences is a new challenge. To this end, question answering (QA) systems such as Alexa and Google Assistant have become quite established in consumer markets but are limited to general information and lack context awareness. In this paper, we propose Context-QA, a light-weight context-aware QA framework, to provide QA experiences on multimedia content. The context awareness is achieved through our innovative Staged QA Controller algorithm that keeps the search for answers in the context most relevant to the question. Our evaluation results show that Context-QA improves the quality of the answers by up to 49% and uses up to 56% less time compared to the conventional QA model. Subjective tests show Context-QA improved results over conventional QA models, with 90% reporting enjoying this new media form.
Kyle Jorgensen, Zhiqun Zhao, Haohong Wang, Mea Wang, Zhihai He
IMX5
2021 Ensemble diversified learning for image classification with noisy labels
Ahmed Ahmed 0003, Hayder Yousif, Zhihai He
Multim. Tools Appl.3
2021 vSocial: a cloud-based system for social virtual reality learning environment applications in special education
Sai Shreya Nuguri, Prasad Calyam, Roland Oruche, Aniket Gulhane, Samaikya Valluripally, Janine Stichter, Zhihai He
Multim. Tools Appl.7
2021 A cascaded registration network RCINet with segmentation mask
Wenlan Zou, Wenming Cao 0001, Zhiquan He, Zhihai He
Neural Comput. Appl.5
2021 Learned Model Composition With Critical Sample Look-Ahead for Semi-Supervised Learning on Small Sets of Labeled Samples
abstract
In this work, we propose to push the performance limit of semi-supervised learning on very small sets of labeled samples by developing a new method called learned model composition with critical sample look-ahead (LMCS). Training efficient deep neural networks on much smaller sets of labeled samples is a challenging problem. With a small labeled set, the initial network suffers from low accuracy. Based on this error-prone network, the subsequent semi-supervised learning process will be fragile and unstable. To address this issue, we propose to introduce a look-ahead master model to identify the correct direction of model evolution to effectively guide the semi-supervised learning process of the student model. Specifically, our proposed LMCS method explores two major ideas. First, it introduces a new learned model composition structure so that we can compose a more efficient master network from student models of past iterations through a network learning process. Second, we develop a new method, called confined maximum entropy search, to discover new critical samples near the model decision boundary and provide the master model with look-ahead access to these samples to enhance its guidance capability. Our extensive experimental results demonstrate that the proposed LMCS method outperforms the state-of-the-art semi-supervised learning methods, especially on small sets of labeled samples. For example, on the CIFAR-10 dataset, with a very small set of 80 labeled samples, our method outperforms Google's MixMatch method, reducing the error rate by more than 10%.
Yang Li 0091, Shichao Kan, Wenming Cao 0001, Zhihai He
IEEE Trans. Circuits Syst. Video Technol.4
2021 Zero-Shot Learning to Index on Semantic Trees for Scalable Image Retrieval
abstract
In this study, we develop a new approach, called zero-shot learning to index on semantic trees (LTI-ST), for efficient image indexing and scalable image retrieval. Our method learns to model the inherent correlation structure between visual representations using a binary semantic tree from training images which can be effectively transferred to new test images from unknown classes. Based on predicted correlation structure, we construct an efficient indexing scheme for the whole test image set. Unlike existing image index methods, our proposed LTI-ST method has the following two unique characteristics. First, it does not need to analyze the test images in the query database to construct the index structure. Instead, it is directly predicted by a network learnt from the training set. This zero-shot capability is critical for flexible, distributed, and scalable implementation and deployment of the image indexing and retrieval services at large scales. Second, unlike the existing distance-based index methods, our index structure is learnt using the LTI-ST deep neural network with binary encoding and decoding on a hierarchical semantic tree. Our extensive experimental results on benchmark datasets and ablation studies demonstrate that the proposed LTI-ST method outperforms existing index methods by a large margin while providing the above new capabilities which are highly desirable in practice.
Shichao Kan, Yi-Gang Cen, Vladimir Mladenovic, Yang Li 0091, Zhihai He
IEEE Trans. Image Process.6
2021 Removing Adversarial Noise via Low-Rank Completion of High-Sensitivity Points
abstract
Deep neural networks are fragile under adversarial attacks. In this work, we propose to develop a new defense method based on image restoration to remove adversarial attack noise. Using the gradient information back-propagated over the network to the input image, we identify high-sensitivity keypoints which have significant contributions to the image classification performance. We then partition the image pixels into the two groups: high-sensitivity and low-sensitivity points. For low-sensitivity pixels, we use a total variation (TV) norm-based image smoothing method to remove adversarial attack noise. For those high-sensitivity keypoints, we develop a structure-preserving low-rank image completion method. Based on matrix analysis and optimization, we derive an iterative solution for this optimization problem. Our extensive experimental results on the CIFAR-10, SVHN, and Tiny-ImageNet datasets have demonstrated that our method significantly outperforms other defense methods which are based on image de-noising or restoration, especially under powerful adversarial attacks.
Zhiqun Zhao, Hengyou Wang, Hao Sun 0024, Jianhe Yuan, Zhongchao Huang, Zhihai He
IEEE Trans. Image Process.6
2021 Snowball: Iterative Model Evolution and Confident Sample Discovery for Semi-Supervised Learning on Very Small Labeled Datasets
abstract
In this work, we develop a joint sample discovery and iterative model evolution method for semi-supervised learning on very small labeled training sets. We propose a master-teacher-student model framework to provide multi-layer guidance during the model evolution process with multiple iterations and generations. The teacher model is constructed by performing an exponential moving average of the student models obtained from past training steps. The master network combines the knowledge of the student and teacher models with additional access to newly discovered samples. The master and teacher models are then used to guide the training of the student network by enforcing the consistency between their predictions of unlabeled samples and evolve all models when more and more samples are discovered. Our extensive experiments demonstrate that the process of discovering confident samples from the unlabeled dataset, once coupled with the master-teacher-student network evolution, can significantly improve the overall semi-supervised learning performance. For example, on the CIFAR-10 dataset, with a small set of 250 labeled samples, our method achieves an error rate of 11.58%, more than 38% lower than Mean-Teacher (49.91%). When coupled with the MixMatch augmentation and loss function, the improvements are also significant.
Yang Li 0091, Zhiqun Zhao, Hao Sun 0024, Yi-Gang Cen, Zhihai He
IEEE Trans. Multim.5
2021 Mask Cross-Modal Hashing Networks
abstract
Due to the rapid development of deep learning, cross-modal retrieval has achieved significant progress in recent years. Moreover, cross-modal hashing has recently attracted considerable attention to multi-modal retrieval applications due to its advantages of low storage costs and fast retrieval speed. However, it is still a challenging problem due to an existing semantic heterogeneity gap between different modalities. In order to further narrow the gap and obtain more effective hash codes, we put forward a novel mask deep cross-modal hashing (MDCH) approach to explore the similarity between inter-modal instances. The main contributions of this paper are that: (1) we attempt to introduce semantic mask information into cross-modal hashing retrieval, (2) we alternately train intra-modal and inter-modal networks to fully mine the semantic relationship between different modalities. The semantic mask can improve the semantic information of the image feature. While inter-modal similarity, explored by inter-modal networks, focuses on enforcing images and their corresponding text tags to have similar hash codes, intra-modal similarity, explored by intra-modal networks, can retain local structural information embedded in each modality to achieve internal similarity. A large number of experiments conducted on three datasets demonstrate that our proposed MDCH approach is superior to several state-of-the-art cross-modal hashing approaches.
Qiubin Lin, Wenming Cao 0001, Zhiquan He, Zhihai He
IEEE Trans. Multim.4
2020 Reciprocal Learning Networks for Human Trajectory Prediction
abstract
We observe that the human trajectory is not only forward predictable, but also backward predictable. Both forward and backward trajectories follow the same social norms and obey the same physical constraints with the only difference in their time directions. Based on this unique property, we develop a new approach, called reciprocal learning, for human trajectory prediction. Two networks, forward and backward prediction networks, are tightly coupled, satisfying the reciprocal constraint, which allows them to be jointly learned. Based on this constraint, we borrow the concept of adversarial attacks of deep neural networks, which iteratively modifies the input of the network to match the given or forced network output, and develop a new method for network prediction, called reciprocal attack for matched prediction. It further improves the prediction accuracy. Our experimental results on benchmark datasets demonstrate that our new method outperforms the state-of-the-art methods for human trajectory prediction.
Zhiqun Zhao, Zhihai He
CVPR3
2020 Ensemble Generative Cleaning With Feedback Loops for Defending Adversarial Attacks
abstract
Effective defense of deep neural networks against adversarial attacks remains a challenging problem, especially under powerful white-box attacks. In this paper, we develop a new method called ensemble generative cleaning with feedback loops (EGC-FL) for effective defense of deep neural networks. The proposed EGC-FL method is based on two central ideas. First, we introduce a transformed deadzone layer into the defense network, which consists of an orthonormal transform and a deadzone-based activation function, to destroy the sophisticated noise pattern of adversarial attacks. Second, by constructing a generative cleaning network with a feedback loop, we are able to generate an ensemble of diverse estimations of the original clean image. We then learn a network to fuse this set of diverse estimations together to restore the original image. Our extensive experimental results demonstrate that our approach improves the state-of-art by large margins in both white-box and black-box attacks. It significantly improves the classification accuracy for white-box PGD attacks upon the second best method by more than 29% on the SVHN dataset and more than 39% on the challenging CIFAR-10 dataset.
Jianhe Yuan, Zhihai He
CVPR2
2020 Unsupervised Deep Metric Learning with Transformed Attention Consistency and Contrastive Clustering Loss
Yang Li 0091, Shichao Kan, Zhihai He
ECCV (11)3
2020 Action Completeness Modeling with Background Aware Networks for Weakly-Supervised Temporal Action Localization
abstract
The state-of-the-art of fully-supervised methods for temporal action localization from untrimmed videos has achieved impressive results. Yet, it remains unsatisfactory for the weakly-supervised temporal action localization, where only video-level action labels are given without the timestamp annotation on when the actions occur. The main reason comes from that, the weakly-supervised networks only focus on the highly discriminative frames, but there are some ambiguous frames in both background and action classes. The ambiguous frames in background class are very similar to the real actions, which may be treated as target actions and result in false positives. On the other hand, the ambiguous frames in action class which possibly contain action instances, are prone to be false negatives by the weakly-supervised networks and result in a coarse localization. To solve these problems, we introduce a novel weakly-supervised Action Completeness Modeling with Background Aware Networks (ACM-BANets). Our Background Aware Network (BANet) contains a weight-sharing two-branch architecture, with an action guided Background aware Temporal Attention Module (B-TAM) and an asymmetrical training strategy, to suppress both highly discriminative and ambiguous background frames to remove the false positives. Our action completeness modeling contains multiple BANets, and the BANets are forced to discover different but complementary action instances to completely localize the action instances in both highly discriminative and ambiguous action frames. In the i-th iteration, the i-th BANet discovers the discriminative features, which are then erased from the feature map. The partially-erased feature map is fed into the (i+1)-th BANet of the next iteration to force this BANet to discover discriminative features different from the i-th BANet. Evaluated on two challenging untrimmed video datasets, THUMOS14 and ActivityNet1.3, our approach outperforms all the current weakly-supervised methods for temporal action localization.
Md. Moniruzzaman 0002, Zhaozheng Yin, Zhihai He, Ruwen Qin, Ming C. Leu
ACM Multimedia3
2020 Semantic deep cross-modal hashing
Qiubin Lin, Wenming Cao 0001, Zhihai He, Zhiquan He
Neurocomputing3
2020 L1-norm low-rank linear approximation for accelerating deep neural networks
Zhiqun Zhao, Hengyou Wang, Hao Sun 0024, Zhihai He
Neurocomputing4
2020 Automated work efficiency analysis for smart manufacturing using human pose tracking and temporal action localization
Guanghan Ning, Zhiqun Zhao, Zhongchao Huang, Zhihai He
J. Vis. Commun. Image Represent.5
2020 CLDA: an adversarial unsupervised domain adaptation method with classifier-level adaptation
Zhihai He, Bo Yang 0011, Chaoxian Chen, Qilin Mu, Zesong Li
Multim. Tools Appl.1
2020 Metric learning-based kernel transformer with triplets and label constraints for feature fusion
Shichao Kan, Linna Zhang, Zhihai He, Yi-Gang Cen, Shiming Chen 0001, Jikun Zhou
Pattern Recognit.3
2020 Multi-Matrices Low-Rank Decomposition With Structural Smoothness for Image Denoising
abstract
In this paper, we propose a multi-matrices lowrank decomposition method for image denoising. In this new method, the total variation (TV) norm is incorporated into lowrank approximation analysis to achieve structural smoothness and to improve quality of the recovered images. Our proposed mathematical framework for multi-matrices low-rank decomposition combines the nuclear norm, TV norm, and L1norm, which allows us to exploit the low-rank property of natural images, enhance the structural smoothness, and detect and remove large sparse noise. Based on the iterative alternating direction method, we develop an algorithm to solve the proposed challenging optimization problem. We conduct extensive experiments and perform evaluations on multi-images denoising and multi-frames video prediction. Our experimental results demonstrate that the proposed method outperforms the state-of-the-art low-rank matrix recovery methods, particularly for images with large sparse noise.
Hengyou Wang, Yang Li 0091, Yi-Gang Cen, Zhihai He
IEEE Trans. Circuits Syst. Video Technol.4
2019 Hybrid representation learning for cross-modal retrieval
Wenming Cao 0001, Qiubin Lin, Zhihai He, Zhiquan He
Neurocomputing3
2019 Supervised Deep Feature Embedding With Handcrafted Feature
abstract
Image representation methods based on deep convolutional neural networks (CNNs) have achieved the state-of-the-art performance in various computer vision tasks, such as image retrieval and person re-identification. We recognize that more discriminative feature embeddings can be learned with supervised deep metric learning and handcrafted features for image retrieval and similar applications. In this paper, we propose a new supervised deep feature embedding with a handcrafted feature model. To fuse handcrafted feature information into CNNs and realize feature embeddings, a general fusion unit is proposed (called Fusion-Net). We also define a network loss function with image label information to realize supervised deep metric learning. Our extensive experimental results on the Stanford online products' data set and the in-shop clothes retrieval data set demonstrate that our proposed methods outperform the existing state-of-the-art methods of image retrieval by a large margin. Moreover, we also explore the applications of the proposed methods in person re-identification and vehicle re-identification; the experimental results demonstrate both the effectiveness and efficiency of the proposed methods.
Shichao Kan, Yi-Gang Cen, Zhihai He, Zhi Zhang 0005, Linna Zhang
IEEE Trans. Image Process.3
2018 Towards a social virtual reality learning environment in high fidelity
abstract
Virtual Learning Environments (VLEs) are spaces designed to educate students remotely via online platforms. Although traditional VLEs such as iSocial have shown promise in educating students, they offer limited immersion that diminishes learning effectiveness. This paper outlines a virtual reality learning environment (VRLE) over a high-speed network, which promotes educational effectiveness and efficiency via our creation of flexible content and infrastructure which meet established VLE standards with improved immersion. This paper further describes our implementation of multiple learning modules developed in High Fidelity, a “social VR” platform. Our experiment results show that the VR mode of content delivery better stimulates the generalization of lessons to the real world than non-VR lessons and provides improved immersion when compared to an equivalent desktop version.
Chiara Zizza, Adam Starr, Devin Hudson, Sai Shreya Nuguri, Prasad Calyam, Zhihai He
CCNC6
2018 Function-Guided Energy-Precision Optimization with Precision-Rate-Complexity Bivariate Models
Hao Liu 0002, Rong Huang 0003, Zhihai He
PRCV (3)3
2018 Object detection from dynamic scene using joint background modeling and fast deep learning classification
Hayder Yousif, Jianhe Yuan, Roland Kays, Zhihai He
J. Vis. Commun. Image Represent.4
2018 Real-Time Moving Object Segmentation and Classification From HEVC Compressed Surveillance Video
abstract
Moving object segmentation and classification from compressed video plays an important role in intelligent video surveillance. Compared with H.264/AVC, High Efficiency Video Coding (HEVC) introduces a host of new coding features that can be further exploited for moving object segmentation and classification. In this paper, we present a real-time approach to segment and classify moving objects using unique features directly extracted from the HEVC compressed domain for video surveillance. In the proposed method, first, motion vector (MV) interpolation for intra-coded prediction unit (PU) and MV outlier removal are employed for preprocessing. Second, blocks with nonzero MVs are clustered into the connected foreground regions using the four-connectivity component labeling algorithm. Third, object region tracking based on temporal consistency is applied to the connected foreground regions to remove the noise regions. The boundary of moving object region is further refined by the coding unit size and PU size. Finally, a person-vehicle classification model using bag of spatial-temporal HEVC syntax words is trained to classify the moving objects, either persons or vehicles. The experimental results demonstrate that the proposed method provides solid performance and can classify moving persons and vehicles accurately.
Liang Zhao 0007, Zhihai He, Wenming Cao 0001, Debin Zhao
IEEE Trans. Circuits Syst. Video Technol.2
2018 Reweighted Low-Rank Matrix Analysis With Structural Smoothness for Image Denoising
abstract
In this paper, we develop a new low-rank matrix recovery algorithm for image denoising. We incorporate the total variation (TV) norm and the pixel range constraint into the existing reweighted low-rank matrix analysis to achieve structural smoothness and to significantly improve quality in the recovered image. Our proposed mathematical formulation of the low-rank matrix recovery problem combines the nuclear norm, TV norm, and norm, thereby allowing us to exploit the low-rank property of natural images, enhance the structural smoothness, and detect and remove large sparse noise. Using the iterative alternating direction and fast gradient projection methods, we develop an algorithm to solve the proposed challenging non-convex optimization problem. We conduct extensive performance evaluations on single-image denoising, hyper-spectral image denoising, and video background modeling from corrupted images. Our experimental results demonstrate that the proposed method outperforms the state-of-the-art low-rank matrix recovery methods, particularly for large random noise. For example, when the density of random sparse noise is 30%, for single-image denoising, our proposed method is able to improve the quality of the restored image by up to 4.21 dB over existing methods.
Hengyou Wang, Yi-Gang Cen, Zhiquan He, Zhihai He, Ruizhen Zhao, Fengzhen Zhang
IEEE Trans. Image Process.4
2017 Rate-coverage analysis and optimization for joint audio-video multimedia retrieval
abstract
In this work, we consider the problem of automatic content retrieval (ACR) using joint audio-video fingerprints. We focus on how to balance the query accuracy and the size of fingerprint, and how to allocate the fingerprint bits to video and audio frames to maximize the query accuracy. By introducing a novel concept called coverage, which is highly correlated to the query accuracy, we are able to construct a rate-coverage model and formulate the joint audio-video fingerprint bit rate allocation into a dynamic programming optimization problem. Our experimental results demonstrate that, compared to existing approaches, our method improves the retrieval accuracy by up to 25% while using 60% of the original fingerprint bit rate.
Guanghan Ning, Zhi Zhang 0005, Xiaobo Ren, Haohong Wang, Zhihai He
ICASSP5
2017 Object segmentation in the deep neural network feature domain from highly cluttered natural scenes
abstract
Deep convolutional neural networks (DCNNs) offer an effective hierarchical representation of images for various vision analysis tasks, including classification and detection. In this paper, we propose to study background modeling and object segmentation from highly cluttered natural scenes in the DCNN feature domain instead of traditional pixel domain. Specifically, we first design and train a DCNN for animal-human-background object classification, which is used to analyze the input image to generate multi-layer feature maps, representing the responses of different image regions to the animal-human-background classifier. From these feature maps, we construct the so-called deep objectness graph for accurate animal-human object segmentation with graph cut. The segmented object regions from each image in the sequence are then verified and fused in the temporal domain using background modeling. Recognizing that the DCNN is very computation-intensive, we explore a fast and efficient design of the DCNN which finds a good trade-off between complexity and the classification-segmentation performance. Our experimental results demonstrate that our proposed method outperforms existing state-of-the-art methods on the camera-trap dataset with highly cluttered natural scenes.
Hayder Yousif, Zhihai He, Roland Kays
ICIP2
2017 Spatially supervised recurrent convolutional neural networks for visual object tracking
abstract
In this paper, we develop a new approach of spatially supervised recurrent convolutional neural networks for visual object tracking. Our recurrent convolutional network exploits the history of locations as well as the distinctive visual features learned by the deep neural networks. Inspired by recent bounding box regression methods for object detection, we study the regression capability of Long Short-Term Memory (LSTM) in the temporal domain, and propose to concatenate high-level visual features produced by convolutional networks with region information. In contrast to existing deep learning based trackers that use binary classification for region candidates, we use regression for direct prediction of the tracking locations both at the convolutional layer and at the recurrent unit. Our experimental results on challenging benchmark video tracking datasets show that our tracker is competitive with state-of-the-art approaches while maintaining low computational cost.
Guanghan Ning, Zhi Zhang 0005, Xiaobo Ren, Haohong Wang, Canhui Cai, Zhihai He
ISCAS7
2017 Fast human-animal detection from highly cluttered camera-trap images using joint background modeling and deep learning classification
abstract
In this paper, we couple effective dynamic background modeling with deep learning classification to develop a fast and accurate scheme for human-animal detection from highly cluttered camera-trap images using joint background modeling and deep learning classification. Specifically, first, we develop an effective background modeling and subtraction scheme to generate region proposals for the foreground objects. We then develop a cross-frame image patch verification to reduce the number of foreground object proposals. Finally, we perform complexity-accuracy analysis of deep convolutional neural networks (DCNN) to develop a fast deep learning classification scheme to classify these region proposals into three categories: human, animals, and background patches. The optimized DCNN is able to maintain high level of accuracy while reducing the computational complexity by 14 times. Our experimental results demonstrate that the proposed method outperforms existing methods on the camera-trap dataset.
Hayder Yousif, Jianhe Yuan, Roland Kays, Zhihai He
ISCAS4
2017 Robust Generalized Low-Rank Decomposition of Multimatrices for Image Recovery
abstract
Low-rank approximation has been successfully used for dimensionality reduction, image noise removal, and image restoration. In existing work, input images are often reshaped to a matrix of vectors before low-rank decomposition. It has been observed that this procedure will destroy the inherent two-dimensional correlation within images. To address this issue, the generalized low-rank approximation of matrices (GLRAM) method has been recently developed, which is able to perform low-rank decomposition of multiple matrices directly without the need for vector reshaping. In this paper, we propose a new robust generalized low-rank matrices decomposition method, which further extends the existing GLRAM method by incorporating rank minimization into the decomposition process. Specifically, our method aims to minimize the sum of nuclear norms and l1-norms. We develop a new optimization method, called alternating direction matrices tri-factorization method, to solve the minimization problem. We mathematically prove the convergence of the proposed algorithm. Our extensive experimental results demonstrate that our method significantly outperforms existing GLRAM methods.
Hengyou Wang, Yi-Gang Cen, Zhihai He, Ruizhen Zhao, Fengzhen Zhang
IEEE Trans. Multim.3
2016 Facial expression recognition based on LLENet
abstract
Facial expression recognition plays an important role in lie detection, and computer-aided diagnosis. Many deep learning facial expression feature extraction methods have a great improvement in recognition accuracy and robutness than traditional feature extraction methods. However, most of current deep learning methods need special parameter tuning and ad hoc fine-tuning tricks. This paper proposes a novel feature extraction model called Locally Linear Embedding Network (LLENet) for facial expression recognition. The proposed LLENet first reconstructs image sets for the cropped images. Unlike previous deep convolutional neural networks that initialized convolutional kernels randomly, we learn multi-stage kernels from reconstructed image sets directly in a supervised way. Also, we create an improved LLE to select kernels, from which we can obtain the most representative feature maps. Furthermore, to better measure the contribution of these kernels, a new distance based on kernel Euclidean is proposed. After the procedure of multi-scale feature analysis, feature representations are finally sent into a linear classifier. Experimental results on facial expression datasets (CK+) show that the proposed model can capture most representative features and thus improves previous results.
Dan Meng 0001, Guitao Cao, Zhihai He, Wenming Cao 0001
BIBM3
2016 HEVC compressed domain moving object detection and classfication
abstract
Compressed domain moving object segmentation and classification plays an important role in many real-time applications, such as video indexing and intelligent video surveillance. Compared with the previous international video coding standards, such as H.264/AVC, HEVC introduces a host of new coding features. Therefore, moving object segmentation and classification directly from HEVC compressed videos represents a new challenge. In this paper, we develop a method for segmenting and classifying moving objects, specifically, persons and vehicles, in the HEVC compression domain. We first train a classifier to determine if an image patch belongs to the foreground objects or background using HEVC syntax features. This will generate a bounding box which locates the object in the video frame. We then train a second classification model to classify the moving objects, either persons or vehicles, using bags of spatial-temporal HEVC syntax words. Our extensive experimental results demonstrate that the approach provides the remarkable performance and can classify moving person and vehicles accurately and robustly.
Liang Zhao 0007, Debin Zhao, Xiaopeng Fan 0001, Zhihai He
ISCAS4
2016 Task-driven progressive part localization for fine-grained recognition
abstract
In this paper we propose a task-driven progressive part localization (TPPL) approach for fine-grained object recognition. Most existing methods follow a two-step approach which first detects salient object parts to suppress the interference from background scenes and then classifies objects based on features extracted from these regions. The part detector and object classifier are often independently designed and trained. In this paper, our major finding is that the part detector should be jointly designed and progressively refined with the object classifier so that the detected regions can provide the most distinctive features for final object recognition. Specifically, we start with a part-based SPP-net (Part-SPP) as our baseline part detector. We then develop a task-driven progressive part localization framework, which takes the predicted boxes of Part-SPP as an initial guess, then examines new regions in the neighborhood, searching for more discriminative image regions to maximize the recognition performance. This procedure is performed in an iterative manner to progressively improve the joint part detection and object classification performance. Experimental results on the Caltech-UCSD-200-2011 dataset demonstrate that our method outperforms state-of-the-art fine-grained categorization methods both in part localization and classification, even without requiring a bounding box during testing.
Zhihai He
WACV2
2016 Constellational contour parsing for deformable object detection
Tony X. Han, Zhihai He, Wenming Cao 0001
J. Vis. Commun. Image Represent.3
2016 Task-Driven Progressive Part Localization for Fine-Grained Object Recognition
abstract
The problem of fine-grained object recognition is very challenging due to the subtle visual differences between different object categories. In this paper, we propose a task-driven progressive part localization (TPPL) approach for fine-grained object recognition. Most existing methods follow a two-step approach that first detects salient object parts to suppress the interference from background scenes and then classifies objects based on features extracted from these regions. The part detector and object classifier are often independently designed and trained. In this paper, our major finding is that the part detector should be jointly designed and progressively refined with the object classifier so that the detected regions can provide the most distinctive features for final object recognition. Specifically, we develop a part-based SPP-net (Part-SPP) as our baseline part detector. We then establish a TPPL framework, which takes the predicted boxes of Part-SPP as an initial guess, and then examines new regions in the neighborhood using a particle swarm optimization approach, searching for more discriminative image regions to maximize the objective function and the recognition performance. This procedure is performed in an iterative manner to progressively improve the joint part detection and object classification performance. Experimental results on the Caltech-UCSD-200-2011 dataset demonstrate that our method outperforms state-of-the-art fine-grained categorization methods both in part localization and classification, even without requiring a bounding box during testing.
Zhihai He, Guitao Cao, Wenming Cao 0001
IEEE Trans. Multim.2
2016 Animal Detection From Highly Cluttered Natural Scenes Using Spatiotemporal Object Region Proposals and Patch Verification
abstract
In this paper, we consider the animal object detection and segmentation from wildlife monitoring videos captured by motion-triggered cameras, called camera-traps. For these types of videos, existing approaches often suffer from low detection rates due to low contrast between the foreground animals and the cluttered background, as well as high false positive rates due to the dynamic background. To address this issue, we first develop a new approach to generate animal object region proposals using multilevel graph cut in the spatiotemporal domain. We then develop a cross-frame temporal patch verification method to determine if these region proposals are true animals or background patches. We construct an efficient feature description for animal detection using joint deep learning and histogram of oriented gradient features encoded with Fisher vectors. Our extensive experimental results and performance comparisons over a diverse set of challenging camera-trap data demonstrate that the proposed spatiotemporal object proposal and patch verification framework outperforms the state-of-the-art methods, including the recent Faster-RCNN method, on animal object detection accuracy by up to 4.5%.
Zhi Zhang 0005, Zhihai He, Guitao Cao, Wenming Cao 0001
IEEE Trans. Multim.2
2015 Scene text detection based on component-level fusion and region-level verification
abstract
In this paper, we present a novel scene text detection method that combines the advantages of component-based methods and region-based methods, while overcoming their inherent limitations. We first extract text regions as candidates, and then aggregate these text components in these regions into words and text lines. To separate non-text components in the background from text components, we perform both character-level filtering and word-level classification with a trained linear SVM (support vector machine) classifier. Our extensive experiments on ICDAR2003 and ICDAR2011 datasets have shown that our method outperforms the state-of-the-art methods in text detection.
Guanghan Ning, Tony X. Han, Zhihai He
ICIP3
2015 Coupled ensemble graph cuts and object verification for animal segmentation from highly cluttered videos
abstract
In this paper, we consider animal object segmentation from wildlife monitoring videos captured by motion-triggered cameras, called camera-traps. This is a very challenging task because the wildlife monitoring scenes are often highly cluttered and dynamic. To address this issue, we propose to explore the ideas of coupled ensemble graph cuts and object verification. We consider video object cut as an ensemble of frame-level background-foreground object classifiers which fuse information across frames and refine their segmentation results in a collaborative and iterative manner. To significantly reduce false positives in foreground animal detection and segmentation, we learn an object verification model to further classify if the segmented image patch belongs to the background or the animal. Our extensive experimental results and performance comparisons over a diverse set of challenging camera-trap data, as well as the new Change Detection 2014 benchmark dataset, demonstrate that the proposed framework outperforms various state-of-the-art algorithms and has the capability to handle even the most challenging objects in a wide variety of video sequences.
Zhi Zhang 0005, Tony X. Han, Zhihai He
ICIP3
2014 Randomized Support Vector Forest
Xutao Lv, Tony X. Han, Zicheng Liu 0001, Zhihai He
BMVC4
2014 Deep convolutional neural network based species recognition for wild animal monitoring
abstract
We proposed a novel deep convolutional neural network based species recognition algorithm for wild animal classification on very challenging camera-trap imagery data. The imagery data were captured with motion triggered camera trap and were segmented automatically using the state of the art graph-cut algorithm. The moving foreground is selected as the region of interests and is fed to the proposed species recognition algorithm. For the comparison purpose, we use the traditional bag of visual words model as the baseline species recognition algorithm. It is clear that the proposed deep convolutional neural network based species recognition achieves superior performance. To our best knowledge, this is the first attempt to the fully automatic computer vision based species recognition on the real camera-trap images. We also collected and annotated a standard camera-trap dataset of 20 species common in North America, which contains 14, 346 training images and 9, 530 testing images, and is available to public for evaluation and benchmark purpose.
Guobin Chen, Tony X. Han, Zhihai He, Roland Kays, Tavis Forrester
ICIP3
2014 Multi-scale embedded descriptor for shape classification
Tony X. Han, Zhihai He
J. Vis. Commun. Image Represent.3
2014 Delay-Bounded Priority-Driven Resource Allocation for Video Transmission Over Multihop Networks
abstract
In this paper we consider the problem of resource allocation for video transmission over mesh networks with delay bound constraints and priority-based packet scheduling. We observe that priority-driven packet scheduling at the intermediate network routers has a direct and significant impact on the queuing behaviors and delay bound violation probabilities of video packets, as well as the overall end-to-end video distortion. Using learning methods, we develop a packet delay bound violation probability model for video transmission over multihop networks with priority-based packet scheduling. With this model, we can successfully predict the probability of packets being dropped due to violation of specified delay bounds. We also observe that the transmission distortion caused by packet drops exhibits a unique exponential behavior with priority-based packet scheduling. With these analysis results, we formulate the resource allocation for multisession video transmission over networks with priority-driven packet scheduling under delay bound constraints as a multiobjective optimization problem. Evolutionary optimization methods based on single- and multiobjective genetic algorithms are proposed to solve the problem and obtain the optimal resource allocation. Extensive experiment results demonstrate the effectiveness of the proposed resource-distortion models and optimization algorithms.
Yongfei Zhang, Shiyin Qin, Bo Li 0006, Zhihai He
IEEE Trans. Circuits Syst. Video Technol.5
2013 Ensemble Video Object Cut in Highly Dynamic Scenes
abstract
We consider video object cut as an ensemble of frame-level background-foreground object classifiers which fuses information across frames and refine their segmentation results in a collaborative and iterative manner. Our approach addresses the challenging issues of modeling of background with dynamic textures and segmentation of foreground objects from cluttered scenes. We construct patch-level bag-of-words background models to effectively capture the background motion and texture dynamics. We propose a foreground salience graph (FSG) to characterize the similarity of an image patch to the bag-of-words background models in the temporal domain and to neighboring image patches in the spatial domain. We incorporate this similarity information into a graph-cut energy minimization framework for foreground object segmentation. The background-foreground classification results at neighboring frames are fused together to construct a foreground probability map to update the graph weights. The resulting object shapes at neighboring frames are also used as constraints to guide the energy minimization process during graph cut. Our extensive experimental results and performance comparisons over a diverse set of challenging videos with dynamic scenes, including the new Change Detection Challenge Dataset, demonstrate that the proposed ensemble video object cut method outperforms various state-of-the-art algorithms.
Xiaobo Ren, Tony X. Han, Zhihai He
CVPR3
2013 Rate-distortion optimized unequal loss protection for video transmission over packet erasure channels
Yongfei Zhang, Shiyin Qin, Bo Li 0006, Zhihai He
Signal Process. Image Commun.4
2012 Histogram of Oriented Normal Vectors for Object Recognition with a Depth Sensor
Xiaoyu Wang 0002, Xutao Lv, Tony X. Han, James Keller 0001, Zhihai He, Marjorie Skubic, Shihong Lao
ACCV (2)6
2012 Semi-supervised learning for robust car windshield tracking and monitoring in live traffic videos
abstract
This paper deals with the problem of car-windshield tracking in live traffic video. To avoid a comprehensive labeled dataset that covers most appearance variations, we aim to appropriately involve unlabeled examples and efficiently update the discriminative model in an online semi-supervised setting. Our approach follows the state-of-the- art “learning by detection” approach, yet different from it in the following aspects. First, instead of assigning hard labels to new added examples, we leave them unlabeled. Second, we focus on exploring the intrinsic manifold structure of data marginal distribution and studying its role in kernel function optimization. The proposed online semi-supervised learning framework involves a 3D mean-shift optimization for windshield localization and is followed by a block-based decision for co-driver detection. The experimental results demonstrate the effectiveness of the proposed method.
Zhongna Zhou, Tony X. Han, Zhihai He
ICIP3
2012 Resource-Distortion Modeling for Video Streaming over Mesh Networks with Priority-Based Packet Scheduling
abstract
Video streaming over mesh networks operates under stringent network resource constraints, with a large number of video sessions competing for limited network resources. In this work, we aim to establish so-called resource-distortion models to characterize the inherent relationship between the allocated network resources to each video session and its end-to-end video distortion or presentation quality at the receiver end. We observed that priority-based packet scheduling has significant impact on such resource-distortion relationship. Using ANN-based learning methods, we develop an end-to-end packet delay bound violation (PDBV) probability model for video streaming over multi-hop networks with priority-based packet scheduling. We then derive a quadratic video transmission distortion model to capture the unique behavior of priority-based packet scheduling and its impact on the end-to-end video distortion. Based on these resource-distortion models, we are able to predict the end-to-end distortion for video streaming over multi-hop networks with priority-based packet scheduling. Our extensive experimental results demonstrate that the proposed method is very accurate.
Yongfei Zhang, Shiyin Qin, Zhihai He
ICME4
2012 Integrated camera and sensor networks: Algorithms, design, and applications
abstract
Recent technological advances in hardware miniaturization of sensors, low-power microprocessor design, and wireless ad hoc networking have enabled the deployment of large-scale wireless sensor networks. A wireless sensor network is a system of geographically distributed sensor nodes that communicate with each other over a dynamic and self-organizing wireless network. In a wireless video sensor network (WVSN), each sensor, equipped with video capture and processing capabilities, is tasked to capture digital video information about the target event or situation, and deliver the video data to the remote control unit for further information analysis and decision making.
Zhihai He
VCIP1
2011 Transmission Distortion-optimized Unequal Loss Protection for video transmission over packet erasure channels
abstract
In this paper, we study the problem of Transmission Distortion-optimized Unequal Loss Protection (TD-ULP) under rate constraints for non-scalable video transmission over packet erasure channels. Based on a packet-level transmission distortion modeling scheme, we estimate the amount of contribution of each video packet to the reconstructed video quality, which defines the priority level of each packet. Unequal amounts of protections are then allocated to different video packets according to their priority levels as well as the dynamic channel conditions. The optimal ULP resource allocation is formulated as a constrained nonlinear optimization problem. An evolutionary algorithm based on Particle Swarm Optimization (PSO) is developed to obtain the optimal resource allocation. Our extensive experimental results demonstrate the effectiveness of the proposed TD-ULP scheme, which outperforms existing methods by up to 2dB gain in reconstructed video quality. 1*
Yongfei Zhang, Shiyin Qin, Zhihai He
ICME3
2011 Joint Source and Flow Optimization for Scalable Video Multirate Multicast Over Hybrid Wired/Wireless Coded Networks
abstract
This paper aims to optimize the overall video quality and traffic performance for multi-rate video multicast over hybrid wired/wireless networks. In order to perform layered utility maximization over tiered networks, we propose a joint source-network flow optimization scheme where individual layers of the scalable video stream are imposed on their optimal multicast paths and associated rates for the highest sustainable layered video quality with minimum costs. It sufficiently guarantees that each destination node accesses progressive layered stream in an incremental order, considers network coding across overlapping paths to destination nodes for decent multicast capacity, and addresses the link contention problem during wireless transmission. We formulate the problem into convex programming with the objective to minimize the total rate-distortion variations between layers. Using primal decomposition and the primal-dual approach, we develop a decentralized algorithm with two levels of optimization. The numerical and packet-level results compare extensive performance under different control conditions over coded and non-coded hybrid networks. It demonstrates that the proposed algorithm could actually achieve the max-flow throughput and provide better video quality with optimal layered access for heterogeneous receivers.
Hongkai Xiong, Junni Zou, Zhihai He
IEEE Trans. Circuits Syst. Video Technol.4
2011 An Error Resilient Video Coding Scheme Using Embedded Wyner-Ziv Description With Decoder Side Non-Stationary Distortion Modeling
abstract
In this paper, we propose a generic error resilient video coding (ERVC) scheme using embedded Wyner-Ziv (WZ) description. At the encoder side, a joint source-channel R-D optimized mode selection (JSC-RDO-MS) algorithm with WZ-coded anchor frames is statistically studied and developed. Given a stationary first-order Markov Gaussian source, the proposed mode optimization is justified by an analysis of the RD impact on the WZ bit-rate. JSC-RDO-MS involves in the estimation of expected rate and distortion of WZ coding with the unavailable side information, and the WZ bit-rate of each coding mode is determined based on the error correction capability of the specific WZ codec. At the decoder side, an online correlation noise model between the source and the side-information is proposed with a mixture of Laplacians whose parameters are attained to reflect the coherence of the motion field of successive frames and the energy of prediction residual. Each mixture component represents the statistical distribution of prediction residuals, and the mixing coefficients represent the amount of errors in motion compensation. The proposed scheme achieves the so-called classification gain by exploiting the spatially non-stationary characteristics of the motion field and texture. Extensive experimental results show that the proposed WZ-ERVC scheme achieves a better overall RD performance than existing ERVC schemes, and the proposed modeling algorithm also significantly outperforms the conventional Laplacian model by up to 2 dB.
Hongkai Xiong, Zhihai He, Songyu Yu, Chang Wen Chen
IEEE Trans. Circuits Syst. Video Technol.3
2011 Reconstruction for Distributed Video Coding: A Context-Adaptive Markov Random Field Approach
abstract
Within the existing reconstruction process of distributed video coding (DVC), there are two major approaches: the maximum probability reconstruction and the minimum mean square error (MMSE) reconstruction. Both of them assume that each node, a pixel in pixel domain DVC or a coefficient in transform domain DVC, is i.i.d., and reconstruct the value of each node independently by only exploiting statistical correlation between source and side-information. These kinds of models produce considerable amount of artifacts in decoded Wyner-Ziv (WZ) frames and degrade the objective performance. In this paper, we propose a context-adaptive Markov random field (MRF) reconstruction algorithm which exploits both the statistical correlation and the spatio-temporal consistency by modeling the corresponding MRF of a generic DVC architecture, and solve the inference by finding its MRF-based maximum a posteriori (MAP) estimate. The energy function of the MRF model consists of two terms: a data term measuring the statistical correlation, and a geometric regularity term enforcing local spatio-temporal structure consistency which is modeled by optical flow estimation with regard to the critical parameters under a wide variety of DVC scenarios. In case the unreliability of the derived local structure, a confidence parameter is introduced to prevent inappropriate penalizing. To find the reconstructed patch assignment with the largest expected probability in the context-adaptive MRF, the energy minimization for the MRF-based MAP estimate of the WZ frames is solved by global optimization and greedy strategies. Compared to the existing maximum probability and MMSE reconstruction with i.i.d. model, a better subjective and objective performance is validated by extensive experiments.
Hongkai Xiong, Zhihai He, Songyu Yu, Chang Wen Chen
IEEE Trans. Circuits Syst. Video Technol.3
2011 Prioritized Flow Optimization With Multi-Path and Network Coding Based Routing for Scalable Multirate Multicasting
abstract
In this paper, we study performance optimization for scalable video coding and multicast over networks. Multi-path video streaming, network coding based routing, and network flow control are jointly optimized to maximize a network utility function defined over heterogeneous receivers. Content priority of video coding layers is considered during the flow routing to determine the optimal multicast paths and associated data rates for each layer. Our optimization scheme attempts to find content distribution meshes with minimum path costs for each video coding layer while satisfying the inter-layer dependency during scalable video coding. Based on primal decomposition and primal-dual analysis, we develop a decentralized algorithm with two optimization levels to solve the performance optimization problem. We also prove the stability and convergence of the proposed iterative algorithm using Lyapunov theories. Extensive experimental results demonstrate that the proposed algorithm not only achieves the max-flow throughput using network coding, but also provides better video quality with balanced layered access for heterogeneous receivers.
Junni Zou, Hongkai Xiong, Li Song 0001, Zhihai He, Tsuhan Chen
IEEE Trans. Circuits Syst. Video Technol.5
2010 Local structure learning and prediction for efficient lossless image compression
abstract
One major challenge in image compression is to efficiently represent and encode high-frequency structure components in images, such as edges, contours, and texture regions. To address this issue for lossy image compression, in our previous work, we proposed a scheme to learn local image structures and efficiently predict image data based on this structure information. In this work, we applied this structure learning and prediction scheme to lossless image compression and developed a lossless image encoder. Our extensive experimental results demonstrate that the lossless image encoder is competitive and even outperforms the state-of-the-art lossless image compression methods.
Xiwen Zhao, Zhihai He
ICASSP2
2010 Reconstruction for distributed video coding: a Markov random field approach with context-adaptive smoothness prior
abstract
An important issue in Wyner-Ziv video coding is the reconstruction of Wyner-Ziv frames with decoded bit-planes. So far, there are two major approaches: the Maximum a Posteriori (MAP) reconstruction and the Minimum Mean Square Error (MMSE) reconstruction algorithms. However, these approaches do not exploit smoothness constraints in natural images. In this paper, we model a Wyner-Ziv frame by Markov random fields (MRFs), and produce reconstruction results by finding an MAP estimation of the MRF model. In the MRF model, the energy function consists of two terms: a data term, MSE distortion metric in this paper, measuring the statistical correlation between side-information and the source, and a smoothness term enforcing spatial coherence. In order to better describe the spatial constraints of images, we propose a context-adaptive smoothness term by analyzing the correspondence between the output of Slepian-Wolf decoding and successive frames available at decoders. The significance of the smoothness term varies in accordance with the spatial variation within different regions. To some extent, the proposed approach is an extension to the MAP and MMSE approaches by exploiting the intrinsic smoothness characteristic of natural images. Experimental results demonstrate a considerable performance gain compared with the MAP and MMSE approaches.
Hongkai Xiong, Zhihai He, Songyu Yu
VCIP3
2010 Adaptive Critic Design for Energy Minimization of Portable Video Communication Devices
abstract
Portable video communication devices operate on batteries with limited energy supply. Video compression is computationally intensive and energy-demanding. Therefore, one critical issue in portable video communication system design is to minimize the energy consumption of video encoding so as to prolong the operational lifetime of portable video devices. In this paper, we explore advanced methods in adaptive system control and develop an online complexity control and energy minimization scheme for real-time video encoding. More specifically, we introduce a set of parameters to control the computational complexity and energy consumption of the video encoder. We consider this video encoder as a nonlinear control system. Based on adaptive critic design, an advanced adaptive control method recently developed in control science, we design an online control scheme which is able to select the best configuration of complexity control parameters to minimize the energy consumption under rate-distortion constraints, or equivalently maximize video quality under rate and energy constraints. Our extensive experimental results demonstrate that the proposed scheme approaches the true optimum performance. The proposed online complexity control and energy minimization scheme will provide an important tool for energy minimization of portable video devices.
Zhihai He
IEEE Trans. Circuits Syst. Video Technol.3
2010 Fine-Granularity Transmission Distortion Modeling for Video Packet Scheduling Over Mesh Networks
abstract
Packet scheduling is a critical component in multi-session video streaming over mesh networks. Different video packets have different levels of contribution to the overall video presentation quality at the receiver side. In this work, we develop a fine-granularity transmission distortion model for the encoder to predict the quality degradation of decoded videos caused by lost video packets. Based on this packet-level transmission distortion model, we propose a content-and-deadline-aware scheduling (CDAS) scheme for multi-session video streaming over multi-hop mesh networks, where content priority, queuing delays, and dynamic network transmission conditions are jointly considered for each video packet. Our extensive experimental results demonstrate that the proposed transmission distortion model and the CDAS scheme significantly improve the performance of multi-session video streaming over mesh networks.
Yongfei Zhang, Shiyin Qin, Zhihai He
IEEE Trans. Multim.3
2010 Multihop Packet Delay Bound Violation Modeling for Resource Allocation in Video Streaming Over Mesh Networks
abstract
Resource allocation plays a critical role in multisession video streaming over mesh networks to maximize the overall video presentation quality under transmission delay and network resource constraints. A critical component in efficient resource allocation is to analyze and model the multihop queuing behavior along the transmission path, estimate the packet loss ratio due to delay bound violation, and predict the amount of video quality degradation after multihop video transmission. In this work, we develop a multihop packet delay bound violation model to predict the packet loss probability and end-to-end distortion for video streaming over multihop networks. To this end, we extract salient features to characterize the input source and network conditions of links along the transmission path and construct a learning-based model using artificial neural network (ANN). Based on this model, we then formulate the resource allocation into a nonconvex optimization problem which aims to minimize the overall video distortion while maintaining fairness between sessions. We solve this optimization problem using Lagrangian duality methods. Extensive experimental results demonstrate that, with this widely-used offline-training-online-estimation mechanism, the proposed model is potentially applicable to almost all network conditions and can provide fairly accurate estimation results as compared with other models with a given sample data set. The proposed optimization algorithm achieves more efficient resource allocation than existing schemes.
Yongfei Zhang, Shixin Sun, S. Y. Qin, Zhihai He
IEEE Trans. Multim.5
2009 Graph Matching Based Side Information Generation for Distributed Multi-View Video Coding
abstract
In this paper, we adopt constrained relaxation for distributed multi-view video coding (DMVC). The novel framework integrates the graph-based segmentation and matching to generate inter-view correlated side information without knowing the camera parameters. Moreover, graph-based representations of multi-view images are incorporated to form more distinctive feature constraints. The sparse data as a good hypothesis space aim for a best matching optimization of inter-view side information with compact syndromes, from inferred relaxed coset. The plausible filling-in from a priori feature constraints between neighboring views could reinforce a promising compensation to inter-view side information generation for joint multi-view decoding. In order to find distinctive feature matching with a more stable approximation, PCA-SIFT and TPS (thin plate spline) are adopted to reduce the dimension of SIFT descriptors and construct a more accurate inter-view motion model. The experimental results validate the high estimation precision and the rate-distortion improvements.
Hui Lu 0001, Hongkai Xiong, Li Song 0001, Zhihai He, Tsuhan Chen
ICC4
2009 Prioritized Flow Optimization with Generalized Routing for Scalable Multirate Multicasting
abstract
This paper addresses the performance optimization for scalable video coding and multicast over networks. Multi-path video streaming, network coding based routing, and network flow control are jointly optimized to maximize a network utility function defined over heterogeneous receivers. Importantly, contextual priors of scalable video layers are imposed on the flow routing optimization problem, seeking to guarantee the transmission cost for each layer in an incremental order and find jointly optimal multicast paths and associated rates. Through a primal decomposition and the primal-dual approach, a decentralized algorithm with two-level optimization update is developed to solve the target convex optimization problem. Numerical and simulation results validate the convergence and network performance of the proposed algorithm.
Junni Zou, Hongkai Xiong, Li Song 0001, Zhihai He, Tsuhan Chen
ICC4
2009 Building recognition using sketch-based representations and spectral graph matching
abstract
In this work, we address the problem of building recognition across two camera views with large changes in scales and viewpoints. The main idea is to construct a semantically rich sketch-based representation for buildings which is invariant under large scale and perspective changes. After multi-scale maximally stable extremal regions (MSER) detection, the proposed approach finds repeated structural components of buildings, such as window, doors, and facades, and extracts semantically rich features, which are organized into a sketch-based representation of buildings. These descriptors are then clustered in association with different planes of the building and matched across video frames using spectral graph analysis. Our experiments demonstrate that the proposed approach outperforms SIFT-based matching schemes, especially for images with large viewpoint changes.
Yu-Chia Chung, Tony X. Han, Zhihai He
ICCV3
2009 Energy-efficient integrated camera and sensor system design for wildlife activity monitoring
abstract
We introduce our recent research effort on energyefficient portable video communication system design for wildlife activity monitoring. The capability of seeing what an animal sees in the field is very important for wildlife activity monitoring and research. We have designed an integrated video and sensor system, called DeerCam and mount it on animals so as to collect important video and sensor data about their activities in the field. We will explain the major design challenges and research issues in this project.
Zhihai He
ICME1
2009 Video-based Activity Monitoring for Indoor Environments
abstract
In this work, we study how continuous video monitoring and intelligent video processing can be used in eldercare to assist the independent living of elders and to improve the efficiency of eldercare practice. More specifically, we construct an advanced silhouette extraction and tracking algorithm for indoor environments. An adaptive learning method was developed to estimate the physical location and moving speed of a person from a single camera view without calibration. Then hierarchical decision tree and dimension reduction methods were used for human action recognition. We extract important ADL (activities of daily living) statistics for automated functional assessment. Our extensive tests over these massive video datasets demonstrate that the proposed automated activity analysis system is very efficient.
Zhongna Zhou, Yu-Chia Chung, Zhihai He, Tony X. Han, James Keller 0001
ISCAS4
2009 Efficient H.264 video coding with a working memory of objects
abstract
In this work, we investigate a working memory approach for efficient temporal prediction in H.264 video coding. After video frames are encoded, objects are extracted, analyzed, and indexed in a dynamic database which acts as a working memory for the H.264 video encoder. During the encoding process, objects with similar spatial characteristics are retrieved from the working memory and used for motion prediction of objects in the current video frame. This approach extends the multiple-frame estimation and provides a more generic framework for spatiotemporal prediction of video data. Our experimental results on surveillance video data demonstrate that the proposed approach is able to save the coding bit rate by up to 35% with a small computational overhead.
Wenqing Dai, Yu Sun 0003, Zhihai He
PCS3
2009 Packet-level transmission distortion analysis for video streaming over mesh networks
abstract
In video transmission over wired or wireless mesh networks, video packets could be lost due to transmission errors, network congestion or delay bound violation. The compressed video data is highly sensitive to packet loss, and the packet loss will cause decoding failure and more importantly, the error propagation would significantly degrade the reconstructed video quality. In this paper, we extend our previous research and propose a simple yet accurate and robust model to estimate the transmission distortion of each Macro Block (MB) and packet. We then use this transmission distortion estimation results for packet scheduling in video streaming over mesh networks. Extensive simulation results demonstrate the effectiveness and robustness of the proposed model under different experiment settings. More importantly, the estimated transmission distortion, which reveals the unequal importance of each MB and packet, enables important applications in content-aware resource allocation and performance optimization in video communication.
Yongfei Zhang, Shiyin Qin, Zhihai He
PCS3
2009 Lossless image compression using super-spatial prediction of structural components
abstract
We recognize that the key challenge in image compression is to efficiently represent and encode high-frequency image structural components, such as edges, patterns, and textures. Existing image compression schemes attempt to predict image data using its spatial neighborhood. In this work, we develop an efficient image compression scheme based on super-spatial prediction of structural units. This so-called super-spatial prediction breaks this neighborhood constraint, attempting to find an optimal prediction of structural components within the whole image domain. We consider only lossless image compression. Our extensive experimental results demonstrate that the proposed scheme is very competitive and even outperforms the state-of-the-art image compression methods.
Xiwen Zhao, Zhihai He
PCS2
2009 Incremental rate control for H.264/AVC video compression
abstract
In this study, the authors propose a new rate-complexity-quantisation model and an incremental rate control algorithm for H.264/AVC video coding. One unique property of this algorithm is that, the picture complexity estimation and rate-quantisation modelling are jointly designed with an incremental rate control for P-frames. In addition, the proposed algorithm also introduces a number of efficient rate control techniques, including accurate rate control for intra-frames, enhanced proportional–integral–derivative (PID) buffer controller, and adaptive quantisation parameter determination for B-frames. The proposed algorithm has low computational complexity while providing robust rate control. Our extensive experimental results demonstrate that the proposed algorithm outperforms the current rate control algorithm adopted in the H.264/AVC reference software JM13.2 by achieving more accurate rate control, reducing frame skipping, depressing quality fluctuation and improving the overall coding quality by up to 2.83 dB.
Yu Sun 0003, Yimin Zhou 0002, Zhidan Feng 0001, Zhihai He, Shixin Sun
IET Image Process.4
2009 Reliability Analysis for Global Motion Estimation
abstract
Global motion estimation (GME) is the enabling step for many important video exploitation tasks. In this work, we focus on indirect GME methods which have low computational complexity. Typically, an indirect GME method has two major steps. The first step is to find point correspondence between frames through local motion search or feature matching. Then, the second step determines global motion parameters using optimal model fitting, such as least mean-squared error (LMSE) fitting or RANSAC. However, due to image noise and inherent ambiguity in point correspondence, local motion estimation often suffers from relatively large errors, which degrade the performance and reliability of GME. In this work, we propose a method to characterize the reliability of local motion estimation results and use this reliability measure as a weighting factor to determine the importance level of each local motion estimation result during global motion estimation. Our simulation results demonstrate that the proposed scheme is able to significantly improve the accuracy and robustness of global motion estimation with a very small computational overhead.
Yu-Chia Chung, Zhihai He
IEEE Signal Process. Lett.2
2008 Adaptive Critic Design for Energy Minimization of Portable Video Communication Devices
abstract
Portable video communication devices operate on batteries with limited energy supply. However, video compression is computationally intensive and energy-demanding. Therefore, one of the central challenging issues in portable video communication system design is to minimize the energy consumption of video encoding so as to prolong the operational lifetime of portable video devices. In this work, we consider a video encoder as a nonlinear system with a number of encoder parameters to its power consumption. We explore the approach of adaptive critic design to control and optimize the power consumption behavior of a portable video encoding system. Our experimental results demonstrate that this approach is very efficiently, being able to achieve the optimum performance accurately and robustly.
Zhihai He
ICCCN3
2008 A novel incremental rate control scheme for H.264 video coding
abstract
In this paper, we propose a new incremental rate control scheme for H.264/AVC video coding, which elegantly resolves the "Chicken and Egg" dilemma by eliminating the need of coding complexity prediction for inter-frames. The proposed scheme introduces a number of new features, including a rate- complexity-quantization model, accurate quantization parameter (QP) estimation for intra-frames, incremental QP calculation for inter-frames, and a basic unit level QP adjustment. Experimental results demonstrate that the proposed scheme outperforms the JVT-G012 solution by providing more accurate rate control, reducing frame skipping, and decreasing quality fluctuations, and finally, improving coding quality by up to 1.85 dB.
Yu Sun 0003, Yimin Zhou 0002, Zhidan Feng 0001, Zhihai He
ICIP4
2008 Side information generation with constrained relaxation for distributed multi-view video coding
abstract
Apart from the existing temporal and inter-view interpolation technique in distributed multi-view video coding (DMVC), this paper is dedicated to not only investigating multiple side information implication at the decoder, but also preserving the constrained relaxation with high-level features matching. We present a novel feature-based Wyner-Ziv coding framework (FWZC) for DMVC, which devotes scale-invariant features extraction and matching to generate inter-view correlated side information without knowing the camera parameters and has a more significant improvement rate-distortion performance. The scale-invariant local features are identified as the most resistant to image deformations and affine distortion between different views of an object or scene. The plausible filling-in from a priori distinctive feature constraints between neighboring views could make a promising compensation to inter-view side information generation for joint multi-view decoding. The experimental results show high precision for objects with high motion and quite significant improvement in the rate-distortion performance.
Hui Lu 0001, Hongkai Xiong, Zhihai He
ISCAS4
2008 Special Issue on Video Surveillance
abstract
The 14 regular papers and two brief papers in this special issue capture some of the state-of-the-art research on video surveillance issues, provide comprehensive overview of existing techniques, and propose novel solutions for important research problems.
Ishfaq Ahmad 0001, Zhihai He, Hong-Yuan Mark Liao, Fernando Pereira 0001, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.2
2008 Energy Minimization of PortableVideo Communication Devices Based on Power-Rate-Distortion Optimization
abstract
Portable video communication devices operate on batteries with limited energy supply. However, video compression is computationally intensive and energy-demanding. Therefore, one of the central challenging issues in portable video communication system design is to minimize the energy consumption of video encoding so as to prolong the operational lifetime of portable video devices. In this work, based on power-rate-distortion (P-R-D) optimization, we develop a new approach for energy minimization by exploring the energy tradeoff between video encoding and wireless communication and exploiting the nonstationary characteristics of input video data. Both analytically and experimentally, we demonstrate that incorporating the third dimension of power consumption into conventional R-D analysis gives us one extra dimension of flexibility in resource allocation and allows us to achieve significant energy saving. Within the P-R-D analysis framework, power is tightly coupled with rate, enabling us to tradebitsforjoulesand perform energy minimization through optimum bit allocation. We analyze the energy saving gain of P-R-D optimization. We develop an adaptive scheme to estimate P-R-D model parameters and perform online resource allocation and energy optimization for real-time video encoding. Our experimental studies show that, for typical videos with nonstationary scene statistics, using the proposed P-R-D optimization technology, the energy consumption of video encoding can be significantly reduced (by up to 50%), especially in delay-tolerant portable video communication applications.
Zhihai He, Wenye Cheng
IEEE Trans. Circuits Syst. Video Technol.1
2008 Activity Analysis, Summarization, and Visualization for Indoor Human Activity Monitoring
abstract
In this work, we study how continuous video monitoring and intelligent video processing can be used in eldercare to assist the independent living of elders and to improve the efficiency of eldercare practice. More specifically, we develop an automated activity analysis and summarization for eldercare video monitoring. At the object level, we construct an advanced silhouette extraction, human detection and tracking algorithm for indoor environments. At the feature level, we develop an adaptive learning method to estimate the physical location and moving speed of a person from a single camera view without calibration. At the action level, we explore hierarchical decision tree and dimension reduction methods for human action recognition. We extract important ADL (activities of daily living) statistics for automated functional assessment. To test and evaluate the proposed algorithms and methods, we deploy the camera system in a real living environment for about a month and have collected more than 200 hours (in excess of 600 G bytes) of activity monitoring videos. Our extensive tests over these massive video datasets demonstrate that the proposed automated activity analysis system is very efficient.
Zhongna Zhou, Yu-Chia Chung, Zhihai He, Tony X. Han, James Keller 0001
IEEE Trans. Circuits Syst. Video Technol.4
2008 Linear Rate Control and Optimum Statistical Multiplexing for H.264 Video Broadcast
abstract
The H.264 video coding standard achieves significantly improved video compression efficiency and finds important applications in digital video broadcast. To enable H.264 video encoding for digital TV broadcast and maximize its broadcast efficiency, there are two important issues that need to be adequately addressed. First, we need to understand the complex coding mechanism of an H.264 video encoder and develop a model to analyze and control its rate-distortion (R-D) behavior in an accurate and robust manner. Second, the R-D behaviors of individual channels in the broadcast system should be jointly controlled and optimized under bandwidth and buffer constraints so as to maximize the overall broadcast quality. In this paper, we develop a linear rate model and a linear rate control scheme for H.264 video coding. We develop an optimum statistical multiplexing system to allocate bits across video programs (each being encoded by an H.264 encoder) and video frames so that the overall video broadcast quality is maximized. We study the bandwidth and buffer constraints in video broadcast and formulate the optimum statistical multiplexing into a constrained mathematical optimization problem. Realizing that it is impossible to find a close-form solution for global optima, we propose a simple yet efficient algorithm to find a near-optimum solution for joint rate allocation under buffer constraints. Our extensive simulation results demonstrate that the proposed statistical multiplexing system achieves about 40–50% saving in bandwidth, provides a smooth video quality change across programs and frames, and maintains robust decoder buffer control.
Zhihai He, Dapeng Oliver Wu
IEEE Trans. Multim.1
2007 Peak Transform - A Nonlinear Transform for Efficient Image Representation and Coding
abstract
In this work, we introduce a nonlinear geometric transform, called peak transform, for efficient image representation and coding. Coupled with wavelet transform and subband decomposition, the peak transform is able to significantly reduce signal energy in high-frequency subbands and achieve a significant transform coding gain. This has important applications in efficient data representation and compression. Based on peak transform (PT), we design an image encoder, called PT encoder, for efficient image compression. Our extensive experimental results demonstrate that, in wavelet-based subband decomposition, the signal energy in high-frequency subbands can be reduced by up to 60% if a peak transform is applied. The PT image encoder outperforms state-of-the-art JPEG2000 and H.264 (INTRA) encoders by up to 2-3 dB in PSNR (peak signal-to-noise ratio), especially for images with a significant amount of high-frequency components.
Zhihai He
ICIP (3)1
2007 Distributed Rate Allocation and Performance Optimization for Video Communication Over Mesh Networks
abstract
Video streaming imposes high rate requirement and stringent constraints on resource limited mesh networks. In this work, we develop a distributed asynchronous particle swarm optimization (DAPSO) algorithm for resource allocation and performance optimization scheme for video communication over large-scale mesh networks. Unlike many network resource allocation performance optimization algorithms in the literatures which are able to handle convex network utility functions, the proposed scheme is able to handle generic nonlinear network utility functions. We will use a specific rate allocation and quality optimization problem for an example to demonstrate the efficiency of the proposed scheme and compare its performance with other algorithms, such as distributed gradient search.
Bo Wang 0005, Zhihai He, Yu Sun 0003
ICIP (6)2
2007 Low-Complexity and Reliable Moving Objects Detection and Tracking for Aerial Video Surveillance with Small UAVS
abstract
Moving objects detection and tracking is the first and enabling step for many high-level UAV surveillance tasks, including cooperative UAV path planning, navigation control, and automated information analysis. In this work, we develop a low-complexity and reliable moving object detection algorithm by exploring the ideas of uncertainty analysis and spatiotemporal activity clustering. More specifically, the authors develop a fast and efficient algorithm to estimate the global vehicle-camera motion. Image regions (blocks) with local motion was detected using statistical hypothesis testing. Using spatiotemporal clustering, the authors group these moving blocks into moving objects with physical meanings, such as moving vehicles or persons. Our extensive experimental results demonstrate the efficiency of the proposed algorithm.
Yu-Chia Chung, Zhihai He
ISCAS2
2007 Peak Transform for Efficient Image Representation and Coding
abstract
In this work, the author introduce a nonlinear geometric transform, called peak transform for efficient image representation and coding. The proposed peak transform is able to convert high-frequency signals into low-frequency ones, making them much easier to be compressed. To maximize the tranform coding gain, the author develop a dynamic programming solution for optimum peak transform design. Based on peak transform (PT), the author design an image encoder, called PT encoder, for efficient image compression. Our extensive experimental results demonstrate that the peak transform is able to reduce the signal energy in high-frequency subbands by up to 60%. The PT image encoder outperforms the state-of-the-art JPEG-2000 and H.264 (INTRA) encoders by up to 2-3 dB in PSNR (peak signal-to-noise ratio), especially for images with a significant amount of high-frequency components.
Zhihai He
ISCAS1
2007 Distributed Optimization Over Wireless Sensor Networks using Swarm Intelligence
abstract
In this work, we study how a group of sensor nodes in a wireless sensor network could collaborate with each other to perform complicated signal processing (e.g. estimation and tracking) and optimization tasks through local communication and distributed computation. We develop a distributed evolutionary optimization framework based on a swarm intelligence principle. During the optimization process, we only use and share local estimation results through communication links. This scheme can significantly reduce the communication energy cost and reach fast convergence. We use target localization as an example to evaluate the performance of the proposed distributed optimization algorithm. Our simulation results demonstrate that it outperforms existing distributed optimization algorithms, such as distributed gradient search.
Bo Wang 0005, Zhihai He
ISCAS2
2007 Image compression using constrained relaxation
abstract
In this work, we develop a new data representation framework, called constrained relaxation for image compression. Our basic observation is that an image is not a random 2-D array of pixels. They have to satisfy a set of imaging constraints so as to form a natural image. Therefore, one of the major tasks in image representation and coding is to efficiently encode these imaging constraints. The proposed data representation and image compression method not only achieves more efficient data compression than the state-of-the-art H.264 Intra frame coding, but also provides much more resilience to wireless transmission errors with an internal error-correction capability.
Zhihai He
VCIP1
2007 Cross-layer optimization for wireless video communication
abstract
With the rapid growth of wireless networks and increasing popularity of portable video devices, wireless video communication is poised to become the enabling technology for many multimedia applications over wireless networks. Real-time wireless video transmission typically has requirements on quality of service (QoS). However, wireless channels are unreliable and the channel capacities are time-varying, which may cause severe degradation to video presentation quality. In addition, for portable devices, video compression and wireless transmission are tightly coupled through the constraints on data rate, power, and delay. These issues make it particularly challenging to design an efficient real-time video compression and wireless transmission system on a portable device. In this paper, we take a cross-layer approach to this problem; our objective is to maximize the video quality under the constraints of resource and delay. Specifically, we minimize the end-to-end video distortion under the constraints of resource and delay, over the parameters in physical, link, and application (video) layers. This formulation is general and capable of capturing the fundamental aspects of the design of wireless video communication systems. Based on this formulation, we study how the resources could be intelligently allocated to maximize the video quality and analyze the performance limits of the wireless video communication system under resource constraints.
Dapeng Oliver Wu, Zhihai He
VCIP2
2007 Peak Transform for Efficient Image Representation and Coding
abstract
In this work, we introduce a nonlinear geometric transform, called peak transform (PT), for efficient image representation and coding. The proposed PT is able to convert high-frequency signals into low-frequency ones, making them much easier to be compressed. Coupled with wavelet transform and subband decomposition, the PT is able to significantly reduce signal energy in high-frequency subbands and achieve a significant transform coding gain. This has important applications in efficient data representation and compression. To maximize the transform coding gain, we develop a dynamic programming solution for optimum PT design. Based on PT, we design an image encoder, called the PT encoder, for efficient image compression. Our extensive experimental results demonstrate that, in wavelet-based subband decomposition, the signal energy in high-frequency subbands can be reduced by up to 60% if a PT is applied. The PT image encoder outperforms state-of-the-art JPEG2000 and H.264 (INTRA) encoders by up to 2-3 dB in peak signal-to-noise ratio (PSNR), especially for images with a significant amount of high-frequency components. Our experimental results also show that the proposed PT is able to efficiently capture and preserve high-frequency image features (e.g., edges) and yields significantly improved visual quality. We believe that the concept explored in this work, designing a nonlinear transform to convert hard-to-compress signals into easy ones, is very useful. We hope this work would motivate more research work along this direction.
Zhihai He
IEEE Trans. Image Process.1
2007 Adaptive motion estimation schemes using maximum mutual information criterion
abstract
Abstract We consider the motion estimation problem in video coding. In our previous work, we proposed a new motion estimation method where motion estimation is formulated as an optimization problem and an adaptive system under the minimum error entropy (MEE) criterion is used for motion estimation. In this paper, we develop an adaptive system under the criterion of maximum mutual information to address the motion estimation problem. Our proposed motion estimation algorithms have very low encoding complexity and hence are ideally suited for wireless video sensor networks where limited bandwidth, restricted computational capability, and limited battery power supply impose stringent constraints on the video encoding system. Copyright © 2007 John Wiley & Sons, Ltd.
Dapeng Oliver Wu, Deniz Erdogmus, Yuguang Fang, Zhihai He
Wirel. Commun. Mob. Comput.5
2006 Adaptive Silhouette Extraction in Dynamic Environments Using Fuzzy Logic
abstract
Extracting a human silhouette from an image is the enabling step for many high-level vision processing tasks, such as human tracking and activity analysis. Although there are a number of silhouette extraction algorithms proposed in the literature, most approaches work efficiently only in constrained environments where the background is relatively simple and static. In a previous paper, we addressed some of the challenges in silhouette extraction and human tracking in a real-world unconstrained environment where the background is complex and dynamic. We extracted features from image regions, accumulated the feature information over time, fused high-level knowledge with low-level features, and built a time-varying background model. A problem with our system is that by adapting the background model, objects moved by a human are difficult to handle. In order to reinsert them into the background, we run the risk of cutting off part of the human silhouette, such as in a quick arm movement. In this paper, we develop a fuzzy logic inference system to detach the silhouette of a moving object from the human body. Our experimental results demonstrate that the fuzzy inference system is very efficient and robust.
Zhihai He, James Keller 0001, Derek Anderson, Marjorie Skubic
FUZZ-IEEE2
2006 Adaptive Silouette Extraction and Human Tracking in Complex and Dynamic Environments
abstract
Extracting a human silhouette from an image is the enabling step for many high-level vision processing tasks, such as human tracking and activity analysis. Although there are a number of silhouette extraction and human tracking algorithms proposed in the literature, most approaches work efficiently only in constrained environments where the background is relatively simple and static. In this work, we propose to address the challenges in silhouette extraction and human tracking in a real-world unconstrained environment where the background is complex and dynamic. We extract features from image regions, accumulate the feature information over time, fuse the high-level knowledge with low-level features, and build a time-varying background model. We develop a fuzzy decision process to detach foreground moving objects from the human body. Our experimental results demonstrate that the algorithm is very efficient and robust.
Zhihai He, Derek Anderson, James Keller 0001, Marjorie Skubic
ICIP2
2006 Doubling of the Operational Lifetime of Portable Video Communication Devices Using Power-Rate-Distortion Analysis and Control
abstract
We consider a portable video device which captures and compresses video data. Video compression is computationally intensive and energy-consuming. However, the portable device, powered with batteries, has limited energy supply for data processing. One of the central challenging issues in portable video communication system design is how to maximize the operational lifetime. Based on our previous work on power-rate-distortion (P-R-D) analysis, in this work, we study how we could trade "bits" for "jouls" (energy) to minimize the energy consumption of video compression using optimum bit allocation. Our results, both theoretically and experimentally, show that using the proposed P-R-D video encoding technology, the operational lifetime of the portable device can be doubled or even tripled. This has a significant impact in energy-efficient portable video communication system design.
Wenye Cheng, Zhihai He
ICIP3
2006 Performance Optimization of Wireless Video Sensor Networks using Swarm Optimization with Convex Mapping
abstract
A wireless video sensor network (WVSN) is a system of spatially distributed video sensors which capture, process and transmit video information over a wireless ad hoc network. The performance optimization in WVSN is a nonlinear high-dimension constrained optimization problem. In this work, we consider the unique characteristics of WVSN and develop an evolutionary optimization scheme using a swarm intelligence principle to solve the WVSN performance optimization problem. We transform the solution space defined by flow balance and energy constraints into a convex region in a low-dimensional space. We then merge the convex condition with the swarm intelligence principle to guide the movement of each particle during the evolutionary optimization process. Our theoretical analysis and experimental results demonstrate that the proposed performance optimization scheme is very efficient.
Bo Wang 0005, Zhihai He
ICIP2
2006 An integrated energy aware wireless transmission system for QoS provisioning in wireless sensor network
Zongkai Yang, Zhihai He, Jianhua He 0001
Comput. Commun.3
2006 A novel frame-level bit allocation based on two-pass video encoding for low bit rate video streaming applications
Jianfei Cai 0001, Zhihai He, Chang Wen Chen
J. Vis. Commun. Image Represent.2
2006 Resource allocation and performance analysis of wireless video sensors
abstract
Wireless video sensor networks (WVSNs) have been envisioned for a wide range of important applications, including battlefield intelligence, security monitoring, emergency response, and environmental tracking. Compared to traditional communication system, the WVSN operates under a set of unique resource constraints, including limitations with respect to energy supply, on-board computational capability, and transmission bandwidth. The objective of this paper is to study the resource utilization behavior of a wireless video sensor and analyze its performance under the resource constraints. More specifically, we develop an analytic power-rate-distortion (P-R-D) model to characterize the inherent relationship between the power consumption of a video encoder and its rate-distortion performance. Based on the P-R-D analysis and a simplified model for wireless transmission power, we study the optimum power allocation between video encoding and wireless transmission and introduce a measure called achievable minimum distortion to quantify the distortion under a total power constraint. We consider two scenarios in wireless video sensing, small-delay wireless video monitoring and large-delay wireless video surveillance, and analyze the performance limit of the wireless video sensor in each scenario. The analysis and results obtained in this paper provide an important guideline for practical wireless video sensor design.
Zhihai He, Dapeng Oliver Wu
IEEE Trans. Circuits Syst. Video Technol.1
2006 Transmission Distortion Analysis for Real-Time Video Encoding and Streaming Over Wireless Networks
abstract
A major challenge in video encoding and transmission over wireless networks is that the channel is error-prone while the compressed video data is highly sensitive to errors. The transmission errors will cause decoding failure at the receiver side. More importantly, the transmission errors introduced in one video frame will propagate to its subsequent frames along the motion prediction path and significantly degrade video presentation quality. This type of picture distortion is called transmission distortion. In this work, we propose a control system approach to transmission distortion modeling. More specifically, we consider the wireless video transmission and decoding system as a linear system with the transmission errors as system input and the transmission distortion as output. We study the fading behavior of the impulse transmission distortion. We analyze the low-pass filtering behavior of video encoders and develop a scheme to estimate the instantaneous transmission distortion. Our extensive simulation results demonstrate that the proposed transmission distortion model is accurate and robust. In addition, it has very low computational complexity. More importantly, because it is a predictive model, it allows the encoder to predict the transmission distortion even before the video is compressed and transmitted. This type of transmission distortion model has important applications in resource allocation and performance optimization in real-time wireless video communication
Zhihai He, Hongkai Xiong
IEEE Trans. Circuits Syst. Video Technol.1
2005 Transmission distortion modeling for wireless video communication
abstract
In video transmission over wireless networks, the channel is error-prone and the compressed video data is highly sensitive to errors. The transmission errors causes decoding failure and distort the pictures. This type of picture distortion is called transmission distortion. In this work, we develop a predictive approach for transmission modeling so that the encoder is able to estimate the average transmission distortion of the decoded video data. Our simulation results demonstrate that the proposed model is robust and accurate. Using this type of predictive transmission distortion model, the encoder is able to intelligently allocate its resources, such as bits and energy, to maximize the decoded video quality at receiver side.
Janak U. Dani, Zhihai He, Hongkai Xiong
GLOBECOM2
2005 Performance analysis of wireless video sensors in video surveillance
abstract
Wireless video sensor networks (WVSN) have been envisioned for a wide range of important applications, including battlefield intelligence, security monitoring, and environmental tracking. Compared to traditional communication systems, the WVSN operates under a set of unique resource constraints, including limitations with respect to energy supply, on-board computational capability, and transmission bandwidth. The objective of this work is to study the resource utilization behavior of a wireless video sensor and analyze its performance under these resource constraints. More specifically, we develop an analytic power-rate-distortion (P-R-D) model to characterize the inherent relationship between the power consumption of a video encoder and its rate-distortion performance. Based on the P-R-D analysis and a simplified model for wireless transmission power, we study the optimum power allocation between video encoding and wireless transmission. We consider an important scenario in wireless video sensing - wireless video surveillance, and analyze the performance limit of the wireless video sensor. The analysis and results obtained in this paper provide an important guideline for practical wireless video sensor design.
Zhihai He, Dapeng Oliver Wu
GLOBECOM1
2005 Accumulative visual information in wireless video sensor network: definition and analysis
abstract
A wireless video sensor network (WVSN) is a system of spatially distributed video sensors which gather and transmit video information over a wireless ad hoc network. To measure, control and optimize the system performance, the following key research problems need to be addressed: 1) how to measure the amount of visual information collected by the video sensor within its operational lifetime; 2) how to quantitatively compare the information sensing efficiency between two video sensors; 3) how to measure the information sensing efficiency of the whole video sensor network; 4) how to control and optimize the information sensing efficiency of the video sensor network. In this work, we introduce the concept of accumulative visual information (AVT), and use it as a measure for the amount of visual information collected the WVSN. Based on the AVI measure and the power-rate-distortion analysis model developed in our previous work, we optimize the efficiency of the WVSN system.
Zhihai He, Dapeng Oliver Wu
ICC1
2005 MPEG-4 video streaming quality evaluation in IEEE 802.11e WLANs
abstract
The IEEE 802.11 working group is currently working on a new standard called IEEE 802.11e to support quality of service (QoS) in WLANs. 802.11e introduces a so-called hybrid coordination function (HCF) containing two medium access mechanisms: enhanced distributed channel access (EDCA) and HCF controlled channel access (HCCA). In the EDCA mechanism, many QoS parameters are introduced including minimum contention window (CWmin), maximum contention window (CWmax), arbitration inter frame space (AIFS) and transmission opportunity limit (TXOPlimit). In this paper, we experimentally assess the MPEG-4 video streaming performance over 802.11e. In particular, we discuss in detail how the human satisfaction of streaming video is affected by the main QoS parameters in IEEE 802.11e WLANs. We measure the level of end user satisfaction together with the network performance and give recommendations regarding the network design and the parameter settings.
Deyun Gao, Jianfei Cai 0001, Paul Bao, Zhihai He
ICIP (1)4
2005 Joint mode selection and unequal error protection for bitplane coded video transmission over wireless channels
Jianfei Cai 0001, Jianhua Wu 0003, King Ngi Ngan, Zhihai He
J. Vis. Commun. Image Represent.4
2005 A model-based adaptive motion estimation scheme using Renyi's entropy for wireless video
Ganesan Ramachandran, Vignesh Krishnan, Dapeng Oliver Wu, Zhihai He
J. Vis. Commun. Image Represent.4
2005 Power-Rate-Distortion Analysis for Wireless Video Communication Under Energy Constraints
abstract
Mobile devices performing video coding and streaming over wireless and pervasive communication networks are limited in energy supply. To prolong the operational lifetime of these devices, an embedded video encoding system should be able to adjust its computational complexity and energy consumption as demanded by the situation and its environment. To analyze, control, and optimize the rate-distortion (R-D) behavior of the wireless video communication system under the energy constraint, we develop a power-rate-distortion (P-R-D) analysis framework, which extends the traditional R-D analysis by including another dimension, the power consumption. Specifically, in this paper, we analyze the encoding mechanism of typical video coding systems, and develop a parametric video encoding architecture which is fully scalable in computational complexity. Using dynamic voltage scaling (DVS), an energy consumption management technology recently developed in CMOS circuits design, the complexity scalability can be translated into the energy consumption scalability of the video encoder. We investigate the R-D behavior of the complexity control parameters and establish an analytic P-R-D model. Both theoretically and experimentally, we show that, using this P-R-D model, the video coding system is able to automatically adjust its complexity control parameters to match the available energy supply of the mobile device while maximizing the picture quality. The P-R-D model provides a theoretical guideline for system design and performance optimization in mobile video communication under energy constraints.
Zhihai He, Yongfang Liang, Lulin Chen, Ishfaq Ahmad 0001, Dapeng Oliver Wu
IEEE Trans. Circuits Syst. Video Technol.1
2005 Low-pass filtering of rate-distortion functions for quality smoothing in real-time video communication
abstract
In variable-bit-rate video coding, the video is preprocessed to collect sequence-level statistics, which are used for global bit allocation in the actual encoding stage to obtain a smoothed video presentation quality. However, in real-time video recording and network streaming, this type of two-pass encoding scheme is not allowed because the access to future frames and global statistics is not available. To address this issue, we introduce the concept of low-pass filtering of rate-distortion functions and develop a smoothed rate control (SRC) framework for real-time video recording and streaming. Theoretically, we prove that, using a geometric averaging filter, the SRC algorithm is able to maintain a smoothed video presentation quality while achieving the target bit rate automatically. We also analyze the buffer requirement of the SRC algorithm in real-time video streaming, and propose a scheme to seamlessly integrate robust buffer control into the SRC framework. The proposed SRC algorithm has very low computational complexity and implementation cost. Our extensive experimental results demonstrate that the SRC algorithm significantly reduces the picture quality variation in the encoded video clips.
Zhihai He, Wenjun Zeng 0001, Chang Wen Chen
IEEE Trans. Circuits Syst. Video Technol.1
2004 Optimal retransmission timeout selection for delay-constrainfd multimedia communications
Jianfei Cai 0001, Wenning Zhan, Zhihai He
ICIP3
2004 Lowpass filtering of rate-distortion functions for quality smoothing for real-time video recording and streaming
abstract
In this work, we introduce the concept of low-pass filtering of rate-distortion (R-D) functions and develop a smoothed rate control (SRC) framework for real-time video recording and streaming. Theoretically, we prove that using a geometric averaging filter the SRC algorithm is able to maintain a very smooth video presentation quality while achieving the target bit rate automatically. The proposed SRC algorithm has very low computational complexity and implementation cost. Our extensive experimental results demonstrate that the proposed SRC algorithm significantly reduces the picture quality variation in the encoded video clips while matching the encoding bit rate target very accurately.
Zhihai He, Chang Wen Chen, Jianfei Cai 0001
ICIP1
2004 A half D1 MPEG-4 encoder on the BSP-15 DSP
abstract
In this paper, we present the work on implementation of a half-D1interlaced MPEG-4 encoder with Equator Technology DSP chip, BSP-15. The BSP-15 DSP consists mainly of a VLIW core, Co-processors, and media I/O interfaces. The encoder utilizes several BSP-15 functional blocks in parallel. In general, the VLIW performs pixel procesing that is computationally intensive. The VLx coprocessor completes variable length coding. Further parallelism is obtained by pre-loading data cache and doubling data buffers. Given the DSP processing power and real time requirements, a complexity control scheme is implemented. A frame-level quantization scheme with quality and rate control is employed. The current implementation for video at 30 fps consumes about 90% of the chip performance at a bit rate ~2Mbps.
Lulin Chen, Zhihai He, Chang Wen Chen, Michael A. Isnardi
VCIP2
2004 Automatic video object detection and mask signal removal for efficient video preprocessing
abstract
In this work, we consider a generic definition of video object, which is a group of pixels with temporal motion coherence. The generic video object (GVO) is the superset of the conventional video objects discussed in the literature. Because of its motion coherence, the GVO can be easily recognized by the human visual system. However, due to its arbitray spatial distribution, the GVO cannot be easily detected by the existing algorithms which often assume the spatial homogeneousness of the video objects. In this work, we introduce the concept of extended optical flow and develop a dynamic programming framework for the GVO detection. Using this mathematical optimization formulation, whose solution is given by the the Viterbi algorithm, the proposed object detection algorithm is able to discover the motion path of the GVO automatically and refine its spatial location progressively. We apply the GVO detection algorithm to extract and remove the so-called "video mask" signals in the video sequence. Our experimental results show that this type of vision-guided video pre-processing significantly improves the compression efficiency.
Zhihai He
VCIP1
2004 Effective quality analysis for video streaming over wireless ad hoc network
abstract
Video encoding and streaming over wireless ad hoc network operates under severe conditions, such as time-varying channel characteristics with bursty errors, limit power for data transmission, and dynamic topology of the self-organized network, stringent time delay for packet delivery, etc. Due to the dynamic topology, complex mechanism, and time-varying nature of the wireless ad hoc network, the network system and the streaming service often exhibit unpredictable behaviors. The ultimate goal in the wireless video streaming service is to provide the end user with the best possible video presentation quality. The video streaming quality, often measured by the end-to-end picture distortion, is affected by the scene coding characteristics S of the input video, and the configuration of the network parameters N which includes bandwidth, transmission power, bit error ratio, delay, etc. This brings up the following important and challenging issue: given an input video with scene characteristics S and a network configuration N, what is the average video streaming quality the receiver could expect the system to provide. To address this issue, in this work, we examine the behaviors and constraints of the major components in the streaming system, and propose an across-layer framework to model, control, and optimize the end-to-end video streaming quality.
Zhihai He, Chang Wen Chen
VCIP1
2004 Power-rate-distortion analysis for wireless video communication under energy constraint
abstract
In video coding and streaming over wireless communication network, the power-demanding video encoding operates on the mobile devices with limited energy supply. To analyze, control, and optimize the rate-distortion (R-D) behavior of the wireless video communication system under the energy constraint, we need to develop a power-rate-distortion (P-R-D) analysis framework, which extends the traditional R-D analysis by including another dimension, the power consumption. Specifically, in this paper, we analyze the encoding mechanism of typical video encoding systems and develop a parametric video encoding architecture which is fully scalable in computational complexity. Using dynamic voltage scaling (DVS), a hardware technology recently developed in CMOS circuits design, the complexity scalability can be translated into the power consumption scalability of the video encoder. We investigate the rate-distortion behaviors of the complexity control parameters and establish an analytic framework to explore the P-R-D behavior of the video encoding system. Both theoretically and experimentally, we show that, using this P-R-D model, the encoding system is able to automatically adjust its complexity control parameters to match the available energy supply of the mobile device while maximizing the picture quality. The P-R-D model provides a theoretical guideline for system design and performance optimization in wireless video communication under energy constraint, especially over the wireless video sensor network.
Zhihai He, Yongfang Liang, Ishfaq Ahmad 0001
VCIP1
2003 Single-pass distortion-smoothing encoding for low bit-rate video streaming applications
abstract
This paper proposes a rate control scheme for smoothing distortion in low bit-rate video coding. We focus on single-pass encoding that is applicable to both live and off-line streaming. The new rate control tends to achieve slow and smooth distortion variation over time. Without the knowledge of future frames, statistics of previously coded frames are used to derive the expected distortion for the current frame. Constraints of decoder buffer size and pre-loading time are considered in the design. The proposed technique is based on the rate and distortion models used in the TMN8. Its advantages have been shown by experiments.
Tao Chen 0044, Zhihai He
ICME2
2002 Optimal bit allocation for low bit rate video streaming applications
abstract
Current rate control schemes in video coding standards do not have efficient frame-level bit allocation because of the inherent constraints in real-time encoding. In this paper, we assume an offline video encoding environment and proposed a rate control scheme based on optimal bit allocation for low bit rate streaming applications. Specifically, we apply a /spl rho/-domain rate-distortion (R-D) model, originally applied at macroblock (MB) level, to frame-level. Based on this frame-level R-D model and a two-pass encoding method, we are able to allocate bits among video frames in an optimal way so that video sequences can be coded at low bit rate with an improved quality. Experimental results demonstrate the proposed scheme is able to achieve not only noticeable reduction in average distortion but also a more consistent and smoother visual quality.
Jianfei Cai 0001, Zhihai He, Chang Wen Chen
ICIP (1)2
2002 Video coding using joint temporal-spatial compensation
abstract
In motion compensation (MC) based video coding system, macroblock data is always predicted by its temporal neighborhood in its previous frame. In this work, we observe that the macroblock data can often be better predicted by its spatial neighborhood than its temporal neighborhood, especially at low coding bit rates. Based on this observation, we propose a joint temporal-spatial compensation (JTSC) scheme for video coding, with the conventional MC being its subset. Our experimental results demonstrate that the coded picture quality is significantly improved by this novel JTSC scheme.
Zhihai He, Chang Wen Chen
ICME (1)1
2002 End-to-end video quality analysis and modeling for video streaming over IP network
abstract
We derive an analytical end-to-end video quality prediction and control model for video streaming over an IP network. This model maps the quality of service (QoS) parameters defined at the connection level to actual video presentation quality at the receiver end. Specifically, we provide an analysis on the effects of packet loss on the decoded video quality and relate the packet loss ratio, one of the most important QoS parameters at the connection level, to the actual video quality. The model also characterizes the rate-distortion (R-D) behavior of the encoder at the video sequence level, predicting the average encoding quality for a given channel bandwidth. Our experimental results demonstrate the accuracy of the proposed video quality model.
Zhihai He, Chang Wen Chen
ICME (1)1
2002 Encoder-based rate smoothing and quality control for low-delay video coding and communication
Zhihai He, Chang Wen Chen
VCIP1
2002 ?-domain optimum bit allocation and accurate rate control for DCT video coding
Zhihai He, Chang Wen Chen
VCIP1
2002 Joint source channel rate-distortion analysis for adaptive mode selection and rate control in wireless video coding
abstract
We first develop a rate-distortion (R-D) model for DCT-based video coding incorporating the macroblock (MB) intra refreshing rate. For any given bit rate and intra refreshing rate, this model is capable of estimating the corresponding coding distortion even before a video frame is coded. We then present a theoretical analysis of the picture distortion caused by channel errors and the subsequent inter-frame propagation. Based on this analysis, we develop a statistical model to estimate such channel errors induced distortion for different channel conditions and encoder settings. The proposed analytic model mathematically describes the complex behavior of channel errors in a video coding and transmission system. Unlike other experimental approaches for distortion estimation reported in the literature, this analytic model has very low computational complexity and implementation cost, which are highly desirable in wireless video applications. Simulation results show that this model is able to accurately estimate the channel errors induced distortion with a minimum delay in processing. Based on the proposed source coding R-D model and the analytic channel-distortion estimation, we derive an analytic solution for adaptive intra mode selection and joint source-channel rate control under time-varying wireless channel conditions. Extensive experimental results demonstrate that this scheme significantly improves the end-to-end video quality in wireless video coding and transmission.
Zhihai He, Jianfei Cai 0001, Chang Wen Chen
IEEE Trans. Circuits Syst. Video Technol.1
2002 . Optimum bit allocation and accurate rate control for video coding via ρ-domain source modeling
abstract
We present a new framework for rate-distortion (R-D) analysis, where the coding rate R and distortion D are considered as functions of /spl rho/ which is the percentage of zeros among the quantized transform coefficients. Previously (see He, Z. et al., Int. Conf. Acoustics, Speech and Sig. Proc., 2001), we observed that, in transform coding of images and videos, the rate function R(/spl rho/) is approximately linear. Based on this linear rate model, a simple and unified rate control algorithm was proposed for all standard video coding systems, such as MPEG-2, H.263, and MPEG-4. We further develop a distortion model and an optimum bit allocation scheme in the /spl rho/ domain. This bit allocation scheme is applied to MPEG-4 video coding to allocate the available bits among different video objects. The bits target of each object is then achieved by our /spl rho/-domain rate control algorithm. When coupled with a macroblock classification scheme, the above bit allocation and rate control scheme can also be applied to other video coding systems, such as H.263, at the macroblock level. Our extensive experimental results show that the proposed algorithm controls the encoder bit rate very accurately and improves the video quality significantly (by up to 1.5 dB).
Zhihai He, Sanjit K. Mitra
IEEE Trans. Circuits Syst. Video Technol.1
2002 A linear source model and a unified rate control algorithm for DCT video coding
abstract
We show that, in any typical transform coding system, there is always a linear relationship between the coding bit rate R and the percentage of zeros among the quantized transform coefficients, denoted by /spl rho/. Based on Shannon's source coding theorem, a theoretical justification is provided for this linear source model. The physical meaning of the model parameter is also discussed. We show that it is directly related to the image content and is a measure of picture complexity. In video coding, we propose an adaptive estimation scheme to estimate this model parameter. Based on the linear source model and the adaptive estimation scheme, a unified rate control algorithm is proposed for various standard video coding systems, such as MPEG-2, H.263, and MPEG-4. Our extensive simulation results show that the proposed rate control outperforms other algorithms reported in the literature by providing much more accurate and robust rate control.
Zhihai He, Sanjit K. Mitra
IEEE Trans. Circuits Syst. Video Technol.1
2001 ρ-domain source modeling and rate control for video coding and transmission
abstract
The coding bit rate, R, is considered as a function of /spl rho/ which is the percentage of zeros among the quantized DCT (discrete cosine transform) coefficients. We discover that the rate function, R(/spl rho/), has some very interesting properties in the /spl rho/-domain. By introducing the new concepts of characteristic rate curves and rate curve decomposition, a novel framework for source modeling is proposed. Using the proposed source model, we can estimate the rate-quantization (R-Q) curve before quantization and coding with relative error less than 5%. Based on the estimated R-Q curve, the output bit rate of the video encoder can be accurately controlled. Our extensive simulation results show that the proposed algorithm outperforms TMN8 (see Ribas-Corbera, J. and Lei, S., IEEE Trans. on Circuits and Systems for Video Technology, vol.9, p.172-85, 1999) and VM7 (see Chiang, T. and Zhang, Y.-Q., IEEE Trans. on Circuits and Systems for Video Technology, vol.7, p.246-50, 1997) rate control algorithms by providing more accurate and robust rate control.
Zhihai He, Yong Kwan Kim, Sanjit K. Mitra
ICASSP1
2001 A novel linear source model and a unified rate control algorithm for H.263/MPEG-2/MPEG-4
abstract
Let /spl rho/ be the percentage of zeros among the quantized transform coefficients. We discover that, in any typical video coding system, there is always a strictly linear relationship between /spl rho/ and the actual coding bit rate R. This linearity leads to a novel and unified source model for different types of source data and different coding systems, such as H.263, MPEG-2, and MPEG-4. The proposed linear source model is much simpler, but much more accurate than other source models reported in the literature. Based on this source model, a unified rate control algorithm is proposed for the above three video coding systems. Despite its extreme simplicity, the proposed algorithm outperforms other rate control algorithms by providing more accurate and robust rate control.
Yong Kwan Kim, Zhihai He, Sanjit K. Mitra
ICASSP2
2001 P-domain bit allocation and rate control for real time video coding
abstract
A novel framework for rate-distortion (RD) analysis is developed. Based on this framework, an optimum bit allocation scheme and an accurate rate control algorithm are proposed for real-time video coding with H.263+ (see Cote, G. et al., IEEE Trans. on Circuits and Systems for Video Technology, vol.8, p.849-66, 1998). With the proposed algorithm, the picture quality is significantly improved. The output bit rate of the video encoder is controlled robustly and accurately according to the network condition.
Zhihai He, Sanjit K. Mitra
ICIP (3)1
2001 Low-delay rate control for DCT video coding via ?-domain source modeling
abstract
By introducing the new concepts of characteristic rate curves and rate curve decomposition, a generic source-modeling framework is developed for transform coding of videos. Based on this framework, the rate-quantization (R-Q) and distortion-quantization (D-Q) functions (collectively called R-D functions in this work) of the video encoder can be accurately estimated with very low computational complexity before quantization and coding. With the accurate estimation of the R-Q function, a frame-level rate control algorithm is proposed for DCT video coding. The proposed algorithm outperforms the TMN8 rate control algorithm by providing more accurate and robust rate regulation and better picture quality. Based on the estimated R-D functions, an encoder-based rate-shape-smoothing algorithm is proposed. With this smoothing algorithm, the output bit stream of the encoder has both a smoothed rate shape and a consistent picture quality, which are highly desirable in practical video coding and transmission.
Zhihai He, Yong Kwan Kim, Sanjit K. Mitra
IEEE Trans. Circuits Syst. Video Technol.1
2001 A unified rate-distortion analysis framework for transform coding
abstract
In our previous work, we have developed a rate-distortion (R-D) modeling framework for H.263 video coding by introducing the new concepts of characteristic rate curves and rate curve decomposition. In this paper, we further show it is a unified R-D analysis framework for all typical image/video transform coding systems, such as embedded zero-tree wavelet (EZW), set partitioning in hierarchical trees (SPIHT) and JPEG image coding; MPEG-2, H.263, and MPEG-4 video coding. Based on this framework, a unified R-D estimation and control algorithm is proposed for all typical transform coding systems. We have also provided a theoretical justification for the unique properties of the characteristic rate curves. A linear rate regulation scheme is designed to further improve the estimation accuracy and robustness, as well as to reduce the computational complexity of the R-D estimation algorithm. Our extensive experimental results show that with the proposed algorithm, we can accurately estimate the R-D functions and robustly control the output bit rate or picture quality of the image/video encoder.
Zhihai He, Sanjit K. Mitra
IEEE Trans. Circuits Syst. Video Technol.1
2000 Optimal Quantization Error Feedback Filters for Wavelet Image Compression
abstract
In wavelet image compression, all information loss occurs during quantization. Based on the property of bi-orthogonal wavelets, optimal 2-D quantization error feedback filters are used to reduce the reconstruction error. With very low complexity with regard to computation and implementation, the error feedback system improves the PSNR of the reconstructed image by about 0.25 dB. In addition, due to its similar structure to the dithered quantizer, it also improves the subjective quality of the reconstructed image by reducing the contouring and ranging effects.
Zhihai He, Sanjit K. Mitra
ICIP1
2000 Blockwise Zero Mapping Image Coding
abstract
A coding algorithm of low addressing and implementation complexity is proposed. It is based on partitioning uniformly quantized wavelet coefficients into multiscale blocks with each block classified as either an all-zero block or a non-zero block. The non-zero blocks are coded by a new method called zero-mapping whose outputs are further compressed by a first-order arithmetic coder. The performance of the proposed coding algorithm compares favorably with that of some well-known coding algorithms, particularly for images with considerable high frequency components.
Zhihai He, Tian-Hu Yu, Sanjit K. Mitra
ICIP1
2000 Simple and Efficient Wavelet Image Compression
abstract
We propose a simple but efficient wavelet image compression algorithm. The proposed coding scheme employs the multi-level dyadic wavelet decomposition, linear quantization with a proper dead zone, 1-D addressing complexity by raster scanning within subbands, variable length block coding, small alphabet representation of 1-D integer sequences, and adaptive arithmetic entropy coding. Despite the simplicity of the proposed coding scheme, the rate-distortion performance of the proposed image compression algorithm is competitive with the best image coders in the literature.
Tian-Hu Yu, Zhihai He, Sanjit K. Mitra
ICIP2