VLDB 2026 Research / reviewers in the wild / expert
King Ngi Ngan
dblp:n/KingNgiNgan · also King N. Ngan
· DBLP profile ↗
288ranked-venue papers
15as first author
25since 2021 · last 2026
0000-0003-1946-3235ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 241 · 12 first-author · 20 since 2021Systems, architecture and hardware · 26 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 24 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4Computer networks · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Zero-shot egocentric action recognition via chain-of-imagination prompts and inertial strengthening adaptor
Mingzhou He, Ruiqian Li, Qingbo Wu 0001, King Ngi Ngan, Fanman Meng, Hongliang Li 0001 |
Pattern Recognit. | 4 |
| 2026 | On the Adversarial Robustness of Learning-Based Image Compression Against Rate-Distortion AttacksabstractDespite demonstrating superior Rate-Distortion (RD) performance, Learning-based Image Compression (LIC) algorithms have been found to be vulnerable to malicious perturbations in recent studies. However, the adversarial attacks considered in existing literature remain divergent from real-world scenarios, both in terms of the attack direction and bitrate. Additionally, existing methods focus solely on empirical observations of the model vulnerability, neglecting to identify the origin of it. These limitations hinder the comprehensive investigation and in-depth understanding of the adversarial robustness of LIC algorithms. To address the aforementioned issues, this paper considers the arbitrary nature of the attack direction and the uncontrollable compression ratio faced by adversaries, and presents two practical rate-distortion attack paradigms,i.e., Specific-ratio Rate-Distortion Attack (SRDA) and Agnostic-ratio Rate-Distortion Attack (ARDA). To the best of our knowledge, we are the first to conduct joint rate-distortion attacks on LIC algorithms. Using the performance variations as indicators, we evaluate the adversarial robustness of eight predominant LIC algorithms against diverse attacks. Furthermore, we propose two novel analytical tools for in-depth analysis,i.e., Entropy Causal Intervention and Layer-wise Distance Magnify Ratio, and reveal thathyperpriorsignificantly increases the bitrate andInverse Generalized Divisive Normalization (IGDN)significantly amplifies input perturbations when under attack. Lastly, we examine the efficacy of adversarial training and introduce the use of online updating for defense. By comparing their advantages and disadvantages, we provide a reference for constructing more robust LIC algorithms against the rate-distortion attacks. Qingbo Wu 0001, Lei Wang 0186, Fanman Meng, King Ngi Ngan, Li Zhuo 0001, Hongliang Li 0001 |
IEEE Trans. Multim. | 6 |
| 2025 | Your Demands Deserve More Bits: Referring Semantic Image Compression at Ultra-low BitrateabstractWith the help of powerful generative models, Semantic Image Compression (SIC) has achieved impressive performance at ultra-low bitrate. However, due to coarse-grained visual-semantic alignment and inherent randomness, the reliability of SIC is seriously concerned for reconstructing completely different object instances, even they are semantically consistent with original images. To tackle this issue, we propose a novel Referring Semantic Image Compression (RSIC) framework to improve the fidelity of user-specified content while retaining extreme compression ratios. Specifically, RSIC consists of three modules: Global Description Encoding (GDE), Referring Guidance Encoding (RGE), and Guided Generative Decoding (GGD). GDE and RGE encode global semantic information and local features, respectively, while GGD handles the non-uniformly guided generative process based on the encoded information. In this way, our RSIC achieves flexible customized compression according to user demands, which better balance the local fidelity, global realism, semantic alignment, and bit overhead. Extensive experiments on three datasets verify the compression efficiency and flexibility of the proposed method. Qingbo Wu 0001, Mingzhou He, King Ngi Ngan, Fanman Meng, Hongliang Li 0001 |
ISCAS | 6 |
| 2025 | High efficiency deep image compression via channel-wise scale adaptive latent representation learning
Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001 |
Signal Process. Image Commun. | 3 |
| 2025 | Learning With Noisy Low-Cost MOS for Image Quality Assessment via Dual-Bias CalibrationabstractLearning-based Image Quality Assessment (IQA) models have obtained impressive performance with the help of reliable subjective quality labels, where Mean Opinion Score (MOS) is the most popular choice. However, in view of the subjective bias of individual annotators, the Labor-Abundant MOS (LA-MOS) typically requires large collections of opinion scores from multiple annotators for each image, which significantly increases the learning cost. In this paper, we aim to learn robust IQA models from Low-Cost MOS (LC-MOS), which only requires very few opinion scores or even a single opinion score for each image. More specifically, we consider the LC-MOS as the noisy observation of LA-MOS and enforce the IQA model learned from LC-MOS to approach the unbiased estimation of LA-MOS. Thus, we represent the subjective bias between LC-MOS and LA-MOS, and the model bias between IQA predictions learned from LC-MOS and LA-MOS (i.e., dual-bias) as two latent variables with unknown parameters. By means of the expectation-maximization-based alternating optimization, we can jointly estimate the parameters of the dual-bias, which suppresses the misleading of LC-MOS via a gated dual-bias calibration (GDBC) module. To the best of our knowledge, this is the first exploration of robust IQA model learning from noisy low-cost labels. Theoretical analysis and extensive experiments on four popular IQA datasets show that the proposed method is robust toward different bias rates and annotation numbers and significantly outperforms the other Learning-based IQA models when only LC-MOS is available. Furthermore, we also achieve comparable performance with respect to the other models learned with LA-MOS. Lei Wang 0029, Qingbo Wu 0001, Desen Yuan, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Robust Real-World Image Dehazing via Knowledge Guided Conditional Diffusion Model FinetuningabstractDue to the domain gap, the dehazing models trained from the synthetic images suffer poor generalization performance on real-world images. To address this issue, we pro-pose a Knowledge guided Conditional Diffusion (KCDiff) model finetuning method, which enables both the domain knowledge adaptation from the synthetic images and general knowledge guidance from the real-world images. More specifically, our KCDiff comprises two modules, i.e., the Conditional Image Generation (CIG) and Dehazing Instruction Generation (DIG). For CIG, we freeze a pre-trained latent diffusion model, add learnable conditioning control layers with Low-Rank Adaptation (LoRA) blocks, and include skip connections with zero-initialized convolutional layers, all of which play a fundamental role in image dehazing. Meanwhile, the DIG utilizes a large vision-language model LLaVA to extract the semantic content of the input hazy image and redescribe it in clear weather, which serves as the control instruction of CIG. To mitigate potential artifacts in CIG caused by misinterpretation of DIG's instructions, we further enforce depth and physical model-based reconstruction consistency constraints on both dehazing and hazy images. In the training phase, CIG is trained with the paired synthetic images to adapt the diffusion prior to the domain knowledge of image dehazing and finetuned with unpaired real-world images to suppress the domain gap with general knowledge guidance from the atmospheric scattering and depth perception. Experiments on real-world databases demonstrate the superiority of the proposed method over many state-of-the-art image dehazing models. Qingbo Wu 0001, Lei Wang 0186, King Ngi Ngan, Fanman Meng, Hongliang Li 0001 |
MMSP | 6 |
| 2024 | IoU-CLIP: IoU-Aware Language-Image Model Tuning for Open Vocabulary Object DetectionabstractOpen vocabulary object detection (OVD), which detects novel categories through detectors trained on base categories, has achieved remarkable advancement attributable to large-scale vision-language models, such as CLIP. The prior OVD works mainly focused on improving the classification accuracy of proposals, ignoring the ability of localization for novel categories. In this work, we propose IoU-aware language-image model tuning (IoU-CLIP) for open vocabulary object detection. Specifically, we construct a region image dataset with different IoU and adopt IoU values as labels to fine-tune the CLIP model to learn IoU-aware and class-agnostic semantic prompts and visual embeddings. The fine-tuned IoU-CLIP can predict IoU scores for proposals, which interact with classification scores. Meanwhile, IoU-aware and class-agnostic visual embeddings are utilized for box regression to enhance the generalization of the localization capability. We evaluate our method on the COCO and LVIS OVD benchmarks, outperforming the baseline (RegionCLIP) by 5.5% AP50and 5.8% AP on novel categories, respectively, achieving state-of-the-art performance. Mingzhou He, Qingbo Wu 0001, King Ngi Ngan, Fanman Meng, Heqian Qiu, Hongliang Li 0001 |
VCIP | 3 |
| 2024 | Robust Unpaired Image Dehazing via Adversarial Deformation ConstraintabstractDue to the flexible training requirement and the appealing generalization ability, unpaired image dehazing has received increasing attention in coping with real-world hazy images. However, most of the existing methods rely on the loose dehazing-hazing cycle constraint, which makes it hard to eliminate poor-quality dehazing results when using a powerful hazing network in the training process. To address this issue, this paper proposes a simple yet efficient Adversarial Deformation Constraint (ADC). More specifically, we sequentially perform two operations, i.e., dehazing and deformation, on a hazy image. In the training process, the dehazing branch is desired to be deformation-unaware, which requires that the output of these two operations remains constant regardless of their performing order. Adversarially, the deformation branch tends to maximize the difference in the outputs of these two operations when their performing orders are different. Through an additive image decomposition model, we verify that the ADC could regularize the solution space to push the dehazing error towards zero. Finally, by incorporating ADC into the common dehazing-hazing cycle constraint, we significantly improve the robustness of unpaired image dehazing. Experiments on multiple benchmark hazy image databases demonstrate the superiority of ADC over many state-of-the-art image dehazing methods. The source code of the proposed ADC-Net will be released on https://github.com/whrws/ADC-Net. Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Heqian Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Continual Cross-Domain Image Compression via Entropy Prior Guided Knowledge Distillation and Scalable DecodingabstractLearning based image compression has achieved impressive rate-distortion performance in recent years. However, due to the disposable learning strategy and rigid network architecture, existing methods perform poorly for compressing the images of different domains when they emerge with the expanding real-world applications, such as, natural, oil painting, medical images and so on. To cope with this open-world challenge, this paper proposes a continual cross-domain image compression method based on entropy prior guided knowledge distillation and scalable decoding network, which perform well in balancing the plasticity, stability and compatibility. Firstly, we generate pseudo-samples of old domains by reusing their entropy priors. These pseudo-samples serve as guides for knowledge distillation in the old domains, ensuring that the bit rate and reconstruction of the new model align with those of the old model. This approach assists the updated model in retaining its capability to compress and reconstruct old images. Secondly, we develop a scalable decoding network via dynamic pruning and masked recovery, which could effectively infer an old entropy decoder from the latestly updated model. It ensures that the updated model could decode image features from binary strings encoded by old entropy encoders. Experiments on five image datasets with different domains demonstrate the effectiveness of the proposed method and its superiority over representative continual learning methods. Code of the proposed method is available athttps://github.com/wuchenhaoo/Continual_Cross-domain_Image_Compression/. Qingbo Wu 0001, Rui Ma 0030, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Heqian Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Cross-Modal Recurrent Semantic Comprehension for Referring Image SegmentationabstractReferring image segmentation aims to segment the target object from the image according to the description of language expression. Due to the diversity of language expressions, word sequences in different orders often express different semantic information. The previous methods focus more on matching different words to different visual regions in the image separately, ignoring the global semantic understanding of language expression based on the sequence structure. To address this problem, we redesign a new recurrent network structure for referring image segmentation, called Cross-Modal Recurrent Semantic Comprehension Network (CRSCNet), to obtain a more comprehensive global semantic understanding through iterative cross-modal semantic reasoning. Specifically, in each iteration, we first propose a Dynamic SepConv to extract relevant visual features guided by language and further propose Language Attentional Feature Modulation to improve the feature discriminability, then propose a Cross-Modal Semantic Reasoning module to perform global semantic reasoning by capturing both linguistic and visual information, and finally updates and corrects the visual features of the predicted object based on semantic information. Moreover, we further propose a Cross-Modal ASPP to capture richer visual information referred to in the global semantics of the language expression from larger receptive fields. Extensive experiments demonstrate that our proposed network significantly outperforms previous state-of-the-art methods on multiple datasets. Chao Shang 0001, Hongliang Li 0001, Heqian Qiu, Qingbo Wu 0001, Fanman Meng, Taijin Zhao, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2023 | Task-Specific Loss for Robust Instance Segmentation With Noisy Class LabelsabstractDeep learning methods have achieved significant progress in the presence of correctly annotated datasets in instance segmentation. However, object classes in large-scale datasets are sometimes ambiguous, which easily causes confusion. Besides, limited experience and knowledge of annotators can lead to mislabeled object semantic classes. To solve this issue, a novel method is proposed in this paper, which considers different roles of noisy class labels in different sub-tasks. Our method is based on two basic observations: firstly, the foreground-background annotation of a sample is correct even though its class label is noisy. Secondly, symmetric loss benefits the model robustness to noisy labels but harms the learning of hard samples, while cross entropy loss is the opposite. Based on the two basic observations, in the foreground-background sub-task, cross entropy loss is used to fully exploit correct gradient guidance. In the foreground-instance sub-task, symmetric loss is used to prevent incorrect gradient guidance provided by noisy class labels. Furthermore, we apply contrastive self-supervised loss to update features of all foreground, to compensate for insufficient guidance provided by partially correct labels especially in the highly noisy setting. Extensive experiments conducted with three popular datasets (i.e., Pascal VOC, Cityscapes and COCO) have demonstrated the effectiveness of our method in a wide range of noisy class label scenarios. Longrong Yang, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Unsupervised Visual Representation Learning via Multi-Dimensional Relationship AlignmentabstractRecently, contrastive learning based on augmentation invariance and instance discrimination has made great achievements, owing to its excellent ability to learn beneficial representations without any manual annotations. However, the natural similarity among instances conflicts with instance discrimination which treats each instance as a unique individual. In order to explore the natural relationship among instances and integrate it into contrastive learning, we propose a novel approach in this paper, Relationship Alignment (RA for abbreviation), which forces different augmented views of current batch instances to main a consistent relationship with other instances. In order to perform RA effectively in existing contrastive learning framework, we design an alternating optimization algorithm where the relationship exploration step and alignment step are optimized respectively. In addition, we add an equilibrium constraint for RA to avoid the degenerate solution, and introduce the expansion handler to make it approximately satisfied in practice. In order to better capture the complex relationship among instances, we additionally propose Multi-Dimensional Relationship Alignment (MDRA for abbreviation), which aims to explore the relationship from multiple dimensions. In practice, we decompose the final high-dimensional feature space into a cartesian product of several low-dimensional subspaces and perform RA in each subspace respectively. We validate the effectiveness of our approach on multiple self-supervised learning benchmarks and get consistent improvements compared with current popular contrastive learning methods. On the most commonly used ImageNet linear evaluation protocol, our RA obtains significant improvements over other methods, our MDRA gets further improvements based on RA to achieve the best performance. The source code of our approach will be released soon. Haoyang Cheng, Hongliang Li 0001, Heqian Qiu, Qingbo Wu 0001, Xiaoliang Zhang 0002, Fanman Meng, King Ngi Ngan |
IEEE Trans. Image Process. | 7 |
| 2023 | Forgetting to Remember: A Scalable Incremental Learning Framework for Cross-Task Blind Image Quality AssessmentabstractRecent years have witnessed the great success of blind image quality assessment (BIQA) in various task-specific scenarios, which present invariable distortion types and evaluation criteria. However, due to the rigid structure and learning framework, they cannot apply to the cross-task BIQA scenario, where the distortion types and evaluation criteria keep changing in practical applications. This paper proposes a scalable incremental learning framework (SILF) that could sequentially conduct BIQA across multiple evaluation tasks with limited memory capacity. More specifically, we develop a dynamic parameter isolation strategy to sequentially update the task-specific parameter subsets, which are non-overlapped with each other. Each parameter subset is temporarily settled toRememberone evaluation preference toward its corresponding task, and the previously settled parameter subsets can be adaptively reused in the following BIQA to achieve better performance based on the task relevance. To suppress the unrestrained expansion of memory capacity in sequential tasks learning, we develop a scalable memory unit by gradually and selectively pruning unimportant neurons from previously settled parameter subsets, which enable us toForgetpart of previous experiences and free the limited memory capacity for adapting to the emerging new tasks. Extensive experiments on eleven IQA datasets demonstrate that our proposed method significantly outperforms the other state-of-the-art methods in cross-task BIQA. The source code of the proposed method is available atgithub.com/maruiperfect/SILF. Rui Ma 0030, Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Efficient Geometry Surface Coding in V-PCCabstractIn recent video-based point cloud compression (V-PCC), 3D point clouds are projected onto 2D images and compressed by High-Efficiency Video Coding (HEVC). However, HEVC was originally designed for natural visual signals, which is a suboptimal framework for point clouds. Therefore, there are still problems in geometry information compression in V-PCC: (1) The distortion based on the sum of squared error (SSE) in the existing rate-distortion optimization (RDO) is inconsistent with the geometric quality measurement; (2) The existing prediction cannot explore the fixed relationship between the corresponding far layer and near layer depth, which means that the far layer depth can be always not less than the corresponding near layer depth. In this paper, we present an efficient geometry surface coding (EGSC) method for V-PCC to address the problems. Firstly, an error projection (EP) model is designed to establish the relationship between the SSE-based distortion and the geometry quality metric. Secondly, an EP-based RDO is employed to improve the geometry information compression by estimating the point normals with gradients. Finally, an occupancy-map driven scheme is proposed to improve the prediction accuracy of merge modes. Experimental results show that the proposed method achieves an average of over 10% bit-rate saving compared with the V-PCC reference software. Jian Xiong 0005, Hao Gao 0005, Miaohui Wang, Hongliang Li 0001, King Ngi Ngan, Weisi Lin |
IEEE Trans. Multim. | 5 |
| 2022 | Instance-level Context Attention Network for instance segmentation
Chao Shang 0001, Hongliang Li 0001, Fanman Meng, Heqian Qiu, Qingbo Wu 0001, Linfeng Xu 0001, King Ngi Ngan |
Neurocomputing | 7 |
| 2022 | Category boundary re-decision by component labels to improve generation of class activation map
Runtong Zhang, Fanman Meng, Hongliang Li 0001, Qingbo Wu 0001, King Ngi Ngan |
Neurocomputing | 5 |
| 2022 | POS-Trends Dynamic-Aware Model for Video CaptionabstractVideo caption aims to generate descriptive sentences about the video, and the most critical problem is how to achieve accurate word prediction with standardized and coherent syntax structure, which requires the model to thoroughly understand video content and precisely map them into corresponding sentence components. Many existing methods usually fuse different video features into a single visual feature for generating sentences. However, they ignore the word dataset prior information in the annotations (such as Part-Of-Speech) and they also ignore the association between sentence components and types of visual features. To solve these problems, we propose a POS-trends dynamic-aware model (PDA) to fully exploit the word dataset prior information in the captions to predict POS tag, so as to assist generating captions. We propose a POS feature extraction (PFE) module to use different filters to extract different POS-trends features, predict POS tags and fuse visual features. Furthermore, we propose a visual-dynamic-aware (VDA) module to dynamically adjust the mapping way of words and supplement the visual information into the local features. The fusion features provide directional visual information to generate correct words, and the predicted POS tags to guide the decoding process to generate a more standardized and coherent syntax structure. A large number of experiments based on MSVD, MSR-VTT and VATEX demonstrated that our method outperforms the state-of-the-art methods in BLEU-4, ROUGE-L, METEOR, CIDEr. Code can be available at:https://github.com/WangLanxiao/PDA-for-video-caption. Lanxiao Wang, Hongliang Li 0001, Heqian Qiu, Qingbo Wu 0001, Fanman Meng, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Objective Object Segmentation Visual Quality Evaluation: Quality Measure and Pooling MethodabstractObjective object segmentation visual quality evaluation is an emergent member of the visual quality assessment family. It aims to develop an objective measure instead of a subjective survey to evaluate the object segmentation quality in agreement with human visual perception. It is an important benchmark for assessing and comparing the performances of object segmentation methods in terms of visual quality. Despite its essential role, sufficient study compared with other visual quality evaluation studies is still lacking. In this article, we propose a novel full-reference objective measure that includes a two-level single object segmentation visual quality measure and a pooling method for multiple object segmentation overall visual quality. The single object segmentation visual quality measure combines a pixel-level sub-measure and a region-level sub-measure for evaluating the similarity of area, shape, and object completeness between the segmentation result and the ground truth in terms of human visual perception. For the proposed multiple object segmentation overall visual quality pooling method, the rank of each object’s segmentation quality as a novel factor is integrated into the weighted harmonic mean to evaluate the overall quality. To evaluate the performance of our proposed measure, we tested it on an object segmentation subjective visual quality assessment database. The experimental results demonstrate that our proposed two-level measure and pooling method with good robustness perform better in matching subjective assessments compared with other state-of-the-art objective measures. King Ngi Ngan, Jian Xiong 0005 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | Remember and Reuse: Cross-Task Blind Image Quality Assessment via Relevance-aware Incremental LearningabstractExisting blind image quality assessment (BIQA) methods have made great progress in various task-specific applications, including the synthetic, authentic, or over-enhanced distortion evaluations. However, limited by the static model and once-for-all learning strategy, they failed to perform the cross-task evaluations in many practical applications, where diverse evaluation criteria and distortion types are constantly emerging. To address this issue, in this paper, we propose a dynamic Remember and Reuse (R&R) network, which efficiently performs the cross-task BIQA based on a novel relevance-aware incremental learning strategy. Given multiple evaluation tasks across different distortion types or databases, our R&R network sequentially updates the parameters for every task one by one. After each update step, part of task-specific parameters is settled, which ensures R&R Remembers their dedicated evaluation preferences. The remaining parameters are pruned for the dynamic usage of the subsequent tasks. To further exploit the correlation between different tasks, we feed the training data of a new task to previously settled parameters. Better prediction accuracy is considered as higher task relevance and vice versa. Then, we selectively Reuse parts of previously settled parameters, whose proportion is adaptively determined by the task relevance. Extensive experiments show that the proposed method efficiently achieves the cross-task BIQA without catastrophic forgetting, and significantly outperforms many state-of-the-art methods. Code is available at https://github.com/maruiperfect/R-R-Net. Rui Ma 0030, Hanxiao Luo, Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001 |
ACM Multimedia | 4 |
| 2021 | Few-Shot Segmentation via Complementary Prototype Learning and Cascaded Refinement
Hanxiao Luo, Hui Li 0080, Qingbo Wu 0001, Hongliang Li 0001, King Ngi Ngan, Fanman Meng, Linfeng Xu 0001 |
PRCV (4) | 5 |
| 2021 | Hierarchical class grouping with orthogonal constraint for class activation map generation
Fanman Meng, Kaixu Huang, Hongliang Li 0001, Qingbo Wu 0001, King Ngi Ngan |
Neural Comput. Appl. | 6 |
| 2021 | High-Quality R-CNN Object Detection Using Multi-Path Detection Calibration NetworkabstractObject proposals are used in two-stage detectors, such as R-CNN, to generate detection results, including category predictions and refined bounding-boxes. As a result, classification scores are assigned to refined bounding-boxes rather than object proposals. However, this procedure ignores the discrepancy of data distribution between object proposals and refined bounding-boxes. We consider this discrepancy could limit the detection accuracy. Specifically, the foreground/background imbalance on object proposals and inaccurate information from low-IoU proposals could hinder the category prediction. In this paper, we propose a detector called the Multi-Path Detection Calibration Network (PDC-Net) to address this problem. The key idea behind PDC-Net is calibrating detection results from R-CNN by considering the statistical discrepancy between object proposals and refined bounding-boxes. PDC-Net is built on Faster R-CNN. The core component in PDC-Net is the multi-path detection head, in which the base detector (from Faster R-CNN) generates detection results from object proposals and multiple calibration detectors fix incorrect outputs from the base detector using refined bounding-boxes. Experiments reveal that PDC-Net can boost detection results. Our method could reach 83.1% and 43.3% mAP respectively on PASCAL VOC and MSCOCO benchmarks, which is comparable to several state-of-the-art methods. Hongliang Li 0001, Qingbo Wu 0001, King Ngi Ngan, Linfeng Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Non-Homogeneous Haze Removal via Artificial Scene Prior and Bidimensional Graph ReasoningabstractDue to the lack of natural scene and haze prior information, it is greatly challenging to completely remove the haze from a single image without distorting its visual content. Fortunately, the real-world haze usually presents non-homogeneous distribution, which provides us with many valuable clues in partial well-preserved regions. In this paper, we propose a Non-Homogeneous Haze Removal Network (NHRN) via artificial scene prior and bidimensional graph reasoning. Firstly, we employ the gamma correction iteratively to simulate artificial multiple shots under different exposure conditions, whose haze degrees are different and enrich the underlying scene prior. Secondly, beyond utilizing the local neighboring relationship, we build a bidimensional graph reasoning module to conduct non-local filtering in the spatial and channel dimensions of feature maps, which models their long-range dependency and propagates the natural scene prior between the well-preserved nodes and the nodes contaminated by haze. To the best of our knowledge, this is the first exploration to remove non-homogeneous haze via the graph reasoning based framework. We evaluate our method on different benchmark datasets. The results demonstrate that our method achieves superior performance over many state-of-the-art algorithms for both the single image dehazing and hazy image understanding tasks. The source code of the proposed NHRN is available on https://github.com/whrws/NHRNet. Qingbo Wu 0001, Hui Li 0080, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Motion Compensated Virtual View Synthesis Using Novel Particle CellabstractDue to the wide interest in advanced multimedia experience, free-viewpoint communication is being greatly developed in recent years. In the free-viewpoint communication, viewers can perceive a view from any angle and any position of a scene. Even though the preferred views are not captured, we can generate the views through virtual view synthesis that synthesizes an arbitrary view from captured reference view(s). For daily use, only one or few cameras in baseline distance are given to capture the scene that makes the virtual view synthesis challenging. The task is more difficult when the camera is continuously moving. In this paper, we propose aparticle cellto model a reference view sequence to a set of moving particles for virtual view synthesis. Using our novel hybrid motion estimation scheme, the projected coordinates of particles in each frame are obtained even they are occluded. The particles are warped to a virtual view and synthesized as a virtual view sequence. Our method is applicable for both dynamic camera setting and static camera setting. The experimental results show our method outperforms the state-of-the-art algorithms in dynamic camera datasets and presents improvement in static camera datasets in general. Chi Ho Cheung, Lu Sheng, King Ngi Ngan |
IEEE Trans. Multim. | 3 |
| 2021 | Query Reconstruction Network for Referring Expression Image SegmentationabstractReferring expression image segmentation aims at segmenting out the object described by a natural language query. Due to the diversity of visual content and language descriptions, it is very challenging to accurately model the correspondence between the vision and language, which inevitably produces some undesired segmentation objects from the queries. In this paper, we propose a query reconstruction network (QRN) to build more consistent corresponding relations between the language queries and object segmentation results. QRN not only generates segmentations from the queries and images but also reversely reconstructs the queries from the segmentations and the images. Through query reconstruction, QRN can confirm the vision-language consistency between the segmentations and queries. In the inference stage, for inconsistent segmentations and queries, we propose an iterative segmentation correction (ISC) method to correct them. ISC takes the difference between the reconstructed and input queries as a loss to optimize the proposed QRN. Then, the proposed QRN can generate new segmentations and queries. By iterative optimization, the segmentations can be gradually corrected. Extensive experiments on four referring expression image segmentation databases demonstrate the effectiveness of the proposed method. Hengcan Shi, Hongliang Li 0001, Qingbo Wu 0001, King Ngi Ngan |
IEEE Trans. Multim. | 4 |
| 2020 | Single Image Dehazing Via Artificial Multiple Shots And Multidimensional ContextabstractThe main challenge for single image dehazing is the lack of effective prior information for restoration. To address this issue, in this paper, we propose to generate artificial multiple shots for simulating the images captured under different haze degrees, and two context reasoning modules are developed to describe the relationship across different spatial regions and artificial shots. It brings two benefits in the inhomogeneous haze distribution. First, within one shot, the regions occluded in one location could be recovered with the help of other clear regions, which share the similar structures. Second, for the same spatial location, the regions distorted in one shot could be restored by means of other shots with clear content. We evaluate our method on different benchmark datasets. The results demonstrate that our method achieves superior performance over many state-of-the-art dehazing algorithms. Qingbo Wu 0001, Hui Li 0080, King Ngi Ngan, Hongliang Li 0001, Fanman Meng |
ICIP | 4 |
| 2020 | Region Adaptive Two-Shot Network For Single Image DehazingabstractExisting single image dehazing methods typically adopt a one-shot strategy by indiscriminately applying the same filters to all local regions, which easily cause under-/over-dehazing across different regions by ignoring the inhomogeneity and asymmetry of illumination and detail distortions. In this paper, we propose a region adaptive two-shot network (RATNet) to address this issue. In the first shot, a lightweight subnetwork is utilized to conduct the regular global filtering, which could remove parts of haze but also distort some image details. In the second shot, a two-branch subnetwork is developed to restore the illumination and details of the initially renovated image respectively. The final dehazed image is obtained by fusing the outputs of the previous two branches, whose region-variant weights are adaptively learned by minimizing the difference between the haze-free image and our fused result. Experiments on four dehazing benchmark datasets show that our RATNet significantly outperforms many state-of-the-art dehazing approaches. Hui Li 0080, Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng |
ICME | 3 |
| 2020 | Language-Aware Fine-Grained Object Representation for Referring Expression ComprehensionabstractReferring expression comprehension expects to accurately locate an object described by a language expression, which requires precise language-aware visual object representations. However, existing methods usually use rectangular object representations, such as object proposal regions and grid regions. They ignore some fine-grained object information like shapes and poses, which are often described in language expressions and important to localize objects. Additionally, rectangular object regions usually contain background contents and irrelevant foreground features, which also decrease the localization performance. To address these problems, we propose a language-aware deformable convolution model (LDC) to learn language-aware fine-grained object representations. Rather than extracting rectangular object representations, LDC adaptively samples a set of key points based on the image and language to represent objects. This type of object representations can capture more fine-grained object information (e.g., shapes and poses) and suppress noises in accordance with language and thus, boosts the object localization performance. Based on the language-aware fine-grained object representation, we next design a bidirectional interaction model (BIM) that leverages a modified co-attention mechanism to build cross-modal bidirectional interactions to further improve the language and object representations. Furthermore, we propose a hierarchical fine-grained representation network (HFRN) to learn language-aware fine-grained object representations and cross-modal bidirectional interactions at local word level and global sentence level, respectively. Our proposed method outperforms the state-of-the-art methods on the RefCOCO, RefCOCO+ and RefCOCOg datasets. Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, Hengcan Shi, Taijin Zhao, King Ngi Ngan |
ACM Multimedia | 7 |
| 2020 | A multi-scale language embedding network for proposal-free referring expression comprehensionabstractReferring expression comprehension (REC) is a task that aims to find the location of an object specified by a language expression. Current solutions for REC can be classified into proposal-based methods and proposal-free methods. Proposal-free methods are popular recently because of its flexibility and lightness. Nevertheless, existing proposal-free works give little consideration to visual context. As REC is a context sensitive task, it is hard for current proposal-free methods to comprehend expressions that describe objects by the relative position with surrounding things. In this paper, we propose a multi-scale language embedding network for REC. Our method adopts the proposal-free structure, which directly feeds fused visual-language features into a detection head to predict the bounding box of the target. In the fusion process, we propose a grid fusion module and a grid-context fusion module to compute the similarity between language features and visual features in different size regions. Meanwhile, we extra add fully interacted vision-language information and position information to strength the feature fusion. This novel fusion strategy can help to utilize context flexibly therefore the network can deal with varied expressions, especially expressions that describe objects by things around. Our proposed method outperforms the state-of-the-art methods on Refcoco, Refcoco+ and Refcocog datasets. Taijin Zhao, Hongliang Li 0001, Heqian Qiu, Qingbo Wu 0001, King Ngi Ngan |
MMAsia | 5 |
| 2020 | Haze-robust image understanding via context-aware deep feature refinementabstractImage understanding under the foggy scene is greatly challenging due to inhomogeneous visibility deterioration. Although various image dehazing methods have been proposed, they usually aim to improve image visibility (such as, PSNR/SSIM) in the pixel space rather than the feature space, which is critical for the perception of computer vision. Due to this mismatch, existing dehazing methods are limited or even adverse in facilitating the foggy scene understanding. In this paper, we propose a generalized deep feature refinement module to minimize the difference between clear images and hazy images in the feature space. It is consistent with the computer perception and can be embedded into existing detection or segmentation backbones for joint optimization. Our feature refinement module is built upon the graph convolutional network, which is favorable in capturing the contextual information and beneficial for distinguishing different semantic objects. We validate our method on the detection and segmentation tasks under foggy scenes. Extensive experimental results show that our method outperforms the state-of-the-art dehazing based pretreatments and the fine-tuning results on hazy images. Hui Li 0080, Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001 |
MMSP | 4 |
| 2020 | A Unified Single Image De-raining Model via Region Adaptive Coupled NetworkabstractSingle image de-raining is quite challenging due to the diversity of rain types and inhomogeneous distributions of rainwater. By means of dedicated models and constraints, existing methods perform well for specific rain type. However, their generalization capability is highly limited as well. In this paper, we propose a unified de-raining model by selectively fusing the clean background of the input rain image and the well restored regions occluded by various rains. This is achieved by our region adaptive coupled network (RACN), whose two branches integrate the features of each other in different layers to jointly generate the spatial-variant weight and restored image respectively. On the one hand, the weight branch could lead the restoration branch to focus on the regions with higher contributions for de-raining. On the other hand, the restoration branch could guide the weight branch to keep off the regions with over-/under-filtering risks. Extensive experiments show that our method outperforms many state-of-the-art de-raining algorithms on diverse rain types including the rain streak, raindrop and rain-mist. Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001 |
VCIP | 3 |
| 2020 | A New Bounding Box based Pseudo Annotation Generation Method for Semantic SegmentationabstractThis paper proposes a fusion-based method to generate pseudo-annotations from bounding boxes for semantic segmentation. The idea is to first generate diverse foreground masks by multiple bounding box segmentation methods, and then combine these masks to generate pseudo-annotations. Existing methods generate foreground masks from bounding boxes by classical segmentation methods driving by low-level features and own local information, which is hard to generate accurate and diverse results for the fusion. Different from the traditional methods, multiple class-agnostic models are modeled to learn the objectiveness cues by using existing labeled pixel-level annotations and then to fuse. Firstly, the classical Fully Convolutional Network (FCN) that densely predicts the pixels' labels is used. Then, two new sparse prediction based class-agnostic models are proposed, which simplify the segmentation task as sparsely predicting the boundary points through predicting the distance from the bounding box border to the object boundary in Cartesian Coordinate System and the Polar Coordinate System, respectively. Finally, a voting-based strategy is proposed to combine these segmentation results to form better pseudo-annotations. We conduct experiments on PASCAL VOC 2012 dataset. The mIoU of the proposed method is 68.7%, which outperforms the state-of-the-art method by 1.9%. Xiaolong Xu 0004, Fanman Meng, Hongliang Li 0001, Qingbo Wu 0001, King Ngi Ngan |
VCIP | 5 |
| 2020 | Mono is Enough: Instance Segmentation from Single Annotated SampleabstractWith the help of various Deep Neural Networks, instance segmentation has achieved significant progress. How-ever, these successes are heavily reliant on large-scale manually annotated samples, which are extremely time-consuming and expensive. To address this issue, we propose a highly efficient anisotropic data augmentation method, which generates high quality training data from a single manually annotated sample. Instead of equivalently modifying foreground and background like traditional data augmentation methods, we focus on enriching the diversities of foreground appearance and positional relation between foreground and background, which are beneficial for the classification and localization sub-tasks respectively. All foreground instances of the source annotated sample undergo various rotation, brightness change, rescale, distortion and frequency-component mixup (FCM). Then, these modified instances are randomly embedded into background, which serve as new training samples. Experiments on Cityscapes dataset show that our method significantly outperforms traditional data augmentation methods. Longrong Yang, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, King Ngi Ngan |
VCIP | 5 |
| 2020 | Mining Larger Class Activation Map with Common Attribute LabelsabstractClass Activation Map (CAM) is the visualization of target regions generated from classification networks. However, classification network trained by class-level labels only has high responses to a few features of objects and thus the network cannot discriminate the whole target. We think that original labels used in classification tasks are not enough to describe all features of the objects. If we annotate more detailed labels like class-agnostic attribute labels for each image, the network may be able to mine larger CAM. Motivated by this idea, we propose and design common attribute labels, which are lower-level labels summarized from original image-level categories to describe more details of the target. Moreover, it should be emphasized that our proposed labels have good generalization on unknown categories since attributes (such as head, body, etc.) in some categories (such as dog, cat, etc.) are common and class-agnostic. That is why we call our proposed labels as common attribute labels, which are lower-level and more general compared with traditional labels. We finish the annotation work based on the PASCAL VOC2012 dataset and design a new architecture to successfully classify these common attribute labels. Then after fusing features of attribute labels into original categories, our network can mine larger CAMs of objects. Our method achieves better CAM results in visual and higher evaluation scores compared with traditional methods. Runtong Zhang, Fanman Meng, Hongliang Li 0001, Qingbo Wu 0001, King Ngi Ngan |
VCIP | 5 |
| 2020 | Hybrid-loss supervision for deep neural network
Qishang Cheng, Hongliang Li 0001, Qingbo Wu 0001, King Ngi Ngan |
Neurocomputing | 4 |
| 2020 | Discriminative deep metric learning for asymmetric discrete hashing
Lei Ma 0004, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, King Ngi Ngan |
Neurocomputing | 5 |
| 2020 | Parametric Deformable Exponential Linear Units for deep neural networks
Qishang Cheng, Hongliang Li 0001, Qingbo Wu 0001, Lei Ma 0004, King Ngi Ngan |
Neural Networks | 5 |
| 2020 | HeadNet: An End-to-End Adaptive Relational Network for Head DetectionabstractHead detection plays an important role in localizing and identifying persons from visual data. Most existing methods treat head detection as a specific form of object detection. Head detection is nontrivial due to the considerable difficulty in building the local and global information under conditions of unconstrained pose and orientation. To address these issues, this paper presents an effective adaptive relational network to capture context information, which is greatly helpful to suppress missed detection. We show that the fundamental contextual properties, such as the global shape priors from different heads and the local adjacent relationship between the head and shoulders, can be systematically quantified by visual operators. Specifically, we propose a two-step search algorithm to quantify the global intergroup conflict with adaptive scale, pose and viewpoint. Meanwhile, a structured feature module is introduced to capture the local relation of intraindividual stability. Finally, the global priors and local relation are integrated seamlessly into a single-stage head detector that is end-to-end trainable. An extensive ablation analysis demonstrates the effectiveness of our approach. We achieve state-of-the-art results on two challenging datasets, i.e., HollywoodHeads and Brainwash. Wei Li 0110, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, Linfeng Xu 0001, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2020 | Subjective and Objective De-Raining Quality Assessment Towards Authentic Rain ImageabstractImages acquired by outdoor vision systems easily suffer poor visibility and annoying interference due to the rainy weather, which brings great challenge for accurately understanding and describing the visual contents. Recent researches have devoted great efforts on the task of rain removal for improving the image visibility. However, there is very few exploration about the quality assessment of de-rained image, even it is crucial for accurately measuring the performance of various de-raining algorithms. In this paper, we first create a de-raining quality assessment (DQA) database that collects 206 authentic rain images and their de-rained versions produced by 6 representative single image rain removal algorithms. Then, a subjective study is conducted on our DQA database, which collects the subject-rated scores of all de-rained images. To quantitatively measure the quality of de-rained image with non-uniform artifacts, we propose a bi-directional feature embedding network (B-FEN) which integrates the features of global perception and local difference together. Experiments confirm that the proposed method significantly outperforms many existing universal blind image quality assessment models. To help the research towards perceptually preferred de-raining algorithm, we will publicly release our DQA database and B-FEN source code on https://github.com/wqb-uestc. Qingbo Wu 0001, Lei Wang 0186, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Linfeng Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Hierarchical Context Features Embedding for Object DetectionabstractPixel-level segmentation has been widely used to improve object detection. Most of the existing methods refine detection features by adding the constraint of the segmentation branch or by simply embedding high-level segmentation features into detection features within the local receptive field. However, noisy segmentation features are unavoidable in real-word applications and can easily cause false positives. To address this problem, we propose a novel hierarchical context embedding module to effectively embed segmentation features into detection features. The idea of this module is to capture hierarchical context information that includes local objects or parts and nonlocal context features by learning multiple attention maps, and subsequently utilize interdependencies between features to recalibrate noisy segmentation features. Furthermore, we use this module in the proposed gated encoder-decoder network that adaptively aggregates feature maps of different resolutions based on the gate mechanism so that we can embed multiscale segmentation feature maps into detection features for more accurate detection of objects of all sizes. Experimental results demonstrate the effectiveness of the proposed method on the Pascal VOC 2012Seg dataset, the Pascal VOC dataset and the MS COCO dataset. Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, Linfeng Xu 0001, King Ngi Ngan, Hengcan Shi |
IEEE Trans. Multim. | 6 |
| 2020 | Rate Constrained Multiple-QP Optimization for HEVCabstractIn High Efficiency Video Coding (HEVC), multiple-QP (quantization parameter) optimization can adapt to a local video content. However, the multiple-QP implementation in the HEVC reference software (HM 16.6) achieves the best QP value for each coding block with a large amount of computational complexity. To address this challenge, we propose a fast rate-constrained multiple-QP optimization approach for the HM platform. We first introduce a template-based transform coefficient selection method which can save the overall complexity of entropy coding. In addition, we model the multiple-QP determination as a new rate-constrained optimization problem, and finally, we get a feasible solution with a lower computation overhead. Experimental results show that our method dramatically reduces the average complexity under the all-intra, low-delay and random-access configuration. Miaohui Wang, Jian Xiong 0005, Long Xu 0001, Wuyuan Xie, King Ngi Ngan, Harry Qin |
IEEE Trans. Multim. | 5 |
| 2019 | MVF-Net: Multi-View 3D Face Morphable Model RegressionabstractWe address the problem of recovering the 3D geometry of a human face from a set of facial images in multiple views. While recent studies have shown impressive progress in 3D Morphable Model (3DMM) based facial reconstruction, the settings are mostly restricted to a single view. There is an inherent drawback in the single-view setting: the lack of reliable 3D constraints can cause unresolvable ambiguities. We in this paper explore 3DMM-based shape recovery in a different setting, where a set of multi-view facial images are given as input. A novel approach is proposed to regress 3DMM parameters from multi-view inputs with an end-to-end trainable Convolutional Neural Network (CNN). Multi-view geometric constraints are incorporated into the network by establishing dense correspondences between different views leveraging a novel self-supervised view alignment loss. The main ingredient of the view alignment loss is a differentiable dense optical flow estimator that can backpropagate the alignment errors between an input view and a synthetic rendering from another input view, which is projected to the target view through the 3D shape to be inferred. Through minimizing the view alignment loss, better 3D shapes can be recovered such that the synthetic projections from one view to another can better align with the observed image. Extensive experiments demonstrate the superiority of the proposed method over other 3DMM methods. Fanzi Wu, Linchao Bao, Yonggen Ling, Yibing Song, Songnan Li, King Ngi Ngan, Wei Liu 0005 |
CVPR | 7 |
| 2019 | Material Segmentation in Hyperspectral Images with a Spatio-spectral Texture DescriptorabstractIn this paper, we address the problem of ground-based hyperspectral image segmentation by combining pixel-level and region-level classification with a region boundary refinement approach. To this end, we represent the spatio-spectral feature of image regions by a descriptor based on Vector of Locally Aggregated Descriptors (VLAD). Further, the region boundaries are refined by minimizing the total region perimeter. Experimental results on a ground-based hyperspectral image dataset clearly demonstrate the advantage of the proposed method over recent prior works, based on several metrics. Yu Zhang 0063, Cong Phuoc Huynh, Nariman Habili, King Ngi Ngan |
ICASSP | 4 |
| 2019 | Beyond Synthetic Data: A Blind Deraining Quality Assessment Metric Towards Authentic Rain ImageabstractDeraining quality assessment (DQA) plays an important role in evaluating and guiding the design of the image deraining algorithm. Due to the absence of rain-free image in the real rainy weather, the existing deraining algorithms are typically tested on several synthetic data by simulating very limited types of rain streaks, which are far from sufficient to measure the practicability of a deraining algorithm. In this paper, we first build a subjective DQA database that collects diverse authentic rain images and their derained versions. Then, a blind quality metric is developed to predict the deraining quality. Since the deraining artifacts are anisotropic and variable, we propose to describe the image via a bi-directional gated fusion network (B-GFN), which adaptively integrates the multi-scale cues of deraining artifact. Experiments confirm the effectiveness of the proposed method and its superiority with respect to many state-of-the-art blind image quality metrics. Qingbo Wu 0001, Lei Wang 0186, King Ngi Ngan, Hongliang Li 0001, Fanman Meng |
ICIP | 3 |
| 2019 | Blind Image Sharpness Assessment And Enhancement via Deep Auxiliary LearningabstractIn this paper, we propose an unified deep auxiliary learning network to train the blind image sharpness assessment (BISA) metric and enhancer simultaneously. Instead of using the BISA as a parameter tuner like existing works, the proposed method aims to exploit the complementary information between two tasks and boost both of their performance. On the one hand, the enhancement subnetwork tries to separate a blurry image into the clear version and disparity map, which provide additional mask effect and blurry degree information for accurate BISA. On the other hand, the BISA subnetwork help determine the enhancement degree by feeding sharpness-aware features to the enhancer, which is helpful for avoiding under-/over-enhancing. Experimental results on three publicly available databases show that the proposed method outperforms many state-of-the-art algorithms in both the BISA and sharpness enhancement tasks. Qingbo Wu 0001, Rui Ma 0030, King Ngi Ngan, Hongliang Li 0001, Fanman Meng |
ICME | 3 |
| 2019 | A New Few-shot Segmentation Network Based on Class RepresentationabstractThis paper studies few-shot segmentation, which is a task of predicting foreground mask of unseen classes by a few of annotations only, aided by a set of rich annotations already existed. The existing methods mainly focus the task on "how to transfer segmentation cues from support images (labeled images) to query images (unlabeled images)", and try to learn efficient and general transfer module that can be easily extended to unseen classes. However, it is proved to be a challenging task to learn the transfer module that is general to various classes. This paper solves few-shot segmentation in a new perspective of "how to represent unseen classes by existing classes", and formulates few-shot segmentation as the representation process that represents unseen classes (in terms of forming the foreground prior) by existing classes precisely. Based on such idea, we propose a new class representation based few-shot segmentation framework, which firstly generates class activation map of unseen class based on the knowledge of existing classes, and then uses the map as foreground probability map to extract the foregrounds from query image. A new two-branch based few-shot segmentation network is proposed. Moreover, a new CAM generation module that extracts the CAM of unseen classes rather than the classical training classes is raised. We validate the effectiveness of our method on Pascal VOC 2012 dataset, the value FB-IoU of one-shot and five-shot arrives at 69.2% and 70.1% respectively, which outperforms the state-of-the-art method. Fanman Meng, Hongliang Li 0001, King Ngi Ngan, Qingbo Wu 0001 |
VCIP | 4 |
| 2019 | Visibility Constrained Generative Model for Depth-Based 3D Facial Pose TrackingabstractIn this paper, we propose a generative framework that unifies depth-based 3D facial pose tracking and face model adaptation on-the-fly, in the unconstrained scenarios with heavy occlusions and arbitrary facial expression variations. Specifically, we introduce a statistical 3D morphable model that flexibly describes the distribution of points on the surface of the face model, with an efficient switchable online adaptation that gradually captures the identity of the tracked subject and rapidly constructs a suitable face model when the subject changes. Moreover, unlike prior art that employed ICP-based facial pose estimation, to improve robustness to occlusions, we propose a ray visibility constraint that regularizes the pose based on the face model's visibility with respect to the input point cloud. Ablation studies and experimental results on Biwi and ICT-3DHP datasets demonstrate that the proposed framework is effective and outperforms completing state-of-the-art depth-based methods. Lu Sheng, Jianfei Cai 0001, Tat-Jen Cham, Vladimir Pavlovic 0001, King Ngi Ngan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2019 | Cascaded regression using landmark displacement for 3D face reconstruction
Fanzi Wu, Songnan Li, King Ngi Ngan, Lu Sheng |
Pattern Recognit. Lett. | 4 |
| 2019 | UHD Video Coding: A Light-Weight Learning-Based Fast Super-Block ApproachabstractThe ultra high-definition (UHD) video format, which has recently become popular, aims to provide high spatial resolution, high temporal frame rate, high sample bit-depth, and wide pixel color gamut. Despite the continued development of global network capacities, it inevitably causes the increased bandwidth cost of catering to the requirement of delivering UHD video services. To address such challenges, this paper presents an improved super coding unit (SCU) method for UHD video coding in High Efficiency Video Coding (HEVC). Initially, the medium coding unit (MCU) is proposed to avoid unnecessary brute-force coding unit (CU) partitions of SCU. Furthermore, the SCU is proposed to be encoded by Direct-MCU and SCU-to-MCU modes: the Direct-MCU mode is intended to better adapt to the texture-rich region, which guarantees the compression efficiency by avoiding extra-size CU partition; the SCU-to-MCU mode is designed for the homogeneous region of UHD content, which saves the encoding time by skipping fine-grained CU partition search. Moreover, a learning-based fast SCU decision approach is proposed to speed up the determination process of Direct-MCU and SCU-to-MCU, where three representative handcrafted features are extracted. Experimental results show that our method achieves an affordable complexity and excellent coding efficiency (up to 7.30% Bjøntegaard Delta rate savings) in UHD video coding compared to recent HEVC reference software. Miaohui Wang, Wuyuan Xie, Xiandong Meng, Huanqiang Zeng, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2019 | Feature Fusion With Predictive Weighting for Spectral Image Classification and SegmentationabstractIn this paper, we propose a spatial-spectral feature fusion model with a predictive feature weighting mechanism and demonstrate its applications to the problems of hyperspectral image classification and segmentation. To address these problems, we learn a set of 1-D convolutional local spectral filters and 2-D spatial-spectral filters that feed features into a fusion module, in an end-to-end fashion. We propose a lightweight predictive feature weighting component embedded in the fusion model and consider four design fusion options, i.e., by adding or concatenating features with equal or predicted weights. For the pixel classification task, the training input consists of image patches with labeled central pixels, whereas for the spatial segmentation task, it includes the label maps of image regions. The proposed networks have been evaluated on the Indian Pines, Pavia University, and Houston University for the classification problem and the SpaceNet data set for the spatial segmentation problem. The quantitative results favor the proposed approach over the state-of-the-art methods across all the four data sets. Yu Zhang 0063, Cong Phuoc Huynh, King Ngi Ngan |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2019 | 3-D Reconstruction of Human Body Shape From a Single Commodity Depth Cameraabstract3-D human body reconstruction is an important research topic in computer vision. A 3-D human body model can be used in sports science, movie industry and personalized entertainment, especially virtual reality games. Most of depth-based 3-D reconstruction algorithms need multiple cameras surrounding the user and require the user to keep a specific pose strictly while capturing depth images. In this paper, we propose an algorithm to reconstruct the 3-D shape of human bodies using a single commodity depth camera. Our algorithm only needs two depth images of the front-facing and back-facing bodies. It also has strong operability since the proposed method is insensitive to the pose variations between the two depth images. We reconstruct 3-D shapes of front-facing and back-facing bodies from the two depth images, respectively, and stitch them together. We also propose a novel registration method, namely, “iterative mid-distance points,” which has fast convergence and robustness to the depth noise. The proposed method enables robust and easy-to-use human body reconstruction, and achieves higher accuracy than state-of-the-art methods. Songnan Li, King Ngi Ngan, Fanzi Wu |
IEEE Trans. Multim. | 3 |
| 2018 | Boosting Scene Parsing Performance via Reliable Scale PredictionabstractSegmenting objects on suitable scales is a key factor to improve the scene parsing performance. Existing methods either simply average multi-scale results or predict scales by weakly-supervised models, due to the lack of scale labels. In this paper, we propose a novel fully-supervised Scale Prediction Model. On one hand, the proposed Scale Prediction Model learns parsing scales by the strong scale supervision, which is automatically generated from the scene parsing ground truth without any extra manually annotation. On the other hand, we explore the relationship between scale and object class, and propose to use the object class information to further improve the reliability of the scale prediction. The proposed Scale Prediction Model improves 23.1%, 20.1% and 29.3% scale prediction accuracies on the NYU Depth v2, PASCAL-Context and SIFT Flow datasets, respectively. Based on the Scale Prediction Model, we design a Scale Parsing Net (SPNet) for scene parsing, which segments each object on the scale predicted by the Scale Prediction Model. Moreover, SPNet leverages the intermediate result (i.e., the object class) to refine the parsing results. The experiment results show that SPNet outperforms many state-of-the-art methods on multiple scene parsing datasets. Hengcan Shi, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, King Ngi Ngan |
ACM Multimedia | 5 |
| 2018 | Multi-task Learning for Deep Semantic HashingabstractDeep learning to hash has emerged as a popular technique for large-scale image retrieval. Existing deep learning to hash methods seek to solve the single retrieval task within one stream framework or jointly solve the retrieval task and the classification task within two stream framework. Consequently, the semantic information is not fully exploited to generate compact and discriminative hash codes. In this paper, we propose a multi-task learning architecture for deep semantic hashing (MLDH), which incorporates the retrieval task and the classification task within one-stream framework. Specifically, we introduce a COCO loss to learn compact binary codes for the classification task. For the retrieval task, we introduce a pairwise loss to learn discriminative binary codes. Finally, these two tasks are investigated into one-stream deep learning framework. Extensive experiments show that MLDH can outperform state-of-the-art methods on benchmark datasets. Lei Ma 0004, Hongliang Li 0001, Qingbo Wu 0001, Chao Shang 0001, King Ngi Ngan |
VCIP | 5 |
| 2018 | Global and local semantics-preserving based deep hashing for cross-modal retrieval
Lei Ma 0004, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, King Ngi Ngan |
Neurocomputing | 5 |
| 2018 | Separable authentication in encrypted HEVC video
Yiqi Tew, Koksheik Wong, Raphael C.-W. Phan, King Ngi Ngan |
Multim. Tools Appl. | 4 |
| 2018 | Interactive object segmentation in two phases
King Ngi Ngan, Songnan Li, Hongliang Li 0001 |
Signal Process. Image Commun. | 2 |
| 2018 | Boundary-Guided Optimization Framework for Saliency RefinementabstractSalient object detection has made a rapid progress in recent years. To improve the quality of initial saliency maps, existing algorithms typically refine them via a neighbor-constrained smoothing model, which assigns similar saliency values to neighboring regions. Since the adjacent regions could also cross the boundary between the salient object and background, these spatial distance-based methods easily cause false detection by involving the background regions that are close to the salient objects. To address this problem, we propose a boundary-guided optimization framework to jointly improve the region smoothness and correct the false detect regions. Specifically, we introduce a latent segmentation variable to regularize the consistency between the refined saliency map and the latent segmentation mask, which penalizes high (low) saliency values of the regions lying outside (inside) the estimated object boundary. To optimize the proposed objective function, we decompose the primary problem into two subproblems: submodular optimization problem and convex optimization problem. The submodular optimization problem can be quickly optimized using the off-the-shelf technique while the convex optimization can be solved with a closed-form solution. The experimental results show that the proposed method consistently improve the performances of eight state-of-the-art salient object detection algorithms on three datasets, including latest deep convolutional neural network-based algorithms. Meanwhile, our method outperforms state-of-the-art saliency refinement algorithms. Liangzhi Tang, Hongliang Li 0001, Qingbo Wu 0001, King Ngi Ngan |
IEEE Signal Process. Lett. | 4 |
| 2018 | An Unsupervised Method to Extract Video Object via Complexity Awareness and Object Local PartsabstractExisting unsupervised video object segmentation generates object information from the whole video, which ignores analysis of the local clips. However, we observe that local clips and their relationships are also useful for the video object segmentation. For example, the simple background clips can be used to improve the segmentation of complex background clips. In this paper, we propose a novel unsupervised segmentation framework to segment the primary object based on two aspects, i.e., the complexity awareness of video clips and their segmentation propagation. The first one is used to select the simple clips with smooth backgrounds and the second one generates an object prior from the simple clips and propagates the object prior to help and improve the segmentation of the complex clips. A complexity awareness method using the static cues and the dynamic cues are proposed to evaluate the complexity of the video frames. A new object prior learning model based on the local part structure is designed and a local part-based prior propagation is proposed for the complex clip segmentation. To verify our method, we collect a new challenging video segmentation data set, in which each video contains diverse backgrounds. Experimental results demonstrate that our method outperforms several state-of-the-art methods both on a classical data set and our new data set. Bing Luo 0003, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | Globally Measuring the Similarity of Superpixels by Binary Edge Maps for Superpixel ClusteringabstractThis paper proposes an edge-based superpixel similarity measurement, which globally evaluates the similarity between superpixels by binary edge maps. The basic idea is to assess whether the superpixels are surrounded by the same edges. To this end, we first describe the edge spatial distributions by directional regions and then use the directional regions to represent the surrounding relationships of superpixels and edges by their traverse relationships, which form the histogram feature. Finally, the similarity is simply calculated by the distances between the features. To verify the proposed similarity measurement, we use our global similarity measurement to perform superpixel clustering. Two clustering methods, the directed graph clustering (DGC) and spectral clustering (ultrametric contour map) are combined to achieve the clustering process. The combination of our global similarity measurement and DGC to form a new three-layer-based superpixel generation method, which can quickly generate the superpixel from edge maps, is highlighted. We verify the global similarity measurement by the BSDS500 dataset. The experimental results demonstrate that the proposed global similarity measurement can improve the clustering accuracy in terms of larger intersection-over-union-criterion-based values. The code can be downloaded from https://github.com/FanmanMeng/Superpixel-Similarity-Measurement. Fanman Meng, Hongliang Li 0001, Qingbo Wu 0001, Bing Luo 0003, Chao Huang 0003, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2018 | Blind Image Quality Assessment Using Local Consistency Aware Retriever and Uncertainty Aware EvaluatorabstractBlind image quality assessment (BIQA) aims to automatically predict the perceptual quality of a digital image without accessing its pristine reference. Previous studies mainly focus on extracting various quality-relevant image features. By contrast, the explorations on highly efficient learning model are still very limited. Motivated by the fact that it is difficult to approximate a complex and large data set via a global parametric model, we propose a novel local learning method for BIQA to improve quality prediction performance. More specifically, we search for the perceptually similar neighbors of a test image to serve as its unique training set. Unlike the widely used k nearest neighbors principle, which only measures the similarity between the testing and training samples, the local consistency of the selected training data is also considered to generate smoother sample space. The image quality is estimated via a sparse Gaussian process. As an additional benefit, the uncertainty of the predicted score is jointly inferred, which can subsequently drive more robust perceptual image processing applications, such as deblocking investigated in this paper. Extensive experiments demonstrate that the proposed learning model leads to consistent quality prediction improvements over many state-of-the-art BIQA algorithms. Qingbo Wu 0001, Hongliang Li 0001, King Ngi Ngan, Kede Ma |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | A Perceptually Weighted Rank Correlation Indicator for Objective Image Quality AssessmentabstractIn the field of objective image quality assessment (IQA), Spearman's ρ and Kendall's τ, which straightforwardly assign uniform weights to all quality levels and assume that each pair of images is sortable, are the two most popular rank correlation indicators. These indicators can successfully measure the average accuracy of an IQA metric for ranking multiple processed images. However, two important perceptual properties are ignored. First, the sorting accuracy (SA) of high-quality images is usually more important than that of poor-quality images in many real-world applications, where only top-ranked images are pushed to the users. Second, due to the subjective uncertainty in making judgments, two perceptually similar images are usually barely sortable, and their ranks do not contribute to the evaluation of an IQA metric. To more accurately compare different IQA algorithms, in this paper, we explore a perceptually weighted rank correlation indicator, which rewards the capability of correctly ranking high-quality images and suppresses the attention towards insensitive rank mistakes. Specifically, we focus on activating a 'valid' pairwise comparison of images whose quality difference exceeds a given sensory threshold (ST). Meanwhile, each image pair is assigned a unique weight that is determined by both the quality level and rank deviation. By modifying the perception threshold, we can illustrate the sorting accuracy with a sophisticated SA-ST curve rather than a single rank correlation coefficient. The proposed indicator offers new insight into interpreting visual perception behavior. Furthermore, the applicability of our indicator is validated for recommending robust IQA metrics for both degraded and enhanced image data. Qingbo Wu 0001, Hongliang Li 0001, Fanman Meng, King Ngi Ngan |
IEEE Trans. Image Process. | 4 |
| 2018 | Generic Proposal Evaluator: A Lazy Learning Strategy Toward Blind Proposal Quality AssessmentabstractExisting detection or recognition systems typically select one state-of-the-art proposal algorithm to produce massive object-covered candidate windows, and a quality metric specifically designed for this algorithm is utilized to single out small amounts of proposals. However, in practice, the accuracies of different proposal algorithms significantly change from one image content to another one. To obtain more robust proposal results, a generic proposal evaluator (GPE) is highly desired, which could choose optimal candidate windows across multiple proposal algorithms. In this paper, we propose a lazy learning strategy to train the GPE, which aims to blindly estimate the quality of each proposal without accessing to its manual annotation. Unlike the traditional end-to-end framework that learns a universal model from all training samples, we try to build query-specific training subset for each given proposal, where only its k-nearest-neighborhoods are collected from all labeled candidate windows. Benefits from the capability of updating the regression parameters for different visual contents, the proposed method delivers a higher quality prediction accuracy even with respect to the deep neural network learned by end-to-end method. Experimental results confirm that the proposed algorithm significantly outperforms many state-of-the-art proposal quality metrics. Qingbo Wu 0001, Hongliang Li 0001, Fanman Meng, King Ngi Ngan |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2018 | Spatio-Temporal Disocclusion Filling Using Novel Sprite CellsabstractDepth image-based rendering is an important technique for virtual view synthesis with limited 3-D data. However, occluded areas result in disocclusions in synthesized images. Filling of the disocclusions in a plausible manner is a critical task in virtual view synthesis. In addition to spatial consistency, temporal consistency of the filled regions also affects the visual quality. In this paper, we propose a novel codebook method called sprite cell for filling the disocclusions with high spatial and temporal consistency. Each codeword consists of a color vector, depth value, frame log, and the confidence score of the corresponding pixel. In contrast with the existing methods that reuse the filling results of previous frames without considering their accuracy, the proposed method estimates the confidence scores of the filling results to prevent temporal continuation of filling errors. Moreover, we introduce a method to correct the luminance of filled disocclusions that compensates for the change of scenes. The experimental results show that the proposed method achieves both objective and subjective improvements over the state-of-the-art methods. The sequences synthesized with the proposed method have higher spatio-temporal consistency. Chi Ho Cheung, King Ngi Ngan, Lu Sheng |
IEEE Trans. Multim. | 2 |
| 2018 | Seeds-Based Part Segmentation by Seeds Propagation and Region Convexity DecompositionabstractObject part segmentation is an important and challenging task in computer vision. The existing supervised part segmentation methods need pixel level training data which leads to a huge workload for the user. In this paper a weakly supervised part segmentation method is proposed which segments part regions from multiple images by only several seeds on an image. Two aspects such as seed propagation among multiple images and part generation from seeds are considered. The first aspect is to generate part seeds in each image in terms of seed propagation which is accomplished by part matching combined with latent object regions. We fuse the local part matching and global shape cosegmentation to avoid the noise propagation. The second aspect is to segment part regions from object regions and part seeds which is formulated as the object shape decomposition model. The shape convexity analysis and seed location are fused to accomplish the decomposition and the final part segmentation. The proposed method is verified on the PASCAL 2010 dataset Bird dataset Cat-Dog dataset and UCF Sports Actions dataset. Experimental results demonstrate the effectiveness of the proposed method with larger intersection over union (IOU) values compared with existing weakly supervised part generation methods. Fanman Meng, Hongliang Li 0001, Qingbo Wu 0001, King Ngi Ngan, Jianfei Cai 0001 |
IEEE Trans. Multim. | 4 |
| 2018 | Hierarchical Parsing Net: Semantic Scene Parsing From Global Scene to ObjectsabstractThis paper proposes a novel Hierarchical Parsing Net (HPN) for semantic scene parsing. Unlike previous methods, which separately classify each object, HPN leverages global scene semantic information and the context among multiple objects to enhance scene parsing. On the one hand, HPN uses the global scene category to constrain the semantic consistency between the scene and each object. On the other hand, the context among all objects is also modeled to avoid incompatible object predictions. Specifically, HPN consists of four steps. In the first step, we extract scene and local appearance features. Based on these appearance features, the second step is to encode a contextual feature for each object, which models both the scene-object context (the context between the scene and each object) and the interobject context (the context among different objects). In the third step, we classify the global scene and then use the scene classification loss and a backpropagation algorithm to constrain the scene feature encoding. In the fourth step, a label map for scene parsing is generated from the local appearance and contextual features. Our model outperforms many state-of-the-art deep scene parsing networks on five scene parsing databases. Hengcan Shi, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, Linfeng Xu 0001, King Ngi Ngan |
IEEE Trans. Multim. | 6 |
| 2017 | A Generative Model for Depth-Based Robust 3D Facial Pose TrackingabstractWe consider the problem of depth-based robust 3D facial pose tracking under unconstrained scenarios with heavy occlusions and arbitrary facial expression variations. Unlike the previous depth-based discriminative or data-driven methods that require sophisticated training or manual intervention, we propose a generative framework that unifies pose tracking and face model adaptation on-the-fly. Particularly, we propose a statistical 3D face model that owns the flexibility to generate and predict the distribution and uncertainty underlying the face model. Moreover, unlike prior arts employing the ICP-based facial pose estimation, we propose a ray visibility constraint that regularizes the pose based on the face models visibility against the input point cloud, which augments the robustness against the occlusions. The experimental results on Biwi and ICT-3DHP datasets reveal that the proposed framework is effective and outperforms the state-of-the-art depth-based methods. Lu Sheng, Jianfei Cai 0001, Tat-Jen Cham, Vladimir Pavlovic 0001, King Ngi Ngan |
CVPR | 5 |
| 2017 | Blind proposal quality assessment via deep objectness representation and local linear regressionabstractThe quality of object proposal plays an important role in boosting the performance of many computer vision tasks, such as, object detection and recognition. Due to the absence of manually annotated bounding-box in practice, the quality metric towards blind assessment of object proposal is highly desirable for singling out the optimal proposals. In this paper, we propose a blind proposal quality assessment algorithm based on the Deep Objectness Representation and Local Linear Regression (DORLLR). Inspired by the hierarchy model of the human vision system, a deep convolutional neural network is developed to extract the objectness-aware image feature. Then, the local linear regression method is utilized to map the image feature to a quality score, which tries to evaluate each individual test window based on its k-nearest-neighbors. Experimental results on a large-scale IoU labeled dataset verify that the proposed method significantly outperforms the state-of-the-art blind proposal evaluation metrics. Qingbo Wu 0001, Hongliang Li 0001, Fanman Meng, King Ngi Ngan, Linfeng Xu 0001 |
ICME | 4 |
| 2017 | Improving object proposals with top-down cues
Wei Li 0110, Hongliang Li 0001, Bing Luo 0003, Hengcan Shi, Qingbo Wu 0001, King Ngi Ngan |
Signal Process. Image Commun. | 6 |
| 2017 | Gaze-Based Object SegmentationabstractThis letter addresses the problem of object segmentation with a gaze map and a group of candidate regions. First, we analyze distribution characteristics of gazes when people look at an image. Then, we summarize different cases of the group of candidate regions. Based on our analysis and summary, we develop three measures to evaluate the likelihood of a candidate region belonging to the target object and a pooling method to create a likelihood map of this object. Finally, the measures and pooling method are integrated with a proposed iterative strategy for generating the segmentation result. Experimental results demonstrate that our method can handle different types of gaze maps and different groups of candidate regions, and the overall performance of our method is better than that of the state-of-the-art method. King Ngi Ngan, Hongliang Li 0001 |
IEEE Signal Process. Lett. | 2 |
| 2017 | Weakly Supervised Part Proposal Segmentation From Multiple ImagesabstractWeakly supervised local part segmentation is challenging, due to the difficulty of modeling multiple local parts from image level prior. In this paper, we propose a new weakly supervised local part proposal segmentation method based on the observation that local parts will keep fixed along the object pose variations. Hence, the local part can be segmented by capturing object pose variations. Based on such observation, a new local part proposal segmentation model is proposed. Three aspects, such as shape similarity-based cosegmentation, shape matching-based part detection and segmentation, and graph matching-based part assignment are considered. A part segmentation energy function is first proposed. Four terms, such as MRF-based single image segmentation term, shape feature-based foreground consistency term, NCuts-based part segmentation term, and two-order graphs matching based part consistency term, are contained. Then, a three sub-minimization-based energy minimization method is proposed to accomplish approximation solution. Finally, we verify our method based on three image data sets (PASCAL VOC 2008 Part data set, UCB Bird data set, and Cat-Dog data set), and one video data set (UCF Sports) data set. The experimental results demonstrate a better segmentation performance compared with the existing object cosegmentation and part proposal generation methods. Fanman Meng, Hongliang Li 0001, Qingbo Wu 0001, Bing Luo 0003, King Ngi Ngan |
IEEE Trans. Image Process. | 5 |
| 2017 | Objective Quality Assessment of Image Retargeting by Incorporating Fidelity Measures and Inconsistency DetectionabstractThe tremendous growth in mobile devices has resulted in huge generation and usage of digital images. Image quality assessment is thus an important issue for mobile media applications. In this paper, we focus on the quality evaluation of images generated by content-aware image retargeting, in which the reference and the distorted images are of different sizes. Through retargeting, many types of deformation inconsistency lead to shape distortion, deformation artifacts, and content information loss, worsening its perceptual quality. The deformation inconsistency occurs on different levels of the retargeted images. Limited by the accuracy of the alignment between the original and retargeted images, previous methods only focus on pixel-level and patch-level fidelity analyses and fail to detect deformation inconsistency. In this paper, we improve the alignment algorithm and propose a three-level representation of the retargeting process. Based on the analysis of this three-level representation, both fidelity measures and inconsistency detection are combined to determine the final retargeting quality. The proposed algorithm is validated on the public data sets RetargetMe and CUHK. Experimental results demonstrate that inconsistency detection contributes to accurately assessing the image retargeting perceptual quality. This inspires us to investigate more about deformation inconsistency to formulate the objective quality of image retargeting. Yichi Zhang 0014, King Ngi Ngan, Lin Ma 0002, Hongliang Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | Learning Efficient Binary Codes From High-Level Feature Representations for Multilabel Image RetrievalabstractDue to the efficiency and effectiveness of hashing technologies, they have become increasingly popular in large-scale image semantic retrieval. However, existing hash methods suppose that the data distributions satisfy the manifold assumption that semantic similar samples tend to lie on a low-dimensional manifold, which will be weakened due to the large intraclass variation. Moreover, these methods learn hash functions by relaxing the discrete constraints on binary codes to real value, which will introduce large quantization loss. To tackle the above problems, this paper proposes a novel unsupervised hashing algorithm to learn efficient binary codes from high-level feature representations. More specifically, we explore nonnegative matrix factorization for learning high-level visual features. Ultimately, binary codes are generated by performing binary quantization in the high-level feature representations space, which will map images with similar (visually or semantically) high-level feature representations to similar binary codes. To solve the corresponding optimization problem involving nonnegative and discrete variables, we develop an efficient optimization algorithm to reduce quantization loss with guaranteed convergence in theory. Extensive experiments show that our proposed method outperforms the state-of-the-art hashing methods on several multilabel real-world image datasets. Lei Ma 0004, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, King Ngi Ngan |
IEEE Trans. Multim. | 5 |
| 2017 | Blind Image Quality Assessment Based on Rank-Order Regularized RegressionabstractBlind image quality assessment (BIQA) aims to estimate the subjective quality of a query image without access to the reference image. Existing learning-based methods typically train a regression function by minimizing the average error between subjective opinion scores and model predictions. However, minimizing average error does not necessarily lead to correct quality rank-orders between the test images, which is a highly desirable property of image quality models. In this paper, we propose a novel rank-order regularized regression model to address this problem. The key idea is to introduce a pairwise rank-order constraint into the maximum margin regression framework, aiming to better preserve the correct perceptual preference. To the best of our knowledge, this is the first attempt to incorporate rank-order constraints into margin-based quality regression model. By combing with a new local spatial structure feature, we achieve highly consistent quality prediction with human perception. Experimental results show that the proposed method outperforms many state-of-the-art BIQA metrics on popular publicly available IQA databases (i.e., LIVE-II, TID2013, VCL@FER, LIVEMD, and ChallengeDB). Qingbo Wu 0001, Hongliang Li 0001, Zhou Wang 0001, Fanman Meng, Bing Luo 0003, Wei Li 0110, King Ngi Ngan |
IEEE Trans. Multim. | 7 |
| 2016 | Material segmentation in hyperspectral images with minimal region perimetersabstractWe propose a supervised approach to the classification and segmentation of material regions in hyperspectral imagery. Our algorithm is a two-stage process, combining a pixelwise classification step with a segmentation step aiming to minimise the total perimeters of the resulting regions. Our algorithm is distinctive in its ability to ensure label consistency within local homogeneous areas and to generate material segments with smooth boundaries. Furthermore, we establish a new hyperspectral benchmark dataset to demonstrate the advantages of the proposed approach over several state-of-the-art methods. Yu Zhang 0004, Cong Phuoc Huynh, Nariman Habili, King Ngi Ngan |
ICIP | 4 |
| 2016 | Fast patch-wise image retargetingabstractContent-aware image retargeting adjusts images to arbitrary sizes and preserves visually salient content. Previous algorithms formulate the problem in terms of either pixel level or mesh level structures, deforming salient objects inconsistently. To improve retargeting quality and reduce complexity, we introduced a patch-wise method to generate sparse image grids based on visual saliency and gradient magnitude. Three energy functions were optimized to warp the grids and generate retargeted images. Experimental results on the public database RetargetMe show that this patch-wise retargeting algorithm has lower complexity and performs slightly better than other algorithms using much denser grids. Yichi Zhang 0014, King Ngi Ngan |
ICIP | 2 |
| 2016 | Model-based face reconstruction using SIFT flow registration and spherical harmonicsabstractIn this paper, we propose a robust method for face reconstruction using a single color image. A 3D morphable model is used to reconstruct a smooth 3D face shape. To find the correspondence between model vertices and image pixels, landmarks are updated using SIFT flow which is illumination and rotation invariant. To reconstruct more detailed information, depth values are refined using a shape from shading method which approximates lighting condition by spherical harmonics. We test the proposed method on a set of real world images and compare reconstructed results with depth maps captured by a depth camera. The average error is around 3.3 mm. Fanzi Wu, Songnan Li, King Ngi Ngan |
ICPR | 4 |
| 2016 | Objective quality assessment of image retargeting based on line distortionabstractWith the proliferation of mobile devices, research on image retargeting is becoming ever more important. However, there is little work on image retargeting quality assessment despite its importance. In this work, we focus on evaluating retargeting quality based on line distortion. Generally, image retargeting results in content loss and shape distortion. Line segments, which are fundamental image structures, are hence discarded or distorted in retargeted images. As a result, we formulate a retargeting quality index consisted of three line distortion measures: line loss, line artifact and line rotation. To test its performance, we have validated it on the public dataset RetargetMe. Experimental results demonstrate that our method outperforms many existent ones and line distortion is a good indicator of retargeting quality. Yichi Zhang 0014, King Ngi Ngan |
SMC | 2 |
| 2016 | Q-DNN: A quality-aware deep neural network for blind assessment of enhanced imagesabstractImage enhancement is widely popular due to its capability of producing "better" visual quality for specific applications. Although many enhancement algorithms have been developed in recent years, the studies towards blind assessment of enhanced images are still very lacking. In this paper, we propose a data-driven blind image quality assessment (BIQA) method based on the quality-aware deep neural network (Q-DNN). Unlike the conventional hand-crafted features designed for measuring the degradation level of specific distortion types, a supervised learning model is utilized in our Q-DNN, which is capable of adaptively updating the feature extractor and quality regressor for describing the visual artifacts caused by different image enhancement tasks. Experimental results on two challenging enhanced image databases show that the proposed method is significantly superior to the state-of-the-art BIQA metrics. Qingbo Wu 0001, Hongliang Li 0001, Fanman Meng, King Ngi Ngan |
VCIP | 4 |
| 2016 | Hybrid human detection and recognition in surveillance
Qiang Liu 0015, Wei Zhang 0021, Hongliang Li 0001, King Ngi Ngan |
Neurocomputing | 4 |
| 2016 | Reorganized DCT-based image representation for reduced reference stereoscopic image quality assessment
Lin Ma 0002, Xu Wang 0006, Qiong Liu 0001, King Ngi Ngan |
Neurocomputing | 4 |
| 2016 | Multi-layer authentication scheme for HEVC video based on embedded statistics
Yiqi Tew, Koksheik Wong, Raphael C.-W. Phan, King Ngi Ngan |
J. Vis. Commun. Image Represent. | 4 |
| 2016 | Perceptual sensitivity-based rate control method for high efficiency video coding
Huanqiang Zeng, Aisheng Yang, King Ngi Ngan, Miaohui Wang |
Multim. Tools Appl. | 3 |
| 2016 | Real-Time Head Pose Tracking with Online Face Template ReconstructionabstractWe propose a real-time method to accurately track the human head pose in the 3-dimensional (3D) world. Using a RGB-Depth camera, a face template is reconstructed by fitting a 3D morphable face model, and the head pose is determined by registering this user-specific face template to the input depth video. Songnan Li, King Ngi Ngan, Raveendran Paramesran, Lu Sheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Blind Image Quality Assessment Based on Multichannel Feature Fusion and Label TransferabstractIn this paper, we propose an efficient blind image quality assessment (BIQA) algorithm, which is characterized by a new feature fusion scheme and a k-nearest-neighbor (KNN)-based quality prediction model. Our goal is to predict the perceptual quality of an image without any prior information of its reference image and distortion type. Since the reference image is inaccessible in many applications, the BIQA is quite desirable in this context. In our method, a new feature fusion scheme is first introduced by combining an image's statistical information from multiple domains (i.e., discrete cosine transform, wavelet, and spatial domains) and multiple color channels (i.e., Y, Cb, and Cr). Then, the predicted image quality is generated from a nonparametric model, which is referred to as the label transfer (LT). Based on the assumption that similar images share similar perceptual qualities, we implement the LT with an image retrieval procedure, where a query image's KNNs are searched for from some annotated images. The weighted average of the KNN labels (e.g., difference mean opinion score or mean opinion score) is used as the predicted quality score. The proposed method is straightforward and computationally appealing. Experimental results on three publicly available databases (i.e., LIVE II, TID2008, and CSIQ) show that the proposed method is highly consistent with human perception and outperforms many representative BIQA metrics. Qingbo Wu 0001, Hongliang Li 0001, Fanman Meng, King Ngi Ngan, Bing Luo 0003, Chao Huang 0003, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2016 | Low-Delay Rate Control for Consistent Quality Using Distortion-Based Lagrange MultiplierabstractVideo quality fluctuation plays a significant role in human visual perception, and hence, many rate control approaches have been widely developed to maintain consistent quality for video communication. This paper presents a novel rate control framework based on the Lagrange multiplier in high-efficiency video coding. With the assumption of constant quality control, a new relationship between the distortion and the Lagrange multiplier is established. Based on the proposed distortion model and buffer status, we obtain a computationally feasible solution to the problem of minimizing the distortion variation across video frames at the coding tree unit level. Extensive simulation results show that our method outperforms the rate control used in HEVC Test Model (HM) by providing a more accurate rate regulation, lower video quality fluctuation, and stabler buffer fullness. The average peak signal-to-noise ratio (PSNR) and PSNR deviation improvements are about 0.37 dB and 57.14% in the low-delay (P and B) video communication, where the complexity overhead is ∼ 4.44% . Miaohui Wang, King Ngi Ngan, Hongliang Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2016 | No-Reference Retargeted Image Quality Assessment Based on Pairwise Rank LearningabstractIn this paper, we propose a novel no-reference image quality assessment method for the retargeted image based on the pairwise rank learning approach. Each retargeted image needs to be first represented as a feature vector, which not only captures the image characteristics but also is sensitive to distortions during the retargeting process. As such, we investigate and examine different image representations for their abilities depicting the perceptual quality of retargeted image. Based on the image representations, we resort to the pairwise rank learning approach to discriminate the perceptual quality between the retargeted image pairs. Experimental results demonstrate that the proposed method can effectively depict the perceptual quality of the retargeted image, which can even perform comparably with the full-reference quality assessment methods. Lin Ma 0002, Long Xu 0001, Yichi Zhang 0014, Yihua Yan, King Ngi Ngan |
IEEE Trans. Multim. | 5 |
| 2016 | Free-Energy Principle Inspired Video Quality Metric and Its Use in Video CodingabstractIn this paper, we extend the free-energy principle to video quality assessment (VQA) by incorporating with the recent psychophysical study on human visual speed perception (HVSP). A novel video quality metric, namely the free-energy principle inspired video quality metric (FePVQ), is therefore developed and applied to perceptual video coding optimization. The free-energy principle suggests that the human visual system (HVS) can actively predict “orderly” information and avoid “disorderly” information for image perception. Basically, “orderly” is associated with the skeletons and edges of objects, and “disorderly” mostly concerns textures in images. Based on this principle, an image is separated into orderly and disorderly regions, and processed differently in image quality assessment. For videos, visual attention, or fixation, is associated with the objects with significant motion according to HVSP, resulting in a motion strength factor in the FePVQ so that the free-energy principle is extended into spatio-temporal domain for VQA. In addition, we investigate the application of the FePVQ in perceptual rate distortion optimization (RDO). For this purpose, the FePVQ is realized with low computational cost by using the relative total variation model and the block-wise motion vectors of video coding to simulate the free-energy principle and the HVSP, respectively. The experimental results indicate that the proposed FePVQ is highly consistent with the HVS perception. The linear correlation coefficient and Spearman's rank-order correlation coefficient are up to 0.8324 and 0.8281 on the LIVE video database. Better perceptual quality of encoded video sequences is achieved by FePVQ-motivated RDO in video coding. Long Xu 0001, Weisi Lin, Lin Ma 0002, Yongbing Zhang 0002, Yuming Fang 0001, King Ngi Ngan, Songnan Li, Yihua Yan |
IEEE Trans. Multim. | 6 |
| 2015 | Optimal bit allocation in HEVC for real-time video communicationsabstractRecently, the Lagrange multiplier λ based rate control has been developed in the latest HEVC (High Efficiency Video Coding) codec. It is revealed that while such a rate control approach successfully guarantees a designated bit-rate, the optimal bit allocation has not been investigated or established in the λ-domain. To enhance the overall performance, we propose a novel linear relationship between distortion and λ. With this distortion model, we deduce a closed-form solution to minimize the distortion while satisfying a given rate regulation at the coding tree unit level. Extensive simulation results show that our method can achieve an average 0.20 dB improvement in the low-delay communications. Miaohui Wang, King Ngi Ngan |
ICIP | 2 |
| 2015 | Region-based image retargeting quality assessmentabstractThe tremendous growth in mobile devices has resulted in huge generation and consumption of digital images. The image quality evaluation is hence an important issue for mobile media applications. In this work, we focus on the quality evaluation of images under content-aware retargeting, in which the sizes of images may differ from each other. Shape distortion and content information loss are two main factors in the evaluation of retargeting quality [1]. We have proposed a region-based framework to evaluate shape distortion and information loss caused by retargeting. The proposed algorithm is validated on the public database RetargetMe [2]. Experimental results demonstrate that this region-based quality assessment algorithm outperforms others based on pixels. Yichi Zhang 0014, King Ngi Ngan |
ICIP | 2 |
| 2015 | Improved block level adaptive quantization for high efficiency video codingabstractAs the concept of block level adaptivity becomes an important feature in recent video CODECs, block level adaptive quantization (BLAQ) is being considered in the High Efficiency Video Coding (HEVC) standard. The BLAQ is based on the assumption that each block should have its own quantization parameter (QP), which can adapt to the local content of video sequences much better, and hence the video encoder with adaptive QP can perform a better perceptual quality. However, in the HEVC reference software, the BLAQ is required to obtain a proper QP for each block by the rate distortion optimization (RDO) scheme and so the computational complexity of the encoder increases significantly. In this paper, an improved BLAQ algorithm is proposed to obtain the adaptive QP for each block. The simulation results show that the proposed method can save more bits as well as require lower computational complexity, compared to the traditional method. Miaohui Wang, King Ngi Ngan, Hongliang Li 0001, Huanqiang Zeng |
ISCAS | 2 |
| 2015 | Rank Learning Based No-Reference Quality Assessment of Retargeted ImagesabstractIn this paper, we first propose a novel no-reference (NR) image quality assessment (IQA) method for retargeted image based on the rank learning approach. Firstly, image features for each retargeted image are extracted, which should not only represent the image characteristics but also be sensitive to the retargeted distortions. Specifically, the image feature should be able to capture the shape distortions, which are the commonly encountered distortions of the retargeted image. Based on the extracted image features, the rank learning method is employed to train a model to discriminate the perceptual quality of the retargeted image. Experimental results demonstrate that the proposed method can effectively depict the perceptual quality of the retargeted image, which can even perform comparably with the full-reference (FR) quality assessment methods. Lin Ma 0002, Long Xu 0001, Yichi Zhang 0014, King Ngi Ngan, Yihua Yan |
SMC | 4 |
| 2015 | No reference image quality assessment metric via multi-domain structural information and piecewise regression
Qingbo Wu 0001, Hongliang Li 0001, Fanman Meng, King Ngi Ngan, Shuyuan Zhu |
J. Vis. Commun. Image Represent. | 4 |
| 2015 | An Efficient Frame-Content Based Intra Frame Rate Control for High Efficiency Video CodingabstractRate control plays an important role in the rapid development of high-fidelity video services. As the High Efficiency Video Coding (HEVC) standard has been finalized, many rate control algorithms are being developed to promote its commercial use. The HEVC encoder adopts a new R-lambda based rate control model to reduce the bit estimation error. However, the R-lambda model fails to consider the frame-content complexity that ultimately degrades the performance of the bit rate control. In this letter, a gradient based R-lambda (GRL) model is proposed for the intra frame rate control, where the gradient can effectively measure the frame-content complexity and enhance the performance of the traditional R-lambda method. In addition, a new coding tree unit (CTU) level bit allocation method is developed. The simulation results show that the proposed GRL method can reduce the bit estimation error and improve the video quality in HEVC all intra frame coding. Miaohui Wang, King Ngi Ngan, Hongliang Li 0001 |
IEEE Signal Process. Lett. | 2 |
| 2015 | Online Temporally Consistent Indoor Depth Video Enhancement via Static StructureabstractIn this paper, we propose a new method to online enhance the quality of a depth video based on the intermediary of a so-called static structure of the captured scene. The static and dynamic regions of the input depth frame are robustly separated by a layer assignment procedure, in which the dynamic part stays in the front while the static part fits and helps to update this structure by a novel online variational generative model with added spatial refinement. The dynamic content is enhanced spatially while the static region is otherwise substituted by the updated static structure so as to favor the long-range spatiotemporal enhancement. The proposed method both performs long-range temporal consistency on the static region and keeps necessary depth variations in the dynamic content. Thus, it can produce flicker-free and spatially optimized depth videos with reduced motion blur and depth distortion. Our experimental results reveal that the proposed method is effective in both static and dynamic indoor scenes and is compatible with depth videos captured by Kinect and time-of-flight camera. We also demonstrate that excellent performance can be achieved by the proposed method in comparison with the existing spatiotemporal approaches. In addition, our enhanced depth videos and static structures can act as effective cues to improve various applications, including depth-aided background subtraction and novel view synthesis, showing satisfactory results with few visual artifacts. Lu Sheng, King Ngi Ngan, Chern-Loon Lim, Songnan Li |
IEEE Trans. Image Process. | 2 |
| 2015 | Visual Quality Evaluation of Image Object Segmentation: Subjective Assessment and Objective MeasureabstractA visual quality evaluation of image object segmentation as one member of the visual quality evaluation family has been studied over the years. Researchers aim at developing the objective measures that can evaluate the visual quality of object segmentation results in agreement with human quality judgments. It is also significant to construct a platform for evaluating the performance of the objective measures in order to analyze their pros and cons. In this paper, first, we present a novel subjective object segmentation visual quality database, in which a total of 255 segmentation results were evaluated by more than thirty human subjects. Then, we propose a novel full-reference objective measure for an object segmentation visual quality evaluation, which involves four human visual properties. Finally, our measure is compared with some state-of-the-art objective measures on our database. The experiment demonstrates that the proposed measure performs better in matching subjective judgments. Moreover, the database is available publicly for other researchers in the field to evaluate their measures. King Ngi Ngan, Songnan Li, Raveendran Paramesran, Hongliang Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2015 | Fast HEVC Inter CU Decision Based on Latent SAD EstimationabstractThe emerging high efficiency video coding (HEVC) standard has improved compression performance significantly in comparison with H.264/AVC. However, more intensive computational complexity has been introduced by adopting a number of new coding tools. In this paper, a fast inter CU decision is proposed based on the latent sum of absolute differences (SAD) estimation. Firstly, a two-layer motion estimation (ME) method is designed to take advantage of the latent SAD cost. The new ME method can obtain the SAD costs for both the upper CU and its sub-CUs. Secondly, a concept of motion compensation rate- distortion (R-D) cost is defined, and an exponential model is proposed to express the relationship between the motion compensation R-D cost and the SAD cost. Then, a fast CU decision approach is designed based on the exponential model. The fast CU decision is implemented by comparing a derived threshold with the SAD cost difference between the upper and sub SAD costs. Experimental results show that the proposed algorithm achieves an average of 52% and 58.4% reductions of the coding time at the cost of 1.61% and 2% bit-rate increases under the low delay and random access conditions, respectively. Jian Xiong 0005, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, King Ngi Ngan |
IEEE Trans. Multim. | 5 |
| 2014 | Accelerating the Distribution Estimation for the Weighted Median/Mode Filters
Lu Sheng, King Ngi Ngan, Tak-Wai Hui |
ACCV (4) | 2 |
| 2014 | Motion-Depth: RGB-D Depth Map Enhancement with Motion and Depth in ComplementabstractLow-cost RGB-D imaging system such as Kinect is widely utilized for dense 3D reconstruction. However, RGB-D system generally suffers from two main problems. The spatial resolution of the depth image is low. The depth image often contains numerous holes where no depth measurements are available. This can be due to bad infra-red reflectance properties of some objects in the scene. Since the spatial resolution of the color image is generally higher than that of the depth image, this paper introduces a new method to enhance the depth images captured by a moving RGB-D system using the depth cues from the induced optical flow. We not only fill the holes in the raw depth images, but also recover fine details of the imaged scene. We address the problem of depth image enhancement by minimizing an energy functional. In order to reduce the computational complexity, we have treated the textured and homogeneous regions in the color images differently. Experimental results on several RGB-D sequences are provided to show the effectiveness of the proposed method. Tak-Wai Hui, King Ngi Ngan |
CVPR | 2 |
| 2014 | Depth enhancement using RGB-D guided filteringabstractDepth maps from low-cost RGB-D system are generally noisy and not accurate enough. Holes often exist in the depth maps. Bilateral filter is commonly utilized to perform depth enhancement. However, it requires high computational time. Its texture transferring property also makes those boundaries between textured and homogeneous regions in the filtered depth map far from satisfactory. In this paper, we present a method to filter raw depth maps using a RGB-D guided filtering in a two-stage framework. Our method not only has a faster computational time than bilateral filter but also avoids the problem of over-texture transfer. We also use RGB-D frames to fill holes in the depth maps. This can effectively prevents depth bleeding artifacts. Tak-Wai Hui, King Ngi Ngan |
ICIP | 2 |
| 2014 | Dense depth map generation using sparse depth data from normal flowabstractIn this paper, we address the problem of dense depth map generation from two successive image frames in a video. We first recover the camera motion from the observable normal flow pattern using our previously proposed apparent flow constraints. Once the camera motion is estimated, sparse depth data can be directly recovered from the flow pattern. We utilize a hierarchical approach to generate an initial dense depth map from the sparse depth data. This depth map is further enhanced through the refinement of the associated optical flow field in a variational framework. Experimental results show that the proposed method can provide high-quality depth maps. We also have a faster computational time than the conventional optical flow approach. Tak-Wai Hui, King Ngi Ngan |
ICIP | 2 |
| 2014 | Screen-camera calibration using a threadabstractIn this paper, we propose a novel screen-camera calibration algorithm which aims to locate the position of the screen in the camera coordinate system. The difficulty comes from the fact that the screen is not directly visible to the camera. Rather than using an external camera or a portable mirror like in previous studies, we propose to use a more accessible and cheaper calibrating object, i.e., a thread. The thread is manipulated so that our algorithm can infer the perspective projections of the four screen corners on the image plane. The 3-dimentional (3D) position of each screen corner is then determined by minimizing the sum of squared projection errors. Experiments show that compared with the previous studies our method can generate similar calibration results without the additional hardware. Songnan Li, King Ngi Ngan, Lu Sheng |
ICIP | 2 |
| 2014 | Temporal depth video enhancement based on intrinsic static structureabstractDepth video enhancement is an essential preprocessing step for various 3D applications. Despite extensive studies of spatial enhancement, effective temporal enhancement that both strengthens temporal consistency and keeps correct depth variation needs further research. In this paper, we propose a novel method to enhance the depth video by blending raw depth frame with the estimated intrinsic static structure, which defines static structure of captured scene and is estimated iteratively by a probabilistic generative model with sequentially incoming depth frames. Our experimental results show that the proposed method is effective both in static and dynamic scene and is compatible with various kinds of depth videos. We will demonstrate that superior performance can be achieved in comparison with existing temporal enhancement approaches. Lu Sheng, King Ngi Ngan, Songnan Li |
ICIP | 2 |
| 2014 | Jaccard index compensation for object segmentation evaluationabstractIn this paper, we propose an objective metric for quality evaluation of individual object segmentation in images. Using the Jaccard Index as a base, additional compensation terms are integrated into our metric. These terms not only allow our metric to combine region-based and boundary-based methods, they also describe human visual tolerance and saturation. This metric can also can adaptively adjust “perception function” under different image sizes. The experiment on our subjective object segmentation quality assessment database demonstrates that the proposed metric performs well in matching subjective ratings. King Ngi Ngan, Songnan Li |
ICIP | 2 |
| 2014 | Using mid-high level cues to detect salient objectabstractThis paper proposes a novel saliency object detection method by using the mid-level and high-level visual cues. In the mid-level objectness evaluation, we generate three complementary saliency maps, such as the multi-scale segmentation cue, the background cue and the spatial color distribution cue. The first cue is used to highlight the objects via the local region segment. The second cue uses the background priors to detect the saliency information. The third cue is to capture the spatial color distribution. For the high-level visual cue, we propose an objectness evaluation model to distinguish the object and the background. All the saliency cues are finally combined to achieve the saliency detection. The experimental results show that the proposed method outperforms the state-of-the-art saliency object detection methods. Hongliang Li 0001, Yurui Xie, Bing Luo 0003, Liangzhi Tang, Bing Zeng 0001, King Ngi Ngan, Fanman Meng |
ICME | 6 |
| 2014 | Cosegmentation from similar backgroundsabstractRecently, the common objects are often required to be extracted from a group of images in many applications, such as video coding and model training. Co-segmentation is a new and efficient method for this requirement. In realistic applications, we observe that the images usually contain similar backgrounds (namely similar scene co-segmentation), such as the city landmark images collected from the web or the key frames sampled from a video. Meanwhile, the existing co-segmentation has not paid so much attention on the similar scene co-segmentation, and the insufficiently accurate segments may be provided by the existing methods. In this paper, we propose an active contours based co-segmentation model to provide foregrounds from the similar backgrounds. We combine the background consistency constraint with the foreground consistency constraint to form the energy function, and use the method of level-set and the calculus of variations to minimize the model. We also speed up the model by the hierarchical structure and the superpixel technique. We test the method on both the image and video dataset. The results show that the proposed model can obtain larger IOU values than the state-of-the-art co-segmentation methods. Fanman Meng, Hongliang Li 0001, King Ngi Ngan, Bing Zeng 0001, Nini Rao |
ISCAS | 3 |
| 2014 | No reference image quality metric via distortion identification and multi-channel label transferabstractIn this paper, we propose a no reference image quality assessment (NR-IQA) algorithm based on distortion identification (DI) and multi-channel label transfer (LT). First, the distortion type classification is used to obtain the query image's probabilities of belonging to each distortion type. Then, the distortion specific label transfer is implemented in multiple distortion category channels. Based on the hypothesis that the similar images share the similar subjective qualities, the label transfer predicts the subjective quality of the query image by pooling the labels of its k-nearest neighbors (KNN) retrieved from the annotated samples. A weighting average of the multi-channel label transfer's outputs is computed to obtain the final perceptual quality score. The weight is the query image's probability that belongs to the corresponding distortion type. The experimental results show that the proposed method outperforms representative NR-IQA approaches and some full-reference metrics. Qingbo Wu 0001, Hongliang Li 0001, King Ngi Ngan, Bing Zeng 0001, Moncef Gabbouj |
ISCAS | 3 |
| 2014 | Efficient H.264/AVC Video Coding with Adaptive TransformsabstractTransform has been widely used to remove spatial redundancy of prediction residuals in the modern video coding standards. However, since the residual blocks exhibit diverse characteristics in a video sequence, conventional transform methods with fixed transform kernels may result in low efficiency. To tackle this problem, we propose a novel content adaptive transform framework for the H.264/AVC-based video coding. The proposed method utilizes pixel rearrangement to dynamically adjust the transform kernels to adapt to the video content. In addition, unlike the traditional adaptive transforms, the proposed method obtains the transform kernels from the reconstructed block, and hence it consumes only one logic indicator for each transform unit. Moreover, a spiral-scanning method is developed to reorder the transform coefficients for better entropy coding. Experimental results on the Key Technical Area (KTA) platform show that the proposed method can achieve an average bitrate reduction of about 7.95% and 7.0% under all-intra and low-delay configurations, respectively. Miaohui Wang, King Ngi Ngan, Long Xu 0001 |
IEEE Trans. Multim. | 2 |
| 2013 | The Objective Evaluation of Image Object Segmentation Quality
King Ngi Ngan, Songnan Li |
ACIVS | 2 |
| 2013 | Depth enhancement based on hybrid geometric hole filling strategyabstractDepth map is a crucial component in various 3D applications. However, current available depth maps usually suffer from low resolution, high noise, random and structural depth missing problems due to theoretical, systematic or hardware limitations. In this paper, we propose a novel method to enhance depth map with the guidance of aligned color image, tackling these problems in a whole framework, where a hybrid strategy on filling hole geometrically by the combination of joint bilateral filtering and segment-based surface structure propagation is introduced. Our experimental results prove the proposed method outperforms existing methods. Lu Sheng, King Ngi Ngan |
ICIP | 2 |
| 2013 | Reduced reference video quality assessment based on spatial HVS mutual masking and temporal motion estimationabstractIn this paper, an effective reduced reference (RR) video quality assessment (VQA) is proposed by depicting both the spatial and temporal statistical characteristics of the video signals. For each video frame, spatial information change (SIC) is employed to depict the energy variation. A novel mutual masking strategy based on the extracted SIC is proposed to accurately simulate the human visual system (HVS) texture masking property. For adjacent video frames, the temporal relationship is depicted by block-based motion estimation (BME). The generalized Gaussian density (GGD) function is employed to depict the histogram natural statistic of the residual frame after BME. The city-block distance (CBD) is used to measure the distance between histograms of the original and distorted video sequence. By pooling the measurements from both spatial and temporal perspectives, an efficient RR VQA is constructed. With the evaluations on the public video quality database, the proposed RR VQA demonstrated to be more effective than the representative RR VQAs and even the full-reference (FR) VQAs, such as peak signal-to-noise ratio (PSNR) and structure similarity index (SSIM) in matching the subjective ratings. Furthermore, the proposed RR VQA demonstrated to be much more effective and efficient, requiring only a very small number of bits for the RR feature representation. Lin Ma 0002, King Ngi Ngan, Long Xu 0001 |
ICME | 2 |
| 2013 | A Head Pose Tracking System Using RGB-D Camera
Songnan Li, King Ngi Ngan, Lu Sheng |
ICVS | 2 |
| 2013 | Object segmentation from wide baseline videoabstractIn this paper, we propose an automatic approach to segment object from stereo videos, for which the viewpoints are widely apart. We first present a novel saliency analysis to emphasize the foreground object. The saliency map is estimated by combining the depth information recovered by feature matching and the boundary information revealed by color segmentation. The object mask is extracted initially based on the saliency map and then refined by graph-cut segmentation, where color and motion information are efficiently incorporated in both data and smoothness terms. Moreover, a background image is gradually reconstructed during video segmentation, based on which an additional constraint is imposed on the data term to further improve the video segmentation. The proposed method is tested on stereo videos with widely separated viewpoints and severe background clutters. Good experimental results demonstrate the feasibility of the proposed method. Chunhui Cui, Qian Zhang 0001, King Ngi Ngan |
ISCAS | 3 |
| 2013 | Overview of quality assessment for visual signals and newly emerged trendsabstractQuality assessment is not only essential on its own for testing, optimizing, benchmarking, monitoring and inspecting related systems and services, but also plays an essential role in the design of virtually all visual signal processing and communication algorithms, as well as various related decision making processes. In the paper, we provide an overview of the quality assessment approaches for traditional visual signals, as well as the newly emerged ones, which covers the subjective quality evaluation and objective quality metrics of scalable and mobile videos, high dynamic range (HDR) images, image segmentation results, 3D images/videos, and retargeted images. Also the challenges for designing effective quality metrics and corresponding applications are discussed. Lin Ma 0002, Chenwei Deng, King Ngi Ngan, Weisi Lin |
ISCAS | 3 |
| 2013 | High quality image construction from multiple low quality copiesabstractIn this paper, the authors proposed to construct a high quality image based on multiple low quality input images. The relationship between one pixel and its neighbourhood should be consistent between different degraded images. Therefore, the reconstruction method is proposed by enforcing the pixel consistency property, which is ensured by estimating the parameters of piecewise image model for each pixel. Subsequently, the reconstructed coefficients are regularized within a reasonable range. Experimental results on multiple images (with different distortions of different levels) have demonstrated that the proposed method can effectively alleviate the noises meanwhile preserve the detailed information. Better quality images in terms of both objective and subjective measurements can be generated. Lin Ma 0002, Long Xu 0001, Qian Zhang 0001, King Ngi Ngan |
MMSP | 4 |
| 2013 | A rate distortion optimized transform for motion compensation residualabstractIn this paper, we propose a rate distortion optimization based content adaptive transform method for motion compensation residuals. The proposed method utilizes pixel rearrangement to dynamically adjust the transform kernels to adapt to the residual content. Comparing with the traditional adaptive transforms, the highlight of this work is that it obtains the transform kernels from the decoded block, and hence it consumes only one overhead bit for each transform unit. Moreover, rate distortion optimization scheme is used to choose the best candidate kernels. Experimental results show that the proposed method achieves an average 0.35 dB gain of PSNR in comparison with the key technical areas (KTA) encoder. Miaohui Wang, King Ngi Ngan, Huanqiang Zeng |
PCS | 2 |
| 2013 | Perceptual adaptive Lagrangian multiplier for high efficiency video codingabstractIn high efficiency video coding (HEVC), Lagrangian rate distortion optimization (RDO) technique is used to optimize the rate distortion (RD) performance. However, the corresponding Lagrangian multiplier does not consider the perceptual characteristic of the input video and thus is not effective for perceptual video coding. To address this problem, an efficient perceptual adaptive Lagrangian multiplier for HEVC is proposed. Based on the human visual system (HVS) observation that the region with less perceptual sensitivity can tolerate more distortion, the Lagrangian multiplier is adaptively adjusted for each coding tree unit (CTU) based on its perceptual sensitivity so that the perceptual quality of the reconstructed video can be improved. The above-mentioned perceptual sensitivity for each CTU is measured according to two perceptual features—spatial energy ratio and temporal motion activity. Experimental results have shown that the proposed method is able to significantly improve the perceptual RD performance, compared with the original HEVC. Huanqiang Zeng, King Ngi Ngan, Miaohui Wang |
PCS | 2 |
| 2013 | Visual quality metric for perceptual video codingabstractThe visual quality assessment (VQA) becomes prevailing in the studies of image and video coding. It assesses the quality of image or video more accurately than mean square error (MSE) with respect to the human visual system (HVS). Toward perceptual video coding, MSE is weighted spatially and temporally to simulate the HVS response to visual signal in this paper. Firstly, the image content is depicted by edge strength to compose spatial weighting factors. Secondly, the motion strength calculated from motion vector of each block gives temporal weighting factors. Thirdly, the motion trajectory based saliency map for video signal is integrated as another weighting factor of MSE. The proposed VQM not only efficiently model HVS but also relate to quantization parameter (QP) capable of guiding perceptual video coding. A perceptual rate distortion optimization (RDO) is established on the proposed VQM. The experimental results indicate that the proposed VQM is consistent well with HVS. In addition, the better rate-distortion efficiency and accurate bit rate control can be achieved by the proposed visual quality control algorithm. Long Xu 0001, Lin Ma 0002, King Ngi Ngan, Weisi Lin, Ying Weng |
VCIP | 3 |
| 2013 | Saliency detection using joint spatial-color constraint and multi-scale segmentation
Linfeng Xu 0001, Hongliang Li 0001, Liaoyuan Zeng, King Ngi Ngan |
J. Vis. Commun. Image Represent. | 4 |
| 2013 | Anaglyph image generation by matching color appearance attributes
Songnan Li, Lin Ma 0002, King Ngi Ngan |
Signal Process. Image Commun. | 3 |
| 2013 | Reduced-reference image quality assessment in reorganized DCT domain
Lin Ma 0002, Songnan Li, King Ngi Ngan |
Signal Process. Image Commun. | 3 |
| 2013 | An efficient framework for image/video inpainting
Miaohui Wang, Bo Yan 0001, King Ngi Ngan |
Signal Process. Image Commun. | 3 |
| 2013 | Consistent Visual Quality Control in Video CodingabstractVisual quality consistency is one of the most important issues in video quality assessment. When people view a sequential video, they may have an unpleasant perceptual experience if the video has an inconsistent visual quality even though the average visual quality of the video is not compromised. Thus, consistent visual quality control is mostly expected in general video encoding with limited channel bandwidth and buffer resources. However, there still has not been enough study on such an issue. In this paper, a new objective visual quality metric (VQM) is proposed first, which can easily be incorporated into video coding for guiding video coding. Second, a VQM-based window model is proposed to handle the tradeoff between visual quality consistency and buffer constraint in video coding. Third, a window-level rate control algorithm is developed to accomplish visual quality control based on the above two proposals. Finally, experimental results prove that consistent visual quality, high rate-distortion efficiency, accurate bit control, and compliant buffer constraint can be achieved by the proposed rate control algorithm. Long Xu 0001, Songnan Li, King Ngi Ngan, Lin Ma 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2013 | Image Cosegmentation by Incorporating Color Reward Strategy and Active Contour ModelabstractThe design of robust and efficient cosegmentation algorithms is challenging because of the variety and complexity of the objects and images. In this paper, we propose a new cosegmentation model by incorporating a color reward strategy and an active contour model. A new energy function corresponding to the curve is first generated with two considerations: the foreground similarity between the image pairs and the background consistency in each of the image pair. Furthermore, a new foreground similarity measurement based on the rewarding strategy is proposed. Then, we minimize the energy function value via a mutual procedure which uses dynamic priors to mutually evolve the curves. The proposed method is evaluated on many images from commonly used databases. The experimental results demonstrate that the proposed model can efficiently segment the common objects from the image pairs with generally lower error rate than many existing and conventional cosegmentation methods. Fanman Meng, Hongliang Li 0001, Guanghui Liu 0001, King Ngi Ngan |
IEEE Trans. Cybern. | 4 |
| 2013 | Global Propagation of Affine Invariant Features for Robust MatchingabstractLocal invariant features have been successfully used in image matching to cope with viewpoint change, partial occlusion, and clutters. However, when these factors become too strong, there will be a lot of mismatches due to the limited repeatability and discriminative power of features. In this paper, we present an efficient approach to remove the false matches and propagate the correct ones for the affine invariant features which represent the state-of-the-art local invariance. First, a pair-wise affine consistency measure is proposed to evaluate the consensus of the matches of affine invariant regions. The measure takes into account both the keypoint location and the region shape, size, and orientation. Based on this measure, a geometric filter is then presented which can efficiently remove the outliers from the initial matches, and is robust to severe clutters and non-rigid deformation. To increase the correct matches, we propose a global match refinement and propagation method that simultaneously finds a optimal group of local affine transforms to relate the features in two images. The global method is capable of producing a quasi-dense set of matches even for the weakly textured surfaces that suffer strong rigid transformation or non-rigid deformation. The strong capability of the proposed method in dealing with significant viewpoint change, non-rigid deformation, and low-texture objects is demonstrated in experiments of image matching, object recognition, and image based rendering. Chunhui Cui, King Ngi Ngan |
IEEE Trans. Image Process. | 2 |
| 2013 | Feature Adaptive Co-Segmentation by Complexity AwarenessabstractIn this paper, we propose a novel feature adaptive co-segmentation method that can learn adaptive features of different image groups for accurate common objects segmentation. We also propose image complexity awareness for adaptive feature learning. In the proposed method, the original images are first ranked according to the image complexities that are measured by superpixel changing cue and object detection cue. Then, the unsupervised segments of the simple images are used to learn the adaptive features, which are achieved using an expectation-minimization algorithm combining l 1-regularized least squares optimization with the consideration of the confidence of the simple image segmentation accuracies and the fitness of the learned model. The error rate of the final co-segmentation is tested by the experiments on different image groups and verified to be lower than the existing state-of-the-art co-segmentation methods. Fanman Meng, Hongliang Li 0001, King Ngi Ngan, Liaoyuan Zeng, Qingbo Wu 0001 |
IEEE Trans. Image Process. | 3 |
| 2013 | Additive Log-Logistic Model for Networked Video Quality AssessmentabstractModeling subjective opinions on visual quality is a challenging problem, which closely relates to many factors of the human perception. In this paper, the additive log-logistic model (ALM) is proposed to formulate such a multidimensional nonlinear problem. The log-logistic model has flexible monotonic or nonmonotonic partial derivatives and thus is suitable to model various uni-type impairments. The proposed ALM metric adds the distortions due to each type of impairment in a log-logistic transformed space of subjective opinions. The features can be evaluated and selected by classic statistical inference, and the model parameters can be easily estimated. Cross validations on five Telecommunication Standardization Sector of International Telecommunication Union (ITU-T) subjectively-rated databases confirm that: 1) based on the same features, the ALM outperforms the support vector regression and the logistic model in quality prediction and, 2) the resultant no-reference quality met-ric based on impairment-relevant video parameters achieves high correlation with a total of 27 216 subjective opinions on 1134 video clips, even compared with existing full-reference quality metrics based on pixel differences. The ALM metric wins the model competition of the ITU-T Study Group 12 (where the validation databases are independent with the training databases) and thus is being put forth into ITU-T Recommendation P.1202.2 for the consent of ITU-T. Fan Zhang 0093, Weisi Lin, Zhibo Chen 0001, King Ngi Ngan |
IEEE Trans. Image Process. | 4 |
| 2013 | Co-Salient Object Detection From Multiple ImagesabstractIn this paper, we propose a novel method to discover co-salient objects from a group of images, which is modeled as a linear fusion of an intra-image saliency (IaIS) map and an inter-image saliency (IrIS) map. The first term is to measure the salient objects from each image using multiscale segmentation voting. The second term is designed to detect the co-salient objects from a group of images. To compute the IrIS map, we perform the pairwise similarity ranking based on an image pyramid representation. A minimum spanning tree is then constructed to determine the image matching order. For each region in an image, we design three types of visual descriptors, which are extracted from the local appearance, e.g., color, color co-occurrence and shape properties. The final region matching problem between the images is formulated as an assignment problem that can be optimized by linear programming. Experimental evaluation on a number of images demonstrates the good performance of the proposed method on co-salient object detection. Hongliang Li 0001, Fanman Meng, King Ngi Ngan |
IEEE Trans. Multim. | 3 |
| 2013 | From Logo to Object SegmentationabstractThis paper proposes a method to segment object from the web images using logo detection. The method consists of three steps. In the first step, the logos are located from the original images by SIFT matching. Based on the logo location and the object shape model, the second step extracts the object boundary from the image. In the third step, we use the object boundary to model the object appearance, which is then used in the MRF based segmentation method to finally achieve the object segmentation. The key of our method is the object boundary extraction, which is achieved by searching a variation of the shape model that best fits the local edge of the image. Affine transform is used to consider the variations among the objects. Meanwhile, the Nelder-Mead simplex method with a simple initial rough search is used to run the boundary search. To verify the proposed method, we collect a LogoSeg dataset from the web such as Flickr and Google. The MOMI dataset is also used for the verification. The experimental results demonstrate that the proposed logo detection based segmentation method can improve the performance of the object segmentation. Fanman Meng, Hongliang Li 0001, Guanghui Liu 0001, King Ngi Ngan |
IEEE Trans. Multim. | 4 |
| 2012 | Overlapping Local Phase Feature (OLPF) for Robust Face Recognition in Surveillance
Qiang Liu 0015, King Ngi Ngan |
ACIVS | 2 |
| 2012 | Study of subjective and objective quality assessment of retargeted imagesabstractThis paper presents the result of a recent large-scale subjective study of image retargeting quality on a collection of images generated by several representative image retargeting methods. Owning to many approaches to image retargeting that are developed, there is a need for a diverse independent public database of the retargeted images and the corresponding subjective scores that is freely available. We build an image retargeting quality database, in which 171 retargeted images (obtained from 57 natural source images of different contents) were generated by several representative image retargeting methods. The perceptual quality of each image is evaluated by at least 30 human subjects and the mean opinion scores (MOS) were recorded. Furthermore, several publicly available quality metrics for the retargeted images are evaluated on the built database. The database is made available [1] to the research community in order to further research on the perceptual quality assessment of the retargeted images. Lin Ma 0002, Weisi Lin, Chenwei Deng, King Ngi Ngan |
ISCAS | 4 |
| 2012 | Depth estimation and view synthesis for narrow-baseline videoabstractIn this paper, we propose the depth estimation and view synthesis approaches for the scene rendering of narrow-baseline videos. Depth estimation is performed by global energy minimization using graph cut technique. A novel smoothness energy term is proposed to enforce powerful smoothness of the estimated depth. In view synthesis, smart hole filling is proposed to efficiently remove the burr artifacts along the object boundary. Experimental results show that compared with the MPEG reference software, our method can generate more accurate depth maps and achieve much better objective (PSNR) and subjective quality in view synthesis. Qian Zhang 0001, Chunhui Cui, King Ngi Ngan, Yu Liu 0041 |
ISCAS | 3 |
| 2012 | Spatial-temporal decorrelation for image/video codingabstractModern image/video compression techniques greatly help to store and transmit digital images and video data. Discrete wavelet transform is used in JPEG2000 because of its scalability and tolerable degradation. In H.264/AVC, predictive coding is employed to remove spatial redundancy before discrete cosine transform. In this paper, we propose a new encoder structure combining both advantages from JPEG2000 [1] and H.264/AVC [2], which firstly utilizes the proposed spatial transform to decompose a single image into several sub-images and then employs motion compensation to convert the conventional spatial decorrelation into temporal decorrelation. Experimental results show that our proposed method outperforms the state-of-the-art image coding algorithms and achieves better rate-distortion performance. Miaohui Wang, King Ngi Ngan, Long Xu 0001 |
PCS | 2 |
| 2012 | Video content dependent directional transform for intra frame codingabstractThe mode-dependent directional transform (MDDT) employed Karhunen-Loève Transform (KLT) for compressing directional residue signal of intra prediction along its direction. The transform bases were derived from the singular value decomposition (SVD) of residue signals coming from all kinds of video sequences, which were expected to be efficient for most of video sequences. However, the advantage of KLT comes from the concept of a “signal content dependent transform”. MDDT and its variants failed to exploit such a concept, so they did not fully exploit the efficiency of KLT. In this paper, a video content feature is firstly defined as the histogram of the residue produced by intra prediction. Secondly, one KLT basis is computed for each feature of each mode from off-line experiments. Thus, multiple KLT bases identified by their features are provided to each mode instead of only one basis in MDDT. One of them is selected during encoding process by matching the feature of signal being processed to the predefined features. The experiments show that the average improvement of 0.17dB PSNR and 2.23% bits saving can be achieved by the proposed video content dependent directional transform (CDDT) comparing to the state-of-the-art MDDT. Long Xu 0001, King Ngi Ngan, Miaohui Wang |
PCS | 2 |
| 2012 | Global salient information maximization for saliency detection
Hongliang Li 0001, Guanghui Liu 0001, King Ngi Ngan |
Signal Process. Image Commun. | 4 |
| 2012 | Two-Layer Directional Transform for High Performance Video CodingabstractThis paper presents a directional transform scheme for coding interprediction errors in block-based hybrid video coding. It proposes a two-layer transform structure, where the first layer uses discrete wavelet transform to compact the residue energy to the LL band and then the second layer uses 2-D nonseparable directional transforms to deal with the arbitrary edge directions in the four subbands. By doing this, the edges in a macroblock are efficiently compacted to a few coefficients and at the same time the overhead used to indicate the transform directions is affordable. Experimental results show that the proposed scheme provides peak signal-to-noise ratio gain up to 0.46 dB, compared with H.264/AVC High Profile. Jie Dong 0001, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2012 | Full-Reference Video Quality Assessment by Decoupling Detail Losses and Additive ImpairmentsabstractVideo quality assessment plays a fundamental role in video processing and communication applications. In this paper, we study the use of motion information and temporal human visual system (HVS) characteristics for objective video quality assessment. In our previous work, two types of spatial distortions, i.e., detail losses and additive impairments, are decoupled and evaluated separately for spatial quality assessment. The detail losses refer to the loss of useful visual information that will affect the content visibility, and the additive impairments represent the redundant visual information in the test image, such as the blocking or ringing artifacts caused by data compression and so on. In this paper, a novel full-reference video quality metric is developed, which conceptually comprises the following processing steps: 1) decoupling detail losses and additive impairments within each frame for spatial distortion measure; 2) analyzing the video motion and using the HVS characteristics to simulate the human perception of the spatial distortions; and 3) taking into account cognitive human behaviors to integrate frame-level quality scores into sequence-level quality score. Distinguished from most studies in the literature, the proposed method comprehensively investigates the use of motion information in the simulation of HVS processing, e.g., to model the eye movement, to predict the spatio-temporal HVS contrast sensitivity, to implement the temporal masking effect, and so on. Furthermore, we also prove the effectiveness of decoupling detail losses and additive impairments for video quality assessment. The proposed method is tested on two subjective quality video databases, LIVE and IVP, and demonstrates the state-of-the-art performance in matching subjective ratings. Songnan Li, Lin Ma 0002, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2012 | Reduced-Reference Video Quality Assessment of Compressed Video SequencesabstractIn this paper, a novel reduced-reference (RR) video quality assessment (VQA) is proposed by exploiting the spatial information loss and the temporal statistical characteristics of the interframe histogram. From the spatial perspective, an energy variation descriptor (EVD) is proposed to measure the energy change of each individual encoded frame, which results from the quantization process. Besides depicting the energy change, EVD can further simulate the texture masking property of the human visual system (HVS). From the temporal perspective, the generalized Gaussian density (GGD) function is employed to capture the natural statistics of the interframe histogram distribution. The city-block distance (CBD) is used to calculate the histogram distance between the original video sequence and the encoded one. For simplicity, the difference image between adjacent frames is employed to characterize the temporal interframe relationship. By combining the spatial EVD together with the temporal CBD, an efficient RR VQA is developed. Evaluation on the subjective quality video database demonstrates that the proposed method outperforms the representative RR video quality metric and the full-reference VQAs, such as peak signal-to-noise ratio and structure similarity index in matching subjective ratings. This means that the proposed metric is more consistent with the HVS perception. Furthermore, as only a small number of RR features are extracted for representing the original video sequence (each frame requires only one parameter for describing EVD and three parameters for recording GGD), the RR features can be embedded into the video sequences or transmitted through the ancillary data channel, which can be used in the video quality monitoring system. Lin Ma 0002, Songnan Li, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2012 | Unsupervised Salient Object Segmentation Based on Kernel Density Estimation and Two-Phase Graph CutabstractIn this paper, we propose an unsupervised salient object segmentation approach based on kernel density estimation (KDE) and two-phase graph cut. A set of KDE models are first constructed based on the pre-segmentation result of the input image, and then for each pixel, a set of likelihoods to fit all KDE models are calculated accordingly. The color saliency and spatial saliency of each KDE model are then evaluated based on its color distinctiveness and spatial distribution, and the pixel-wise saliency map is generated by integrating likelihood measures of pixels and saliency measures of KDE models. In the first phase of salient object segmentation, the saliency map based graph cut is exploited to obtain an initial segmentation result. In the second phase, the segmentation is further refined based on an iterative seed adjustment method, which efficiently utilizes the information of minimum cut generated using the KDE model based graph cut, and exploits a balancing weight update scheme for convergence of segmentation refinement. Experimental results on a dataset containing 1000 test images with ground truths demonstrate the better segmentation performance of our approach. Zhi Liu 0003, Liquan Shen, Yinzhu Xue, King Ngi Ngan, Zhaoyang Zhang 0002 |
IEEE Trans. Multim. | 5 |
| 2012 | Object Co-Segmentation Based on Shortest Path Algorithm and Saliency ModelabstractSegmenting common objects that have variations in color, texture and shape is a challenging problem. In this paper, we propose a new model that efficiently segments common objects from multiple images. We first segment each original image into a number of local regions. Then, we construct a digraph based on local region similarities and saliency maps. Finally, we formulate the co-segmentation problem as the shortest path problem, and we use the dynamic programming method to solve the problem. The experimental results demonstrate that the proposed model can efficiently segment the common objects from a group of images with generally lower error rate than many existing and conventional co-segmentation methods. Fanman Meng, Hongliang Li 0001, Guanghui Liu 0001, King Ngi Ngan |
IEEE Trans. Multim. | 4 |
| 2011 | Motion trajectory based visual saliency for video quality assessmentabstractIn this paper, we propose a novel visual saliency detection method for video sequences by considering the object motion trajectories. Firstly, each frame of the video sequence is described in a new Quaternion Representation (QR), which comprises the spatial image content and the temporal motion characteristics. Based on the QR, Quaternion Fourier Transform (QFT) is employed to construct the visual salien-cy of the video sequence. Finally, the detected visual salien-cy map is incorporated with several video quality metrics. Compared with other visual saliency models, the proposed method can improve the performances of video quality metrics. It further confirms that the proposed visual saliency model can accurately depict the Human Vision System (HVS) properties. Lin Ma 0002, Songnan Li, King Ngi Ngan |
ICIP | 3 |
| 2011 | Adaptive pre-interpolation filter for motion-compensated predictionabstractThe proposed interpolation filter comprises two concatenating fiiters, adaptive pre-interpolation filter (APIF) and the normative interpolation filter in H.264/AVC. The former is applied only to the integer pixels in the reference frames; the latter generates all the sub-position samples, supported by the output of APIF. The convolution of APIF and the standard filter minimizes the motion prediction error on a frame basis. APIF preserves the merits of the adaptive interpolation filter (AIF) and the adaptive loop filter (ALF) in the key technical area (KTA) software and overcomes their drawbacks. The experimental results show that APIF has comparable or even better performance compared with the joint use of AIF and ALF. Jie Dong 0001, King Ngi Ngan |
ISCAS | 2 |
| 2011 | Perceptual image compression via adaptive block- based super-resolution directed down-samplingabstractIn this paper, we propose a novel perceptual image coding scheme via adaptive block-based super-resolution directed down-sampling. At the encoder side, for each macroblock of a given image, Rate Distortion Optimization (RDO) determines whether it is encoded at the original or down-sampled resolution. The down-sampling process is directed by super-resolution, which generates the down-sampled block by minimizing the reconstruction errors between the original macroblock and the one restored by the corresponding super-resolution method. At the decoder side, in order to reduce the complexity, the super-resolution method reconstructs the full-resolution macroblock in the DCT domain together with the inverse DCT. Experimental results have demonstrated that the proposed method can produce higher quality images in terms of both PSNR and visual quality compared with the existing methods. Lin Ma 0002, Songnan Li, King Ngi Ngan |
ISCAS | 3 |
| 2011 | Adaptive pre-interpolation filter for high efficiency video coding
Jie Dong 0001, King Ngi Ngan |
J. Vis. Commun. Image Represent. | 2 |
| 2011 | Automatic body segmentation with graph cut and self-adaptive initialization level set (SAILS)
Qiang Liu 0015, Hongliang Li 0001, King Ngi Ngan |
J. Vis. Commun. Image Represent. | 3 |
| 2011 | Adaptive Block-size Transform based Just-Noticeable Difference model for images/videos
Lin Ma 0002, King Ngi Ngan, Fan Zhang 0093, Songnan Li |
Signal Process. Image Commun. | 2 |
| 2011 | Learning to Extract Focused Objects From Low DOF ImagesabstractThis paper proposes an approach to extract focused objects (i.e., attention objects) from low depth-of-field images. To recognize the focused object, we decompose the image into multiple regions, which are described by using three types of visual descriptors. Each descriptor is extracted from a representation of some aspects of local appearance, e.g., a spatially localized texture, color, or geometrical property. Therefore, the focus detection of a region can be achieved by the classification of extracted visual descriptors based on a binary classifier. We employ a boosting algorithm to learn the classifier with a cascade of decision structure. Given a test image, initial segmentation can be achieved using obtained classification results. Finally, we apply a post-processing technique to improve the results by incorporating region grouping and pixel-level segmentation. Experimental evaluation on a number of images demonstrates the performance advantages of the proposed method, when compared with state-of-the-art methods. Hongliang Li 0001, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | Scale- and Affine-Invariant Fan FeatureabstractMost existing feature detectors assume no surface discontinuity within the keypoints' support regions and, hence, have little chance to match the keypoints located on or near the surface boundaries. These keypoints, though not many, are salient and representative. In this paper, we show that they can be successfully matched by using the proposed scale- and affine-invariant Fan features. Specifically, the image neighborhood of a keypoint is depicted by multiple fan subregions, namely Fan features, to provide robustness to surface discontinuity and background change. These Fan features are made scale-invariant by using the automatic scale selection method based on the Fan Laplacian of Gaussian (FLOG). Affine invariance is further introduced to the Fan features based on the affine shape diagnosis of the mirror-predicted surface patch. The Fan features are then described by Fan-SIFT, which is an extension of the famous scale-invariant feature transform (SIFT) descriptor. Experimental results of quantitative comparisons show that the proposed Fan feature has good repeatability that is comparable to the state-of-the-art features for general structured scenes. Moreover, by using Fan features, we can successfully match image structures near surface discontinuities despite significant scale, viewpoint, and background changes. These structures are complementary to those found by the traditional methods and are especially useful for describing weakly textured scenes, which is demonstrated in our experiments on image matching and object rendering. Chunhui Cui, King Ngi Ngan |
IEEE Trans. Image Process. | 2 |
| 2011 | Composite Model-Based DC Dithering for Suppressing Contour Artifacts in Decompressed VideoabstractBecause of the outstanding contribution in improving compression efficiency, block-based quantization has been widely accepted in state-of-the-art image/video coding standards. However, false contour artifacts are introduced, which result in reducing the fidelity of the decoded image/video especially in terms of subjective quality. In this paper, a block-based decontouring method is proposed to reduce the false contour artifacts in the decoded image/video by automatically dithering its direct current (DC) value according to a composite model established between gradient smoothness and block-edge smoothness. Feature points on the model with the corresponding criteria in suppressing contour artifacts are compared to show a good consistency between the model and the actual processing effects. Discrete cosine transform (DCT)-based block level contour artifacts detection mechanism ensures the blocks within the texture region are not affected by the DC dithering. Both the implementation method and the algorithm complexity are analyzed to present the feasibility in integrating the proposed method into an existing video decoder on an embedded platform or system-on-chip (SoC). Experimental results demonstrate the effectiveness of the proposed method both in terms of subjective quality and processing complexity in comparison with the previous methods. Xin Jin 0002, Satoshi Goto, King Ngi Ngan |
IEEE Trans. Image Process. | 3 |
| 2011 | A Co-Saliency Model of Image PairsabstractIn this paper, we introduce a method to detect co-saliency from an image pair that may have some objects in common. The co-saliency is modeled as a linear combination of the single-image saliency map (SISM) and the multi-image saliency map (MISM). The first term is designed to describe the local attention, which is computed by using three saliency detection techniques available in literature. To compute the MISM, a co-multilayer graph is constructed by dividing the image pair into a spatial pyramid representation. Each node in the graph is described by two types of visual descriptors, which are extracted from a representation of some aspects of local appearance, e.g., color and texture properties. In order to evaluate the similarity between two nodes, we employ a normalized single-pair SimRank algorithm to compute the similarity score. Experimental evaluation on a number of image pairs demonstrates the good performance of the proposed method on the co-saliency detection task. Hongliang Li 0001, King Ngi Ngan |
IEEE Trans. Image Process. | 2 |
| 2011 | Spread Spectrum Image Watermarking Based on Perceptual Quality MetricabstractEfficient image watermarking calls for full exploitation of the perceptual distortion constraint. Second-order statistics of visual stimuli are regarded as critical features for perception. This paper proposes a second-order statistics (SOS)-based image quality metric, which considers the texture masking effect and the contrast sensitivity in Karhunen-Loève transform domain. Compared with the state-of-the-art metrics, the quality prediction by SOS better correlates with several subjectively rated image databases, in which the images are impaired by the typical coding and watermarking artifacts. With the explicit metric definition, spread spectrum watermarking is posed as an optimization problem: we search for a watermark to minimize the distortion of the watermarked image and to maximize the correlation between the watermark pattern and the spread spectrum carrier. The simple metric guarantees the optimal watermark a closed-form solution and a fast implementation. The experiments show that the proposed watermarking scheme can take full advantage of the distortion constraint and improve the robustness in return. Fan Zhang 0093, Wenyu Liu 0001, Weisi Lin, King Ngi Ngan |
IEEE Trans. Image Process. | 4 |
| 2011 | Segmentation and Tracking Multiple Objects Under Occlusion From Multiview VideoabstractIn this paper, we present a multiview approach to segment the foreground objects consisting of a group of people into individual human objects and track them across the video sequence. Depth and occlusion information recovered from multiple views of the scene is integrated into the object detection, segmentation, and tracking processes. Adaptive background penalty with occlusion reasoning is proposed to separate the foreground regions from the background in the initial frame. Multiple cues are employed to segment individual human objects from the group. To propagate the segmentation through video, each object region is independently tracked by motion compensation and uncertainty refinement, and the motion occlusion is tackled as layer transition. The experimental results implemented on both our sequences and other's sequence have demonstrated the algorithm's efficiency in terms of subjective performance. Objective comparison with a state-of-the-art algorithm validates the superior performance of our method quantitatively. Qian Zhang 0001, King Ngi Ngan |
IEEE Trans. Image Process. | 2 |
| 2011 | Guided Face Cartoon SynthesisabstractIn this paper, we propose a new method, called guided synthesis, to synthesize a face cartoon from a face photo. The guided synthesis is defined as a local linear model, which generates a cartoon image by incorporating the content of guidance images taken from the training set. Our synthesis operation is achieved based on four weight functions. The first is a photo-photo weight that aims to measure the similarity between an input photo patch and a training photo patch. The second is defined as a photo-cartoon weight, which is used to compute the likelihood by computing the similarity between a cartoon patch and an input photo patch. The third weight is defined in the synthesized photos, which is to set a smoothness constraint between neighboring synthesized patches. The final weight is designed to evaluate the similarity of a synthesized patch to an input patch based on the spatial distance. Experimental evaluation on a number of face photos demonstrates the good performance of the proposed method on the face cartoon synthesis. Hongliang Li 0001, Guanghui Liu 0001, King Ngi Ngan |
IEEE Trans. Multim. | 3 |
| 2011 | Image Quality Assessment by Separately Evaluating Detail Losses and Additive ImpairmentsabstractIn the research field of image processing, mean squared error (MSE) and peak signal-to-noise ratio (PSNR) are extensively adopted as the objective visual quality metrics, mainly because of their simplicity for calculation and optimization. However, it has been well recognized that these pixel-based difference measures correlate poorly with the human perception. Inspired by existing works, in this paper we propose a novel algorithm which separately evaluates detail losses and additive impairments for image quality assessment. The detail loss refers to the loss of useful visual information which affects the content visibility, and the additive impairment represents the redundant visual information whose appearance in the test image will distract viewer's attention from the useful contents causing unpleasant viewing experience. To separate detail losses and additive impairments, a wavelet-domain decoupling algorithm is developed which can be used for a host of distortion types. Two HVS characteristics, i.e., the contrast sensitivity function and the contrast masking effect, are taken into account to approximate the HVS sensitivities. We propose two simple quality measures to correlate detail losses and additive impairments with visual quality, respectively. Based on the findings inthat observers judge low-quality images in terms of the ability to interpret the content, the outputs of the two quality measures are adaptively combined to yield the overall quality index. By conducting experiments based on five subjectively-rated image databases, we demonstrate that the proposed metric has a better or similar performance in matching subjective ratings when compared with the state-of-the-art image quality metrics. Songnan Li, Fan Zhang 0093, Lin Ma 0002, King Ngi Ngan |
IEEE Trans. Multim. | 4 |
| 2011 | Reduced-Reference Image Quality Assessment Using Reorganized DCT-Based Image RepresentationabstractIn this paper, a novel reduced-reference (RR) image quality assessment (IQA) is proposed by statistical modeling of the discrete cosine transform (DCT) coefficient distributions. In order to reduce the RR data rates and further exploit the identical nature of the coefficient distributions between adjacent DCT subbands, the DCT coefficients are reorganized into a three-level coefficient tree. Subsequently, generalized Gaussian density (GGD) is employed to model the coefficient distribution of each reorganized DCT subband. The city-block distance is employed to measure the difference between the two images. Experimental results demonstrate that only a small number of RR features is sufficient for representing the image perceptual quality. The proposed method outperforms the RR WNISM and even the full-reference (FR) quality metric PSNR. Lin Ma 0002, Songnan Li, Fan Zhang 0093, King Ngi Ngan |
IEEE Trans. Multim. | 4 |
| 2011 | Practical Image Quality Metric Applied to Image CodingabstractPerceptual image coding requires an effective image quality metric, yet most of the existing metrics are complex and can hardly guide the compression effectively. This paper proposes a practical full-reference metric with consideration of the texture masking effect and contrast sensitivity function. The metric is capable of evaluating typical image impairments in real-world applications and can achieve the comparable performance as the state-of-the-art metrics on the publicly available subjectively-rated image databases. Due to its simplicity, the metric is embedded into JPEG image coding to ensure a better perceptual rate-distortion performance. Fan Zhang 0093, Lin Ma 0002, Songnan Li, King Ngi Ngan |
IEEE Trans. Multim. | 4 |
| 2010 | Dense Stereo Matching from Separated Views of Wide-Baseline Images
Qian Zhang 0001, King Ngi Ngan |
ACIVS (1) | 2 |
| 2010 | A novel geometric filter for affine invariant featuresabstractInvariant local image features have proven to be very successful in computer vision tasks involving partial occlusion and various image deformations. Even though the image features can be extracted in a high repeatability, their local appearance alone usually does not bring enough discriminative power to support a reliable matching, resulting in a relatively high number of outliers in the correspondence set. To reject these mismatches, various geometric filters have been proposed for different image features. In this paper, we present a novel and efficient geometric filter for the state-of-the-art affine invariant features. The proposed method detects the mismatches by examining the consistency of local affine geometry between neighboring matches of affine invariant features. Experimental results show that the proposed geometric filter not only achieves a higher inlier ratio than the standard Hough clustering, but also presents superior robustness to severe clutters, significant viewpoint changes and non-rigid deformation. Chunhui Cui, King Ngi Ngan |
ICIP | 2 |
| 2010 | Video Quality Assessment based on Adaptive Block-size Transform Just-Noticeable Difference modelabstractIn this paper, we propose a full reference Video Quality Assessment (VQA) algorithm based on the Adaptive Block-size Transform Just-Noticeable Difference (ABT-JND) model. Firstly, ABT-JND is introduced for its efficiency of modeling the Human Vision System (HVS) characteristics. Based on the ABT-JND model, the full reference VQA is developed, by capturing HVS responses of spatio-temporal distortions over different block-size transforms. Experimental results have demonstrated that the proposed VQA outperforms other VQA methods, while slightly poorer than MOVIE. However, it maintains a very simple formulation. Since the proposed VQA performs on transform domain, it could be easily applied on many related applications, such as video compression, watermarking, and so on. Lin Ma 0002, Fan Zhang 0093, Songnan Li, King Ngi Ngan |
ICIP | 4 |
| 2010 | Perceptual video coding: Challenges and approachesabstractInvestigation on the human perception can play an important role in video signal processing. Recently, there has been great interest in incorporating the human perception in video coding systems to enhance the perceptual quality of the represented visual signal. However, the limited understanding of the human visual system and high complexity of computational models of human visual system make it a challenging task. Furthermore, the hybrid video coding structure brings difficulties to integrate computational models with coding components to fulfill the requirements. In this paper, we review the physiological characteristics of human perception and address the most relevant aspects to video coding applications. Moreover, we discuss the computational models and metrics which guide the design and implementation of the video coding system, as well as the recent advances in perceptual video coding. To introduce this overview with the latest technologies and most promising directions in perceptual video coding, we focus on three key areas. Specifically, we cover 1) visual attention and sensitivity modeling, with which we concentrate on the computational models of bottom-up and top-down attention, contrast sensitivity functions and masking effects, and fovea based manipulations; 2) perceptual quality optimization for constrained video coding, with which we discuss how to achieve maximum perceptual quality whilst satisfying various constraints; and 3) the impact of the human perception on advanced video applications, including emerging immersive multimedia services, and compression of high dynamic range video content and 3D video. For each aspect, we discuss the major challenges, highlight significant approaches, and outline future research directions. Zhenzhong Chen 0001, Weisi Lin, King Ngi Ngan |
ICME | 3 |
| 2010 | Learn to segment attention object from low DoF imageabstractIn this paper, a novel segmentation algorithm is proposed to extract attention object (i.e., focus object) from Low depth of field image. In order to recognize the focus object, we first decompose the image into multiple segments that are described by visual words. Each visual word is computed from a filter bank to represent the high frequency components. The boosting method is then used to generate a strong classifier for each training image. Given a test image, we employ the voting algorithm to achieve the attention decision according to obtained strong classifiers. To extract focus objects from the test image, two-level segmentation method is proposed, which includes region and pixel levels segmentation. Experimental evaluation on test images shows that the proposed method is capable of segmenting the attention object quite effectively. Hongliang Li 0001, Guanghui Liu 0001, King Ngi Ngan |
ISCAS | 3 |
| 2010 | Subtractive impairment, additive impairment and image visual qualityabstractIn this paper, we propose an engineering-based image quality metric which distinguishes subtractive impairment from additive impairment. Since the amount of subtractive impairment is up-bounded by the total details within the reference image but the same limitation can't be applied to additive impairment, intuitively visual quality degradation due to the two types of impairments should be measured differently. In the proposed metric, subtractive and additive impairments are separated and represented in the wavelet domain, and their influences to image visual quality is measured by different equations. We tested the proposed metric on five subjectively-rated databases and proved its effectiveness in objective image quality assessment. Songnan Li, King Ngi Ngan |
ISCAS | 2 |
| 2010 | Adaptive block-size transform based just-noticeable difference profile for videosabstractIn this paper, we propose a novel adaptive block-size transform (ABT) based just-noticeable difference (JND) model for videos. Firstly, the ABT-based spatial JND profile is extended to spatial-temporal JND model for videos by considering temporal contrast sensitivity function (TCSF), eye movement, and the motion information of the objects in video sequence. Furthermore, a metric named motion characteristics distance (MCD) is proposed to depict the motion characteristics similarity between a macroblock and its corresponding sub-blocks. Based on the proposed MCD and the obtained spatial image content information, a novel balanced strategy is proposed to determine which transform size is employed to generate the resulting JND model. Experimental results have demonstrated that our proposed scheme could tolerate more distortions while preserving better perceptual quality than other JND profiles, which means that the proposed model consists well with human vision system (HVS). Moreover, for the balanced strategy, experiments have shown that temporal motion characteristics accord very well with the spatial image content information, which has demonstrated the efficiency of our proposed balanced strategy. Lin Ma 0002, King Ngi Ngan |
ISCAS | 2 |
| 2010 | Temporal inconsistency measure for video quality assessmentabstractVisual quality assessment plays a crucial role in many vision-related signal processing applications. In the literature, more efforts have been spent on spatial visual quality measure. Although a large number of video quality metrics have been proposed, the methods to use temporal information for quality assessment are less diversified. In this paper, we propose a novel method to measure the temporal impairments. The proposed method can be incorporated into any image quality metric to extend it into a video quality metric. Moreover, it is easy to apply the proposed method in video coding system to incorporate with MSE for rate-distortion optimization. Songnan Li, Lin Ma 0002, Fan Zhang 0093, King Ngi Ngan |
PCS | 4 |
| 2010 | Limitation and challenges of image quality measurementabstractSubjectively-rated image databases have become increasingly popular in the evaluation of image quality measurement algorithms. Several groups recently have improved their metrics' performance in matching these databases, using particular HVS (human visual system) properties or image statistical models. However, it is difficult to know whether these improvements are due to progress towards mimicking the perceptual properties, or are due to matching some characteristics of the databases. This paper demonstrates an inherent limitation in using such databases, showing that our very simple metric, built on the contrast masking effect, is able to perform as good as many state-of-the-art metrics. It is also argued that existent databases neither contain enough images with particularly biased distortions to test the significance of single HVS property, nor cover diverse distortion types to reflect the requirement of emerging applications. Fan Zhang 0093, Songnan Li, Lin Ma 0002, King Ngi Ngan |
VCIP | 4 |
| 2010 | Multi-view video based multiple objects segmentation using graph cut and spatiotemporal projections
Qian Zhang 0001, King Ngi Ngan |
J. Vis. Commun. Image Represent. | 2 |
| 2010 | Plane-based external camera calibration with accuracy measured by relative deflection angle
Chunhui Cui, King Ngi Ngan |
Signal Process. Image Commun. | 2 |
| 2010 | Visual Horizontal Effect for Image Quality AssessmentabstractIn this paper, an image quality metric is proposed by modeling the visual Horizontal Effect (HE) and saliency property over structural distortions. Specifically, Structrue SIMilarity (SSIM) is firstly performed to obtain the structural distortion map. Subsequently, the obtained distortion map is refined by the visual HE model, which depicts visual sensitivities of oriented stimuli over different oriented contents. Finally, in order to describe the local Human Visual System (HVS) conspicuities, a saliency pooling strategy is proposed to generate the resulting image quality index. The experimental results have demonstrated that the proposed method outperforms SSIM and Visual Information Fidelity (VIF), which indicates that the obtained similarity index is more consistent with the perceptual evaluation of image quality. Lin Ma 0002, Songnan Li, King Ngi Ngan |
IEEE Signal Process. Lett. | 3 |
| 2010 | Parametric Interpolation Filter for HD Video CodingabstractRecently, adaptive interpolation filter (AIF) for motion-compensated prediction (MCP) has received increasing attention. This letter studies the existing AIF techniques, and points out that making tradeoff between the two conflicting aspects: the accuracy of coefficients and the size of side information, is the major obstacle to improving the performance of the AIF techniques that code the filter coefficients individually. To overcome this obstacle, parametric interpolation filter (PIF) is proposed for MCP, which represents interpolation filters by a function determined by five parameters instead of by individual coefficients. The function is designed based on the fact that high frequency energies of HD video source are mainly distributed along the vertical and horizontal directions; the parameters are calculated to minimize the energy of prediction error. The experimental results show that PIF outperforms the existing AIF techniques and approaches the efficiency of the optimal filter. Jie Dong 0001, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Real-Time De-Interlacing for H.264-Coded HD VideosabstractIt is very challenging to de-interlace high-definition videos in real time, as both good visual quality and low complexity should be fulfilled, which, however, are conflicting. To resolve the conflict, a syntax-based de-interlacer is proposed specially for H.264-coded videos, by which the values of the syntax elements (SE) in H.264 bitstreams, such as macroblock-type, intra-prediction modes, and motion vectors (MV), are used to detect edge and motion, estimate the MV, and select proper local de-interlacing algorithms accordingly. A verification mechanism is also proposed to make sure the referred SE values are reliable. The experimental results show that the proposed de-interlacer provides better visual quality than common ones and can de-interlace 1080isequences in real time on PCs. Jie Dong 0001, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | Automatic scale selection for corners and junctionsabstractIn this paper, we aim to address the problem of automatic scale selection for image features like corners and junctions. Image neighborhood of these features usually contain the background and multiple foreground surfaces, thus can not be correctly described by a single scale. We show that the proposed Fan Laplacian-of-Gaussian (FLOG) kernel is capable to select the appropriate scales for independent image partitions that can represent meaningful physical surfaces attached to the corner or junction. Support for the proposed method is given in terms of theoretical investigation and experiments on real images. Chunhui Cui, King Ngi Ngan |
ICIP | 2 |
| 2009 | Parametric interpolation filter for motion compensated predictionabstractRecently, adaptive interpolation filter (AIF) has received increasing attention for motion-compensated prediction (MCP). The existing methods code the filter coefficients individually and the accuracy of coefficients and the size of side information are conflicting. This paper studies the effect of making trade-off between the two conflicting aspects and proposes the parametric interpolation filter (PIF), which represents filters by five parameters instead of individual coefficients and approximate the optimal filter by tuning the parameters. The experimental results show that PIF approaches the efficiency of the optimal filter and outperforms the related work. Jie Dong 0001, King Ngi Ngan |
ICIP | 2 |
| 2009 | Composite modeling of optical flow for artifacts reductionabstractBecause of the outstanding contribution in removing pixel correlation, block-based transform and quantization has been widely accepted in the state-of-art image/video coding standards. However, artifacts are introduced, which result in reducing the subjective quality of the decoded image/video. In this paper, a composite model of the optical flow velocity is mathematically derived and proved to further reduce the blocking artifacts by compensating the DC surface of the decoded image according to the model. The functional relationship among the optical flow velocity of the DC surface of the decoded image, the value discontinuity across the block boundary in the decoded image and the quantity of DC surface adjustment is analyzed and mathematically modeled as a composite function with second-degree polynomial. Four feature points on the model with the corresponding criteria in artifacts reduction are compared to show a good consistency between the model and the actual filtering effects. Finding a point on the model with a better trade-off between the optical flow smoothness and the edge smoothness is promising in producing a more pleasing and natural looking image/video. Xin Jin 0002, Satoshi Goto, King Ngi Ngan |
ICME | 3 |
| 2009 | Fully Scalable Multiview Wavelet Video CodingabstractThis paper presents a 2-D pipeline-based locally adaptive inter-view-temporal lifting scheme for scalable multi-view video coding. Instead of global temporal and inter-view correlation analysis based on the whole picture, the proposed method adopts locally adaptive inter-view-temporal correlation analysis based on the macroblocks. In addition, in order to reduce the memory requirement and computational complexity and remove both of temporal and view boundary effects, a new 2-D pipeline-based lifting scheme is proposed. Experimental results show that the proposed scheme always outperforms the other tested schemes in coding efficiency, while providing several advantages, such as view scalability, pipeline-based processing, and moderate complexity. Yu Liu 0041, King Ngi Ngan |
ISCAS | 2 |
| 2009 | Enhancing Compression Rate by Just-noticeable Distortion Model for H.264/AVCabstractIn this paper, we proposed an algorithm to enhance compression rate for the H.264/AVC. The properties of human visual system are utilized to reduce the bit rate without introducing noticeable visual distortion. The just-noticeable distortion (JND) threshold for every block in a frame is computed, and non-zero coefficients are removed if their magnitudes are less than the JND thresholds. Experimental results show that bit rate can be reduced by as much as 50% without degradation in visual quality. Chun-Man Mak, King Ngi Ngan |
ISCAS | 2 |
| 2009 | The Perceptually Transparent Coding for ImageabstractIn this paper, an improved transparent image coding algorithm is proposed, where the quantization factor can be tuned for each block according to the just-noticeable distortion (JND). The experimental results show that the images compressed by the proposed method are hardly distinguished from the original images. Therefore, the proposed method can achieve perceptually transparent quality for images. Moreover, the bitrate of the proposed algorithm is less than that of the state-of-the-arts lossless and near-lossless codecs. King Ngi Ngan |
ISCAS | 2 |
| 2009 | Optical flow based DC surface compensation for artifacts reductionabstractThe block-based prediction, transform and quantization reduce the redundancy in video data efficiently. However, the blocking artifacts are introduced because of quantization error and prediction scheme. Even if the de-blocking filter is applied like that designed in H.264 and AVS, the blocking artifacts are still visible especially in videos with smooth area coded at low bit-rate. In this paper, an optical flow based technique is proposed to further reduce the blocking artifacts by compensating the DC surfaces of the decoded images using joint optimization of optical flow and edge difference. The proposed algorithm automatically classifies the macroblocks in the decoded images according to their texture feature and endures the smooth area with smooth brightness variation by minimizing the joint function of optical flow magnitude and edge smoothness. It produces a more pleasing image without over-filtering, which can be further applied as an in loop filter for video coding. Xin Jin 0002, Satoshi Goto, King Ngi Ngan |
PCS | 3 |
| 2009 | Special issue on AVS and its applications: Guest editorial
Wen Gao 0001, King Ngi Ngan, Lu Yu 0003 |
Signal Process. Image Commun. | 2 |
| 2009 | Platform-independent MB-based AVS video standard implementation
Xin Jin 0002, Songnan Li, King Ngi Ngan |
Signal Process. Image Commun. | 3 |
| 2009 | A two-pass rate control algorithm for H.264/AVC high definition video coding
King Ngi Ngan, Zhenzhong Chen 0001 |
Signal Process. Image Commun. | 2 |
| 2009 | 2-D Order-16 Integer Transforms for HD Video CodingabstractIn this paper, the spatial properties of high-definition (HD) videos are investigated based on a large set of HD video sequences. Compared with lower resolution videos, the prediction errors of HD videos have higher correlation. Hence, we propose using 2-D order-16 transforms for HD video coding, which are expected to be more efficient to exploit this spatial property, and specifically propose two types of 2-D order-16 integer transforms, nonorthogonal integer cosine transform (ICT) and modified ICT. The former resembles the discrete cosine transform (DCT) and is approximately orthogonal, of which the transform error introduced by the nonorthogonality is proven to be negligible. The latter modifies the structure of the DCT matrix and is inherently orthogonal, no matter what the values of the matrix elements are. Both types allow selecting matrix elements more freely by releasing the orthogonality constraint and can provide comparable performance with that of the DCT. Each type is integrated into the audio and video coding standard (AVS) Enhanced Profile (EP) and the H.264 high profile (HP), respectively, and used adaptively as an alternative to the 2-D order-8 transform according to local activities. At the same time, many efforts have been devoted to further reducing the complexity of the 2-D order-16 transforms and specially for the modified ICT, a fast algorithm is developed and extended to a universal approach. Experimental results show that 2-D order-16 transforms provide significant performance improvement for both AVS enhanced profile and H.264 high profile, which means they can be efficient coding tools especially for HD video coding. Jie Dong 0001, King Ngi Ngan, Chi-Keung Fong, Wai-kuen Cham |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | Pre- and Post-Shift Filtering for Blocking Removing in Downsizing TranscodingabstractWhen high-quality video is delivered via a bandwidth-limited network, one solution is to reduce the spatial resolution of the bitstream before transmission. Downsizing transcoding is adopted as an efficient tool to reduce the spatial resolution of the precoded bitstream. Since the downsizing process and the coding process both apply a low-pass filtering operation on the picture, the reconstructed video may result in serious blocking artifact. In this paper, pre- and post-shift filters are considered in the transcoding process. The block boundary is shifted between the downsizing process and the coding process. This reduces the blocking effect in downsizing transcoding. Experimental results show that the proposed scheme can efficiently remove the blocking effect without introducing high computation. Haiyan Shu, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | Spatio-Temporal Just Noticeable Distortion Profile for Grey Scale Image/Video in DCT DomainabstractIn image and video processing field, an effective compression algorithm should remove not only the statistical redundancy information but also the perceptually insignificant component from the pictures. Just-noticeable distortion (JND) profile is an efficient model to represent those perceptual redundancies. Human eyes are usually not sensitive to the distortion below the JND threshold. In this paper, a DCT based JND model for monochrome pictures is proposed. This model incorporates the spatial contrast sensitivity function (CSF), the luminance adaptation effect, and the contrast masking effect based on block classification. Gamma correction is also considered to compensate the original luminance adaptation effect which gives more accurate results. In order to extend the proposed JND profile to video images, the temporal modulation factor is included by incorporating the temporal CSF and the eye movement compensation. Moreover, a psychophysical experiment was designed to parameterize the proposed model. Experimental results show that the proposed model is consistent with the human visual system (HVS). Compared with the other JND profiles, the proposed model can tolerate more distortion and has much better perceptual quality. This model can be easily applied in many related areas, such as compression, watermarking, error protection, perceptual distortion metric, and so on. King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | FaceSeg: Automatic Face Segmentation for Real-Time VideoabstractSegmenting human faces automatically is very important for face recognition and verification, security system, and computer vision. In this paper, we present an accurate segmentation system for cutting human faces out from video sequences in real-time. First, a learning based face detector is developed to rapidly find human faces. To speed up the detection process, a face rejection cascade is constructed to remove most of negative samples while retaining all the face samples. Then, we develop a coarse-to-fine segmentation approach to extract the faces based on a min-cut optimization. Finally, a new matting algorithm is proposed to estimate the alpha-matte based on an adaptive trimap generation method. Experimental results demonstrate the effectiveness and robustness of our proposed method that can compete with the well-known interactive methods in real-time. Hongliang Li 0001, King Ngi Ngan, Qiang Liu 0015 |
IEEE Trans. Multim. | 2 |
| 2008 | A temporal just-noticeble distortion profile for video in DCT domainabstractJust-noticeable distortion (JND) profile is an efficient model to represent the perceptual redundancies. Human eyes usually are not sensitive to the distortion below the JND threshold. In order to extend the spatial JND profile to video images by considering the temporal effects, a temporal JND modulation factor is introduced in this paper, which not only incorporates the temporal contrast sensitivity function (CSF) and the smooth pursuit eye movement (SPEM), but also considers the directionality of the motion in video sequences. This factor can be easily combined with other existing subband spatial JND models to achieve a complete spatio-temporal JND profile. Experimental results show that the proposed model is consistent with the human visual system (HVS). Compared with the other JND profiles, the proposed model can tolerate more distortion and has much better perceptual quality at the same time. Since the proposed JND model is DCT based, it can be easily applied in many related areas, such as compression, watermarking, error protection, perceptual distortion metric, and so on. King Ngi Ngan |
ICIP | 2 |
| 2008 | An adaptive and parallel scheme for hd video de-interlacingabstractIt is very challenging to de-interlace HD videos in real time, as both high efficiency and low complexity should be fulfilled, which, however, are conflicting. This paper presents a de-interlacer to resolve the conflict specially for H.264 coded videos. It adapts to spatially and temporally local activities by making full use of the syntax element (SE) values in bit-streams, which give many hints of the motions and textures of video sequences. Accuracy analysis is also introduced to deal with the disparity between the SE values and the real motions and textures. The experimental results show the proposed de-interlacer provides better visual quality than common ones and can de-interlace 1080isequences in real time on PCs. Jie Dong 0001, King Ngi Ngan |
ICME | 2 |
| 2008 | Spatial just noticeable distortion profile for image in DCT domainabstractIn this paper, a DCT based JND model for monochrome pictures is proposed. This model incorporates the spatial contrast sensitivity function (CSF), the luminance adaptation effect and the contrast masking effect based on block classification. Gamma correction is also considered to compensate the original luminance adaptation effect which gives more accurate results. Moreover, a psychophysical experiment was designed to parameterize our model. Experimental results show that the proposed model is consistent with the human visual system. Compared with the other JND profiles, the proposed model can tolerate more distortion and has much better perceptual quality. The proposed JND model can be easily applied in many related areas, such as compression, watermarking, error protection, perceptual distortion metric, and so on. King Ngi Ngan |
ICME | 2 |
| 2008 | 3-D direction aligned wavelet transform for scalable video codingabstractThis paper presents a novel 3-D direction aligned wavelet transform for scalable video coding. A new generalized 3- D directional threading technique is seamlessly incorporated into 3-D weighted adaptive lifting (WAL)-based wavelet transform to exploit the spatio-temporal correlation inside the video cube along the 3-D directional trajectory, leading to a new class of algorithm called 3-D direction aligned wavelet transform (DAWT). Experimental results show that the proposed 3-D WAL-based DAWT achieves better coding performance than the conventional 3-D discrete wavelet transform (3-D DWT). Yu Liu 0041, King Ngi Ngan, Feng Wu 0001 |
ISCAS | 2 |
| 2008 | Constant distortion rate control for H.264/AVC high definition videos with scene changeabstractIn this paper, we propose a constant distortion rate control algorithm for H.264/AVC high definition videos with scene change. With the rate and distortion information of each frame, the frame scene complexity is modeled and the scene change is detected in the first pass encoding. In the second-pass encoding, the expected distortion of each GOP under the constraint of target bits is derived based on the first-pass statistics. And the quantization step (Q-step) of each frame can be obtained with a linear distortion-quantizer (D-Q) model by considering the parameter update of this model at scene change. The experimental results show that our proposed algorithm can provide more constant quality among different video scenes, compared to traditional joint model (JM) rate control algorithm. Zhenzhong Chen 0001, King Ngi Ngan |
ISCAS | 3 |
| 2008 | Saliency model-based face segmentation and tracking in head-and-shoulder video sequences
Hongliang Li 0001, King Ngi Ngan |
J. Vis. Commun. Image Represent. | 2 |
| 2008 | Adaptive partition size temporal error concealment for H.264 using weighted double-sided EBME minimization
King Ngi Ngan |
Signal Process. Image Commun. | 2 |
| 2008 | An efficient intra-mode selection algorithm for H.264 based on edge classification and rate-distortion estimation
King Ngi Ngan, Hongliang Li 0001 |
Signal Process. Image Commun. | 2 |
| 2008 | Fast and Efficient Method for Block Edge Classification and Its Application in H.264/AVC Video CodingabstractEdge is an important feature in video classification which finds applications in video representation and coding. In H.264/AVC, intra-prediction mode decision (a computationally intensive process) is based on the orientation of the edges in the macroblock. In this paper, we first investigate the difference properties derived from three coefficients in the non-normalized Haar transform (NHT) domain and present a fast and efficient method to classify block edge using these properties. The proposed method significantly reduces the number of computational operations in the edge models determination with no multiplications and less addition operations. The use of edge classification for fast intra-prediction mode decision in H.264/AVC video coding is then presented. The experimental results show the effectiveness of the proposed method. Hongliang Li 0001, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | 3-D Shape-Adaptive Directional Wavelet Transform for Object-Based Scalable Video CodingabstractThis paper presents a 3-D shape-adaptive directional wavelet coding technique for object-based scalable video coding. This technique includes 3-D object-based directional threading and extensions of weighted adaptive lifting (WAL) scheme. The 3-D object-based directional threading, which unifies the concept of temporal motion threading and 2-D spatial directional threading, provides the opportunities to align a series of video object planes to form a 3-D video object and exploit the spatio-temporal correlation inside the 3-D video object. The WAL scheme, which is extended from 2-D frame-based image coding to 3-D object-based video coding, is employed to decompose the 3-D video object into a 3-D multispatio-temporal resolution video object pyramid for object-based scalable video coding. Experimental results show that the proposed 3-D shape-adaptive directional wavelet coding technique consistently outperforms MPEG-4 and other wavelet-based schemes for coding arbitrarily shaped video objects. Yu Liu 0041, King Ngi Ngan, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Weighted Adaptive Lifting-Based Wavelet Transform for Image CodingabstractIn this paper, a new weighted adaptive lifting (WAL)-based wavelet transform is presented. The proposed WAL approach is designed to solve the problems existing in the previous adaptive directional lifting (ADL) approach, such as mismatch between the predict and update steps, interpolation favoring only horizontal or vertical direction, and invariant interpolation filter coefficients for all images. The main contribution of the proposed approach consists of two parts: one is the improved weighted lifting, which maintains the consistency between the predict and update steps as far as possible and preserves the perfect reconstruction at the same time; another is the directional adaptive interpolation, which improves the orientation property of the interpolated image and adapts to statistical property of each image. Experimental results show that the proposed WAL-based wavelet transform for image coding outperforms the conventional lifting-based wavelet transform up to 3.06 dB in PSNR and significant improvement in subjective quality is also observed. Compared with the ADL-based wavelet transform, up to 1.22-dB improvement in PSNR is reported. Yu Liu 0041, King Ngi Ngan |
IEEE Trans. Image Process. | 2 |
| 2007 | Unsupervised Multiple Object Segmentation of Multiview Images
King Ngi Ngan |
ACIVS | 2 |
| 2007 | A Fast Rate-Distortion Optimization Algorithm for H.264/AVCabstractIn H.264/AVC video coding standard, rate distortion optimization (RDO) is a very efficient technique to decide the coding mode of the macroblock and improves the coding performance greatly. But at the same time, RDO introduces a great deal of computation complexity because it needs to do the transform and entropy coding to get accurate distortion and bit-rate to calculate the Lagrange cost. In order to avoid the expensive computational cost, in this paper , a simple and accurate bit estimation model is proposed and used in RDO scheme. Combined with distortion measure method in transform domain and the fast intra mode RDO method, a lot of computation of RDO is reduced. Experimental results prove that our proposed fast RDO algorithm can reduce about 60% total encoding time and save about 76% computation time of RDO module with little degradation in coding performance. King Ngi Ngan |
ICASSP (1) | 2 |
| 2007 | Weighted Adaptive Lifting-Basedwavelet TransformabstractIn this paper, we propose a new weighted adaptive lifting (WAL)-based wavelet transform that is designed to solve the problems existing in the previous adaptive directional lifting (ADL) approach. The proposed approach uses the weighted function to make sure that the prediction and update stages are consistent, the directional interpolation to improve the orientation property of interpolated image, and adaptive interpolation filter to adjust to statistical property of each image. Experimental results show that the proposed WAL-based wavelet transform for image coding outperforms the conventional lifting-based wavelet transform up to 3.02 dB in PSNR and significant improvement in subjective quality is also observed. Compared with the ADL approach, up to 1.18 dB improvement in PSNR is reported. Yu Liu 0041, King Ngi Ngan |
ICIP (3) | 2 |
| 2007 | Pre-And Post-Shift Filtering for Removing Blocking Effects in Downsizing TranscodingabstractWhen the high definition video is delivered via a limited bandwidth network, one solution is to reduce the spatial resolution of the bitstream. Downsizing transcoding is adopted as an efficient tool to reduce the spatial resolution of the pre-coded bitstream. Since the downsizing process and the coding process apply a low pass filter to the picture, the reconstructed video may present serious blocking artifact. In this paper, pre-and post-shift filters are considered in the transcoding process. The block boundary is shifted between the downsizing process and the coding process. This reduces the blocking effect in downsizing transcoding. Experimental results show that the proposed scheme can efficiently reduce the blocking effect without introducing high computation. Haiyan Shu, King Ngi Ngan |
ICME | 2 |
| 2007 | Enhancement Techniques for Intra Block MatchingabstractThe intra prediction techniques employed in H.264 predict the intensities of the pixels of the current block from the neighboring pixels. Coding of intra slices requires a large number of bits. Therefore, to improve the coding efficiency, another intra prediction technique is applied with several proposed enhancement techniques including best match prediction, multiple best matches, novel padding method, skip mode, etc. Experimental results show that the coding efficiency can be improved significantly, and more than 1 dB and 0.6 dB gain in PSNR can be achieved for intra and hybrid coding, respectively. Kai Lam Tang, King Ngi Ngan |
ICME | 2 |
| 2007 | A Universal Approach to Developing Fast Algorithm for Simplified Order-16 ICTabstractSimplified order-16 integer cosine transform (ICT) has been proved to be an efficient coding tool especially for high-definition (HD) video coding and is much simpler than ICT and discrete cosine transform (DCT). To further reduce the computational complexity, a universal approach to developing fast algorithm mainly for, but not restricted to, simplified order-16 ICT is proposed in this paper. The fast algorithm developed by the proposed approach involves additions and shiftings only and can save about 90% of the computational time compared with matrix multiplication. Jie Dong 0001, King Ngi Ngan, Chi-Keung Fong, Wai-kuen Cham |
ISCAS | 2 |
| 2007 | 3D Object-based Scalable Wavelet Video Coding with Boundary Effect SuppressionabstractThis paper extends the lifting-based motion threading technique from the frame-based coding to the object-based coding, attracted by the unique advantages of the object-based coding that do not exist in other coding schemes. The boundary effects, which exist in spatial and temporal wavelet transforms due to the manner of object-based coding, are well suppressed by 3D shape adaptive discrete wavelet transform (3D SA-DWT) via the lifting implementation in a unified framework. Based on the proposed object-based motion threading technique, a 3D object-based scalable wavelet video coder is developed. Experimental results show that the proposed coder achieves good performance compared with the existing object-based coders, and provides the scalability functionality at the same time. Yu Liu 0041, Feng Wu 0001, King Ngi Ngan |
ISCAS | 3 |
| 2007 | Quality Enhancement in H.264 Transform Domain DownsizingabstractIn the H.264 coding standard, an integer cosine transform (ICT) is adopted for transform coding. As the ICT is an integer approximation to the discrete cosine transform (DCT), its performance is different from that of the DCT. When the scheme of discarding high frequency coefficients is employed in the ICT domain downsizing, the output quality is worse than when done in the DCT domain. In this paper, the authors propose a scheme to enhance the output quality by converting the downsizing operation back to the DCT domain. The experimental results show that the proposed schemes present quality enhancement over the ICT domain downsizing. Haiyan Shu, King Ngi Ngan |
ISCAS | 2 |
| 2007 | An Efficient Intra Mode Selection Algorithm For H.264 Based On Fast Edge ClassificationabstractThe H.264/AVC is the newest video coding standard recommended by ITU-T and MPEG. Compared with all existing video coding standards, H.264 can achieve superior performance by using many advanced techniques. Intra mode selection is an important feature in H.264 standard and can reduce the spatial redundancy in intra frame significantly. An efficient rate distortion optimization (RDO) technique is employed in H.264 to choose the best mode for each MB, but the computational cost increases drastically. In this paper, a fast intra mode selection algorithm is introduced. By using a fast edge detection method which is based on non-normalized Haar transform (NHT), edge for each sub-block can be extracted. Based on the local edge information, only few intra modes are chosen as mode candidates. A fast RDO algorithm is also proposed in this paper. By combing these two methods, computational load is reduced remarkably. Experimental results show that this fast intra mode selection scheme can lessen about 80% encoding time with little loss of bit-rate and visual quality. Hongliang Li 0001, King Ngi Ngan |
ISCAS | 3 |
| 2007 | External Calibration of Multi-camera System Based on Efficient Pair-wise Estimation
Chunhui Cui, King Ngi Ngan |
PSIVT | 3 |
| 2007 | A novel statistical learning-based rate distortion analysis approach for multiscale binary shape codingabstractIn this paper, we propose a statistical learning based approach to analyze the rate-distortion characteristics of multiscale binary shape coding. We employ the polynomial kernel function and incorporate rate-distortion related features for our support vector regression. ε-Insensitive loss function is chosen to improve the estimation robustness. The parameter tuning is also studied. Moreover, we discuss the feature selection which helps to improve the estimation accuracy. Comparing to the traditional method, our proposed framework provides better rate distortion estimation not only on simple shapes but also on complex shapes. Zhenzhong Chen 0001, King Ngi Ngan |
VCIP | 2 |
| 2007 | Enhanced SAD reuse fast motion estimationabstractH.264 is a new video coding standard which outperforms the previous video coding standards. It uses many advanced video coding techniques to improve the coding performance. Variable block size motion estimation (ME) is one of the techniques that contributes to the excellent performance of H.264 but it is computational intensive. In this paper, a fast SAD reuse ME algorithm is proposed which reuses SAD within the same macroblock and uses pattern-based ME and refinement search to reduce the computational complexity of variable block size ME with a little degradation of coding performance in terms of PSNR and bitrate. Experimental results show that the proposed algorithm reduces the ME time by more than 90% on average with only a little degradation of coding performance when comparing with that of the Fast Full Search (FFS). Kai Lam Tang, King Ngi Ngan |
VCIP | 2 |
| 2007 | Recent advances in rate control for video coding
Zhenzhong Chen 0001, King Ngi Ngan |
Signal Process. Image Commun. | 2 |
| 2007 | Fast multiresolution motion estimation algorithms for wavelet-based scalable video coding
Yu Liu 0041, King Ngi Ngan |
Signal Process. Image Commun. | 2 |
| 2007 | Towards Rate-Distortion Tradeoff in Real-Time Color Video CodingabstractIn this paper, we address the key problem in real-time video coding, the rate-distortion (R-D) tradeoff. As most video coding applications employ color images, we analyze the R-D modeling problem of the color image sequence. We investigate that the R-D characteristics of color video signal in traditional transform-based video coding systems should be modeled for the luminance and chrominance components separately. So we propose separable R-D models for color video coding. The feedback from the encoder buffer is analyzed by a control-theoretic adaptation approach to avoid buffer overflow and underflow. To achieve smooth video quality and satisfy the delay constraints in real-time applications, a novel R-D tradeoff controller is designed. In the proposed R-D tradeoff framework, both the quality variation and buffer safety are considered. Extensive simulation results demonstrate that the proposed approach can achieve smooth video quality without sacrificing overall quality while efficiently utilize the bandwidth without buffer overflow or underflow Zhenzhong Chen 0001, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | A Unified Approach of Bit-Rate Control for Binary and Gray-Level Shape Sequences CodingabstractThis paper describes a unified framework of bit-rate control for binary and gray-level shape coding. An exponential rate-distortion model is proposed for binary shape and used in the bit-rate control. An approximation method is applied for finer distortion scale adjustment by adjusting the conversion ratio at the binary alpha block level. Since the gray-level shape is encoded as its support information and the alpha plane information, the rate-quantizer relationship of the alpha plane is modeled. The bit-rate control is then able to adjust the quantizer of the alpha plane for the bit-rate control purpose. The extension for video object (VO) coding is also studied. In this bit-rate control approach, a buffer feedback controller is employed to determine the target bit rates for different sources, i.e., binary shape, gray-level shape, and texture. Accordingly, the distortion scales of different sources are adjusted based on their rate-distortion relationships to maintain a stable encoder buffer. The advantages of the proposed method are demonstrated through various tests for different scenarios. Overall, the algorithm can efficiently encode the shape information and shows a good performance for VO coding. Zhenzhong Chen 0001, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | Optimized Cross-Layer Design for Scalable Video Transmission Over the IEEE 802.11e NetworksabstractA cross-layer design for optimizing 3-D wavelet scalable video transmission over the IEEE 802.11e networks is proposed. A thorough study on the behavior of the IEEE 802.11e protocol is conducted. Based on our findings, all timescales rate control is developed featuring a unique property of soft capacity support for multimedia delivery. The design consists of a macro timescale and a micro timescale rate control schemes residing at the application layer and the network sublayer respectively. The macro rate control uses bandwidth estimation to achieve optimal bit allocation with minimum distortion. The micro rate control employs an adaptive mapping of packets from video classifications to appropriate network priorities which preemptively drops less important video packets to maximize the transmission protection to the important video packets. The performance is investigated by simulations highlighting advantages of our cross-layer design. Chuan Heng Foh, Yu Zhang 0004, Zefeng Ni, Jianfei Cai 0001, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2007 | Unsupervized Video Segmentation With Low Depth of FieldabstractIn this paper, a novel segmentation algorithm based on matting model is proposed to extract the focused objects in low depth-of-field (DoF) video images. The proposed algorithm is fully automatic and can be used to partition the video image into focused objects and defocused background. This method consists of three stages. The first stage is to generate a saliency map of the input image by the reblurring model. In the second stage, bilateral and morphological filtering are employed to smooth and accentuate the salient regions. Then a trimap with three regions is calculated by an adaptive thresholding method. The third stage involves the proposed adaptive error control matting scheme to extract the boundaries of the focused objects accurately. Experimental evaluation on test sequences shows that the proposed method is capable of segmenting the focused region effectively and accurately. Hongliang Li 0001, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | 3-D Object-Based Scalable Wavelet Video Coding With Boundary Effect SuppressionabstractThis paper extends the lifting-based motion threading (MTh) technique from the frame-based coding to the object-based coding, attracted by the unique advantages of the object-based coding that do not exist in other coding schemes. The boundary effects, which exist in spatial and temporal wavelet transforms due to the manner of object-based coding, are well suppressed by 3-D shape-adaptive discrete wavelet transform (3D SA-DWT) via the lifting implementation in a unified framework. Based on the proposed object-based MTh technique, a 3-D object-based scalable wavelet video coder is developed. Experimental results show that the proposed coder achieves good performance compared with the existing object-based coders and provides the scalability functionality at the same time Yu Liu 0041, Feng Wu 0001, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2007 | A Rate and Distortion Analysis of Multiscale Binary Shape Coding Based on Statistical LearningabstractIn this paper, we propose a statistical learning-based approach to analyze the rate-distortion characteristics of MPEG-4 multiscale binary shape coding. We employ the polynomial kernel function and epsiv-insensitive loss function for our support vector regression. To improve the accuracy of the estimation, rate and distortion related features are incorporated in the statistical learning framework. Our experimental results show that the proposed approach can achieve good performance, e.g., modelling the rate-distortion curves accurately. Zhenzhong Chen 0001, King Ngi Ngan |
IEEE Trans. Multim. | 2 |
| 2006 | Unsupervised Segmentation of Defocused Video Based on Matting ModelabstractIn this paper, an unsupervised segmentation algorithm based on matting model is proposed to extract the focused objects in the low depth of field (DOF) video images. The proposed algorithm is fully automatic and can be used to partition the video image into focused objects and defocused background. This method consists of three stages. The first stage is to generate the saliency map from the input image. In the second stage, bilateral and morphological filtering are employed to smooth and lift the saliency regions. Then a trimap with three regions is calculated by an adaptive thresholding method. The third stage involves the Poisson matting scheme to extract the boundaries of the focused objects accurately. Experimental evaluation on test sequences shows that the proposed method is capable of segmenting the focused region quite effectively and accurately. Hongliang Li 0001, King Ngi Ngan |
ICIP | 2 |
| 2006 | Towards rate-distortion tradeoff in real-time color video codingabstractIn this paper, we address to the key problem in real-time video coding, the rate-distortion tradeoff. We found that the distribution of the integer transform coefficients in H.264 color video coding can be approximated by Laplacian distribution but the rate-distortion characteristics should be modelled for the luminance and chrominance components separately. So we propose separable rate-distortion models for color video coding. The feedback from the encoder buffer is analyzed by a control-theoretic adaptation approach. To achieve smooth video quality and satisfy the delay constraints in real-time applications, a novel rate distortion tradeoff controller is employed. Simulations results demonstrate that the proposed approach can achieve smooth video quality without sacrificing overall quality while efficiently utilize the bandwidth without buffer overflow or underflow. Zhenzhong Chen 0001, King Ngi Ngan |
ISCAS | 2 |
| 2006 | Face segmentation in head-and-shoulder video sequences based on facial saliency mapabstractIn this paper, a novel face segmentation algorithm is proposed based on facial saliency map (FSM) for head-and-shoulder type video application. This method consists of three stages. The first stage is to generate the saliency map of input video image by our proposed facial attention model. In the second stage, a geometric model and a eye-map built from chrominance components are employed to localize face region according to the saliency map. The third stage involves the adaptive boundary correction and the final face contour extraction. Experimental evaluation on test sequences shows that the proposed method is capable of segmenting and tracking the face area quite effectively. Hongliang Li 0001, King Ngi Ngan |
ISCAS | 2 |
| 2006 | Fast lossless multi-resolution motion estimation for scalable wavelet video codingabstractIn this paper, we propose a fast lossless multi-resolution motion estimation algorithm in shift-invariant wavelet domain by using wavelet matching error characteristic based partial distortion elimination (MR-WMEC-PDE). Due to its multi-resolution nature, the proposed approach can be applied to scalable video coding. The proposed algorithm can achieve the same estimate accuracy as full search algorithm (FSA), while needing much less computation requirement than FSA Yu Liu 0041, King Ngi Ngan |
ISCAS | 2 |
| 2006 | Fast and efficient method for block edge classificationabstractAdvanced multimedia applications will have to provide the user with the flexibility to rapidly access and manipulate the multimedia data. In order to achieve this goal, especially in limited computing environments, such as mobile-phone, the computational cost of extracting the visual features for image segmentation must be greatly reduced. In this paper, we investigate difference properties from three coefficients in the non-normalised Haar transform (NHT) domain and present a fast and efficient method to classify block edge using these properties. The proposed method significantly reduces the number of computational operations in the edge models determination with no multiplications and less addition operations. The experimental results are presented to show the effectiveness of the proposed method. Hongliang Li 0001, King Ngi Ngan |
IWCMC | 2 |
| 2006 | Efficient intra- and inter-mode selection algorithms for H.264/ AVC
Andy C. Yu, King Ngi Ngan, Graham R. Martin |
J. Vis. Commun. Image Represent. | 2 |
| 2006 | Distortion variation minimization in real-time video coding
Zhenzhong Chen 0001, King Ngi Ngan |
Signal Process. Image Commun. | 2 |
| 2006 | Embedded wavelet packet object-based image coding based on context classification and quadtree ordering
Yu Liu 0041, King Ngi Ngan |
Signal Process. Image Commun. | 2 |
| 2006 | Unsupervised extraction of visual attention objects in color imagesabstractThis paper proposes a generic model for unsupervised extraction of viewer's attention objects from color images. Without the full semantic understanding of image content, the model formulates the attention objects as a Markov random field (MRF) by integrating computational visual attention mechanisms with attention object growing techniques. Furthermore, we describe the MRF by a Gibbs random field with an energy function. The minimization of the energy function provides a practical way to obtain attention objects. Experimental results on 880 real images and user subjective evaluations by 16 subjects demonstrate the effectiveness of the proposed approach. Junwei Han 0001, King Ngi Ngan, Mingjing Li, HongJiang Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2006 | Dynamic Programming-Based Reverse Frame Selection for VBR Video Delivery Under Constrained ResourcesabstractIn this paper, we investigate optimal frame-selection algorithms based on dynamic programming for delivering stored variable bit rate (VBR) video under both bandwidth and buffer size constraints. Our objective is to find a feasible set of frames that can maximize the video's accumulated motion values without violating any constraint. It is well known that dynamic programming has high complexity. In this research, we propose to eliminate nonoptimal intermediate frame states, which can effectively reduce the complexity of dynamic programming. Moreover, we propose a reverse frame selection (RFS) algorithm, where the selection starts from the last frame and ends at the first frame. Compared with the conventional dynamic programming-based forward frame selection, the RFS is able to find all of the optimal results for different preloads in one round. We further extend the RFS scheme to solve the problem of frame selection for VBR channels. In particular, we first perform the RFS algorithm offline, and the complexity is modest and scalable with the aids of frame stuffing and nonoptimal state elimination. During online streaming, we only need to retrieve the optimal frame-selection path from the pregenerated offline results, and it can be applied to any VBR channels as long as the VBR channels can be modeled as piecewise CBR channels. Experimental results show good performance of our proposed algorithms Dayong Tao, Jianfei Cai 0001, Haoran Yi, Deepu Rajan, Liang-Tien Chia, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2006 | 4-D Wavelet-Based Multiview Video CodingabstractThe conventional multiview video coding (MVC) schemes, utilizing both neighboring temporal frames and view frames as possible references, have only shown a slight gain over those using temporal frames alone in terms of coding efficiency. The reason for this is that the neighboring temporal frames exhibit stronger correlation with the current frame and the view frames often fail to be selected as references. This paper proposes an elegant MVC framework using high dimensional wavelet, which rightly matches the inherent high dimension property of multiview video. It also makes a better usage of both temporal and view correlations thanks to the hierarchical decomposition. Besides the proposed framework, this paper also investigates MVC coding from the following aspects. First, a disparity-compensated view filter (DCVF) with pixel alignment is proposed, which can accommodate both global and local view disparities among view frames. The proposed DCVF and the existing motion-compensated temporal filter (MCTF) unify the view and temporal decompositions as a generic lifting transform. Second, an adaptive decomposition structure based on the analysis of the temporal and view correlations is proposed. A Lagrangian cost function is derived to determine the optimum decomposition structure. Third, the major components of the proposed MVC coding are figured out, including macroblock type design, subband coefficient coding, and rate allocation. Extensive experiments are carried out on the MPEG 3DAV test sequences and the superior performance of the proposed MVC coding is demonstrated. In addition, the proposed MVC framework can easily support temporal, spatial, SNR, as well as view scalabilities Yan Lu 0001, Feng Wu 0001, Jianfei Cai 0001, King Ngi Ngan, Shipeng Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2006 | An MPEG-4-compatible stereoscopic/multiview video coding schemeabstractIn this paper, we propose an efficient codec for multiview video coding, which is compatible with the MPEG-4 video standard. The main views of the multiview video are encoded using an MPEG-4 encoder and the auxiliary views are encoded by joint disparity and motion compensation. An edge-preserving regularization scheme that jointly calculates disparity and motion vectors is performed on the VOP basis. The output of the encoder contains one bitstream for each view, and the main view bitstreams can be decoded by a standard MPEG-4 decoder. In addition, in the case of five-view encoding, we compare four different prediction structures in order to find the best one under certain scenarios. To evaluate the proposed encoder, the MPEG-2 multiview profile (MVP) is implemented on the MPEG-4 platform for fair comparison, which is referred to as MPEG-4 MVP in this paper. Experimental results prove that the proposed encoder achieves a higher image quality at similar bit rate than the conventional scheme and is very promising for the applications including videoconferencing and three-dimensional telepresence. King Ngi Ngan, Jianfei Cai 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2006 | Dynamic Bit Allocation for Multiple Video Object CodingabstractIn MPEG-4, a visual scene may be treated as a composition of video objects and coded at object level. Such a flexible video coding framework makes it possible to code different video objects with different priority according to human perceptual characteristics. In this paper, we introduce a novel dynamic bit allocation framework to improve the subjective quality in such an object-based video coding system. We incorporate the rate distortion models with the dynamic priorities of the video objects and jointly encode video objects to minimize the weighted distortion within the bit budget constraint. We guarantee the human-interested video objects a better reconstructed quality by using the weighted bit allocation strategy in favour of the video objects with higher priority. To obtain the priority automatically, we apply a visual attention model. Comparing with traditional bit allocation algorithms, the objective quality of the object with higher priority is significantly improved under this framework. These results demonstrate the usefulness of this dynamic bit allocation framework. Zhenzhong Chen 0001, Junwei Han 0001, King Ngi Ngan |
IEEE Trans. Multim. | 3 |
| 2005 | Joint mode selection and unequal error protection for bitplane coded video transmission over wireless channels
Jianfei Cai 0001, Jianhua Wu 0003, King Ngi Ngan, Zhihai He |
J. Vis. Commun. Image Represent. | 3 |
| 2005 | Special issue on visual communication in the ubiquitous era
Chang Wen Chen, Mohammed Ghanbari 0001, King Ngi Ngan |
J. Vis. Commun. Image Represent. | 3 |
| 2005 | VLC/FLC data partitioning with intra AC prediction disabled
Chengji Zhao, King Ngi Ngan |
J. Vis. Commun. Image Represent. | 3 |
| 2005 | A real-time video transport system for the best-effort Internet
Zefeng Ni, Zhenzhong Chen 0001, King Ngi Ngan |
Signal Process. Image Commun. | 3 |
| 2005 | Joint Motion and Disparity Fields Estimation for Stereoscopic Video Sequences
King Ngi Ngan, JeongEun Lim, Kwanghoon Sohn |
Signal Process. Image Commun. | 2 |
| 2005 | Joint texture-shape optimization for MPEG-4 multiple video objectsabstractIn this paper, we present a uniform framework for optimal multiple video object bit allocation in MPEG-4. We combine the rate-distortion (R-D) models for the texture and shape information of arbitrarily shaped video objects to develop the joint texture-shape R-D models. The dynamic programming technique is applied to optimize the bit allocation for the multiple video objects. The simulation results demonstrate that the proposed joint texture-shape optimization algorithm outperforms the MPEG-4 verification model on the decoded picture quality. Zhenzhong Chen 0001, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2005 | A memory learning framework for effective image retrievalabstractMost current content-based image retrieval systems are still incapable of providing users with their desired results. The major difficulty lies in the gap between low-level image features and high-level image semantics. To address the problem, this study reports a framework for effective image retrieval by employing a novel idea of memory learning. It forms a knowledge memory model to store the semantic information by simply accumulating user-provided interactions. A learning strategy is then applied to predict the semantic relationships among images according to the memorized knowledge. Image queries are finally performed based on a seamless combination of low-level features and learned semantics. One important advantage of our framework is its ability to efficiently annotate images and also propagate the keyword annotation from the labeled images to unlabeled images. The presented algorithm has been integrated into a practical image retrieval system. Experiments on a collection of 10,000 general-purpose images demonstrate the effectiveness of the proposed framework. Junwei Han 0001, King Ngi Ngan, Mingjing Li, HongJiang Zhang |
IEEE Trans. Image Process. | 2 |
| 2004 | Multiple feature clustering algorithm for automatic video object segmentationabstractIn this paper, we present an automatic video segmentation algorithm for object-based coding, based on a k-medians clustering algorithm and 2D binary model. Firstly, the k-medians algorithm is employed to partition an image into a set of homogeneous regions. Then, a 2D binary model of the moving object is set up, which, combined with temporal and spatial information, guides the extraction process of the video object planes from the video sequence. The performance of the segmentation algorithm is illustrated by simulations carried out on standard video sequences. King Ngi Ngan, Nariman Habili |
ICASSP (3) | 2 |
| 2004 | MPEG-4 based stereoscopic video sequences encoderabstractIn this paper, we propose an object-based MPEG-4 compatible stereoscopic video sequence encoder. We aim at efficient stereoscopic video compression for videoconferencing and 3D telepresence systems. The stereoscopic video sequence includes one main view and one auxiliary view. The main view is encoded using an MPEG-4 encoder and the auxiliary view is encoded by joint motion and disparity compensation. After the input sequences are balanced to compensate for lighting conditions and camera differences, the joint disparity and motion regularization is performed on a VOP basis. The output of the encoder contains two bitstreams, a main bitstream, which can be decoded by a standard MPEG-4 decoder, and an auxiliary bitstream. Simulation results show that the joint estimation and regularization of disparity and motion fields provide more accurate vector fields and efficient compression for the auxiliary stream. The proposed system achieves high image quality at lower bitrate than existing stereoscopic video encoders. King Ngi Ngan |
ICASSP (3) | 2 |
| 2004 | Optimal bit allocation for MPEG-4 multiple video objectsabstractIn this paper, we present an optimal bit allocation algorithm for multiple video objects (MVO's) in MPEG-4 video coding. We combine the rate-distortion models for texture and shape information to achieve the accurate rate control. To optimize the bit allocation for multiple video objects, we use the dynamic programming (DP). The simulation results demonstrate that the proposed algorithm outperforms the MPEG-4 verification model not only in terms of buffer stability but the objective quality. Zhenzhong Chen 0001, King Ngi Ngan |
ICIP | 2 |
| 2004 | Towards unsupervised attention object extraction by integrating visual attention and object growingabstractContent-related functionalities of image/video applications call for efficient tools that can automatically extract meaningful objects from images. However, traditional methods generally fail to capture objects of user interest because they totally neglect human visual attention perception. Aiming to address this problem, this study proposes a generic model for unsupervised extraction of viewer's attention objects from color images. We formulate the attention objects as a Markov random field (MRF). Then, the MRF is expressed in the form of a Gibbs random field with an energy function. The energy minimization that integrates visual attention and object growing provides a practical way to obtain attention objects. The proposed model works in a manner analogous to humans and has great promise to be a basic tool for content-based image/video applications. Experimental results show the effectiveness of the proposed model. Junwei Han 0001, King Ngi Ngan, Mingjing Li, HongJiang Zhang |
ICIP | 2 |
| 2004 | Integration of motion and image features for automatic video object segmentationabstractThe paper presents an automatic video segmentation algorithm concerning multiple features. The k-median algorithm is employed to partition an image into a set of homogenous regions. The location of moving objects is determined by change detection and tracked by using region descriptors. Experiments have been carried out on several video sequences and results have shown the efficiency of this approach. King Ngi Ngan |
ICIP | 2 |
| 2004 | Learning semantic concepts from user feedback log for image retrievalabstractTo improve the performance of image retrieval systems, the well-known semantic gap needs to be bridged. Relevance feedback provides a strategy for learning semantic concepts from visual features. This paper reports a novel framework to learn semantic concepts from the accumulated user feedback log. The semantic concepts consist of two categories: explicit semantics and implicit semantics. The former can be directly estimated by analyzing the user-provided feedback log. The latter is learned according to the obtained explicit semantics. Finally, both explicit and implicit semantics are applied to an image retrieval system. Experiments on 10,000 images show the superiority of the proposed method. Junwei Han 0001, King Ngi Ngan, Mingjing Li, HongJiang Zhang |
ICME | 2 |
| 2004 | A multiview sequence CODEC with view scalability
JeongEun Lim, King Ngi Ngan, Kwanghoon Sohn |
Signal Process. Image Commun. | 2 |
| 2004 | Linear rate-distortion models for MPEG-4 shape codingabstractIn this letter, we explore the rate-distortion (R-D) characteristics of binary shape information in the MPEG-4 standard. The shape coding scheme is a block-based context-based arithmetic encoding approach and distortion is introduced by down and up sampling. Generally, the shape bit rate and nonnormalized distortion increase in proportion with the number of border blocks. At any given resolution scale, the more complex an object, the more distortion that is introduced. Consequently, we propose an R-D model based on the parameters derived from the border blocks and a block-based shape complexity for the video object. The computational complexity is much lower than other existing methods. Experimental results show that it can accurately predict the bit rate and distortion of the binary shape for rate control purposes. Zhenzhong Chen 0001, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2003 | Improved single video object rate control for MPEG-4abstractRate control plays a central role in constant bit-rate (CBR) video coding applications using MPEG-4. The paper considers single video object (SVO) rate control for MPEG-4 and presents a new rate control algorithm based on the quadratic rate-distortion model. Based on a new measure for the encoding complexity, the constraints of the model parameter estimation and a novel quantizer control strategy are proposed. Simulation results show that the MPEG-4 coder, using the proposed algorithm, can achieve a higher PSNR than a coder using the conventional rate control algorithm. Zhenzhong Chen 0001, King Ngi Ngan |
ICASSP (3) | 2 |
| 2003 | Improved rate control for MPEG-4 video transport over wireless channel
Zhenzhong Chen 0001, King Ngi Ngan, Chengji Zhao |
VCIP | 2 |
| 2003 | Automatic video object segmentation for MPEG-4
King Ngi Ngan |
VCIP | 2 |
| 2003 | Guest Editorial
Yücel Altunbasak, Chang Wen Chen, M. Reha Civanlar, King Ngi Ngan |
Signal Process. Image Commun. | 4 |
| 2003 | Improved rate control for MPEG-4 video transport over wireless channel
Zhenzhong Chen 0001, King Ngi Ngan, Chengji Zhao |
Signal Process. Image Commun. | 2 |
| 2003 | High accuracy flashlight scene determination for shot boundary detection
Wei Jyh Heng, King Ngi Ngan |
Signal Process. Image Commun. | 2 |
| 2003 | Improved single-video-object rate control for MPEG-4abstractRate control plays a central role in constant bit-rate video coding applications using MPEG-4. This paper considers single-video-object rate control for MPEG-4 and presents a new rate-control algorithm based on the quadratic rate-distortion (R-D) model. The major innovations are a novel constraint for the least-mean-square estimation of the model parameters of the R-D function, a new measure for the encoding complexity, novel quantizer control and an efficient frame skipping strategy. Simulation results show that the MPEG-4 coder, using the proposed algorithm, can achieve a higher PSNR than a coder using the conventional rate-control algorithm. King Ngi Ngan, Thomas Meier 0001, Zhenzhong Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2002 | Object-based disparity estimation for stereoscopic imagesabstractDisparity estimation has been a critical procedure in stereoscopic or multiview video processing. In this paper, we present an object-based disparity estimation algorithm capable of providing an accurate disparity map from a stereo pair effectively. The algorithm involves a preprocessing step, which identifies the foreground object, and a predictive hierarchical block matching with different parameters for foreground and background. Left-to-right disparity map is calculated, the right image is used as a reference and the left image is reconstructed at the receiver end. The algorithm has been tested and the quality of the reconstruction using the disparity map demonstrates its effectiveness. King Ngi Ngan |
ICARCV | 2 |
| 2002 | Special Issue on Recent Advances in Wireless Video SIGNAL PROCESSING: IMAGE COMMUNICATION
Yücel Altunbasak, Chang Wen Chen, M. Reha Civanlar, King Ngi Ngan |
Signal Process. Image Commun. | 4 |
| 2002 | Shot boundary refinement for long transition in digital video sequenceabstractShot boundary detection, or scene change detection, is a technique used in the initial step of video indexing. With the recent introduction of object-based shot boundary detection technique, the deficiency of traditional detection techniques in handling long transitions has been overcome. However, such a technique merely detects the existence of a transition without giving accurate positions of the transition start and end for long transitions. Shot boundary refinement fills this gap by pinpointing the exact location of the shot boundary regardless of the transition types. This paper presents the theories behind the transition and classifies existing transition types. A novel shot boundary refinement algorithm is then constructed which works independently of the contents and transition length, and is accurate under most transition types. The results of testing using different indicators and different transition types show that the boundaries of long transitions are detected with minimal errors. The success of shot boundary refinement brings us one step closer in developing an accurate, fully automatic indexing system. Wei Jyh Heng, King Ngi Ngan |
IEEE Trans. Multim. | 2 |
| 2001 | Long-transition analysis for post shot-boundary detection
Wei Jyh Heng, King Ngi Ngan |
VCIP | 2 |
| 2001 | An Object-Based Shot Boundary Detection Using Edge Tracing and Tracking
Wei Jyh Heng, King Ngi Ngan |
J. Vis. Commun. Image Represent. | 2 |
| 2000 | Foreground/Background Bit Allocation for Region-of-Interest CodingabstractTwo bit allocation strategies, namely, maximum bit transfer (MET) and joint bit assignment (JBA) are proposed for region-of-interest coding. The MBT strategy uses a pair of quantizers to facilitate maximum bit transfer from the background to foreground image region. It assigns the highest quantization parameter to the background quantizer, and then determines the finest value of foreground quantizer that can be used without increasing the bit rate. In this approach, the background is always encoded with the coarsest quantization level, but this is not always desirable. Therefore, the JBA strategy can be used instead. It has two modes: user defined and automatic. In the user defined mode, the user can adjust bit consumption using a scale that ranges from "not coding foreground" to "coding foreground only". In the automatic mode, bit allocation is based on the characteristics of the image regions, which include size, motion and priority. The coding results showed that improved image quality was achieved at the same bit rate. Douglas Chai, King Ngi Ngan, Abdesselam Bouzerdoum |
ICIP | 2 |
| 2000 | Improved single VO rate control for constant bit-rate applications using MPEG-4
Thomas Meier 0001, King Ngi Ngan |
VCIP | 2 |
| 2000 | Guest editorial
King Ngi Ngan, Michael G. Strintzis, Masayuki Tanimoto, Yao Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1999 | Integrated Shot Boundary Detection Using Object-Based Technique
Wei Jyh Heng, King Ngi Ngan |
ICIP (3) | 2 |
| 1999 | A Flexible Bayesian Framework for Image SegmentationabstractThis paper presents a new Bayesian framework for image segmentation. The major contribution is a novel optimization strategy that can be applied to any cost function derived from the MAP criterion. Classical Bayesian techniques normally minimize this cost using ICM together with a K-label approach that assigns each pixel a label m/spl isin/{0,1,...,K-1}. Several shortcomings of this approach are pointed out. Our proposed method first extracts initial seeds that represent the interior of regions. The boundary location is then determined by a modified HCF method that labels pixels in the order of decreasing confidence. There is no need for an initial estimate of the segmentation, and no parameter K is required. Moreover, the presented framework can be viewed as a combination of the elegant morphological segmentation approach with the spatial continuity constraints inherent to Markov random fields in Bayesian techniques. Experimental results demonstrate the significant improvements achieved by our optimization strategy. Thomas Meier 0001, King Ngi Ngan |
ICIP (3) | 2 |
| 1999 | A combined source-channel video coding scheme for mobile channels
Chi W. Yap, King Ngi Ngan, Ranjith Liyanapathirana |
Signal Process. Image Commun. | 2 |
| 1999 | Face segmentation using skin-color map in videophone applicationsabstractThis paper addresses our proposed method to automatically segment out a person's face from a given image that consists of a head-and-shoulders view of the person and a complex background scene. The method involves a fast, reliable, and effective algorithm that exploits the spatial distribution characteristics of human skin color. A universal skin-color map is derived and used on the chrominance component of the input image to detect pixels with skin-color appearance. Then, based on the spatial distribution of the detected skin-color pixels and their corresponding luminance values, the algorithm employs a set of novel regularization processes to reinforce regions of skin-color pixels that are more likely to belong to the facial regions and eliminate those that are not. The performance of the face-segmentation algorithm is illustrated by some simulation results carried out on various head-and-shoulders test images. The use of face segmentation for video coding in applications such as videotelephony is then presented. We explain how the face-segmentation results can be used to improve the perceptual quality of a videophone sequence encoded by the H.261-compliant coder. Douglas Chai, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1999 | Video segmentation for content-based codingabstractTo provide multimedia applications with new functionalities, the new video coding standard MPEG-4 relies on a content-based representation. This requires a prior decomposition of sequences into semantically meaningful, physical objects. We formulate this problem as one of separating foreground objects from the background based on motion information. For the object of interest, a 2D binary model is derived and tracked throughout the sequence. The model points consist of edge pixels detected by the Canny operator. To accommodate rotation and changes in shape of the tracked object, the model is updated every frame. These binary models then guide the actual video object plane (VOP) extraction. Thanks to our new boundary postprocessor and the excellent edge localization properties of the Canny operator, the resulting VOP contours are very accurate. Both the model initialization and update stages exploit motion information. The main assumption underlying our approach is the existence of a dominant global motion that can be assigned to the background. Areas that do not follow this background motion indicate the presence of independently moving physical objects. Two alternative methods to identify such objects are presented. The first one employs a morphological motion filter with a new filter criterion, which measures the deviation of the locally estimated optical flow from the corresponding global motion. The second method computes a change detection mask by taking the difference between consecutive frames. The first version is more suitable for sequences with little motion, whereas the second version is better at dealing with faster moving or changing objects. Experimental results demonstrate the performance of our algorithm. Thomas Meier 0001, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1999 | Reduction of blocking artifacts in image and video codingabstractThe discrete cosine transform (DCT) is the most popular transform for image and video compression. Many international standards such as JPEG, MPEG, and H.261 are based on a block-DCT scheme. High compression ratios are obtained by discarding information about DCT coefficients that is considered to be less important. The major drawback is visible discontinuities along block boundaries, commonly referred to as blocking artifacts. These often limit the maximum compression ratios that can be achieved. Various postprocessing techniques have been published that reduce these blocking effects, but most of them introduce unnecessary blurring, ringing, or other artifacts. In this paper, a novel postprocessing algorithm based on Markov random fields (MRFs) is proposed. It efficiently removes blocking effects while retaining the sharpness of the image and without introducing new artifacts. The degraded image is first segmented into regions, and then each region is enhanced separately to prevent blurring of dominant edges. A novel texture detector allows the segmentation of images containing both texture and monotone areas. It finds all texture regions in the image before the remaining monotone areas are segmented by an MRF segmentation algorithm that has a new edge component incorporated to detect dominant edges more reliably. The proposed enhancement stage then finds the maximum a posteriori estimate of the unknown original image, which is modeled by an MRF and is therefore Gibbs distributed. A very efficient implementation is presented. Experiments demonstrate that our proposed postprocessor gives excellent results compared to other approaches, from both a subjective and an objective viewpoint. Furthermore, it will be shown that our technique also works for wavelet encoded images, which typically contain ringing artifacts. Thomas Meier 0001, King Ngi Ngan, Gregory A. Crebbin |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1999 | Special Issue On Representation And Coding Of Images And Video II
King Ngi Ngan, Sethuraman Panchanathan, Thomas Sikora, Ming-Ting Sun |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1998 | Locating Facial Region of a Head-and-Shoulders Color Image
Douglas Chai, King Ngi Ngan |
FG | 2 |
| 1998 | Video Object Plane Segmentation using a Morphological Motion Filter and Hausdorff Object TrackingabstractThis paper considers the decomposition of video sequences into so-called video object planes, which is required for the content-based representation of visual objects in MPEG-4. A new segmentation algorithm is described that identifies physical objects using a morphological motion filter. For the object of interest, a two-dimensional binary model is derived based on areas in the scene that are moving differently from the background. This model is updated each frame to pick up possible rotation and changes in shape of the object. Temporal correspondence is established by a Hausdorff object tracker: The binary model sequences guide the actual video object plane extraction. Since the model points correspond to edges detected by the Canny operator; a high object boundary location is achieved. Experimental results demonstrate that our proposed technique can successfully extract the physical object from video sequences. Thomas Meier 0001, King Ngi Ngan |
ICIP (2) | 2 |
| 1998 | Disparity map coding based on adaptive triangular surface modelling
Hongyin Fan, King Ngi Ngan |
Signal Process. Image Commun. | 2 |
| 1998 | Automatic segmentation of moving objects for video object plane generationabstractThe new video coding standard MPEG-4 is enabling content-based functionalities. It takes advantage of a prior decomposition of sequences into video object planes (VOPs) so that each VOP represents one moving object. A comprehensive review summarizes some of the most important motion segmentation and VOP generation techniques that have been proposed. Then, a new automatic video sequence segmentation algorithm that extracts moving objects is presented. The core of this algorithm is an object tracker that matches a two-dimensional (2-D) binary model of the object against subsequent frames using the Hausdorff distance. The best match found indicates the translation the object has undergone, and the model is updated every frame to accommodate for rotation and changes in shape. The initial model is derived automatically, and a new model update method based on the concept of moving connected components allows for comparatively large changes in shape. The proposed algorithm is improved by a filtering technique that removes stationary background. Finally, the binary model sequence guides the extraction objects of the VOPs from the sequence. Experimental results demonstrate the performance of our algorithm. Thomas Meier 0001, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1998 | Guest Editorial
King Ngi Ngan, Sethuraman Panchanathan, Thomas Sikora, Ming-Ting Sun |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1998 | Special Issue On Representation And Coding Of Images And Video I [Guest Editorial]
King Ngi Ngan, Sethuraman Panchanathan, Thomas Sikora, Ming-Ting Sun |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1997 | A Robust Markovian Segmentation Based on Highest Confidence First (HCF)abstractA new robust method to segment images based on Markov random fields (MRF) is presented. The algorithm does not require the number of classes or regions K as input, which is normally difficult to determine in advance. There is also no need for an initial estimate obtained by an algorithm such as K-means. Further, each region is connected during the whole segmentation process leading to more reliable estimates of the regions' mean gray levels and to fewer wrong detected boundaries. In addition, a novel way to incorporate edge information into the segmentation process is proposed resulting in a better detection of small objects. Experimental results demonstrate the performance of our technique. Thomas Meier 0001, King Ngi Ngan, Gregory A. Crebbin |
ICIP (1) | 2 |
| 1997 | A Rate-Distortion Function for Vector Quantization with a Variable Block-Size Classification Model
Michael H. Lee, King Ngi Ngan, Gregory A. Crebbin |
J. Vis. Commun. Image Represent. | 2 |
| 1996 | Very low bit rate video coding using H.263 coderabstractRecently, ITU-T established a draft standard, H.263 to he used as a possible standard for very low bit rate video coding at less than 64 kb/s. It is based on the ITU-T H.261 standard for videoconferencing. This paper provides an analysis of the performance of H.263 coder with the aim of establishing its operational parameters and limitations when operating at very low bit rates. Its performance is then optimized by selecting the desired set of operational parameters. The H.263 coder suffers from "blocking" effects when operating at very low bit rates. To improve the subjective quality of the coded images, the discrete cosine transform (DCT) coefficients are weighted using a visual quantization matrix before quantization. Simulation results showed that the proposed method increases the peak signal-to-noise ratio (PSNR) values at all bit rates tested, but more importantly, the perceived visual quality is improved significantly by reducing the annoying "blocking" effects. King Ngi Ngan, Douglas Chai, Andrew Millin |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1996 | Classified perceptual coding with adaptive quantizationabstractA new technique of adaptively classifying the scene content of an image block has been developed in the proposed perceptual coder. It measures the texture masking energy of an image block and classifies it into one of four perceptual classes: flat, edge, texture, and fine-texture. Each class has an associated factor to adapt the quantizer with the aim of achieving constant quality across an image. A second feature of the perceptual coder is the visual thresholding, a process that reduces bit rate by discarding subthreshold discrete cosine transform (DCT) coefficients without degrading the image perceived quality. Finally, further quality gain is achieved by an improved reference model 8 (RM8) intramode decision, which removes sticking noise artifacts from newly uncovered background found in H.261 coded sequences. Subjective viewing tests, guided by Rec. 500-5, were conducted with 30 subjects. Subjective results confirm the efficacy of the proposed classified coder over the RMS based H.261 coder in two ways: (i) it consistently produces better quality sequences (with a mean opinion score, MOS, of approximately 2.0) when comparing at any fix bit rate; and (ii) it achieves a bit rate saving of 35% when measuring at the same picture quality (i.e., same MOS). Soon Hie Tan, Khee K. Pang, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 1994 | 3D subband coder for very low bit ratesabstractA video coder based on 3D subband coding system (SBC) targetted for very low bit rate applications is developed and simulated. It employs QMF subband analysis to decompose an input sequence into 4 temporal bands. A low-complexity block-based motion detection scheme is then applied to all bands followed by motion classification and reduced motion search on the baseband which is H.261-like coded and the high temporal bands are vector-quantised (VQ) with a bank of codebooks (switched codebook VQ). Simulations are performed at 9.6 kbit/sand 14.4 kbit/s at 5 frame/s for QCIF format and results showed that it achieved better performance than the ITU-T short-term reference model SlM3 coder at similar bit rates. Subjective evaluation favours the 3D SBC coder because it does not suffer from 'blocking' artifacts.> Weng Leong Chooi, King Ngi Ngan |
ICASSP (5) | 2 |
| 1994 | Very low bit rate video coding using 3D subband approachabstractA video coder for very low bit rate applications is designed and simulated. It employs subband analysis technique to first split the video signal into four temporal subbands. Motion detection is performed to classify the co-sited macroblocks of the subbands into temporal activity (TA) macroblocks or no temporal activity (NTA) macroblocks. For the TA macroblocks, motion classification is carried out by further dividing the four temporal subbands into 16 spatiotemporal subbands. The base temporal subband (LL/sub t/) is coded using a H.261-like (European COST 211 ter SIM3) coder with the motion estimation replaced by a reduced-motion search algorithm. The higher temporal subbands (LH/sub t/, HL/sub t/, HH/sub t/) are vector-quantized. The NTA-macroblocks (except in the baseband) are not coded at all. Simulation results are presented at 9.6 kbits per second (kbps) at five frames per second (fps) for a typical videoconferencing sequence. Comparisons are made with results obtained by the SIM3 coder.> King Ngi Ngan, Weng Leong Chooi |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1994 | A frequency scalable coding scheme employing pyramid and subband techniquesabstractA flexible layered coding scheme using block-based DCT domain decimation has been simulated in software. A conditional replacement scheme which switches adaptively between the pyramid and subband approaches at the coefficient level was studied. Experimental results comparing the adaptive scheme against the pyramid approach show that this new approach increases the overall coding efficiency of the scalable coder by up to 1 dB.> Thiow K. Tan, Khee K. Pang, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 1994 | Cell-loss concealment techniques for layered video codecs in an ATM networkabstractA layered video coding scheme with its inherent cell loss resilience has been considered as a means for transporting reliably integrated video services over an asynchronous transfer mode (ATM) based network such as the broadband-ISDN. This paper presents some data concealment techniques that can be implemented in the coding of video data at the encoder, in the ATM adaptation layer (AAL) functionality of network realization and at the decoder to improve the performance of a layered codec under different conditions of video packet loss. The performance of these techniques are verified by software simulations. Liem H. Kieu, King Ngi Ngan |
IEEE Trans. Image Process. | 2 |
| 1993 | Layered video coder with self-concealed capability using frequency scanning techniqueabstractIn this paper, the frequency scanning and the Modified Universal Variable Length Coding (MUVLC) technique is examined as a means to improve the cell loss resilience of video codecs. An appropriate implementation of this technique for effective use in a pyramid layered coding scheme is described. Simulation results are presented which show the superior performance of this slice based coding technique in comparison to the conventional block scanning and Variable Length Coding (VLC) technique having similar coding efficiency. Liem H. Kieu, King Ngi Ngan |
VCIP | 2 |
| 1993 | Motion analysis in 3D subband coderabstractA block-based motion model capable of determining velocity and orientation of motion of a sequence is developed in this paper. Based on the model, three classes of motion orientation, namely horizontal, vertical and diagonal motion in the velocity range of 1 - 8 pixels/frame, are identified. The classification is valuable in coding whereby a reduced motion search is adequate to obtain the best prediction. A low-complexity motion detection scheme is also considered here which takes advantage of the 3D subband decomposition architecture employed. King Ngi Ngan, Weng Leong Chooi, Khok Khee Pang |
VCIP | 1 |
| 1993 | HDTV coding using hybrid MRVQ/DCTabstractThe authors develop a technique to compress HDTV images to below 135 Mb/s at transparent quality, i.e., visually indistinguishable from the original. This is at least equal to the contribution quality. The proposed scheme consists of a lossy and a lossless coder. The lossy coder combines the advantages of two image coding methods, namely, mean/residue vector quantization (MRVQ) and discrete cosine transform (DCT) coding, and the lossless coder is an arithmetic coder (AC), which has been shown to outperform the commonly used Huffman code . The authors describe the various components of the hybrid HDTV image coding scheme and how they are implemented.> King Ngi Ngan, Kok K. Sin, Hee C. Koh |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1992 | Predictive classified vector quantizationabstractA vector quantization scheme based on the classified vector quantization (CVQ) concept, called predictive classified vector quantization (PCVQ), is presented. Unlike CVQ where the classification information has to be transmitted, PCVQ predicts it, thus saving valuable bit rate. Two classifiers, one operating in the Hadamard domain and the other in the spatial domain, were designed and tested. The classification information was predicted in the spatial domain. The PCVQ schemes achieved bit rate reductions over the CVQ ranging from 20 to 32% for two commonly used color test images while maintaining the same acceptable image quality. Bit rates of 0.70-0.93 bits per pixel (bpp) were obtained depending on the image and PCVQ scheme used. King Ngi Ngan, Hee C. Koh |
IEEE Trans. Image Process. | 1 |
| 1990 | Fuzzy quaternion approach to object recognition incorporating Zernike moment invariantsabstractA novel approach to 3-D object recognition based on fuzzy subset theory is described. This method uses Zernike moment invariants of the silhouette of the unknown object to form a set of fuzzy-weighted quantities called fuzzy quaternions. These are matched against those of known objects at predetermined viewpoints. The determination of the Zernike moment invariants can be faster if the equivalent contour integrals are calculated instead. By employing a novel rho -correction scheme, errors due to the digitization are reduced. To speed up the recognition process, a modified Nelder-Mead simplex method is used. Preliminary results demonstrate the potential of the fuzzy quaternion as a viable basis for discrimination. It is concluded that the primary merits of this approach are the ease of model formation, the simplicity of the recognition scheme, and the speed of object recognition. Its disadvantages include the inability to recognize occluded objects and a poor object recognition rate for high perspective distortion of objects.> King Ngi Ngan, Sing B. Kang |
ICPR (1) | 1 |
| 1989 | Geometric modelling of IC die bonds for inspection
King Ngi Ngan, Sing B. Kang |
Pattern Recognit. Lett. | 1 |
| 1988 | IC wire bond inspection using elliptical model approximationabstractAn algorithm that has been developed to inspect the integrity of wire bonds on an IC die is described. The algorithm uses of an elliptical model to approximate the shape of the bond so that any aberrations from the specified dimensions can be easily identified. Missing bonds and double bonds can also be detected. > King Ngi Ngan, Sing B. Kang |
ICRA | 1 |
| 1986 | A simple fast adaptive array based on a null steering beamformerabstractThis paper describes and analyses a simple adaptive array based on the use of perturbation algorithms on a null steering beamformer proposed by Davies [1]. The convergence time constants and misadjustments of using these algorithms are derived. It is shown that, from choosing the feedback factors appropriately, the time constants can be roughly equalized, resulting in faster convergence behaviour. Chi Chong Ko, Yong Ching Lim, King Ngi Ngan |
ICASSP | 3 |
| 1986 | Two-dimensional transform domain decimation techniquesabstractDecimation is normally carried out in two passes; lowpass filtering and subsampling where the latter is normally performed in the time domain. This paper describes a technique whereby both the operations can be combined in the transform domain. The two-dimensional decimation scheme is first implemented in the discrete Fourier transform domain and then extended to the discrete cosine transform domain. It is further applied to a non-sinusoidal i.e. Hadamard transform domain. The effect of using ideal filter transfer function in transform domain decimation on the quality of the decimated images is investigated. This approach results in greater computational efficiency as the constraint of filter length to meet certain specifications is removed permitting the use of smaller transform block sizes. King Ngi Ngan |
ICASSP | 1 |
| 1982 | Enhancement of PCM and DPCM Images Corrupted by Transmission ErrorsabstractSimple methods for enhancing PCM and DPCM monochrome pictures corrupted by transmission bit errors are described. For PCM signals the enhancement is achieved by operating exclusively on the corrupted image, without recourse to any sideinformation from the transmitter. The DPCM signals have simple protection codewords that increase the bit rate by 12 percent (if majority protection coding is used). Photographs of the original, corrupted, and enhanced images are presented, together with SNR as a function of percentage bit error rate (BER) for the corrupted and enhanced picture. For 0.1 < BER < 1.0, the SNR gains for PCM and DPCM are ≃ 10 and 18 dB, respectively. For BER <0.05 percent there is imperceptible picture degradation. King Ngi Ngan, R. Steele |
IEEE Trans. Commun. | 1 |