EDBT 2026 Demo / reviewers in the wild / expert
Lai-Man Po
dblp:77/4317 · also Lai Man Po
· DBLP profile ↗
125ranked-venue papers
19as first author
31since 2021 · last 2026
0000-0002-5185-1492ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 101 · 16 first-author · 18 since 2021Artificial intelligence and machine learning · 22 · 17 since 2021Systems, architecture and hardware · 4 · 1 first-authorComputer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CSBoRA: A continual learning method for large language models with true orthogonality and reduced forgetting
Lai-Man Po, Farrell Hung, Zhuohan Wang, Haoxuan Wu, Kun Li 0015, Xuyuan Xu, Kwok-Wai Cheung 0002 |
Pattern Recognit. | 2 |
| 2026 | RSTFA: Efficient Training-Free Human-Preference Alignment via Rejection Sampling for Text-to-Image Diffusion ModelsabstractGiven a text-to-image diffusion model pretrained on large-scale text-image pairs, can we align the model with human pReferences without further fine-tuning? In this paper, we analyze the effect of alignment tuning in diffusion models by comparing the diffusion denoising trajectory between base and aligned models. Our findings reveal that alignment tuning primarily affects superficial stylistic aspects during denoising, rather than fundamental content, suggesting superficial alignment behaviors. Based on this discovery, we introduce a novel, training-free alignment approach (RSTFA) that leverages rejection sampling at specific stylistic timesteps, ensuring human preference alignment without fine-tuning or heavy inference overhead. We provide a theoretical analysis and derive a bias bound for our rejection-sampling alignment scheme. Empirically, we show that RSTFA better preserves sample diversity than reinforcement-learning-based tuning methods. Extensive experiments on Pick-a-Pic, COCO, HPD V2, and PartiPrompts show that our method not only achieves superior alignment with human preferences compared to state-of-the-art methods, but also reduces computational demands, establishing efficient, human-centered diffusion model alignment. Hongzheng Yang, Jason Chun Lok Li, Kun Li 0015, Wenao Ma, Mingjie Xu, Yuzhi Zhao, Lai-Man Po |
IEEE Trans. Image Process. | 7 |
| 2025 | AesBiasBench: Evaluating Bias and Alignment in Multimodal Language Models for Personalized Image Aesthetic AssessmentabstractMultimodal Large Language Models (MLLMs) are increasingly applied in Personalized Image Aesthetic Assessment (PIAA) as a scalable alternative to expert evaluations.However, their predictions may reflect subtle biases influenced by demographic factors such as gender, age, and education.In this work, we propose AesBiasBench, a benchmark designed to evaluate MLLMs along two complementary dimensions: (1) stereotype bias, quantified by measuring variations in aesthetic evaluations across demographic groups; and (2) alignment between model outputs and genuine human aesthetic preferences.Our benchmark covers three subtasks (Aesthetic Perception, Assessment, Empathy) and introduces structured metrics (IFD, NRD, AAS) to assess both bias and alignment.We evaluate 19 MLLMs, including proprietary models (e.g., GPT-4o, Claude-3.5-Sonnet)and open-source models (e.g., InternVL-2.5, Qwen2.5-VL).Results indicate that smaller models exhibit stronger stereotype biases, whereas larger models align more closely with human preferences.Incorporating identity information often exacerbates bias, particularly in emotional judgments.These findings underscore the importance of identity-aware evaluation frameworks in subjective vision-language tasks. Kun Li 0015, Lai-Man Po, Hongzheng Yang, Xuyuan Xu, Kangcheng Liu, Yuzhi Zhao |
EMNLP | 2 |
| 2025 | Comprehensive regional guidance for attention map semantics in text-to-image diffusion models
Haoxuan Wu, Lai-Man Po, Xuyuan Xu, Kun Li 0015 |
Comput. Vis. Image Underst. | 2 |
| 2025 | RMP-adapter: A region-based Multiple Prompt Adapter for multi-concept customization in text-to-image diffusion modelabstractThis paper introduces a novel framework for multi-concept customization in text-to-image diffusion models . At its core is a Multiple Prompt Adapter (MP-Adapter) capable of processing multiple image prompts in parallel, extracting features from target concepts and projecting them into the same latent space as the text prompt. This enables simultaneous handling of multiple concepts using just one reference image per concept. To address challenges in fusing multiple concepts with complex interactions, we propose a Region-based Denoising Framework (RDF) that dynamically generates concept-specific regions of interest during inference, allowing spatially decoupled injection of concept features. By integrating the MP-Adapter and RDF, our end-to-end pipeline enables multi-concept customization with intricate occlusions and interactions while preserving concept identities. This approach surpasses current methods by resolving concept conflicts, identity degradation, and occlusion issues, allowing flexible customization without concept-specific retraining. Both qualitative and quantitative evaluations demonstrate that our framework outperforms state-of-the-art approaches in multi-concept customization tasks, while ablation studies validate the effectiveness of each proposed component. This work significantly advances text-to-image generation capabilities for complex, user-defined concept combinations. Code and models will be released at https://github.com/baojudezeze/RMP-Adapter . Lai-Man Po, Xuyuan Xu, Yexin Wang, Haoxuan Wu, Kun Li 0015 |
Expert Syst. Appl. | 2 |
| 2025 | Multi-SBoRA: regional and non-overlapping weight updates for multi-concept customization of diffusion models
Haoxuan Wu, Lai-Man Po, Wing Yin Yu, Kun Li 0015 |
Multim. Syst. | 2 |
| 2025 | Modeling Dual-Exposure Quad-Bayer Patterns for Joint Denoising and DeblurringabstractImage degradation caused by noise and blur remains a persistent challenge in imaging systems, stemming from limitations in both hardware and methodology. Single-image solutions face an inherent tradeoff between noise reduction and motion blur. While short exposures can capture clear motion, they suffer from noise amplification. Long exposures reduce noise but introduce blur. Learning-based single-image enhancers tend to be over-smooth due to the limited information. Multi-image solutions using burst mode avoid this tradeoff by capturing more spatial-temporal information but often struggle with misalignment from camera/scene motion. To address these limitations, we propose a physical-model-based image restoration approach leveraging a novel dual-exposure Quad-Bayer pattern sensor. By capturing pairs of short and long exposures at the same starting point but with varying durations, this method integrates complementary noise-blur information within a single image. We further introduce a Quad-Bayer synthesis method (B2QB) to simulate sensor data from Bayer patterns to facilitate training. Based on this dual-exposure sensor model, we design a hierarchical convolutional neural network called QRNet to recover high-quality RGB images. The network incorporates input enhancement blocks and multi-level feature extraction to improve restoration quality. Experiments demonstrate superior performance over state-of-the-art deblurring and denoising methods on both synthetic and real-world datasets. The code, model, and datasets are publicly available at https://github.com/zhaoyuzhi/QRNet. Yuzhi Zhao, Lai-Man Po, Yongzhe Xu, Qiong Yan |
IEEE Trans. Image Process. | 2 |
| 2024 | FedRepOpt: Gradient Re-parametrized Optimizers in Federated Learning
Kin Wai Lau, Yasar Abbas Ur Rehman, Pedro Porto Buarque de Gusmão, Lai-Man Po |
ACCV (8) | 4 |
| 2024 | Motion Transfer-Driven Intra-Class Data Augmentation for Finger Vein RecognitionabstractFinger vein recognition (FVR) has emerged as a secure biometric technique because of the confidentiality of vascular bio-information. Recently, deep learning-based FVR has gained increased popularity and achieved promising performance. However, the limited size of public vein datasets has caused overfitting issues and greatly limits the recognition performance. Although traditional data augmentation can partially alleviate this data shortage issue, it cannot capture the real finger posture variations due to the rigid label-preserving image transformations, bringing limited performance improvement. To address this issue, we propose a novel motion transfer (MT) model for finger vein image data augmentation via modeling the actual finger posture and rotational movements. The proposed model first utilizes a key point detector to extract the key point and pose map of the source and drive finger vein images. We then utilize a dense motion module to estimate the motion optical flow, which is fed to an image generation module for generating the image with the target pose. Experiments conducted on three public finger vein databases demonstrate that the proposed motion transfer model can generate realistic intra-class augmented samples and effectively improve the recognition accuracy. Xiu-Feng Huang, Lai-Man Po, Weifeng Ou |
ICASSP | 2 |
| 2024 | Joint Learning of Identity and Vein Features for Enhanced Representations in Vascular BiometricsabstractVascular biometrics have shown great promise for secure authentication applications and have received increased attention in recent years. This paper proposes a novel framework for joint identity and segmentation feature learning to enrich representations and improve verification performance. The framework utilizes an encoder-decoder architecture, where the encoder is trained under metric learning supervision to extract discriminative identity features. Concurrently, the decoder is trained with vein mask segmentation supervision to extract vein pattern features. By jointly learning high-level identity features and low-level vein features in an end-to-end manner, the representations are enriched. We further design a bi-feature matching scheme utilizing score fusion to integrate both features for identity verification. Experiments conducted on public finger and palm vein datasets reveal that the proposed approach significantly improves verification accuracy, while introducing reasonable complexity overhead. Weifeng Ou, Lai-Man Po, Xiu-Feng Huang |
ICASSP | 2 |
| 2024 | Large Separable Kernel Attention: Rethinking the Large Kernel Attention design in CNN
Kin Wai Lau, Lai-Man Po, Yasar Abbas Ur Rehman |
Expert Syst. Appl. | 2 |
| 2024 | AudioRepInceptionNeXt: A lightweight single-stream architecture for efficient audio recognition
Kin Wai Lau, Yasar Abbas Ur Rehman, Lai-Man Po |
Neurocomputing | 3 |
| 2024 | Self-Calibration Flow Guided Denoising Diffusion Model for Human Pose TransferabstractThe human pose transfer task aims to generate synthetic person images that preserve the style of reference images while accurately aligning them with the desired target pose. However, existing methods based on generative adversarial networks (GANs) struggle to produce realistic details and often face spatial misalignment issues. On the other hand, methods relying on denoising diffusion models require a large number of model parameters, resulting in slower convergence rates. To address these challenges, we propose a self-calibration flow-guided module (SCFM) to establish precise spatial correspondence between reference images and target poses. This module facilitates the denoising diffusion model in predicting the noise at each denoising step more effectively. Additionally, we introduce a multi-scale feature fusing module (MSFF) that enhances the denoising U-Net architecture through a cross-attention mechanism, achieving better performance with a reduced parameter count. Our proposed model outperforms state-of-the-art methods on the DeepFashion and Market-1501 datasets in terms of both the quantity and quality of the synthesized images. Our code is publicly available at https://github.com/zylwithxy/SCFM-guided-DDPM. Lai-Man Po, Wing Yin Yu, Haoxuan Wu, Xuyuan Xu, Kun Li 0015 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | HSGAN: Hyperspectral Reconstruction From RGB Images With Generative Adversarial NetworkabstractHyperspectral (HS) reconstruction from RGB images denotes the recovery of whole-scene HS information, which has attracted much attention recently. State-of-the-art approaches often adopt convolutional neural networks to learn the mapping for HS reconstruction from RGB images. However, they often do not achieve high HS reconstruction performance across different scenes consistently. In addition, their performance in recovering HS images from clean and real-world noisy RGB images is not consistent. To improve the HS reconstruction accuracy and robustness across different scenes and from different input images, we present an effective HSGAN framework with a two-stage adversarial training strategy. The generator is a four-level top-down architecture that extracts and combines features on multiple scales. To generalize well to real-world noisy images, we further propose a spatial-spectral attention block (SSAB) to learn both spatial-wise and channel-wise relations. We conduct the HS reconstruction experiments from both clean and real-world noisy RGB images on five well-known HS datasets. The results demonstrate that HSGAN achieves superior performance to existing methods. Please visit https://github.com/zhaoyuzhi/HSGAN to try our codes. Yuzhi Zhao, Lai-Man Po, Tingyu Lin 0002, Qiong Yan, Wei Liu 0004, Pengfei Xian |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Bidirectionally Deformable Motion Modulation For Video-based Human Pose TransferabstractVideo-based human pose transfer is a video-to-video generation task that animates a plain source human image based on a series of target human poses. Considering the difficulties in transferring highly structural patterns on the garments and discontinuous poses, existing methods often generate unsatisfactory results such as distorted textures and flickering artifacts. To address these issues, we propose a novel Deformable Motion Modulation (DMM) that utilizes geometric kernel offset with adaptive weight modulation to simultaneously perform feature alignment and style transfer. Different from normal style modulation used in style transfer, the proposed modulation mechanism adaptively reconstructs smoothed frames from style codes according to the object shape through an irregular receptive field of view. To enhance the spatio-temporal consistency, we leverage bidirectional propagation to extract the hidden motion information from a warped image sequence generated by noisy poses. The proposed feature propagation significantly enhances the motion prediction ability by forward and backward propagation. Both quantitative and qualitative experimental results demonstrate superiority over the state-of-the-arts in terms of image fidelity and visual continuity. The source code is publicly available at github.com/rocketappslab/bdmm. Wing Yin Yu, Lai-Man Po, Ray C. C. Cheung, Yuzhi Zhao, Kun Li 0015 |
ICCV | 2 |
| 2023 | CSRNet: Cascaded Selective Resolution Network for real-time semantic segmentation
Jingjing Xiong, Lai-Man Po, Wing Yin Yu, Chang Zhou 0008, Pengfei Xian, Weifeng Ou |
Expert Syst. Appl. | 2 |
| 2023 | SVCNet: Scribble-Based Video Colorization Network With Temporal AggregationabstractIn this paper, we propose a scribble-based video colorization network with temporal aggregation called SVCNet. It can colorize monochrome videos based on different user-given color scribbles. It addresses three common issues in the scribble-based video colorization area: colorization vividness, temporal consistency, and color bleeding. To improve the colorization quality and strengthen the temporal consistency, we adopt two sequential sub-networks in SVCNet for precise colorization and temporal smoothing, respectively. The first stage includes a pyramid feature encoder to incorporate color scribbles with a grayscale frame, and a semantic feature encoder to extract semantics. The second stage finetunes the output from the first stage by aggregating the information of neighboring colorized frames (as short-range connections) and the first colorized frame (as a long-range connection). To alleviate the color bleeding artifacts, we learn video colorization and segmentation simultaneously. Furthermore, we set the majority of operations on a fixed small image resolution and use a Super-resolution Module at the tail of SVCNet to recover original sizes. It allows the SVCNet to fit different image resolutions at the inference. Finally, we evaluate the proposed SVCNet on DAVIS and Videvo benchmarks. The experimental results demonstrate that SVCNet produces both higher-quality and more temporally consistent videos than other well-known video colorization approaches. The codes and models can be found at https://github.com/zhaoyuzhi/SVCNet. Yuzhi Zhao, Lai-Man Po, Kangcheng Liu, Wing Yin Yu, Pengfei Xian, Yujia Zhang 0002, Mengyang Liu |
IEEE Trans. Image Process. | 2 |
| 2023 | Distortion Map-Guided Feature Rectification for Efficient Video Semantic SegmentationabstractTo leverage the strong cross-frame relations of videos, many video semantic segmentation methods tend to explore feature reuse and feature warping based on motion clues. However, since the video dynamics are too complex to model accurately, some warped feature values may be invalid. Moreover, the warping errors can accumulate across frames, thereby resulting in degraded segmentation performance. To tackle this problem, we present an efficient distortion map-guided feature rectification method for video semantic segmentation, specifically targeting the feature updating and correction on the distorted regions with unreliable optical flow. The updated features for the distorted regions are extracted from a light correction network (CoNet). A distortion map serves as the weighted attention to guide the feature rectification by aggregating the warped features and the updated features. The generation of the distortion map is simple yet effective in predicting the distorted areas in the warped features, i.e., moving boundaries, thin objects, and occlusions. In addition, we propose an auxiliary edge-semantics loss to implement the distorted region supervision with classes. Our network is trained in an end-to-end manner and highly modular. Comprehensive experiments on Cityscapes and CamVid datasets demonstrate that the proposed method has achieved state-of-the-art performance by weighing accuracy, inference speed, and temporal consistency on video semantic segmentation. Jingjing Xiong, Lai-Man Po, Wing Yin Yu, Yuzhi Zhao, Kwok-Wai Cheung 0002 |
IEEE Trans. Multim. | 2 |
| 2023 | ChildPredictor: A Child Face Prediction Framework With Disentangled LearningabstractThe appearances of children are inherited from their parents, which makes it feasible to predict them. Predicting realistic children's faces may help settle many social problems, such as age-invariant face recognition, kinship verification, and missing child identification. It can be regarded as an image-to-image translation task. Existing approaches usually assume domain information in the image-to-image translation can be interpreted by “style”, i.e., the separation of image content and style. However, such separation is improper for the child face prediction, because the facial contours between children and parents are not the same. To address this issue, we propose a new disentangled learning strategy for children's face prediction. We assume that children's faces are determined by genetic factors (compact family features, e.g., face contour), external factors (facial attributes irrelevant to prediction, such as moustaches and glasses), and variety factors (individual properties for each child). On this basis, we formulate predictions as a mapping from parents’ genetic factors to children's genetic factors, and disentangle them from external and variety factors. In order to obtain accurate genetic factors and perform the mapping, we propose a ChildPredictor framework. It transfers human faces to genetic factors by encoders and back by generators. Then, it learns the relationship between the genetic factors of parents and children through a mapping function. To ensure the generated faces are realistic, we collect a large Family Face Database to train ChildPredictor and evaluate it on the FF-Database validation set. Experimental results demonstrate that ChildPredictor is superior to other well-known image-to-image translation methods in predicting realistic and diverse child faces. Implementation codes can be found athttps://github.com/zhaoyuzhi/ChildPredictor. Yuzhi Zhao, Lai-Man Po, Qiong Yan, Wei Shen 0002, Yujia Zhang 0002, Wei Liu 0004, Chun Kit Wong, Chiu-Sing Pang, Weifeng Ou, Wing Yin Yu, Buhua Liu |
IEEE Trans. Multim. | 2 |
| 2023 | VCGAN: Video Colorization With Hybrid Generative Adversarial NetworkabstractWe propose a Video Colorization with Hybrid Generative Adversarial Network (VCGAN), an improved approach to video colorization using end-to-end learning and recurrent architecture. The VCGAN addresses two prevalent issues in the video colorization domain: Temporal consistency and the unification of colorization network and refinement network into a single architecture. To enhance colorization quality and spatiotemporal consistency, the mainstream of the generator in VCGAN is assisted by two additional networks,i.e.,global feature extractor and placeholder feature extractor, respectively. The global feature extractor encodes the global semantics of grayscale input to enhance colorization quality, whereas the placeholder feature extractor serves as a feedback connection to encode the semantics of the previous colorized frame in order to maintain spatiotemporal consistency. If changing the input for placeholder feature extractor as grayscale input, the hybrid VCGAN also has the potential to colorize single images. To improve the color consistency of far frames, we propose a dense long-term loss that minimizes the temporal disparity of every two remote frames. Trained with colorization and temporal losses jointly, VCGAN strikes a good balance between video color vividness and spatiotemporal continuity. Experimental results demonstrate that VCGAN produces higher-quality and temporally more consistent colorful videos than existing approaches. Yuzhi Zhao, Lai-Man Po, Wing Yin Yu, Yasar Abbas Ur Rehman, Mengyang Liu, Yujia Zhang 0002, Weifeng Ou |
IEEE Trans. Multim. | 2 |
| 2022 | Contrastive Spatio-Temporal Pretext Learning for Self-Supervised Video RepresentationabstractSpatio-temporal representation learning is critical for video self-supervised representation. Recent approaches mainly use contrastive learning and pretext tasks. However, these approaches learn representation by discriminating sampled instances via feature similarity in the latent space while ignoring the intermediate state of the learned representations, which limits the overall performance. In this work, taking into account the degree of similarity of sampled instances as the intermediate state, we propose a novel pretext task - spatio-temporal overlap rate (STOR) prediction. It stems from the observation that humans are capable of discriminating the overlap rates of videos in space and time. This task encourages the model to discriminate the STOR of two generated samples to learn the representations. Moreover, we employ a joint optimization combining pretext tasks with contrastive learning to further enhance the spatio-temporal representation learning. We also study the mutual influence of each component in the proposed scheme. Extensive experiments demonstrate that our proposed STOR task can favor both contrastive learning and pretext tasks and the joint optimization scheme can significantly improve the spatio-temporal representation in video understanding. The code is available at https://github.com/Katou2/CSTP. Yujia Zhang 0002, Lai-Man Po, Xuyuan Xu, Mengyang Liu, Yexin Wang, Weifeng Ou, Yuzhi Zhao, Wing Yin Yu |
AAAI | 2 |
| 2022 | D2HNet: Joint Denoising and Deblurring with Hierarchical Network for Robust Night Image Restoration
Yuzhi Zhao, Yongzhe Xu, Qiong Yan, Dingdong Yang, Lai-Man Po |
ECCV (7) | 6 |
| 2022 | Pixel Voting Decoder: A novel decoder that regresses pixel relationships for segmentation
Pengfei Xian, Lai-Man Po, Jingjing Xiong, Chang Zhou 0008, Yuzhi Zhao, Wing Yin Yu, Weifeng Ou, Yujia Zhang 0002, Xiaori Zhang |
Expert Syst. Appl. | 2 |
| 2022 | ShaTure: Shape and Texture Deformation for Human Pose and Attribute TransferabstractIn this paper, we present a novel end-to-end pose transfer framework to transform a source person image to an arbitrary pose with controllable attributes. Due to the spatial misalignment caused by occlusions and multi-viewpoints, maintaining high-quality shape and texture appearance is still a challenging problem for pose-guided person image synthesis. Without considering the deformation of shape and texture, existing solutions on controllable pose transfer still cannot generate high-fidelity texture for the target image. To solve this problem, we design a new image reconstruction decoder - ShaTure which formulates shape and texture in a braiding manner. It can interchange discriminative features in both feature-level space and pixel-level space so that the shape and texture can be mutually fine-tuned. In addition, we develop a new bottleneck module - Adaptive Style Selector (AdaSS) Module which can enhance the multi-scale feature extraction capability by self-recalibration of the feature map through channel-wise attention. Both quantitative and qualitative results show that the proposed framework has superiority compared with the state-of-the-art human pose and attribute transfer methods. Detailed ablation studies report the effectiveness of each contribution, which proves the robustness and efficacy of the proposed framework. Wing Yin Yu, Lai-Man Po, Jingjing Xiong, Yuzhi Zhao, Pengfei Xian |
IEEE Trans. Image Process. | 2 |
| 2022 | Angular Deep Supervised Vector Quantization for Image RetrievalabstractMost of the deep quantization methods adopt unsupervised approaches, and the quantization process usually occurs in the Euclidean space on top of the deep feature and its approximate value. When this approach is applied to the retrieval tasks, since the internal product space of the retrieval process is different from the Euclidean space of quantization, minimizing the quantization error (QE) does not necessarily lead to a good performance on the maximum inner product search (MIPS). To solve these problems, we treat Softmax classification as vector quantization (VQ) with angular decision boundaries and propose angular deep supervised VQ (ADSVQ) for image retrieval. Our approach can simultaneously learn the discriminative feature representation and the updatable codebook, both lying on a hypersphere. To reduce the QE between centroids and deep features, two regularization terms are proposed as supervision signals to encourage the intra-class compactness and inter-class balance, respectively. ADSVQ explicitly reformulates the asymmetric distance computation in MIPS to transform the image retrieval process into a two-stage classification process. Moreover, we discuss the extension of multiple-label cases from the perspective of quantization with binary classification. Extensive experiments demonstrate that the proposed ADSVQ has excellent performance on four well-known image data sets when compared with the state-of-the-art hashing methods. Chang Zhou 0008, Lai-Man Po, Weifeng Ou |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Spatial Content Alignment for Pose TransferabstractDue to unreliable geometric matching and content misalignment, most conventional pose transfer algorithms fail to generate fine-trained person images. In this paper, we propose a novel framework – Spatial Content Alignment GAN (SCA-GAN) which aims to enhance the content consistency of garment textures and the details of human characteristics. We first alleviate the spatial misalignment by transferring the edge content to the target pose in advance. Secondly, we introduce a new Content-Style DeBlk which can progressively synthesize photo-realistic person images based on the appearance features of the source image, the target pose heatmap and the prior transferred content in edge domain. We compare the proposed framework with several state-of-the-art methods to show its superiority in quantitative and qualitative analysis. Moreover, detailed ablation study results demonstrate the efficacy of our contributions. Codes are publicly available at github.com/rocketappslab/SCA-GAN. Wing Yin Yu, Lai-Man Po, Yuzhi Zhao, Jingjing Xiong, Kin Wai Lau |
ICME | 2 |
| 2021 | Legacy Photo Editing with Learned Noise PriorabstractThere are quite a number of photographs captured under undesirable conditions in the last century. Thus, they are of-ten noisy, regionally incomplete, and grayscale formatted. Conventional approaches mainly focus on one point so that those restoration results are not perceptually sharp or clean enough. To solve these problems, we propose a noise prior learner NEGAN to simulate the noise distribution of real legacy photos using unpaired images. It mainly focuses on matching high-frequency parts of noisy images through discrete wavelet transform (DWT) since they include most of noise statistics. We also create a large legacy photo dataset for learning noise prior. Using learned noise prior, we can easily build valid training pairs by degrading clean images. Then, we propose an IEGAN framework performing image editing including joint denoising, inpainting and colorization based on the estimated noise prior. We evaluate the proposed system and compare it with state-of-the-art image enhancement methods. The experimental results demonstrate that it achieves the best perceptual quality. Please see the webpage https://github.com/zhaoyuzhi/Legacy-Photo-Editing-with-Learned-Noise-Prior for the codes and the proposed LP dataset. Yuzhi Zhao, Lai-Man Po, Tingyu Lin 0002, Kangcheng Liu, Yujia Zhang 0002, Wing Yin Yu, Pengfei Xian, Jingjing Xiong |
WACV | 2 |
| 2021 | Fusion loss and inter-class data augmentation for deep finger vein feature learning
Weifeng Ou, Lai-Man Po, Chang Zhou 0008, Yasar Abbas Ur Rehman, Pengfei Xian, Yujia Zhang 0002 |
Expert Syst. Appl. | 2 |
| 2021 | Deep triplet residual quantization
Chang Zhou 0008, Lai-Man Po, Weifeng Ou, Pengfei Xian, Kwok-Wai Cheung 0002 |
Expert Syst. Appl. | 2 |
| 2021 | FEANet: Foreground-edge-aware network with DenseASPOC for human parsing
Wing Yin Yu, Lai-Man Po, Yuzhi Zhao, Yujia Zhang 0002, Kin Wai Lau |
Image Vis. Comput. | 2 |
| 2021 | SCGAN: Saliency Map-Guided Colorization With Generative Adversarial NetworkabstractGiven a grayscale photograph, the colorization system estimates a visually plausible colorful image. Conventional methods often use semantics to colorize grayscale images. However, in these methods, only classification semantic information is embedded, resulting in semantic confusion and color bleeding in the final colorized image. To address these issues, we propose a fully automatic Saliency Map-guided Colorization with Generative Adversarial Network (SCGAN) framework. It jointly predicts the colorization and saliency map to minimize semantic confusion and color bleeding in the colorized image. Since the global features from pre-trained VGG-16-Gray network are embedded to the colorization encoder, the proposed SCGAN can be trained with much less data than state-of-the-art methods to achieve perceptually reasonable colorization. In addition, we propose a novel saliency map-based guidance method. Branches of the colorization decoder are used to predict the saliency map as a proxy target. Moreover, two hierarchical discriminators are utilized for the generated colorization and saliency map, respectively, in order to strengthen visual perception performance. The proposed system is evaluated on ImageNet validation set. Experimental results show that SCGAN can generate more reasonable colorized images than state-of-the-art techniques. Yuzhi Zhao, Lai-Man Po, Kwok-Wai Cheung 0002, Wing Yin Yu, Yasar Abbas Ur Rehman |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | SLNet: Stereo face liveness detection via dynamic disparity-maps and convolutional neural network
Yasar Abbas Ur Rehman, Lai-Man Po, Mengyang Liu |
Expert Syst. Appl. | 2 |
| 2020 | Data-level information enhancement: Motion-patch-based Siamese Convolutional Neural Networks for human activity recognition in videos
Yujia Zhang 0002, Lai-Man Po, Mengyang Liu, Yasar Abbas Ur Rehman, Weifeng Ou, Yuzhi Zhao |
Expert Syst. Appl. | 2 |
| 2020 | Enhancing deep discriminative feature maps via perturbation for face presentation attack detection
Yasar Abbas Ur Rehman, Lai-Man Po, Jukka Komulainen |
Image Vis. Comput. | 2 |
| 2019 | Deep Hashing with Triplet Labels and Unification Binary Code Selection for Fast Image Retrieval
Chang Zhou 0008, Lai-Man Po, Mengyang Liu, Wilson Y. F. Yuen, Peter Hon-Wah Wong, Hon-Tung Luk, Kin Wai Lau, Hok Kwan Cheung |
MMM (1) | 2 |
| 2019 | Face liveness detection using convolutional-features fusion of real and deep network generated face images
Yasar Abbas Ur Rehman, Lai-Man Po, Mengyang Liu, Zijie Zou, Weifeng Ou, Yuzhi Zhao |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | Video copy detection by conducting fast searching of inverted files
Mengyang Liu, Lai-Man Po, Yasar Abbas Ur Rehman, Xuyuan Xu, Litong Feng |
Multim. Tools Appl. | 2 |
| 2019 | A Novel Patch Variance Biased Convolutional Neural Network for No-Reference Image Quality AssessmentabstractDeep convolutional neural networks (CNNs) have been successfully applied on no-reference image quality assessment (NR-IQA) with respect to human perception. Most of these methods deal with small image patches and use the average score of the test patches for predicting the whole image quality. We discovered that image patches from homogenous regions are unreliable for both neural network training and final image quality score estimation. In addition, image patches with complex structures have much higher chances of achieving better image quality prediction. Based on these findings, we enhanced the conventional CNN-based NR-IQA algorithm to avoid homogenous patches for the network training and quality score estimation. Moreover, we also use a variance-based weighting average to bias the final image quality score to the patches with complex structure. The experimental results show that this simple approach can achieve state-of-the-art performance compared with well-known NR-IQA algorithms. Lai-Man Po, Mengyang Liu, Wilson Y. F. Yuen, Xuyuan Xu, Chang Zhou 0008, Peter Hon-Wah Wong, Kin Wai Lau, Hon-Tung Luk |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | LiveNet: Improving features generalization for face liveness detection using convolution neural networks
Yasar Abbas Ur Rehman, Lai-Man Po, Mengyang Liu |
Expert Syst. Appl. | 2 |
| 2018 | Block-based adaptive ROI for remote photoplethysmography
Lai-Man Po, Litong Feng, Xuyuan Xu, Chun-Ho Cheung, Kwok-Wai Cheung 0002 |
Multim. Tools Appl. | 1 |
| 2016 | Face liveness detection and recognition using shearlet based feature descriptorsabstractFace recognition is a widely used biometric technology due to its convenience but it is vulnerable to spoofing attacks made by non-real faces such as a photograph or video of valid user. Face liveness detection is a core technology to make sure that the input face is a live person. However, this is still very challenging using conventional liveness detection approaches of texture analysis and motion detection. The aim of this paper is to develop a multifunctional feature descriptor and an efficient framework which can be used to deal with both face liveness detection and recognition. In this framework, new feature descriptors are defined using a multiscale directional transform (shearlet transform). Then, stacked autoencoders and softmax classifier are concatenated to detect face liveness and identify person. We evaluated this approach using CASIA Face Anti-Spoofing Database and the results show that our approach performs better than state-of-the-art techniques following the provided evaluation protocols of this database, and is possible to significantly enhance the security of face recognition biometric system. Lai-Man Po, Xuyuan Xu, Litong Feng |
ICASSP | 2 |
| 2016 | Integration of image quality and motion cues for face anti-spoofing: A neural network approach
Litong Feng, Lai-Man Po, Xuyuan Xu, Chun-Ho Cheung, Kwok-Wai Cheung 0002 |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | No-Reference Video Quality Assessment With 3D Shearlet Transform and Convolutional Neural NetworksabstractIn this paper, we propose an efficient general-purpose no-reference (NR) video quality assessment (VQA) framework that is based on 3D shearlet transform and convolutional neural network (CNN). Taking video blocks as input, simple and efficient primary spatiotemporal features are extracted by 3D shearlet transform, which are capable of capturing natural scene statistics properties. Then, CNN and logistic regression are concatenated to exaggerate the discriminative parts of the primary features and predict a perceptual quality score. The resulting algorithm, which we name shearlet- and CNN-based NR VQA (SACONVA), is tested on well-known VQA databases of Laboratory for Image & Video Engineering, Image & Video Processing Laboratory, and CSIQ. The testing results have demonstrated that SACONVA performs well in predicting video quality and is competitive with current state-of-the-art full-reference VQA methods and general-purpose NR-VQA algorithms. Besides, SACONVA is extended to classify different video distortion types in these three databases and achieves excellent classification accuracy. In addition, we also demonstrate that SACONVA can be directly applied in real applications such as blind video denoising. Lai-Man Po, Chun-Ho Cheung, Xuyuan Xu, Litong Feng, Kwok-Wai Cheung 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | Dynamic ROI based on K-means for remote photoplethysmographyabstractRemote imaging photoplethysmography (RIPPG) can achieve contactless human vital signs monitoring. Though the remote operation mode brings a great convenience for RIPPG applications, the RIPPG signal quality is limited by the remote nature. Improving the RIPPG signal quality becomes an essential task in the clinical application of RIPPG. Since the region of interest (ROI) of the RIPPG transforms from a point to an area, there is a new approach to improving the RIPPG signal quality through refining the ROI. In this paper, we propose a dynamic ROI for RIPPG, which can automatically select the skin regions corresponding to good quality RIPPG signals. First, a fixed ROI is divided into non-overlapped blocks. Then two features are proposed to perform no-reference quality assessment for RIPPG signals from different blocks. After that, K-means clustering operates in a two dimensional feature space. A dynamic ROI can be selected for a video segment based on the clustering result, updated every two seconds. Nineteen healthy subjects were enrolled to test the proposed ROI selection method on both the facial region and the palmar region. Experimental results of heart rate measurement show that the proposed dynamic ROI method for RIPPG can effectively improve the RIPPG signal quality, compared with the state-of-the-art ROI methods for RIPPG. Litong Feng, Lai-Man Po, Xuyuan Xu, Chun-Ho Cheung, Kwok-Wai Cheung 0002 |
ICASSP | 2 |
| 2015 | No-reference image quality assessment using shearlet transform and stacked autoencodersabstractIn this work, we describe an efficient generalpurpose no-reference (NR) image quality assessment (IQA) algorithm that is based on a new multiscale directional transform (shearlet transform) with a strong ability to localize distributed discontinuities. The algorithm relies on utilizing the sum of subband coefficient amplitudes (SSCA) as primary features to describe the behavior of natural images and distorted images. Then, stacked autoencoders are applied to exaggerate the discriminative parts of the primary features. Finally, by translating the NR-IQA problem into classification problem, the differences of evolved features are identified by softmax classifier. The resulting algorithm, which we name SESANIA (ShEarlet and Stacked Autoencoders based No-reference Image quality Assessment), is tested on several databases (LIVE, Multiply Distorted LIVE and TID2008) and shown to be suitable to many common distortions, consistent with subjective assessment and comparable to full-reference IQA methods and state-of-the-art general purpose NR-IQA algorithms. Lai-Man Po, Xuyuan Xu, Litong Feng, Chun-Ho Cheung, Kwok-Wai Cheung 0002 |
ISCAS | 2 |
| 2015 | Frame adaptive ROI for photoplethysmography signal extraction from fingertip video captured by smartphoneabstractPhotoplethysmography (PPG) has been widely used in clinical applications for monitoring vital signs especially heart rate by pulse oximeter. Recent researches have demonstrated the possibility of using fingertip video based PPG approach to estimate heart rate by smartphones. However, due to the variation of camera sensor characteristics in difference smartphones, the conventional fixed region-of-interest (ROI) for PPG signal extraction technique is not reliable. In this paper, a novel frame adaptive ROI method is proposed to detour the color saturation or cut-off distortion in the fingertip video capturing process for improving the reliability due to variation and limited dynamic range of the camera sensors in different smartphone models. Experimental results demonstrate that the proposed method can produce good pulsatile waveform and achieve high heart rate estimation accuracy using different smartphone models as compared with a FDA (U.S. Food and Drug Administration) approved commercial pulse oximeter. Lai-Man Po, Xuyuan Xu, Litong Feng, Kwok-Wai Cheung 0002, Chun-Ho Cheung |
ISCAS | 1 |
| 2015 | No-reference image quality assessment with shearlet transform and deep neural networks
Lai-Man Po, Xuyuan Xu, Litong Feng, Chun-Ho Cheung, Kwok-Wai Cheung 0002 |
Neurocomputing | 2 |
| 2015 | Motion-Resistant Remote Imaging Photoplethysmography Based on the Optical Properties of SkinabstractRemote imaging photoplethysmography (RIPPG) can achieve contactless monitoring of human vital signs. However, the robustness to a subject's motion is a challenging problem for RIPPG, especially in facial video-based RIPPG. The RIPPG signal originates from the radiant intensity variation of human skin with pulses of blood and motions can modulate the radiant intensity of the skin. Based on the optical properties of human skin, we build an optical RIPPG signal model in which the origins of the RIPPG signal and motion artifacts can be clearly described. The region of interest (ROI) of the skin is regarded as a Lambertian radiator and the effect of ROI tracking is analyzed from the perspective of radiometry. By considering a digital color camera as a simple spectrometer, we propose an adaptive color difference operation between the green and red channels to reduce motion artifacts. Based on the spectral characteristics of photoplethysmography signals, we propose an adaptive bandpass filter to remove residual motion artifacts of RIPPG. We also combine ROI selection on the subject's cheeks with speeded-up robust features points tracking to improve the RIPPG signal quality. Experimental results show that the proposed RIPPG can obtain greatly improved performance in accessing heart rates in moving subjects, compared with the state-of-the-art facial video-based RIPPG methods. Litong Feng, Lai-Man Po, Xuyuan Xu, Ruiyi Ma |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | Adaptive block truncation filter for MVC depth image enhancementabstractIn Multiview Video plus Depth (MVD) format, virtual views are generated from decoded texture videos with decoded depth images through Depth Image based Rendering (DIBR). 3DV-ATM is a reference model for H.264/AVC based Multiview Video Coding (MVC) and aims at achieving high coding efficiency for 3D video in MVD format. Depth images are first downsampled then coded by 3DV-ATM. However, sharp object boundary characteristic of depth images does not well match with the transform coding of 3DV-ATM. Depth boundaries are often blurred with ringing artifacts in the decoded depth images that result in noticeable artifacts in synthesized views. This paper presents a low complexity adaptive block truncation filter to recover the sharp object boundaries of depth images using adaptive block repositioning and expansion for increasing the depth values refinement accuracy. This new approach is very efficient and can avoid false depth boundary refinement when block boundaries lie around the depth edge regions and ensure sufficient information within the processing block for depth layers classification. Experimental results show that sharp depth edges can be recovered using the proposed filter and boundary artifacts in the synthesized views can be removed. The proposed method can provide improvement up to 3.25dB in the depth map enhancement and bitrate reduction of 3.06% in the synthesized views. Xuyuan Xu, Lai-Man Po, Chun-Ho Cheung, Litong Feng, Kwok-Wai Cheung 0002, Chi-Wang Ting, Ka-Ho Ng |
ICASSP | 2 |
| 2014 | An improved hybrid fast mode decision method for H.264/AVC intra coding with local information
Changnian Chen, Jiazhong Chen, Zengwei Ju, Lai-Man Po |
Multim. Tools Appl. | 5 |
| 2014 | No-reference image quality assessment using statistical characterization in the shearlet domain
Lai-Man Po, Xuyuan Xu, Litong Feng |
Signal Process. Image Commun. | 2 |
| 2014 | Adaptive depth truncation filter for MVC based compressed depth image
Xuyuan Xu, Lai-Man Po, Chun-Ho Cheung, Kwok-Wai Cheung 0002, Litong Feng, Chi-Wang Ting, Ka-Ho Ng |
Signal Process. Image Commun. | 2 |
| 2013 | Watershed based depth map misalignment correction and foreground biased dilation for DIBR view synthesisabstractThe quality of the synthesized views by Depth Image Based Rendering (DIBR) highly depends on the accuracy of the depth map, especially the alignment of object boundaries of texture image. In practice, the misalignment of sharp depth map edges is the major cause of the annoying artifacts at the disoccluded regions of the synthesized views. In this paper, a new depth map preprocessing method using Watershed misalignment correction and dilation filter is proposed to align the foreground depth edges to cover the whole transitional color edge regions. This approach can handle the sharp depth map edges lying inside or outside the object boundaries in 2D sense. The quality of the disoccluded regions of the synthesized views can be significantly improved. Experimental results show that the proposed method achieves superior performance for view synthesis by DIBR especially for generating large baseline virtual views. Xuyuan Xu, Lai-Man Po, Kwok-Wai Cheung 0002, Litong Feng, Chun-Ho Cheung |
ICIP | 2 |
| 2013 | An adaptive background biased depth map hole-filling method for KinectabstractThe launch of Kinect provides a convenient way to access the depth information in real time. However the depth map quality still needs to be enhanced for 3D visual applications. In this paper, an adaptive background biased depth map hole-filling method is proposed. First, depth holes caused by abnormal reflection are filled by color similarity in-painting, and a soft decision for color similarity checking is performed by the use of probabilities in random walks color segmentation. Afterwards it is assumed that the lost information in the rest of depth holes belongs to the background. The background depth information is extracted by automatic thresholding in the neighborhood of each hole. Depth holes are in-painted with the background information in their local neighborhood. Combination of color similarity in-painting and background biased in-painting is able to perform depth map hole-filling adaptively for different kinds of depth holes for Kinect. The hole-filling results and virtual view synthesis results show that the Kinect depth map quality can be improved significantly by the proposed method. Litong Feng, Lai-Man Po, Xuyuan Xu, Ka-Ho Ng, Chun-Ho Cheung, Kwok-Wai Cheung 0002 |
IECON | 2 |
| 2013 | Depth-aided exemplar-based hole filling for DIBR view synthesisabstractQuality of synthesized view by Depth-Image-Based Rendering (DIBR) highly depends on hole filling, especially for synthesized view with large disocclusion. Many hole filling methods are proposed to improve the synthesized view quality and inpainting is the most popular approach to recover the disocclusions. However, the conventional inpainting either makes the hole regions blurred via diffusion or propagates the foreground information to the disoclusion regions. Annoying artifacts are created in the synthesized virtual views. This paper proposes a depth-aided exemplar-based inpainting method for recovering large disoclusion. It consists of two processes, warped depth map filling and warped color image filling. Since depth map can be considered as a grey-scale image without texture, it is much easier to be filled. Disoccluded regions of color image are predicted based on its associated filled depth map information. Regions with texture lying around the background have higher priority to be filled than other regions and disoccluded regions are filled by propagating the background texture through the exemplar-based inpainting. Thus artifacts created by diffusion or using foreground information for prediction can be eliminated. Experimental results show texture can be recovered in large disocclusions and the proposed method has better visual quality compared to existing methods. Xuyuan Xu, Lai-Man Po, Chun-Ho Cheung, Litong Feng, Ka-Ho Ng, Kwok-Wai Cheung 0002 |
ISCAS | 2 |
| 2013 | Depth map misalignment correction and dilation for DIBR view synthesis
Xuyuan Xu, Lai-Man Po, Ka-Ho Ng, Litong Feng, Kwok-Wai Cheung 0002, Chun-Ho Cheung, Chi-Wang Ting |
Signal Process. Image Commun. | 2 |
| 2012 | A new motion compensation method using superimposed inter-frame signalsabstractA new MCP method called Neighbor Predicted Superimposed Search (NPSS) algorithm that uses superimposed inter-frame signals to achieve higher prediction accuracy is proposed in this paper. It outperforms other Multi-Hypothesis MCP (MHMCP) methods as it does not require the transmission of multiple motion vectors. The proposed method has better prediction quality and yet having comparable computational complexity as conventional block-based MCP with no extra side-information overhead. Ka-Ho Ng, Lai-Man Po, Kwok-Wai Cheung 0002, Xuyuan Xu, Ka-Man Wong |
ICASSP | 2 |
| 2012 | A foreground biased depth map refinement method for DIBR view synthesisabstractThe performance of view synthesis using depth image based rendering (DIBR) highly depends on the accuracy of depth map. Inaccurate boundary alignment between texture image and depth map especially for large depth discontinuities always cause annoying artifacts in disocclusion regions of the synthesized view. Pre-filtering approach and reliability-based approach have been proposed to tackle this problem. However, pre-filtering approach blurs the depth map with drawback of degradation of the depth map and may also cause distortion in non-hole region. Reliability-based approach uses reliable warping information from other views to fill up holes and is not suitable for the view synthesis with single texture video such as video-plus-depth based DIBR applications. This paper presents a simple and efficient depth map preprocessing method with use of texture edge information to refine depth pixels around the large depth discontinuities. The refined depth map can make the whole texture edge pixels assigned with foreground depth values. It can significantly improve the quality of the synthesized view by avoiding incorrect use of foreground texture information in hole filling. The experimental results show the proposed method achieves superior performance for view synthesis by DIBR especially for large baseline. Xuyuan Xu, Lai-Man Po, Kwok-Wai Cheung 0002, Ka-Ho Ng, Ka-Man Wong, Chi-Wang Ting |
ICASSP | 2 |
| 2012 | Horizontal Scaling and Shearing-Based Disparity-Compensated Prediction for Stereo Video CodingabstractIn multiview video coding (MVC), disparity-compensated prediction (DCP) exploits the correlation among different views. A common approach is to use block-based motion-compensated prediction (MCP) tools to predict the disparity effect among different views. However, some regions in different views may have various deformations due to nonconstant depth. Thus, performance of DCP is not satisfactory with the simple translational model assumed in conventional block-based MCP tools. Previous attempts to achieve better disparity prediction were usually too complex for practical use. In this paper, horizontal scaling and shearing (HSS) effects are investigated to increase interview prediction accuracy for stereo video. HSS deformations are common among images of horizontally aligned views, due to horizontal and vertical flat surfaces that are not parallel with projection image planes. To achieve HSS-based DCP with minimal complexity, an efficient subsampled block-matching technique is adopted and integrated into MVC extension of H.264/AVC in stereo profile. Affine parameters estimation and additional frame buffers are not required and the overall increase of computational complexity and memory requirements are moderate. Experimental results show that the new technique can achieve up to 5.25% bitrate reduction in interview prediction using JM17.0 reference software implementation. Ka-Man Wong, Lai-Man Po, Kwok-Wai Cheung 0002, Chi-Wang Ting, Ka-Ho Ng, Xuyuan Xu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | Adaptive Quantization in DCT Domain for Distributed Video CodingabstractSummary form only given. In DCT domain distributed video coding (DVC), the uniform symmetric scalar quantization has been widely implemented, where, coefficients organized into 16 DCT bands are quantized with the given quantization levels. For each sequence has its own contents, the optimal quantization levels should vary with frames. In this paper, a low complexity algorithm for designing adaptive quantization levels (AQL) according to the characteristic of the bands is proposed. Chun-Ling Yang, Dong-Qin Xiao, Lai-Man Po, Wang-Hua Mo |
DCC | 3 |
| 2011 | Stretching, compression and shearing disparity compensated prediction techniques for stereo and multiview video codingabstractIn multiview video coding, disparity compensated prediction exploits the correlation among different views. A common approach is to use the conventional motion compensated prediction to predict disparity effect among different views. However, the same object in different views usually has deformation of different extents and, thus, accurate disparity prediction cannot be achieved with such simple translational motion model. Previous attempts to achieve more accurate disparity prediction are usually too complex for practical implementation. In this paper, stretching, compression and shearing (SCSH) effects are investigated to better model the disparity effect in disparity compensated prediction. To achieve SCSH effects with minimal computation, an efficient disparity compensated prediction using subsampled block-matching technique is proposed. No affine parameters estimation or additional frame buffers is required and the overall increase in memory requirement and computational complexity is moderate. Experimental results show that the new technique can achieve up to 4.84% bitrate reduction in inter-view prediction using JM17.0 reference software implementation. Ka-Man Wong, Lai-Man Po, Kwok-Wai Cheung 0002, Ka-Ho Ng, Xuyuan Xu |
ICASSP | 2 |
| 2011 | A new multidirectional extrapolation hole-filling method for Depth-Image-Based RenderingabstractDepth-Image-Based Rendering (DIBR) is widely used to generate virtual view of a scene from a known view with associated depth map in 3D video applications. However, disocclusion arises in image warping of DIBR. Many hole-filling methods have been proposed such as constant color, horizontal interpolation, horizontal extrapolation, and variational inpainting, but they cause different types of annoying artifact for large holes with complex texture background. In this paper, a novel multidirectional extrapolation hole-filling method is proposed to enhance visual quality for large hole-filling with complex texture background. The proposed method uses neighbor pixels' texture features to estimate hole-filling direction in a pixel-by-pixel manner. Experimental results demonstrated that the proposed method could provide better visual quality compared with conventional methods for virtual views synthesis with high-quality depth map. Lai-Man Po, Shihang Zhang, Xuyuan Xu, Yuesheng Zhu |
ICIP | 1 |
| 2010 | A novel watermarking scheme with compensation in bit-stream domain for H.264/AVCabstractCurrently, most of the watermarking algorithms for H.264/AVC video coding standard are encoder-based due to their high perceptual quality. However, for the compressed video, they increase the computational burden to decode the video, embed the watermark, and then re-encode it. Obviously it is a bottleneck for real-time applications. Conventional watermarking algorithms in the bit-stream domain can reduce the computation complexity but result in the error propagation and PSNR loss. In this paper, a new fast watermarking scheme with compensation in bit-stream compressed video domain is proposed, in which the watermark is directly embedded into the quantized residual coefficients. A texture-based perceptual model is employed to decide whether a 4×4-block is appropriate for watermarking or not. A secret key is used to decide the actual location to be embedded in a 4×4-block. With the proposed compensation method, the video watermarking scheme can achieve high robustness and good visual quality without much bit-rate increase. The simulation results have shown that a high PSNR is achieved with little increase in bit rate, and the watermarks can resist common attacks. Yuesheng Zhu, Lai-Man Po |
ICASSP | 3 |
| 2010 | Distance-based weighted prediction for Adaptive Intra Mode Bit Skip in H.264/AVCabstractAdaptive Intra Mode Bit Skip (AIMBS) technique using boundary pixels smoothness has been shown to achieve coding efficiency improvement for H.264/AVC's Intra_4×4 coding in relatively large QPs. However, the DC mode in the Multiple-Prediction of the AIMBS becomes much less effective. To tackle this problem and further improve the coding efficiency, distance-based weighted prediction (DWP) is proposed to replace DC mode in Multiple-Prediction for predicting blocks without directional preferences. The proposed method is named as AIMBS-DWP that can enhance the robustness of AIMBS in much larger range of QPs and achieve higher rate-distortion performance. Experimental results show that an average bitrate reduction of 3.79% with lower computational requirement can be obtained by AIMBS-DWP as compared with H.264/AVC. The improvement is especially obvious in high visual quality configurations with small QPs. Lai-Man Po, Liping Wang 0009, Kwok-Wai Cheung 0002, Ka-Man Wong, Ka-Ho Ng, Shenyuan Li, Chi-Wang Ting |
ICIP | 1 |
| 2010 | Subsampled Block-Matching for Zoom Motion Compensated PredictionabstractMotion compensated prediction plays a vital role in achieving enormous video compression efficiency in advanced video coding standards. Most practical motion compensated prediction techniques implicitly assume pure translational motions in the video contents for effective operation. Some attempts aiming at more general motion models are usually too complex requiring parameter estimation in practical implementation. In this paper, zoom motion compensation is investigated to extend the assumed model to support both zoom and translation motions. To accomplish practical complexity, a novel and efficient subsampled block-matching zoom motion estimation technique is proposed which makes use of the interpolated reference frames for subpixel motion estimation in a conventional hybrid video coding structure. Specially designed subsampling patterns in block matching are used to realize the translation and zoom motion estimation and compensation. No zoom parameters estimation and additional frame buffers are required in the encoder implementation. The complexity of the decoder is similar to the conventional hybrid video codec that supports subpixel motion compensation. The overall increase in memory requirement and computational complexity is moderate. Experimental results show that the new technique can achieve up to 9.89% bitrate reduction using KTA2.2r1 reference software implementation. Lai-Man Po, Ka-Man Wong, Kwok-Wai Cheung 0002, Ka-Ho Ng |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2009 | Block-Matching Translation And Zoom Motion-Compensated Prediction by sub-samplingabstractIn modern video coding standards, motion compensated prediction (MCP) plays a key role to achieve video compression efficiency. Most of them make use of block matching techniques and assume the motions are pure translational. Some attempts toward a more general motion model usually too complex to be practical in near future. In this paper, a new Block-Matching Translation and Zoom Motion-Compensated Prediction (BTZMP) is proposed to extend the pure translational model to a more general model with zooming in a practical way. It adopts the camera zooming and object motions that becomes zooming while projected on the video frames. The proposed BTZMP significantly improve motion compensated prediction. Experimental results show that BTZMP can give prediction gain up to 1.09dB compared to conventional sub-pixel block-matching MCP. In addition, BTZMP can be incorporated with Multiple Reference Frames (MRF) technique to give extra improvement, evidentially by the prediction gain ranging up to 2.08dB in the empirical simulations. Ka-Man Wong, Lai-Man Po, Kwok-Wai Cheung 0002, Ka-Ho Ng |
ICIP | 2 |
| 2009 | A novel weighted cross prediction for H.264 intra codingabstractIn this paper, a novel weighted cross prediction (WCP) mode is proposed to replace DC mode in Intra_4times4 prediction of H.264/AVC. In the proposed scheme, the upper right part of one 4times4 block mainly employs vertical prediction while the lower left part mainly uses horizontal prediction, predicting both in vertical and horizontal directions in one block. This scheme uses simple prediction equations with fixed weighting coefficients. Experimental results show that WCP has improvement compared to H.264 and it is very competitive while comparing to other Intra_4times4 prediction algorithms. Liping Wang 0009, Lai-Man Po, Y. M. S. Uddin, Ka-Man Wong, Shenyuan Li |
ICME | 2 |
| 2009 | Block-Matching Translation and Zoom Motion-Compensated PredictionabstractIn modern video coding standards, motion compensated prediction (MCP) plays a key role to achieve video compression efficiency. Most of them make use of block matching techniques and assume the motions are pure translational. Attempts toward a more general motion model are usually too complex to be practical in near future. In this paper, a new Block-Matching Translation and Zoom Motion-Compensated Prediction (BTZMP) is proposed to extend the pure translational model to a more general model with zooming. It adopts the camera zooming and object motions that becomes zooming while projected on video frames. Experimental results show that BTZMP can give prediction gain up to 2.25dB for various sequences compared to conventional block-matching MCP. BTZMP can also be incorporated with multiple reference frames technique to give extra improvement, evidentially by the prediction gain ranging from 2.03 to 3.68dB in the empirical simulations. Ka-Man Wong, Lai-Man Po, Kwok-Wai Cheung 0002 |
ICME | 2 |
| 2009 | Six-Digit Stroke-based Chinese Input MethodabstractDuring the last three decades, more than one thousand Chinese input methods have been developed. However, people are still looking for better input methods in terms of easy to use, easy to remember, high input speed and small keypad implementation on handheld devices. The well-known stroke-based Chinese input method using only five basic stroke types could achieve low learning curve and small numeric keypad implementation but its input speed is limited for complex Chinese characters with a lot of strokes. To tackle this problem, simplified stroke-based Chinese character and phrase coding methods using (3+3) rules are proposed in this paper. The proposed method only uses the first 3 stroke codes and the last 3 stroke codes to represent the first and last radical information of the character for achieving lower average code length and higher hit rate of first character on the candidate list. To further enhance the input speed, a very user-friendly (3+3) phrase coding rule is also proposed for inputting Chinese phrases in terms of 2-character, 3-character and long-character phrases. Three special key assignment designs are developed for practical implementation of the proposed Chinese character and phrase input method using conventional QWERTY keyboard, PC's numeric keypad and mobile phone 12-key keypad. Experimental results have shown that the proposed character coding can achieve lower average code length and higher hit rate of first character as compared with conventional stroke-based method and some well-known Chinese input methods. The proposed coding rules are also very easy to use and remember. Lai-Man Po, Chi-Kwan Wong, Yiu-Ki Au, Ka-Ho Ng, Ka-Man Wong |
SMC | 1 |
| 2009 | The training of Karhunen-Loève transform matrix and its application for H.264 intra coding
Jiazhong Chen, Shengsheng Yu, Jingli Zhou, Lai-Man Po |
Multim. Tools Appl. | 5 |
| 2009 | A Search Patterns Switching Algorithm for Block Motion EstimationabstractCenter-biased fast motion estimation algorithms, e.g., block-based gradient descent search and diamond search, can perform much better than coarse-to-fine search algorithms, such as 2-D logarithmic search and three-step search. The latter type of algorithms, however, is more suitable for handling large motion content. To combine the advantages of both types of algorithms, an adaptive algorithm performing search patterns switching (SPS) is proposed in this paper. The proposed SPS algorithm classifies the motion content of a block using a simple yet efficient motion content classifier callederrordescentrate. Unlike other classifiers with heavy overhead, this classifier requires only the searching of a few points in the search window and then a division operation. Experimental results show that the proposed SPS algorithm is very robust. Ka-Ho Ng, Lai-Man Po, Ka-Man Wong, Chi-Wang Ting, Kwok-Wai Cheung 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | Novel Directional Gradient Descent Searches for Fast Block Motion EstimationabstractSearch point pattern-based fast block motion estimation algorithms provide significant speedup for motion estimation but usually suffer from being easily trapped in local minima. This may lead to low robustness in prediction accuracy particularly for video sequences with complex motions. This problem is especially serious in one-at-a-time search (OTS) and block-based gradient descent search (BBGDS), which provide very high speedup ratio. A multipath search using more than one search path has been proposed to improve the robustness of BBGDS but the computational requirement is much increased. To tackle this drawback, a novel directional gradient descent search (DGDS) algorithm using multiple OTSs and gradient descent searches on the error surface in eight directions is proposed in this letter. The search point patterns in each stage depend on the minima found in these eight directions, and thus the global minimum can be traced more efficiently. In addition, a fast version of the DGDS (FDGDS) algorithm is also described to further improve the speed of DGDS. Experimental results show that DGDS reduces computation load significantly compared with the well-known fast block motion estimation algorithms. Moreover, FDGDS can achieve faster speedup compared with the UMHexagonS algorithm in H.264/AVC implementation while maintaining very similar rate-distortion performance. Lai-Man Po, Ka-Ho Ng, Kwok-Wai Cheung 0002, Ka-Man Wong, Y. M. S. Uddin, Chi-Wang Ting |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2008 | Discrete wavelet transform-based structural similarity for image quality assessmentabstractImage quality assessment method plays a major role in image processing. Structural similarity (SSIM) is a novel image quality assessment method, and attracts a lot of attentions for its good performance and simple calculation. But it is proposed in pixel domain, a large computation load is introduced when using it to guide image processing algorithm in discrete wavelet transform (DWT) domain. And image processing in DWT domain has been very common. This paper proposes a discrete wavelet transform-based structural similarity (DWT-SSIM) method for image quality assessment. As it highlights human eyes’ sensitive frequency bands, the proposed measure has a better correlation with the judgment of human observers than SSIM. Furthermore, the method is easy to implement and embed in DWT domain processing algorithms. Chun-Ling Yang, Wen-Rui Gao, Lai-Man Po |
ICIP | 3 |
| 2008 | Fast sum of absolute transformed difference based 4×4 intra-mode decision of H.264/AVC video coding standard
Mohammed Golam Sarwer, Lai-Man Po, Q. M. Jonathan Wu |
Signal Process. Image Commun. | 2 |
| 2007 | Dominant Color Structure Descriptor for Image RetrievalabstractA new dominant color structure descriptor (DCSD) is proposed in this paper. It is designed to provide an efficient way to represent both color and spatial structure information with single compact descriptor. The descriptor combines the compactness of dominant color descriptor (DCD) and the retrieval accuracy of color structure descriptor (CSD) to enhance the retrieval performance in a highly efficient manner. The feature extraction and similarity measure of the descriptor are designed to address the problems of the existing descriptors while utilize the advantages of them. Experimental results show that DCSD has a significant improvement on both retrieval performance and descriptor size over DCD. An eight-color DCSD (DCSD 8) gives an averaged normalized modified retrieval rate (ANMRR) of 0.0993 using MPEG-7 common color dataset, outperforming compact configurations of scalable color descriptor and color structure descriptor with smaller descriptor size. Ka-Man Wong, Lai-Man Po, Kwok-Wai Cheung 0002 |
ICIP (6) | 2 |
| 2007 | A New Partial Codeword Updating Scheme Based on Rate-Distortion Optimization for Adaptive Vector QuantizationabstractIn this paper, we propose a new adaptive vector quantization (AVQ) algorithm based on the rate-distortion optimization. This algorithm employs a new partial codeword updating (PCU) scheme which achieves rate-distortion performance superior to that of the conventional AVQ algorithms using the full codeword updating (FCU) scheme. The PCU-AVQ only updates the codeword's components with the quantization error higher than an optimal threshold instead of replacing the whole codeword. Additionally, the mathematical relation between the Lagrangian multiplier and the approximate optimal threshold is devised to reduce the rate-distortion cost computation. The experimental results show that the proposed PCU-AVQ algorithm indeed improves the rate-distortion performance without much computational complexity penalty. The PCU-AVQ can be combined with transform coding and entropy coding for higher compression ratio, and it can be widely implemented in specific AVQ algorithms for image, video and speech coding. Lai-Man Po |
ICME | 2 |
| 2007 | Search Patterns Switching for Motion Estimation using Rate of Error DescentabstractMost of the fast motion estimation algorithms based on search-point pattern are only good at handling videos with small motions, for example, block-based gradient descent search, diamond search and hexagonal-based search. An adaptive motion estimation algorithm which can switch between search patterns for different video contents should work better than a single search pattern algorithm. In this paper, a simple classifier based on error descent rate (EDR) is proposed. This classifier uses a very few number of search points to predict whether the global minimum is far away or near the center of the search window. If it is far away, a search pattern which is good at searching large motions is used. Otherwise, a pattern good at searching small motions is applied. The proposed search patterns switching (SPS) algorithm performs well for all kinds of video contents. Ka-Ho Ng, Lai-Man Po, Ka-Man Wong |
ICME | 2 |
| 2007 | Bit Rate Estimation for Cost Function of 4×4 Intra Mode Decision of H.264/AVCabstractH.264/AVC is a newest international video coding standard that can achieve considerably higher coding efficiency than previous standards. This comes at the cost of the complex mode decision procedure using the rate-distortion optimization, which makes real-time encoding difficult. To reduce the complexity of rate-distortion cost, we propose a bit rate estimation technique to avoid the entropy coding method during mode decision of intra prediction. The estimation method is based on the properties of context-based variable length coding (CAVLC). Simulation results demonstrate that the proposed estimation method achieves up to 53 % reduced encoding time of intra coding with ignorable degradation of coding performance. Mohammed Golam Sarwer, Lai-Man Po |
ICME | 2 |
| 2007 | A Compact and Efficient Color Descriptor for Image RetrievalabstractAn important problem in color based image retrieval is the lack of efficient way to represent both the color and the spatial structure information with single descriptor. To solve this problem, a new dominant color structure descriptor (DCSD), is proposed. The descriptor combines the compactness of dominant color descriptor (DCD) and the accuracy of color structure descriptor (CSD) to enhance the retrieval performance in a highly efficient manner. The feature extraction and similarity measure of the descriptor are designed to address the problems of the existing descriptors such as color inaccuracy of DCD and redundancy of CSD. Experimental results show that DCSD has a significant improvement in retrieval performance and descriptor size over DCD. An eight-color DCSD (DCSD 8) gives an averaged normalized modified retrieval rate (ANMRR) of 0.0993 using MPEG-7 common color dataset, outperforming compact configurations of scalable color descriptor and color structure descriptor with smaller descriptor size. Ka-Man Wong, Lai-Man Po, Kwok-Wai Cheung 0002 |
ICME | 2 |
| 2007 | Transform-Domain Fast Sum of the Squared Difference Computation for H.264/AVC Rate-Distortion OptimizationabstractIn H.264/AVC, the rate-distortion optimization for mode decision plays a significant role to achieve its outstanding performance in terms of both compression efficiency and video quality. However, this mode decision process also introduces extremely high complexity in the encoding process especially the computation of the sum of squared differences (SSD) between the original and reconstructed image blocks. In this paper, fast SSD (FSSD) algorithms are proposed to reduce the complexity of the rate-distortion cost function implementation. The proposed FSSD algorithm is based on the theoretical equivalent of the SSDs in spatial and transform domains and determines the distortion in integer cosine transform domain using an iterative table-lookup quantization process. This approach could avoid the inverse quantization/transform and pixel reconstructions processes with nearly no rate-distortion performance degradation. In addition, the FSSD can also be used with efficient bit rate estimation algorithms to further reduce the cost function complexity. Experimental results show that the new FSSD can save up to 15% of total encoding time with less than 0.1% coding performance degradation and it can save up to 30% with ignorable performance degradation when combining with conventional bit rate estimation algorithm Lai-Man Po |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | Fast Bit Rate Estimation for Mode Decision of H.264/AVCabstractTo achieve the highest coding efficiency, H.264/AVC uses rate-distortion optimization technique. This means that the encoder has to code the video by exhaustively trying all the mode combinations including the different intra- and inter-prediction modes. Therefore, the complexity and computation load of video coding in H.264/AVC increase drastically compared to any previous standards. To reduce the complexity of rate-distortion cost computation, we propose a fast bit rate estimation technique to avoid the entropy coding method during intra- and inter-mode decision of H.264/AVC. The estimation method is based on the properties of context-based variable length coding (CAVLC). The proposed rate model predicts the rate of a 4 times 4 quantized residual block using five different tokens of CAVLC. Experimental results demonstrate that the proposed estimation method reduces about 47% of total encoding time on using intra-modes only and saves about 34% of total encoding time on using both inter- and intra-modes with ignorable degradation of coding performance when the fast motion search algorithm is used. When full search motion estimation algorithm is used, the proposed algorithm reduces about 17% of total encoding time. Mohammed Golam Sarwer, Lai-Man Po |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | Novel Point-Oriented Inner Searches for Fast Block Motion EstimationabstractRecently, an enhanced hexagon-based (EHS) search algorithm was proposed to speedup the original hexagon-based search (HS) using a 6-side-based fast inner search. However, this 6-side-based method is quite irregular by inspecting the distance between the inner search points and the coarse search points that would lower prediction accuracy. In this paper, a new point-oriented grouping strategy is proposed to develop fast inner search techniques for speeding up the HS and diamond search (DS) algorithms. Experimental results show that the new HS and DS using point-oriented inner searches are faster than their original algorithms up to 30% with negligible peak signal-to-noise ratio degradation Lai-Man Po, Chi-Wang Ting, Ka-Man Wong, Ka-Ho Ng |
IEEE Trans. Multim. | 1 |
| 2006 | Edge-Based Structural Similarity for Image Quality AssessmentabstractObjective quality assessment has been widely used in image processing for decades and many researchers have been studying the objective quality assessment method based on Human Visual System (HVS). Recently the Structural Similarity (SSIM) is proposed, under the assumption that the HVS is highly adapted for extracting structural information from a scene, and simulation results have proved that it is better than PSNR (or MSE). By deeply studying the SSIM, we find it fails in measuring the badly blurred images. Based on this, we develop an improved method which is called Edge-based Structural Similarity (ESSIM). Experiment results show that ESSIM is more consistent with HVS than SSIM and PSNR especially for the blurred images. Guan-Hao Chen, Chun-Ling Yang, Lai-Man Po, Shengli Xie 0001 |
ICASSP (2) | 3 |
| 2006 | A Novel Motion Estimation Method Based on Structural Similarity for H.264 Inter PredictionabstractIn the motion estimation of H.264, the best matching blocks and the best prediction modes are chosen by Lagrange cost function whose distortion metric is the sum of absolute (transformed) differences [SA(T)D] which has similar meaning with MSE or PSNR. Recently a new image measurement called Structural Similarity (SSIM) based on the degradation of structural information was brought forward. It is proved that the SSIM can provide a better approximation for the perceived image distortion than the currently used PSNR (or MSE). In this paper, we propose an improved motion estimation method based on SSIM(MEBSS) for H.264 inter coding. Experiment results show that the MEBSS can reduce average 20% bit rate and 2% encoding time while maintaining the same perceptual video quality, and the maximum reduction in bitrate is more than 50%. Zhi-Yi Mai, Chun-Ling Yang, Kai-Zhi Kuang, Lai-Man Po |
ICASSP (2) | 4 |
| 2006 | Enhanced Diamond Search Using Four-Corner-Based Inner Search For Fast Block Motion EstimationabstractThis paper proposes a very efficient and simple fineresolution inner search for the well-known diamond search (DS) algorithm. Based on the local unimodal error surface assumption, a new 4-corner ased inner search technique is proposed to speedup the original DS in both large and small motion environments by exploiting the group-distortion information of some evaluated points. Experimental results show that the enhanced diamond search (EDS) is faster than the original DS up to 30% with negligible PSNR degradation. Lai-Man Po, Chi-Wang Ting, Ka-Ho Ng |
ICASSP (2) | 1 |
| 2006 | Point Oriented Hexagonal Inner Search For Fast Block Motion EstimationabstractRecently, an enhanced hexagon-based search (EHS) algorithm was proposed to speedup the original hexagon-based search (HS) by exploiting the group-distortion information of some evaluated points. In this paper, a second version of the EHS is proposed with a new point-oriented inner search technique which can further speedup the HS in both large and small motion environments. Experimental results show that the enhanced hexagon-based search version-2 (EHS2) is faster than the HS up to 34% with negligible PSNR degradation. Lai-Man Po, Chi-Wang Ting, Ka-Ho Ng |
ICIP | 1 |
| 2005 | A New Rate-Distortion Optimization Using Structural Information in H.264 I-Frame Encoder
Zhi-Yi Mai, Chun-Ling Yang, Lai-Man Po, Shengli Xie 0001 |
ACIVS | 3 |
| 2005 | Novel cross-diamond-hexagonal search algorithms for fast block motion estimationabstractWe propose two cross-diamond-hexagonal search (CDHS) algorithms, which differ from each other by their sizes of hexagonal search patterns. These algorithms basically employ two cross-shaped search patterns consecutively in the very beginning steps and switch using diamond-shaped patterns. To further reduce the checking points, two pairs of hexagonal search patterns are proposed in conjunction with candidates found located at diamond corners. Experimental results show that the proposed CDHSs perform faster than the diamond search (DS) by about 144% and the cross-diamond search (CDS) by about 73%, whereas similar prediction quality is still maintained. Chun-Ho Cheung, Lai-Man Po |
IEEE Trans. Multim. | 2 |
| 2004 | A novel kite-cross-diamond search algorithm for fast video coding and videoconferencing applicationsabstractWe propose a kite-cross-diamond search (KCDS) algorithm, which is an improved version of the well-known cross-diamond search (CDS) algorithm and small cross-diamond search (SCDS) algorithm. Unlike traditional search patterns in block matching algorithms, such as square, diamond or cross, which are all in vertical and horizontal symmetric shapes, the KCDS algorithm adopts a novel asymmetric kite-shaped search pattern to keep similar distortion while the speed of the motion estimation for stationary or quasi-stationary blocks is further boosted. Experimental results show that the KCDS algorithm could achieve 39% search point reduction as compared with CDS, whereas there is similar and even better prediction accuracy in low-motion sequences. Simulations show that KCDS is the fastest algorithm and it performs more accurately in some kinds of sequences. This algorithm is especially suitable for videoconferencing applications. Chi-Wai Lam, Lai-Man Po, Chun-Ho Cheung |
ICASSP (3) | 2 |
| 2004 | MPEG-7 dominant color descriptor based relevance feedback using merged palette histogramabstractThis paper proposes techniques to improve the effectiveness of relevance feedback (RF) based image retrieval using MPEG-7 dominant color descriptor (DCD). The conventional RF query point moving techniques could not be directly applied on DCD due to the difference of the color spaces used in each DCD histogram. In order to tackle this problem, a new merged palette histogram (MPH) technique is proposed for generating a new query DCD histogram from the relevant image set. The new method has been implemented on a MPEG-7 XM with use of an image database containing 1000 images. The effectiveness of this new MPH-RF technique using different similarity measures is demonstrated by experimental results in the terms of averaged normalized modified retrieval rate (ANMRR) and visual retrieved image comparison. Ka-Man Wong, Lai-Man Po |
ICASSP (3) | 2 |
| 2004 | A new palette histogram similarity measure for mpeg-7 dominant color descriptor
Lai-Man Po, Ka-Man Wong |
ICIP | 1 |
| 2004 | Fast block-matching motion estimation bv recent-biased search for multiple reference frames
Chi-Wang Ting, Hong Lam, Lai-Man Po |
ICIP | 3 |
| 2004 | A fast H.264 intra prediction algorithm using macroblock propertiesabstractIn the newest video coding standard of H.264, intra frame prediction is conducted in the spatial domain. There are two types of intra prediction for luma - Intra/spl I.bar/16/spl times/16 does prediction for the whole 16/spl times/16 luma macroblock and Intra/spl I.bar/4/spl times/4 does prediction for 4/spl times/4 block. Four and nine modes are supported by Intra/spl I.bar/16/spl times/16 and Intra/spl I.bar/4/spl times/4, respectively. The full search algorithm is used in the JVT reference software to choose the best modes, however it is very computationally expensive. In this paper, a new fast intra prediction algorithm using macroblock properties (FIPAMP) is proposed. Experimental results show that the proposed fast algorithm can achieve 10% to 40% computation reduction while maintaining similar PSNR and bit rate performance of H.264 codes. Chun-Ling Yang, Lai-Man Po, Wing-Hong Lam |
ICIP | 2 |
| 2004 | Enhanced hexagonal search for fast block motion estimationabstractFast block motion estimation normally consists of low-resolution coarse search and the following fine-resolution inner search. Most motion estimation algorithms developed attempt to speed up the coarse search without considering accelerating the focused inner search. On top of the hexagonal search method recently developed, an enhanced hexagonal search algorithm is proposed to further improve the performance in terms of reducing number of search points and distortion, where a novel fast inner search is employed by exploiting the distortion information of the evaluated points. Our experimental results substantially justify the merits of the proposed algorithm. Ce Zhu, Xiao Lin 0001, Lap-Pui Chau, Lai-Man Po |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2003 | A novel histogram-biasing factor for fast sorted histogram-based measurement in large image database retrieval systemabstractThe exhaustive histogram matching is usually the most computationally intensive part for any query in most large image database retrieval systems. In this paper, we introduce a histogram-biasing factor (HBF) to measure the biased-behavior of ordered-bins in a sorted histogram. The proposed HBF can be used to increase the early rejection rate of unreliable or impossible candidate reference images based on one of the sorted histograms. Moreover, it can be treated as a color-histogram descriptor. Only images with very closed HBF are taken into account, searching speed can thus be increased without loss of accuracy. Experimental results show that the proposed factor results in up to 13 times speedup meanwhile providing the exhaustive retrieval performance. Chun-Ho Cheung, Lai-Man Po |
ICASSP (3) | 2 |
| 2003 | Adjustable partial distortion search algorithm for fast block motion estimationabstractThe quality control for video coding usually absents from many traditional fast block motion estimators. A novel block-matching algorithm for fast motion estimation named the adjustable partial distortion search algorithm (APDS) is proposed. It is a new normalized partial distortion comparison method capable of adjusting the prediction accuracy against searching speed by a quality factor k. With adjustability, APDS could act as the normalized partial distortion search algorithm (NPDS) when k is equal to 0, and the conventional partial distortion search algorithm (PDS) when k is equal to 1. In addition, it uses a halfway-stop technique with progressive partial distortions (PPD) to increase early rejection rate of impossible candidate motion vectors at very early stages. Simulations with PPD reduce computations up to 38 times with less than 0.50-dB degradation in PSNR performance, as compared to the full-search algorithm (FS). Experimental results show that APDS could provide peak signal-to-noise ratio performance very close to that of FS with speedup ratios of 7 to 16 times, and close to that of NPDS from 22 to 32 times, respectively, as compared to FS. Chun-Ho Cheung, Lai-Man Po |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2002 | A novel rood-diamond search algorithm for fast block motion estimationabstractSearch patterns and the center-biased characteristics of motion vector distribution (MVD) have large impact on both searching speed and quality of block motion estimation. In this paper, we propose a novel algorithm using a rood-shaped search pattern as the initial step and large/small diamond search patterns as the subsequent steps for fast block motion estimation (BME). The rood-shaped pattern is to fit the rood-center-biased MVD characteristics of the real-world sequences by evaluating the 9 relatively higher probable candidates located as a rood-shaped pattern at the search-grid center. The proposed rood-diamond search algorithm (RDS) employs halfway-stop technique and could find small motion vectors with fewer points than the diamond search algorithm (DS) while maintains similar or even better quality. The speedup improvement of RDS over DS can be up to 40% Simulations show that RDS is much more robust. provides faster searching speed and smaller distortions than other fast algorithms. Chun-Ho Cheung, Lai-Man Po |
ICASSP | 2 |
| 2002 | Spatial coefficient partitioning for lossless wavelet image codingabstractA novel coefficient partitioning algorithm is introduced for splitting the coefficients into two sets using spatial orientation tree data structure. By splitting the coefficients, the overall theoretical entropy is reduced due to the different probability distribution for the two coefficient sets. In spatial domain, it is equivalent to identifying smooth regions of the image. A lossless coder based on this spatial coefficient partitioning is described. Experimental results show that the new algorithm has a better coding performance than other wavelet based lossless image coder such as S+P and JPEG-2000. Kwok-Wai Cheung 0002, Lai-Man Po |
ICASSP | 2 |
| 2002 | A novel small-cross-diamond search algorithm for fast video coding and videoconferencing applicationsabstractSearch patterns and the center-biased characteristics of motion vector distribution (MVD) have a large impact on both searching speed and quality of block motion estimation. We propose a novel algorithm using two cross-shaped search patterns as the first two initial steps and large/small diamond-shaped patterns as the subsequent steps for fast block motion estimation (BME). The first small cross-shaped pattern is to fit the cross-center-biased MVD characteristics of the real-world sequences by evaluating the 5 relatively higher probable candidates located as a cross-shaped pattern at the search-grid center. The proposed small-cross-diamond search algorithm (SCDS) employs a halfway-stop technique and could find small motion vectors with much fewer points than the diamond search algorithm (DS) while maintains similar or even better quality. The speedup improvement of SCDS over DS can be up to 146%, i.e. 2.46 times faster than DS. Simulations show that SCDS is much more robust, provides faster searching speed and smaller distortions than other fast algorithms, typically very suitable for videoconferencing applications. Chun-Ho Cheung, Lai-Man Po |
ICIP (1) | 2 |
| 2002 | Merged-color histogram for color image retrievalabstractConventional histogram-based image retrieval algorithms usually find only intersecting areas of the color-component distributions of images, and thus work well in matching images with exact colors instead of similar colors, especially for computer generated pictures. This could be greatly affected by overall variations, such as intensity changes. A novel merged-color histogram (MCH) method for color image retrieval is proposed to facilitate matching between similar colors by means of color quantization and palette merging. Color quantization compacts the color information and matches each color instead of color components, and matching of similar colors is accomplished using palette merging. Experimental results show that the proposed MCH method is about 11-32% more precise in the first 20 retrievals for the same image query, and is able to recall 14%-23% more relevant images in the first 100 retrievals, as compared to the conventional RGB-based histogram method. Ka-Man Wong, Chun-Ho Cheung, Lai-Man Po |
ICIP (3) | 3 |
| 2002 | A novel cross-diamond search algorithm for fast block motion estimationabstractIn block motion estimation, search patterns with different shapes or sizes and the center-biased characteristics of motion-vector distribution have a large impact on the searching speed and quality of performance. We propose a novel algorithm using a cross-search pattern as the initial step and large/small diamond search (DS) patterns as the subsequent steps for fast block motion estimation. The initial cross-search pattern is designed to fit the cross-center-biased motion vector distribution characteristics of the real-world sequences by evaluating the nine relatively higher probable candidates located horizontally and vertically at the center of the search grid. The proposed cross-diamond search (CDS) algorithm employs the halfway-stop technique and finds small motion vectors with fewer search points than the DS algorithm while maintaining similar or even better search quality. The improvement of CDS over DS can be up to a 40% gain on speedup. Experimental results show that the CDS is much more robust, and provides faster searching speed and smaller distortions than other popular fast block-matching algorithms. Chun-Ho Cheung, Lai-Man Po |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2001 | Generalized partial distortion search algorithm for fast block motion estimationabstractQuality versus speed control for real-time video applications, such as speed-oriented videoconferencing or high quality video entertainment, usually is absent from many traditional fast block motion estimators. A novel block-matching algorithm for fast motion estimation named generalized partial distortion search algorithm (GPDS) is proposed. It uses a halfway-stop technique with progressive partial distortion (PPD) to increase the chance of early rejection of impossible candidate motion vectors at very early stages. Simulations on PPD show 28 to 38 times computational reduction with only 0.45-0.50 dB PSNR performance degradation as compared to the full search algorithm. In addition, a new normalized partial distortion comparison method is also proposed for enabling control of searching speed against prediction quality by a speedup factor k. This method also generalizes the conventional partial distortion search algorithm when k is equal to 1, and the normalized partial distortion search algorithm (NPDS) when k is equal to infinity. Experimental results show that GPDS with use of PPD could provide PSNR performance very close to the full search algorithm and NPDS with 7 to 17 times and 22 to 33 times speedup, respectively, as compared to the full search algorithm. Chun-Ho Cheung, Lai-Man Po |
ICASSP | 2 |
| 2001 | Generalized partial distortion search algorithm for block-matching motion estimationabstractQuality against speed control for real-time video applications, such as low-bit-rate video conferencing or high quality video entertainment, is usually absent from many traditional fast block motion estimators. A novel block-matching algorithm for fast motion estimation, named generalized partial distortion search algorithm (GPDS), is proposed. It uses a halfway-stop technique with progressive partial distortion (PPD) to increase the chance of early rejection of impossible candidate motion vectors at very early stages. Simulations on PPD show 28 to 38 times computational reduction with only 0.45-0.50 dB PSNR performance degradation as compared to the full search algorithm. In addition, a new normalized partial distortion comparison method is proposed for enabling the control of searching speed against prediction accuracy by a quality factor k. This method also generalizes the conventional partial distortion search algorithm when k=1, and the normalized partial distortion search algorithm (NPDS) when k=/spl infin/. Experimental results show that GPDS with use of PPD could provide PSNR performance very close to the full search algorithm with 7 to 17 times speedup, and to NPDS with 22 to 33 times speedup, respectively, as compared to the full search algorithm. Chun-Ho Cheung, Lai-Man Po |
ICIP (3) | 2 |
| 2000 | Normalized partial distortion search algorithm for block motion estimationabstractMany fast block-matching algorithms reduce computations by limiting the number of checking points. They can achieve high computation reduction, but often result in relatively higher matching error compared with the full-search algorithm. A novel fast block-matching algorithm named normalized partial distortion search is proposed. The proposed algorithm reduces computations by using a halfway-stop technique in the calculation of the block distortion measure. In order to increase the probability of early rejection of non-possible candidate motion vectors, the proposed algorithm normalized the accumulated partial distortion and the current minimum distortion before comparison. Experimental results show that the proposed algorithm can maintain its mean square error performance very close to the full-search algorithm while achieving an average computation reduction of 12-13 times, with respect to the full-search algorithm. Chok-Kwan Cheung, Lai-Man Po |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1999 | A lossless multi-partitioning successive zero coder for wavelet-based progressive image transmissionabstractThis paper proposed an embedded image compression algorithm called lossless multi-partitioning successive zero coder (LMP-SZC) using the integer wavelet transform for progressive image transmission (PIT). By dynamically adjusting the partitions based on the space-frequency domain coefficients, our algorithm can achieve lower complexity and superior coding efficiency as compared with other well-known embedded lossless wavelet-based coder even without the zerotree analysis. Chun-Ho Cheung, Sheung-Yeung Wang, Kwok-Wai Cheung 0002, Lai-Man Po |
ICASSP | 4 |
| 1999 | A Novel Multiwavelet-Based Integer Transform for Lossless Image CodingabstractInteger Haar wavelet transform or S-transform is used as the basic building block for many exiting integer wavelet transform. As an alternative, a new integer multiwavelet transform and its associated integer prefilter are designed based on box-and-slope multi-scaling system. Both the transform and prefilter can be implemented with simple integer Haar transform requiring only addition and bit shift operations. Since the new integer transform is an approximation to nontruncated transform with higher vanishing moment than that of Haar transform, better approximation accuracy is expected and verified experimentally. The transform is successfully applied to lossless image coding with results outperforming that of lossless JPEG and S-transform. Kwok-Wai Cheung 0002, Chun-Ho Cheung, Lai-Man Po |
ICIP (1) | 3 |
| 1999 | Embedded Lossless Wavelet Coder Using Multi-Partitioning Algorithm with Bit-TrunkingabstractIn this paper, an embedded lossless wavelet image coder using multi-partitioning algorithm with bit-trunking (LMP-BT) is described. Without the conventional use of zerotrees analysis, our coder still gives a considerable progressive image transmission (PIT) performance and a superior coding efficiency over others in the literature. Chun-Ho Cheung, Sheung-Yeung Wang, Kwok-Wai Cheung 0002, Lai-Man Po |
ICIP (4) | 4 |
| 1999 | Successive Partition Zero Coder for Embedded Lossless Wavelet-Based Image CodingabstractA new embedded lossless wavelet-based image coding algorithm called successive partition zero coder (SPZC) which uses both horizontal and vertical bit scanning is proposed. Experimental results show that SPZC outperforms other state-of-the-art coders such as LJEG, SPIHT and CREW etc. In term of coding efficiency by successive partition the wavelet coefficients in the space-frequency domain and noncausal adaptive context modeling. Sheung-Yeung Wang, Chun-Ho Cheung, Kwok-Wai Cheung 0002, Lai-Man Po |
ICIP (4) | 4 |
| 1999 | Embedded lossless wavelet-based image coding algorithm with successive partitioning and hybrid bit scanningabstractA new embedded lossless wavelet-based image coding algorithm called successive partition zero coder (SPZC) which use hybrid bit scanning is proposed. By successive partition the wavelet coefficients in the space-frequency domain and non-causal adaptive context modeling, SPZC outperforms other state-of-the-art coders such as SPIHT, CREW and LJPEG etc. in terms of coding efficiency even without zerotree analysis. Sheung-Yeung Wang, Chun-Ho Cheung, Kwok-Wai Cheung 0002, Lai-Man Po |
MMSP | 4 |
| 1999 | Variable tree size fractal compression for wavelet pyramid image coding
Lai-Man Po |
Signal Process. Image Commun. | 2 |
| 1999 | Adaptive motion tracking block matching algorithms for video codingabstractIn most block-based video coding systems, the fast block matching algorithms (BMAs) use the origin as the initial search center, which may not track the motion very well. To improve the accuracy of the fast BMAs, a new adaptive motion tracking search algorithm is proposed. Based on the spatial correlation of motion blocks, a predicted starting search point, which reflects the motion trend of the current block, is adaptively chosen. This predicted search center is found closer to the global minimum, and thus the center-biased BMAs can be used to find the motion vector more efficiently. Experimental results show that the proposed algorithm enhances the accuracy of the fast center-biased BMAs, such as the new three-step search, the four-step search, and the block-based gradient descent search, as well as reduces their computational requirement. Jie-Bin Xu, Lai-Man Po, Chok-Kwan Cheung |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1998 | A novel subtree partitioning algorithm for wavelet-based fractal image codingabstractA novel wavelet subtree partitioning algorithm is proposed, which divides a subtree into scalar quantized wavelet coefficients and a fractal coded sub-subtree. Based on this new technique, a variable size wavelet subtree fractal coding scheme for still image compression is developed. Experimental results show that the new scheme can achieve a nearly optimal partition of the wavelet subtree with substantially computational reduction as compared with Davis' (see Proc. SPIE Wavelet Applications II, Orlando, vol.2491, p.141-152, 1995) scheme. Lai-Man Po, Kwok-Wai Cheung 0002, Chun-Ho Cheung |
ICASSP | 1 |
| 1998 | Minimax partial distortion competitive learning for optimal codebook designabstractThe design of the optimal codebook for a given codebook size and input source is a challenging puzzle that remains to be solved. The key problem in optimal codebook design is how to construct a set of codevectors efficiently to minimize the average distortion. A minimax criterion of minimizing the maximum partial distortion is introduced in this paper. Based on the partial distortion theorem, it is shown that minimizing the maximum partial distortion and minimizing the average distortion will asymptotically have the same optimal solution corresponding to equal and minimal partial distortion. Motivated by the result, we incorporate the alternative minimax criterion into the on-line learning mechanism, and develop a new algorithm called minimax partial distortion competitive learning (MMPDCL) for optimal codebook design. A computation acceleration scheme for the MMPDCL algorithm is implemented using the partial distance search technique, thus significantly increasing its computational efficiency. Extensive experiments have demonstrated that compared with some well-known codebook design algorithms, the MMPDCL algorithm consistently produces the best codebooks with the smallest average distortions. As the codebook size increases, the performance gain becomes more significant using the MMPDCL algorithm. The robustness and computational efficiency of this new algorithm further highlight its advantages. Ce Zhu, Lai-Man Po |
IEEE Trans. Image Process. | 2 |
| 1997 | Text-Driven Automatic Frame Generation Using MPEG-4 Synthetic/Natural Hybrid Coding for 2-D Head-and-Shoulder SceneabstractIn this paper, we propose a facial modeling technique based on the MPEG-4 synthetic/natural hybrid coding for automating frame sequence generation of a talking head. With the definition and animation parameters on a generic face object, the shape, textures and expressions of an adapted frontal face can generally be controlled and synchronized by the phonemes transcribed from plain text. By this developed facial modeling technique, it increases the intelligibility of an non-verbal facial communication for potential audiovisual lip-synch application on news reporting, lip-reading for the hearing-impaired or the deaf, virtual meeting through Internet, and story teller on demand (STOD). Chun-Ho Cheung, Lai-Man Po |
ICIP (2) | 2 |
| 1997 | Preprocessing for Discrete Multiwavelet Transform of Two-Dimensional SignalsabstractMultiwavelet transform can be implemented by a tree-structured matrix filter bank, which operates on vector sequence input instead of scalar ones. Therefore, unlike in a scalar wavelet system, preprocessing is usually required to extract vector sequence input from the input signal for better performance. A 2-D approximation-based preprocessing scheme for GHM (Geronimo, Hardin and Massopust) discrete multiwavelet transform of 2-D signals is proposed. Compared with the recent 1-D approximation-based method, the proposed scheme reduces the preprocessing computational complexity by 65% while maintaining similar energy compression ability. Kwok-Wai Cheung 0002, Lai-Man Po |
ICIP (2) | 2 |
| 1997 | A hierarchical block motion estimation algorithm using partial distortion measureabstractMost of the hierarchical block matching algorithms reduce computation by matching only some of the locations inside the search area. In this paper, we propose a novel hierarchical partial distortion search (HPDS) algorithm which reduces the computation of each distortion measure instead of the number of checking points by using partial distortion measure. Experimental results show that the proposed algorithm provides 6 times speed up as compared with full search while maintains the MSE performance very close to full search. Moreover, it can also achieve 20 times speed up when the MSE performance is kept close to the well-known three-step search. With appropriate parameter selections, the proposed algorithm can fit HDTV and MPEG-2 applications where high motion estimation accuracy is required, or real-time video conferencing applications where encoding speed is critical. Chok-Kwan Cheung, Lai-Man Po |
ICIP (3) | 2 |
| 1997 | A new prediction model search algorithm for fast block motion estimationabstractMost of the fast block matching algorithms (BMAs) use the origin as the initial center of the search. To improve the accuracy of the fast BMAs, a new prediction model (PM) search algorithm is proposed. Based on the spatial correlation of motion and the fast center-biased BMA, the four causal neighbour blocks are chosen to predict the starting search point and then the center-biased BMA is used to find the final motion vector as the global minimum is closer to the predicted center. Experimental results show that the proposed prediction model improves the performance of most fast BMAs such as the new three-step search and the four-step search, in terms of the motion compensation errors. In addition, the average computational requirement is also reduced. Jie-Bin Xu, Lai-Man Po, Chok-Kwan Cheung |
ICIP (3) | 2 |
| 1997 | Wavelet Transform Based Variable Tree Size Fractal Video CodingabstractA new wavelet video compression scheme using adaptive fractal codi ng i s pr esented i n t his paper. Through pyramidal w avelet t ransform, each vi deo f rame i s decomposed into multiresolution subbands and then organized into a set of wavelet subtrees to represent motion activities. After mu ltiresolution mo tion d etection, th ese wavelet subtrees are classified into motion and non-motion subtress. The codi ng of non-mot ion s ubtrees i s straightforward and s imple, while the motion subtress are adaptively separated and then encoded us ing variable tree size fractal coding. Experimental r esults s how t hat t he proposed scheme provides a superior performance in terms of PSNR as well as the subjective quality at ¦ low bit-rates. I. Lai-Man Po, Ying-Lin Yu |
ICIP (2) | 2 |
| 1996 | A novel four-step search algorithm for fast block motion estimationabstractBased on the real world image sequence's characteristic of center-biased motion vector distribution, a new four-step search (4SS) algorithm with center-biased checking point pattern for fast block motion estimation is proposed in this paper. A halfway-stop technique is employed in the new algorithm with searching steps of 2 to 4 and the total number of checking points is varied from 17 to 27. Simulation results show that the proposed 4SS performs better than the well-known three-step search and has similar performance to the new three-step search (N3SS) in terms of motion compensation errors. In addition, the 4SS also reduces the worst-case computational requirement from 33 to 27 search points and the average computational requirement from 21 to 19 search points, as compared with N3SS. Lai-Man Po, Wing-Chung Ma |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1995 | A new center-biased search algorithm for block motion estimationabstractExperimental results show that the block motion field of a real world image sequence is usually gentle, smooth, and varies slowly, as a result in a center-biased global minimum motion vector distribution instead of an uniform distribution. Based on this characteristic of the image sequence a new four-step search (4SS) algorithm with a center-biased checking point pattern for fast block motion estimation is proposed. A variable searching-step technique is employed in the proposed algorithm with a minimum of 2 searching steps and a maximum of 4. The total number of checking points is varied from 17 to 27. Simulation results show that, as compared to the well-known three-step search, the proposed 4SS produces smaller motion compensation errors with a smaller computational requirement. The 4SS also possesses hardware-oriented features regularity and simplicity. Lai-Man Po, Wing-Chung Ma |
ICIP | 1 |
| 1995 | Fractal color image compression using vector distortion measureabstractA fractal monochrome image compression technique proposed by Jacquin is investigated for 24-bit true color image by encoding the RGB components' images independently. To exploit the spectral redundancy in RGB components, the root mean square error distortion measure in grayscale space is extended to 3-dimensional color space for fractal-based color image coding. Experimental results show that a 1.5 compression ratio improvement can be obtained using the vector distortion measure in fractal coding with fixed image partition as compared to separate fractal coding in RGB images. In addition, adaptive fractal color image coding based on quadtree partition is also proposed which can obtain a good trade-off between compression ratio and image fidelity. Lai-Man Po |
ICIP (3) | 2 |
| 1994 | Address predictive color quantization image compression for multimedia applicationsabstractMost modem workstations and personal computers are using 8-bit color mapped graphics architecture for displaying color images. To display a decompressed color image of the conventional DCT-based compression algorithms on these palette-based display systems, the decoded image must first be color quantized. To avoid the computational intensive color quantization process and provide a fast decoding process, new address predictive color quantization image compression schemes for multimedia applications are proposed in this paper. A closest pairs color palette ordering technique is also proposed for effectively exploited the redundancy of the palettized image. Computer simulation results are given to show the performance of the new compression algorithms in terms of bit rate and signal to noise ratio.> Lai-Man Po, Wen-Tao Tan, Chi-Ho Chan |
ICASSP (5) | 1 |
| 1994 | Adaptive dimensionality reduction techniques for tree-structured vector quantizationabstractIt is well-known that the exorbitant complexity of vector quantization's codebook design and encoding process have hindered the vector quantization based coding systems for real-time implementations. The-structured vector quantization (TSVQ) is one of the most efficient techniques to solve these problems. However, the encoder's memory requirement has increased as compared with the conventional full search vector quantization. To tackle this drawback and further reduce the searching complexity of TSVQ, new multisubspace tree-structured vector quantization design and encoding techniques are proposed in this paper. During the tree codebook generation process, each nonterminal node of the tree is associated with a partition (or region) and the child nodes are generated by further partitioning of these regions individually. Thus, different distortion measures can be used in different partitions for the node splitting. Based on this idea, the proposed multisubspace TSVQ design algorithms perform the vector quantization in spatial domain while using specially designed subspace distortions in transform domain as cost functions in the optimization process. Due to the energy compaction property of orthonormal transforms, dimensionality reduction is permitted by disregarding the components of minimal energy. The conventional Euclidean distortion is, therefore, adaptively replaced by dimension reduced transform domain subspace distortions based on the local statistics of the partition. Experimental results show that exceptionally low subspace dimension can be used in multisubspace TSVQ based on a fixed transform domain to obtain a similar performance as the conventional TSVQ or single subspace TSVQ. For example, a 256-level and 4x4 image multisubspace TSVQ using binary tree and 1-dimensional subspace distortions achieved almost 63 times computational reduction and 5 times storage reduction when compared with conventional full search vector quantization. In addition, the proposed generalized multisubspace TSVQ design algorithm is a general TSVQ design algorithm and which can also be utilized as a fast codebook design algorithm. Lai-Man Po, Chok-Ki Chan |
IEEE Trans. Commun. | 1 |
| 1993 | Multi-subspace tree-structured vector quantizer design algorithms
Lai-Man Po |
ICASSP (5) | 1 |
| 1992 | A complexity reduction technique for image vector quantizationabstractA technique for reducing the complexity of spatial-domain image vector quantization (VQ) is proposed. The conventional spatial domain distortion measure is replaced by a transform domain subspace distortion measure. Due to the energy compaction properties of image transforms, the dimensionality of the subspace distortion measure can be reduced drastically without significantly affecting the performance of the new quantizer. A modified LBG algorithm incorporating the new distortion measure is proposed. Unlike conventional transform domain VQ, the codevector dimension is not reduced and a better image quality is guaranteed. The performance and design considerations of a real-time image encoder using the techniques are investigated. Compared with spatial domain a speed up in both codebook design time and search time is obtained for mean residual VQ, and the size of fast RAM is reduced by a factor of four. Degradation of image quality is less than 0.4 dB in PSNR. Chok-Ki Chan, Lai-Man Po |
IEEE Trans. Image Process. | 2 |