Lap-Pui Chau

dblp:03/5597 · DBLP profile ↗
← Back
171ranked-venue papers
8as first author
50since 2021 · last 2026
0000-0003-4932-0593ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 111 · 7 first-author · 25 since 2021Systems, architecture and hardware · 32 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 30 · 22 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Human-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Egocentric human-object interaction detection: A new benchmark and method
Kunyuan Deng, Yi Wang 0068, Lap-Pui Chau
Expert Syst. Appl.3
2026 CaRe-Ego: Contact-aware relationship modeling for egocentric interactive hand-object segmentation
Yuejiao Su, Yi Wang 0068, Lap-Pui Chau
Expert Syst. Appl.3
2026 LaSSM: Efficient Semantic-Spatial Query Decoding via Local Aggregation and State Space Models for 3D Instance Segmentation
abstract
Query-based 3D scene instance segmentation from point clouds has attained notable performance. However, existing methods suffer from the query initialization dilemma due to the sparse nature of point clouds and rely on computationally intensive attention mechanisms in query decoders. We accordingly introduceLaSSM, prioritizing simplicity and efficiency while maintaining competitive performance. Specifically, we propose a hierarchical semantic-spatial query initializer to derive the query set from superpoints by considering both semantic cues and spatial distribution, achieving comprehensive scene coverage and accelerated convergence. We further present a coordinate-guided state space model (SSM) decoder that progressively refines queries. The novel decoder features a local aggregation scheme that restricts the model to focus on geometrically coherent regions and a spatial dual-path SSM block to capture underlying dependencies within the query set by integrating associated coordinates information. Our design enables efficient instance prediction, avoiding the incorporation of noisy information and reducing redundant computation. LaSSM ranksfirst placeon the latest ScanNet++ V2 leaderboard, outperforming the previous best method by 2.5% mAP with only 1/3 FLOPs, demonstrating its superiority in challenging large-scale scene instance segmentation. LaSSM also achieves competitive performance on ScanNet V2, ScanNet200, S3DIS and ScanNet++ V1 benchmarks with less computational cost. Extensive ablation studies and qualitative results validate the effectiveness of our design. The code and weights are available at https://github.com/RayYoh/LaSSM.
Yi Wang 0068, Yawen Cui, Moyun Liu, Lap-Pui Chau
IEEE Trans. Circuits Syst. Video Technol.5
2026 Fuzzy-Aware Loss for Source-Free Domain Adaptation in Visual Emotion Recognition
abstract
Source-free domain adaptation in visual emotion recognition (SFDA-VER) is a highly challenging task that re quires adapting VER models to the target domain without relying on source data, which is of great significance for data privacy protection. However, due to the unignorable disparities between visual emotion data and traditional image classification data, existing SFDA methods perform poorly on this task. In this paper, we investigate the SFDA-VER task from a fuzzy perspective and identify two key issues: fuzzy emotion labels and fuzzy pseudo-labels. These issues arise from the inherent uncertainty of emotion annotations and the potential mispredictions in pseudo labels. To address these issues, we propose a novel fuzzy aware loss (FAL) to enable the VER model to better learn and adapt to new domains under fuzzy labels. Specifically, FAL modifies the standard cross entropy loss and focuses on adjusting the losses of non-predicted categories, which prevents a large number of uncertain or incorrect predictions from overwhelming the VER model during adaptation. In addition, we provide a theoretical analysis of FAL and prove its robustness in handling the noise in generated pseudo-labels. Extensive experiments on 26 domain adaptation sub-tasks across three benchmark datasets demonstrate the effectiveness of our method. Code is available at: https://github.com/zhengyinghit/FAL.
Ying Zheng 0009, Yiyi Zhang 0001, Yi Wang 0068, Lap-Pui Chau
IEEE Trans. Fuzzy Syst.4
2026 EVA02-AT: Egocentric Video-Language Understanding With Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization
abstract
Egocentric video-language understanding demands both high efficiency and accurate spatial-temporal modeling. Existing approaches face three key challenges: 1) Excessive pre-training cost arising from multi-stage pre-training pipelines, 2) Ineffective spatial-temporal encoding due to manually split 3D rotary positional embeddings that hinder feature interactions, and 3) Imprecise learning objectives in soft-label multi-instance retrieval, which neglect negative pair correlations. In this paper, we introduce EVA02-AT, a suite of EVA02-based video-language foundation models tailored to egocentric video understanding tasks. EVA02-AT first efficiently transfers an image-based CLIP model into a unified video encoder via a single-stage pretraining. Second, instead of applying rotary positional embeddings to isolated dimensions, we introduce spatial-temporal rotary positional embeddings along with joint attention, which can effectively encode both spatial and temporal information on the entire hidden dimension. This joint encoding of spatial-temporal features enables the model to learn cross-axis relationships, which are crucial for accurately modeling motion and interaction in videos. Third, focusing on multi-instance video-language retrieval tasks, we introduce the Symmetric Multi-Similarity (SMS) loss and a novel training framework that advances all soft labels for both positive and negative pairs, providing a more precise learning objective. Extensive experiments on Ego4D, EPIC-Kitchens-100, and Charades-Ego under zero-shot and fine-tuning settings demonstrate that EVA02-AT achieves state-of-the-art performance across diverse egocentric video-language tasks with fewer parameters. Models with our SMS loss also show significant performance gains on multi-instance retrieval benchmarks. Our code and models are publicly available at https://github.com/xqwang14/EVA02-AT.
Yi Wang 0068, Lap-Pui Chau
IEEE Trans. Image Process.3
2026 PromptSR: Cascade Prompting for Lightweight Image Super-Resolution
abstract
Although the lightweight Vision Transformer has significantly advanced image super-resolution (SR), it faces the inherent challenge of a limited receptive field due to the window-based self-attention modeling. The quadratic computational complexity relative to window size restricts its ability to use a large window size for expanding the receptive field while maintaining low computational costs. To address this challenge, we propose PromptSR, a novel prompt-empowered lightweight image SR method. The core component is the proposed cascade prompting block (CPB), which enhances global information access and local refinement via three cascaded prompting layers: a global anchor prompting layer (GAPL) and two local prompting layers (LPLs). The GAPL leverages downscaled features as anchors to construct low-dimensional anchor prompts (APs) through cross-scale attention, significantly reducing computational costs. These APs, with enhanced global perception, are then used to provide global prompts, efficiently facilitating long-range token connections. The two LPLs subsequently combine category-based self-attention and window-based self-attention to refine the representation in a coarse-to-fine manner. They leverage attention maps from the GAPL as additional global prompts, enabling them to perceive features globally at different granularities for adaptive local refinement. In this way, the proposed CPB effectively combines global priors and local details, significantly enlarging the receptive field while maintaining the low computational costs of our PromptSR. The experimental results demonstrate the superiority of our method, which outperforms state-of-the-art lightweight SR methods in quantitative, qualitative, and complexity evaluations. Our code will be released at https://github.com/wenyang001/PromptSR.
Wenyang Liu, Jianjun Gao 0005, Kejun Wu, Yi Wang 0068, Kim-Hui Yap, Lap-Pui Chau
IEEE Trans. Multim.7
2025 ANNEXE: Unified Analyzing, Answering, and Pixel Grounding for Egocentric Interaction
abstract
Egocentric interaction perception is one of the essential branches in investigating human-environment interaction, which lays the basis for developing next-generation intelligent systems. However, existing egocentric interaction understanding methods cannot yield coherent textual and pixel-level responses simultaneously according to user queries, which lacks flexibility for varying downstream application requirements. To comprehend egocentric interactions exhaustively, this paper presents a novel task named Egocentric Interaction Reasoning and pixel Grounding (Ego-IRG). Taking an egocentric image with the query as input, Ego-IRG is the first task that aims to resolve the interactions through three crucial steps: analyzing, answering, and pixel grounding, which results in fluent textual and fine-grained pixel-level responses. Another challenge is that existing datasets cannot meet the conditions for the Ego-IRG task. To address this limitation, this paper creates the Ego-IRGBench dataset based on extensive manual efforts, which includes over 20k egocentric images with 1.6 million queries and corresponding multimodal responses about interactions. Moreover, we design a unified ANNEXE model to generate text- and pixel-level outputs utilizing multimodal large language models, which enables a comprehensive interpretation of egocentric interactions. The experiments on the Ego-IRGBench exhibit the effectiveness of our ANNEXE model compared with other works.
Yuejiao Su, Yi Wang 0068, Qiongyang Hu, Chuang Yang 0003, Lap-Pui Chau
CVPR5
2025 OccProphet: Pushing the Efficiency Frontier of Camera-Only 4D Occupancy Forecasting with an Observer-Forecaster-Refiner Framework
abstract
Predicting variations in complex traffic environments is crucial for the safety of autonomous driving. Recent advancements in occupancy forecasting have enabled forecasting future 3D occupied status in driving environments by observing historical 2D images. However, high computational demands make occupancy forecasting less efficient during training and inference stages, hindering its feasibility for deployment on edge agents. In this paper, we propose a novel framework, \textit{i.e.}, OccProphet, to efficiently and effectively learn occupancy forecasting with significantly lower computational requirements while improving forecasting accuracy. OccProphet comprises three lightweight components: Observer, Forecaster, and Refiner. The Observer extracts spatio-temporal features from 3D multi-frame voxels using the proposed Efficient 4D Aggregation with Tripling-Attention Fusion, while the Forecaster and Refiner conditionally predict and refine future occupancy inferences. Experimental results on nuScenes, Lyft-Level5, and nuScenes-Occupancy datasets demonstrate that OccProphet is both training- and inference-friendly. OccProphet reduces 58\%$\sim$78\% of the computational cost with a 2.6$\times$ speedup compared with the state-of-the-art Cam4DOcc. Moreover, it achieves 4\%$\sim$18\% relatively higher forecasting accuracy. Code and models are publicly available at https://github.com/JLChen-C/OccProphet.
Huaiyuan Xu, Yi Wang 0068, Lap-Pui Chau
ICLR4
2025 Restoration of Bitstream-Corrupted Images: A Mamba-based Thumbnail-guided Network
abstract
This paper investigates the real-world JPEG image restoration problem with bit errors on the compressed bitstream. To mimic the effect of bit errors encountered in real images, we automatically inject various bit errors to generate damaged images, thereby simulating the bitstream-corrupted JPEG images in real situations. The image restoration problem is proposed to recover these images caused by bit errors that conventional decoders cannot perfectly decode. Typically, when a bit stream containing bit errors is decoded by the robust decoder, the resulting image exhibits two distinct characteristics: color casts and block shifts. To solve those problems, we propose a Mamba-based thumbnail-guided network to address the impact of color casts and block shifts on the image. The proposed framework is structurally divided into three blocks. Firstly, we use a feature aggregation (FA) block to reassemble the information from the corrupted image and the thumbnail image into a more acceptable input format for the subsequent networks, allowing for self-adjustment of the input format. Then, we design a point-to-point restoration (PPR) block as a decoder to parse the features of inputs and thumbnails to generate coarse images. Finally, with the guidance of coarse images, a pyramid fusion (PF) block is used to generate the refined images. Extensive experimental results demonstrate our model outperforms state-of-the-art methods. Ablation studies and comparisons with super-resolution methods illustrate the effectiveness of our approach. The code will be available at https://github.com/HU1qy/MambaThumbnail.
Qiongyang Hu, Yi Wang 0068, Lap-Pui Chau
ISCAS3
2025 Probabilistic Mixture of Hyperbolic Mamba for Few-Shot Class-Incremental Learning
abstract
Few-shot class-incremental learning (FSCIL) grapples with the dual challenge of learning new classes from minimal labeled training data while alleviating catastrophic forgetting of previous learned classes. Compared with previous methods employing static adaptation on specific parameters, current works verify that dynamic weights and sequence modeling in Selective State Space Models (SSMs) can capture distinctive feature drifts in FSCIL. However, the flattening operation in SSMs fragments the latent semantic relationship, where the resulting task isolation and representation degeneration are detrimental to FSCIL. Toward this issue, this paper presents a novel framework named Probabilistic Mixture of Hyperbolic State Space Experts (PmH-SSE) for FSCIL. First, since SSMs rely on scanning as an alternative to self-attention, the Hyperbolic state space model with multi-scale hybrid scan is built to facilitate few-shot learning by providing an extra Hyperbolic geometry that encodes hierarchical relationships. Moreover, we propose the probabilistic mixture of Mamba to increase the model's flexibility in handling non-stationary data streams in FSCIL and enhance the stability of high-parameter models in few-shot conditions. Finally, under the same experimental conditions, the proposed PmH-SSE demonstrates superior performance in comprehensive experiments. The codes are available at https://github.com/yawencui/PmH-SSE.
Yawen Cui, Wenbin Zou, Huiping Zhuang, Yi Wang 0068, Lap-Pui Chau
ACM Multimedia5
2025 Towards Blind Bitstream-corrupted Video Recovery: A Visual Foundation Model-driven Framework
abstract
Video signals are vulnerable in multimedia communication and storage systems, as even slight bitstream-domain corruption can lead to significant pixel-domain degradation. To recover faithful spatio-temporal content from corrupted inputs, bitstream-corrupted video recovery has recently emerged as a challenging and understudied task. However, existing methods require time-consuming and labor-intensive annotation of corrupted regions for each corrupted video frame, resulting in a large workload in practice. In addition, high-quality recovery remains difficult as part of the local residual information in corrupted frames may mislead feature completion and successive content recovery. In this paper, we propose the first blind bitstream-corrupted video recovery framework that integrates visual foundation models with recovery model, which is adapted to different types of corruption and bitstream-level prompts. Within the framework, the proposed Detect Any Corruption (DAC) model leverages the rich priors of the visual foundation model while incorporating bitstream and corruption knowledge to enhance corruption localization and blind recovery. Additionally, we introduce a novel Corruption-aware Feature Completion (CFC) module, which adaptively processes residual contributions based on high-level corruption understanding. With VFM-guided hierarchical feature augmentation and high-level coordination in a mixture-of-residual-experts (MoRE) structure, our method suppresses artifacts and enhances informative residuals. Comprehensive evaluations show that the proposed method achieves outstanding performance in bitstream-corrupted video recovery without requiring a manually labeled mask sequence. The demonstrated effectiveness will help to realize improved user experience, wider application scenarios, and more reliable multimedia communication and storage systems.
Kejun Wu, Yi Wang 0068, Kim-Hui Yap, Lap-Pui Chau
ACM Multimedia6
2025 GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
abstract
The significance of informative and robust point representations has been widely acknowledged for 3D scene understanding. Despite existing self-supervised pre-training counterparts demonstrating promising performance, the model collapse and structural information deficiency remain prevalent due to insufficient point discrimination difficulty, yielding unreliable expressions and suboptimal performance. In this paper, we present GaussianCross, a novel cross-modal self-supervised 3D representation learning architecture integrating feed-forward 3D Gaussian Splatting (3DGS) techniques to address current challenges. GaussianCross seamlessly converts scale-inconsistent 3D point clouds into a unified cuboid-normalized Gaussian representation without missing details, enabling stable and generalizable pre-training. Subsequently, a tri-attribute adaptive distillation splatting module is incorporated to construct a 3D feature field, facilitating synergetic feature capturing of appearance, geometry, and semantic cues to maintain cross-modal consistency. To validate GaussianCross, we perform extensive evaluations on various benchmarks, including ScanNet, ScanNet200, and S3DIS. In particular, GaussianCross shows a prominent parameter and data efficiency, achieving superior performance through linear probing (<0.1% parameters) and limited data training (1% of scenes) compared to state-of-the-art methods. Furthermore, GaussianCross demonstrates strong generalization capabilities, improving the full fine-tuning accuracy by 9.3% mIoU and 6.1% AP50 on ScanNet200 semantic and instance segmentation tasks, respectively, supporting the effectiveness of our approach. The code, weights, and visualizations are publicly available at https://rayyoh.github.io/GaussianCross/.
Yi Wang 0068, Moyun Liu, Lap-Pui Chau
ACM Multimedia5
2025 Semantic Representation Attack against Aligned Large Language Models
abstract
Large Language Models (LLMs) increasingly employ alignment techniques to prevent harmful outputs. Despite these safeguards, attackers can circumvent them by crafting prompts that induce LLMs to generate harmful content. Current methods typically target exact affirmative responses, suffering from limited convergence, unnatural prompts, and high computational costs. We introduce semantic representation attacks, a novel paradigm that fundamentally reconceptualizes adversarial objectives against aligned LLMs. Rather than targeting exact textual patterns, our approach exploits the semantic representation space that can elicit diverse responses that share equivalent harmful meanings. This innovation resolves the inherent trade-off between attack effectiveness and prompt naturalness that plagues existing methods. Our Semantic Representation Heuristic Search (SRHS) algorithm efficiently generates semantically coherent adversarial prompts by maintaining interpretability during incremental search. We establish rigorous theoretical guarantees for semantic convergence and demonstrate that SRHS achieves unprecedented attack success rates (89.4% averaged across 18 LLMs, including 100% on 11 models) while significantly reducing computational requirements. Extensive experiments show that our method consistently outperforms existing approaches.
Jiawei Lian, Jianhong Pan, Lefan Wang, Yi Wang 0068, Shaohui Mei, Lap-Pui Chau
NeurIPS6
2025 REAL: Representation enhanced analytic learning for exemplar-free class-incremental learning
Run He, Di Fang 0004, Yizhu Chen, Kai Tong, Cen Chen 0002, Yi Wang 0068, Lap-Pui Chau, Huiping Zhuang
Knowl. Based Syst.7
2025 PADetBench: Towards benchmarking texture- and patch-based physical attacks against object detection
Jiawei Lian, Jianhong Pan, Lefan Wang, Yi Wang 0068, Shaohui Mei, Lap-Pui Chau
Knowl. Based Syst.6
2025 Foundation model-assisted interpretable vehicle behavior decision making
Shiyu Meng, Yi Wang 0068, Yawen Cui, Lap-Pui Chau
Knowl. Based Syst.4
2025 SGIFormer: Semantic-Guided and Geometric-Enhanced Interleaving Transformer for 3D Instance Segmentation
abstract
In recent years, transformer-based models have exhibited considerable potential in point cloud instance segmentation. Despite the promising performance achieved by existing methods, they encounter challenges such as instance query initialization problems and excessive reliance on stacked layers, rendering them incompatible with large-scale 3D scenes. This paper introduces a novel method, named SGIFormer, for 3D instance segmentation, which is composed of the Semantic-guided Mix Query (SMQ) initialization and the Geometric-enhanced Interleaving Transformer (GIT) decoder. Specifically, the principle of our SMQ initialization scheme is to leverage the predicted voxel-wise semantic information to implicitly generate the scene-aware query, yielding adequate scene prior and compensating for the learnable query set. Subsequently, we feed the formed overall query into our GIT decoder to alternately refine instance query and global scene features for further capturing fine-grained information and reducing complex design intricacies simultaneously. To emphasize geometric property, we consider bias estimation as an auxiliary task and progressively integrate shifted point coordinates embedding to reinforce instance localization. SGIFormer attains state-of-the-art performance on ScanNet V2, ScanNet200, S3DIS datasets, and the challenging high-fidelity ScanNet ++ benchmark, striking a balance between accuracy and efficiency. The code, weights, and demo videos are publicly available athttps://rayyoh.github.io/SGIFormer/.
Yi Wang 0068, Moyun Liu, Lap-Pui Chau
IEEE Trans. Circuits Syst. Video Technol.4
2025 SignEye: Traffic Sign Interpretation From Vehicle First-Person View
abstract
Traffic signs play a key role in assisting autonomous driving systems (ADS) by enabling the assessment of vehicle behavior in compliance with traffic regulations and providing navigation instructions. However, current works are limited to basic sign understanding without considering the egocentric vehicle’s spatial position, which fails to support further regulation assessment and direction navigation. Following the above issues, we introduce a new task: traffic sign interpretation from the vehicle’s first-person view, referred to asTSI-FPV. Meanwhile, we develop a traffic guidance assistant (TGA) scenario application to re-explore the role of traffic signs in ADS as a complement to popular autonomous technologies (such as obstacle perception). Notably, TGA is not a replacement for electronic map navigation; rather, TGA can be an automatic tool for updating it and complementing it in situations such as offline conditions or temporary sign adjustments. Lastly, a spatial and semantic logic-aware stepwise reasoning pipeline (SignEye) is constructed to achieve the TSI-FPV and TGA, and an application-specific dataset (Traffic-CN) is built. Experiments show that TSI-FPV and TGA are achievable via our SignEye trained on Traffic-CN. The results also demonstrate that the TGA can provide complementary information to ADS beyond existing popular autonomous technologies.
Chuang Yang 0003, Xu Han 0019, Tao Han 0002, Yuejiao Su, Junyu Gao 0001, Hongyuan Zhang 0001, Yi Wang 0068, Lap-Pui Chau
IEEE Trans. Intell. Transp. Syst.8
2025 ByteNet: Rethinking Multimedia File Fragment Classification Through Visual Perspectives
abstract
Multimedia file fragment classification (MFFC) aims to identify file fragment types, e.g., image/video, audio, and text without system metadata. It is of vital importance in multimedia storage and communication. Existing MFFC methods typically treat fragments as 1D byte sequences and emphasize the relations between separate bytes (interbytes) for classification. However, the more informative relations inside bytes (intrabytes) are overlooked and seldom investigated. By looking inside bytes, the bit-level details of file fragments can be accessed, enabling a more accurate classification. Motivated by this, we first proposeByte2Image, a novel visual representation model that incorporates previously overlooked intrabyte information into file fragments and reinterprets these fragments as 2D grayscale images. This model involves a sliding byte window to reveal the intrabyte information and a rowwise stacking of intrabyte n-grams for embedding fragments into a 2D space. Thus, complex interbyte and intrabyte correlations can be mined simultaneously using powerful vision networks. Additionally, we propose an end-to-end dual-branch networkByteNetto enhance robust correlation mining and feature representation. ByteNet makes full use of the raw 1D byte sequence and the converted 2D image through a shallow byte branch feature extraction (BBFE) and a deep image branch feature extraction (IBFE) network. In particular, the BBFE, composed of a single fully-connected layer, adaptively recognizes the co-occurrence of several some specific bytes within the raw byte sequence, while the IBFE, built on a vision Transformer, effectively mines the complex interbyte and intrabyte correlations from the converted image. Experiments on the two representative benchmarks, including 14 cases, validate that our proposed method outperforms state-of-the-art approaches on different cases by up to 12.2%.
Wenyang Liu, Kejun Wu, Yi Wang 0068, Kim-Hui Yap, Lap-Pui Chau
IEEE Trans. Multim.6
2025 3DGeoDet: General-Purpose Geometry-Aware Image-Based 3D Object Detection
abstract
This paper proposes 3DGeoDet, a novel geometry-aware 3D object detection approach that effectively handles single- and multi-view RGB images in indoor and outdoor environments, showcasing its general-purpose applicability. The key challenge for image-based 3D object detection tasks is the lack of 3D geometric cues, which leads to ambiguity in establishing correspondences between images and 3D representations. To tackle this problem, 3DGeoDet generates efficient 3D geometric representations in both explicit and implicit manners based on predicted depth information. Specifically, we utilize the predicted depth to learn voxel occupancy and optimize the voxelized 3D feature volume explicitly through the proposed voxel occupancy attention. To further enhance 3D awareness, the feature volume is integrated with an implicit 3D representation, the truncated signed distance function (TSDF). Without requiring supervision from 3D signals, we significantly improve the model's comprehension of 3D geometry by leveraging intermediate 3D representations and achieve end-to-end training. Our approach surpasses the performance of state-of-the-art image-based methods on both single- and multi-view benchmark datasets across diverse environments, achieving a 9.3 [email protected] improvement on the SUN RGB-D dataset, a 3.3 [email protected] improvement on the ScanNetV2 dataset, and a 0.19$\text{AP}_{\text{3D}}[email protected] improvement on the KITTI dataset. The project page is available at:https://cindy0725.github.io/3DGeoDet/
Yi Wang 0068, Yawen Cui, Lap-Pui Chau
IEEE Trans. Multim.4
2024 SinSR: Diffusion-Based Image Super-Resolution in a Single Step
abstract
While super-resolution (SR) methods based on diffusion models exhibit promising results, their practical application is hindered by the substantial number of required inference steps. Recent methods utilize the degraded images in the initial state, thereby shortening the Markov chain. Nevertheless, these solutions either rely on a precise formulation of the degradation process or still necessitate a relatively lengthy generation path (e.g., 15 iterations). To enhance inference speed, we propose a simple yet effective method for achieving single-step SR generation, named SinSR. Specifically, we first derive a deterministic sampling process from the most recent state-of-the-art (SOTA) method for accelerating diffusion-based SR. This allows the mapping between the input random noise and the generated high-resolution image to be obtained in a reduced and acceptable number of inference steps during training. We show that this deterministic mapping can be distilled into a student model that performs SR within only one inference step. Additionally, we propose a novel consistency-preserving loss to simultaneously leverage the ground-truth image during the distillation process, ensuring that the performance of the student model is not solely bound by the feature manifold of the teacher model, resulting in further performance improvement. Extensive experiments conducted on synthetic and real-world datasets demonstrate that the proposed method can achieve comparable or even superior performance compared to both previous SOTA methods and the teacher model, in just one sampling step, resulting in a remarkable up to × 10 speedup for inference. Our code will be released at https://github.com/wyf0912/SinSR/.
Yufei Wang 0006, Wenhan Yang, Yaohui Wang 0001, Lanqing Guo, Lap-Pui Chau, Ziwei Liu 0002, Yu Qiao 0001, Alex Chichung Kot, Bihan Wen
CVPR6
2024 Weakly-Supervised Crowd Counting with Token Attention and Fusion: A Simple and Effective Baseline
abstract
Conventional crowd counting methods exploit a large number of point annotations to train regression-based neural networks for density map estimation. However, laborious point annotations of human heads (strong supervision) are required in training. This paper presents a simple and effective crowd counting method with only image-level count annotations, i.e., the number of people in an image (weak supervision). Specifically, we first investigate three backbone networks and find the significance of the global information extracted by self-attention for weakly-supervised crowd counting. Then, we propose an effective network composed of a Transformer backbone and token channel attention module (T-CAM) in the counting head, where the attention in channels of tokens can compensate for the self-attention between tokens of the Transformer. Finally, a simple token fusion is proposed to obtain global information. Experimental results on two representative crowd counting benchmarks show the superiority of the proposed method, with an average 10% relative improvement compared with baselines. The code is publicly available at https://github.com/WangyiNTU/WSCC_TAF.
Yi Wang 0068, Qiongyang Hu, Lap-Pui Chau
ICASSP3
2024 Depth-powered Moving-obstacle Segmentation Under Bird-eye-view for Autonomous Driving
abstract
Sensing the moving obstacles accurately under birdeye view (BEV) is the foundation for reliable autonomous driving, providing straightforward information for the downstream tasks. However, accurately segmenting moving obstacles only through monocular camera views is extremely difficult due to the lack of depth information. It can easily generate the projected depth information from point clouds, but its sparsity provides incomplete depth information. Therefore, in this paper, we propose a dense depth-powered framework, dubbed DPMoSeg, to generate dense moving-obstacle segmentation observations under BEV space. To better represent the depth prediction, we design a sparse-dense attention module to fully combine the knowledge across non- homogeneous and homogeneous representations. The experimental results demonstrate the effectiveness and superiority of our proposed framework.
Shiyu Meng, Yi Wang 0068, Lap-Pui Chau
ISCAS3
2024 Few-shot Class-agnostic Counting with Occlusion Augmentation and Localization
abstract
Most existing few-shot class-agnostic counting (FCAC) methods follow the extract-and-compare pipeline to count all instances of an arbitrary category in the query image given a few exemplars. However, these methods generate the density map rather than the exact instance location for counting, which is less intuitive and accurate than the latter. Besides, how to alleviate the problem of occlusion is ignored in most existing work. To solve the above problems, this paper proposes an Occlusion-Augmented Localization Network (OALNet), which extracts multiple occluded features of exemplars for comparison and utilizes the precise position of instances for more accurate and confident counting results. Specifically, the OALNet is in an extract-and-attention manner. It includes an Occluded Feature Generation module to deal with the occlusion problem in query images. Besides, the OALNet adopts the Feature Attention module to improve the extracted feature by self-attention and model the relationship between the exemplar features and query features by cross-attention. Compared with other FCAC methods, experimental results demonstrate that the proposed OALNet achieves superior performance.
Yuejiao Su, Yi Wang 0068, Lap-Pui Chau
ISCAS4
2024 F-OAL: Forward-only Online Analytic Learning with Fast Training and Low Memory Footprint in Class Incremental Learning
abstract
Online Class Incremental Learning (OCIL) aims to train models incrementally, where data arrive in mini-batches, and previous data are not accessible. A major challenge in OCIL is Catastrophic Forgetting, i.e., the loss of previously learned knowledge. Among existing baselines, replay-based methods show competitive results but requires extra memory for storing exemplars, while exemplar-free (i.e., data need not be stored for replay in production) methods are resource friendly but often lack accuracy. In this paper, we propose an exemplar-free approach—Forward-only Online Analytic Learning (F-OAL). Unlike traditional methods, F-OAL does not rely on back-propagation and is forward-only, significantly reducing memory usage and computational time. Cooperating with a pre-trained frozen encoder with Feature Fusion, F-OAL only needs to update a linear classifier by recursive least square. This approach simultaneously achieves high accuracy and low resource consumption. Extensive experiments on bench mark datasets demonstrate F-OAL’s robust performance in OCIL scenarios. Code is available at: https://github.com/liuyuchen-cz/F-OAL
Huiping Zhuang, Yuchen Liu 0001, Run He, Kai Tong, Ziqian Zeng, Cen Chen 0002, Yi Wang 0068, Lap-Pui Chau
NeurIPS8
2024 Image compressive sensing reconstruction via nonlocal low-rank residual-based ADMM framework
Junhao Zhang 0005, Kim-Hui Yap, Lap-Pui Chau, Ce Zhu
Comput. Vis. Image Underst.3
2024 Beyond Learned Metadata-Based Raw Image Reconstruction
Yufei Wang 0006, Yi Yu 0011, Wenhan Yang, Lanqing Guo, Lap-Pui Chau, Alex Chichung Kot, Bihan Wen
Int. J. Comput. Vis.5
2024 Intra- and inter-sector contextual information fusion with joint self-attention for file fragment classification
abstract
File fragment classification (FFC) aims to identify the file type of file fragments in memory sectors, which is of great importance in memory forensics and information security. Existing works focused on processing the bytes within sectors separately and ignoring contextual information between adjacent sectors. In this paper, we introduce a joint self-attention network (JSANet) for FFC to learn intra-sector local features and inter-sector contextual features. Specifically, we propose an end-to-end network with the byte, channel, and sector self-attention modules. Byte self-attention adaptively recognizes the intra-sector significant bytes, and channel self-attention re-calibrates the features between channels. Based on the insight that adjacent memory sectors are most likely to store a file fragment, sector self-attention leverages contextual information in neighboring sectors to enhance inter-sector feature representation. Extensive experiments on seven FFC benchmarks show the superiority of our method compared with state-of-the-art methods. Moreover, we construct VFF-16, a variable-length file fragment dataset to reflect file fragmentation. Integrated with sector self-attention, our method improves accuracy by more than 16.3% against the baseline on VFF-16, and the runtime achieves 5.1 s/GB with GPU acceleration. In addition, we extend our model to malware detection and show its applicability.
Yi Wang 0068, Wenyang Liu, Kejun Wu, Kim-Hui Yap, Lap-Pui Chau
Knowl. Based Syst.5
2024 A three-stream fusion and self-differential attention network for multi-modal crowd counting
Haihan Tang, Yi Wang 0068, Zhiping Lin 0001, Lap-Pui Chau, Huiping Zhuang
Pattern Recognit. Lett.4
2024 PEM: Perception Error Model for Virtual Testing of Autonomous Vehicles
abstract
Even though virtual testing of Autonomous Vehicles (AVs) has been well recognized as essential for safety assessment, AV simulators are still undergoing active development. One particular challenge is the problem of including the Sensing and Perception (S&P) subsystem into the virtual simulation loop in an efficient and effective manner. In this article, we define Perception Error Models (PEM), a virtual simulation component that can enable the analysis of the impact of perception errors on AV safety, without the need to model the sensors themselves. We propose a generalized data-driven procedure towards parametric modeling and evaluate it using Apollo, an open-source driving software, and nuScenes, a public AV dataset. Additionally, we implement PEMs in SVL, an open-source vehicle simulator. Furthermore, we demonstrate the usefulness of PEM-based virtual tests, by evaluating camera, LiDAR, and camera-LiDAR setups. Our virtual tests highlight limitations in the current evaluation metrics, and the proposed approach can help study the impact of perception errors on AV safety.
Andrea Piazzoni, Jim Cherian, Justin Dauwels, Lap-Pui Chau
IEEE Trans. Intell. Transp. Syst.4
2024 Portrait matting using an attention-based memory network
Shufeng Song, Lap-Pui Chau
Vis. Comput.2
2023 Bitstream-Corrupted JPEG Images are Restorable: Two-stage Compensation and Alignment Framework for Image Restoration
abstract
In this paper, we study a real-world JPEG image restoration problem with bit errors on the encrypted bitstream. The bit errors bring unpredictable color casts and block shifts on decoded image contents, which cannot be resolved by existing image restoration methods mainly relying on pre-defined degradation models in the pixel domain. To address these challenges, we propose a robust JPEG decoder, followed by a two-stage compensation and alignment framework to restore bitstream-corrupted JPEC images. Specifically, the robust JPEC decoder adopts an error-resilient mechanism to decode the corrupted JPEG bitstream. The two-stage framework is composed of the self-compensation and alignment (SCA) stage and the guided-compensation and alignment (GCA) stage. The SCA adaptively performs block-wise image color compensation and alignment based on the estimated color and block offsets via image content similarity. The GCA leverages the extracted low-resolution thumbnail from the JPEG header to guide full-resolution pixel-wise image restoration in a coarse-to-fine manner. It is achieved by a coarse-guided pix2pix network and a refine-guided bi-directional Laplacian pyramid fusion network. We conduct experiments on three benchmarks with varying degrees of bit error rates. Experimental results and ablation studies demonstrate the superiority of our proposed method. The code will be released at https://github.com/wenyang001/Two-ACIR.
Wenyang Liu, Yi Wang 0068, Kim-Hui Yap, Lap-Pui Chau
CVPR4
2023 Raw Image Reconstruction with Learned Compact Metadata
abstract
While raw images exhibit advantages over sRGB images (e.g., linearity and fine-grained quantization level), they are not widely used by common users due to the large storage requirements. Very recent works propose to compress raw images by designing the sampling masks in the raw image pixel space, leading to suboptimal image representations and redundant metadata. In this paper, we propose a novel framework to learn a compact representation in the latent space serving as the metadata in an end-to-end manner. Furthermore, we propose a novel sRGB-guided context model with the improved entropy estimation strategies, which leads to better reconstruction quality, smaller size of metadata, and faster speed. We illustrate how the proposed raw image compression scheme can adaptively allocate more bits to image regions that are important from a global perspective. The experimental results show that the proposed method can achieve superior raw image reconstruction results using a smaller size of the metadata on both uncompressed sRGB images and JPEG images. The code will be released at https://github.com/wyf0912/R2LCM
Yufei Wang 0006, Yi Yu 0011, Wenhan Yang, Lanqing Guo, Lap-Pui Chau, Alex Chichung Kot, Bihan Wen
CVPR5
2023 A Spatial-Focal Error Concealment Scheme for Corrupted Focal Stack Video
abstract
Focal stack image sequences can be regarded as successive frames of videos, which are densely captured by focusing on a stack of focal planes. This type of data is able to provide focus cues for display technologies. Before the displays on the user side, focal stack video is possibly corrupted during compression, storage and transmission chains, generating error frames on the decoder side. The error regions are difficult to be recovered due to the focal changes among frames. Conventional error concealment methods result in sharpness inconsistency between recovered regions and their spatial adjacent regions. Motivated by this, in this paper, we propose a spatial-focal error concealment scheme specialized for focal stack videos. The spatial adjacent regions around an error region are employed to reveal the prediction relations between error frame and focal adjacent frames. Gaussian blur filtering and Lucy-Richardson deblur filtering are applied to simulate the video focal changes. In this way, the error regions can be well recovered by exploiting the spatial-focal information. Experiment results show that the proposed scheme can achieve the highest objective quality in terms of PSNR and SSIM. It can also obtain the best subjective quality with sharpness consistency in recovered regions and without block effect.
Kejun Wu, Yi Wang 0068, Wenyang Liu, Kim-Hui Yap, Lap-Pui Chau
DCC5
2023 ExposureDiffusion: Learning to Expose for Low-light Image Enhancement
abstract
Previous raw image-based low-light image enhancement methods predominantly relied on feed-forward neural networks to learn deterministic mappings from low-light to normally-exposed images. However, they failed to capture critical distribution information, leading to visually undesirable results. This work addresses the issue by seamlessly integrating a diffusion model with a physics-based exposure model. Different from a vanilla diffusion model that has to perform Gaussian denoising, with the injected physics-based exposure model, our restoration process can directly start from a noisy image instead of pure noise. As such, our method obtains significantly improved performance and reduced inference time compared with vanilla diffusion models. To make full use of the advantages of different intermediate steps, we further propose an adaptive residual layer that effectively screens out the side-effect in the iterative refinement when the intermediate results have been already well-exposed. The proposed framework can work with both real-paired datasets, SOTA noise models, and different backbone networks. We evaluate the proposed method on various public benchmarks, achieving promising results with consistent improvements using different exposure models and backbones. Besides, the proposed method achieves better generalization capacity for unseen amplifying ratios and better performance than a larger feedforward neural model when few parameters are adopted. The code is released at https://github.com/wyf0912/ExposureDiffusion.
Yufei Wang 0006, Yi Yu 0011, Wenhan Yang, Lanqing Guo, Lap-Pui Chau, Alex Chichung Kot, Bihan Wen
ICCV5
2023 Nonlocal Low-Rank Residual Modeling for Image Compressive Sensing Reconstruction
abstract
The nonlocal low-rank (LR) modeling has proven to be an effective approach in image compressive sensing (CS) reconstruction, which starts by clustering similar patches using the nonlocal self-similarity (NSS) prior into nonlocal image groups and then imposes an L-R penalty on each nonlocal image group. However, most existing methods only approximate the LR matrix directly from the degraded nonlocal image group, which may lead to suboptimal LR matrix approximation and thus obtain unsatisfactory reconstruction results. This paper proposes a novel nonlocal low-rank residual (NLRR) approach for image CS reconstruction, which progressively approximates the underlying LR matrix by minimizing the LR residual. To do this, we first use the NSS prior to obtain a good estimate of the original nonlocal image group, and then the LR residual between the degraded nonlocal image group and the estimated nonlocal image group is minimized to derive a more accurate LR matrix. To ensure the optimization is both feasible and reliable, we employ an alternative direction multiplier method (ADMM) to solve the NLRR-based image CS reconstruction problem. Our experimental results show that the proposed NLRR algorithm achieves superior performance against many popular or state-of-the-art image CS reconstruction methods, both in objective metrics and subjective perceptual quality.
Junhao Zhang 0005, Kim-Hui Yap, Lap-Pui Chau, Ce Zhu
ICIP3
2023 Image Representation and Deep Inception-Attention for File-type and Malware Classification
abstract
File-type classification aims to recognize the file types of files/fragments without file-system metadata, which is essential for memory forensics and data recovery. In this paper, we introduce an image representation and deep inception-attention manner for file-type classification. Specifically, we consider file-type classification as an image classification problem. Raw data sequences in the memory block are converted to 2D binary images, enriching the representation ability and visualization while retaining the completeness of the bitstream. With binary images as inputs, we propose a deep inception-attention network to extract discriminate horizontal features and re-calibrate the weights of feature maps, and finally, predict file types. Experiments on a large-scale benchmark show the superiority of the proposed model. Moreover, our method can be extended to a similar application, like malware classification, and achieve outstanding performance.
Yi Wang 0068, Kejun Wu, Wenyang Liu, Kim-Hui Yap, Lap-Pui Chau
ISCAS5
2023 Removing Image Artifacts From Scratched Lens Protectors
abstract
A protector is placed in front of the camera lens for mobile devices to avoid damage, while the protector itself can be easily scratched accidentally, especially for plastic ones. The artifacts appear in a wide variety of patterns, making it difficult to see through them clearly. Removing image artifacts from the scratched lens protector is inherently challenging due to the occasional flare artifacts and the co-occurring interference within mixed artifacts. Though different methods have been proposed for some specific distortions, they seldom consider such inherent challenges. In our work, we consider the inherent challenges in a unified framework with two cooperative modules, which facilitate the performance boost of each other. We also collect a new dataset from the real world to facilitate training and evaluation purposes. The experimental results demonstrate that our method outperforms the baselines qualitatively and quantitatively. The code and datasets will be released at https://github.com/wyf0912/flare-removal
Yufei Wang 0006, Renjie Wan, Wenhan Yang, Bihan Wen, Lap-Pui Chau, Alex Chichung Kot
ISCAS5
2023 Bitstream-Corrupted Video Recovery: A Novel Benchmark Dataset and Method
abstract
The past decade has witnessed great strides in video recovery by specialist technologies, like video inpainting, completion, and error concealment. However, they typically simulate the missing content by manual-designed error masks, thus failing to fill in the realistic video loss in video communication (e.g., telepresence, live streaming, and internet video) and multimedia forensics. To address this, we introduce the bitstream-corrupted video (BSCV) benchmark, the first benchmark dataset with more than 28,000 video clips, which can be used for bitstream-corrupted video recovery in the real world. The BSCV is a collection of 1) a proposed three-parameter corruption model for video bitstream, 2) a large-scale dataset containing rich error patterns, multiple corruption levels, and flexible dataset branches, and 3) a new video recovery framework that serves as a benchmark. We evaluate state-of-the-art video inpainting methods on the BSCV dataset, demonstrating existing approaches' limitations and our framework's advantages in solving the bitstream-corrupted video recovery problem. The benchmark and dataset are released at https://github.com/LIUTIGHE/BSCV-Dataset.
Kejun Wu, Yi Wang 0068, Wenyang Liu, Kim-Hui Yap, Lap-Pui Chau
NeurIPS6
2023 Moving Towards Centers: Re-Ranking With Attention and Memory for Re-Identification
abstract
Re-ranking utilizes contextual information to optimize the initial ranking list of person or vehicle re-identification (re-ID), which boosts the retrieval performance at post-processing steps. This paper proposes a re-ranking network to predict the correlations between the probe and top-ranked neighbor samples. Specifically, all the feature embeddings of query and gallery images are expanded and enhanced by a linear combination of their neighbors, with the correlation prediction serving as discriminative combination weights. The combination process is equivalent to moving independent embeddings toward the identity centers, improving cluster compactness. For correlation prediction, we first aggregate the contextual information for probe's$k$-nearest neighbors via the Transformer encoder. Then, we distill and refine the probe-related features into the Contextual Memory cell via attention mechanism. Like humans that retrieve images by not only considering probe images but also memorizing the retrieved ones, the Contextual Memory produces multi-view descriptions for each instance. Finally, the neighbors are reconstructed with features fetched from the Contextual Memory, and a binary classifier predicts their correlations with the probe. Experiments on six widely-used person and vehicle re-ID benchmarks demonstrate the effectiveness of the proposed method. Especially, our method surpasses the state-of-the-art re-ranking approaches on large-scale datasets by a significant margin, i.e., with an average 4.83% CMC@1 and 14.83% mAP improvements on VERI-Wild, MSMT17, and VehicleID datasets.
Yunhao Zhou, Yi Wang 0068, Lap-Pui Chau
IEEE Trans. Multim.3
2022 Low-Light Image Enhancement with Normalizing Flow
abstract
To enhance low-light images to normally-exposed ones is highly ill-posed, namely that the mapping relationship between them is one-to-many. Previous works based on the pixel-wise reconstruction losses and deterministic processes fail to capture the complex conditional distribution of normally exposed images, which results in improper brightness, residual noise, and artifacts. In this paper, we investigate to model this one-to-many relationship via a proposed normalizing flow model. An invertible network that takes the low-light images/features as the condition and learns to map the distribution of normally exposed images into a Gaussian distribution. In this way, the conditional distribution of the normally exposed images can be well modeled, and the enhancement process, i.e., the other inference direction of the invertible network, is equivalent to being constrained by a loss function that better describes the manifold structure of natural images during the training. The experimental results on the existing benchmark datasets show our method achieves better quantitative and qualitative results, obtaining better-exposed illumination, less noise and artifact, and richer colors.
Yufei Wang 0006, Renjie Wan, Wenhan Yang, Haoliang Li, Lap-Pui Chau, Alex Chichung Kot
AAAI5
2022 TAFNet: A Three-Stream Adaptive Fusion Network for RGB-T Crowd Counting
abstract
In this paper, we propose a three-stream adaptive fusion network named TAFNet, which uses paired RGB and thermal images for crowd counting. Specifically, TAFNet is divided into one main stream and two auxiliary streams. We combine a pair of RGB and thermal images to constitute the input of main stream. Two auxiliary streams respectively exploit RGB image and thermal image to extract modality-specific features. Besides, we propose an Information Improvement Module (IIM) to fuse the modality-specific features into the main stream adaptively. Experiment results on RGBT-CC dataset show that our method achieves more than 20% improvement on mean average error and root mean squared error compared with state-of-the-art method. The source code will be publicly available at https://github.com/TANGHAIHAN/TAFNet.
Haihan Tang, Yi Wang 0068, Lap-Pui Chau
ISCAS3
2022 Deep Spatial-Angular Regularization for Light Field Imaging, Denoising, and Super-Resolution
abstract
Coded aperture is a promising approach for capturing the 4-D light field (LF), in which the 4-D data are compressively modulated into 2-D coded measurements that are further decoded by reconstruction algorithms. The bottleneck lies in the reconstruction algorithms, resulting in rather limited reconstruction quality. To tackle this challenge, we propose a novel learning-based framework for the reconstruction of high-quality LFs from acquisitions via learned coded apertures. The proposed method incorporates the measurement observation into the deep learning framework elegantly to avoid relying entirely on data-driven priors for LF reconstruction. Specifically, we first formulate the compressive LF reconstruction as an inverse problem with an implicit regularization term. Then, we construct the regularization term with a deep efficient spatial-angular separable convolutional sub-network in the form of local and global residual learning to comprehensively explore the signal distribution free from the limited representation ability and inefficiency of deterministic mathematical modeling. Furthermore, we extend this pipeline to LF denoising and spatial super-resolution, which could be considered as variants of coded aperture imaging equipped with different degradation matrices. Extensive experimental results demonstrate that the proposed methods outperform state-of-the-art approaches to a significant extent both quantitatively and qualitatively, i.e., the reconstructed LFs not only achieve much higher PSNR/SSIM but also preserve the LF parallax structure better on both real and synthetic LF benchmarks. The code will be publicly available at https://github.com/MantangGuo/DRLF.
Mantang Guo, Junhui Hou, Jing Jin 0006, Jie Chen 0026, Lap-Pui Chau
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Weakly-Supervised Part-Attention and Mentored Networks for Vehicle Re-Identification
abstract
Vehicle re-identification (Re-ID) aims to retrieve images with the same vehicle ID across different cameras. Current part-level feature learning methods typically detect vehicle parts via uniform division, outside tools, or attention modeling. However, such part features often require expensive additional annotations and cause sub-optimal performance in case of unreliable part mask predictions. In this paper, we propose a weakly-supervised Part-Attention Network (PANet) and Part-Mentored Network (PMNet) for Vehicle Re-ID. Firstly, PANet localizes vehicle parts via part-relevant channel recalibration and cluster-based mask generation without vehicle part supervisory information. Secondly, PMNet leverages teacher-student guided learning to distill vehicle part-specific features from PANet and performs multi-scale global-part feature extraction. During inference, PMNet can adaptively extract discriminative part features without part localization by PANet, preventing unstable part mask predictions. We address this Re-ID issue as a multi-task problem and adopt Homoscedastic Uncertainty to learn the optimal weighing of ID losses. Experiments are conducted on two public benchmarks, showing that our approach outperforms recent methods, which require no extra annotations by an average increase of 3.0% in CMC@5 on VehicleID and over 1.4% in mAP on VeRi776. Moreover, our method can extend to the occluded vehicle Re-ID task and exhibits good generalization ability.
Lisha Tang, Yi Wang 0068, Lap-Pui Chau
IEEE Trans. Circuits Syst. Video Technol.3
2022 Attention-Guided Progressive Neural Texture Fusion for High Dynamic Range Image Restoration
abstract
High Dynamic Range (HDR) imaging via multi-exposure fusion is an important task for most modern imaging platforms. In spite of recent developments in both hardware and algorithm innovations, challenges remain over content association ambiguities caused by saturation, motion, and various artifacts introduced during multi-exposure fusion such as ghosting, noise, and blur. In this work, we propose an Attention-guided Progressive Neural Texture Fusion (APNT-Fusion) HDR restoration model which aims to address these issues within one framework. An efficient two-stream structure is proposed which separately focuses on texture feature transfer over saturated regions and multi-exposure tonal and texture feature fusion. A neural feature transfer mechanism is proposed which establishes spatial correspondence between different exposures based on multi-scale VGG features in the masked saturated HDR domain for discriminative contextual clues over the ambiguous image areas. A progressive texture blending module is designed to blend the encoded two-stream features in a multi-scale and progressive manner. In addition, we introduce several novel attention mechanisms, i.e., the motion attention module detects and suppresses the content discrepancies among the reference images; the saturation attention module facilitates differentiating the misalignment caused by saturation from those caused by motion; and the scale attention module ensures texture blending consistency between different coder/decoder scales. We carry out comprehensive qualitative and quantitative evaluations and ablation studies, which validate that these novel modules work coherently under the same framework and outperform state-of-the-art methods.
Jie Chen 0026, Zaifeng Yang, Tsz Nam Chan, Hui Li 0029, Junhui Hou, Lap-Pui Chau
IEEE Trans. Image Process.6
2022 Rethinking and Designing a High-Performing Automatic License Plate Recognition Approach
abstract
In this paper, we propose a real-time and accurate automatic license plate recognition (ALPR) approach. Our study illustrates the outstanding design of ALPR with four insights: (1) the resampling-based cascaded framework is beneficial to both speed and accuracy; (2) the highly efficient license plate recognition should abundant additional character segmentation and recurrent neural network (RNN), but adopt a plain convolutional neural network (CNN); (3) in the case of CNN, taking advantage of vertex information on license plates improves the recognition performance; and (4) the weight-sharing character classifier addresses the lack of training images in small-scale datasets. Based on these insights, we propose a novel ALPR approach, termed VSNet. Specifically, VSNet includes two CNNs, i.e., VertexNet for license plate detection and SCR-Net for license plate recognition, integrated in a resampling-based cascaded manner. In VertexNet, we propose an efficient integration block to extract the spatial features of license plates. With vertex supervisory information, we propose a vertex-estimation branch in VertexNet such that license plates can be rectified as the input images of SCR-Net. In SCR-Net, we introduce a horizontal encoding technique for left-to-right feature extraction and propose a weight-sharing classifier for character recognition. Experimental results show that the proposed VSNet outperforms state-of-the-art methods by more than 50% relative improvement on error rate, achieving >99% recognition accuracy on CCPD and AOLP datasets with 149 FPS inference speed. Moreover, our method illustrates an outstanding generalization capability when evaluated on the unseen PKUData and CLPD datasets.
Yi Wang 0068, Zhen-Peng Bian, Yunhao Zhou, Lap-Pui Chau
IEEE Trans. Intell. Transp. Syst.4
2022 Soft Warping Based Unsupervised Domain Adaptation for Stereo Matching
abstract
Stereo matching is a practical method to estimate depth information and retrieve 3D world in robot perception and autonomous driving scenarios. With the development of convolution neural networks (CNNs), deep-learning based stereo matching algorithms have significantly improved the accuracy and dominated most of the online benchmarks. However, limited labels in real world, especially in challenging weather conditions, still hinder the technology from practical usage. In this paper, we propose a new unsupervised learning mechanism for stereo matching, utilizing adversarial iterative learning and novel soft warping loss to promote the effectiveness of the networks in unseen environments. The experiments transferring the stereo matching module from synthetic domain to real-world domain demonstrate the superiority of our proposed method. Extensive experiments in challenging weathers further prove that our method shows great practical potential in strait environments.
Lap-Pui Chau, Danwei Wang
IEEE Trans. Multim.2
2021 Embracing the Dark Knowledge: Domain Generalization Using Regularized Knowledge Distillation
abstract
Though convolutional neural networks are widely used in different tasks, lack of generalization capability in the absence of sufficient and representative data is one of the challenges that hinders their practical application. In this paper, we propose a simple, effective, and plug-and-play training strategy named Knowledge Distillation for Domain Generalization (KDDG) which is built upon a knowledge distillation framework with the gradient filter as a novel regularization term. We find that both the "richer dark knowledge" from the teacher network, as well as the gradient filter we proposed, can reduce the difficulty of learning the mapping which further improves the generalization ability of the model. We also conduct experiments extensively to show that our framework can significantly improve the generalization capability of deep neural networks in different tasks including image classification, segmentation, reinforcement learning by comparing our method with existing state-of-the-art domain generalization techniques. Last but not the least, we propose to adopt two metrics to analyze our proposed method in order to better understand how our proposed method benefits the generalization capability of deep neural networks.
Yufei Wang 0006, Haoliang Li, Lap-Pui Chau, Alex Chichung Kot
ACM Multimedia3
2021 A Self-Training Approach for Point-Supervised Object Detection and Counting in Crowds
abstract
In this article, we propose a novel self-training approach named Crowd-SDNet that enables a typical object detector trained only with point-level annotations (i.e., objects are labeled with points) to estimate both the center points and sizes of crowded objects. Specifically, during training, we utilize the available point annotations to supervise the estimation of the center points of objects directly. Based on a locally-uniform distribution assumption, we initialize pseudo object sizes from the point-level supervisory information, which are then leveraged to guide the regression of object sizes via a crowdedness-aware loss. Meanwhile, we propose a confidence and order-aware refinement scheme to continuously refine the initial pseudo object sizes such that the ability of the detector is increasingly boosted to detect and count objects in crowds simultaneously. Moreover, to address extremely crowded scenes, we propose an effective decoding method to improve the detector's representation ability. Experimental results on the WiderFace benchmark show that our approach significantly outperforms state-of-the-art point-supervised methods under both detection and counting tasks, i.e., our method improves the average precision by more than 10% and reduces the counting error by 31.2%. Besides, our method obtains the best results on the crowd counting and localization datasets (i.e., ShanghaiTech and NWPU-Crowd) and vehicle counting datasets (i.e., CARPK and PUCPR+) compared with state-of-the-art counting-by-detection methods. The code will be publicly available at https://github.com/WangyiNTU/Point-supervised-crowd-detection.
Yi Wang 0068, Junhui Hou, Xinyu Hou, Lap-Pui Chau
IEEE Trans. Image Process.4
2021 Convolutional Neural Networks With Dynamic Regularization
abstract
Regularization is commonly used for alleviating overfitting in machine learning. For convolutional neural networks (CNNs), regularization methods, such as DropBlock and Shake-Shake, have illustrated the improvement in the generalization performance. However, these methods lack a self-adaptive ability throughout training. That is, the regularization strength is fixed to a predefined schedule, and manual adjustments are required to adapt to various network architectures. In this article, we propose a dynamic regularization method for CNNs. Specifically, we model the regularization strength as a function of the training loss. According to the change of the training loss, our method can dynamically adjust the regularization strength in the training procedure, thereby balancing the underfitting and overfitting of CNNs. With dynamic regularization, a large-scale model is automatically regularized by the strong perturbation, and vice versa. Experimental results show that the proposed method can improve the generalization capability on off-the-shelf network architectures and outperform state-of-the-art regularization methods.
Yi Wang 0068, Zhen-Peng Bian, Junhui Hou, Lap-Pui Chau
IEEE Trans. Neural Networks Learn. Syst.4
2020 Deep Spatial-Angular Regularization for Compressive Light Field Reconstruction over Coded Apertures
Mantang Guo, Junhui Hou, Jing Jin 0006, Jie Chen 0026, Lap-Pui Chau
ECCV (2)5
2020 RSAN: A Retinex based Self Adaptive Stereo Matching Network for Day and Night Scenes
abstract
It is essential in many robot tasks to retrieve depth information, while it still remains a challenging problem to get robust depth in unfavorable conditions such as night or rainy environments. With the development of convolutional neural networks (CNNs), a large number of algorithms have emerged to tackle the problem of dark image enhancement and depth estimation, but there are few works focus on recovering depth map in dark environments and normal light condition. To meet this demand, we proposed a neural network which takes the paired stereo images in all light conditions as input and estimates the fully scaled depth map. The network contains a novel feature extractor and a stereo matching module which follows a light-weight manner to guarantee this work practical for real robotic applications. We introduced the Retinex Theory into depth estimation and trained the decomposition module with LOL dataset. Then it is adapted into depth estimation by fusing the decompose module into stereo matching algorithm. The whole network is then trained in an end-to-end manner. To demonstrate the robustness and effectiveness of our proposed method, we perform various studies and compare our results to the state-of-the-art algorithms in depth estimation as well as direct combination of image enhancement and stereo matching algorithm. We also collect stereo images in real night environments and present the improved performance of our network.
Lap-Pui Chau, Danwei Wang
ICARCV2
2020 Surface Consistent Light Field Extrapolation Over Stratified Disparity And Spatial Granularities
abstract
The light field captures both the spatial and angular configurations of the scene, which facilitates a wide range of imaging possibilities. In this work, we propose an LF view extrapolation algorithm which renders high quality novel LF views far outside the range of given angular baselines. A stratified synthesis strategy is adopted which projects the scene content based on stratified disparity layers and across varying scales of spatial granularities. Such a stratified methodology proves to help preserve scene structures over large angular shifts, and provide informative clues for inferring the contents of occluded regions. A generative-adversarial network model is further adopted for parallax correction and occlusion completion conditioned on surface consistent feature. Experiments show that our proposed model can provide more reliable novel view extrapolation quality at large baseline extension ratios compared with state-of-the-art LF synthesis algorithms.
Jie Chen 0026, Lap-Pui Chau, Junhui Hou
ICME2
2020 Haze Removal with Fusion of Local and Non-Local Statistics
abstract
Most of the outdoor images suffer from contrast degradation caused by fog and haze. Two statistical frameworks have been proposed in recent years that exploit local (dark channel prior) and non-local (haze-lines) characteristics of hazy images for the estimation of scene configurations and the restoration of scene albedo. Both frameworks show intrinsic limitations due to the basic assumptions they rely on. In this paper we propose a novel dehazing method that combines the advantages of local and non-local dehazing methods. Exploiting their complementary statistical properties, we use the local features to regulate the estimation of non-local haze-lines for a better final restoration at challenging regions. Both quantitative and qualitative results validate the effectiveness of our proposed method over state-of-the-art frameworks.
Jie Chen 0026, Cheen-Hau Tan, Lap-Pui Chau
ISCAS3
2020 Remote detection of idling cars using infrared imaging and deep networks
Muhammet Bastan, Kim-Hui Yap, Lap-Pui Chau
Neural Comput. Appl.3
2019 Vehicle Tracking Using Deep SORT with Low Confidence Track Filtering
abstract
Multi-object tracking (MOT) becomes an attractive topic due to its wide range of usability in video surveillance and traffic monitoring. Recent improvements on MOT has focused on tracking-by-detection manner. However, as a relatively complicated and integrated computer vision mission, state-of-the-art tracking-by-detection techniques are still suffering from issues such as a large number of false-positive tracks. To reduce the effect of unreliable detections on vehicle tracking, in this paper, we propose to incorporate a low confidence track filtering into the Simple Online and Realtime Tracking with a Deep association metric (Deep SORT) algorithm. We present a self-generated UA-DETRAC vehicle re-identification dataset which can be used to train the convolutional neural network of Deep SORT for data association. We evaluate our proposed tracker on UA-DETRAC test dataset. Experimental results show that the proposed method can improve the original Deep SORT algorithm with a significant margin. Our tracker outperforms the state-of-the-art online trackers and is comparable with batch-mode trackers.
Xinyu Hou, Yi Wang 0068, Lap-Pui Chau
AVSS3
2019 Object Counting in Video Surveillance Using Multi-scale Density Map Regression
abstract
In this paper, we present an effective convolutional neural network (CNN) for object counting in video surveillance, namely multi-scale density map regressor (MSDMR). In contrast to existing CNN-based methods that achieve high accuracy by means of empirically increasing the model capacity with more complex structures/layers, we focus on a compact CNN. Specifically, the MSDMR is mainly designed with the supervision of multi-scale outputs, in which two CNN stacks estimate coarse- and fine-scale density maps, respectively. The integral of the fine density map provides the count of objects. The two stacks are connected in a cascaded manner and jointly trained such that the overall model can learn discriminative and complementary features to produce expressive performance. Experimental results show that the proposed MSDMR can achieve higher accuracy compared with state-of-the-art methods on the surveillance datasets.
Yi Wang 0068, Junhui Hou, Lap-Pui Chau
ICASSP3
2019 Convolutional Three-Stream Network Fusion for Driver Fatigue Detection from Infrared Videos
abstract
We propose a convolutional three-stream network architecture for driver fatigue detection from infrared videos that are available both in the daytime and in the night time. Specifically, the convolutional three-stream network architecture incorporates current-infrared-frame-based spatial information, optical-flows-based short-term temporal information of two consecutive infrared frames and optical flow-motion history image-based (OF-MHI-based) temporal information within the infrared video sequence. And then these three networks are fused at the last convolutional layer by 3D CNN. Besides, an estimation method to evaluate the current driver fatigue level is proposed based on the fatigue detection results from previous frames, which helps to generate alerts properly in real-life driving applications. We show that the proposed method achieves state-of-the-art performance, 94.68% accuracy, in our driver behavior dataset using the infrared data.
Xiaoxi Ma, Lap-Pui Chau, Kim-Hui Yap, Guiju Ping
ISCAS2
2019 Airtight Estimation Based on Distant Region Segmentation
abstract
Natural images suffer from bad weather conditions, such as haze or fog, which decreases the contrast and degrades the color of observed images. Haze removal aims to recover haze-free images by the image degradation model. The global atmospheric light (airlight) estimation is an essential step for haze removal. With an assumption that the airlight exists in the infinite distance, we propose a novel learning-based framework for airlight estimation. Our framework is mainly composed of two steps: i) the airlight is initially determined by distant region segmentation based on U-Net; ii) the final airlight can be obtained by the weighted sum of the pixel values inside the distant region. Owing to lack of ground-truth airlight, we present a method to synthesize outdoor training examples. The proposed framework not only perform well on synthetic images but also has a good generalization ability for natural images. Experimental results demonstrate that our proposed approach can achieve more accurate estimate of airlight than state-of-the-art methods on both synthetic and natural images.
Yi Wang 0068, Lap-Pui Chau, Xiaoxi Ma
ISCAS2
2019 Deepsea video descattering
Hui Liu 0032, Lap-Pui Chau
Multim. Tools Appl.2
2019 Light Field Image Compression Based on Bi-Level View Compensation With Rate-Distortion Optimization
abstract
Compared with conventional color images, light field images (LFIs) contain richer scene information, which allows a wide range of interesting applications. However, such additional information is obtained at the cost of generating substantially more data, which poses challenges to both data storage and transmission. In this paper, we propose a new hybrid framework for effective compression of LFIs. The proposed framework takes the particular characteristics of LFIs into account so that the inter- and intra-view correlations of LFIs can be more efficiently exploited to produce better compression performance. Specifically, the proposed scheme partitions sub-aperture images (SAIs) of an LFI into two groups, namely, key SAIs and non-key SAIs. Bi-level view compensation is proposed to exploit the inter-view correlation: first, based on the group of selected key SAIs, learning-based angular super-resolution is performed to compensate non-key SAIs in pixel-wise, during which heterogeneous inter-view correlation between the non-key SAIs is efficiently removed; second, the two groups of SAIs are respectively reorganized as pseudo-sequences, and block-wise motion compensation is carried out with a standard video encoder, during which the homogeneous inter-view correlation is subsequently exploited. The video encoder also helps to remove the intra-view correlation of the SAIs and finally generates the encoded bitstream. Moreover, the bits allocated to each group are optimally determined via model-based rate distortion optimization. Extensive experimental evaluations and comparisons demonstrate the advantage of the proposed framework over existing methods in terms of rate-distortion performance.
Junhui Hou, Jie Chen 0026, Lap-Pui Chau
IEEE Trans. Circuits Syst. Video Technol.3
2018 Robust Video Content Alignment and Compensation for Rain Removal in a CNN Framework
abstract
Rain removal is important for improving the robustness of outdoor vision based systems. Current rain removal methods show limitations either for complex dynamic scenes shot from fast moving cameras, or under torrential rain fall with opaque occlusions. We propose a novel derain algorithm, which applies superpixel (SP) segmentation to decompose the scene into depth consistent units. Alignment of scene contents are done at the SP level, which proves to be robust towards rain occlusion and fast camera motion. Two alignment output tensors, i.e., optimal temporal match tensor and sorted spatial-temporal match tensor, provide informative clues for rain streak location and occluded background contents to generate an intermediate derain output. These tensors will be subsequently prepared as input features for a convolutional neural network to restore high frequency details to the intermediate output for compensation of mis-alignment blur. Extensive evaluations show that up to 5dB reconstruction PSNR advantage is achieved over state-of-the-art methods. Visual inspection shows that much cleaner rain removal is achieved especially for highly dynamic scenes with heavy and opaque rainfall from a fast moving camera.
Jie Chen 0026, Cheen-Hau Tan, Junhui Hou, Lap-Pui Chau
CVPR4
2018 Idling Car Detection with ConvNets in Infrared Image Sequences
abstract
We propose a system to detect and localize idling cars in infrared (IR) image sequences for law enforcement to reduce vehicular emission. To this end, we leverage the differences in spatio-temporal heat signatures of idling and stopped cars and monitor car temperatures with a long-wavelength IR camera. We collected a dataset by recording IR image sequences of cars in car parks and trained a ConvNet-based car detector to localize stationary cars in the IR sequences, by utilizing transfer learning and models pre-trained on regular RGB/grayscale images. Then, we used ConvNets with a 3D stack of cropped frames as input to model the spatio-temporal evolution of car temperature over time and detect idling cars. We present promising experimental results on our IR image dataset.
Muhammet Bastan, Kim-Hui Yap, Lap-Pui Chau
ISCAS3
2018 Improved Network Robustness with Adversary Critic
abstract
Ideally, what confuses neural network should be confusing to humans. However, recent experiments have shown that small, imperceptible perturbations can change the network prediction. To address this gap in perception, we propose a novel approach for learning robust classifier. Our main idea is: adversarial examples for the robust classifier should be indistinguishable from the regular data of the adversarial target. We formulate a problem of learning robust classifier in the framework of Generative Adversarial Networks (GAN), where the adversarial attack on classifier acts as a generator, and the critic network learns to distinguish between regular and adversarial images. The classifier cost is augmented with the objective that its adversarial examples should confuse the adversary critic. To improve the stability of the adversarial mapping, we introduce adversarial cycle-consistency constraint which ensures that the adversarial mapping of the adversarial examples is close to the original. In the experiments, we show the effectiveness of our defense. Our method surpasses in terms of robustness networks trained with adversarial training. Additionally, we verify in the experiments with human annotators on MTurk that adversarial examples are indeed visually confusing.
Alexander Matyasko, Lap-Pui Chau
NeurIPS2
2018 Light Field Denoising via Anisotropic Parallax Analysis in a CNN Framework
abstract
Light field (LF) cameras provide perspective information of scenes by taking directional measurements of the focusing light rays. The raw outputs are usually dark with additive camera noise, which impedes subsequent processing and applications. We propose a novel LF denoising framework based on anisotropic parallax analysis (APA). Two convolutional neural networks are jointly designed for the task: first, the structural parallax synthesis network predicts the parallax details for the entire LF based on a set of anisotropic parallax features. These novel features can efficiently capture the high-frequency perspective components of a LF from noisy observations. Second, the view-dependent detail compensation network restores non-Lambertian variation to each LF view by involving view-specific spatial energies. Extensive experiments show that the proposed APA LF denoiser provides a much better denoising performance than state-of-the-art methods in terms of visual quality and in preservation of parallax details.
Jie Chen 0026, Junhui Hou, Lap-Pui Chau
IEEE Signal Process. Lett.3
2018 Simultaneous Spatial and Spectral Low-Rank Representation of Hyperspectral Images for Classification
abstract
Arising from various environmental and atmos- pheric conditions and sensor interference, spectral variations are inevitable during hyperspectral remote sensing, which degrade the subsequent hyperspectral image analysis significantly. In this paper, we propose simultaneous spatial and spectral low-rank representation (S3LRR) that can effectively suppress the within-class spectral variations for classification purposes. The S3LRR recovers an intrinsic component with the same dimension as the original image, in which both spatial and spectral low-rank priors are adopted to regularize the intrinsic component simultaneously and compensate to each other, together with robust modeling of spectral variations. Compared with existing methods that explore only the spectral low-rank prior, the novel spatial low-rank prior (i.e., low-rank prior in band-wise) can take the spatial structure information of hyperspectral images into account, which has demonstrated to be very useful. Technically, we formulate S3LRR as a constrained convex optimization problem, and solve it using the efficient inexact augmented Lagrangian multiplier method. The resulting intrinsic component is less interfered by within-class spectral variations, and more discriminatory to offer higher classification accuracy. Comprehensive experiments on benchmark data sets demonstrate that the proposed S3LRR improves classification accuracy significantly, which outperforms state-of-the-art methods.
Shaohui Mei, Junhui Hou, Jie Chen 0026, Lap-Pui Chau, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2018 Light Field Compression With Disparity-Guided Sparse Coding Based on Structural Key Views
abstract
Recent imaging technologies are rapidly evolving for sampling richer and more immersive representations of the 3D world. One of the emerging technologies is light field (LF) cameras based on micro-lens arrays. To record the directional information of the light rays, a much larger storage space and transmission bandwidth are required by an LF image as compared with a conventional 2D image of similar spatial dimension. Hence, the compression of LF data becomes a vital part of its application. In this paper, we propose an LF codec with disparity guided Sparse Coding over a learned perspective-shifted LF dictionary based on selected Structural Key Views (SC-SKV). The sparse coding is based on a limited number of optimally selected SKVs; yet the entire LF can be recovered from the coding coefficients. By keeping the approximation identical between encoder and decoder, only the residuals of the non-key views, disparity map, and the SKVs need to be compressed into the bit stream. An optimized SKV selection method is proposed such that most LF spatial information can be preserved. To achieve optimum dictionary efficiency, the LF is divided into several coding regions, over which the reconstruction works individually. Experiments and comparisons have been carried out over benchmark LF data set, which show that the proposed SC-SKV codec produces convincing compression results in terms of both rate-distortion performance and visual quality compared with Joint Exploration Model: with 37.9% BD-rate reduction and 1.17-dB BD-PSNR improvement achieved on average, especially with up to 6-dB improvement for low bit rate scenarios.
Jie Chen 0026, Junhui Hou, Lap-Pui Chau
IEEE Trans. Image Process.3
2018 Accurate Light Field Depth Estimation With Superpixel Regularization Over Partially Occluded Regions
abstract
Depth estimation is a fundamental problem for light field photography applications. Numerous methods have been proposed in recent years, which either focus on crafting cost terms for more robust matching, or on analyzing the geometry of scene structures embedded in the epipolar-plane images. Significant improvements have been made in terms of overall depth estimation error; however, current state-of-the-art methods still show limitations in handling intricate occluding structures and complex scenes with multiple occlusions. To address these challenging issues, we propose a very effective depth estimation framework which focuses on regularizing the initial label confidence map and edge strength weights. Specifically, we first detect partially occluded boundary regions (POBR) via superpixel-based regularization. Series of shrinkage/reinforcement operations are then applied on the label confidence map and edge strength weights over the POBR. We show that after weight manipulations, even a low-complexity weighted least squares model can produce much better depth estimation than the state-of-the-art methods in terms of average disparity error rate, occlusion boundary precision-recall rate, and the preservation of intricate visual features.
Jie Chen 0026, Junhui Hou, Yun Ni, Lap-Pui Chau
IEEE Trans. Image Process.4
2018 Multimodal Recurrent Neural Networks With Information Transfer Layers for Indoor Scene Labeling
abstract
This paper proposes a new method called multimodal recurrent neural networks (RNNs) for RGB-D scene semantic segmentation. It is optimized to classify image pixels given two input sources: RGB color channels and depth maps. It simultaneously performs training of two RNNs that are crossly connected through information transfer layers, which are learnt to adaptively extract relevant cross-modality features. Each RNN model learns its representations from its own previous hidden states and transferred patterns from the other RNNs previous hidden states; thus, both model-specific and cross-modality features are retained. We exploit the structure of quad-directional 2D-RNNs to model the short- and long-range contextual information in the 2D input image. We carefully designed various baselines to efficiently examine our proposed model structure. We test our multimodal RNNs method on popular RGB-D benchmarks and show how it outperforms previous methods significantly and achieves competitive results with other state-of-the-art works.
Abrar H. Abdulnabi, Bing Shuai, Zhen Zuo, Lap-Pui Chau, Gang Wang 0012
IEEE Trans. Multim.4
2017 Sparse representation for colors of 3D point cloud via virtual adaptive sampling
abstract
Sparse signal representation has proven to be an extremely powerful tool in a wide range of engineering applications. However, most of the existing techniques are designed for regular data (such as audio signals and images/videos) that uniformly lies in regular Euclidian spaces. This paper aims at extending sparse representation for irregular data (such as colors of 3D point clouds) that is defined on irregular domains embedded in Euclidean spaces. Dealing with the irregular structure of such data via a virtual adaptive sampling process, we formulate sparse representation as an ℓ0-norm regularized optimization problem. Experimental results show that the proposed algorithm outperforms the state-of-the-art algorithm to a large extent: with the same number of nonzero coefficients, we improve the reconstruction quality up to 5 dB; conversely, fixing the reconstruction quality, our method uses only 55% coefficients. Using compressive sensing theory, we provide an intuitive explanation on how and why our algorithm works well in practice.
Junhui Hou, Lap-Pui Chau, Ying He 0001, Philip A. Chou
ICASSP2
2017 Margin maximization for robust classification using deep learning
abstract
Deep neural networks have achieved significant success for image recognition problems. Despite the wide success, recent experiments demonstrated that neural networks are sensitive to small input perturbations, or adversarial noise. The lack of robustness is intuitively undesirable and limits neural networks applications in adversarial settings, and for image search and retrieval problems. Current approaches consider augmenting training dataset using adversarial examples to improve robustness. However, when using data augmentation, the model fails to anticipate changes in an adversary. In this paper, we consider maximizing the geometric margin of the classifier. Intuitively, a large margin relates to classifier robustness. We introduce novel margin maximization objective for deep neural networks. We theoretically show that the proposed objective is equivalent to the robust optimization problem for a neural network. Our work seamlessly generalizes SVM margin objective to deep neural networks. In the experiments, we extensively verify the effectiveness of the proposed margin maximization objective to improve neural network robustness and to reduce overfitting on MNIST and CIFAR-10 dataset.
Alexander Matyasko, Lap-Pui Chau
IJCNN2
2017 Reflection removal based on single light field capture
abstract
Photography through reflective surfaces suffers from the obstruction of reflections, which deteriorates the visibility of background targets and causes challenges for subsequent computer vision applications. In this paper, we propose a novel reflection removal algorithm using light field (LF) cameras. Unlike conventional cameras, LF cameras capture extra directional information of incoming rays which enable our algorithm to remove reflections with only a single shot. We analyze the optical geometry of the background and reflection imagery in a LF camera, and generalize a set of rules that could facilitate our algorithm to differentiate the edges from different optical sources. We show that the proposed method produces significantly better reflection removal results based on the LF data as compared to traditional methods based on multiple-shot image sequences as input.
Yun Ni, Jie Chen 0026, Lap-Pui Chau
ISCAS3
2017 Single underwater image restoration using attenuation-curve prior
abstract
Underwater images suffer from low contrast and color distortion due to the existence of dust-like particles and light attenuation. Some previous works using the patch-based priors, e.g. adaptations of the dark channel prior, cannot achieve satisfactory results in both contrast enhancement and color restoration in the underwater environment. In this paper, we propose a novel underwater image restoration method based on a non-local prior, termed an attenuation-curve prior. This prior relies on the observation that colors of a clear image can be well approximated by several hundred distinct color clusters and the pixels in the same color cluster will form a power function curved line in RGB space after their colors are attenuated by water. Our work mainly contains two steps. Firstly, we estimate the waterlight based on its smoothness properties and the different attenuation coefficient of light. Secondly, we estimate the transmission map using the attenuation-curve prior. Once the waterlight and transmission are obtained, the clear underwater image can be restored. Experimental results demonstrate that our proposed method can achieve better results when comparing with state-of-the-art approaches.
Yi Wang 0068, Hui Liu 0032, Lap-Pui Chau
ISCAS3
2017 Lattice-Support repetitive local feature detection for visual search
Dipu Manandhar, Kim-Hui Yap, Zhenwei Miao, Lap-Pui Chau
Pattern Recognit. Lett.4
2017 Light Field Compressed Sensing Over a Disparity-Aware Dictionary
abstract
Light field (LF) acquisition faces the challenge of extremely bulky data. Available hardware solutions usually compromise the sensor resource between spatial and angular resolutions. In this paper, a compressed sensing framework is proposed for the sampling and reconstruction of a high-resolution LF based on a coded aperture camera. First, an LF dictionary based on perspective shifting is proposed for the sparse representation of the highly correlated LF. Then, two separate methods, i.e., subaperture scan and normalized fluctuation, are proposed to acquire/calculate the scene disparity, which will be used during the LF reconstruction with the proposed disparity-aware dictionary. At last, a hardware implementation of the proposed LF acquisition/reconstruction scheme is carried out. Both quantitative and qualitative evaluation show that the proposed methods produce the state-of-the-art performance in both reconstruction quality and computation efficiency.
Jie Chen 0026, Lap-Pui Chau
IEEE Trans. Circuits Syst. Video Technol.2
2017 Sparse Low-Rank Matrix Approximation for Data Compression
abstract
Low-rank matrix approximation (LRMA) is a powerful technique for signal processing and pattern analysis. However, its potential for data compression has not yet been fully investigated. In this paper, we propose sparse LRMA (SLRMA), an effective computational tool for data compression. SLRMA extends conventional LRMA by exploring both the intra and inter coherence of data samples simultaneously. With the aid of prescribed orthogonal transforms (e.g., discrete cosine/wavelet transform and graph transform), SLRMA decomposes a matrix into a product of two smaller matrices, where one matrix is made up of extremely sparse and orthogonal column vectors and the other consists of the transform coefficients. Technically, we formulate SLRMA as a constrained optimization problem, i.e., minimizing the approximation error in the least-squares sense regularized by the $\ell _{0}$ -norm and orthogonality, and solve it using the inexact augmented Lagrangian multiplier method. Through extensive tests on real-world data, such as 2D image sets and 3D dynamic meshes, we observe that: 1) SLRMA empirically converges well; 2) SLRMA can produce approximation error comparable to LRMA but in a much sparse form; and 3) SLRMA-based compression schemes significantly outperform the state of the art in terms of rate-distortion performance.
Junhui Hou, Lap-Pui Chau, Nadia Magnenat-Thalmann, Ying He 0001
IEEE Trans. Circuits Syst. Video Technol.2
2016 Robust laplacian matrix learning for smooth graph signals
abstract
We propose a new method for robust learning Laplacian matrices from observed smooth graph signals in the presence of both Gaussian noise and random-valued impulse noise (i.e., outliers). Using the recently developed factor analysis model for representing smooth graph signals in [1], we formulate our learning process as a constrained optimization problem, and adopt the £i-norm for measuring the data fidelity in order to improve robustness. Computational results on three types of synthetic graphs demonstrate that the proposed method outperforms the state-of-the-art methods in terms of commonly used information retrieval metrics, such as F-measure, precision, recall and normalized mutual information. In particular, we observed that F-measure is improved by up to 16%.
Junhui Hou, Lap-Pui Chau, Ying He 0001, Huanqiang Zeng
ICIP2
2016 Sparse two-dimensional singular value decomposition
abstract
In this paper, we propose a new data-driven transform, called sparse two-dimensional singular value decomposition (S2DSVD). By leveraging the advantages of discrete cosine transform and the conventional 2D SVD, we decompose a set of matrices into transform coefficient matrices with sparse and orthogonal basis functions. Such sparsity characteristic can significantly reduce their overhead, hence being beneficial to data compression. We formulate S2DSVD as a constrained optimization problem and solve it via alternative iteration. We demonstrate the efficacy of S2DSVD on image and video datasets, and observe that it can produce results with error comparable to 2D SVD whereas its space complexity is significantly smaller than 2D SVD.
Junhui Hou, Jie Chen 0026, Lap-Pui Chau, Ying He 0001
ICME3
2016 Low-latency compression of mocap data using learned spatial decorrelation transform
abstract
Due to the growing needs of motion capture (mocap) in movie, video games , sports, etc., it is highly desired to compress mocap data for efficient storage and transmission. Unfortunately, the existing compression methods have either high latency or poor compression performance , making them less appealing for time-critical applications and/or network with limited bandwidth . This paper presents two efficient methods to compress mocap data with low latency. The first method processes the data in a frame-by-frame manner so that it is ideal for mocap data streaming. The second one is clip-oriented and provides a flexible trade-off between latency and compression performance . It can achieve higher compression performance while keeping the latency fairly low and controllable. Observing that mocap data exhibits some unique spatial characteristics , we learn an orthogonal transform to reduce the spatial redundancy . We formulate the learning problem as the least square of reconstruction error regularized by orthogonality and sparsity , and solve it via alternating iteration. We also adopt a predictive coding and temporal DCT for temporal decorrelation in the frame- and clip-oriented methods, respectively. Experimental results show that the proposed methods can produce higher compression performance at lower computational cost and latency than the state-of-the-art methods. Moreover, our methods are general and applicable to various types of mocap data.
Junhui Hou, Lap-Pui Chau, Nadia Magnenat-Thalmann, Ying He 0001
Comput. Aided Geom. Des.2
2016 Facial Position and Expression-Based Human-Computer Interface for Persons With Tetraplegia
abstract
A human-computer interface (namely Facial position and expression Mouse system, FM) for the persons with tetraplegia based on a monocular infrared depth camera is presented in this paper. The nose position along with the mouth status (close/open) is detected by the proposed algorithm to control and navigate the cursor as computer user input. The algorithm is based on an improved Randomized Decision Tree, which is capable of detecting the facial information efficiently and accurately. A more comfortable user experience is achieved by mapping the nose motion to the cursor motion via a nonlinear function. The infrared depth camera enables the system to be independent of illumination and color changes both from the background and on human face, which is a critical advantage over RGB camera-based options. Extensive experimental results show that the proposed system outperforms existing assistive technologies in terms of quantitative and qualitative assessments.
Zhen-Peng Bian, Junhui Hou, Lap-Pui Chau, Nadia Magnenat-Thalmann
IEEE J. Biomed. Health Informatics3
2015 Heavy haze removal in a learning framework
abstract
Extreme weather hazards happens more often these days due to climate changes and increased human industrial activities, and one of most notorious of them is haze. State-of-the-art haze removal methods generally work well with light haze conditions, however when haze gets heavier, the physical model tend to produce over-shadowed, noisy, and color distorted restorations. A new physical model has been proposed in this paper for heavy haze weathers. An airlight vector map has been proposed to address the problem caused by uneven aerosol distribution w.r.t. altitude variation. A Random Decision Forest model has been adopted to deal with the additional light attenuation and transmission map underestimation problem caused by heavy haze. Experiment shows the proposed model produces much better visual restoration for heavy haze weathers compared to state-of-the-art methods in terms of colour fidelity, noise reduction, and overall contrast.
Jie Chen 0026, Lap-Pui Chau
ISCAS2
2015 Reordering-based transform for compressing human motion capture data
abstract
This paper presents a simple yet effective algorithm for compressing human motion capture (mocap) data. With a reordering-based discrete wavelet transform and the standard discrete cosine transform, our method can effectively reduce the spatial and temporal correlation in mocap data. Our method is conceptually simple and easy to implement. Experimental results show that our method can achieve better compression performance with lower latency, compared to the state-of-the-art methods.
Junhui Hou, Lap-Pui Chau, Ying He 0001, Nadia Magnenat-Thalmann
ISCAS2
2015 A linear dependent rate-quantization model for scalable video enhancement layer encoding
abstract
In this paper, we propose a linear dependent rate-quantization model for video enhancement layers encoding in H.264/AVC based scalable video coding (SVC). It is noted that the proposed model is applicable for different scalable structures, such as temporal, quality, spatial and combined scalability. Leveraging the base layer information (such as bitrate and quantization parameter), proposed model can accurately predict the number of bit required for the enhancement layer encoding. Such linear model demonstrates the high accuracy for bitrate estimation at enhancement layers, with the average prediction accuracy over 94%. It has the noticeable improvement from the existing works, without requiring additional complexity increase. Meanwhile, proposed model is applied to do the rate control for enhancement layers encoding. Experimental results show that the average bitrate mismatch error can be significantly reduced compared with the existing algorithms.
Junhui Hou, Shuai Wan, Lap-Pui Chau
ISCAS4
2015 Multiscale Dictionary Learning via Cross-Scale Cooperative Learning and Atom Clustering for Visual Signal Processing
abstract
For sparse signal representation, the sparsity across the scales is a promising yet underinvestigated direction. In this paper, we aim to design a multiscale sparse representation scheme to explore such potential. A multiscale dictionary (MD) structure is designed. A cross-scale matching pursuit algorithm is proposed for multiscale sparse coding. Two dictionary learning methods, cross-scale cooperative learning and cross-scale atom clustering, are proposed each focusing on one of the two important attributes of an efficient MD: the similarity and uniqueness of corresponding atoms in different scales. We analyze and compare their different advantages in the application of image denoising under different noise levels, where both methods produce state-of-the-art denoising results.
Jie Chen 0026, Lap-Pui Chau
IEEE Trans. Circuits Syst. Video Technol.2
2015 Compressing 3-D Human Motions via Keyframe-Based Geometry Videos
abstract
This paper presents keyframe-based geometry video (KGV), a novel framework for compressing 3-D human motion data by using geometry videos. Given a motion data encoded in a geometry video (GV) format, our method extracts the keyframes and produces a reconstruction matrix. Then it applies the video compression technique (e.g., H.264/Advanced Video Coding) to the reordered keyframes, which can significantly reduce the spatial and temporal redundancy in the KGV. We develop a rate distortion-based optimization algorithm to determine the parameters (i.e., the number of keyframes and quantization parameter) leading to optimal performance. Experimental results show that the proposed KGV framework significantly outperforms the existing GV techniques in terms of both the rate distortion performance and visual quality. Besides, the computational cost of the KGV is rather low at the decoder, making it highly desirable for power-constrained devices. Last but not least, our method can be easily extended to progressive compression with heterogeneous communication network.
Junhui Hou, Lap-Pui Chau, Nadia Magnenat-Thalmann, Ying He 0001
IEEE Trans. Circuits Syst. Video Technol.2
2015 Fall Detection Based on Body Part Tracking Using a Depth Camera
abstract
The elderly population is increasing rapidly all over the world. One major risk for elderly people is fall accidents, especially for those living alone. In this paper, we propose a robust fall detection approach by analyzing the tracked key joints of the human body using a single depth camera. Compared to the rivals that rely on the RGB inputs, the proposed scheme is independent of illumination of the lights and can work even in a dark room. In our scheme, a pose-invariant randomized decision tree algorithm is proposed for the key joint extraction, which requires low computational cost during the training and test. Then, the support vector machine classifier is employed to determine whether a fall motion occurs, whose input is the 3-D trajectory of the head joint. The experimental results demonstrate that the proposed fall detection method is more accurate and robust compared with the state-of-the-art methods.
Zhen-Peng Bian, Junhui Hou, Lap-Pui Chau, Nadia Magnenat-Thalmann
IEEE J. Biomed. Health Informatics3
2015 Human Motion Capture Data Tailored Transform Coding
abstract
Human motion capture (mocap) is a widely used technique for digitalizing human movements. With growing usage, compressing mocap data has received increasing attention, since compact data size enables efficient storage and transmission. Our analysis shows that mocap data have some unique characteristics that distinguish themselves from images and videos. Therefore, directly borrowing image or video compression techniques, such as discrete cosine transform, does not work well. In this paper, we propose a novel mocap-tailored transform coding algorithm that takes advantage of these features. Our algorithm segments the input mocap sequences into clips, which are represented in 2D matrices. Then it computes a set of data-dependent orthogonal bases to transform the matrices to frequency domain, in which the transform coefficients have significantly less dependency. Finally, the compression is obtained by entropy coding of the quantized coefficients and the bases. Our method has low computational cost and can be easily extended to compress mocap databases. It also requires neither training nor complicated parameter setting. Experimental results demonstrate that the proposed scheme significantly outperforms state-of-the-art algorithms in terms of compression performance and speed.
Junhui Hou, Lap-Pui Chau, Nadia Magnenat-Thalmann, Ying He 0001
IEEE Trans. Vis. Comput. Graph.2
2015 Motion capture data recovery using skeleton constrained singular value thresholding
Cheen-Hau Tan, Junhui Hou, Lap-Pui Chau
Vis. Comput.3
2014 Low-rank based compact representation of motion capture data
abstract
In this paper, we propose a practical, elegant and effective scheme for compact mocap data representation. Guided by our analysis of the unique properties of mocap data, the input mocap sequence is optimally segmented into a set of subsequences. Then, we project the subsequences onto a pair of computational orthogonal matrices to explore strong low-rank characteristic within and among the subsequences. The experimental results show that the proposed scheme is much more effective for reducing the data size, compared with the existing techniques.
Junhui Hou, Lap-Pui Chau, Ying He 0001, Nadia Magnenat-Thalmann
ICIP2
2014 Restoring corrupted motion capture data via jointly low-rank matrix completion
abstract
Motion capture (mocap) technology is widely used in various applications. The acquired mocap data usually has missing data due to occlusions or ambiguities. Therefore, restoring the missing entries of the mocap data is a fundamental issue in mocap data analysis. Based on jointly low-rank matrix completion, this paper presents a practical and highly efficient algorithm for restoring the missing mocap data. Taking advantage of the unique properties of mocap data (i.e, strong correlation among the data), we represent the corrupted data as two types of matrices, where both the local and global characteristics are taken into consideration. Then we formulate the problem as a convex optimization problem, where the missing data is recovered by solving the two matrices using the alternating direction method of multipliers algorithm. Experimental results demonstrate that the proposed scheme significantly outperforms the state-of-the-art algorithms in terms of both the quality and computational cost.
Junhui Hou, Zhen-Peng Bian, Lap-Pui Chau, Nadia Magnenat-Thalmann, Ying He 0001
ICME3
2014 A fast adaptive guided filtering algorithm for light field depth interpolation
abstract
Light field camera provides 4D information of the light rays, from which the scene depth information can be inferred. The disparity/depth maps calculated from light field data are always noisy with missing and false entries in homogeneous regions or areas where view-dependant effects are present. In this paper we proposed an adaptive guided filtering (AGF) algorithm to get an optimized output disparity/depth map. A guidance image is used to provide the image contour and texture information, the filter is able to preserve the disparity edges, smooth the regions without influence of the image texture, and reject the data entries with low confidence during coefficients regression. Experiment shows AGF is much faster in implementation as compared to other variational or hierarchical based optimization algorithms, and produces competitive visual results.
Jie Chen 0026, Lap-Pui Chau
ISCAS2
2014 A novel compression framework for 3D time-varying meshes
abstract
Compression of 3D time-varying meshes (TVMs) plays a critical role in the storage and transmission of 3D contents. In this paper, we propose a novel framework for compressing 3D TVMs. In our framework, 3D TVMs are parameterized and represented by the geometry videos (GVs) through polycube parameterization. By considering the low-rank characteristic of dynamic meshes, we decompose GVs into a sequence with small frames namely EigenGV and the computed reconstruction matrix. We further apply 2D video encoder to eliminate spatial and temporal redundancy among the EigenGV. Experimental results demonstrate that the proposed method significantly outperforms the existing compression schemes in terms of both the rate distortion performance and visual quality. Besides, the proposed method naturally achieves progressive form, which is very suitable for error prone channel transmission.
Junhui Hou, Lap-Pui Chau, Ying He 0001, Nadia Magnenat-Thalmann
ISCAS2
2014 Human Computer Interface for Quadriplegic People Based on Face Position/gesture Detection
abstract
This paper proposes a human computer interface using a single depth camera for quadriplegic people. The nose position is employed to control the cursor along with the commands provided by mouth's status. The detection of nose position and mouth's status is based on randomized decision tree algorithm.The experimental results show that the proposed interface is comfortable, easy to use, robust, and outperforms the existing assistive technology.
Zhen-Peng Bian, Junhui Hou, Lap-Pui Chau, Nadia Magnenat-Thalmann
ACM Multimedia3
2014 Scalable and Compact Representation for Motion Capture Data Using Tensor Decomposition
abstract
Motion capture (mocap) technology is widely used in movie and game industries. Compact representation of the mocap data is critical to efficient storage and transmission. In this letter, we propose a novel tensor decomposition based scheme for compact and progressive representation of the mocap data. Our method segments and stacks the mocap sequence locally, and generates a 3rd-order tensor, which has strong correlation within and across slices of the tensor. Then, our method iteratively applies tensor decomposition in a multi-layer structure to explore the correlation characteristic. Experimental results demonstrate that the proposed scheme significantly outperforms existing algorithms in terms of scalability and storage requirement.
Junhui Hou, Lap-Pui Chau, Nadia Magnenat-Thalmann, Ying He 0001
IEEE Signal Process. Lett.2
2014 Low Power Motion Estimation Based on Probabilistic Computing
abstract
As CMOS technology driven by Moore's law has approached device sizes in the range of 5-20 nm, noise immunity of such future technology nodes is predicted to decrease considerably, eventually affecting the reliability of computations through them. A shift in the design paradigm is expected from 100% accurate computations to probabilistic computing with accuracy dependent on the target application or circuit specifications. One model developed for CMOS technology that emulates the erroneous behavior predicted is termed probabilistic CMOS (PCMOS). In this paper, we propose a PCMOS-based architecture implementation for traditional motion estimation algorithms and show that up to 57% energy savings are possible for different existing motion estimation algorithms. Furthermore, algorithmic modifications are proposed that can enhance the energy savings to 70% with a PCMOS architectural implementation. About 1.8-5 dB improvement in peak signal-to-noise ratio under energy savings of 57% to 70% for two different motion estimation algorithms is shown, establishing the resilience of the proposed algorithm to probabilistic computing over the comparable conventional algorithm.
Charvi Dhoot, Lap-Pui Chau, Shubhajit Roy Chowdhury, Vincent John Mooney III
IEEE Trans. Circuits Syst. Video Technol.2
2014 A Highly Efficient Compression Framework for Time-Varying 3-D Facial Expressions
abstract
The rapid recent development of 3-DTV technology has led to an increase in studies on mesh-based 3-D scene representation. Compressing 3-D time-varying meshes is critical for the storage and transmission of 3-D contents. This paper proposes a highly efficient framework for compressing time-varying 3-D facial expressions. We use the near-isometric property of human facial expressions to parameterize the 3-D dynamic faces into an expression-invariant 2-D canonical domain that will naturally generate 2-D geometry videos (GVs). Considering the intrinsic properties of GVs, we apply low-rank and sparse matrix decomposition (LRSMD) separately to three dimensions of GVs (namely, \(X, Y,\) and \(Z\) ). Based on our high precision rate and distortion models for GVs, we further compress the components from LRSMD using a video encoder in which bitrates of all components are assigned optimally according to the target bitrate. Experimental results show that the proposed scheme can significantly improve compression performance in terms of rate-distortion performance and visual quality compared with the state-of-the-art algorithms.
Junhui Hou, Lap-Pui Chau, Minqi Zhang, Nadia Magnenat-Thalmann, Ying He 0001
IEEE Trans. Circuits Syst. Video Technol.2
2014 A Rain Pixel Recovery Algorithm for Videos With Highly Dynamic Scenes
abstract
Rain removal is a very useful and important technique in applications such as security surveillance and movie editing. Several rain removal algorithms have been proposed these years, where photometric, chromatic, and probabilistic properties of the rain have been exploited to detect and remove the rainy effect. Current methods generally work well with light rain and relatively static scenes, when dealing with heavier rainfall in dynamic scenes, these methods give very poor visual results. The proposed algorithm is based on motion segmentation of dynamic scene. After applying photometric and chromatic constraints for rain detection, rain removal filters are applied on pixels such that their dynamic property as well as motion occlusion clue are considered; both spatial and temporal informations are then adaptively exploited during rain pixel recovery. Results show that the proposed algorithm has a much better performance for rainy scenes with large motion than existing algorithms.
Jie Chen 0026, Lap-Pui Chau
IEEE Trans. Image Process.2
2013 An enhanced window-variant dark channel prior for depth estimation using single foggy image
abstract
The dark channel prior is a simple yet efficient way to estimate the scene depth information using one single foggy image. However the prior fails for pixels with low colour saturation. Based on the observation that areas with dramatic colour changes tend to belong to similar depth, a window variation mechanism is proposed in this paper based on the neighbourhood scene complexity and colour saturation rate to achieve an ideal compromise between depth resolution and precision. The proposed method greatly alleviates the intrinsic drawbacks of the original dark channel prior. Experiments show the proposed method produces more accurate depth estimation in most of the scenes than the original prior.
Jie Chen 0026, Lap-Pui Chau
ICIP2
2013 Human motion capture data recovery via trajectory-based sparse representation
abstract
Motion capture is widely used in sports, entertainment and medical applications. An important issue is to recover motion capture data that has been corrupted by noise and missing data entries during acquisition. In this paper, we propose a new method to recover corrupted motion capture data through trajectory-based sparse representation. The data is firstly represented as trajectories with fixed length and high correlation. Then, based on the sparse representation theory, the original trajectories can be recovered by solving the sparse representation of the incomplete trajectories through the OMP algorithm using a dictionary learned by K-SVD. Experimental results show that the proposed algorithm achieves much better performance, especially when significant portions of data is missing, than the existing algorithms.
Junhui Hou, Lap-Pui Chau, Ying He 0001, Jie Chen 0026, Nadia Magnenat-Thalmann
ICIP2
2013 Rain removal from dynamic scene based on motion segmentation
abstract
Rain removal technique has been intensively studied over these years, the photometric, chromatic, and probabilistic properties of the rain have been exploited to remove the rainy effect. However, current available algorithms only work well with light rain and static scenes, when dealing with heavier rain fall in dynamic scenes, obvious visual degradation will occur especially in motion intensive areas. The proposed algorithm is based on motion segmentation of dynamic scenes. Photometric and chromatic constraints are used for rain detection, motion occlusion information are involved in the adaptive prediction of the rain pixels' original value, using both spatial and temporal neighbor information. Results show the proposed algorithm has a much better performance for rainy scenes with large motion than existing algorithms.
Jie Chen 0026, Lap-Pui Chau
ISCAS2
2013 Expression-invariant and sparse representation for mesh-based compression for 3-D face models
abstract
Compression of mesh-based 3-D models has been an important issue, which ensures efficient storage and transmission. In this paper, we present a very effective compression scheme specifically for expression variation 3-D face models. Firstly, 3-D models are mapped into 2-D parametric domain and corresponded by expression-invariant parameterizaton, leading to 2-D image format representation namely geometry images, which simplifies the 3-D model compression into 2-D image compression. Then, sparse representation with learned dictionaries via K-SVD is applied to each patch from sliced GI so that only few coefficients and their indices are needed to be encoded, leading to low datasize. Experimental results demonstrate that the proposed scheme provides significant improvement in terms of compression performance, especially at low bitrate, compared with existing algorithms.
Junhui Hou, Lap-Pui Chau, Ying He 0001, Nadia Magnenat-Thalmann
VCIP2
2013 Rate-Distortion Model Based Bit Allocation for 3-D Facial Compression Using Geometry Video
abstract
With the extensive applications of 3-D multimedia technology, 3-D content compression has been an important issue, which ensures its smooth transmission on the network with constrained bandwidth. In this letter, we propose a new compression framework for dynamic 3-D facial expressions. Taking advantage of the near-isometric property of human facial expressions, we parameterize the dynamic 3-D faces into an expression-invariant canonical domain, which naturally generates 2-D geometry videos and allows us to apply the well-studied video compression techniques. Due to the difference from natural videos, each dimension (i.e., X, Y and Z, respectively) of the geometry video is regarded as a video sequence and encoded separately. Meanwhile, a model-based joint bit allocation scheme is designed to allocate reasonable bitrate to each dimension by detailed analysis of rate-distortion model for geometry videos, to obtain optimal results under given target bitrate. Experimental results show that up to 25% improvement in terms of bitrate reduction can be achieved, compared to existing algorithms.
Junhui Hou, Lap-Pui Chau, Ying He 0001, Minqi Zhang, Nadia Magnenat-Thalmann
IEEE Trans. Circuits Syst. Video Technol.2
2012 Joint rate allocation for statistical multiplexing of SVC
abstract
This paper presents a joint rate allocation scheme for statistical multiplexing of multiprogram video coding in broadcasting systems. The scheme is based on a scalable video coding (SVC) platform that does not require computationally expensive re-encoding or transcoding to adjust the bit-rate of each video program. A piecewise linear model is used to estimate the rate-distortion relationship in SVC enhancement layers. Based on the model, a joint rate allocation scheme is developed to dynamically allocate the available bandwidth by taking into consideration both inter-program fairness and intraprogram smoothness constraints. Experiments have been carried out to compare the performance of existing methods with our proposed scheme. Results demonstrate that the proposed scheme achieves a fine balance in picture quality across all statistical multiplexed programs as well as within each program.
Wei Yao 0001, Lap-Pui Chau, Susanto Rahardja
ICIP2
2012 Keyframe selection for motion capture using motion activity analysis
abstract
Motion capture data acquired from high definition cameras creates accurate human motion representation but introduces many redundant frames which pose a problem in data storage and motion retrieval purposes. In this paper, a keyframing approach is proposed to reduce the motion data by extracting keyframes using motion analysis approach in sampling windows. Motion changes in sampling windows for original motion without frame skipping and with frame skipping are computed. The difference in the motion changes is the main aspect in deciding whether the frames in sampling windows are possible candidates for keyframe selection. Simulation results showed that the proposed method is able to achieve an overall good visual quality for different types of motion. It also gives an improvement of up to 52% in terms of mean square error measurement, as compared to the existing keyframe extraction method, which is curve simplification method.
Ming-Hwa Kim, Lap-Pui Chau, Wan-Chi Siu
ISCAS2
2012 Image-driven simplification with single viewpoint
abstract
Image-driven simplification was proposed as a simplification method which generates models with high fidelity and intrinsically factors in the error from mesh appearance properties. This approach is very time consuming since it requires repeated render and image captures during simplification. Compared to the twenty viewpoints used in the original method, we propose a method which uses only a single viewpoint. The proposed single viewpoint simplification runs an estimated twenty-five times faster than the original method at comparable to better quality.
Cheen-Hau Tan, Lap-Pui Chau
ISCAS2
2011 A discriminative learning technique for mobile landmark recognition
abstract
This paper proposes a discriminative learning bags-of-words (BoW) approach for mobile landmark recognition at patch and image levels. Conventional methods often treat the local patches and images equally important for recognition and do not differentiate their different importance. Although there exist several works that consider the patches' discrimination information, they mainly focus on which patches are to be retained for training and do not incorporate this information when generating the BoW histograms. In view of this, this paper proposes to learn the discriminative information for each landmark category at two levels: local patches and images. At patch level, the patches' discrimination information for each landmark is first discovered using an iterative learning approach. This information is then incorporated into the quantization process to generate the BoW histogram. At image level, the different importance of training images is estimated through a non-parametric density estimator. Finally, fuzzy SVM is used to train the classifier for each category. Experimental results on a landmark database consisting of 3622 training images and 534 testing images show that the proposed method is effective in mobile landmark recognition.
Tao Chen 0003, Kim-Hui Yap, Lap-Pui Chau
ICIP3
2011 From universal bag-of-words to adaptive bag-of-phrases for mobile scene recognition
abstract
This paper proposes an adaptive bag-of-phrases (BoP) algorithm for mobile scene recognition based on bag-of- words approach. Conventional BoW methods do not consider the dependence and pairwise relationship among different codewords. However, these contextual relations between pairwise codewords play an important role for users to recognize an image. In light of this problem, this paper proposes an effective BoP technique to integrate both the spatial and contextual information between visual words for scene recognition. It first uses hierarchical k-means algorithm to construct a universal codebook for all categories. The contextual (dependence) relationship between pairwise words is then mined for each category based on the mutual information they contain. Subsequently, a visual phrase vocabulary is constructed which is then used to generate a BoP histogram through a proposed quantization method. Finally, support vector machine (SVM) is used to train these histograms into a classifier. Experimental results on the Scene 15 dataset show that the proposed method is effective for mobile scene recognition.
Tao Chen 0003, Kim-Hui Yap, Lap-Pui Chau
ICIP3
2011 Image based approach with k-mean clustering for the compression of human motion sequences
abstract
In this paper, an image format named Virtual Character Animation Image (VCAI) is presented for providing an efficient form of representation for humanoid motion data. By mapping the VCA motion information as 2-D images, characteristics of joint's correlation for the skeletal avatar and temporal coherence within the motion data are jointly reflected as spatial correlation of an image to aid compression. Since the VCA is now encoded as an image, the use of image processing tools and image delivery techniques are now possible. Lastly, a modified motion filter (MMF) is proposed to minimize the visual discontinuity in VCA's motion due to the quantization and transmission noise at high compression rate. The MMF helps to remove high frequency noise components and smoothen the motion signal providing perceptually improved VCA with reduction in distortion. Simulation results demonstrate the effectiveness of the proposed scheme ensuring the minor degradation of VCA quality measured by objective error metric and perceptual loss to the VCA for highly compressed motion stream.
Boon-Seng Chew, Lap-Pui Chau, Kim-Hui Yap
ISCAS2
2011 Fault tolerant design for low power hierarchical search motion estimation algorithms
abstract
Highly scaled CMOS devices are predicted to show probabilistic behavior due to process variations or the presence of noise sources such as thermal noise. Past research dealing with characterizing CMOS devices with probabilistic behavior has shown that computing via these devices, termed probabilistic computing, can help realize highly efficient circuits in terms of energy consumption. In this paper, we explore low power motion estimation, specifically low power hierarchical search algorithms for motion estimation, in the context of probabilistic computing. With the fault tolerant algorithm design (MC-TSS) proposed in this paper, we show that energy savings that can be realized with probabilistic computing increase to 70% versus 57% with the conventional algorithm (TSS), with minor impact on the quality of motion estimation. Furthermore, a 1.8 dB improvement in PSNR under the same energy savings of 70% for both algorithms is shown establishing the superior resilience of the proposed algorithm to probabilistic computing over the conventional algorithm.
Charvi Dhoot, Vincent John Mooney III, Shubhajit Roy Chowdhury, Lap-Pui Chau
VLSI-SoC4
2011 Integrated Content and Context Analysis for Mobile Landmark Recognition
abstract
This paper proposes a new approach for mobile landmark recognition based on integrated content and context analysis. Conventional scene/landmark recognition methods focus mainly on nonmobile desktop/PC platform, where content analysis alone is used to perform landmark recognition. These nonmobile systems, however, do not take unique features of mobile devices into consideration, e.g., limited computational power and fast response time requirement of mobile users. On the contrary, most existing context-aware content mobile landmark recognition methods mainly rely on global positioning system location information for context analysis. In view of this, this paper proposes an effective method that employs an integration of content and context analysis to perform landmark recognition using mobile devices. A new bags-of-words (BoW) framework is developed to perform content analysis. It is then integrated with context analysis involving fusion of location and direction information to perform mobile landmark recognition. Experimental results based on the NTU50Landmark database show that the proposed method can achieve good recognition performance in mobile landmark recognition.
Tao Chen 0003, Kim-Hui Yap, Lap-Pui Chau
IEEE Trans. Circuits Syst. Video Technol.3
2011 A Fuzzy Clustering Algorithm for Virtual Character Animation Representation
abstract
The use of realistic humanoid animations generated through motion capture (MoCap) technology is widespread across various 3-D applications and industries. However, the existing compression techniques for such representation often do not consider the implicit coherence within the anatomical structure of a human skeletal model and lacks portability for transmission consideration. In this paper, a novel concept virtual character animation image (VCAI) is proposed. Built upon a fuzzy clustering algorithm, the data similarity within the anatomy structure of a virtual character (VC) model is jointly considered with the temporal coherence within the motion data to achieve efficient data compression. Since the VCA is mapped as an image, the use of image processing tool is possible for efficient compression and delivery of such content across dynamic network. A modified motion filter (MMF) is proposed to minimize the visual discontinuity in VCA's motion due to the quantization and transmission error. The MMF helps to remove high frequency noise components and smoothen the motion signal providing perceptually improved VCA with lessened distortion. Simulation results show that the proposed algorithm is competitive in compression efficiency and decoded VCA quality against the state-of-the-art VCA compression methods, making it suitable for providing quality VCA animation to low-powered mobile devices.
Boon-Seng Chew, Lap-Pui Chau, Kim-Hui Yap
IEEE Trans. Multim.2
2010 Bit allocation for scalable video coding of multiple video programs
abstract
The problem of joint rate allocation for scalable video coding (SVC) of multiple video programs is addressed in this paper. Most of the existing approaches are based on non-scalable video coding platforms, where computationally expensive encoding or transcoding is demanded to adjust the bit-rate of each video program. Different from all these works, we develop a new statistical multiplexing system, where the scalable video coding technique is applied to compress the video programs. Experiments are carried out to verify the performance of the proposed scheme by comparing it with existing methods. The results demonstrate the merit of the proposed scheme and the variation of quality between different video programs is significantly reduced.
Yu Wang 0006, Lap-Pui Chau, Kim-Hui Yap
ICIP2
2010 Synchronized partial-body motion graphs
abstract
Motion graphs are regarded as a promising technique for interactive applications. However, the graphs are generated based on the distance metric of whole body, which produce a limit set of possible transitions. In this paper, we present an automatic method to construct a new data structure that specifies transitions and correlations between partial-body motions, called Synchronized Partial-body Motion Graphs (SPbMGs). We exploit the similarity between lower-body motions to create synchronization conditions with upper-body motions. Under these conditions, we generate all possible transitions between partial-body motions. The proposed graph representation not only maximizes the reusability of motion data, but also increases the connectivity of motion graphs while retaining the quality of motion.
William Wai-Lam Ng, Clifford S. T. Choy, Daniel Pak-Kong Lun, Lap-Pui Chau
SIGGRAPH ASIA (Sketches)4
2010 Adaptive resynchronization approach for scalable video over wireless channel
Yu Wang 0006, Lap-Pui Chau, Kim-Hui Yap
J. Vis. Commun. Image Represent.2
2010 Joint Rate Allocation for Multiprogram Video Coding Using FGS
abstract
In this paper, we address the problem of joint rate allocation for scalable video coding (SVC) of multiple video programs using the fine granularity scalability (FGS), which is not specified by any current H.264/AVC profile. Most of the existing approaches are based on non-SVC platforms, where computationally expensive encoding or transcoding is demanded to adjust the bit-rate of each video program. Different from all these works, we develop a new statistical multiplexing system, where FGS is applied to compress the video programs. First, we propose an efficient look-ahead approach to distribute the base layer coding bit-rate. Second, a piecewise linear model is applied to accurately estimate the rate-distortion relationship in the FGS layers. Based on this model, a novel algorithm is designed to dynamically allocate the available bandwidth to different programs for rate adaptation in order to minimize the variation of quality of different video programs. Experiments are carried out to verify the performance of the proposed scheme by comparing it with existing methods. The results demonstrate the superiority of the proposed scheme and the quality difference between different programs is greatly reduced.
Yu Wang 0006, Lap-Pui Chau, Kim-Hui Yap
IEEE Trans. Circuits Syst. Video Technol.2
2009 Progressive Transmission of Motion Capture Data for Scalable Virtual Character Animation
abstract
In this paper, a technique for transmitting level of details motion sequences for virtual character animation is proposed. We have demonstrated that by using progressive level of details (LOD) scheme on the motion capture information, efficient LOD representation of the skeleton motion can be obtained. The proposed virtual character encoder is scalable in nature providing a form of flexibility for the bit-stream to be partially decodable at any bit rate within the bit rate range to address the dynamic bandwidth constraint of a heterogeneous wireless network. Based on the structural characteristic of the virtual character, the packet will be delivered across dynamic channel condition to provide the viewer with the best reconstructed virtual character's motion quality.
Boon-Seng Chew, Lap-Pui Chau, Kim-Hui Yap
ISCAS2
2009 A Learning Approach for Single-frame Face Super-resolution
abstract
This paper presents a new learning approach for single-frame face super-resolution (SR). The aim of face SR is to estimate the missing high-resolution (HR) information from a single low-resolution (LR) face image by learning from training samples in the database. A commonly encountered issue in conventional face SR methods is that when the given LR image is a new face significantly different from those in the database, the quality of the reconstructed HR face is usually unsatisfactory. To alleviate this difficulty, we develop a new method to perform face SR based on principal component analysis (PCA) and locally linear embedding (LLE). The reconstructed HR face is able to preserve standard facial features and detailed local information through a residue prediction method using manifold learning. Experimental results show that the proposed method is effective in performing single-frame face SR.
Kim-Hui Yap, Lap-Pui Chau
ISCAS3
2009 Efficient Inter Mode Decision for H.263 to H.264 Video Transcoding using Support Vector Machines
abstract
This paper presents an efficient mode decision algorithm for H.263 to H.264 interframe transcoding. The proposed scheme uses a support vector machines (SVMs) approach to investigate the relationship between data extracted from H.263 decoding stage and the optimal coding mode in H.264 re-encoding process. Based on the SVM classifier, the transcoder only enables a subset of candidate modes for each macroblock in the rate-distortion-optimized mode decision in H.264. The objective is to eliminate unlikely modes in earlier stages in order to achieve computation saving. Simulation results show that the proposed method can reduce the computational complexity of interframe transcoding by up to 77% while maintaining similar rate-distortion performance.
Xuan Jing, Wan-Chi Siu, Lap-Pui Chau, Anthony G. Constantinides
ISCAS3
2009 Broadcast of Scalable Video over Wireless Networks
abstract
Scalable video coding technique has been desired for many years to realize a reliable transmission of video over heterogeneous networks. In this paper, we present a new system of scalable video broadcasting over wireless networks. For each group of pictures, a number of quality layers are produced using scalable video coding. We design the channel protection schemes using the forward error correction codes for the base layer and the enhancement layers, respectively. Given the clients' distribution, we propose a novel algorithm to determine both the source coding bit-rate and the channel coding bit-rate for each layer to maximize a system-defined utility function. Experimental results can demonstrate the superior of the proposed scheme to other schemes and the improvement is up to 2 dB.
Yu Wang 0006, Lap-Pui Chau, Kim-Hui Yap
ISCAS2
2009 Streaming 3D meshes using spectral geometry images
abstract
The transmission of 3D models in the form of Geometry Images (GI) is an emerging and appealing concept due to the reduction in complexity from R3 to image space and wide availability of mature image processing tools and standards. However, geometry images often suffer from the artifacts and error during compression and transmission. Thus, there is a need to address the artifact reduction, error resilience and protection of such data information during the transmission across an error prone network. In this paper, we introduce a new concept, called Spectral Geometry Images (SGI), which naturally combines the powerful spectral analysis with geometry images. We show that SGI is more effective than GI to generate visually pleasing shapes at high compression rates. Furthermore, by coupling SGI to the proposed error protection scheme, we are able to ensure the smooth delivery of 3D model across error networks for different packet loss rate simulated using the two-state Markov model.
Ying He 0001, Boon-Seng Chew, Steven C. H. Hoi, Lap-Pui Chau
ACM Multimedia5
2009 A soft MAP framework for blind super-resolution image reconstruction
Kim-Hui Yap, Li Chen 0011, Lap-Pui Chau
Image Vis. Comput.4
2009 A Nonlinear L 1 -Norm Approach for Joint Image Registration and Super-Resolution
abstract
This letter proposes a nonlinearL1-norm approach for joint image registration and super-resolution (SR). Image SR is the fusion of multiple low-resolution (LR) images to produce a high-resolution (HR) image. Conventional SR algorithms are sensitive to the initial registration error and outliers in the LR images. In view of this, we present a new SR method to address these problems usingL1-norm optimization in joint image registration and HR image reconstruction. Experimental results show that the proposed method is effective in handling these issues in the HR image reconstruction.
Kim-Hui Yap, Yushuang Tian, Lap-Pui Chau
IEEE Signal Process. Lett.4
2008 A new color image regularization scheme for blind image deconvolution
abstract
This paper proposes a new regularization scheme to address blind color image deconvolution. Conventional blind monochromatic image deconvolution algorithms handle each color channel independently, thereby ignoring the inter-channel correlation present in the color images. Further, most existing blind color deconvolution algorithms do not take the parametric information of the blurs into consideration. In view of these, a regularization scheme is proposed to perform blind color image deconvolution. A new regularization operator is developed in the blur domain. A reinforcement blur modeling scheme is adopted to evaluate the relevance of manifold parametric blur structures, and the information is integrated into the deconvolution scheme. In addition, a regularization scheme for image is developed to recover edges of color images and reduce color artifacts. Experimental results show that the method is able to achieve satisfactory restored color images under noisy environment.
Kim-Hui Yap, Li Chen 0011, Lap-Pui Chau
ICASSP4
2008 Spatial resolution decision in scalable bitstream extraction for network and receiver aware adaptation
abstract
Reliable transmission of video over heterogeneous networks requires efficient coding, as well as scalability to different client capabilities, system resources and network conditions. Scalable video coding can provide a full scalability comprising temporal scalability, spatial scalability and quality scalability to increase its adaptability to network and client conditions. It encodes the original video at a full resolution, but enables extracting partial streams to reconstruct the video depending on the specific rate and resolution required by a certain application. This paper addresses the problem of scalable bitstream extraction. Given the bandwidth constraint and the display resolution of the end user, the proposed algorithm will decide the spatial resolution of the sub-stream to be extracted based on the analysis of the content information to maximize the perceptual video quality. Experimental results demonstrate the efficiency of the proposed algorithm.
Yu Wang 0006, Lap-Pui Chau, Kim-Hui Yap
ICME2
2008 Frame Complexity-Based Rate-Quantization Model for H.264/AVC Intraframe Rate Control
abstract
In this letter, we present an adaptive intraframe rate-quantization (R-Q) model for H.264/AVC video coding. The proposed method aims at selecting accurate quantization parameters (QP) for intra-coded frames according to the target bit rate. By taking gradient-based frame complexity measure into consideration, the model parameters can be adaptively updated. Experimental results show that when employing our proposed R-Q model, the intraframe target bits mismatch ratio can be reduced by up to 75% as compared to the traditional Cauchy-density-based model. Hence, this is extremely useful for H.264/AVC rate control applications.
Xuan Jing, Lap-Pui Chau, Wan-Chi Siu
IEEE Signal Process. Lett.2
2008 Intra/Inter Macroblock Mode Decision for Error-Resilient Transcoding
abstract
When transmitting the precoded bitstream over an error-prone network, error-resilient transcoding is adopted to convert the bitstream to a resilient format for robust delivery. Intra refreshment is an efficient tool to reduce the dependency between frames and stop the channel error propagation. In the conventional scheme, the rate-distortion optimized macroblock mode decision is employed to adaptively determine the coding mode of each macroblock. However, this scheme only considers the channel error propagated from the previous frames to the current frame. As opposed to this traditional algorithm, this paper proposes a method which considers consecutive two frames in a sequence, thus taking the error propagation to the following frame into account. This enhances the overall robustness of the transcoded bitstream against the packet loss. Considering the availability of the next frame information, two cases are discussed respectively. Experimental results show that the proposed methods present quality improvement when compared with the conventional rate-distortion optimized error-resilient coding scheme under different test environments, and the PSNR improvement can reach as high as 0.9 dB.
Haiyan Shu, Lap-Pui Chau
IEEE Trans. Multim.2
2008 A Novel Hybrid Model Framework to Blind Color Image Deconvolution
abstract
This paper presents a new hybrid model framework to address blind color image deconvolution. Blind color image deconvolution is a challenging problem due to the limited information on the blurring function. Conventional methods based on the single-input single-output (SISO) model experience suboptimal results as each color channel is processed independently. On the other hand, there are limitations on the practicality of using a multiinput multioutput (MIMO) model in solving this problem as the color channels are usually highly correlated. In view of these constraints, this paper proposes a novel framework to solve blind color image deconvolution by first decomposing the color channels into wavelet subbands, and performing image deconvolution using a hybrid of SISO and single-input multioutput models. The proposed method utilizes the correlation information among different color channels to alleviate the constraints imposed by the MIMO systems. Experimental results show that the method is able to achieve satisfactory restored images under different noise and blurring environments.
Kim-Hui Yap, Li Chen 0011, Lap-Pui Chau
IEEE Trans. Syst. Man Cybern. Part A4
2007 Joint Image Registration and Super-Resolution using Nonlinear Least Squares Method
abstract
This paper proposes a new algorithm to integrate image registration into image super-resolution (SR) by fusing multiple blurred low-resolution (LR) images to render a high-resolution (HR) image. Conventional super-resolution (SR) image reconstruction algorithms assume either the estimated motion (displacement) errors by existing registration methods are negligible or the displacement is known a priori. This assumption, however, is impractical as the performance of existing registration algorithms is still less than perfect. In view of this, we present a new estimation framework that performs joint image registration and HR reconstruction. An iterative scheme based on nonlinear least squares method is developed to estimate the motion shift (displacement) and HR image progressively. The motion model that is considered in this work includes both translation as well as rotation. Experimental results show that the proposed method is effective in performing image super-resolution.
Kim-Hui Yap, Li Chen 0011, Lap-Pui Chau
ICASSP (1)4
2007 A Motion-Based Selective Error Protection Method for Scalable Video Over Error-Prone Channel
abstract
Video transmission over unreliable networks introduces new challenges in video coding. Due to the predictive coding techniques, the effect of channel errors on the decoded video can be extremely severe when the compressed video is transmitted over error-prone channel. In this paper, the problem of scalable video transmission over error-prone channel is addressed. It is proposed to selectively add forward error correction (FEC) codes to partial information of the compressed bit-stream based on the motion activity of the input video. In addition, unequal error protection (UEP) is applied on the selected data of different temporal layers in a group of pictures (GOP), where the channel rates are optimally allocated. It is shown from the experimental results that our proposed method has a good performance and the improvement is up to 1.2 dB.
Yu Wang 0006, Lap-Pui Chau, Kim-Hui Yap
ICME2
2007 Improved Frame Level MAD Prediction and Bit Allocation Scheme for H.264/AVC Rate Control
abstract
In this paper, we present an improved frame level mean absolute difference (MAD) prediction model for H.264/AVC rate control. Based on the histogram information of successive frames, the MAD of current frame can be accurately estimated. Instead of only using buffer status in the frame target bits allocation process, we also take into consideration the frame complexity which depends on the improved predicted frame MAD value. The objective of this frame complexity based bit allocation is to follow the non-stationary characteristics of video source and to provide smoother visual quality. Simulation results show that by using our proposed scheme, smaller frame target bits mismatch can be achieved. In addition, the average peak signal-to-noise ratio (PSNR) for the reconstructed video is improved by up to 0.5 dB with up to 62% reduction in picture quality variation.
Xuan Jing, Lap-Pui Chau
ISCAS2
2007 Partial Distortion Search Algorithm Using Predictive Search Area for Fast Full-Search Motion Estimation
abstract
In this letter, a fast partial distortion search algorithm for motion estimation is presented. The proposed method is based on the observation that when normalized partial distortion is utilized, the false rejection of impossible candidates most likely occurs within a small area adjacently located near the true motion vector. Our objective is to enhance the prediction accuracy in this small predictive search area and further save computations outside this area. Experimental results show that the proposed algorithm achieves an average 42 times speedup ratio as compared to full-search algorithm with only 0.05 dB degradation in PSNR performance.
Xuan Jing, Lap-Pui Chau
IEEE Signal Process. Lett.2
2007 A Resizing Algorithm With Two-Stage Realization for DCT-Based Transcoding
abstract
In video transcoding, arbitrary resizing is always desired for end users with different display capabilities. In this paper, a two-stage resizing structure is proposed for arbitrary resizing. It gives a general solution for the realization of arbitrary resizing in the transform domain. Two constraints are given for the selection of parameter set to benefit the anti-aliasing and picture quality. Illustrative examples show that our proposed two-stage resizing structure, with the qualified parameter sets, can provide an appropriate solution for arbitrary resizing
Haiyan Shu, Lap-Pui Chau
IEEE Trans. Circuits Syst. Video Technol.2
2007 A Nonlinear Least Square Technique for Simultaneous Image Registration and Super-Resolution
abstract
This paper proposes a new algorithm to integrate image registration into image super-resolution (SR). Image SR is a process to reconstruct a high-resolution (HR) image by fusing multiple low-resolution (LR) images. A critical step in image SR is accurate registration of the LR images or, in other words, effective estimation of motion parameters. Conventional SR algorithms assume either the estimated motion parameters by existing registration methods to be error-free or the motion parameters are known a priori. This assumption, however, is impractical in many applications, as most existing registration algorithms still experience various degrees of errors, and the motion parameters among the LR images are generally unknown a priori. In view of this, this paper presents a new framework that performs simultaneous image registration and HR image reconstruction. As opposed to other current methods that treat image registration and HR reconstruction as disjoint processes, the new framework enables image registration and HR reconstruction to be estimated simultaneously and improved progressively. Further, unlike most algorithms that focus on the translational motion model, the proposed method adopts a more generic motion model that includes both translation as well as rotation. An iterative scheme is developed to solve the arising nonlinear least squares problem. Experimental results show that the proposed method is effective in performing image registration and SR for simulated as well as real-life images.
Kim-Hui Yap, Li Chen 0011, Lap-Pui Chau
IEEE Trans. Image Process.4
2007 Two-Dimensional Channel Coding Scheme for MCTF-Based Scalable Video Coding
abstract
The motion-compensated temporal filtering (MCTF)-based scalable video coding (SVC) provides a full scalability including spatial, temporal and signal-to-noise ratio (SNR) scalability with fine granularity, each of which may result in different visual effect. This paper addresses a novel approach of two-dimensional unequal error protection (2D UEP) for the scalable video with a combined temporal and quality (SNR) scalability over packet-erasure channel. The bit-stream is divided into scalable subbitstreams based on the structure of MCTF. Each subbitstream is further divided into several quality layers. Unequal quantities of bits are allocated to protect different layers to obtain acceptable quality video with smooth degradation under different transmission error conditions. Experimental results are presented to show the advantage of the proposed 2D UEP scheme over the traditional one-dimensional unequal error protection (1D UEP) scheme. Comparing the proposed method with the 1D UEP scheme on SNR layers, our method gives up to 0.81-dB improvement for some video sequences
Yu Wang 0006, Lap-Pui Chau, Kim-Hui Yap
IEEE Trans. Multim.3
2006 Blind Super-Resolution Image Reconstruction using a Maximum a Posteriori Estimation
abstract
This paper proposes a new algorithm to address blind image super-resolution by fusing multiple blurred low-resolution (LR) images to render a high-resolution (HR) image. Conventional super-resolution (SR) image reconstruction algorithms assume either the blurring during the image formation process is negligible or the blurring function is known a priori. This assumption, however, is impractical as it is difficult to eliminate blurring completely in some applications or characterize the blurring function fully. In view of this, we present a new maximum a posteriori (MAP) estimation framework that performs joint blur identification and HR image reconstruction. An iterative scheme based on alternating minimization is developed to estimate the blur and HR image progressively. A blur prior that incorporates the soft parametric blur information and smoothness constraint is introduced in the proposed method. Experimental results show that the new method is effective in performing blind SR image reconstruction where there is limited information about the blurring function.
Kim-Hui Yap, Li Chen 0011, Lap-Pui Chau
ICIP4
2006 A Novel Resynchronization Method for Scalable Video Over Wireless Channel
abstract
A scalable video coder generates scalable compressed bit-stream, which can provide different types of scalability depend on different requirements. This paper proposes a novel resynchronization method for the scalable video with combined temporal and quality (SNR) scalability. The main purpose is to improve the robustness of the transmitted video. In the proposed scheme, the video is encoded into scalable compressed bit-stream with combined temporal and quality scalability. The significance of each enhancement layer unit is estimated properly. A novel resynchronization method is proposed where joint group of picture (GOP) level and picture level insertion of resynchronization marker approach is applied to insert different amount of resynchronization markers in different enhancement layer units for reliable transmission of the video over error-prone channels. It is demonstrated from the experimental results that the proposed method can perform graceful degradation under a variety of error conditions and the improvement can be up to 1 dB compared with the conventional method
Yu Wang 0006, Lap-Pui Chau, Kim-Hui Yap
ICME2
2006 Region-Based Image Retrieval using Radial Basis Function Network
abstract
This paper presents a new framework that integrates relevance feedback into region-based image retrieval (RBIR) systems based on radial basis function network (RBFN). A modified unsupervised subtractive clustering algorithm is proposed for RBFN center selection according to the characteristics of region-based image representation. A new kernel function of RBFN is introduced for image similarity comparison under region-based representation. The underlying network parameters (weight and width) are then optimized using a supervised gradient-descent training strategy. Experimental results using a database of 10,000 images demonstrate the effectiveness of the proposed hybrid learning approach
Kui Wu 0002, Kim-Hui Yap, Lap-Pui Chau
ICME3
2006 A novel intra-rate estimation method for H.264 rate control
abstract
In this paper, we present a novel intra-rate estimation method for H.264 rate control. The proposed method establishes the intra-frame rate-quantization (R-Q) estimation model by using gradient-based picture complexity measure. The proposed model aims at selecting accurate quantization parameters for intra-coded frames which can fully comply with buffer constraints. By adaptively updating this R-Q model, the performance of the model can be guaranteed regardless of the changing of video frame characteristics. Simulation results show that by using our proposed scheme, better rate control for intra-frames and frames at scene changes can be achieved. As a result, the number of skipped frames is significantly reduced thus it provides improved visual quality of the reconstructed pictures.
Xuan Jing, Lap-Pui Chau
ISCAS2
2006 Fast mode decision for spatial scalable video coding
abstract
Scalable video coding (SVC) is an on-going standard and the current working draft (WD) is an extension of H.264/AVC. It provides scalability at the bit stream level with good compression efficiency, and allowing free combinations of spatial, temporal and SNR scalability. In the WD, exhaustive search technique is employed to select the best coding mode for each macroblock. This technique achieves highest possible coding efficiency, but it results in higher computational complexity. To overcome this, we propose a novel fast mode decision scheme for spatial scalability in SVC. In this scheme, the mode distribution relationship between base layer and enhancement layers is employed to reduce the candidate mode set at enhancement layers. The experimental results show that the proposed scheme provides significant reduction in computational complexity without any noticeable coding loss.
Zhengguo Li, Changyun Wen, Lap-Pui Chau
ISCAS4
2006 Generalized arbitrary resizing for video transcoding
abstract
In video transcoding, arbitrary resizing is always desired for end users with different display capabilities. In this paper, a two-stage resizing structure is proposed for arbitrary resizing. It gives a general solution for the realization of arbitrary resizing in the transform domain. Two constraints are given for the selection of parameter set to benefit the anti-aliasing and the picture quality. Illustrative examples show that our proposed two-stage resizing structure, with qualified parameter set, can provide a satisfactory solution for arbitrary resizing.
Haiyan Shu, Lap-Pui Chau
ISCAS2
2006 Two-dimensional channel rate allocation for SVC over error-prone channel
abstract
The motion compensated temporal filtering (MCTF) based scalable video coding (SVC) provides a full scalability including spatial, temporal and signal-to-noise ratio (SNR) scalability with fine granularity, each of which may result in different visual effect. This paper addresses a novel approach of two-dimensional unequal error protection (2D UEP) for the scalable video with a combined temporal and quality (SNR) scalability over packet-erasure channel. The bit-stream is divided into scalable sub-bit-streams based on the structure of MCTF. Each sub-bit-stream is further divided into several quality layers. Unequal quantities of bits are allocated to protect different layers to obtain acceptable quality video with smooth degradation under different transmission error conditions. Experimental results are presented to show the advantage of the proposed 2D UEP scheme over the traditional one-dimensional unequal error protection (1D UEP) scheme.
Yu Wang 0006, Lap-Pui Chau, Kim-Hui Yap
ISCAS3
2006 The realization of arbitrary downsizing video transcoding
abstract
In order to transmit the precoded bit stream through bandwidth constrained network, transcoding technique has been introduced to reduce the overall bit rate generated. Downsizing transcoding is one of the commonly used methods since it can reach very low bit rate without losing any motion information. Many efforts have been devoted to this research area. With the proposal of arbitrary downsizing, fine gradual reduction of bit rate and video quality becomes feasible. However, there is little research focusing on frame size selection for arbitrary downsizing transcoding. In this letter, the realization of arbitrary downsizing transcoding is discussed. Rate estimation scheme is applied on video sequences to choose the suitable frame size for the target bit rate. An efficient method is proposed to estimate the number of bits allocated to residue in requantization process. Experimental results show that, it can provide an accurate estimation and make a reliable decision in frame size selection.
Haiyan Shu, Lap-Pui Chau
IEEE Trans. Circuits Syst. Video Technol.2
2006 GOP-based channel rate allocation using genetic algorithm for scalable video streaming over error-prone networks
abstract
In this paper, we address the problem of unequal error protection (UEP) for scalable video transmission over wireless packet-erasure channel. Unequal amounts of protection are allocated to the different frames (I- or P-frame) of a group-of-pictures (GOP), and in each frame, unequal amounts of protection are allocated to the progressive bit-stream of scalable video to provide a graceful degradation of video quality as packet loss rate varies. We use a genetic algorithm (GA) to quickly get the allocation pattern, which is hard to get with other conventional methods, like hill-climbing method. Theoretical analysis and experimental results both demonstrate the advantage of the proposed algorithm.
Lap-Pui Chau
IEEE Trans. Image Process.2
2005 Blind color image deconvolution based on wavelet decomposition
abstract
This paper presents a new framework to address blind color image deconvolution based on wavelet decomposition. Blind color image deconvolution is a challenging problem due to the lack of information available. Conventional methods based on single-input single-output (SISO) model experience significant color artifacts in the restored images. On the other hand, there are limitations on the practicality of using multi-input multi-output (MIMO) model in solving this problem as the color channels are usually highly correlated. In view of this, the paper proposes a new framework to solve blind color image deconvolution by first decomposing the color channels into wavelet subbands, and performing image deconvolution using a combination of SISO and single-input multi-output (SIMO) models. Experimental results show that the proposed method is able to achieve satisfactory restored images.
Kim-Hui Yap, Li Chen 0011, Lap-Pui Chau
ICIP (2)4
2005 Efficient content-based resynchronization approach for wireless video
abstract
Recent advances in technology have caused a significant growth in wireless communications, which have resulted in a strong demand for reliable transmission of video data. The challenge of robust video transmission is to protect the compressed data against hostile channel conditions while bringing little impact on bandwidth efficiency. In this paper, using results from a simplified macroblock-based segmentation algorithm, we propose a framework called content-based resynchronization for the effective positioning of resynchronization markers such that the image quality of foreground can be improved at the expense of sacrificing unimportant background. We do this because, in applications such as video telephony and video conferencing, foreground is typically the most important image region for viewers. Experimental results demonstrate that this scheme significantly improve the perceptual quality of video sequences for robust video transmission.
Lap-Pui Chau
IEEE Trans. Multim.2
2005 An error-resilient GOP structure for robust video transmission
abstract
Recent advances in technology have resulted in a significant growth in wireless communications and widespread access to information via the Internet, which have resulted in a strong demand for reliable transmission of video data. The challenge of robust video transmission is to protect the compressed data against hostile channel conditions while bringing little impact on bandwidth efficiency. In motion-compensated video-coding schemes, such as MPEG-1 or MPEG-2, an I frame normally is followed by several P frames and possibly B frames in a group-of-pictures (GOP). In error-prone environments, error happening in the previous frames in a GOP may propagate to all the following frames until the next I frame, which is the beginning of the next GOP. In this paper, we propose a novel GOP structure for robust transmission of MPEG video bitstream. By selecting the optimal position of the I frame in a GOP, robustness can be achieved without reducing any coding efficiency. Another advantage of the proposed GOP structure is also analyzed: compared with the conventional GOP structure, it provides reverse-play operation for MPEG video streaming with much less requirement on the network bandwidth. Experimental results demonstrate both the robustness of the proposed GOP structure and the efficient reverse-play functionality it leads to.
Lap-Pui Chau
IEEE Trans. Multim.2
2005 Efficient motion vector recovery algorithm for H.264 based on a polynomial model
abstract
In this paper, we propose an efficient motion vector recovery algorithm for the new coding standard H.264, which is based on a polynomial model. To achieve better coding efficiency, the motion estimation scheme used in H.264 is different from previous coding standards. In H.264, a 16/spl times/16 macroblock can be divided into different block shapes for motion estimation. Each macroblock contains more motion vectors than previous coding standards. For nature video, the blocks within a small area likely belong to the same object, hence the motion vectors of neighboring blocks are highly correlated. Based on the correlation of neighboring motion vectors, we can use the motion vectors that are adjacent to the lost motion vectors to constitute a polynomial model, which can describe the change tendency of motion vectors within a small area. Through this model, the lost motion vectors can be predicted and the lost macroblocks can be reconstructed. Different video sequences are used to test the performance of proposed method. The simulation results show that the quality of corrupted video can be obviously improved by proposed algorithm.
Jinghong Zheng 0001, Lap-Pui Chau
IEEE Trans. Multim.2
2004 An efficient resynchronization technique for perceptual quality enhancement for robust video transmission
abstract
In this paper, using results from a simplified macroblock (MB)-based segmentation algorithm, we propose a framework called content-based resynchronization (CBR) for the effective positioning of resynchronization markers such that the image quality of the foreground can be improved at the expense of sacrificing an unimportant background. We do this because, in applications such as video telephony and video conferencing, the foreground is typically the most important image region for viewers. Experimental results demonstrate that this scheme significantly improves the subjective quality of video sequence for robust video transmission.
Lap-Pui Chau
ICASSP (3)1
2004 Reducing drift for FGS coding based on multiframe motion compensation [video coding]
abstract
Fine granularity scalability (FGS) in MPEG-4 and some enhanced FGS coding techniques have been studied recently for video streaming. However, a problem of drifting may emerge due to the difference of reference frames used in the encoder and decoder for those enhanced FGS coding schemes. In this paper, we propose incorporating multiframe-based motion compensation into an FGS coding scheme to further improve the coding efficiency and error robustness. It has been shown that the multiple frames based FGS coding scheme has better coding efficiency than the conventional one-frame based FGS video coding, and more importantly, the new one can alleviate the drifting problem significantly. The proposed approach is implemented within a one-loop FGS coding scheme based on an H.26L framework.
Ce Zhu, Lap-Pui Chau
ICASSP (3)3
2004 Content-based periodic macroblock for error-resilient transmission of H.264 video
abstract
For the compressed video, the transmission error in one frame will propagate to the following frames and the qualities of successive frames will drop seriously. In this paper, we present a new error-resilience scheme to alleviate the effect of error propagation in video transmission for the new coding standard H.264. In this new coding standard, multiple reference frame is adopted to achieve better coding efficiency than previous coding standard. By making use of the reference frame buffer in encoder, we can make some macroblocks in every p inter frame to reference the frame that is p frames interval away and these macroblocks are named as periodic macroblocks. The periodic macroblock can efficiently alleviate the error propagation between the frames that contain periodic macroblocks. The selection of periodic macroblock is based on the distortion expectation of each macroblock in every p frame. The number of periodic macroblock in every p frame can be determined by the transmission bandwidth, as the periodic macroblock will consume a little more bits. The simulation results prove that the periodic macroblocks can obviously improve the quality of video with different macroblock loss rate.
Jinghong Zheng 0001, Lap-Pui Chau
ICIP2
2004 An efficient inter mode decision approach for H.264 video coding
abstract
Variable size block motion estimation is a very important technique for video coding. The new H.264 standard employs 7 different size block types which can significantly improve the coding performance compared with the previous video coding standards. On the other hand, the computational complexity of H.264 encoder increases dramatically due to the various coding modes used. An efficient inter mode decision approach is presented. The objective is to reduce the number of candidate block types in the motion estimation while maintaining the coding efficiency. Experimental results show that the proposed method can save the computation cost by up to 42% at the same PSNR and bitrate.
Xuan Jing, Lap-Pui Chau
ICME2
2004 Efficient inner search for faster diamond search
Ce Zhu, Xiao Lin 0001, Lap-Pui Chau, Hock-Ann Ang, Choo-Yin Ong
Signal Process.3
2004 An efficient arbitrary downsizing algorithm for video transcoding
abstract
When delivering video over communication networks, it is required to transcode the precoded video content to meet the demands of a broad range of end users with different bandwidths and resource constraints. One solution to transmit video over bandwidth-constrained channels is to reduce the spatial resolution of video frames and transmit a low-resolution version of the video as a tradeoff for the bit rate. In this paper, an arbitrary downsizing algorithm is proposed. This arbitrary downsizing algorithm is processed directly in the discrete cosine transform domain. Experimental results show that the proposed method can achieve a satisfactory performance. Compared to the existing methods, this algorithm is not only applicable to an intracoding frame but also to an intercoding frame. This is one of the crucial features of this method.
Haiyan Shu, Lap-Pui Chau
IEEE Trans. Circuits Syst. Video Technol.2
2004 Enhanced hexagonal search for fast block motion estimation
abstract
Fast block motion estimation normally consists of low-resolution coarse search and the following fine-resolution inner search. Most motion estimation algorithms developed attempt to speed up the coarse search without considering accelerating the focused inner search. On top of the hexagonal search method recently developed, an enhanced hexagonal search algorithm is proposed to further improve the performance in terms of reducing number of search points and distortion, where a novel fast inner search is employed by exploiting the distortion information of the evaluated points. Our experimental results substantially justify the merits of the proposed algorithm.
Ce Zhu, Xiao Lin 0001, Lap-Pui Chau, Lai-Man Po
IEEE Trans. Circuits Syst. Video Technol.3
2004 An efficient three-step search algorithm for block motion estimation
abstract
The three-step search algorithm has been widely used in block matching motion estimation due to its simplicity and effectiveness. The sparsely distributed checking points pattern in the first step is very suitable for searching large motion. However, for stationary or quasistationary blocks it will easily lead the search to be trapped into a local minimum. In this paper we propose a modification on the three-step search algorithm which employs a small diamond pattern in the first step, and the unrestricted search step is used to search the center area. Experimental results show that the new efficient three-step search performs better than new three-step search in terms of MSE and requires less computation by up to 15% on average.
Xuan Jing, Lap-Pui Chau
IEEE Trans. Multim.2
2004 Error-concealment algorithm for H.26L using first-order plane estimation
abstract
In this paper, we propose a new error-concealment algorithm for the forthcoming video coding standard H.26L, which makes use of a first-order plane to estimate motion vectors. In H.26L, a 16/spl times/16 inter macroblok can be divided into variant block shape for motion prediction, and there are up to sixteen sets of motion vector in one macroblock. For nature image, the motions within a small area are likely to move in the same direction. By using the motion vectors that are next to the vertices of lost macroblock, we can constitute a first-order plane that indicates the movement tendency in this small area, and estimate the motion vector of vertices. Then we use the motion vectors of vertices to interpolate motion vector for each pixel separately. The interpolation function we selected makes the motion change smoothly within the lost macroblock. The simulation results show that our method can efficiently improve the video quality over different macroblock lost rate.
Jinghong Zheng 0001, Lap-Pui Chau
IEEE Trans. Multim.2
2003 Efficient three-step search algorithm for block motion estimation in video coding
abstract
The three-step search algorithm has been widely used in block matching motion estimation due to its simplicity and effectiveness. The sparsely distributed checking points pattern in the first step is very suitable for searching large motion. However, for quasi-stationary blocks it will easily lead the search to be trapped into a local minimum. In this paper we propose a modification on the three-step search algorithm which employs a small diamond pattern in the first step, and the unrestricted search step is used to search the center area. Experimental results show that the proposed algorithm performs better than new three-step search in terms of MSE and requires less computation by up to 15% on average.
Lap-Pui Chau, Xuan Jing
ICASSP (3)1
2003 Smooth constrained block matching criterion for motion estimation
abstract
In this paper, a novel and efficient criterion for block matching motion estimation is presented. The proposed criterion is to enhance the conventional mean absolute difference (MAD) scheme with a new smoothness constraint on the residue block. The objective is to reduce the bit rate for encoding the residue image without any degradation of the reconstructed image quality. Simulation results show that by applying the new criterion in motion estimation, both the improvement in peak signal-to-noise ratio (PSNR) and the reduction in bit rate up to 4.3% can be achieved compared to MAD.
Xuan Jing, Ce Zhu, Lap-Pui Chau
ICASSP (3)3
2003 A fast arbitrary downsizing algorithm for video transcoding
abstract
Video transmission requires the delivery of video content to a broad range of end users with different bandwidth and resource constraints. One solution to transmitting video over bandwidth-constrained channels is to transmit a low-resolution version of video as a trade off for the bit-rate. In this paper, we proposed an arbitrary downsizing algorithm. This algorithm can downsize the frame to arbitrary size and process it directly in DCT domain. Experimental results show that the proposed method can achieve a comparable performance. Compared with the existing methods, this algorithm is not only applicable to intra coding frame but also inter coding frame. This is one of the crucial features of this method.
Haiyan Shu, Lap-Pui Chau
ICIP (1)2
2003 A fast octagon-based search algorithm for motion estimation
Lap-Pui Chau, Ce Zhu
Signal Process.1
2003 Smooth constrained motion estimation for video coding
Xuan Jing, Ce Zhu, Lap-Pui Chau
Signal Process.3
2003 Efficient multiplier structure for realization of the discrete cosine transform
Lap-Pui Chau, Wan-Chi Siu
Signal Process. Image Commun.1
2002 Arbitrary downsizing video transcoding using fast motion vector reestimation
abstract
We propose a video transcoding method for arbitrarily downsizing a precoded video by reestimating from its original motion vectors the new motion vectors required to code the downsized video. Compared with the existing methods for downsizing a precoded video by an integral factor, the main advantage of the proposed method is that the spatial resolution of the precoded video can be freely adjusted to meet different bandwidth and device requirements. Experimental results show that the proposed method can obtain arbitrarily downsized video with good perceptual quality while reducing the computational complexity of the process.
Yongqing Liang 0002, Lap-Pui Chau, Yap-Peng Tan
IEEE Signal Process. Lett.2
2002 Hexagon-based search pattern for fast block motion estimation
abstract
In block motion estimation, a search pattern with a different shape or size has a very important impact on search speed and distortion performance. A square-shaped search pattern is adopted in many popular fast algorithms. Recently, a diamond-shaped search pattern was introduced in fast block motion estimation and has exhibited a faster search speed. Based on an in-depth examination of the influence of the search pattern on speed performance, we propose a novel algorithm using a hexagon-based search pattern to achieve further improvement. The hexagon-based search pattern is investigated in comparison with diamond search pattern and demonstrates significant speedup gain over the diamond-based search. Analysis shows that a speed improvement rate of the hexagon-based search (HEXBS) algorithm over the diamond search (DS) algorithm can be over 80% for locating some motion vectors in certain scenarios. In short, the proposed HEXBS algorithm can find the same motion vector with fewer search points than the DS algorithm. Generally speaking, the larger the motion vector, the more search points the. HEXBS algorithm can save, which is further justified by experimental results.
Ce Zhu, Xiao Lin 0001, Lap-Pui Chau
IEEE Trans. Circuits Syst. Video Technol.3
2001 A novel hexagon-based search algorithm for fast block motion estimation
abstract
In block motion estimation, search patterns with different shape or size have a very important impact on search speed and distortion performance. In this paper, we propose a novel algorithm using a hexagon-based search (HEXBS) pattern for fast block motion estimation. The proposed HEXBS algorithm may find any motion vector with fewer search points than the diamond search (DS) algorithm. The speedup gain of the HEXBS method over the DS algorithm is more striking for finding large motion vectors. Experimental results substantially justify the fastest performance of the HEXBS algorithm compared with several other popular fast algorithms.
Ce Zhu, Xiao Lin 0001, Lap-Pui Chau, Keng-Pang Lim, Hock-Ann Ang, Choo-Yin Ong
ICASSP3
2000 Recursive algorithm for the realization of the discrete cosine transform
abstract
Recursive algorithms have been found very effective for realization using software and VLSI techniques. Recently, some recursive algorithms have been proposed for the realization of the discrete cosine transform (DCT). In this paper, an efficient recursive algorithm for the computation of the DCT is proposed. By using some appropriate iterative techniques, the formulation of the arbitrary length forward DCT (FDCT) and inverse DCT (IDCT) can be effectively implemented by recursive equations, and the hardware complexity is further reduced as compared to approaches in the literature. The proposed algorithm is suitable for both software and VLSI implementations.
Lap-Pui Chau, Wan-Chi Siu
ISCAS1
2000 Efficient recursive algorithm for the inverse discrete cosine transform
abstract
Recursive algorithms have been found very effective for realization using software and very large scale integrated circuit (VLSI) techniques. Previously, some recursive algorithms have been proposed for the realization of the inverse discrete cosine transform (IDCT). In this paper, an efficient recursive algorithm for the IDCT with arbitrary length is presented. By using some appropriate iterative techniques, the formulation of the IDCT can be implemented effectively using recursive equations, and the hardware complexity is further reduced as compared with the approaches in the literature.
Lap-Pui Chau, Wan-Chi Siu
IEEE Signal Process. Lett.1
1999 An Mpeg-4 Real-Time Video Decoder Software
abstract
This paper gives a description of our real-time software MPEG-4 video decoder. The performance of the decoder and the fast algorithms used for real-time decoding are discussed. Our techniques used to speed up four computationally intensive modules, inverse discrete cosine transform, context-based arithmetic decoding, padding, and variable length decoding, are also presented in this paper. Our decoding speed is improved so that software real-time video decoding is possible.
Lap-Pui Chau, Nam Ling, Gunnar Hovden, Hui Lan, Hon-Cheong Ng, Keng-Pang Lim
ICIP (1)1
1995 New 2n discrete cosine transform algorithm using recursive filter structure
abstract
The discrete cosine transform (DCT) is widely used in digital signal processing. It is always desirable to look for more efficient algorithms for the realization of the DCT. We generalize a formulation for converting a length-2/sup n/ DCT into n groups of equations, then apply a novel technique for its implementation. The sizes of the groups are 2/sup m/, for m=n-1,...,0. While their structures are extremely regular. The realization can then be converted into the simplest recursive filter form, which is of particularly simple for practical implementation. The filter structure is numerically stable, since it involves no division at all.
Wan-Chi Siu, Yuk-Hee Chan, Lap-Pui Chau
ICASSP3
1994 Efficient formulation for the realization of discrete cosine transform using recursive structure
abstract
Effective formulations for the conversion of the discrete Fourier transform (DFT) into a recursive structure are available and have been very effective for the realization using software, hardware and VLSI techniques. Little research work has been reported on the effective way to convert the discrete cosine transform (DCT) into a recursive form and the related realization. We propose a new method to convert a prime length DCT into a recursive structure. A trivial approach is to convert the DCT into the DFT and to apply Goertzel's (1958) algorithm for the rest of the realization. However, this method is inefficient. In our approach, we convert a prime length DCT into suitable transforms with half of the original length to effect fast realization. The number of operations is greatly reduced and the structure is extremely regular.>
Lap-Pui Chau, Wan-Chi Siu
ICASSP (3)1
1994 Efficient implementation of discrete cosine transform using recursive filter structure
abstract
We generalize a formulation for converting a length-2/sup n/ discrete cosine transform into n groups of equations, then apply a novel technique for its implementation. The sizes of the groups are 2/sup n-1/, 2/sup n-2/, ...2/sup 0/ respectively, while their structures are extremely regular. The realization can then be converted into recursive filter form, which is particularly simple for practical implementation.>
Yuk-Hee Chan, Lap-Pui Chau, Wan-Chi Siu
IEEE Trans. Circuits Syst. Video Technol.2