VLDB 2026 Research / reviewers in the wild / expert
Kin-Man Lam 0001
dblp:16/1994 · also Kenneth Kin-Man Lam
· DBLP profile ↗
241ranked-venue papers
5as first author
93since 2021 · last 2026
0000-0002-0422-8454ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 143 · 1 first-author · 52 since 2021Artificial intelligence and machine learning · 90 · 4 first-author · 32 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 1 first-author · 13 since 2021Systems, architecture and hardware · 6 · 1 first-authorSecurity and privacy · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SfM-free 3D Gaussian Splatting from extremely sparse view
Zongqi He, Hanmin Li, Kin-Chung Chan, Yushen Zuo, Zhe Xiao 0001, Jun Xiao 0010, Xiaoyang Bai, Kin-Man Lam 0001 |
Comput. Graph. | 10 |
| 2026 | NeRF-UAVeL: Unified attention-driven volumetric learning for robust NeRF-based 3D object detection
Hana Lebeta Goshu, Tadesse G. Wakjira, Meklit M. Atlaw, Kin-Chung Chan, Songjiang Lai, Kin-Man Lam 0001 |
Neurocomputing | 6 |
| 2026 | Vision-language model guided image restoration
Cuixin Yang, Rongkang Dong, Kin-Man Lam 0001 |
Image Vis. Comput. | 3 |
| 2026 | AFFusion: Atmospheric scattering enhancement and frequency integrated spatial-channel attention for infrared and visible image fusion
Jiwei Hu, Qiwen Jin, Kin-Man Lam 0001 |
Pattern Recognit. | 4 |
| 2026 | Degradation-Aware Prompt Learning With Cross-Modal Compensation for Adverse Weather RemovalabstractAdverse weather causes diverse and complex image degradations, severely compromising the reliability of computer vision systems. Existing all-in-one restoration models attempt to address multiple degradation types within a unified framework, but often lack explicit spatial and semantic modeling of degradation characteristics, limiting their adaptability to diverse weather conditions. To address this limitation, we propose a Degradation-Aware Cross-Modal Prompt Compensation Network (DCMPC-Net) that leverages cross-modal degradation cues from a pre-trained vision-language model to condition restoration features within a unified backbone. Specifically, our DCMPC-Net mainly consists of the Cross-Modal Prompt Generator (CMPG), Prompt-Guided Attention Alignment Module (PGAAM), and Dual Feature Compensation Module (DFCM). The CMPG integrates textual embeddings with visual features to produce degradation-aware prompts that encode degradation-related semantic and contextual cues. These prompts are injected into the decoder via a PGAAM, which adaptively aligns semantic information with degraded regions to facilitate context-aware restoration. To further enhance structural fidelity, DFCM is introduced that disentangles degradation artifacts from scene structures, thereby improving the reconstruction of fine textures and detailed content. By integrating cross-modal semantic guidance with spatial alignment and structural enhancement, DCMPC-Net achieves robust and perceptually consistent restoration across diverse weather conditions. Extensive experiments show that DCMPC-Net outperforms state-of-the-art methods in both task-specific and unified settings, achieving superior accuracy and visual fidelity. The code is available at https://github.com/fanamber831/DCMPC-Net. Wanshu Fan, Yunzhe Zhang, Jing Qin 0007, Kin-Man Lam 0001, Cong Wang 0018, Jinshan Pan |
IEEE Trans. Image Process. | 6 |
| 2026 | An Episode Memory-Guided Dual-Stage Framework for Long-Form Video Temporal GroundingabstractVideo temporal grounding (VTG) aims to localize video moments that are semantically related to a given natural language query. In spite of recent progress in short-form videos, research on VTG in long-form videos (e.g., hours long) remains highly demanded yet underexplored. Existing methods predominantly adopt sliding window-based or multi-scale anchor-based strategies to generate temporal proposals, which require time-consuming post-processing or are independent of video content, thereby limiting their performance and efficiency. To address this dilemma, in this paper, we propose an episode memory-prompted (EMP) two-stage framework for temporal grounding in long-form videos. Specifically, the first stage generates a set of dynamic episode memories, which explicitly summarize various activities occurring throughout the lengthy video. An unsupervised memory learning paradigm is formulated by imposing discriminability and diversity constraints, eliminating the reliance on additional activity-instance annotations. Then, in the second stage, based on the supplement of frame-level detailed content and the guidance of a language query, the augmented memory prompts function as anchors for efficiently regressing the refined boundaries of the target video moment. Extensive experimental results on two public long-form video data sets, i.e., MAD and Ego4d, validate that the proposed EMP framework saves more than 8.5% trainable parameters and 13.9% FLOPs, while still achieving comparable performance with existing methods. Tianshan Liu, Bing-Kun Bao, Kin-Man Lam 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | See In Detail: Enhancing Sparse-view 3D Gaussian Splatting with Local Depth and Semantic Regularizationabstract3D Gaussian Splatting (3DGS) has shown remarkable performance in novel view synthesis. However, its rendering quality deteriorates with sparse inphut views, leading to distorted content and reduced details. This limitation hinders its practical application. To address this issue, we propose a sparse-view 3DGS method. Given the inherently ill-posed nature of sparse-view rendering, incorporating prior information is crucial. We propose a semantic regularization technique, using features extracted from the pretrained DINO-ViT model, to ensure multi-view semantic consistency. Additionally, we propose local depth regularization, which constrains depth values to improve generalization on unseen views. Our method outperforms state-of-the-art novel view synthesis approaches, achieving up to 0.4dB improvement in terms of PSNR on the LLFF dataset, with reduced distortion and enhanced visual quality. Zongqi He, Zhe Xiao 0001, Kin-Chung Chan, Yushen Zuo, Jun Xiao 0010, Kin-Man Lam 0001 |
ICASSP | 6 |
| 2025 | Safeguarding Vision-Language Models: Mitigating Vulnerabilities to Gaussian Noise in Perturbation-Based AttacksabstractVision-Language Models (VLMs) extend the capabilities of Large Language Models (LLMs) by incorporating visual information, yet they remain vulnerable to jailbreak attacks, especially when processing noisy or corrupted images. Although existing VLMs adopt security measures during training to mitigate such attacks, vulnerabilities associated with noise-augmented visual inputs are overlooked. In this work, we identify that missing noise-augmented training causes critical security gaps: many VLMs are susceptible to even simple perturbations such as Gaussian noise. To address this challenge, we propose Robust-VLGuard, a multimodal safety dataset with aligned / misaligned image-text pairs, combined with noise-augmented fine-tuning that reduces attack success rates while preserving functionality of VLM. For stronger optimization-based visual perturbation attacks, we propose DiffPure-VLM, leveraging diffusion models to convert adversarial perturbations into Gaussian-like noise, which can be defended by VLMs with noise-augmented safety fine-tuning. Experimental results demonstrate that the distribution-shifting property of diffusion model aligns well with our fine-tuned VLMs, significantly mitigating adversarial perturbations across varying intensities. The dataset and code are available at https://github.com/JarvisUSTC/DiffPure-RobustVLM. Yushen Zuo, Yuanjun Chai, Yicheng Fu, Yichun Feng, Kin-Man Lam 0001 |
ICCV | 7 |
| 2025 | NDFormer: A Mixed-Scale Transformer with Enhanced Nonlinearity for Nighttime Image Deraining
Zhirui Liu, Shangquan Sun, Yuning Cui 0001, Dehong Kong, Wenqi Ren, Kin-Man Lam 0001 |
PRCV (9) | 7 |
| 2025 | Multi-kernel feature extraction with dynamic fusion and downsampled residual feature embedding for predicting rice RNA N6-methyladenine sitesabstractRNA N$^{6}$-methyladenosine (m$^{6}$A) is a critical epigenetic modification closely related to rice growth, development, and stress response. m$^{6}$A accurate identification, directly related to precision rice breeding and improvement, is fundamental to revealing phenotype regulatory and molecular mechanisms. Faced on rice m$^{6}$A variable-length sequence, to input into the model, the maximum length padding and label encoding usually adapt to obtain the max-length padded sequence for prediction. Although this can retain complete sequence information, resulting in sparse information and invalid padding, reducing feature extraction accuracy. Simultaneously, existing rice-specific m$^{6}$A prediction methods are still at an early stage. To address these issues, we develop a new end-to-end deep learning framework, MFDm$^{6}$ARice, for predicting rice m$^{6}$A sites. In particular, to alleviate sparseness, we construct a multi-kernel feature fusion module to mine essential information in max-length padded sequences by multi-kernel feature extraction function and effectively transfer information through global-local dynamic fusion function. Concurrently, considering the complexity and computational efficiency of high-dimensional features caused by invalid padding, we design a downsampling residual feature embedding module to optimize feature space compression and achieve accurate feature expression and efficient computational performance. Experiments show that MFDm$^{6}$ARice outperforms comparison methods in cross-validation, same- and cross-species independent test sets, demonstrating good robustness and generalization. The application on maize m$^{6}$A indicates the MFDm$^{6}$ARice's scalability. Further investigations have shown that combining different kernel features, focusing on global channel-local spatial, and employing reasonable downsampling and residual connections can improve feature representation and extraction, ensure effective information transfer, and significantly enhance model performance. Zhigang Zeng, Kin-Man Lam 0001 |
Briefings Bioinform. | 4 |
| 2025 | A Memory-Assisted Knowledge Transferring Framework with Curriculum Anticipation for Weakly Supervised Online Activity Detection
Tianshan Liu, Kin-Man Lam 0001, Bing-Kun Bao |
Int. J. Comput. Vis. | 2 |
| 2025 | Underwater Camera: Improving Visual Perception Via Adaptive Dark Pixel Prior and Color Correction
Jingchun Zhou, Qiuping Jiang, Wenqi Ren, Kin-Man Lam 0001, Weishi Zhang |
Int. J. Comput. Vis. | 5 |
| 2025 | A dual-domain mutual compensation network for multi-modality image fusion
Jiwei Hu, Ping Lou, Kin-Man Lam 0001, Qiwen Jin |
Neurocomputing | 4 |
| 2025 | Revisiting One-Stage Deep Uncalibrated Photometric Stereo via Fourier EmbeddingabstractThis paper introduces a one-stage deep uncalibrated photometric stereo (UPS) network, namely Fourier Uncalibrated Photometric Stereo Network (FUPS-Net), for non-Lambertian objects under unknown light directions. It departs from traditional two-stage methods that first explicitly learn lighting information and then estimate surface normals. Two-stage methods were deployed because the interplay of lighting with shading cues presents challenges for directly estimating surface normals without explicit lighting information. However, these two-stage networks are disjointed and separately trained so that the error in explicit light calibration will propagate to the second stage and cannot be eliminated. In contrast, the proposed FUPS-Net utilizes an embedded Fourier transform network to implicitly learn lighting features by decomposing inputs, rather than employing a disjointed light estimation network. Our approach is motivated from observations in the Fourier domain of photometric stereo images: lighting information is mainly encoded in amplitudes, while geometry information is mainly associated with phases. Leveraging this property, our method "decomposes" geometry and lighting in the Fourier domain as guidance, via the proposed Fourier Embedding Extraction (FEE) block and Fourier Embedding Aggregation (FEA) block, which generate lighting and geometry features for the FUPS-Net to implicitly resolve the geometry-lighting ambiguity. Furthermore, we propose a Frequency-Spatial Weighted (FSW) block that assigns weights to combine features extracted from the frequency domain and those from the spatial domain for enhancing surface reconstructions. FUPS-Net overcomes the limitations of two-stage UPS methods, offering better training stability, a concise end-to-end structure, and avoiding accumulated errors in disjointed networks. Experimental results on synthetic and real datasets demonstrate the superior performance of our approach, and its simpler training setup, potentially paving the way for a new strategy in deep learning-based UPS methods. Yakun Ju, Boxin Shi, Bihan Wen, Kin-Man Lam 0001, Xudong Jiang 0001, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Long Short-Term Fusion by Multi-Scale Distillation for Screen Content Video Quality EnhancementabstractDifferent from natural videos, where artifacts distributed evenly, the artifacts of compressed screen content videos mainly occur in the edge areas. Besides, these videos often exhibit abrupt scene switches, resulting in noticeable distortions in video reconstruction. Existing multiple-frame models using a fixed range of neighbor frames face challenges in effectively enhancing frames during scene switches and lack efficiency in reconstructing high-frequency details. To address these limitations, we propose a novel method that effectively handles scene switches and reconstructs high-frequency information. In the feature extraction part, we develop long-term and short-term feature extraction streams, in which the long-term feature extraction stream learns the contextual information, and the short-term feature extraction stream extracts more related information from shorter input to assist the long-term stream to handle fast motion and scene switches. To further enhance the frame quality during scene switches, we incorporate a similarity-based neighbor frame selector before feeding frames into the short-term stream. This selector identifies relevant neighbor frames, aiding in the efficient handling of scene switches. To dynamically fuse the short-term feature and long-term features, the muti-scale feature distillation focuses on adaptively recalibrating channel-wise feature responses to achieve effective feature distillation. In the reconstruction part, a high-frequency reconstruction block is proposed for guiding the model to restore the high-frequency components. Experimental results demonstrate the significant advancements achieved by our proposed Long Short-term Fusion by Multi-Scale Distillation (LSFMD) method in enhancing the quality of compressed screen content videos, surpassing the current state-of-the-art methods. Ziyin Huang, Yui-Lam Chan, Ngai-Wing Kwong, Sik-Ho Tsang, Kin-Man Lam 0001, Bingo Wing-Kuen Ling |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Geometric Distortion Guided Transformer for Omnidirectional Image Super-ResolutionabstractAs virtual and augmented reality applications gain popularity, omnidirectional image (ODI) super-resolution has become increasingly important. Unlike 2D plain images that are formed on a plane, ODIs are projected onto spherical surfaces. Applying established image super-resolution methods to ODIs, therefore, requires performing equirectangular projection (ERP) to map the ODIs onto a plane. ODI super-resolution needs to take into account geometric distortion resulting from ERP. However, without considering such geometric distortion of ERP images, previous methods only utilize a limited range of pixels and may easily miss self-similar textures for reconstruction. In this paper, we introduce a novel Geometric Distortion Guided Transformer for Omnidirectional image Super-Resolution (GDGT-OSR). Specifically, a distortion modulated rectangle-window selfattention mechanism, integrated with deformable self-attention, is proposed to better perceive the distortion and thus involve more self-similar textures. Distortion modulation is achieved through a newly devised distortion guidance generator that produces guidance for the rectangular windows by exploiting the variability of distortion across latitudes. Furthermore, we propose a dynamic feature aggregation scheme to adaptively fuse the features from different self-attention modules. We present extensive experimental results on public datasets and show that the new GDGT-OSR outperforms methods in existing literature. Cuixin Yang, Rongkang Dong, Jun Xiao 0010, Kin-Man Lam 0001, Fei Zhou 0001, Guoping Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Multi-Frame Spatiotemporal Feature and Hierarchical Learning Approach for No-Reference Screen Content Video Quality AssessmentabstractThe rapid adoption of remote work, online conferencing, and shared-screen collaboration has significantly increased the usage of screen content videos (SCVs), creating a growing need for reliable quality assessment to maintain excellent quality of service. While several full-reference SCV quality assessment (SCVQA) methods have been proposed, their practical application is often limited by the unavailability of reference videos. Existing no-reference SCVQA (NR-SCVQA) methods rely on handcrafted features and focus solely on specific distortions and features, potentially limiting their generalization ability. Moreover, they fail to explore the underlying spatiotemporal information of SCVs, which could hinder their performance. In this work, we propose a novel deep learning-based NR-SCVQA model specifically tailored to capture the comprehensive spatiotemporal features of SCVs to overcome these issues and challenges posed by the SCVQA task. Our approach incorporates a dual-channel spatiotemporal convolutional neural network (DCST-CNN) module to extract both content-aware and edge-aware spatiotemporal quality features, which enables an effective spatiotemporal quality feature representation learning for the downstream SCVQA task. Building upon the DCST-CNN, we further propose a Temporal Pyramid Transformer (TPT) module to fuse spatiotemporal features across multiple temporal scales, enabling the model to capture both short-term and long-term temporal dependencies within an SCV for hierarchical learning. The proposed DCST-CNN and TPT modules work together to provide a robust and accurate NR-SCVQA framework. We conduct experiments on SCVQA databases to validate the effectiveness of our model, which outperforms existing state-of-the-art NR-SCVQA method. The results demonstrate the strength and applicability of our approach in real-world SCVQA tasks. Ngai-Wing Kwong, Yui-Lam Chan, Sik-Ho Tsang, Ziyin Huang, Kin-Man Lam 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | AMSP-UOD: When Vortex Convolution and Stochastic Perturbation Meet Underwater Object DetectionabstractIn this paper, we present a novel Amplitude-Modulated Stochastic Perturbation and Vortex Convolutional Network, AMSP-UOD, designed for underwater object detection. AMSP-UOD specifically addresses the impact of non-ideal imaging factors on detection accuracy in complex underwater environments. To mitigate the influence of noise on object detection performance, we propose AMSP Vortex Convolution (AMSP-VConv) to disrupt the noise distribution, enhance feature extraction capabilities, effectively reduce parameters, and improve network robustness. We design the Feature Association Decoupling Cross Stage Partial (FAD-CSP) module, which strengthens the association of long and short range features, improving the network performance in complex underwater environments. Additionally, our sophisticated post-processing method, based on non-maximum suppression with aspect-ratio similarity thresholds, optimizes detection in dense scenes, such as waterweed and schools of fish, improving object detection accuracy. Extensive experiments on the URPC and RUOD datasets demonstrate that our method outperforms existing state-of-the-art methods in terms of accuracy and noise immunity. AMSP-UOD proposes an innovative solution with the potential for real-world applications. Our code is available at https://github.com/zhoujingchun03/AMSP-UOD. Jingchun Zhou, Zongxin He, Kin-Man Lam 0001, Yudong Wang 0002, Weishi Zhang, Chunle Guo, Chongyi Li |
AAAI | 3 |
| 2024 | Towards Progressive Multi-Frequency Representation for Image WarpingabstractImage warping, a classic task in computer vision, aims to use geometric transformations to change the appearance of images. Recent methods learn the resampling kernels for warping through neural networks to estimate missing values in irregular grids, which, however, fail to capture local variations in deformed content and produce images with distortion and less high-frequency details. To address this issue, this paper proposes an effective method, namely MFR, to learn Multi-Frequency Representations from in-put images for image warping. Specifically, we propose a progressive filtering network to learn image representations from different frequency subbands and generate deformable images in a coarse-to-fine manner. Furthermore, we employ learnable Gabor wavelet filters to improve the model's capability to learn local spatial-frequency representations. Comprehensive experiments, including homography trans-formation, equirectangular to perspective projection, and asymmetric image super-resolution, demonstrate that the proposed MFR significantly outperforms state-of-the-art image warping methods. Our method also showcases superior generalization to out-of-distribution domains, where the generated images are equipped with rich details and less distortion, thereby high visual quality. The source code is available at https://github.com/junxiao01/MFR. Jun Xiao 0010, Zihang Lyu, Yakun Ju, Changjian Shui, Kin-Man Lam 0001 |
CVPR | 6 |
| 2024 | Learning Equilibrium Transformation for Gamut Expansion and Color Restoration
Jun Xiao 0010, Changjian Shui, Kin-Man Lam 0001 |
ECCV (71) | 5 |
| 2024 | Hierarchical Vertex-Wise Intensification Graph Convolution for Skeleton-Based Activity RecognitionabstractGraph convolutional networks (GCNs), which can effectively captures the spatial and temporal relationships between skeleton joints through graph topology, have shown promising performances in skeleton-based activity recognition in recent years. These methods typically learn the semantic features of the vertices of a skeleton and the associated adjacency matrix. However, how to efficiently establish relationships between vertices still remains a substantial problem. To solve this problem, we propose a novel Hierarchical Vertex-wise Intensification Graph Convolution Network (HVI-GCN) for skeleton-based action recognition. The proposed module dilates input features into higher dimensions to broaden the temporal horizon, and builds a vertex-wise topology based on self-adaptively learned attention. With the adjacency matrix, features from other positions can be collected to aid the prediction of the current position. The proposed module provides a better receptive field and semantic understanding of both the spatial and temporal domains than related methods. Experiments were mainly conducted on the at NTU-RGB-D, NTU-GRB-D 120, and NW-UCLA datasets with joint and bone integrated with motion sequences. Experimental results show that HVI-GCN can improve accuracy by up to 1.1% on the RGB-D 120 dataset. Meanwhile, the accuracy on RGB-D 60 dataset and NW-UCLA dataset can be boosted by 1.4% and 1.2%, respectively. Jun Xiao 0010, Tianshan Liu, Kin-Man Lam 0001 |
ICIP | 6 |
| 2024 | AI-Generated Image Detection With Wasserstein Distance Compression and Dynamic AggregationabstractWith the rapid advancement of generative models, image detectors for AI-generated content have become an increasingly necessary technology in computer vision, attracting significant attention from researchers. This technology aims to detect whether an image is naturally generated by imaging systems (e.g., digital cameras) or generated by advanced AI techniques. Despite the promising performance achieved by recent fake detection methods, they are typically trained on millions of redundant images with similar characteristics, leading to inefficient training. Furthermore, the performances of existing detectors often deteriorate when the training datasets are imbalanced. To address these challenges, we propose a novel AI-generated image detector based on dynamic aggregation and information compression with the Wasserstein distance. Experimental results show that our proposed method significantly outperforms state-of-the-art models that generalize across different generative models, with an increase of $\mathbf{+ 1. 8 6 \%}$ average accuracy and $\mathbf{+ 0. 1 4 \%}$ average precision, while substantially reducing the training time. On imbalanced datasets, our proposed method leads to a $\mathbf{+ 1 4. 4 6 \%}$ accuracy improvement, clearly demonstrating its robustness on imbalanced datasets. Zihang Lyu, Jun Xiao 0010, Kin-Man Lam 0001 |
ICIP | 4 |
| 2024 | Point Cloud Densification for 3D Gaussian Splatting from Sparse Input ViewsabstractThe technique of 3D Gaussian splatting (3DGS) has demonstrated its effectiveness and efficiency in rendering photo-realistic images for novel view synthesis. However, 3DGS requires a high density of camera coverage, and its performance inevitably degrades with sparse training views, which significantly limits its applicability in real-world scenarios. In recent years, many researchers have explored the use of depth information to alleviate this problem, but the performance of their methods is sensitive to the accuracy of depth estimation. To this end, we propose an efficient method to enhance the performance of 3DGS with sparse training views. Specifically, instead of applying depth maps for regularization, we propose a densification method that generates high-quality point clouds, providing a superior initialization for 3D Gaussians. Furthermore, we propose Systematically Angle of View Sampling (SAOVS), which employs Spherical Linear Interpolation (SLERP) and linear interpolation for side view sampling, to determine unseen views outside the training data for semantic pseudo-label regularization. Experiments show that our proposed method significantly outperforms other leading 3D rendering models on the ScanNet dataset and the LLFF dataset. In particular, compared with the conventional 3DGS method, our proposed method achieves performance gains of up to 1.71dB in PSNR and 0.07 in SSIM. In addition, the novel view synthesis produced by our method demonstrates the highest visual quality with minimal distortions. Kin-Chung Chan, Jun Xiao 0010, Hana Lebeta Goshu, Kin-Man Lam 0001 |
ACM Multimedia | 4 |
| 2024 | Label Text-aided Hierarchical Semantics Mining for Panoramic Activity RecognitionabstractPanoramic activity recognition is a comprehensive yet challenging task in crowd scene understanding, which aims to concurrently identify multi-grained human behaviors, including individual actions, social group activities, and global activities. Previous studies tend to capture cross-granularity activity-semantics relations from solely the video input, thus ignoring the intrinsic semantic hierarchy in label-text space. To this end, we propose a label text-aided hierarchical semantics mining (THSM) framework, which explores multi-level cross-modal associations by learning hierarchical semantic alignment between visual content and label texts. Specifically, a hierarchical encoder is first constructed to encode the visual and text inputs into semantics-aligned representations at different granularities. To fully exploit the cross-modal semantic correspondence learned by the encoder, a hierarchical decoder is further developed, which progressively integrates the lower-level representations with the higher-level contextual knowledge for coarse-to-fine action/activity recognition. Extensive experimental results on the public JRDB-PAR benchmark validate the superiority of the proposed THSM framework over state-of-the-art methods. Tianshan Liu, Kin-Man Lam 0001, Bing-Kun Bao |
ACM Multimedia | 2 |
| 2024 | Frame Similarity-Based Screen Content Video Quality Enhancement via Adaptive Long Short-Term FusionabstractCompressed screen content videos often exhibit artifacts in edge areas and suffer from distortions during scene switches, where content abruptly changes between frames. Existing multi-frame models, which use a fixed range of neighbor frames, struggle with these switches. To address this, we propose a novel method that effectively handles scene switches. Our approach utilizes Long-term Feature Extraction (LFE) to capture contextual information, while the Frame Similarity-based Short-term Feature Extraction (FSFE) focuses on texture information to manage fast motion and scene switches. In FSFE, a Similarity-based Neighbor Frame Selector (SNFS) is designed to choose relevant neighbor frames for the short-term stream, enhancing the quality of scene switch frames. To fuse short-term and long-term features adaptively, we introduce a local-spatial and global-channel attention module, which recalibrates spatial and channel-wise feature responses. Experimental results show that our Frame Similarity-Based via Adaptive Long Short-Term Fusion (FSLST) method significantly improves the quality of compressed videos, outperforming current state-of-the-art methods. Ziyin Huang, Yui-Lam Chan, Ngai-Wing Kwong, Sik-Ho Tsang, Kin-Man Lam 0001, Bingo Wing-Kuen Ling |
VCIP | 5 |
| 2024 | AFSRNet: learning local descriptors with adaptive multi-scale feature fusion and symmetric regularization
Dong Li 0028, Haowen Liang, Kin-Man Lam 0001 |
Appl. Intell. | 3 |
| 2024 | Contrastive decoupling global and local features for pavement crack detection
Ching-Chi Yeung, Kin-Man Lam 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | HCLR-Net: Hybrid Contrastive Learning Regularization with Locally Randomized Perturbation for Underwater Image EnhancementabstractUnderwater image enhancement presents a significant challenge due to the complex and diverse underwater environments that result in severe degradation phenomena such as light absorption, scattering, and color distortion. More importantly, obtaining paired training data for these scenarios is a challenging task, which further hinders the generalization performance of enhancement models. To address these issues, we propose a novel approach, the Hybrid Contrastive Learning Regularization (HCLR-Net). Our method is built upon a distinctive hybrid contrastive learning regularization strategy that incorporates a unique methodology for constructing negative samples. This approach enables the network to develop a more robust sample distribution. Notably, we utilize non-paired data for both positive and negative samples, with negative samples are innovatively reconstructed using local patch perturbations. This strategy overcomes the constraints of relying solely on paired data, boosting the model’s potential for generalization. The HCLR-Net also incorporates an Adaptive Hybrid Attention module and a Detail Repair Branch for effective feature extraction and texture detail restoration, respectively. Comprehensive experiments demonstrate the superiority of our method, which shows substantial improvements over several state-of-the-art methods in terms of quantitative metrics, significantly enhances the visual quality of underwater images, establishing its innovative and practical applicability. Our code is available at: https://github.com/zhoujingchun03/HCLR-Net . Jingchun Zhou, Chongyi Li, Qiuping Jiang, Man Zhou 0003, Kin-Man Lam 0001, Weishi Zhang, Xianping Fu |
Int. J. Comput. Vis. | 6 |
| 2024 | Correction: HCLR-Net: Hybrid Contrastive Learning Regularization with Locally Randomized Perturbation for Underwater Image Enhancement
Jingchun Zhou, Chongyi Li, Qiuping Jiang, Man Zhou 0003, Kin-Man Lam 0001, Weishi Zhang, Xianping Fu |
Int. J. Comput. Vis. | 6 |
| 2024 | Deep progressive feature aggregation network for multi-frame high dynamic range imaging
Jun Xiao 0010, Tianshan Liu, Kin-Man Lam 0001 |
Neurocomputing | 5 |
| 2024 | Spatio-temporal feature learning for enhancing video quality based on screen content characteristics
Ziyin Huang, Yui-Lam Chan, Sik-Ho Tsang, Ngai-Wing Kwong, Kin-Man Lam 0001, Bingo Wing-Kuen Ling |
J. Vis. Commun. Image Represent. | 5 |
| 2024 | Spatiotemporal feature learning for no-reference gaming content video quality assessment
Ngai-Wing Kwong, Yui-Lam Chan, Sik-Ho Tsang, Ziyin Huang, Kin-Man Lam 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2024 | Integrally Mixing Pyramid Representations for Anchor-Free Object Detection in Aerial ImageryabstractAnchor-free object detectors have recently received increasing research attention in the field of aerial scene object detection, due to their high flexibility and practicality. Anchor-free detectors typically depend on the feature pyramid network (FPN) to alleviate the challenge of significant variations in object scales in aerial contexts. Despite establishing a multi-scale feature pyramid, existing FPN-based methods treat each aerial object as an indivisible entity solely managed by a single-scale representation. However, they fail to take into account the distinct characteristics of various components within an instance. To this end, this letter proposes a novel anchor-free detector, namely IMPR-Det, which can integrally mix multi-scale pyramid representations for different components of an instance, thus boosting the fine-grained object representation capability. Specifically, IMPR-Det fundamentally introduces a more advanced detection head with an adaptive routing mechanism for pixel-level multi-scale feature assignment, instead of previous instance-level assignment. Experimental results demonstrate the superiority of the proposed method over its counterparts, in terms of both accuracy and efficiency, for object detection in aerial images. Jun Xiao 0010, Cuixin Yang, Jingchun Zhou, Kin-Man Lam 0001, Qi Wang 0009 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | Deep Learning Methods for Calibrated Photometric Stereo and BeyondabstractPhotometric stereo recovers the surface normals of an object from multiple images with varying shading cues, i.e., modeling the relationship between surface orientation and intensity at each pixel. Photometric stereo prevails in superior per-pixel resolution and fine reconstruction details. However, it is a complicated problem because of the non-linear relationship caused by non-Lambertian surface reflectance. Recently, various deep learning methods have shown a powerful ability in the context of photometric stereo against non-Lambertian surfaces. This paper provides a comprehensive review of existing deep learning-based calibrated photometric stereo methods utilizing orthographic cameras and directional light sources. We first analyze these methods from different perspectives, including input processing, supervision, and network architecture. We summarize the performance of deep learning photometric stereo models on the most widely-used benchmark data set. This demonstrates the advanced performance of deep learning-based photometric stereo methods. Finally, we give suggestions and propose future research trends based on the limitations of existing models. Yakun Ju, Kin-Man Lam 0001, Wuyuan Xie, Huiyu Zhou 0001, Junyu Dong, Boxin Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Deep multi-scale feature mixture model for image super-resolution with multiple-focal-length degradation
Jun Xiao 0010, Rui Zhao 0012, Kin-Man Lam 0001, Kao Wan |
Signal Process. Image Commun. | 4 |
| 2024 | Bi-Center Loss for Compound Facial Expression RecognitionabstractCompound facial expressions involve combinations of basic emotions, posing challenges to automatic facial expression recognition research. The focus of existing studies in facial expression recognition remains primarily on classifying basic or single expressions, which limits its application to compound facial expression recognition. Moreover, some compound facial expression recognition methods rely on laboratory-controlled expression data, narrowing their generalizability to real-world scenarios. In this work, we carry out the task of compound facial expression recognition in unconstrained environments. To this end, we devise a new loss function, specifically for compound facial expression recognition, called bi-center loss, which is built upon center loss. Unlike center loss that considers all categories, bi-center loss enables deep neural networks to learn compound emotion features by leveraging basic emotion centers. Additionally, we introduce a basic-center regularization term, based on the variance among the basic centers, to ensure appropriate discriminative capabilities of the learned features. Experiments conducted on unconstrained compound expression datasets demonstrate the effectiveness of the proposed method over the baselines, achieving state-of-the-art performance in compound facial expression recognition. Rongkang Dong, Kin-Man Lam 0001 |
IEEE Signal Process. Lett. | 2 |
| 2024 | Flow-Edge-Net: Video Saliency Detection Based on Optical Flow and Edge-Weighted Balance LossabstractOptical flow networks have been widely utilized for video saliency detection (VSD) due to their effective performance in capturing the motion of objects. However, the use of optical flow blurs the edges of salient objects and leads to the problems of poorly defined object boundaries. To address this issue, we propose an optical flow-based edge-weighted loss function, to train a network called Flow-Edge-Net, which can balance the weights of the foreground and background information at the edges of video frames. It has achieved superior performance in detecting salient boundaries. Specifically, we propose two complementary encoding and decoding networks based on the concept of decoupling. That is, the optical flow network focuses on moving objects, while the edge network, based on the encoder-decoder structure, focuses on edge information. As the two networks output features of the same dimension and are from the same input, our proposed self-designed adaptive weighted feature fusion module can compare and integrate the edge information and location information from the two networks through adaptive weighting. The proposed method has been evaluated on five widely used databases. Experiment results demonstrate the superior performance of the proposed Flow-Edge-Net in locating salient objects, with accurate and refined edges. The proposed method achieves superior performance over the state-of-the-art methods in detecting salient objects in videos. Muwei Jian, Xiangwei Lu, Yakun Ju, Hui Yu 0001, Kin-Man Lam 0001 |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2024 | UniFRD: A Unified Method for Facial Image Restoration Based on Diffusion Probabilistic ModelabstractThis paper presents a Unified Facial image and video Restoration method based on the Diffusion probabilistic model (UniFRD), designed to effectively address both single- and multi-type image degradation. The noise predictor in UniFRD consists of a ViT-based encoder and a novel Separation Fusion Decoding Module (SFDM). The flexible feature optimization strategy allows for decoding complex conditional noise without being limited by degradation patterns. Specifically, SFDM adjusts and refines the channel correlation and expressive power of high-dimensional features step by step, enabling the network to more accurately perceive and enhance the interaction between posterior probabilities and conditional inputs. This process is crucial for improving the visual quality and stability of the restoration results. Extensive experiments demonstrate that even when facial images suffer from both pixel-level and image-level degradation, UniFRD can still guarantee the restoration of rich details and maintain attribute consistency. In summary, compared to existing methods, the solution proposed in this study for facial restoration tasks offers greater generality and adaptability. Moreover, it has high practical value for applications involving faces in complex and unconstrained outdoor scenarios. Muwei Jian, Rui Wang 0199, Feng Xu 0005, Hui Yu 0001, Kin-Man Lam 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Estimating High-Resolution Surface Normals via Low-Resolution Photometric Stereo ImagesabstractAcquiring high-resolution 3D surface structures is a crucial task in computer vision as it provides more detailed surface textures and clearer structures. Photometric stereo can measure per-pixel surface normals of a 3D object using various shading cues. However, obtaining high-resolution images in a linear response photometric stereo imaging system can be challenging. Additionally, photometric stereo, as a per-pixel reconstruction method, requires higher-resolution surface normal maps to accurately depict complex surface structures, particularly in regions that demand more attention and precise reconstruction. Therefore, measuring high-resolution surface normals via low-resolution photometric stereo images is of great importance. Motivated by these, we propose a Super-resolution Photometric Stereo Network, namely SR-PSN. In order to address the issues of measuring the high-resolution surface normals from low-resolution photometric images, we mainly (1) apply a dual-position threshold normalization pre-processing scheme to effectively handle the spatially-varying reflectance of non-Lambertian surfaces, (2) adopt a local affinity feature module to learn the rich structural representation by explicitly revealing the neighbor relationships, (3) employ a parallel multi-scale feature extractor, which preserves high-resolution representations and deep feature extraction, and (4) propose a shared-weight regressor to handle the multi-scale features, to prevent the model collapsing into learning non-important features related to a certain fixed scale. Extensive ablation experiments validate the effectiveness of our proposed modules. Furthermore, quantitative experiments conducted on public benchmarks demonstrate that SR-PSN outperforms state-of-the-art calibrated photometric stereo methods. Notably, SR-PSN achieves superior results while utilizing photometric stereo images with only half the resolution of other methods. It effectively restores the structure of complex surfaces, producing a high-resolution normal map. Yakun Ju, Muwei Jian, Cong Wang 0018, Junyu Dong, Kin-Man Lam 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Attention-ConvNet Network for Ocean-Front Prediction via Remote Sensing SST ImagesabstractOcean front is one typical geophysical phenomenon acting as oases in the ocean for fishes and marine mammals. Accurate ocean-front prediction is critical for fishery and navigation safety. However, the formation and evolution of ocean fronts are inherently nonlinear and are influenced by various factors such as ocean currents, wind fields, and temperature changes, making ocean-front prediction a considerable challenge. This study proposes a temporal-sensitive network named Attention-ConvNet to address this challenge. Ocean fronts exhibit significant multiscale characteristics, requiring analysis and prediction across various temporal and spatial scales. The proposed network designs a hierarchical attention mechanism (HAM) that efficiently prioritizes relevant spatial and temporal information to meet the specific requirement. What is more, the proposed network uses a complex hierarchical branching convolutional network (HBCNet) architecture, which allows our network to leverage the complementary strengths of spatial and temporal information, effectively capturing the dynamic and complex variations in ocean fronts. In general, the network prioritizes and focuses on the most relevant information of front dynamics, which ensures its ability to effectively predict the ocean front. External experiments demonstrate that our network significantly outperforms conventional methods, confirming its capability for precise ocean-front prediction. The codes will be publicly available athttps://github.com/yuhudeyue/Ocean-Front-Prediction-Model. Yuting Yang 0001, Xin Sun 0003, Junyu Dong, Kin-Man Lam 0001, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Structured Adversarial Self-Supervised Learning for Robust Object Detection in Remote Sensing ImagesabstractObject detection plays a crucial role in scene understanding and has extensive practical applications. In the field of remote sensing object detection, both detection accuracy and robustness are of significant concern. Existing methods heavily rely on sophisticated adversarial training strategies that tend to improve robustness at the expense of accuracy. However, detection robustness is not always indicative of improved accuracy. Therefore, in this paper, we research how to enhance robustness, while still preserving high accuracy, or even improve both simultaneously, with simple vanilla adversarial training or even in the absence thereof. In pursuit of a solution, we first conduct an exploratory investigation by shifting our attention from adversarial training, referred to as adversarial fine-tuning, to adversarial pretraining. Specifically, we propose a novel pretraining paradigm, namely structured adversarial self-supervised (SASS) pretraining, to strengthen both clean accuracy and adversarial robustness for object detection in remote sensing images. At a high level, SASS pretraining aims to unify adversarial learning and self-supervised learning into pretraining and encode structured knowledge into pretrained representations for powerful transferability to downstream detection. Moreover, to fully explore the inherent robustness of vision Transformers and facilitate their pretraining efficiency, by leveraging the recent masked image modeling (MIM) as the pretext task, we further instantiate SASS pretraining into a concise end-to-end framework, named structured adversarial MIM (SA-MIM). SA-MIM consists of two pivotal components, structured adversarial attack and structured MIM (S-MIM). The former establishes structured adversaries for the context of adversarial pretraining, while the latter introduces a structured local-sampling global-masking strategy to adapt to hierarchical encoder architectures. Comprehensive experiments on three different datasets have demonstrated the significant superiority of the proposed pretraining paradigm over previous counterparts for remote sensing object detection. More importantly, regardless of with or without adversarial fine-tuning, it enables simultaneous improvements on detection accuracy and robustness as expected, promisingly alleviating the dependence on complicated adversarial fine-tuning. Kin-Man Lam 0001, Tianshan Liu, Yui-Lam Chan, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | IACC: Cross-Illumination Awareness and Color Correction for Underwater Images Under Mixed Natural and Artificial LightingabstractEnhancing underwater images captured under mixed artificial and natural lighting conditions presents two critical challenges. Existing methods lack a unified luminance feature extraction paradigm for mixed lighting scenes, leading to imbalance in luminance features, and consequent local overexposure or underexposure. Additionally, some color correction methods, through the fusion of features across multiple color spaces neglect the information loss due to the absence of feature alignment in cross-space fusion. To address these challenges, we propose a specialized method, namely IACC, which unifies the luminance features of underwater images under mixed lighting and guides consistent enhancement across similar luminance regions. Furthermore, complementary colors are introduced to globally guide the correction of color discrepancies, preserving the structural consistency and mitigating potential structural information loss during the original image feature extraction. Extensive experiments on various underwater datasets demonstrate the superiority of our method, which outperforms state-of-the-art methods in both machine and human visual perception. Our code is available athttps://github.com/zhoujingchun03/IACC. Jingchun Zhou, Qilin Gai, Dehuan Zhang, Kin-Man Lam 0001, Weishi Zhang, Xianping Fu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Injecting Text Clues for Improving Anomalous Event Detection From Weakly Labeled VideosabstractVideo anomaly detection (VAD) aims at localizing the snippets containing anomalous events in long unconstrained videos. The weakly supervised (WS) setting, where solely video-level labels are available during training, has attracted considerable attention, owing to its satisfactory trade-off between the detection performance and annotation cost. However, due to lack of snippet-level dense labels, the existing WS-VAD methods still get easily stuck on the detection errors, caused by false alarms and incomplete localization. To address this dilemma, in this paper, we propose to inject text clues of anomaly-event categories for improving WS-VAD, via a dedicated dual-branch framework. For suppressing the response of confusing normal contexts, we first present a text-guided anomaly discovering (TAG) branch based on a hierarchical matching scheme, which utilizes the label-text queries to search the discriminative anomalous snippets in a global-to-local fashion. To facilitate the completeness of anomaly-instance localization, an anomaly-conditioned text completion (ATC) branch is further designed to perform an auxiliary generative task, which intrinsically forces the model to gather sufficient event semantics from all the relevant anomalous snippets for completely reconstructing the masked description sentence. Furthermore, to encourage the cross-branch knowledge sharing, a mutual learning strategy is introduced by imposing a consistency constraint on the anomaly scores of these two branches. Extensive experimental results on two public benchmarks validate that the proposed method achieves superior performance over the competing methods. Tianshan Liu, Kin-Man Lam 0001, Bing-Kun Bao |
IEEE Trans. Image Process. | 2 |
| 2024 | Landmark Localization From Medical Images With Generative Distribution PriorabstractIn medical image analysis, anatomical landmarks usually contain strong prior knowledge of their structural information. In this paper, we propose to promote medical landmark localization by modeling the underlying landmark distribution via normalizing flows. Specifically, we introduce the flow-based landmark distribution prior as a learnable objective function into a regression-based landmark localization framework. Moreover, we employ an integral operation to make the mapping from heatmaps to coordinates differentiable to further enhance heatmap-based localization with the learned distribution prior. Our proposed Normalizing Flow-based Distribution Prior (NFDP) employs a straightforward backbone and non-problem-tailored architecture (i.e., ResNet18), which delivers high-fidelity outputs across three X-ray-based landmark localization datasets. Remarkably, the proposed NFDP can do the job with minimal additional computational burden as the normalizing flows module is detached from the framework on inferencing. As compared to existing techniques, our proposed NFDP provides a superior balance between prediction accuracy and inference speed, making it a highly efficient and effective approach. The source code of this paper is available at https://github.com/jacksonhzx95/NFDP. Zixun Huang, Rui Zhao 0012, Frank H. F. Leung, Sunetra Banerjee, Kin-Man Lam 0001, Sai-Ho Ling |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Point Clouds are Specialized Images: A Knowledge Transfer Approach for 3D UnderstandingabstractSelf-supervised representation learning (SSRL) has gained increasing attention in point cloud understanding, in addressing the challenges posed by 3D data scarcity and high annotation costs. This paper presents PCExpert, a novel SSRL approach that reinterprets point clouds as “specialized images”. This conceptual shift allows PCExpert to leverage knowledge derived from large-scale image modality in a more direct and deeper manner, via extensively sharing the parameters with a pre-trained image encoder in a multi-way Transformer architecture. The parameter sharing strategy, combined with an additional pretext task for pre-training, i.e., transformation estimation, empowers PCExpert to outperform the state of the arts in a variety of tasks, with a remarkable reduction in the number of trainable parameters. Notably, PCExpert's performance underLINEARfine-tuning (e.g., yielding a 90.02% overall accuracy on ScanObjectNN) has already closely approximated the results obtained withFULLmodel fine-tuning (92.66%), demonstrating its effective representation capability. Jiachen Kang, Wenjing Jia, Xiangjian He, Kin-Man Lam 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Distilling Privileged Knowledge for Anomalous Event Detection From Weakly Labeled VideosabstractWeakly supervised video anomaly detection (WS-VAD) aims to identify the snippets involving anomalous events in long untrimmed videos, with solely text video-level binary labels. A typical paradigm among the existing text WS-VAD methods is to employ multiple modalities as inputs, e.g., RGB, optical flow, and audio, as they can provide sufficient discriminative clues that are robust to the diverse, complicated real-world scenes. However, such a pipeline has high reliance on the availability of multiple modalities and is computationally expensive and storage demanding in processing long sequences, which limits its use in some applications. To address this dilemma, we propose a privileged knowledge distillation (KD) framework dedicated to the WS-VAD task, which can maintain the benefits of exploiting additional modalities, while avoiding the need for using multimodal data in the inference phase. We argue that the performance of the privileged KD framework mainly depends on two factors: 1) the effectiveness of the multimodal teacher network and 2) the completeness of the useful information transfer. To obtain a reliable teacher network, we propose a text cross-modal interactive learning strategy and an anomaly normal discrimination loss, which target learning task-specific cross-modal features and encourage the separability of anomalous and normal representations, respectively. Furthermore, we design both representation- and text logits-level distillation loss functions, which force the unimodal student network to distill abundant privileged knowledge from the text well-trained multimodal teacher network, in a snippet-to-video fashion. Extensive experimental results on three public benchmarks demonstrate that the proposed privileged KD framework can train a lightweight yet effective detector, for localizing anomaly events under the supervision of video-level annotations. Tianshan Liu, Kin-Man Lam 0001, Jun Kong 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Holistic-Guided Disentangled Learning With Cross-Video Semantics Mining for Concurrent First-Person and Third-Person Activity RecognitionabstractThe popularity of wearable devices has increased the demands for the research on first-person activity recognition. However, most of the current first-person activity datasets are built based on the assumption that only the human-object interaction (HOI) activities, performed by the camera-wearer, are captured in the field of view. Since humans live in complicated scenarios, in addition to the first-person activities, it is likely that third-person activities performed by other people also appear. Analyzing and recognizing these two types of activities simultaneously occurring in a scene is important for the camera-wearer to understand the surrounding environments. To facilitate the research on concurrent first- and third-person activity recognition (CFT-AR), we first created a new activity dataset, namely PolyU concurrent first- and third-person (CFT) Daily, which exhibits distinct properties and challenges, compared with previous activity datasets. Since temporal asynchronism and appearance gap usually exist between the first- and third-person activities, it is crucial to learn robust representations from all the activity-related spatio-temporal positions. Thus, we explore both holistic scene-level and local instance-level (person-level) features to provide comprehensive and discriminative patterns for recognizing both first- and third-person activities. On the one hand, the holistic scene-level features are extracted by a 3-D convolutional neural network, which is trained to mine shared and sample-unique semantics between video pairs, via two well-designed attention-based modules and a self-knowledge distillation (SKD) strategy. On the other hand, we further leverage the extracted holistic features to guide the learning of instance-level features in a disentangled fashion, which aims to discover both spatially conspicuous patterns and temporally varied, yet critical, cues. Experimental results on the PolyU CFT Daily dataset validate that our method achieves the state-of-the-art performance. Tianshan Liu, Rui Zhao 0012, Wenqi Jia 0001, Kin-Man Lam 0001, Jun Kong 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | GR-PSN: Learning to Estimate Surface Normal and Reconstruct Photometric Stereo ImagesabstractIn this paper, we propose a novel method, namely GR-PSN, which learns surface normals from photometric stereo images and generates the photometric images under distant illumination from different lighting directions and surface materials. The framework is composed of two subnetworks, named GeometryNet and ReconstructNet, which are cascaded to perform shape reconstruction and image rendering in an end-to-end manner. ReconstructNet introduces additional supervision for surface-normal recovery, forming a closed-loop structure with GeometryNet. We also encode lighting and surface reflectance in ReconstructNet, to achieve arbitrary rendering. In training, we set up a parallel framework to simultaneously learn two arbitrary materials for an object, providing an additional transform loss. Therefore, our method is trained based on the supervision by three different loss functions, namely the surface-normal loss, reconstruction loss, and transform loss. We alternately input the predicted surface-normal map and the ground-truth into ReconstructNet, to achieve stable training for ReconstructNet. Experiments show that our method can accurately recover the surface normals of an object with an arbitrary number of inputs, and can re-render images of the object with arbitrary surface materials. Extensive experimental results show that our proposed method outperforms those methods based on a single surface recovery network and shows realistic rendering results on 100 different materials. Yakun Ju, Boxin Shi, Yang Chen 0036, Huiyu Zhou 0001, Junyu Dong, Kin-Man Lam 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2023 | Efficient Feature Fusion for Learning-Based Photometric StereoabstractHow to handle an arbitrary number for input images is a fundamental problem of learning-based photometric stereo methods. Existing approaches adopt max-pooling or observation map to fuse an arbitrary number of extracted features. However, these methods discard a large amount of the features from the input images, impacting the utilization and accuracy, or ignore the constraints from the intra-image spatial domain. In this paper, we explore how to efficiently fuse features from a variable number of input images. First, we propose a bilateral extraction module, which categorizes features into positive and negative, to maximally keep the useful feature in the fusion stage. Second, we adopt a top-k pooling to both the bilateral information, which selects the k maximum response value from all features. These two modules proposed are "plug-and-play" and can be used in different fusion tasks. We further propose a hierarchical photometric stereo network, namely HPS-Net, to handle bilateral extraction and top-k pooling for multiscale features. Experiments in the widely used benchmark illustrate the improvement of our proposed framework in the conventional max-pooling method and the proposed HPS-Net outperforms existing learning-based photometric stereo methods. Yakun Ju, Kin-Man Lam 0001, Jun Xiao 0010, Cuixin Yang, Junyu Dong |
ICASSP | 2 |
| 2023 | Improving Robustness of Single Image Super-Resolution Models with Monte Carlo MethodabstractDeep learning-based methods have achieved promising results in single image super-resolution (SISR). However, the performance of existing deep SISR methods is very sensitive to image degradation. In addition, these methods are deterministic and do not introduce any uncertainty to the generated images, so we have no way of knowing the reliability of these generated images. To address these two challenging issues, we propose a model-agnostic approach for existing deep SISR networks to improve their robustness under various degradations. Our proposed method follows a probabilistic framework and applies Monte Carlo dropout to existing deep SISR methods. Instead of performing point estimation, the proposed method predicts the posterior distribution of super-resolved images. Based on this, we can determine the uncertainty of the generated images. Experiment results show that the proposed method can effectively improve the robustness of existing deep SISR methods, leading to state-of-the-art performance when applied to images having different degradations. The code is available at https://github.com/YangTracy/MCD-SR. Cuixin Yang, Jun Xiao 0010, Yakun Ju, Guoping Qiu, Kin-Man Lam 0001 |
ICIP | 5 |
| 2023 | Pyramid Masked Image Modeling for Transformer-Based Aerial Object DetectionabstractTwo obstacles, the scarcity of annotated samples and the difficulty in preserving multi-scale hierarchical representations, hinder the advancement of vision Transformer-based aerial object detection. The emergence of self-supervised learning has inspired some solutions to the first issue. However, most solutions focus on single-scale features, conflicting with solving the second issue. To bridge this gap, this paper proposes a novel pyramid masked image modeling (MIM) framework, termed PyraMIM, for self-supervised pretraining in aerial scenarios. Without manual annotation, PyraMIM enables establishing pyramid representations during pretraining, which can be seamlessly adapted to downstream aerial object detection for performance improvement. Experimental results demonstrate the effectiveness and superiority of our method. Tianshan Liu, Yakun Ju, Kin-Man Lam 0001 |
ICIP | 4 |
| 2023 | Learning Deep Photometric Stereo Network with Reflectance PriorsabstractPhotometric stereo recovers the surface normals of an object from images with varying shading cues. Conventional photometric stereo methods attempt to use handcrafted reflectance models to approximate surface normals, while deep learning-based networks have shown a much more powerful ability to handle non-Lambertian objects. However, none of the existing deep learning methods explores how prior reflectance information can be used to optimize surface-normal prediction. In this paper, we first present the introduction of reflectance prior to deep photometric stereo models. Our explorations include how the reflectance prior can simplify the optimization of deep networks by reparametrizing the weights, and (2) eliminate the impacts of surfaces with spatially varying reflectance for all-pixel input photometric stereo methods. To achieve these goals, we propose a residual fusion module (RFM) in our method, which explicitly extracts features useful for surface-normal recovery and removes those features influenced by reflectance. Additionally, we design a shading extractor with multi-scale and global-local feature fusion operations, which can fuse features with different receptive fields and better utilize the non-maximum features missing in the max-pooling operation. Experiments and ablation studies verify the accuracy and effectiveness of the proposed reflectance prior network on a widely used benchmark. Yakun Ju, Songsong Huang, Yuan Rao 0001, Kin-Man Lam 0001 |
ICME | 5 |
| 2023 | Restoration of Multiple Image Distortions using a Semi-dynamic Deep Neural NetworkabstractRestoring multiple image distortions with a single model is difficult because different distortions require fundamentally different processing mechanisms, e.g., deblurring requires high-pass filtering, while denoising requires low-pass filtering operations. This paper presents a dynamic universal image restoration (DUIR) system capable of simultaneously processing multiple distortions. The new model features several innovative designs: (i) a distortion embedding module (DEM) to automatically encode the distortion information of an input, (ii) a distortion attention module (DAM) that uses a bi-directional long short-term memory (LSTM) to encode the distortion into a sequence of forward and backward interdependent modulating signals, and (iii) a dynamically adaptive image restoration deep convolutional neural network (DAIR-DCNN) featuring unique semi-dynamic layers (SDLs) in which part of their parameters are dynamically modulated by the distortion signals. DEM, DAM, and SDLs together make DAIR-DCNN adaptive to the distortions of the current input, which in turn equips the DUIR system with the capability of simultaneously processing multiple image distortions with a single trained model. We present extensive experimental results to show that the new technique achieves superior performance to state-of-the-art models on both synthetic and real data. We further demonstrate that a trained DUIR system can simultaneously handle different distortions, including those with conflicting demands, such as denoising, deblurring, and compression artifact removal. Hongming Luo, Fei Zhou 0001, Zehong Zhou, Kin-Man Lam 0001, Guoping Qiu |
ACM Multimedia | 4 |
| 2023 | Spectrum-irrelevant fine-grained representation for visible-infrared person re-identification
Jiahao Gong, Sanyuan Zhao, Kin-Man Lam 0001, Xin Gao 0001, Jianbing Shen |
Comput. Vis. Image Underst. | 3 |
| 2023 | Causal diffused graph-transformer network with stacked early classification loss for efficient stream classification of rumours
Tsun-Hin Cheung, Kin-Man Lam 0001 |
Knowl. Based Syst. | 2 |
| 2023 | Robust seed selection of foreground and background priors based on directional blocks for saliency-detection system
Muwei Jian, Ruihong Wang, Hui Yu 0001, Junyu Dong, Gongfa Li, Yilong Yin, Kin-Man Lam 0001 |
Multim. Tools Appl. | 8 |
| 2023 | Dual autoencoder based zero shot learning in special domain
Eric Rigall, Xin Sun 0003, Kin-Man Lam 0001, Junyu Dong |
Pattern Anal. Appl. | 4 |
| 2023 | Self-embedding reversible color-to-grayscale conversion with watermarking feature
S. K. Felix Yu, Yuk-Hee Chan, Kin-Man Lam 0001, Daniel Pak-Kong Lun |
Signal Process. Image Commun. | 3 |
| 2023 | Geometry-Aware Facial Expression Recognition via Attentive Graph Convolutional NetworksabstractLearning discriminative representations with good robustness from facial observations serves as a fundamental step towards intelligent facial expression recognition (FER). In this article, we propose a novel geometry-aware FER framework to boost the FER performance based on both the geometric and appearance knowledge. Specifically, we propose an encoding strategy for facial landmarks, and adopt a graph convolutional network (GCN) to fully explore the structural information of the facial components behind different expressions. A convolutional neural network (CNN) is further applied to the whole facial observation to learn the global characteristics of different expressions. The features from these two networks are fused into a comprehensive high-semantic representation, which promotes the FER reasoning from both visual and structural perspectives. Moreover, to facilitate the networks to concentrate on the most informative facial regions and components, we introduce multi-level attention mechanisms into the proposed framework, which enhance the reliability of the learned representations for effective FER. Experiments on two challenging FER benchmarks demonstrate that the attentive graph-based learning on the facial geometry boosts the FER accuracy. Furthermore, the insensitivity of the geometric information to the appearance variations also improves the generalization of the proposed framework. Rui Zhao 0012, Tianshan Liu, Zixun Huang, Daniel Pak-Kong Lun, Kin-Man Lam 0001 |
IEEE Trans. Affect. Comput. | 5 |
| 2023 | Spatial-Temporal Graphs Plus Transformers for Geometry-Guided Facial Expression RecognitionabstractFacial expression recognition (FER) is of great interest to the current studies of human-computer interaction. In this paper, we propose a novel geometry-guided facial expression recognition framework, based on graph convolutional networks and transformers, to perform effective emotion recognition from videos. Specifically, we detect and utilize facial landmarks to construct a spatial-temporal graph, based on both the landmark coordinates and local appearance, for representing a facial expression sequence. The graph convolutional blocks and transformer modules are employed to produce high-semantic emotion-related representations from the structured facial graphs, which facilitate the framework to establish both the local and non-local dependency between the vertices. Moreover, spatial and temporal attention mechanisms are introduced into graph-based learning to promote FER reasoning, via the emphasis on the most informative facial components and frames. Extensive experiments demonstrate that the proposed framework achieves promising performance for geometry-based FER and shows great generalization and robustness in real-world applications. Rui Zhao 0012, Tianshan Liu, Zixun Huang, Daniel Pak-Kong Lun, Kin-Man Lam 0001 |
IEEE Trans. Affect. Comput. | 5 |
| 2023 | CoF-Net: A Progressive Coarse-to-Fine Framework for Object Detection in Remote-Sensing ImageryabstractObject detection in remote-sensing images is a crucial task in the fields of Earth observation and computer vision. Despite impressive progress in modern remote-sensing object detectors, there are still three challenges to overcome: 1) complex background interference; 2) dense and cluttered arrangement of instances; and 3) large-scale variations. These challenges lead to two key deficiencies, namely, coarse features and coarse samples, which limit the performance of existing object detectors. To address these issues, in this article, a novel coarse-to-fine framework (CoF-Net) is proposed for object detection in remote-sensing imagery. CoF-Net mainly consists of two parallel branches, namely, coarse-to-fine feature adaptation (CoF-FA) and coarse-to-fine sample assignment (CoF-SA), which aim to progressively enhance feature representation and select stronger training samples, respectively. Specifically, CoF-FA smoothly refines the original coarse features into multispectral nonlocal fine features with discriminative spatial–spectral details and semantic relations. Meanwhile, CoF-SA dynamically considers samples from coarse to fine by progressively introducing geometric and classification constraints for sample assignment during training. Comprehensive experiments on three public datasets demonstrate the effectiveness and superiority of the proposed method. Kin-Man Lam 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Efficient Inductive Vision Transformer for Oriented Object Detection in Remote Sensing ImageryabstractObject detection is a fundamental task in remote sensing image analysis and scene understanding. Previous remote sensing object detectors are typically based on convolutional neural networks (CNNs), whose performance is significantly limited by the intrinsic locality of convolution operations. The emergence of vision Transformers brings potential solutions to this problem, which have the capability to be a solid alternative to CNNs. However, three crucial obstacles hinder the application and performance of Transformers in the task of remote sensing object detection, i.e., 1) high computational complexity, especially for high-resolution remote sensing images, 2) training-and sample-inefficiency caused by lack of inductive bias, and 3) difficulty in learning arbitrary orientation knowledge of geospatial objects. To address these issues, in this paper, a novel efficient inductive vision Transformer framework is proposed for oriented object detection in remote sensing imagery. This framework follows the hierarchical feature pyramid structure and makes threefold contributions, as follows. 1) Spatial redundancy in remote sensing images is fully explored and an adaptive multi-grained routing mechanism is proposed to facilitate token sparsity in Transformers, which can dramatically reduce the computational cost without comprising the accuracy. 2) A compact dual-path encoding architecture, where both global long-range dependencies and local semantic relations are jointly and complementarily captured, is proposed to enhance inductive bias in Transformers. 3) An angle tokenization technique is proposed to promote the encoding, embedding, and learning of direction knowledge for oriented objects in remote sensing scenarios. In this work, the above three contributions are instantiated in an advanced Transformer-based object detector, namely EIA-PVT. Comprehensive experiments on two publicly available datasets have demonstrated its effectiveness and superiority for oriented object detection in remote sensing images. Jingran Su, Yakun Ju, Kin-Man Lam 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Decouple and Resolve: Transformer-Based Models for Online Anomaly Detection From Weakly Labeled VideosabstractAs one of the vital topics in intelligent surveillance, weakly supervised online video anomaly detection (WS-OVAD) aims to identify the ongoing anomalous events moment-to-moment in streaming videos, trained with only video-level annotations. Previous studies tended to utilize a unified single-stage framework, which struggled to simultaneously address the issues of online constraints and weakly supervised settings. To solve this dilemma, in this paper, we propose a two-stage-based framework, namely “decouple and resolve” (DAR), which consists of two modules, i.e., temporal proposal producer (TPP) and online anomaly localizer (OAL). With the supervision of video-level binary labels, the TPP module targets fully exploiting hierarchical temporal relations among snippets for generating precise snippet-level pseudo-labels. Then, given fine-grained supervisory signals produced by TPP, the Transformer-based OAL module is trained to aggregate both the useful cues retrieved from historical observations and anticipated future semantics, for making predictions at the current time step. Both the TPP and OAL modules are jointly trained to share the beneficial knowledge in a multi-task learning paradigm. Extensive experimental results on three public data sets validate the superior performance of the proposed DAR framework over the competing methods. Tianshan Liu, Kin-Man Lam 0001, Jun Kong 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | Online Video Super-Resolution With Convolutional Kernel Bypass GraftsabstractDeep learning-based models have achieved remarkable performance in video super-resolution (VSR) in recent years, but most of these models are less applicable to online video applications. These methods solely consider the distortion quality and ignore crucial requirements for online applications, e.g., low latency and low model complexity. In this paper, we focus on online video transmission in which VSR algorithms are required to generate high-resolution video sequences frame by frame in real time. To address such challenges, we propose an extremely low-latency VSR algorithm based on a novel kernel knowledge transfer method, named the convolutional kernel bypass graft (CKBG). First, we design a lightweight network structure that does not require future frames as inputs and saves extra time for caching these frames. Then, our proposed CKBG method enhances this lightweight base model by bypassing the original network with “kernel grafts”, which are extra convolutional kernels containing the prior knowledge of the external pretrained image SR models. During the testing phase, we further accelerate the grafted multibranch network by converting it into a simple single-path structure. The experimental results show that our proposed method can process online video sequences up to 110 FPS with very low model complexity and competitive SR performance. Jun Xiao 0010, Xinyang Jiang, Ningxin Zheng, Huan Yang 0005, Yifan Yang 0004, Yuqing Yang 0001, Dongsheng Li 0002, Kin-Man Lam 0001 |
IEEE Trans. Multim. | 8 |
| 2022 | A Hybrid Egocentric Activity Anticipation Framework via Memory-Augmented Recurrent and One-shot Representation ForecastingabstractEgocentric activity anticipation involves identifying the interacted objects and target action patterns in the near future. A standard activity anticipation paradigm is re-currently forecasting future representations to compensate the missing activity semantics of the unobserved sequence. However, the limitations of current recursive prediction models arise from two aspects: (i) The vanilla recurrent units are prone to accumulated errors in relatively long periods of anticipation. (ii) The anticipated representations may be insufficient to reflect the desired semantics of the target activity, due to lack of contextual clues. To address these issues, we propose “HRO ”, a hybrid framework that integrates both the memory-augmented recurrent and one-shot representation forecasting strategies. Specifically, to solve the limitation (i), we introduce a memory-augmented contrastive learning paradigm to regulate the process of the recurrent representation forecasting. Since the external memory bank maintains long-term prototypical activity semantics, it can guarantee that the anticipated representations are reconstructed from the discriminative activity prototypes. To further guide the learning of the memory bank, two auxiliary loss functions are designed, based on the diversity and sparsity mechanisms, respectively. Furthermore, to resolve the limitation (ii), a one-shot transferring paradigm is proposed to enrich the forecasted representations, by distilling the holistic activity semantics after the target anticipation moment, in the offline training. Extensive experimental results on two large-scale data sets validate the effectiveness of our proposed HRO method. Tianshan Liu, Kin-Man Lam 0001 |
CVPR | 2 |
| 2022 | Interaction and Alignment for Visible-Infrared Person Re-IdentificationabstractVisible-Infrared Person Re-Identification (VI-ReID) is a challenging person matching problem and is also a practical solution for intelligent surveillance systems at night. Due to the heterogeneity between visible and infrared modalities, the retrieval performance is seriously damaged. To address the issue of the discrepancy of the information between visible and infrared modalities, many works have been proposed. However, the relationship between cross-modality samples has rarely been mined. In this paper, we propose a Cross-modality Interaction and Alignment (CIA) module to solve the discrepancy problem. Through transforming the information between different modalities, the module guides the network to capture the modality-shared feature, which is beneficial to address the cross-modality discrepancy. Meanwhile, to better supervise the network, an enhanced contrastive loss is introduced. Contributed by the further optimization in the distance between intra-class samples, the network gains more effective supervision. Extensive experiments on two benchmark datasets show that our method achieves an excellent performance in VI-ReID. Jiahao Gong, Sanyuan Zhao, Kin-Man Lam 0001 |
ICPR | 3 |
| 2022 | Angle Tokenization Guided Multi-Scale Vision Transformer for Oriented Object Detection in Remote Sensing ImageryabstractIn this paper, an angle tokenization guided multi-scale Trans-former framework is proposed for oriented object detection in remote sensing images. Different from existing detectors that are based on convolutional neural networks (CNNs), our proposed method is based on a pyramid Transformer architecture with a compact and flexible angle tokenization module (ATM) to efficiently learn the orientation knowledge for rotated geospatial objects. The Transformer structure can progressively render long-range dependencies and multi-scale spatial details required for accurate localization, while the ATM provides robust guidance on feature refinement for angle prediction, jointly achieving end-to-end orientation de-tection. To the best of our knowledge, this is the first work to adapt Vision Transformers to remote sensing oriented object detection. Experimental results demonstrate the effectiveness and superiority of our method. Tianshan Liu, Kin-Man Lam 0001 |
IGARSS | 3 |
| 2022 | AGTGAN: Unpaired Image Translation for Photographic Ancient Character GenerationabstractThe study of ancient writings has great value for archaeology and philology. Essential forms of material are photographic characters, but manual photographic character recognition is extremely time-consuming and expertise-dependent. Automatic classification is therefore greatly desired. However, the current performance is limited due to the lack of annotated data. Data generation is an inexpensive but useful solution to data scarcity. Nevertheless, the diverse glyph shapes and complex background textures of photographic ancient characters make the generation task difficult, leading to unsatisfactory results of existing methods. To this end, we propose an unsupervised generative adversarial network called AGTGAN in this paper. By explicitly modeling global and local glyph shape styles, followed by a stroke-aware texture transfer and an associate adversarial learning mechanism, our method can generate characters with diverse glyphs and realistic textures. We evaluate our method on photographic ancient character datasets, e.g., OBC306 and CSDD. Our method outperforms other state-of-the-art methods in terms of various metrics and performs much better in terms of the diversity and authenticity of generated samples. With our generated images, experiments on the largest photographic oracle bone character dataset show that our method can achieve a significant increase in classification accuracy, up to 16.34%. The source code is available at https://github.com/Hellomystery/AGTGAN. Hongxiang Huang, Daihui Yang, Gang Dai 0002, Zhen Han 0003, Yuyi Wang 0001, Kin-Man Lam 0001, Fan Yang 0082, Shuangping Huang, Yongge Liu, Mengchao He |
ACM Multimedia | 6 |
| 2022 | Restoration of User Videos Shared on Social MediaabstractUser videos shared on social media platforms usually suffer from degradations caused by unknown proprietary processing procedures, which means that their visual quality is poorer than that of the originals. This paper presents a new general video restoration framework for the restoration of user videos shared on social media platforms. In contrast to most deep learning-based video restoration methods that perform end-to-end mapping, where feature extraction is mostly treated as a black box, in the sense that what role a feature plays is often unknown, our new method, termed Video restOration through adapTive dEgradation Sensing (VOTES), introduces the concept of a degradation feature map (DFM) to explicitly guide the video restoration process. Specifically, for each video frame, we first adaptively estimate its DFM to extract features representing the difficulty of restoring its different regions. We then feed the DFM to a convolutional neural network (CNN) to compute hierarchical degradation features to modulate an end-to-end video restoration backbone network, such that more attention is paid explicitly to potentially more difficult to restore areas, which in turn leads to enhanced restoration performance. We will explain the design rationale of the VOTES framework and present extensive experimental results to show that the new VOTES method outperforms various state-of-the-art techniques both quantitatively and qualitatively. In addition, we contribute a large scale real-world database of user videos shared on different social media platforms. Codes and datasets are available at https://github.com/luohongming/VOTES.git Hongming Luo, Fei Zhou 0001, Kin-Man Lam 0001, Guoping Qiu |
ACM Multimedia | 3 |
| 2022 | MGF6mARice: prediction of DNA N6-methyladenine sites in rice by exploiting molecular graph feature and residual blockabstractDNA N6-methyladenine (6mA) is produced by the N6 position of the adenine being methylated, which occurs at the molecular level, and is involved in numerous vital biological processes in the rice genome. Given the shortcomings of biological experiments, researchers have developed many computational methods to predict 6mA sites and achieved good performance. However, the existing methods do not consider the occurrence mechanism of 6mA to extract features from the molecular structure. In this paper, a novel deep learning method is proposed by devising DNA molecular graph feature and residual block structure for 6mA sites prediction in rice, named MGF6mARice. Firstly, the DNA sequence is changed into a simplified molecular input line entry system (SMILES) format, which reflects chemical molecular structure. Secondly, for the molecular structure data, we construct the DNA molecular graph feature based on the principle of graph convolutional network. Then, the residual block is designed to extract higher level, distinguishable features from molecular graph features. Finally, the prediction module is used to obtain the result of whether it is a 6mA site. By means of 10-fold cross-validation, MGF6mARice outperforms the state-of-the-art approaches. Multiple experiments have shown that the molecular graph feature and residual block can promote the performance of MGF6mARice in 6mA prediction. To the best of our knowledge, it is the first time to derive a feature of DNA sequence by considering the chemical molecular structure. We hope that MGF6mARice will be helpful for researchers to analyze 6mA sites in rice. Zhigang Zeng, Kin-Man Lam 0001 |
Briefings Bioinform. | 4 |
| 2022 | NormAttention-PSN: A High-frequency Region Enhanced Photometric Stereo Network with Normalized Attention
Yakun Ju, Boxin Shi, Muwei Jian, Lin Qi 0004, Junyu Dong, Kin-Man Lam 0001 |
Int. J. Comput. Vis. | 6 |
| 2022 | Crossmodal bipolar attention for multimodal classification on social media
Tsun-Hin Cheung, Kin-Man Lam 0001 |
Neurocomputing | 2 |
| 2022 | Visual-semantic graph neural network with pose-position attentive learning for group activity recognition
Tianshan Liu, Rui Zhao 0012, Kin-Man Lam 0001, Jun Kong 0001 |
Neurocomputing | 3 |
| 2022 | Deep low-rank feature learning and encoding for cross-age face recognition
M. Saad Shakeel, Kin-Man Lam 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2022 | Corn-Plant Counting Using Scare-Aware Feature and Channel InterdependenceabstractCorn-plant counting is an important process for predicting corn yield and analyzing corn-plant phenotypes. In this letter, an effective corn-plant counting method is proposed, which is based on utilizing the scale-aware (SA) contextual feature and channel interdependence (CI). Given the Visual Geometry Group (VGG) Network features, the SA features are extracted by spatial pyramid pooling to derive multiscale context information. In order to utilize the channel interdependent information, the VGG features are integrated via a channel attention module. Moreover, an encoder–decoder structure is constructed to fuse the SA features and the CI-based features. Considering the sparsity of a corn plant, a hybrid loss function is adopted to train the network, by considering a density map loss function and an absolute count loss function. Experimental results demonstrate the effectiveness of the proposed method for corn-plant counting. Yong-Yang Ma, Zhigang Zeng, Kin-Man Lam 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Visual saliency detection via combining center prior and U-Net
Xiangwei Lu, Muwei Jian, Xing Wang 0002, Hui Yu 0001, Junyu Dong, Kin-Man Lam 0001 |
Multim. Syst. | 6 |
| 2022 | Facial Expressions of Comprehension (FEC)abstractWhile the relationship between facial expressions and emotion has been a productive area of inquiry, research is only recently exploring whether a link exists between facial expressions and cognitive processes. Using findings from psychology and neuroscience to guide predictions of affectation during a cognitive task, this article aimed to study facial dynamics as a mean to understand comprehension. We present a new multimodal facial expression database, named Facial Expressions of Comprehension (FEC), consisting of the videos recorded during a computer-mediated task in which each trial consisted of reading, answering, and feedback to general knowledge true and false statements. To identify the level of engagement with the corresponding stimuli, we present a new methodology using animation units (AnUs) from the Kinect v2 device to explore the changes in facial configuration caused by an event: Event-Related Intensities (ERIs). To identify dynamic facial configurations, we used ERIs in statistical analyses with generalized additive models. To identify differential facial dynamics linked to knowing vs. guessing and true vs. false responses, we employed an SVM classifier with facial appearance information extracted using LPQ-TOP. Results of ERIs in sentence comprehension show that facial dynamics are promising to help understand affective and cognitive states of the mind. Cigdem Turan, Karl David Neergaard, Kin-Man Lam 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2022 | Enhanced Attention Tracking With Multi-Branch Network for Egocentric Activity RecognitionabstractThe emergence of wearable devices has opened up new potentials for egocentric activity recognition. Although some methods integrate attention mechanisms into deep neural networks to capture fine-grained human-object interactions in a weak-supervision manner, they either ignore exploiting the temporal consistency or generate attention based on considering appearance cues only. To address these limitations, in this paper, we propose an enhanced attention-tracking method, combined with multi-branch network (EAT-MBNet), for egocentric activity recognition. Specifically, we propose class-aware attention maps (CAAMs) by employing a self-attention-based module to refine the class activation maps (CAMs). Our proposed method can enhance the semantic dependency between the activity categories and the feature maps. To highlight the discriminative features from the regions of interest across frames, we propose a flow-guided attention-tracking (F-AT) module, by simultaneously leveraging historical attention and motion patterns. Furthermore, we propose a cross-modality modeling branch based on an interactive GRU module, which captures the time-synchronized long-term relationships between the appearance and motion branches. Experimental results on four egocentric activity benchmarks demonstrate that the proposed method achieves state-of-the-art performance. Tianshan Liu, Kin-Man Lam 0001, Rui Zhao 0012, Jun Kong 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Deep Cross-Modal Representation Learning and Distillation for Illumination-Invariant Pedestrian DetectionabstractIntegrating multispectral data has been demonstrated to be an effective solution for illumination-invariant pedestrian detection, in particular, RGB and thermal images can provide complementary information to handle light variations. However, most of the current multispectral detectors fuse the multimodal features by simple concatenation, without discovering their latent relationships. In this paper, we propose a cross-modal feature learning (CFL) module, based on a split-and-aggregation strategy, to explicitly explore both the shared and modality-specific representations between paired RGB and thermal images. We insert the proposed CFL module into multiple layers of a two-branch-based pedestrian detection network, to learn the cross-modal representations in diverse semantic levels. By introducing a segmentation-based auxiliary task, the multimodal network is trained end-to-end by jointly optimizing a multi-task loss. On the other hand, to alleviate the reliance of existing multispectral pedestrian detectors on thermal images, we propose a knowledge distillation framework to train a student detector, which only receives RGB images as input and distills the cross-modal representations guided by a well-trained multimodal teacher detector. In order to facilitate the cross-modal knowledge distillation, we design different distillation loss functions for the feature, detection and segmentation levels. Experimental results on the public KAIST multispectral pedestrian benchmark validate that the proposed cross-modal representation learning and distillation method achieves robust performance. Tianshan Liu, Kin-Man Lam 0001, Rui Zhao 0012, Guoping Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Implicit and Explicit Feature Purification for Age-Invariant Facial Representation LearningabstractThis paper presents a new method, named implicit and explicit feature purification (IEFP), for age-invariant face recognition. Facial features extracted from a face image contain the information about the identity, age, and other attributes. For age-invariant face recognition, it is important to remove the irrelevant information, and retain the identity information only, in the facial features. Through the two proposed feature purification mechanisms, our framework can produce facial-feature embeddings that preserve identity information as much as possible and are insensitive to age variations. Specifically, on the one hand, a special network module is devised to implicitly purify the original facial features obtained from a face encoder. On the other hand, to obtain purer facial feature representations for age-invariant face recognition, irrelevant information within the implicitly purified features, such as the age, is further removed. This is realized by using a regularizer, based on information theory, to explicitly minimize the correlation between identity-related features and age-related features. Comprehensive ablation studies show that these two feature purification schemes can work independently, as well as collaboratively, to achieve better performance. Extensive evaluations on several benchmark data sets show that the IEFP method is on par with those competitors learned on far more favorable training samples, and it achieves the best performance in a fair comparison. Furthermore, we provide mathematical interpretation to explain the effectiveness of our approach, and find that it tends to generate low-rank, yet high-dimensional, representations for age-invariant face recognition. Jiucheng Xie, Chi-Man Pun, Kin-Man Lam 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | Joint Spine Segmentation and Noise Removal From Ultrasound Volume Projection Images With Selective Feature SharingabstractVolume Projection Imaging from ultrasound data is a promising technique to visualize spine features and diagnose Adolescent Idiopathic Scoliosis. In this paper, we present a novel multi-task framework to reduce the scan noise in volume projection images and to segment different spine features simultaneously, which provides an appealing alternative for intelligent scoliosis assessment in clinical applications. Our proposed framework consists of two streams: i) A noise removal stream based on generative adversarial networks, which aims to achieve effective scan noise removal in a weakly-supervised manner, i.e., without paired noisy-clean samples for learning; ii) A spine segmentation stream, which aims to predict accurate bone masks. To establish the interaction between these two tasks, we propose a selective feature-sharing strategy to transfer only the beneficial features, while filtering out the useless or harmful information. We evaluate our proposed framework on both scan noise removal and spine segmentation tasks. The experimental results demonstrate that our proposed method achieves promising performance on both tasks, which provides an appealing approach to facilitating clinical diagnosis. Zixun Huang, Rui Zhao 0012, Frank H. F. Leung, Sunetra Banerjee, Timothy Tin-Yan Lee, De Yang, Daniel Pak-Kong Lun, Kin-Man Lam 0001, Sai-Ho Ling |
IEEE Trans. Medical Imaging | 8 |
| 2021 | Feature Redundancy Mining: Deep Light-Weight Image Super-Resolution ModelabstractDespite the great success achieved by deep convolutional neural network (CNN)-based models in the single image super-resolution (SISR) problem, the requirement of high computational complexity, accompanied with the deep CNN models, makes it less applicable in embedded devices, e.g., mobile phones. Recently, deep light-weight models for the SISR problem have been in demand for industrial applications, and have caught the attention of many researchers. The strategies of cascading several small networks and multi-path feature extraction have shown their effectiveness in most of the existing methods. In this paper, by considering the correlation and redundancy of feature maps, we propose a feature information mining network to efficiently investigate the features, for the SISR problem. Experiment results show that our proposed model achieves the best balance between the performance and the model size, compared with other competitive deep SR models. Jun Xiao 0010, Wenqi Jia 0001, Kin-Man Lam 0001 |
ICASSP | 3 |
| 2021 | Structure-Enhanced Attentive Learning For Spine Segmentation From Ultrasound Volume Projection ImagesabstractAutomatic spine segmentation, based on ultrasound volume projection imaging (VPI), is of great value in clinical applications to diagnose scoliosis in teenagers. In this paper, we propose a novel framework to improve the segmentation accuracy on spine images via structure-enhanced attentive learning. Since the spine bones contain strong prior knowledge of their shapes and positions in ultrasound VPI images, we propose to encode this information into the semantic representations in an attentive manner. We first revisit the self-attention mechanism in representation learning, and then present a strategy to introduce the structural knowledge into the key representation in self-attention. By this means, the network explores both the contextual and structural information in the learned features, and consequently improves the segmentation accuracy. We conduct various experiments to demonstrate that our proposed method achieves promising performance on spine image segmentation, which shows great potential in clinical diagnosis. Rui Zhao 0012, Zixun Huang, Tianshan Liu, Frank H. F. Leung, Sai-Ho Ling, De Yang, Timothy Tin-Yan Lee, Daniel Pak-Kong Lun, Kin-Man Lam 0001 |
ICASSP | 10 |
| 2021 | A New Semi-automatic Annotation Model via Semantic Boundary Estimation for Scene Text Detection
Zhenzhou Zhuang, Zonghao Liu, Kin-Man Lam 0001, Shuangping Huang, Gang Dai 0002 |
ICDAR (3) | 3 |
| 2021 | Multimodal-Semantic Context-Aware Graph Neural Network for Group Activity RecognitionabstractGroup activities in videos involve visual interaction contexts in multiple modalities between actors, and co-occurrence between individual action labels. However, most of the current group activity recognition methods either model actor-actor relations based on the single RGB modality, or ignore exploiting the label relationships. To capture these rich visual and semantic contexts, we propose a multimodal-semantic context-aware graph neural network (MSCA-GNN). Specifically, we first build two visual sub-graphs based on the appearance cues and motion patterns extracted from RGB and optical-flow modalities, respectively. Then, two attention-based aggregators are proposed to refine each node, by gathering representations from other nodes and heterogeneous modalities. In addition, a semantic graph is constructed based on linguistic embeddings to model label relationships. We employ a bi-directional mapping learning strategy to further integrate the information from both multimodal visual and semantic graphs. Experimental results on two group activity benchmarks show the effectiveness of the proposed method. Tianshan Liu, Rui Zhao 0012, Kin-Man Lam 0001 |
ICME | 3 |
| 2021 | Self-feature Learning: An Efficient Deep Lightweight Network for Image Super-resolutionabstractDeep learning-based models have achieved unprecedented performance in single image super-resolution (SISR). However, existing deep learning-based models usually require high computational complexity to generate high-quality images, which limits their applications in edge devices, e.g., mobile phones. To address this issue, we propose a dynamic, channel-agnostic filtering method in this paper. The proposed method not only adaptively generates convolutional kernels based on the local information of each position, but also can significantly reduce the cost of computing the inter-channel redundancy. Based on this, we further propose a simple, yet effective, deep lightweight model for SISR. Experiment results show that our proposed model outperforms other state-of-the-art deep lightweight SISR models, leading to the best trade-off between the performance and the number of model parameters. Jun Xiao 0010, Rui Zhao 0012, Kin-Man Lam 0001, Kao Wan |
ACM Multimedia | 4 |
| 2021 | Progressive and Selective Fusion Network for High Dynamic Range ImagingabstractThis paper considers the problem of generating an HDR image of a scene from its LDR images. Recent studies employ deep learning and solve the problem in an end-to-end fashion, leading to significant performance improvements. However, it is still hard to generate a good quality image from LDR images of a dynamic scene captured by a hand-held camera, e.g., occlusion due to the large motion of foreground objects, causing ghosting artifacts. The key to success relies on how well we can fuse the input images in their feature space, where we wish to remove the factors leading to low-quality image generation while performing the fundamental computations for HDR image generation, e.g., selecting the best-exposed image/region. We propose a novel method that can better fuse the features based on two ideas. One is multi-step feature fusion; our network gradually fuses the features in a stack of blocks having the same structure. The other is the design of the component block that effectively performs two operations essential to the problem, i.e., comparing and selecting appropriate images/regions. Experimental results show that the proposed method outperforms the previous state-of-the-art methods on the standard benchmark tests. Jun Xiao 0010, Kin-Man Lam 0001, Takayuki Okatani |
ACM Multimedia | 3 |
| 2021 | Constructing an efficient and adaptive learning model for 3D object generationabstractAbstract Studying representation learning and generative modelling has been at the core of the 3D learning domain. By leveraging the generative adversarial networks and convolutional neural networks for point‐cloud representations, we propose a novel framework, which can directly generate 3D objects represented by point clouds. The novelties of the proposed method are threefold. First, the generative adversarial networks are applied to 3D object generation in the point‐cloud space, where the model learns object representation from point clouds independently. In this work, we propose a 3D spatial transformer network, and integrate it into a generation model, whose ability for extracting and reconstructing features for 3D objects can be improved. Second, a point‐wise approach is developed to reduce the computational complexity of the proposed network. Third, an evaluation system is proposed to measure the performance of our model by employing various categories and methods, and the error, considered as the difference between synthesized objects and raw objects are quantitatively compared, is less than 2.8%. Extensive experiments on benchmark dataset show that this method has a strong ability to generate 3D objects in the point‐cloud space, and the synthesized objects have slight differences with man‐made 3D objects. Jiwei Hu, Wupeng Deng, Quan Liu 0001, Kin-Man Lam 0001, Ping Lou |
IET Image Process. | 4 |
| 2021 | Face hallucination based on cluster consistent dictionary learningabstractAbstract Face hallucination is a super‐resolution technique specially designed to reconstruct high‐resolution faces from low‐resolution faces. Most state‐of‐the‐art algorithms leverage position‐patch prior knowledge of human faces to better super‐resolve face images. However, most of them assume the training face dataset is sufficiently large, well cropped or aligned. This paper, proposes a novel example‐based face hallucination method, based on cluster consistent dictionary learning with the assumption that human faces have similar facial structures. In this method, the paired face image patches are firstly labelled as face areas including eyes, nose, mouth and other parts, as well as non‐face areas without requiring the training face images cropped and aligned. Then, the training patches are clustered according their labels and textures. The cluster consistent dictionary is learned to represent the low‐resolution patches and the high‐resolution patches. Finally, the high‐resolution patches of the input low‐resolution face image can be efficiently generated by using the adjusted anchored neighbourhood regression. As utilizing the labelled facial parts prior knowledge, the proposed method represents more details in the reconstruction. Experimental results demonstrate that the authors' algorithm outperforms many state‐of‐the‐art techniques for face hallucination under different datasets. Minqi Li, Xiangjian He, Kin-Man Lam 0001, Kaibing Zhang, Junfeng Jing |
IET Image Process. | 3 |
| 2021 | Balanced distortion and perception in single-image super-resolution based on optimal transport in wavelet domain
Jun Xiao 0010, Tianshan Liu, Rui Zhao 0012, Kin-Man Lam 0001 |
Neurocomputing | 4 |
| 2021 | Bayesian sparse hierarchical model for image denoising
Jun Xiao 0010, Rui Zhao 0012, Kin-Man Lam 0001 |
Signal Process. Image Commun. | 3 |
| 2021 | 3D Shape Estimation With an Enhanced Sparse Representation ApproachabstractIn this paper, an enhanced sparse representation approach is proposed to estimate the 3D shapes of objects in 2D image sequences. In the proposed method, the unknown 3D shape is estimated via a two-stage scheme, namely the main 3D shape estimation stage and the compensatory 3D shape estimation stage. Moreover, a reweighted sparse representation model is constructed to extract the shape bases for each estimation stage. In the sparse model, a reweighted constraint is enforced to enhance the coefficient sparsity of the shape bases. Experimental results on the well-known CMU image sequences demonstrate the effectiveness and feasibility of the proposed approach. Jiaxiang Wang 0001, Zhigang Zeng, Kin-Man Lam 0001 |
IEEE Signal Process. Lett. | 4 |
| 2021 | Invertible Image DecolorizationabstractInvertible image decolorization is a useful color compression technique to reduce the cost in multimedia systems. Invertible decolorization aims to synthesize faithful grayscales from color images, which can be fully restored to the original color version. In this paper, we propose a novel color compression method to produce invertible grayscale images using invertible neural networks (INNs). Our key idea is to separate the color information from color images, and encode the color information into a set of Gaussian distributed latent variables via INNs. By this means, we force the color information lost in grayscale generation to be independent of the input color image. Therefore, the original color version can be efficiently recovered by randomly re-sampling a new set of Gaussian distributed variables, together with the synthetic grayscale, through the reverse mapping of INNs. To effectively learn the invertible grayscale, we introduce the wavelet transformation into a UNet-like INN architecture, and further present a quantization embedding to prevent the information omission in format conversion, which improves the generalizability of the framework in real-world scenarios. Extensive experiments on three widely used benchmarks demonstrate that the proposed method achieves a state-of-the-art performance in terms of both qualitative and quantitative results, which shows its superiority in multimedia communication and storage systems. Rui Zhao 0012, Tianshan Liu, Jun Xiao 0010, Daniel Pak-Kong Lun, Kin-Man Lam 0001 |
IEEE Trans. Image Process. | 5 |
| 2020 | NTGAN: Learning Blind Image Denoising without Clean Reference
Rui Zhao 0012, Daniel Pak-Kong Lun, Kin-Man Lam 0001 |
BMVC | 3 |
| 2020 | Flow-guided Spatial Attention Tracking for Egocentric Activity RecognitionabstractThe popularity of wearable cameras has opened up a new dimension for egocentric activity recognition. While some methods introduce attention mechanisms into deep learning networks to capture fine-grained hand-object interactions, they often neglect exploring the spatio-temporal relationships. Generating spatial attention, without adequately exploiting temporal consistency, will result in potentially sub-optimal performance in the video-based task. In this paper, we propose a flow-guided spatial attention tracking (F-SAT) module, which is based on enhancing motion patterns and inter-frame information, to highlight the discriminative features from regions of interest across a video sequence. A new form of input, namely the optical-flow volume, is presented to provide informative cues from moving parts for spatial attention tracking. The proposed F-SAT module is deployed to a two-branch-based deep architecture, which fuses complementary information for egocentric activity recognition. Experimental results on three egocentric activity benchmarks show that the proposed method achieves state-of-the-art performance. Tianshan Liu, Kin-Man Lam 0001 |
ICPR | 2 |
| 2020 | Deep Multi-task Learning for Facial Expression Recognition and Synthesis Based on Selective Feature SharingabstractMulti-task learning is an effective learning strategy for deep-learning-based facial expression recognition tasks. However, most existing methods take into limited consideration the feature selection, when transferring information between different tasks, which may lead to task interference when training the multi-task networks. To address this problem, we propose a novel selective feature-sharing method, and establish a multi-task network for facial expression recognition and facial expression synthesis. The proposed method can effectively transfer beneficial features between different tasks, while filtering out useless and harmful information. Moreover, we employ the facial expression synthesis task to enlarge and balance the training dataset to further enhance the generalization ability of the proposed method. Experimental results show that the proposed method achieves state-of-the-art performance on those commonly used facial expression recognition benchmarks, which makes it a potential solution to real-world facial expression recognition problems. Rui Zhao 0012, Tianshan Liu, Jun Xiao 0010, Daniel Pak-Kong Lun, Kin-Man Lam 0001 |
ICPR | 5 |
| 2020 | Pay Attention to Devils: A Photometric Stereo Network for Better DetailsabstractWe present an attention-weighted loss in a photometric stereo neural network to improve 3D surface recovery accuracy in complex-structured areas, such as edges and crinkles, where existing learning-based methods often failed. Instead of using a uniform penalty for all pixels, our method employs the attention-weighted loss learned in a self-supervise manner for each pixel, avoiding blurry reconstruction result in such difficult regions. The network first estimates a surface normal map and an adaptive attention map, and then the latter is used to calculate a pixel-wise attention-weighted loss that focuses on complex regions. In these regions, the attention-weighted loss applies higher weights of the detail-preserving gradient loss to produce clear surface reconstructions. Experiments on real datasets show that our approach significantly outperforms traditional photometric stereo algorithms and state-of-the-art learning-based methods. Yakun Ju, Kin-Man Lam 0001, Yang Chen 0036, Lin Qi 0004, Junyu Dong |
IJCAI | 2 |
| 2020 | Learning Photometric Stereo via Manifold-based MappingabstractThree-dimensional reconstruction technologies are fundamental problems in computer vision. Photometric stereo recovers the surface normals of a 3D object from varying shading cues, prevailing in its capability for generating fine surface normal. In recent years, deep learning-based photometric stereo methods are capable of improving the surface-normal estimation under general non-Lambertian surfaces, due to its powerful fitting ability on the non-Lambertian surface. These state-of-the-art methods however usually regress the surface normal directly from the high-dimensional features, without exploring the embedded structural information. This results in the underutilization of the information available in the features. Therefore, in this paper, we propose an efficient manifold-based framework for learning-based photometric stereo, which can better map combined high-dimensional feature spaces to low-dimensional manifolds. Extensive experiments show that our method, learning with the low-dimensional manifolds, achieves more accurate surface-normal estimation, outperforming other state-of-the-art methods on the challenging DiLiGenT benchmark dataset. Yakun Ju, Muwei Jian, Junyu Dong, Kin-Man Lam 0001 |
VCIP | 4 |
| 2020 | Two-dimensional multi-scale perceptive context for scene text recognition
Daihui Yang, Shuangping Huang, Kin-Man Lam 0001, Zhenzhou Zhuang |
Neurocomputing | 4 |
| 2020 | Output based transfer learning with least squares support vector machine and its application in bladder cancer prognosis
Guanjin Wang, Guangquan Zhang 0001, Kup-Sze Choi, Kin-Man Lam 0001, Jie Lu 0001 |
Neurocomputing | 4 |
| 2020 | Deeply learned pore-scale facial features with a large pore-to-pore correspondences dataset
Xianxian Zeng, Xiaodong Wang 0018, Kairui Chen, Dong Li 0028, Yun Zhang 0001, Kin-Man Lam 0001 |
Pattern Recognit. Lett. | 6 |
| 2020 | Progressive Motion Representation Distillation With Two-Branch Networks for Egocentric Activity RecognitionabstractVideo-based egocentric activity recognition involves fine-grained spatio-temporal human-object interactions. State-of-the-art methods, based on the two-branch-based architecture, rely on pre-calculated optical flows to provide motion information. However, this two-stage strategy is computationally intensive, storage demanding, and not task-oriented, which hampers it from being deployed in real-world applications. Albeit there have been numerous attempts to explore other motion representations to replace optical flows, most of the methods were designed for third-person activities, without capturing fine-grained cues. To tackle these issues, in this letter, we propose a progressive motion representation distillation (PMRD) method, based on two-branch networks, for egocentric activity recognition. We exploit a generalized knowledge distillation framework to train a hallucination network, which receives RGB frames as input and produces motion cues guided by the optical-flow network. Specifically, we propose a progressive metric loss, which aims to distill local fine-grained motion patterns in terms of each temporal progress level. To further enforce the proposed distillation framework to concentrate on those informative frames, we integrate a temporal attention mechanism into the metric loss. Moreover, a multi-stage training procedure is employed for the efficient learning of the hallucination network. Experimental results on three egocentric activity benchmarks demonstrate the state-of-the-art performance of the proposed method. Tianshan Liu, Rui Zhao 0012, Jun Xiao 0010, Kin-Man Lam 0001 |
IEEE Signal Process. Lett. | 4 |
| 2020 | Learning the Traditional Art of Chinese Calligraphy via Three-Dimensional Reconstruction and AssessmentabstractThe traditional art of Chinese calligraphy, reflecting the wisdom of the grass-roots community, is the soul of Chinese culture. Just like many other types of craftsmanship, it is part of the historical heritage and is worth conserving, from generation to generation. Since the movements of an ink brush are in a 3D style when Chinese calligraphy is written, they embody “The Power of Beauty,” comprising various reflectance properties and rough-surface geometry. To truly understand the powerful significance and beauty of the art of Chinese calligraphy, in this paper, a 3D calligraphy reconstruction method, based on Photometric Stereo, is designed to capture the detailed appearance of the calligraphy's 3D surface geometry. For assessment, an Iterative Closest Point (ICP) algorithm is applied for registration of 3D intrinsic shapes between the Chinese calligraphy and the calligraphy fans' handwriting. Through matching these two sets of calligraphy characters, the designed system can give a score to the handwriting of a user. Experiments have been performed on Chinese calligraphy from different historical dynasties to evaluate the effectiveness of the proposed scheme, and experimental results show that the developed system is useful and provides a convenient method of calligraphy appreciation and assessment. Muwei Jian, Junyu Dong, Maoguo Gong, Hui Yu 0001, Liqiang Nie, Yilong Yin, Kin-Man Lam 0001 |
IEEE Trans. Multim. | 7 |
| 2019 | Improving Object Detection with Relation Graph InferenceabstractMany classic object detection approaches have proven that detection performance can be improved by adding the object's context information. However, only a few methods have attempted to exploit the object-to-object relationship during learning. The reason for this is that objects may appear at different locations in an image, with an arbitrary size and scale. This makes it difficult to model the objects in a unified way within a network. Inspired by Graph Convolutional Network (GCN), we propose a detection algorithm that can infer the relationship among multiple objects during the inference, achieved by constructing a relation graph dynamically with a self-adopted attention mechanism. The relation graph encodes both the geometric and visual relationship between objects. This can enrich the object feature by aggregating the information from the object and its relevant neighbors. Experiments show that our proposed module can efficiently improve the detection performance of existing object detectors. Chen-Hang He, Shun-Cheung Lai, Kin-Man Lam 0001 |
ICASSP | 3 |
| 2019 | Low-Resolution Face Recognition Based on Identity-Preserved Face HallucinationabstractThe state-of-the-art Convolutional Neural Network (CNN)-based methods have achieved promising recognition performance on human face images. However, the accuracy cannot be retained when face images are at very low resolution (LR). In this paper, we propose a novel loss function, called identity-preserved loss, which combines with the image-content loss to jointly supervise CNNs, for performing face hallucination and recognition simultaneously. Therefore, the trained network is able to perform face hallucination and identity preservation, even if the query face is of very low resolution. More importantly, experimental results show that our proposed method can preserve the identities for the LR images from unknown subjects, who are not included in the training set. The source code of our proposed method is available at: https://github.com/johnnysclai/SR_LRFR. Shun-Cheung Lai, Chen-Hang He, Kin-Man Lam 0001 |
ICIP | 3 |
| 2019 | High-Resolution Face Recognition Via Deep Pore-Feature MatchingabstractBecause of the advancement of capturing devices, both image resolution and image quality have been significantly improved. Efficiently utilizing facial information is beneficial in enhancing the performance of face recognition methods. For high-resolution face images, pore-scale facial features can be observed. The positions and local patterns of pore features are biologically discriminative, so they can be explored for face identification. In this paper, we extend the previous work on pore-scale features, by proposing a new learning-based descriptor, namely PoreNet. Experiment results show that our proposed descriptor achieves an excellent performance on two high-resolution face datasets, namely Bosphorus and MultiPIE. More importantly, our proposed method significantly outperforms the state-of-the-art Convolutional Neural Network (CNN)-based face recognition method, when query faces are highly occluded. The code of our proposed method is available at: https://github.com/johnnysclai/PoreNet. Shun-Cheung Lai, Minna Kong, Kin-Man Lam 0001, Dong Li 0028 |
ICIP | 3 |
| 2019 | Deep Progressive Convolutional Neural Network for Blind Super-Resolution With Multiple DegradationsabstractBlind super-resolution (SR) of blurry and noisy low-resolution (LR) images is still a challenging problem in single image super-resolution (SISR). The performance of most existing convolutional neural network (CNN)-based models is inevitably degraded when LR images are corrupted by both blur and noise. For those blind SR methods based on kernel estimation, accurate estimation is barely attained under complex degradations and this gives rise to poor-quality results. To address these problems, we propose a deep progressive network under a probabilistic framework and a novel up-sampling method for blind super-resolution with multiple degradations, which effectively utilizes image priors across scales. Experimental results show that the proposed method achieves promising performance on images with multiple degradations. Jun Xiao 0010, Rui Zhao 0012, Shun-Cheung Lai, Wenqi Jia 0001, Kin-Man Lam 0001 |
ICIP | 5 |
| 2019 | Enhancement of a CNN-Based Denoiser Based on Spatial and Spectral AnalysisabstractConvolutional neural network (CNN)-based image denoising methods have been widely studied recently, because of their high-speed processing capability and good visual quality. However, most of the existing CNN-based denoisers learn the image prior from the spatial domain, and suffer from the problem of spatially variant noise, which limits their performance in real-world image denoising tasks. In this paper, we propose a discrete wavelet denoising CNN (WDnCNN), which restores images corrupted by various noise with a single model. Since most of the content or energy of natural images resides in the low-frequency spectrum, their transformed coefficients in the frequency domain are highly imbalanced. To address this issue, we present a band normalization module (BNM) to normalize the coefficients from different parts of the frequency spectrum. Moreover, we employ a band discriminative training (BDT) criterion to enhance the model regression. We evaluate the proposed WDnCNN, and compare it with other state-of-the-art denoisers. Experimental results show that WDnCNN achieves promising performance in both synthetic and real noise reduction, making it a potential solution to many practical image denoising applications. Rui Zhao 0012, Kin-Man Lam 0001, Daniel Pak-Kong Lun |
ICIP | 2 |
| 2019 | Multi-scale Capsule Attention-Based Salient Object Detection with Multi-crossed Layer ConnectionsabstractWith the popularization of convolutional networks being used for saliency models, saliency detection performance has achieved significant improvement. However, how to integrate accurate and crucial features for modeling saliency is still underexplored. In this paper, we present CapSalNet, which includes a multi-scale Capsule attention module and multi-crossed layer connections for Salient object detection. We first propose a novel capsule attention model, which integrates multi-scale contextual information with dynamic routing. Then, our model adaptively learns to aggregate multi-level features by using multi-crossed skip-layer connections. Finally, the predicted results are efficiently fused to generate the final saliency map in a coarse-to-fine manner. Comprehensive experiments on four benchmark datasets demonstrate that our proposed algorithm outperforms existing state-of-the-art approaches. Sanyuan Zhao, Jianbing Shen, Kin-Man Lam 0001 |
ICME | 4 |
| 2019 | Deep-network based method for joint image deblocking and super-resolutionabstractMany pieces of research have been conducted on image‐restoration techniques to recover high‐quality images from their low‐quality versions, but they usually aim to handle a single degraded factor. However, captured images usually suffer from various degradation factors, such as low resolution and compression distortion, in the procedures of image acquisition, compression, and transmission simultaneously. Ignoring the correlation of different degraded factors may result in the limited efficiency of the existing image‐restoration methods for captured images. A joint deep‐network‐based image‐restoration algorithm is proposed to establish a restoration framework for image deblocking and super‐resolution. The proposed convolutional neural network is made up of two stages. A deblocking network is constructed with two cascade deblocking subnets first, then, super‐resolution is performed by a very deep network with skipping links. Cascading these two stages forms a novel deep network. An end‐to‐end training scheme is developed, which makes the two stages be trained jointly so as to achieve better performance. Intensive evaluations have been conducted to measure the performance of the authors’ method both in general images and face images. Experimental results on several datasets demonstrate that the proposed method outperforms other state‐of‐the‐art methods, in terms of both subjective and objective performances. Kin-Man Lam 0001, Li Zhuo 0001, Jiafeng Li 0001 |
IET Image Process. | 3 |
| 2019 | Learning sparse discriminant low-rank features for low-resolution face recognition
M. Saad Shakeel, Kin-Man Lam 0001, Shun-Cheung Lai |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | Cascaded one-vs-rest detection network for fine-grained recognition without part annotations
Long Chen 0019, Shengke Wang, Kin-Man Lam 0001, Huiyu Zhou 0001, Muwei Jian, Junyu Dong |
Multim. Tools Appl. | 3 |
| 2019 | Deep-feature encoding-based discriminative model for age-invariant face recognition
M. Saad Shakeel, Kin-Man Lam 0001 |
Pattern Recognit. | 2 |
| 2019 | Image super-resolution via feature-augmented random forest
Kin-Man Lam 0001, Miaohui Wang |
Signal Process. Image Commun. | 2 |
| 2019 | A CSF-Based CNR Approach for Small-Size Image SequencesabstractFor non-rigid structure from motion (NRSFM), the performance of most traditional approaches may decrease significantly when the frame number of the image sequence is relatively small. In this letter, a column space fitting (CSF) based consensus of non-rigid (CNR) reconstruction approach is proposed to deal with the 3D structure estimation problem for small-size image sequences. In the proposed method, a set of trajectory groups are first extracted by utilizing the distance weight of the pairwise points. In order to improve the estimation accuracy, an adaptive rank selection strategy is designed to choose the approximately optimal rank parameter. Corresponding to the trajectory group, the z-coordinates of the observation matrix are estimated by the CSF algorithm due to its good performance. After obtaining the outputs of the CSF-based weak estimators, the final 3D shape is derived by combining the outputs via the alternating directional method of multipliers. Experimental results on several widely used image sequences demonstrate the effectiveness and feasibility of the proposed algorithm. Jiaxiang Wang 0001, Xia Chen 0008, Kin-Man Lam 0001, Zhigang Zeng |
IEEE Signal Process. Lett. | 4 |
| 2019 | Error Concealment for Cloud-Based and Scalable Video Coding of HD VideosabstractThe encoding of HD videos faces two challenges: requirements for a strong processing power and a large storage space. One time-efficient solution addressing these challenges is to use a cloud platform and to use a scalable video coding technique to generate multiple video streams with varying bit-rates. Packet-loss is very common during the transmission of these video streams over the Internet and becomes another challenge. One solution to address this challenge is to retransmit lost video packets, but this will create end-to-end delay. Therefore, it would be good if the problem of packet-loss can be dealt with at the user's side. In this paper, we present a novel system that encodes and stores the videos using the Amazon cloud computing platform, and recover lost video frames on user side using a new Error Concealment (EC) technique. To efficiently utilize the computation power of a user's mobile device, the EC is performed based on a multiple-thread and parallel process. The simulation results clearly show that, on average, our proposed EC technique outperforms the traditional Block Matching Algorithm (BMA) and the Frame Copy (FC) techniques. Muhammad Usman 0015, Xiangjian He, Kin-Man Lam 0001, Min Xu 0001, Syed Mohsin Matloob Bokhari, Jinjun Chen, Mian Ahmad Jan |
IEEE Trans. Cloud Comput. | 3 |
| 2018 | Pyramid Dilated Deeper ConvLSTM for Video Salient Object Detection
Hongmei Song, Wenguan Wang, Sanyuan Zhao, Jianbing Shen, Kin-Man Lam 0001 |
ECCV (11) | 5 |
| 2018 | Fast Vehicle Detection with Lateral Convolutional Neural NetworkabstractIn this paper, we propose a fast vehicle detector for traffic surveillance. We first explore using different feature layers from a deep residual network to perform vehicle detection. Experiment results show that the high-resolution features from earlier feature layers contain more structural information, which is good to achieve fine-grained localization but yields low recall rates. The low-resolution features in the deep layers contain semantically strong information, which is good to represent the objectness but too coarse to achieve accurate localization. Therefore, we decouple the localization and objectness prediction from a single layer. Instead, we employ a lateral network that takes the features from earlier layers as input and outputs the localization residual. Our proposed detector can achieve fast detection at a rate of 28 frames/s, and a mean average precision (mAP) of 67.25% in the DETRAC vehicle detection benchmark. Chen-Hang He, Kin-Man Lam 0001 |
ICASSP | 2 |
| 2018 | Can a machine have two systems for recognition, like human beings?
Jiwei Hu, Kin-Man Lam 0001, Ping Lou, Quan Liu 0001, Wupeng Deng |
J. Vis. Commun. Image Represent. | 2 |
| 2018 | Integrating QDWD with pattern distinctness and local contrast for underwater saliency detection
Muwei Jian, Qiang Qi, Junyu Dong, Yilong Yin, Kin-Man Lam 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2018 | Saliency detection based on background seeds by object proposals and extended random walk
Muwei Jian, Runxia Zhao, Xin Sun 0003, Hanjiang Luo, Wenyin Zhang, Huaxiang Zhang 0001, Junyu Dong, Yilong Yin, Kin-Man Lam 0001 |
J. Vis. Commun. Image Represent. | 9 |
| 2018 | Histogram-based local descriptors for facial expression recognition (FER): A comprehensive study
Cigdem Turan, Kin-Man Lam 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2018 | Saliency detection using quaternionic distance based weber local descriptor and level priors
Muwei Jian, Qiang Qi, Junyu Dong, Xin Sun 0003, Yujuan Sun, Kin-Man Lam 0001 |
Multim. Tools Appl. | 6 |
| 2018 | Content-based image retrieval via a hierarchical-local-feature extraction scheme
Muwei Jian, Yilong Yin, Junyu Dong, Kin-Man Lam 0001 |
Multim. Tools Appl. | 4 |
| 2018 | Age-invariant face recognition based on identity inference from appearance ageabstractFace recognition across age progression remains one of the area's most challenging tasks, as the aging process affects both the shape and texture of a face. One possible solution is to apply a probabilistic model to represent a face simultaneously with its identity variable, which is stable through time, and its aging variable, which changes with time. However, as the aging process varies for different people, a person may look younger or older than another person, even though their ages are the same. Consequently, using the ‘real’ age labels given by existing face datasets for age-invariant face recognition will inevitably introduce ambiguity to learning algorithms. In this paper, an identity-inference model, based on age-subspace learning from appearance-age labels, is proposed. We first model human identity and aging variables simultaneously using Probabilistic Linear Discriminant Analysis (PLDA). Then, the aging subspace is learnt independently with the appearance-age labels, and the identity subspace is then determined iteratively with the Expectation-Maximization (EM) algorithm. We found that the learned aging subspace is insensitive to the training face images used, and is independent of the identity model. Consequently, the recognition of aging faces becomes simpler as identity inference no longer needs to consider age labels. Furthermore, in our algorithm, different identity features learnt from the identity model are further combined using Canonical Correlation Analysis (CCA), where their correlations are maximized for face recognition. A thorough experimental analysis of face recognition is performed on three public domain face-aging datasets: FGNET, MORPH, and CACD. Experiment results show that the proposed framework can achieve a comparable, or even better, performance against other state-of-the-art methods, especially when the age range is large. Huiling Zhou, Kin-Man Lam 0001 |
Pattern Recognit. | 2 |
| 2018 | A Joint Framework for QoS and QoE for Video Transmission over Wireless Multimedia Sensor NetworksabstractWith the emergence of Wireless Multimedia Sensor Networks (WMSNs), the distribution of multimedia contents have now become a reality. Without proper management, the transmission of multimedia data over WMSNs affects the performance of networks due to excessive packet-drop. The existing studies on Quality of Service (QoS) mostly deal with simple Wireless Sensor Networks (WSNs) and as such do not account for an increasing number of sensor nodes and an increasing volume of data. In this paper, we propose a novel framework to support QoS in WMSNs along with a light-weight Error Concealment (EC) scheme. The EC schemes play a vital role to enhance Quality of Experience (QoE) by maintaining an acceptable quality at the receiving ends. The main objectives of the proposed framework are to maximize the network throughput and to cover-up the effects produced by dropped video packets. To control the data-rate, Scalable High efficiency Video Coding (SHVC) is applied at multimedia sensor nodes with variable Quantization Parameters (QPs). Multi-path routing is exploited to support real-time video transmission. Experimental results show that the proposed framework can efficiently adjust large volumes of video data under certain network distortions and can effectively conceal lost video frames by producing better objective measurements. Muhammad Usman 0015, Ning Yang 0003, Mian Ahmad Jan, Xiangjian He, Min Xu 0001, Kin-Man Lam 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2017 | Constructing a hierarchical tree for image annotationabstractImage annotation is always an easy task for humans but a tough task for machines. Inspired by human's thinking mode, there is an assumption that the computer has double systems. Each of the systems can handle the task individually and in parallel. In this paper, we introduce a new hierarchical model for image annotation, based on constructing a novel, hierarchical tree, which consists of exploring the relationships between the labels and the features used, and dividing labels into several hierarchies for efficient and accurate labeling. Jiwei Hu, Kin-Man Lam 0001, Ping Lou, Quan Liu 0001 |
ICME | 2 |
| 2017 | The OUC-vision large-scale underwater image databaseabstractIn this paper, a large-scale underwater image database for underwater salient object detection or saliency detection is presented in detail. This database is called the OUC-VISION underwater image database, which contains 4400 underwater images of 220 individual objects. Each object is captured with four pose variations (the frontal-, the opposite-, the left-, and the right-views of each underwater object) and five spatial locations (the underwater object is located at the top-left corner, the top-right corner, the center, the bottom-left corner, and the bottom-right corner) to obtain 20 images. Meanwhile, this publicly available OUC-VISION database also provides relevant industrial fields, and academic researchers with underwater images under different sources of variations, especially pose, spatial location, illumination, turbidity of water, etc. Ground-truth information is also manually labelled for this database. The OUC-VISION database can not only be widely used to assess and evaluate the performance of the state-of-the-art salient-object detection and saliency-detection algorithms for general images, but also will particularly benefit the development of underwater vision technology in the future. Muwei Jian, Qiang Qi, Junyu Dong, Yinlong Yin, Wenyin Zhang, Kin-Man Lam 0001 |
ICME | 6 |
| 2017 | Efficient likelihood Bayesian constrained local modelabstractThe constrained local model (CLM) proposes a paradigm that the locations of a set of local landmark detectors are constrained to lie in a subspace, spanned by a shape point distribution model (PDM). Fitting the model to an object involves two steps. A response map, which represents the likelihood of locations for a landmark, is first computed for each landmark using local-texture detectors. Then, an optimal PDM is determined by jointly maximizing all the response maps simultaneously, with a global-shape constraint. This global optimization can be considered a Bayesian inference problem, where the posterior distribution of the shape parameters, as well as the pose parameters, can be inferred using maximum a posteriori (MAP). In this paper, based on the CLM model, we present a novel CLM variant, which employs random-forest regressors to estimate the location of each landmark, as a likelihood term, efficiently. This novel CLM framework is called efficient likelihood Bayesian constrained local model (elBCLM). Furthermore, in each stage of the regressors, the PDM local non-rigid parameters, i.e. the shape parameters, of the previous stage can work as shape clues for training the regressors for the current stage. To further improve the efficiency, we also propose a feature-switching scheme used in the cascaded framework. Experimental results on benchmark datasets show our approach achieves about 3 to 5 times speed-up, when compared with the existing CLM models, and improves by around 10% on fitting accuracy, when compared with the other regression-based models. Kin-Man Lam 0001, Man-Yau Chiu, Kangheng Wu, Zhibin Lei |
ICME | 2 |
| 2017 | A joint deep-network-based image restoration algorithm for multi-degradationsabstractIn the procedures of image acquisition, compression, and transmission, captured images usually suffer from various degradations, such as low-resolution and compression distortion. Although there have been a lot of research done on image restoration, they usually aim to deal with a single degraded factor, ignoring the correlation of different degradations. To establish a restoration framework for multiple degradations, a joint deep-network-based image restoration algorithm is proposed in this paper. The proposed convolutional neural network is composed of two stages. Firstly, a de-blocking subnet is constructed, using two cascaded neural network. Then, super-resolution is carried out by a 20-layer very deep network with skipping links. Cascading these two stages forms a novel deep network. Experimental results on the Set5, Setl4 and BSD100 benchmarks demonstrate that the proposed method can achieve better results, in terms of both the subjective and objective performances. Li Zhuo 0001, Kin-Man Lam 0001, Jiafeng Li 0001 |
ICME | 4 |
| 2017 | An output-based knowledge transfer approach and its application in bladder cancer predictionabstractMany medical applications face a situation that the on-hand data cannot fully fit an existing predictive model or on-line tool, since these models or tools only use the most common predictors and the other valuable features collected in the current scenario are not considered altogether. On the other hand, the training data in the current scenario is not sufficient to learn a predictive model effectively yet. In order to overcome these problems and construct an efficient classifier, for these real situations in medical fields, in this work we present an approach based on the least squares support vector machine (LS-SVM), which utilizes a transfer learning framework to make maximum use of the data and guarantee its enhanced generalization capability. The proposed approach is capable of effectively learning a target domain with limited samples by relying on the probabilistic outputs from the other previously learned model using a heterogeneous method in the source domain. Moreover, it autonomously and quickly decides how much output knowledge to transfer from source domain to the target one using a fast leave-one-out cross validation strategy. This approach is applied on a real-world clinical dataset to predict 5-year mortality of bladder cancer patients after radical cystectomy, and the experimental results indicate that the proposed method can achieve better performances compared to traditional machine learning methods, consistently showing the potential of the proposed method under the circumstances with insufficient data. Guanjin Wang, Guangquan Zhang 0001, Kup-Sze Choi, Kin-Man Lam 0001, Jie Lu 0001 |
IJCNN | 4 |
| 2017 | A novel correspondence-based face-hallucination method
Zhuo Hui, Kin-Man Lam 0001 |
Image Vis. Comput. | 3 |
| 2017 | Person re-identification by unsupervised video matching
Xiatian Zhu, Shaogang Gong, Xudong Xie, Jianming Hu, Kin-Man Lam 0001, Yisheng Zhong |
Pattern Recognit. | 6 |
| 2016 | Shape-appearance-correlated active appearance modelabstractAmong the challenges faced by current active shape or appearance models, facial-feature localization in the wild, with occlusion in a novel face image, i.e. in a generic environment, is regarded as one of the most difficult computer-vision tasks. In this paper, we propose an Active Appearance Model (AAM) to tackle the problem of generic environment. Firstly, a fast face-model initialization scheme is proposed, based on the idea that the local appearance of feature points can be accurately approximated with locality constraints. Nearest neighbors, which have similar poses and textures to a test face, are retrieved from a training set for constructing the initial face model . To further improve the fitting of the initial model to the test face, an orthogonal CCA (oCCA) is employed to increase the correlation between shape features and appearance features represented by Principal Component Analysis (PCA). With these two contributions, we propose a novel AAM, namely the shape-appearance-correlated AAM (SAC-AAM), and the optimization is solved by using the recently proposed fast simultaneous inverse compositional (Fast-SIC) algorithm. Experiment results demonstrate a 5–10% improvement on controlled and semi-controlled datasets, and with around 10% improvement on wild face datasets in terms of fitting accuracy compared to other state-of-the-art AAM models. Huiling Zhou, Kin-Man Lam 0001, Xiangjian He |
Pattern Recognit. | 2 |
| 2016 | A Level Set Approach to Image Segmentation With Intensity InhomogeneityabstractIt is often a difficult task to accurately segment images with intensity inhomogeneity, because most of representative algorithms are region-based that depend on intensity homogeneity of the interested object. In this paper, we present a novel level set method for image segmentation in the presence of intensity inhomogeneity. The inhomogeneous objects are modeled as Gaussian distributions of different means and variances in which a sliding window is used to map the original image into another domain, where the intensity distribution of each object is still Gaussian but better separated. The means of the Gaussian distributions in the transformed domain can be adaptively estimated by multiplying a bias field with the original signal within the window. A maximum likelihood energy functional is then defined on the whole image region, which combines the bias field, the level set function, and the piecewise constant function approximating the true image signal. The proposed level set method can be directly applied to simultaneous segmentation and bias correction for 3 and 7T magnetic resonance images. Extensive evaluation on synthetic and real-images demonstrate the superiority of the proposed method over other representative algorithms. Kaihua Zhang 0001, Lei Zhang 0006, Kin-Man Lam 0001, David Zhang 0001 |
IEEE Trans. Cybern. | 3 |
| 2016 | Frame Interpolation for Cloud-Based Mobile Video StreamingabstractCloud-based High Definition (HD) video streaming is becoming popular day by day. On one hand, it is important for both end users and large storage servers to store their huge amount of data at different locations and servers. On the other hand, it is becoming a big challenge for network service providers to provide reliable connectivity to the network users. There have been many studies over cloud-based video streaming for Quality of Experience (QoE) for services like YouTube. Packet losses and bit errors are very common in transmission networks, which affect the user feedback over cloud-based media services. To cover up packet losses and bit errors, Error Concealment (EC) techniques are usually applied at the decoder/receiver side to estimate the lost information. This paper proposes a time-efficient and quality-oriented EC method. The proposed method considers H.265/HEVC based intra-encoded videos for the estimation of whole intra-frame loss. The main emphasis in the proposed approach is the recovery of Motion Vectors (MVs) of a lost frame in real-time. To boost-up the search process for the lost MVs, a bigger block size and searching in parallel are both considered. The simulation results clearly show that our proposed method outperforms the traditional Block Matching Algorithm (BMA) by approximately 2.5 dB and Frame Copy (FC) by up to 12 dB at a packet loss rate of 1%, 3%, and 5% with different Quantization Parameters (QPs). The computational time of the proposed approach outperforms the BMA by approximately 1788 seconds. Muhammad Usman 0015, Xiangjian He, Kin-Man Lam 0001, Min Xu 0001, Syed Mohsin Matloob Bokhari, Jinjun Chen |
IEEE Trans. Multim. | 3 |
| 2015 | Survey of Error Concealment techniques: Research directions and open issuesabstractError Concealment (EC) techniques use either spatial, temporal or a combination of both types of information to recover the data lost in transmitted video. In this paper, existing EC techniques are reviewed, which are divided into three categories, namely Intra-frame EC, Inter-frame EC, and Hybrid EC techniques. We first focus on the EC techniques developed for the H.264/AVC standard. The advantages and disadvantages of these EC techniques are summarized with respect to the features in H.264. Then, the EC algorithms are also analyzed. These EC algorithms have been recently adopted in the newly introduced H.265/HEVC standard. A performance comparison between the classic EC techniques developed for H.264 and H.265 is performed in terms of the average PSNR. Lastly, open issues in the EC domain are addressed for future research consideration. Muhammad Usman 0015, Xiangjian He, Min Xu 0001, Kin-Man Lam 0001 |
PCS | 4 |
| 2015 | Saliency detection based on singular value decomposition
Xudong Xie, Kin-Man Lam 0001, Jianming Hu, Yisheng Zhong |
J. Vis. Commun. Image Represent. | 3 |
| 2015 | Efficient saliency analysis based on wavelet transform and entropy theory
Xudong Xie, Kin-Man Lam 0001, Yisheng Zhong |
J. Vis. Commun. Image Represent. | 3 |
| 2015 | Design and learn distinctive features from pore-scale facial keypoints
Dong Li 0028, Kin-Man Lam 0001 |
Pattern Recognit. | 2 |
| 2015 | Simultaneous Hallucination and Recognition of Low-Resolution Faces Based on Singular Value DecompositionabstractIn video surveillance, the captured face images are usually of low resolution (LR). Thus, a framework based on singular value decomposition (SVD) for performing both face hallucination and recognition simultaneously is proposed in this paper. Conventionally, LR face recognition is carried out by super-resolving the LR input face first, and then performing face recognition to identify the input face. By considering face hallucination and recognition simultaneously, the accuracy of both the hallucination and the recognition can be improved. In this paper, singular values are first proved to be effective for representing face images, and the singular values of a face image at different resolutions have approximately a linear relation. In our algorithm, each face image is represented using SVD. For each LR input face, the corresponding LR and high-resolution (HR) face-image pairs can then be selected from the face gallery. Based on these selected LR-HR pairs, the mapping functions for interpolating the two matrices in the SVD representation for the reconstruction of HR face images can be learned more accurately. Therefore, the final estimation of the high-frequency details of the HR face images will become more reliable and effective. The experimental results demonstrate that our proposed framework can achieve promising results for both face hallucination and recognition. Muwei Jian, Kin-Man Lam 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | Visual-Patch-Attention-Aware Saliency DetectionabstractThe human visual system (HVS) can reliably perceive salient objects in an image, but, it remains a challenge to computationally model the process of detecting salient objects without prior knowledge of the image contents. This paper proposes a visual-attention-aware model to mimic the HVS for salient-object detection. The informative and directional patches can be seen as visual stimuli, and used as neuronal cues for humans to interpret and detect salient objects. In order to simulate this process, two typical patches are extracted individually and in parallel from the intensity channel and the discriminant color channel, respectively, as the primitives. In our algorithm, an improved wavelet-based salient-patch detector is used to extract the visually informative patches. In addition, as humans are sensitive to orientation features, and as directional patches are reliable cues, we also propose a method for extracting directional patches. These two different types of patches are then combined to form the most important patches, which are called preferential patches and are considered as the visual stimuli applied to the HVS for salient-object detection. Compared with the state-of-the-art methods for salient-object detection, experimental results using publicly available datasets show that our produced algorithm is reliable and effective. Muwei Jian, Kin-Man Lam 0001, Junyu Dong, LinLin Shen |
IEEE Trans. Cybern. | 2 |
| 2015 | High-Resolution Face Verification Using Pore-Scale Facial FeaturesabstractFace recognition methods, which usually represent face images using holistic or local facial features, rely heavily on alignment. Their performances also suffer a severe degradation under variations in expressions or poses, especially when there is one gallery per subject only. With the easy access to high-resolution (HR) face images nowadays, some HR face databases have recently been developed. However, few studies have tackled the use of HR information for face recognition or verification. In this paper, we propose a pose-invariant face-verification method, which is robust to alignment errors, using the HR information based on pore-scale facial features. A new keypoint descriptor, namely, pore-Principal Component Analysis (PCA)-Scale Invariant Feature Transform (PPCASIFT)-adapted from PCA-SIFT-is devised for the extraction of a compact set of distinctive pore-scale facial features. Having matched the pore-scale features of two-face regions, an effective robust-fitting scheme is proposed for the face-verification task. Experiments show that, with one frontal-view gallery only per subject, our proposed method outperforms a number of standard verification methods, and can achieve excellent accuracy even the faces are under large variations in expression and pose. Dong Li 0028, Huiling Zhou, Kin-Man Lam 0001 |
IEEE Trans. Image Process. | 3 |
| 2014 | Segmentation-enhanced saliency detection model based on distance transform and center biasabstractSaliency detection is one of the extraordinary abilities of the human visual system (HVS), and also provides a powerful tool for predicting where humans tend to focus in the free-viewing process. In this paper, we propose a novel method for computing image saliency. At first, an image is subject to L0smoothing to characterize its fundamental constituents while diminishing insignificant details. Distance-transform-based saliency detection is then applied to the smoothed image, to extract the general salient regions and form a rough saliency map. Next, the segmentation information generated by normalized cuts is used to improve the saliency detection performance by averaging the saliency values in each segmented block. Finally, we employ the center-bias mechanism to further improve the saliency model. The proposed method is compared with six existing saliency models, and achieves the best performance in terms of the area under the ROC curve (AUC). Hongyun Gao 0001, Kin-Man Lam 0001 |
ICASSP | 2 |
| 2014 | From quaternion to octonion: Feature-based image saliency detectionabstractA novel computational model for detecting salient regions in color images is proposed by utilizing early visual features and performing spectral normalization in the octonion algebra framework, which can accommodate more feature channels than quaternions can. Firstly, feature maps based on edge intensity, the black-white, red-green, and blue-yellow color opponents, as well as the Gabor features with four directions, are incorporated into the eight channels of the octonion image. Then, spectral normalization is achieved by preserving the phase information of the octonion image. Finally, saliency maps are generated at different scales using Gaussian pyramids, and are combined to form the final saliency map. The integration of frequency normalization into the octonion image and saliency-map pyramids exploits the benefits from both the spectral domain and the spatial domain. Experimental results on the MSRA dataset demonstrate that our proposed method outperforms five existing saliency detection models. Hongyun Gao 0001, Kin-Man Lam 0001 |
ICASSP | 2 |
| 2014 | Salient object detection using octonion with Bayesian inferenceabstractA novel computational model for detecting salient regions in color images is proposed, based on a two-stage coarse-to-fine framework. Firstly, different early visual feature maps - including the edge intensity; the black-white, red-green, and blue-yellow color opponents; and the Gabor features with four directions - are incorporated into the eight channels of an octonion image. Spectral normalization is achieved with the octonion Fourier transform by preserving the phase information of the octonion image. Then, with mean-shift segmentation, the saliency values in each segment are averaged to form a coarse saliency map. Finally, the coarse saliency map is subject to Bayesian inference to further refine the salient regions. The integration of frequency normalization, spatial segmentation and Bayesian inference exploits the benefits from both the spectral domain and the spatial domain. Experimental results show the superiority of the proposed method compared to several existing methods. Hongyun Gao 0001, Kin-Man Lam 0001 |
ICIP | 2 |
| 2014 | Salient-region detection in a multi-level framework of image smoothing with over-segmentationabstractSaliency detection is one of the extraordinary abilities of the human visual system; it also provides a powerful tool for predicting where people tend to focus in the free-viewing process. In this paper, we propose a novel salient-object detection method which applies an over-segmentation-based saliency detection algorithm to multi-level smoothed images. The original image is initially subjected to smoothing based on multi-level L0gradient minimization; this can characterize its fundamental constituents while diminishing the insignificant details. Then, segment-based saliency computation is applied to the multi-level smoothed images to produce a series of intermediate saliency maps. The final saliency map is generated by combining the intermediate saliency maps. The proposed method is compared with six existing saliency models, and achieves the best performance in terms of Precision, Recall and F-measure, as well as in terms of the area under the ROC curve (AUC). Hongyun Gao 0001, Kin-Man Lam 0001 |
ICIP | 2 |
| 2014 | Region-based feature fusion for facial-expression recognitionabstractIn this paper, we propose a feature-fusion method based on Canonical Correlation Analysis (CCA) for facial-expression recognition. In our proposed method, features from the eye and the mouth windows are extracted separately, which are correlated with each other in representing a facial expression. For each of the windows, two effective features, namely the Local Phase Quantization (LPQ) and the Pyramid of Histogram of Oriented Gradients (PHOG) descriptors, are employed to form low-level representations of the corresponding windows. The features are then represented in a coherent subspace by using CCA in order to maximize the correlation. In our experiments, the Extended Cohn-Kanade dataset is used; its face images span seven different emotions, namely anger, contempt, disgust, fear, happiness, sadness, and surprise. Experiment results show that our method can achieve excellent accuracy for facial-expression recognition. Cigdem Turan, Kin-Man Lam 0001 |
ICIP | 2 |
| 2014 | An Adaptive-Profile Active Shape Model for Facial-Feature DetectionabstractIn this paper, a novel algorithm based on the Active Shape Model (ASM) for locating landmarks on human faces is proposed. A challenge for detecting facial features is that faces may be under different poses, this makes the local appearance of each facial landmark vary greatly. To account for these variations, we propose an adaptive-profile scheme for ASM so that facial landmarks can be detected reliably and accurately under different poses. In our algorithm, a 2D profile is used for each landmark, and the 2D profiles of each landmark of the training face images are grouped to form a number of clusters. The corresponding shape vector for each of the clusters is then learned. For a query face image, the profiles to be used to locate the respective facial landmarks will be selected according to the face-shape vector in the current iteration. In other words, adaptive profiles are used in the search for landmarks. Face images from two subsets of the IMM Face Database are used for training, and the other two subsets are used for testing. The performance of our proposed algorithm is also evaluated using another dataset, namely the Bosphorus Dataset. Experiment results show that our proposed Adaptive-Profile Active Shape Model (APASM) can locate facial landmarks accurately under different face shapes, expressions, and poses. Huiling Zhou, Kin-Man Lam 0001 |
ICPR | 3 |
| 2014 | Saliency detection based on adaptive DoG and distance transformabstractA novel computational model for detecting salient regions in color images is proposed, based on adaptive difference of Gaussian (DoG) filtering and distance transform. In our method, we first transform an image into the frequency domain, and perform adaptive DoG filtering, whose parameters are determined by the energy spectrum of the image. Then, the edge information is extracted from the DoG filtering output, and the distance transform is applied to the edge map. Finally, the Gaussian pyramids are used to enhance the distance transform performance. Our proposed method achieves spectral domain filtering as well as spatial domain edge extraction, thus exploiting the benefits from both the spatial domain and the spectral domain for saliency detection. We compare our proposed method with five existing saliency detection methods in terms of precision, recall, and F-measure. Experiments on the MSRA dataset show the outperformance of the proposed method over those saliency algorithms. Hongyun Gao 0001, Kin-Man Lam 0001 |
ISCAS | 2 |
| 2014 | Using artificial neural network to predict mortality of radical cystectomy for bladder cancerabstractSurgical removal of bladder, i.e. radical cystectomy, is a standard treatment option for muscle invasive bladder cancer. Unfortunately, the treatment is associated with significant morbidities and mortalities. Many studies have been conducted to predict the morbidities and mortalities of radical cystectomy based on statistical analysis. In this paper, an artificial neural network is employed to predict 5-year mortality of radical cystectomy. The clinico-pathological data from a urology unit of a district hospital in Hong Kong were used to train and test the model. The outcome of the surgery was computed by an artificial neural network based on the risk factors identified by a conventional statistical method. It was found that the best overall accuracy of the neural network model was 77.8% and the 5-year mortality predicted by the model was comparable to that achieved by conventional statistical methods. The results of this study reflect that artificial intelligence has great development potential in medicine. Kin-Man Lam 0001, Xuejian He, Kup-Sze Choi |
SMARTCOMP | 1 |
| 2014 | Facial-feature detection and localization based on a hierarchical scheme
Muwei Jian, Kin-Man Lam 0001, Junyu Dong |
Inf. Sci. | 2 |
| 2014 | Illumination-insensitive texture discrimination based on illumination compensation and enhancement
Muwei Jian, Kin-Man Lam 0001, Junyu Dong |
Inf. Sci. | 2 |
| 2014 | Image classification without segmentation using a hybrid pyramid kernel
Wai-Shing Cho, Kin-Man Lam 0001 |
Multim. Tools Appl. | 2 |
| 2014 | Face hallucination based on sparse local-pixel structureabstractIn this paper, we propose a face-hallucination method, namely face hallucination based on sparse local-pixel structure. In our framework, a high resolution (HR) face is estimated from a single frame low resolution (LR) face with the help of the facial dataset. Unlike many existing face-hallucination methods such as the from local-pixel structure to global image super-resolution method (LPS-GIS) and the super-resolution through neighbor embedding, where the prior models are learned by employing the least-square methods, our framework aims to shape the prior model using sparse representation. Then this learned prior model is employed to guide the reconstruction process. Experiments show that our framework is very flexible, and achieves a competitive or even superior performance in terms of both reconstruction error and visual quality. Our method still exhibits an impressive ability to generate plausible HR facial images based on their sparse local structures. Cheng Cai, Guoping Qiu, Kin-Man Lam 0001 |
Pattern Recognit. | 4 |
| 2014 | Multi-resolution feature fusion for face recognition
Kuong-Hon Pong, Kin-Man Lam 0001 |
Pattern Recognit. | 2 |
| 2014 | A biased selection strategy for information recycling in Boosting cascade visual-object detectors
Chensheng Sun, Jiwei Hu, Kin-Man Lam 0001 |
Pattern Recognit. Lett. | 3 |
| 2014 | Face-image retrieval based on singular values and potential-field representation
Muwei Jian, Kin-Man Lam 0001 |
Signal Process. | 2 |
| 2013 | Color facial image denoising based on rpca and noisy pixel detectionabstractIn this paper, a novel approach for color facial-image denoising based on robust principal component analysis (RPCA) [1] in the L*a*b* color space and noisy pixel detection is proposed. Firstly, RPCA is employed for color facial-image recovery in the L*a*b* space. Then, the reconstructed image is used for noisy pixel detection. Finally, the denoised facial-image can be obtained. Experiments are conducted based on the AR database, where our proposed method is compared with several state-of-the-art image-denoising methods. Experimental results show that our method can achieve a better performance in terms of both quantitatively evaluation and visual quality. Zhaojun Yuan, Xudong Xie, Kin-Man Lam 0001 |
ICASSP | 4 |
| 2013 | Multi-view face hallucination based on sparse representationabstractIn this paper, we propose a novel method to generate the hallucinated multi-views of faces using the sparse-representation model. In order to render a faithful virtual view, we introduce centralized constraints into a variation framework for optimization. The constraints are formulated based on an attempt to minimize the difference between the sparse-coding coefficients derived for two distinct views. In our algorithm, sift optical-flow method is employed to formulate the constraints. An input face is firstly sparsely coded over a given dictionary, and then the sparse-coding coefficients for the input face are refined through an optimization framework with the centralized constraints. Intensive experimental results demonstrate that our proposed method can perform well in terms of both reconstruction accuracy and visual quality. Hui Zhuo, Kin-Man Lam 0001 |
ICASSP | 2 |
| 2013 | A Mass Spectra-Based Compound-Identification Approach with a Reduced Reference Library
Kin-Man Lam 0001 |
ICIC (2) | 2 |
| 2013 | A Novel Ensemble Algorithm for Tumor Classification
Han Wang 0001, Wai-Shing Lau, Gerald Seet, Danwei Wang, Kin-Man Lam 0001 |
ISNN (2) | 6 |
| 2013 | An efficient two-stage framework for image annotation
Jiwei Hu, Kin-Man Lam 0001 |
Pattern Recognit. | 2 |
| 2013 | A novel face-hallucination scheme based on singular value decomposition
Muwei Jian, Kin-Man Lam 0001, Junyu Dong |
Pattern Recognit. | 2 |
| 2013 | Multiple-Kernel, Multiple-Instance Similarity Features for Efficient Visual Object DetectionabstractWe propose to use the similarity between the sample instance and a number of exemplars as features in visual object detection. Concepts from multiple-kernel learning and multiple-instance learning are incorporated into our scheme at the feature level by properly calculating the similarity. The similarity between two instances can be measured by various metrics and by using the information from various sources, which mimics the use of multiple kernels for kernel machines. Pooling of the similarity values from multiple instances of an object part is introduced to cope with alignment inaccuracy between object instances. To deal with the high dimensionality of the multiple-kernel multiple-instance similarity feature, we propose a forward feature-selection technique and a coarse-to-fine learning scheme to find a set of good exemplars, hence we can produce an efficient classifier while maintaining a good performance. Both the feature and the learning technique have interesting properties. We demonstrate the performance of our method using both synthetic data and real-world visual object detection data sets. Chensheng Sun, Kin-Man Lam 0001 |
IEEE Trans. Image Process. | 2 |
| 2013 | Depth Estimation of Face Images Using the Nonlinear Least-Squares ModelabstractIn this paper, we propose an efficient algorithm to reconstruct the 3D structure of a human face from one or more of its 2D images with different poses. In our algorithm, the nonlinear least-squares model is first employed to estimate the depth values of facial feature points and the pose of the 2D face image concerned by means of the similarity transform. Furthermore, different optimization schemes are presented with regard to the accuracy levels and the training time required. Our algorithm also embeds the symmetrical property of the human face into the optimization procedure, in order to alleviate the sensitivities arising from changes in pose. In addition, the regularization term, based on linear correlation, is added in the objective function to improve the estimation accuracy of the 3D structure. Further, a model-integration method is proposed to improve the depth-estimation accuracy when multiple nonfrontal-view face images are available. Experimental results on the 2D and 3D databases demonstrate the feasibility and efficiency of the proposed methods. Kin-Man Lam 0001, Qingwei Gao |
IEEE Trans. Image Process. | 2 |
| 2012 | Totally-corrective boosting using continuous-valued weak learnersabstractThe Boosting algorithm has two main variants: the gradient Boosting and the totally-corrective column-generation Boosting. Recently, the latter has received increasing attention since it exhibits a better convergence property, thus resulting in more efficient strong learners. In this work, we point out that the totally-corrective column-generation Boosting is equivalent to the gradient-descent method for the gradient Boosting in the weak-learner selection criterion, but uses additional totally-corrective updates for the weak-learner weights. Therefore, other techniques for the gradient Boosting that produce continuous-valued weak learners, e.g. step-wise direct minimization and Newtons method, may also be used in combination with the totally-corrective procedure. In this work we take the well known AdaBoost algorithm as an example, and show that employing the continuous-valued weak learners improves the performance when used with the totally-corrective weak-learner weight update. Chensheng Sun, Sanyuan Zhao, Jiwei Hu, Kin-Man Lam 0001 |
ICASSP | 4 |
| 2012 | An efficient local-structure-based face-hallucination methodabstractIn this paper, we propose a novel patch-based face-hallucination algorithm, which is based on the local structure kernels established via the relation between interpolated low-resolution (LR) images and their corresponding high-resolution counterparts. In our algorithm, the local linear embedding (LLE) algorithm is used to extract local structures, and the kernels are then constructed based on non-overlapped patches in the interpolated LR images. The information about local structures as described by the kernels is propagated to the corresponding regions of the HR images. Sub-pixel distortions are refined by solving a constrained problem at pixel level via iterative procedures. Experimental results show that our proposed method can provide a good performance in terms of reconstruction errors and visual quality. Hui Zhuo, Kin-Man Lam 0001 |
ICASSP | 2 |
| 2012 | Combination of global and local baseline-independent features for offline Arabic handwriting recognition
Xudong Xie, Kin-Man Lam 0001 |
ICPR | 4 |
| 2012 | An efficient method for occluded face recognition
Xudong Xie, Kin-Man Lam 0001 |
ICPR | 3 |
| 2012 | Dynamic textures indexing and retrieval based on intrinsic propertiesabstractThe indexing and retrieval of dynamic textures remains one of the most challenging tasks in both computer vision and computer graphics. This paper first introduces the concept of intrinsic properties of dynamic textures, including Static Texture Features (STF) and Dynamic Texture Features (DTF), and then proposes an efficient method that utilizes these intrinsic properties for the indexing and retrieval of dynamic textures. Our proposed method is evaluated and compared to existing methods based on a variety of video texture samples. Experimental results show that our scheme can produce promising results. Muwei Jian, Kin-Man Lam 0001, Junyu Dong |
ISCAS | 2 |
| 2012 | Eigentransformation-based face super-resolution in the wavelet domain
Hui Zhuo, Kin-Man Lam 0001 |
Pattern Recognit. Lett. | 2 |
| 2011 | Saliency Modulated High Dynamic Range Image Tone MappingabstractThis paper presents a new high dynamic range image tone mapping technique - saliency modulated tone mapping (SMTM). The HDR image is not directly viewable and dynamic range compression will unavoidably loose information. A saliency map analyzes the visual importance of the regions and can therefore direct the tone mapping operators to preserve the visual conspicuity of the regions that should more likely attract visual attention. In SMTM, we have developed a very fast algorithm to first compute the visual saliency map of the high dynamic range radiance map and then directly use the saliency of the local regions to control the local tone mapping curve such that highly salient regions will have their details and contrast better protected so as to remain salient and attract visual attention in the tone mapped display. We present experimental results to show that SMTM provides competitive performances to state of the art tone mapping techniques in rending visually pleasing low dynamic range displays. We also show that SMTM is better able to preserve the visual saliency of the HDR image and that SMTM renders high saliency regions to stand out to attract observers attention. Yujie Mei, Guoping Qiu, Kin-Man Lam 0001 |
ICIG | 3 |
| 2011 | Color correction via robust reference selection and recovery using a low-rank matrix modelabstractIn this paper, we propose a method that can handle the color correction of a large collection of photographs simultaneously and automatically via robust reference selection. The method does not use any particular model to handle the errors on the photographs, but corrects all kinds of errors caused by changes of viewpoint, large illumination variations, gross pixel corruptions, and partial occlusions under a low-rank matrix model. Furthermore, our method uses the image pixel values directly in vector form, which preserves the spatial information, to obtain the matrix for color correction, unlike other statistics-based image-representation methods such as color histograms. Experiments verify that our method can achieve consistent and promising results on uncontrolled real photographs acquired from the Internet. Dong Li 0028, Xudong Xie, Kin-Man Lam 0001 |
ICIP | 3 |
| 2011 | Feature subset selection for efficient AdaBoost trainingabstractWorking with a very large feature set is a challenge in the current machine learning research. In this paper, we address the feature-selection problem in the context of training AdaBoost classifiers. The AdaBoost algorithm embeds a feature selection mechanism based on training a classifier for each feature. Learning the single-feature classifiers is the most time consuming part of AdaBoost training, especially when large number of features are available. To solve this problem, we generate a working feature subset using a novel feature subset selection method based on the partial least square regression, and then train and select from this feature subset. The partial least square method is capable of selecting high-dimensional and highly redundant features. The experiments show that the proposed PLS-based feature-selection method generates sensible feature subsets for AdaBoost in a very efficient way. Chensheng Sun, Jiwei Hu, Kin-Man Lam 0001 |
ICME | 3 |
| 2011 | A novel kernel-based framework for facial-image hallucination
Kin-Man Lam 0001, Tingzhi Shen, Weijiang Wang |
Image Vis. Comput. | 2 |
| 2011 | Video-object segmentation and 3D-trajectory estimation for monocular video sequences
Feng Xu 0005, Kin-Man Lam 0001, Qionghai Dai |
Image Vis. Comput. | 2 |
| 2011 | Depth Estimation of Face Images Based on the Constrained ICA ModelabstractIn this paper, we propose a novel and efficient algorithm to reconstruct the 3-D structure of a human face from one or a number of its 2-D images with different poses. In our proposed algorithm, the rotation and translation process from a frontal-view face image to a nonfrontal-view face image is at first formulated as a constrained independent component analysis (cICA) model. Then, the overcomplete ICA problem is converted into a normal ICA problem by incorporating a prior from the CANDIDE 3-D face model. Furthermore, the CANDIDE model is employed to construct a reference signal that is used in both the initialization and the objective function of the cICA model. Moreover, a model-integration method is proposed to improve the depth-estimation accuracy when multiple nonfrontal-view face images are available. An important advantage of the proposed algorithm is that no frontal-view face image is required for the estimation of the corresponding 3-D face structure. Experimental results on a real 3-D face image database demonstrate the feasibility and efficiency of the proposed method. Kin-Man Lam 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2011 | From Local Pixel Structure to Global Image Super-Resolution: A New Face Hallucination FrameworkabstractWe have developed a new face hallucination framework termed from local pixel structure to global image super-resolution (LPS-GIS). Based on the assumption that two similar face images should have similar local pixel structures, the new framework first uses the input low-resolution (LR) face image to search a face database for similar example high-resolution (HR) faces in order to learn the local pixel structures for the target HR face. It then uses the input LR face and the learned pixel structures as priors to estimate the target HR face. We present a three-step implementation procedure for the framework. Step 1 searches the database for K example faces that are the most similar to the input, and then warps the K example images to the input using optical flow. Step 2 uses the warped HR version of the K example faces to learn the local pixel structures for the target HR face. An effective method for learning local pixel structures from an individual face, and an adaptive procedure for fusing the local pixel structures of different example faces to reduce the influence of warping errors, have been developed. Step 3 estimates the target HR face by solving a constrained optimization problem by means of an iterative procedure. Experimental results show that our new method can provide good performances for face hallucination, both in terms of reconstruction error and visual quality; and that it is competitive with existing state-of-the-art methods. Kin-Man Lam 0001, Guoping Qiu, Tingzhi Shen |
IEEE Trans. Image Process. | 2 |
| 2010 | Canny Edge Detection Using Bilateral Filter on Real Hexagonal Structure
Xiangjian He, Daming Wei, Kin-Man Lam 0001, Wenjing Jia, Qiang Wu 0001 |
ACIVS (1) | 3 |
| 2010 | A hierarchical algorithm for image multi-labelingabstractThis paper presents an efficient two-stage method for multi-class image labeling. We first propose a simple label-filtering algorithm (LFA), which can remove most of the irrelevant labels for a query image while the potential labels are maintained. With a small population of potential labels left, we then apply the Naive-Bayes Nearest-Neighbor (NBNN) classifier as the second stage of our algorithm to identify the labels for the query image. This approach has been evaluated on the Corel database, and compared to existing algorithms. Experiment results show that our proposed algorithm can achieve a promising result, as it outperforms existing algorithms. Jiwei Hu, Kin-Man Lam 0001, Guoping Qiu |
ICIP | 2 |
| 2010 | Learning local pixel structure for face hallucinationabstractIn this paper, we present a novel learning-based face hallucination method based on the assumption that similar faces will have similar local pixel structures. We use the low- resolution (LR) input face to search a database for K example faces that are the most similar to the input and align them with the input accordingly. The local pixel structures of the target high-resolution (HR) image are learned from those warped HR example faces in a neighbor embedding manner, and a total variation (TV) constraint is employed to aid the learning of all pixels' embedding weights. The learned local pixel structures are then used as constraints to reconstruct a HR version of the input face. Experimental results show that the method performs well in terms of both reconstruction error and visual quality. Kin-Man Lam 0001, Guoping Qiu, Tingzhi Shen |
ICIP | 2 |
| 2010 | Tone mapping HDR images using optimization: A general frameworkabstractThis paper presents a novel tone mapping framework. First, we introduce a tone mapping fidelity principle which explicitly stipulates that tone-mapped image data should not only be visually enhanced but should also stay faithful to the original image. Second, this principle naturally translates tone mapping into a constrained optimization problem where a two-term cost function, one measures the difference between the tone mapped image and a visually enhanced version of the image, and the other measures the difference between the tone mapped image and the original image, is optimized. The relative weightings of the two terms in the cost function not only offers an insightful and simple mechanism to control the appearance of the tone mapped image but also enables the introduction of spatially varying or uniform weighting functions thus unifying local and global tone mapping in a single framework. We present results of tone mapping high dynamic range (HDR) images and low dynamic range JPEG images to demonstrate the effectiveness of the new tone mapping framework. Guoping Qiu, Yujie Mei, Kin-Man Lam 0001 |
ICIP | 3 |
| 2010 | An image restoration method based on PDEs and a new gradient modelabstractThe goal of image restoration is to eliminate noise and insignificant details from a blurred or noise-affected image, without blurring important semantic structures such as edges. In this paper, we propose a new approach based on partial differential equations (PDEs) by adjusting the threshold when iteration proceeds, and a new gradient model based on the idea of bilateral filtering, which can retain more visual details than other methods. With this new approach, the noise can be removed efficiently while preserving important semantic structures such as edges. Xiaoling Zhang 0001, Kin-Man Lam 0001 |
ICIP | 3 |
| 2010 | Complexity scalable control for H.264 motion estimation and mode decision under energy constraints
Xuejuan Gao, Kin-Man Lam 0001, Li Zhuo 0001, Lansun Shen |
Signal Process. | 2 |
| 2009 | Partially occluded face completion and recognitionabstractThis paper proposes a spectral graph based algorithm for face image repairing, which can improve the recognition performance on occluded faces. Our algorithm is called `guided label-learning', so named from graphical models, and can achieve a high-quality repairing of damaged or occluded faces. We apply our face repairing algorithm in order to produce completed faces, and then use face recognition to evaluate the performance of our algorithm. Experiment results show that, at most, a nearly 30%-increase in the recognition rate can be achieved for occluded faces with the use of our algorithm. Yue Deng 0001, Dong Li 0028, Xudong Xie, Kin-Man Lam 0001, Qionghai Dai |
ICIP | 4 |
| 2009 | Elastic block set reconstruction for face recognitionabstractIn this paper, a novel face recognition algorithm named elastic block set reconstruction (EBSR) is proposed. In our method, the EBSR face is used to represent a set of training faces and to simulate different factors in a query image. An EBSR face is constructed by using the blocks from the training face images which best match to the blocks of the query image at the corresponding locations. The elastic local reconstruction (ELR) error is then used to evaluate how well a block pair matches, and the query image is classified based on the accumulated reconstruction error. The proposed method can effectively explore local information in the training set and deal with various conditions well. Also, the reconstruction error can be considered as a kind of dissimilarity measure, which gives a new approach to designing the training set so as to maximize robustness of recognition. Experiments show that consistent and promising results are obtained. Dong Li 0028, Xudong Xie, Kin-Man Lam 0001, Zhigang Jin |
ICIP | 3 |
| 2009 | Region-based Eigentransformation for Face Image HallucinationabstractIn this paper, we propose a region-based eigentransformation method for human face hallucination. The algorithm first uses the snake method to locate the targeted regions with high-frequency features in an edge-smoothed face mask, and then employs the eigentransformation to reconstruct the corresponding high-resolution counterpart. The resulting high-resolution face can be obtained by combining all of the reconstructed regions. As the boundaries of the respective high-frequency regions are also smooth, so the merging process at the boundaries will not produce visual artifacts. In addition, by segmenting the facial features into different regions, the eigentransformation can achieve a better performance level in reconstructing the high-frequency information. Experiments show that our proposed algorithm outperforms the original eigentransformation for face hallucination. Tingzhi Shen, Kin-Man Lam 0001 |
ISCAS | 3 |
| 2009 | Example-based image super-resolution with class-specific predictors
Kin-Man Lam 0001, Guoping Qiu, Lansun Shen, Suyu Wang |
J. Vis. Commun. Image Represent. | 2 |
| 2009 | Facial expression recognition based on shape and texture
Xudong Xie, Kin-Man Lam 0001 |
Pattern Recognit. | 2 |
| 2009 | Face detection using simplified Gabor features and hierarchical regions in a cascade of classifiers
Kin-Man Lam 0001, Lansun Shen, Jiliu Zhou |
Pattern Recognit. Lett. | 2 |
| 2009 | Efficient Edge Detection Using Simplified Gabor WaveletsabstractGabor wavelets (GWs) have been commonly used for extracting local features for various applications, such as recognition, tracking, and edge detection. However, extracting the Gabor features is computationally intensive, so the features may be impractical for real-time applications. In this paper, we propose a set of simplified version of GWs (SGWs) and an efficient algorithm for extracting the features for edge detection. Experimental results show that our SGW-based edge-detection algorithm can achieve a similar performance level to that using GWs, while the runtime required for feature extraction using SGWs is faster than that with GWs with the use of the fast Fourier transform. When compared to the Canny and other conventional edge-detection methods, our proposed method can achieve a better performance in the terms of detection accuracy and computational complexity. Kin-Man Lam 0001, Tingzhi Shen |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2008 | Efficient face recognition with a large databaseabstractFace recognition has a wide range of applications, and most of the current face recognition algorithms can achieve a high level of accuracy. However, when the face database is very large, the amount of computation required to search the faces in that database becomes an important concern. In this paper, we propose to use a number of vantage objects to form an efficient indexing structure for searching a huge database. These vantage objects are constructed using the discriminative features extracted from Gabor wavelets. The training faces in the database are ranked either in ascending or descending order with reference to each of the vantage objects, and hence each vantage object can form one or more ranked lists. With a query face, it is ranked with reference to each vantage object, and is positioned in each of the ranked lists accordingly. Then, the neighboring training faces to the query face in the respective ranked lists are selected to form a much smaller database, which is called a condensed database. Experimental results show that, with a database of more than 2000 distinct faces, the probabilities of a query face being selected in a condensed database of 36%, 26%, 11%, and 6% of the original database size are 99%, 98%, 95%, and 90%, respectively. To search a face in the much smaller condensed database, a more computational and accurate recognition algorithm can then be adopted. Siu-Hong Tse, Kin-Man Lam 0001 |
ICARCV | 2 |
| 2008 | Image magnification based on a blockwise adaptive Markov random field model
Xiaoling Zhang 0001, Kin-Man Lam 0001, Lansun Shen |
Image Vis. Comput. | 2 |
| 2008 | Simplified Gabor wavelets for human face recognition
Wing-Pong Choi, Siu-Hong Tse, Kwok-Wai Wong, Kin-Man Lam 0001 |
Pattern Recognit. | 4 |
| 2008 | Elastic shape-texture matching for human face recognition
Xudong Xie, Kin-Man Lam 0001 |
Pattern Recognit. | 2 |
| 2008 | Face recognition using elastic local reconstruction based on a single face image
Xudong Xie, Kin-Man Lam 0001 |
Pattern Recognit. | 2 |
| 2008 | Recovering the 3D shape and poses of face images based on the similarity transform
Hei-Sheung Koo, Kin-Man Lam 0001 |
Pattern Recognit. Lett. | 2 |
| 2008 | An Efficient Scene-Break Detection Method Based on Linear Prediction With Bayesian Cost FunctionsabstractThis paper describes an efficient approach to scene-break detection, which can detect cuts, dissolves, and wipes reliably and effectively by means of temporally linear prediction models. In our algorithm, two linear prediction models are adopted to predict a current frame: one for dissolves, and the other for stationary scenes. The predicted frames, derived based on the two models, are compared with the original frames, and cuts and dissolves are then determined based on Bayesian cost functions. For the detection, our algorithm requires the setting of a single threshold only. In wipe detection, our linear prediction models are employed to detect areas of change between two successive frames. By accumulating the changed areas and the overlap of the changed areas over the successive frames, wipes of an arbitrary shape and direction are detected. Experimental results show that our algorithm can achieve a high level of precision even if a video contains object motion and camera motion. The detection time required to analyze a 38-min video is no more than several seconds. Cheng Cai, Kin-Man Lam 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | An Effective Promoter Detection Method using the Adaboost Algorithm
Xudong Xie, Shuanhu Wu, Kin-Man Lam 0001, Hong Yan 0001 |
APBC | 3 |
| 2007 | An adaptive algorithm for the display of high-dynamic range images
Kin-Man Lam 0001, Lansun Shen |
J. Vis. Commun. Image Represent. | 2 |
| 2006 | Semi-supervised Learning based on Bayesian Networks and Optimization for Interactive Image RetrievalabstractIn this paper, we present a novel interactive image retrieval technique using semi-supervised learning. Recently, Guan and Qiu [8, 9] have shown that by constructing a Bayesian Network where the nodes represent the (continuous) class membership scores and arcs represent the dependence relations of the data points, the (semi-supervised) classification problem can be formulated as a quadratic optimization problem; and by using the labeled data as linear constraints, the optimization problem yields a large, sparse system of linear equations which can be solved very efficiently using standard methods. In this work, we show that this semi-supervised learning method can be naturally adopted as a computational tool to incorporate users feedbacks for interactive image retrieval. We present experimental results to show the effectiveness of our new interactive image retrieval method. We also show that semisupervised learning can have advantages over supervised and unsupervised learning in image retrieval applications. 1 Mai Yang, Jian Guan 0004, Guoping Qiu, Kin-Man Lam 0001 |
BMVC | 4 |
| 2006 | PromoterExplorer: an effective promoter identification method based on the AdaBoost algorithmabstractMOTIVATION: Promoter prediction is important for the analysis of gene regulations. Although a number of promoter prediction algorithms have been reported in literature, significant improvement in prediction accuracy remains a challenge. In this paper, an effective promoter identification algorithm, which is called PromoterExplorer, is proposed. In our approach, we analyze the different roles of various features, that is, local distribution of pentamers, positional CpG island features and digitized DNA sequence, and then combine them to build a high-dimensional input vector. A cascade AdaBoost-based learning procedure is adopted to select the most 'informative' or 'discriminating' features to build a sequence of weak classifiers, which are combined to form a strong classifier so as to achieve a better performance. The cascade structure used for identification can also reduce the false positive. RESULTS: PromoterExplorer is tested based on large-scale DNA sequences from different databases, including the EPD, DBTSS, GenBank and human chromosome 22. Experimental results show that consistent and promising performance can be achieved. Xudong Xie, Shuanhu Wu, Kin-Man Lam 0001, Hong Yan 0001 |
Bioinform. | 3 |
| 2006 | An efficient illumination normalization method for face recognition
Xudong Xie, Kin-Man Lam 0001 |
Pattern Recognit. Lett. | 2 |
| 2006 | Channel-adaptive error protection for streaming stored MPEG-4 FGS over error-prone environmentsabstractIn this paper, we investigate an adaptive channel error-protection scheme for streaming stored MPEG-4 fine granular scalability (FGS) bitstreams over error-prone environments. The human eye is sensitive to the variations in reconstructed image quality. Therefore, instead of using only the minimal average frame distortion as the optimization criterion, we propose a rate-distortion (R-D)-based bit-allocation method to determine the source coding rate and the channel coding rate under a given channel condition so as to minimize both the average image distortion and quality variation. Based on our proposed piecewise model of the MPEG-4 FGS enhancement layer, a balance between the average frame quality and quality variation will be obtained. Experimental results show that, compared with other existing methods, our proposed method can achieve less quality variation and comparable average frame distortion under various channel conditions. Li Zhuo 0001, Kin-Man Lam 0001, Lansun Shen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2006 | A faster converging snake algorithm to locate object boundariesabstractA different contour search algorithm is presented in this paper that provides a faster convergence to the object contours than both the greedy snake algorithm (GSA) and the fast greedy snake (FGSA) algorithm. This new algorithm performs the search in an alternate skipping way between the even and odd nodes (snaxels) of a snake with different step sizes such that the snake moves to a likely local minimum in a twisting way. The alternative step sizes are adjusted so that the snake is less likely to be trapped at a pseudo-local minimum. The iteration process is based on a coarse-to-fine approach to improve the convergence. The proposed algorithm is compared with the FGSA algorithm that employs two alternating search patterns without altering the search step size. The algorithm is also applied in conjunction with the subband decomposition to extract face profiles in a hierarchical way. Mustafa Sakalli, Kin-Man Lam 0001, Hong Yan 0001 |
IEEE Trans. Image Process. | 2 |
| 2006 | Gabor-based kernel PCA with doubly nonlinear mapping for face recognition with a single face imageabstractIn this paper, a novel Gabor-based kernel principal component analysis (PCA) with doubly nonlinear mapping is proposed for human face recognition. In our approach, the Gabor wavelets are used to extract facial features, then a doubly nonlinear mapping kernel PCA (DKPCA) is proposed to perform feature transformation and face recognition. The conventional kernel PCA nonlinearly maps an input image into a high-dimensional feature space in order to make the mapped features linearly separable. However, this method does not consider the structural characteristics of the face images, and it is difficult to determine which nonlinear mapping is more effective for face recognition. In this paper, a new method of nonlinear mapping, which is performed in the original feature space, is defined. The proposed nonlinear mapping not only considers the statistical property of the input features, but also adopts an eigenmask to emphasize those important facial feature points. Therefore, after this mapping, the transformed features have a higher discriminating power, and the relative importance of the features adapts to the spatial importance of the face images. This new nonlinear mapping is combined with the conventional kernel PCA to be called "doubly" nonlinear mapping kernel PCA. The proposed algorithm is evaluated based on the Yale database, the AR database, the ORL database and the YaleB database by using different face recognition methods such as PCA, Gabor wavelets plus PCA, and Gabor wavelets plus kernel PCA with fractional power polynomial models. Experiments show that consistent and promising results are obtained. Xudong Xie, Kin-Man Lam 0001 |
IEEE Trans. Image Process. | 2 |
| 2005 | Illumination invariant face recognition
Dang-Hui Liu, Kin-Man Lam 0001, Lan-Sun Shen |
Pattern Recognit. | 2 |
| 2005 | Face recognition under varying illumination based on a 2D face shape model
Xudong Xie, Kin-Man Lam 0001 |
Pattern Recognit. | 2 |
| 2005 | An accurate active shape model for facial feature extraction
Kwok-Wai Wan, Kin-Man Lam 0001, Kit-Chong Ng |
Pattern Recognit. Lett. | 2 |
| 2005 | A new key frame representation for video segment retrievalabstractIn this paper, we propose an optimal key frame representation scheme based on global statistics for video shot retrieval. Each pixel in this optimal key frame is constructed by considering the probability of occurrence of those pixels at the corresponding pixel position among the frames in a video shot. Therefore, this constructed key frame is called temporally maximum occurrence frame (TMOF), which is an optimal representation of all the frames in a video shot. The retrieval performance of this representation scheme is further improved by considering the k pixel values with the largest probabilities of occurrence and the highest peaks of the probability distribution of occurrence at each pixel position for a video shot. The corresponding schemes are called k-TMOF and k-pTMOF, respectively. These key frame representation schemes are compared to other histogram-based techniques for video shot representation and retrieval. In the experiments, three video sequences in the MPEG-7 content set were used to evaluate the performances of the different key frame representation schemes. Experimental results show that our proposed representations outperform the alpha-trimmed average histogram for video retrieval. Kin-Wai Sze, Kin-Man Lam 0001, Guoping Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2004 | An efficient illumination compensation scheme for face recognitionabstractThis paper proposes a novel illumination compensation algorithm, which can compensate for the uneven illuminations on human faces and reconstruct face images of normal lighting conditions. A simple yet effective local contrast enhancement method, namely block-based histogram equalization (BHE), is first proposed. The resulting image processed by BHE is then compared with the original face image processed using histogram equalization (HE) to estimate the lighting category. Based on the category identified, a corresponding lighting compensation model is used to reconstruct an image that would visually be under normal illumination. In order to eliminate the influence of uneven illumination while retaining the shape information, a 2D face shape model is used. Experimental results show that, with the use of principal component analysis (PCA) for face recognition, the recognition rate can be improved by 53.3% to 56,1% when our proposed algorithm for lighting compensation is used. Xudong Xie, Kin-Man Lam 0001 |
ICARCV | 2 |
| 2004 | Mean-shift based mixture model for face detection in color imageabstractHuman face detection is a challenging task under different lighting conditions. We propose an efficient and reliable algorithm to detect human faces in an image. Our algorithm uses a region-based approach to identify skin-colored pixels under various lighting conditions. Within the detected skin-color regions, a ratio method is proposed to determine possible eye candidates. Two eye candidates form a possible face region, which is then verified by means of a two-stage procedure with an eigenmask. Experimental results based on the HHI MPEG-7 face database show that this face detection algorithm is efficient and reliable under different lighting conditions. Tze-Yin Chow, Kin-Man Lam 0001 |
ICIP | 2 |
| 2004 | Rate control for MPEG-4 FGS coded video using piecewise rate distortion modelabstractThree efficient rate allocation schemes for the MPEG-4 fine granular scalability (FGS) enhancement layer with constant quality are proposed. With the distortion measured using the mean squared error (MSE) or peak signal-to-noise ratio (PSNR) metric, the rate distortion (R-D) curve of the MPEG-4 FGS enhancement layer exhibits piecewise characteristics. For each piecewise region, a first-order or a second-order model is adopted to estimate the actual R-D curve. Based on these piecewise models, three optimal rate allocation schemes are proposed to allocate an appropriate number of bits to the enhancement layer of each frame. Experimental results show that, compared to the average rate allocation method and the exponential model method, our proposed algorithms can achieve a more constant quality in the reconstructed video. Kin-Man Lam 0001, Lansun Shen |
ICME | 2 |
| 2004 | Optimal sampling of Gabor features for face recognition
Dang-Hui Liu, Kin-Man Lam 0001, Lan-Sun Shen |
Pattern Recognit. Lett. | 2 |
| 2003 | Maximal disk based histogram for shape retrievalabstractWe propose a robust and efficient representation scheme for shape retrieval, which is based on the normalized maximal disks used to represent the shape of an object. The maximal disks are extracted by means of a fast skeletonization technique with a pruning algorithm. The logarithm of the radii of the normalized maximal disks is used to construct a histogram to represent the shape. The retrieval performance of this maximal disk based histogram approach is compared to other methods, including moment invariants, Zernike moments, and curvature scale-space. Experimental results show that our proposed representation scheme outperforms the other methods under affine transformation and different noise levels. Wai-Pak Choi, Kin-Man Lam 0001, Wan-Chi Siu |
ICASSP (3) | 2 |
| 2003 | Scene cut detection using the colored pattern appearance modelabstractIn this paper, we propose to use the colored pattern appearance model (CPAM) as a content representation for video scene break detection. This model represents a scene by means of global statistics of the local visual appearance, and was originally motivated by studies in human color vision. The performance of this method is compared to several histogram-based approaches. An adaptive thresholding technique, namely entropic thresholding, is applied to determine the respective optimal threshold values for each of the approaches. In the experiments, the two video sequences in the MPEG-7 content set are used to evaluate the performances of the CPAM and the histogram-based methods. Experimental results show that our proposed model outperforms other histogram-based approaches in scene break detection. Kin-Wai Sze, Kin-Man Lam 0001, Guoping Qiu |
ICIP (2) | 2 |
| 2003 | Extraction of the Euclidean skeleton based on a connectivity criterion
Wai-Pak Choi, Kin-Man Lam 0001, Wan-Chi Siu |
Pattern Recognit. | 2 |
| 2003 | Spatially eigen-weighted Hausdorff distances for human face recognition
Kwan-Ho Lin, Kin-Man Lam 0001, Wan-Chi Siu |
Pattern Recognit. | 2 |
| 2003 | Human face recognition based on spatially weighted Hausdorff distance
Baofeng Guo, Kin-Man Lam 0001, Kwan-Ho Lin, Wan-Chi Siu |
Pattern Recognit. Lett. | 2 |
| 2003 | A robust scheme for live detection of human faces in color images
Kwok-Wai Wong, Kin-Man Lam 0001, Wan-Chi Siu |
Signal Process. Image Commun. | 2 |
| 2003 | A fast fractal image coding based on kick-out and zero contrast conditionsabstractA fast algorithm for fractal image coding based on a single kick-out condition and the zero contrast prediction is proposed in this paper. The single kick-out condition can avoid a large number of range-domain block matches when finding the best matched domain block. An efficient method for zero contrast prediction is also proposed, which can determine whether the contrast factor for a domain block is zero or not, and compute the corresponding difference between the range block and the transformed domain block efficiently and exactly. The proposed algorithm can achieve the same reconstructed image quality as the exhaustive search, and can greatly reduce the required computation or runtime. In addition, this algorithm does not need any pre-processing step or additional memory for its implementation, and can combine with other fast fractal algorithms to further improve the speed. Experimental results show that the runtime is reduced by about 50% of that of the exhaustive search method. When combined with the DCT Inner Product algorithm, the required runtime for the algorithm can be further reduced by about 50%. The proposed algorithm was also compared to two other fast fractal algorithms. Experimental results also show that our algorithm achieves a better efficiency and requires a much smaller amount of memory for implementation. Cheung-Ming Lai, Kin-Man Lam 0001, Wan-Chi Siu |
IEEE Trans. Image Process. | 2 |
| 2003 | Frequency layered color indexing for content-based image retrievalabstractImage patches of different spatial frequencies are likely to have different perceptual significance as well as reflect different physical properties. Incorporating such concept is helpful to the development of more effective image retrieval techniques. We introduce a method which separates an image into layers, each of which retains only pixels in areas with similar spatial frequency characteristics and uses simple low-level features to index the layers individually. The scheme associates indexing features with perceptual and physical significance thus implicitly incorporating high level knowledge into low level features. We present a computationally efficient implementation of the method, which enhances the power and at the same time retains the simplicity and elegance of basic color indexing. Experimental results are presented to demonstrate the effectiveness of the method. Guoping Qiu, Kin-Man Lam 0001 |
IEEE Trans. Image Process. | 2 |
| 2002 | A new approach using modified Hausdorff distances with eigenface for human face recognitionabstractHausdorff distance is an efficient measure of the similarity of two point sets. In this paper, we propose two new spatially weighted Hausdorff distance measures for human face recognition, namely, spatially eigen-weighted Hausdorff distance (SEWHD) and spatially eigen-weighted 'doubly' Hausdorff distance (SEW2HD). These new Hausdorff distances incorporate the information about the location of important facial features so that distances at those regions will be emphasized. The weighting function used in the Hausdorff distance measure is based on an eigenface, which has a large value at locations of important facial features and can reflect the face structure more effectively. Experimental results based on a combination of the ORL, MIT, and Yale face databases show that SEW2HD can achieve recognition rates of 83%, 90% and 92% for the first one, the first three and the first five likely matched faces, respectively, while the corresponding recognition rates of SEWHD are 80%, 83% and 88%, respectively. Kwan-Ho Lin, Kin-Man Lam 0001, Wan-Chi Siu |
ICARCV | 2 |
| 2002 | An efficient algorithm for the extraction of a Euclidean skeletonabstractThe skeleton is essential for general shape representation but the discrete representation of an image presents a lot of problems that may influence the process of skeleton extraction. Some of the methods are memory-intensive and computationally intensive, and require a complex data structure. In this paper, we propose a fast, efficient and accurate skeletonization method for the extraction of a well-connected Euclidean skeleton based on a signed sequential Euclidean distance map. A connectivity criterion that can be used to determine whether a given pixel inside an object is a skeleton point is proposed. The criterion is based on a set of points along the object boundary, which are the nearest contour points to the pixel under consideration and its 8 neighbors. The extracted skeleton is of single-pixel width without requiring a linking algorithm or iteration process. Experiments show that the runtime of our algorithm is faster than. those of using the distance transformation and is linearly proportional to the number of pixels of an image. Wai-Pak Choi, Kin-Man Lam 0001, Wan-Chi Siu |
ICASSP | 2 |
| 2002 | Robust Hausdorff distance for shape matching
Wai-Pak Choi, Kin-Man Lam 0001, Wan-Chi Siu |
VCIP | 2 |
| 2001 | Locating the human eye using fractal dimensionsabstractA new method for locating eye pairs based on valley field detection and measurement of fractal dimensions is proposed. Fractal dimension is an efficient representation of the texture of facial features. Possible eye candidates to an image with a complex background are identified by valley field detection. The eye candidates are then grouped to form eye pairs if their local properties for eyes are satisfied. Two eyes are matched if they have similar roughness and orientation as represented by fractal dimensions. We propose a modified approach to estimate the fractal dimensions that are less sensitive to lighting conditions and provide information about the orientation of an image under consideration. Possible eye pairs are further verified by comparing the fractal dimensions of the eye-pair window and the corresponding face region with the respective means of the fractal dimensions of the eye-pair windows and the face regions. The means of the fractal dimensions are obtained based on a number of facial images in a database. Experiments have shown that this approach is fast and reliable. This shows that the texture of the eyes can be represented very well by fractal surfaces. Kwan-Ho Lin, Kin-Man Lam 0001, Wan-Chi Siu |
ICIP (3) | 2 |
| 2001 | An adaptive active contour model for highly irregular boundaries
Wai-Pak Choi, Kin-Man Lam 0001, Wan-Chi Siu |
Pattern Recognit. | 2 |
| 2001 | A novel approach for human face detection from color images under complex background
Kwok-Wai Wong, Kin-Man Lam 0001, Wan-Chi Siu |
Pattern Recognit. | 2 |
| 2001 | An efficient low bit-rate video-coding algorithm focusing on moving regionsabstractBlock-based motion estimation and compensation are the most popular techniques for video coding. However, as the shape and the structure of an object in a picture are arbitrary, the performance of such conventional block-based methods may not be satisfactory. In this paper, a very low bit-rate video coding algorithm that focuses on moving regions is proposed. The aim is to improve the coding performance, which gives better subjective and objective quality than that of the conventional coding methods at the same bit rate. Eight patterns are pre-defined to approximate the moving regions in a macroblock. The patterns are then used for motion estimation and compensation to reduce the prediction errors. Furthermore, in order to increase the compression performance, the residual errors of a macroblock are rearranged into a block with no significant increase of high-order DCT coefficients. As a result, both the prediction efficiency and the compression efficiency are improved. Kwok-Wai Wong, Kin-Man Lam 0001, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2000 | An Efficient and Accurate Algorithm for Extracting a SkeletonabstractIn this paper, a non-iterative method is proposed, which is fast, efficient, and more importantly, is robust to boundary noise and rotation. Unnecessary branches and hairs can be reduced by the adjustment of the residual distance and the skeleton can be represented in a hierarchical manner. The reconstruction error can also be estimated using the residual distance. A new definition of a skeleton and the criteria of being a skeleton point are introduced. The effect of boundary noise and curved boundary are investigated and compared to other skeletonization algorithms. Finally, the reconstruction of an object using its skeleton and the associated radii of the maximal disks, is illustrated. Wai-Pak Choi, Kin-Man Lam 0001, Wan-Chi Siu |
ICPR | 2 |
| 1998 | Shivering Greedy Snakes Gradient-Guided in Wavelet DomainabstractThe convergence speed of snakes is increased with the switching step size of local minima search and they are incorporated with the local gradient information in the wavelet domain to provide extra speed of convergence to their final contours. This is important in model based compression applications using wavelet decomposition. The directional gradient information is extracted from the subbands provided for the three directions at each stage of the decomposition. The advantages of this approach include, increased speed of convergence due to the switching step size of search and due to the decreased number of pixels to be searched and decreased burden of calculations for pre-processing due to the decreased size of subbands and due to the decreased dimension of the Gaussian filters in the pre-processing stage. Mustafa Sakalli, Kin-Man Lam 0001, Hong Yan 0001 |
ICIP (2) | 2 |
| 1998 | Model-based multi-stage compression of human face imagesabstractThis paper describes a multi-stage compression scheme of human face images. Snakes are employed in localisation of facial features and biorthogonal spline filters are used for the decomposition of segmented and normalised face images. Wavelet coefficients are vector quantized in different number of cascaded stages depending on their contribution to the subjective quality of the image. Mustafa Sakalli, Hong Yan 0001, Kin-Man Lam 0001, Toshiaki Kondo |
ICPR | 3 |
| 1998 | Human Face Image Recognition: An Evidence Aggregation Approach
Ali Reza Mirhosseini, Hong Yan 0001, Kin-Man Lam 0001, Tuan D. Pham |
Comput. Vis. Image Underst. | 3 |
| 1998 | An Analytic-to-Holistic Approach for Face Recognition Based on a Single Frontal ViewabstractWe propose an analytic-to-holistic approach which can identify faces at different perspective variations. The database for the test consists of 40 frontal-view faces. The first step is to locate 15 feature points on a face. A head model is proposed, and the rotation of the face can be estimated using geometrical measurements. The positions of the feature points are adjusted so that their corresponding positions for the frontal view are approximated. These feature points are then compared with the feature points of the faces in a database using a similarity transform. In the second step, we set up windows for the eyes, nose, and mouth. These feature windows are compared with those in the database by correlation. Results show that this approach can achieve a similar level of performance from different viewing directions of a face. Under different perspective variations, the overall recognition rates are over 84 percent and 96 percent for the first and the first three likely matched faces, respectively. Kin-Man Lam 0001, Hong Yan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1997 | A Hierarchical and Adaptive Deformable Model for Mouth Boundary DetectionabstractAn automatic algorithm to extract mouth boundaries in human face images is proposed. The algorithm is based on a hierarchical model adaptation scheme using deformable models. The knowledge about the shape of the object is used to define its initial deformable template. Each mouth boundary curve is initially formed based on three control points whose locations are found through an optimization process using a suitable cost functional. The cost functional captures the essential knowledge about the shape for perceptual organization. Two control points are the mouth corners, which are used as the initial location of the mouth after an approximate mouth window is found based on locating the head boundary. The model is hierarchically improved in the second stage of the algorithm. Each boundary curve is finely tuned using more control points. An old model is adaptively replaced by a new model only if a secondary cost is further reduced. The results show that the model adaptation technique satisfactorily enhances the mouth boundary model in an automated fashion. Ali Reza Mirhosseini, Hong Yan 0001, Kin-Man Lam 0001 |
ICIP (2) | 3 |
| 1996 | An Improved Method for Locating and Extracting the Eye in Human Face ImagesabstractIn this paper, the head boundary is first located in a head-and-shoulders image. The approximate positions of the eyes are estimated by means of average anthropometric measures. Corners, the salient features of the eyes, are detected. The corner detection scheme introduced in this paper can provide accurate information about the corners. The shape of the eye is then extracted by a new scheme, based on snakes and on the deformable template. This new representation is in a form similar to snakes, but the energy functional to be minimised is based on some new energy terms and those terms used by the deformable template. A fast algorithm based on the greedy algorithm for active contour modelling is presented for the optimisation step. Experiments show that the execution time required for the whole procedure to extract the eye features is less than one second on a Sun workstation. Kin-Man Lam 0001, Hong Yan 0001 |
ICPR | 1 |
| 1996 | Locating and extracting the eye in human face images
Kin-Man Lam 0001, Hong Yan 0001 |
Pattern Recognit. | 1 |
| 1995 | Generation of moment invariants and their uses for character recognition
Wai-Hong Wong, Wan-Chi Siu, Kin-Man Lam 0001 |
Pattern Recognit. Lett. | 3 |
| 1993 | Transform-based fine-corase vector quantization
Kin-Man Lam 0001, Wan-Chi Siu, Kai-Ming Tse |
ISCAS | 1 |
| 1993 | Automatic generation of moment invariants and the use of higher order moments for character recognition
Wai-Hong Wong, Wan-Chi Siu, Kin-Man Lam 0001 |
ISCAS | 3 |