EDBT 2026 Demo / reviewers in the wild / expert
Alex Chichung Kot
dblp:k/AlexChiChungKot · also A. C. Kot, Alex C. Kot
· DBLP profile ↗
314ranked-venue papers
1as first author
128since 2021 · last 2026
0000-0001-6262-8125ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 185 · 1 first-author · 76 since 2021Artificial intelligence and machine learning · 110 · 64 since 2021Security and privacy · 28 · 12 since 2021Computer networks · 15 · 2 since 2021Systems, architecture and hardware · 13 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 4 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Theory of computation · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early RevisionabstractLarge Vision-Language Models (LVLMs) recently achieve significant breakthroughs in understanding complex visual-textual contexts. However, hallucination issues still limit their real-world applicability. Although previous mitigation methods effectively reduce hallucinations in photographic images, they largely overlook the potential risks posed by stylized images, which play crucial roles in critical scenarios such as game scene understanding, art education, and medical analysis. In this work, we first construct a dataset comprising photographic images and their corresponding stylized versions with carefully annotated caption labels. We then conduct head-to-head comparisons on both discriminative and generative tasks by benchmarking 13 advanced LVLMs on the collected datasets. Our findings reveal that stylized images tend to induce significantly more hallucinations than their photographic counterparts. To address this issue, we propose Style-Aware Visual Early Revision (SAVER), a novel mechanism that dynamically adjusts LVLMs' final outputs based on the token-level visual attention patterns, leveraging early-layer feedback to mitigate hallucinations caused by stylized images. Extensive experiments demonstrate that SAVER achieves state-of-the-art performance in hallucination mitigation across various models, datasets, and tasks. Zhaoxu Li, Chenqi Kong, Yi Yu 0011, Qiangqiang Wu, Xinghao Jiang, Ngai-Man Cheung, Bihan Wen, Alex Chichung Kot, Xudong Jiang 0001 |
AAAI | 8 |
| 2026 | From Pretrain to Pain: Adversarial Vulnerability of Video Foundation Models Without Task KnowledgeabstractLarge-scale Video Foundation Models (VFMs) have significantly advanced various video-related tasks, either through task-specific models or Multi-modal Large Language Models (MLLMs). However, the open accessibility of VFMs also introduces critical security risks, as adversaries can exploit full knowledge of the VFMs to launch potent attacks. This paper investigates a novel and practical adversarial threat scenario: attacking downstream models or MLLMs fine-tuned from open-source VFMs, without requiring access to the victim task, training data, model query, and architecture. In contrast to conventional transfer-based attacks that rely on task-aligned surrogate models, we demonstrate that adversarial vulnerabilities can be exploited directly from the VFMs. To this end, we propose the Transferable Video Attack (TVA), a temporal-aware adversarial attack method that leverages the temporal representation dynamics of VFMs to craft effective perturbations. TVA integrates a bidirectional contrastive learning mechanism to maximize the discrepancy between the clean and adversarial features, and introduces a temporal consistency loss that exploits motion cues to enhance the sequential impact of perturbations. TVA avoids the need to train expensive surrogate models or access to domain-specific data, thereby offering a more practical and efficient attack strategy. Extensive experiments across 24 video-related tasks demonstrate the efficacy of TVA against downstream models and MLLMs, revealing a previously underexplored security vulnerability in the deployment of video models. Yi Yu 0011, Song Xia, Deepu Rajan, Boon Poh Ng, Alex Chichung Kot, Xudong Jiang 0001 |
AAAI | 7 |
| 2026 | Active Adversarial Noise Suppression for Image Forgery LocalizationabstractRecent advances in deep learning have significantly propelled the development of image forgery localization. However, existing models remain highly vulnerable to adversarial attacks: imperceptible noise added to forged images can severely mislead these models. In this paper, we address this challenge with an Adversarial Noise Suppression Module (ANSM) that generates a defensive perturbation to suppress the attack effect of adversarial noise. We observe that forgery-relevant features extracted from adversarial and original forged images exhibit distinct distributions. To bridge this gap, we introduce Forgery-relevant Features Alignment (FFA) as a first-stage training strategy, which reduces distributional discrepancies by minimizing the channel-wise Kullback-Leibler divergence between these features. To further refine the defensive perturbation, we design a second-stage training strategy, termed Mask-guided Refinement (MgR), which incorporates a dual-mask constraint. MgR ensures that the defensive perturbation remains effective for both adversarial and original forged images, recovering forgery localization accuracy to their original level. Extensive experiments across various attack algorithms demonstrate that our method significantly restores the forgery localization model's performance on adversarial images. Notably, when ANSM is applied to original forged images, the performance remains nearly unaffected. To our best knowledge, this is the first report of adversarial defense in image forgery localization tasks. Rongxuan Peng, Shunquan Tan, Xianbo Mo, Alex Chichung Kot, Jiwu Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | See what you seek: Semantic contextual integration for cloth-changing person re-identification
Wenxin Huang, Xian Zhong, Jingling Yuan, Alex Chichung Kot |
Pattern Recognit. | 5 |
| 2026 | Open-Set Deepfake Detection: A Parameter-Efficient Adaptation Method With Forgery Style MixtureabstractOpen-set face forgery detection poses significant security threats and presents substantial challenges for existing detection models. These detectors primarily have two limitations: they cannot generalize across unknown forgery domains or inefficiently adapt to new data. To address these issues, we introduce an approach that is both general and parameter-efficient for face forgery detection. Our method builds on the assumption that different forgery source domains exhibit distinct style statistics. Specifically, we design a forgery-style-mixture formulation that augments the diversity of forgery source domains, enhancing the model’s generalizability across unseen domains. In addition, previous methods typically require fully fine-tuning pretrained networks, consuming substantial time and computational resources. Drawing on recent advancements in vision transformers (ViT) for face forgery detection, we develop a parameter-efficient ViT-based detection model that includes lightweight forgery feature extraction modules and enables the model to extract global and local forgery clues simultaneously. We only optimize the inserted lightweight modules during training, maintaining the original ViT structure with its pre-trained weights. This training strategy effectively preserves the informative pre-trained knowledge while flexibly adapting the model to the task of Deepfake detection. Extensive experimental results demonstrate that the designed model achieves state-of-the-art generalizability with significantly reduced trainable parameters, representing an important step toward open-set Deepfake detection in the wild. Chenqi Kong, Anwei Luo, Peijun Bao, Haoliang Li, Renjie Wan, Zengwei Zheng, Anderson Rocha 0001, Alex Chichung Kot |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2026 | MoE-FFD: Mixture of Experts for Generalized and Parameter-Efficient Face Forgery DetectionabstractDeepfakes have recently raised significant trust issues and security concerns among the public. Compared to CNN-based face forgery detectors, ViT-based methods take advantage of the expressivity of transformers, achieving superior detection performance. However, these approaches still exhibit the following limitations: (1) Fully fine-tuning ViT-based models from ImageNet weights demands substantial computational and storage resources; (2) ViT-based methods struggle to capture local forgery clues, leading to model bias; (3) These methods limit their scope on only one or few face forgery features, resulting in limited generalizability. To tackle these challenges, this work introduces Mixture-of-Experts modules for Face Forgery Detection (MoE-FFD), a generalized yet parameter-efficient ViT-based approach. MoE-FFD only updates lightweight Low-Rank Adaptation (LoRA) and Adapter layers while keeping the ViT backbone frozen, thereby achieving parameter-efficient training. Moreover, MoE-FFD leverages the expressivity of transformers and local priors of CNNs to simultaneously extract global and local forgery clues. Additionally, novel MoE modules are designed to scale the model's capacity and smartly select optimal forgery experts, further enhancing forgery detection performance. Our proposed learning scheme can be seamlessly adapted to various transformer backbones in a plug-and-play manner. Extensive experimental results demonstrate that the proposed method achieves state-of-the-art face forgery detection performance with significantly reduced parameter overhead in cross-dataset, cross-manipulation, and robustness evaluations. Our ablation studies further validate the effectiveness of the designed components and the proposed learning scheme. The code is available at: https://github.com/LoveSiameseCat/MoE-FFD. Chenqi Kong, Anwei Luo, Peijun Bao, Yi Yu 0011, Haoliang Li, Zengwei Zheng, Shiqi Wang 0001, Alex Chichung Kot |
IEEE Trans. Dependable Secur. Comput. | 8 |
| 2026 | Proactive Image Manipulation Detection and Tracing in Fake NewsabstractThe pervasive spread of fake news, particularly through manipulated images, presents a consequential negative impact on society. To prevent fake news images from misleading the public, existing methods focus on verifying the authenticity of news images but ignore source traceability, leaving a gap in creating a complete forensic chain for reliable fake news detection. To simultaneously achieve the goals of authenticity verification and source tracing, we propose a proactive image tagging approach based on a design of Disentangled Invertible Neural Networks (DINN). It can simultaneously embed the dual-tags,i.e., authenticable tag and traceable tag, into each news image prior to publication, allowing for separate extraction for authenticity verification and source tracing. Within the proposed DINN, we design a parallel Feature Aware Projection Module (FAPM) to assist DINN in preserving essential tag information, thereby improving extraction accuracy. In addition, we introduce a Distance Metric-Guided Module (DMGM) that learns asymmetric one-class representations, enabling the dual-tags to exhibit different robustness performances under malicious manipulations. Extensive experiments on diverse datasets and unseen manipulations demonstrate that the proposed tagging approach achieves promising performances on both authenticity verification and source tracing for reliable fake news detection and outperforms the prior works. Ruohan Meng, Siyuan Yang 0001, Zhili Zhou 0001, Kwok-Yan Lam, Zengwei Zheng, Alex Chichung Kot |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2026 | DCD-UIE: Decoupled Chromatic Diffusion Model for Underwater Image EnhancementabstractColor distortion and structural degradation in underwater images are classic challenges in underwater image enhancement. The core goal is to restore degraded images to high-quality images with both color and structure that conform to visual perception. However, in the traditional RGB space, these two issues are highly coupled, resulting in existing enhancement methods often neglecting one over the other. To address this challenge, we propose a guided diffusion model based on the principle of decoupling. Our key insight is that in perceptual color spaces such as HSV, color (H, S) and structure (V) are naturally separated. To exploit this property, we first design an adaptive perceptual guidance module, which analyzes the degraded HSV image and generates two orthogonal guidance signals: a color guide and a structure guide, which guide the denoising process of the diffusion model. To ensure that this decoupled guidance is faithfully implemented, we propose a corresponding decoupled loss optimization module, which uses independent loss functions to supervise the final output color and structure. By combining the forward decoupled guidance with the backward decoupled supervision, we construct a closed-loop optimization framework. This framework enables the model to collaboratively optimize color and structure under various degradation scenarios. Extensive experiments demonstrate that our proposed method outperforms existing state-of-the-art approaches in a variety of underwater scenes, particularly those degraded by color casts and haze. Furthermore, it exhibits superior performance on no-reference image quality assessment metrics. The source code is available at https://github.com/zy-world/DCD-UIE. Jingchun Zhou, Yakun Ju, Guang-Yong Chen, Jinjiang Li 0001, Alex Chichung Kot |
IEEE Trans. Image Process. | 7 |
| 2026 | Open-Set Anomaly Segmentation in Complex ScenariosabstractPrecise segmentation of out-of-distribution (OoD) objects, herein referred to as anomalies, is crucial for the reliable deployment of semantic segmentation models in open-set, safety-critical applications, such as autonomous driving. Current anomalous segmentation benchmarks predominantly focus on favorable weather conditions, resulting in untrustworthy evaluations that overlook the risks posed by diverse meteorological conditions in open-set environments, such as low illumination, dense fog, and heavy rain. To bridge this gap, this paper introduces the ComSAmy, a Complex Scenarios Anomaly segmentation benchmark. ComSAmy encompasses a wide spectrum of adverse weather conditions, dynamic driving environments, and diverse anomaly types to comprehensively evaluate the model performance in realistic open-world scenarios. Our extensive evaluation of several state-of-the-art anomalous segmentation models reveals that existing methods demonstrate significant deficiencies in such challenging scenarios, highlighting their serious safety risks for real-world deployment. To solve that, we propose a novel energy-entropy learning (EEL) strategy that integrates the complementary information from energy and entropy to bolster the robustness of anomaly segmentation under complex open-world environments. Additionally, a diffusion-based anomalous training data synthesizer is proposed to generate diverse and high-quality anomalous images to enhance the existing copy-paste training data synthesizer. Extensive experimental results on both public and ComSAmy benchmarks demonstrate that our proposed diffusion-based synthesizer with energy and entropy learning (DiffEEL) framework serves as an effective and generalizable plug-and-play method to enhance existing models, yielding an average improvement of around 4.96% in AUPRC and 9.87% in $\rm {FPR}_{95}$ . Song Xia, Yi Yu 0011, Henghui Ding, Wenhan Yang, Shifei Liu, Alex Chichung Kot, Xudong Jiang 0001 |
IEEE Trans. Image Process. | 6 |
| 2026 | Fine-Grained Lexical-Centric Semantic Network for Coherent Video Paragraph CaptioningabstractVideo paragraph captioning (VPC) aims to generate coherent, detailed narratives that accurately reflect a video's content. However, existing methods typically depend on coarse-grained event correlations and neglect the nuanced spatio-temporal interactions critical for comprehensive understanding. Refined verbs and prepositions, encoding actions and spatial relations, are essential for clear, consistent descriptions. To address these issues, we propose the Fine-Grained Lexical-Centric Semantic Network (FLS-Net), which emphasizes verbs and prepositions linked to salient objects to improve spatio-temporal coherence across events. FLS-Net integrates a multi-lexical synergy mechanism, leveraging nouns obtained via multi-modal matching, and employs a Verb-Guided Event Consistency Module (VECM) alongside a Preposition-Driven Relation Representation Module (PRRM). A cyclic encoder-decoder architecture further enforces event consistency, significantly boosting VPC performance. Extensive experiments onActivityNet CaptionsandYouCook2demonstrate FLS-Net's superiority over state-of-the-art approaches. The source code is available athttps://github.com/yangxingrui/FLS. Shuqin Chen, Xian Zhong, Xingrui Yang 0003, Bin Sheng 0001, Alex Chichung Kot |
IEEE Trans. Multim. | 6 |
| 2025 | Backdoor Attacks Against No-Reference Image Quality Assessment Models via a Scalable TriggerabstractNo-Reference Image Quality Assessment (NR-IQA), responsible for assessing the quality of a single input image without using any reference, plays a critical role in evaluating and optimizing computer vision systems, e.g., low-light enhancement. Recent research indicates that NR-IQA models are susceptible to adversarial attacks, which can significantly alter predicted scores with visually imperceptible perturbations. Despite revealing vulnerabilities, these attack methods have limitations, including high computational demands, untargeted manipulation, limited practical utility in white-box scenarios, and reduced effectiveness in black-box scenarios. To address these challenges, we shift our focus to another significant threat and present a novel poisoning-based backdoor attack against NR-IQA (BAIQA), allowing the attacker to manipulate the IQA model's output to any desired target value by simply adjusting a scaling coefficient alpha for the trigger. We propose to inject the trigger in the discrete cosine transform (DCT) domain to improve the local invariance of the trigger for countering trigger diminishment in NR-IQA models due to widely adopted data augmentations. Furthermore, the universal adversarial perturbations (UAP) in the DCT space are designed as the trigger, to increase IQA model susceptibility to manipulation and improve attack effectiveness. In addition to the heuristic method for poison-label BAIQA (P-BAIQA), we explore the design of clean-label BAIQA (C-BAIQA), focusing on alpha sampling and image data refinement, driven by theoretical insights we reveal. Extensive experiments on diverse datasets and various NR-IQA models demonstrate the effectiveness of our attacks. Yi Yu 0011, Song Xia, Xun Lin, Wenhan Yang, Shijian Lu, Yap-Peng Tan, Alex Chichung Kot |
AAAI | 7 |
| 2025 | Reconciling Stochastic and Deterministic Strategies for Zero-shot Image Restoration using Diffusion Model in DualabstractPlug-and-play (PnP) methods offer an iterative strategy for solving image restoration (IR) problems in a zero-shot manner, using a learned discriminative denoiser as the implicit prior. More recently, a sampling-based variant of this approach, which utilizes a pre-trained generative diffusion model, has gained great popularity for solving IR problems through stochastic sampling. The IR results using PnP with a pre-trained diffusion model demonstrate distinct advantages compared to those using discriminative denoisers, i.e.,improved perceptual quality while sacrificing the data fidelity. The unsatisfactory results are due to the lack of integration of these strategies in the IR tasks. In this work, we propose a novel zero-shot IR scheme, dubbed Reconciling Diffusion Model in Dual (RDMD), which leverages only a single pre-trained diffusion model to construct two complementary regularizers. Specifically, the diffusion model in RDMD will iteratively perform deterministic denoising and stochastic sampling, aiming to achieve highfidelity image restoration with appealing perceptual quality. RDMD also allows users to customize the distortion-perception tradeoff with a single hyperparameter, enhancing the adaptability of the restoration process in different practical scenarios. Extensive experiments on several IR tasks demonstrate that our proposed method could achieve superior results compared to existing approaches on both the FFHQ and ImageNet datasets. Code is available at https://github.com/chongwang1024/rdmd. Chong Wang 0011, Lanqing Guo, Zixuan Fu, Siyuan Yang 0001, Hao Cheng 0016, Alex Chichung Kot, Bihan Wen |
CVPR | 6 |
| 2025 | Pay Attention to the Foreground in Object-Centric LearningabstractThe slot attention-based method is widely used in unsupervised object-centric learning, aiming to decompose scenes into interpretable objects and associate them with slots. However, complex backgrounds in the real images can disrupt the model’s focus, leading it to excessively segment background stuff into different regions based on low-level information such as color or texture variations. As a result, the detailed segmentation of foreground objects, which requires shape or geometric information, is often neglected. To address this issue, we introduce a contrastive learning-based indicator designed to differentiate between foreground and background. Integrating this indicator into the existing slot attention-based method enables the model to focus more on segmenting foreground objects while minimizing background distractions. During the testing phase, we utilize a spectral clustering mechanism to refine the results based on the similarity between the slots. Experimental results show that incorporating our method with various state-of-the-art models significantly improves their performance on both simulated data and real-world datasets. Furthermore, multiple sets of ablation experiments confirm the effectiveness of each proposed component. The source code is available at https://github.com/sjyjs09/FG-BG_Indicator. Pinzhuo Tian, Shengjie Yang, Hang Yu 0006, Alex Chichung Kot |
CVPR | 4 |
| 2025 | Theoretical Insights in Model Inversion Robustness and Conditional Entropy Maximization for Collaborative Inference SystemsabstractBy locally encoding raw data into intermediate features, collaborative inference enables end users to leverage powerful deep learning models without exposure of sensitive raw data to cloud servers. However, recent studies have revealed that these intermediate features may not sufficiently preserve privacy, as information can be leaked and raw data can be reconstructed via model inversion attacks (MIAs). Obfuscation-based methods, such as noise corruption, adversarial representation learning, and information filters, enhance the inversion robustness by obfuscating the task-irrelevant redundancy empirically. However, methods for quantifying such redundancy remain elusive, and the explicit mathematical relation between this redundancy minimization and inversion robustness enhancement has not yet been established. To address that, this work first theoretically proves that the conditional entropy of inputs given intermediate features provides a guaranteed lower bound on the reconstruction mean square error (MSE) under any MIA. Then, we derive a differentiable and solvable measure for bounding this conditional entropy based on the Gaussian mixture estimation and propose a conditional entropy maximization (CEM) algorithm to enhance the inversion robustness. Experimental results on four datasets demonstrate the effectiveness and adaptability of our proposed CEM; without compromising feature utility and computing efficiency, plugging the proposed CEM into obfuscation-based defense mechanisms consistently boosts their inversion robustness, achieving average gains ranging from 12.9% to 48.2%. Code is available at https://github.com/xiasong0501/CEM. Song Xia, Yi Yu 0011, Wenhan Yang, Meiwen Ding, Zhuo Chen 0006, Ling-Yu Duan, Alex Chichung Kot, Xudong Jiang 0001 |
CVPR | 7 |
| 2025 | Vid-Group: Temporal Video Grounding Pretraining from Unlabeled Videos in the Wild
Peijun Bao, Chenqi Kong, Siyuan Yang 0001, Zihao Shao, Xinghao Jiang, Boon Poh Ng, Meng Hwa Er, Alex Chichung Kot |
ICCV | 8 |
| 2025 | Temporal Unlearnable Examples: Preventing Personal Video Data from Unauthorized Exploitation by Object TrackingabstractWith the rise of social media, vast amounts of user-uploaded videos (e.g., YouTube) are utilized as training data for Visual Object Tracking (VOT). However, the VOT community has largely overlooked video data-privacy issues, as many private videos have been collected and used for training commercial models without authorization. To alleviate these issues, this paper presents the first investigation on preventing personal video data from unauthorized exploitation by deep trackers. Existing methods for preventing unauthorized data use primarily focus on image-based tasks (e.g., image classification), directly applying them to videos reveals several limitations, including inefficiency, limited effectiveness, and poor generalizability. To address these issues, we propose a novel generative framework for generating Temporal Unlearnable Examples (TUEs), and whose efficient computation makes it scalable for usage on large-scale video datasets. The trackers trained w/ TUEs heavily rely on unlearnable noises for temporal matching, ignoring the original data structure and thus ensuring training video data-privacy. To enhance the effectiveness of TUEs, we introduce a temporal contrastive loss, which further corrupts the learning of existing trackers when using our TUEs for training. Extensive experiments demonstrate that our approach achieves state-of-the-art performance in video data-privacy protection, with strong transferability across VOT models, datasets, and temporal matching tasks. Qiangqiang Wu, Yi Yu 0011, Chenqi Kong, Ziquan Liu, Jia Wan 0001, Haoliang Li, Alex Chichung Kot, Antoni B. Chan |
ICCV | 7 |
| 2025 | Towards Effective and Robust Unlearnable Examples Against Object DetectionabstractObject detection has become crucial due to its extensive applications across various industries. However, the data used to train these models is often sensitive and proprietary, raising significant concerns about its security and unauthorized usage. Unlearnable examples (UEs) represent a promising strategy to safeguard proprietary datasets by embedding imperceptible perturbations that degrade model performance when such data is used during training. This paper explores UEs tailored specifically for object detection tasks, which pose unique challenges due to the multi-task nature of object detection. We propose a novel framework that generates robust and effective UEs, and significantly degrades object detector performance while maintaining imperceptibility. Comprehensive experiments demonstrate the resilience of the proposed UEs against various countermeasures, underscoring their potential as a practical solution for protecting data in object detection. Chenyu Yi, Ruohan Meng, Haohang Peng, Bingquan Shen, Alex Chichung Kot |
ICIP | 5 |
| 2025 | MTL-UE: Learning to Learn Nothing for Multi-Task LearningabstractMost existing unlearnable strategies focus on preventing unauthorized users from training single-task learning (STL) models with personal data. Nevertheless, the paradigm has recently shifted towards multi-task data and multi-task learning (MTL), targeting generalist and foundation models that can handle multiple tasks simultaneously. Despite their growing importance, MTL data and models have been largely neglected while pursuing unlearnable strategies. This paper presents MTL-UE, the first unified framework for generating unlearnable examples for multi-task data and MTL models. Instead of optimizing perturbations for each sample, we design a generator-based structure that introduces label priors and class-wise feature embeddings which leads to much better attacking performance. In addition, MTL-UE incorporates intra-task and inter-task embedding regularization to increase inter-class separation and suppress intra-class variance which enhances the attack robustness greatly. Furthermore, MTL-UE is versatile with good supports for dense prediction tasks in MTL. It is also plug-and-play allowing integrating existing surrogate-dependent unlearnable methods with little adaptation. Extensive experiments show that MTL-UE achieves superior attacking performance consistently across 4 MTL datasets, 3 base UE methods, 5 model backbones, and 5 MTL task-weighting strategies. Code is available at https://github.com/yuyi-sd/MTL-UE. Yi Yu 0011, Song Xia, Siyuan Yang 0001, Chenqi Kong, Wenhan Yang, Shijian Lu, Yap-Peng Tan, Alex Chichung Kot |
ICML | 8 |
| 2025 | HRHuman: Tuning-Free Higher-Resolution Human Image Generation via Template KnowledgeabstractHigh-resolution human-centric image generation offers significant potential across various industries, such as entertainment, media, and fashion. Diffusion models for text-to-image generation have significantly improved the quality of human image synthesis. However, when scaling to higher resolutions (2K, 4K, and above), they often encounter issues such as object repetition and structural distortion, which appear especially unnatural in human images. To address these challenges, we propose HRHuman, a tuning-free framework for Higher-Resolution Human Image Generation. By leveraging an open-source large human vision model that incorporates rich template knowledge as prior, we first introduce the Prompt Discretization scheme to discretize user-input text prompts, mapping them to image elements and human body parts. Additionally, we implement a Template-guided Prompt Filtering mechanism to align these discretized prompts with regional image semantics, ensuring fine-grained prompt guidance. Extensive experiments demonstrate that HRHuman achieves state-of-the-art performance in human-centeric higher-resolution image generation, significantly addressing both issues of object repetition and structural distortion. Ling Li 0012, Lanqing Guo, Siyuan Yang 0001, Yakun Ju, Weisi Lin, Alex Chichung Kot |
ISCAS | 7 |
| 2025 | Towards Data-Centric Face Anti-spoofing: Improving Cross-Domain Generalization via Physics-Based Data Synthesis
Rizhao Cai, Cecelia Soh, Zitong Yu, Haoliang Li, Wenhan Yang, Alex Chichung Kot |
Int. J. Comput. Vis. | 6 |
| 2025 | Mining Generalized Multi-timescale Inconsistency for Detecting Deepfake Videos
Yang Yu 0039, Siyuan Yang 0001, Yu Ni, Yao Zhao 0001, Alex Chichung Kot |
Int. J. Comput. Vis. | 6 |
| 2025 | Face reconstruction with detailed skin features via three selfie images
Yakun Ju, Bandara Dissanayake, Rachel Ang, Ling Li 0012, Dennis Sng, Alex Chichung Kot |
J. Vis. Commun. Image Represent. | 6 |
| 2025 | Rehearsal-Free and Efficient Continual Learning for Cross-Domain Face Anti-SpoofingabstractFace Anti-Spoofing (FAS) is constantly challenged by new attack types and mediums, and thus it is crucial for a FAS model to not only mitigate Catastrophic Forgetting (CF) of previously learned spoofing knowledge on the training data during continual learning but also enhance the model's generalization ability to potential spoofing attacks. In this paper, we first highlight that current strategies for catastrophic forgetting are not well-suited to the imperceptible nature of spoofing information in FAS and lack the focus on improving generalization capability. Then, the instance-wise dynamic central difference convolutional adapter module with the weighted ensemble strategy for Vision Transformer (ViT) is proposed for efficiently fine-tuning with low-shot data by extracting generalized spoofing texture information. Furthermore, we find that catastrophic forgetting in FAS can be reflected through the inconsistent attention matrices of ViT between different continual sessions, as the attention matrices embody relationships of spoofing clues between different patch tokens. Hence, we introduce attention consistency regularization by learning and reusing attention matrices to alleviate catastrophic forgetting. Finally, we devise new protocols and conduct extensive experiments to validate the superior performance of alleviating catastrophic forgetting and generalization on unseen domains. Rizhao Cai, Yawen Cui, Zitong Yu, Xun Lin, Changsheng Chen 0001, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Revisiting One-Stage Deep Uncalibrated Photometric Stereo via Fourier EmbeddingabstractThis paper introduces a one-stage deep uncalibrated photometric stereo (UPS) network, namely Fourier Uncalibrated Photometric Stereo Network (FUPS-Net), for non-Lambertian objects under unknown light directions. It departs from traditional two-stage methods that first explicitly learn lighting information and then estimate surface normals. Two-stage methods were deployed because the interplay of lighting with shading cues presents challenges for directly estimating surface normals without explicit lighting information. However, these two-stage networks are disjointed and separately trained so that the error in explicit light calibration will propagate to the second stage and cannot be eliminated. In contrast, the proposed FUPS-Net utilizes an embedded Fourier transform network to implicitly learn lighting features by decomposing inputs, rather than employing a disjointed light estimation network. Our approach is motivated from observations in the Fourier domain of photometric stereo images: lighting information is mainly encoded in amplitudes, while geometry information is mainly associated with phases. Leveraging this property, our method "decomposes" geometry and lighting in the Fourier domain as guidance, via the proposed Fourier Embedding Extraction (FEE) block and Fourier Embedding Aggregation (FEA) block, which generate lighting and geometry features for the FUPS-Net to implicitly resolve the geometry-lighting ambiguity. Furthermore, we propose a Frequency-Spatial Weighted (FSW) block that assigns weights to combine features extracted from the frequency domain and those from the spatial domain for enhancing surface reconstructions. FUPS-Net overcomes the limitations of two-stage UPS methods, offering better training stability, a concise end-to-end structure, and avoiding accumulated errors in disjointed networks. Experimental results on synthetic and real datasets demonstrate the superior performance of our approach, and its simpler training setup, potentially paving the way for a new strategy in deep learning-based UPS methods. Yakun Ju, Boxin Shi, Bihan Wen, Kin-Man Lam 0001, Xudong Jiang 0001, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Pixel-Inconsistency Modeling for Image Manipulation LocalizationabstractDigital image forensics plays a crucial role in image authentication and manipulation localization. Despite the progress powered by deep neural networks, existing forgery localization methodologies exhibit limitations when deployed to unseen datasets and perturbed images (i.e., lack of generalization and robustness to real-world applications). To circumvent these problems and aid image integrity, this paper presents a generalized and robust manipulation localization model through the analysis of pixel inconsistency artifacts. The rationale is grounded on the observation that most image signal processors (ISP) involve the demosaicing process, which introduces pixel correlations in pristine images. Moreover, manipulating operations, including splicing, copy-move, and inpainting, directly affect such pixel regularity. We, therefore, first split the input image into several blocks and design masked self-attention mechanisms to model the global pixel dependency in input images. Simultaneously, we optimize another local pixel dependency stream to mine local manipulation clues within input forgery images. In addition, we design novel Learning-to-Weight Modules (LWM) to combine features from the two streams, thereby enhancing the final forgery localization performance. To improve the training process, we propose a novel Pixel-Inconsistency Data Augmentation (PIDA) strategy, driving the model to focus on capturing inherent pixel-level artifacts instead of mining semantic forgery traces. This work establishes a comprehensive benchmark integrating 16 representative detection models across 12 datasets. Extensive experiments show that our method successfully extracts inherent pixel-inconsistency forgery fingerprints and achieve state-of-the-art generalization and robustness performances in image manipulation localization. Chenqi Kong, Anwei Luo, Shiqi Wang 0001, Haoliang Li, Anderson Rocha 0001, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Reliable and Balanced Transfer Learning for Generalized Multimodal Face Anti-SpoofingabstractFace Anti-Spoofing (FAS) is essential for securing face recognition systems against presentation attacks. Recent advances in sensor technology and multimodal learning have enabled the development of multimodal FAS systems. However, existing methods often struggle to generalize to unseen attacks and diverse environments due to two key challenges: (1) Modality unreliability, where sensors such as depth and infrared suffer from severe domain shifts, impairing the reliability of cross-modal fusion; and (2) Modality imbalance, where over-reliance on a dominant modality weakens the model's robustness against attacks that affect other modalities. To overcome these issues, we propose MMDG++, a multimodal domain-generalized FAS framework built upon the vision-language model CLIP. In MMDG++, we design the Uncertainty-Guided Cross-Adapter++ (U-Adapter++) to filter out unreliable regions within each modality, enabling more reliable multimodal interactions. Additionally, we introduce Rebalanced Modality Gradient Modulation (ReGrad) for adaptive gradient modulation to balance modality convergence. To further enhance generalization, propose Asymmetric Domain Prompts (ADPs) that leverage CLIP's language priors to learn generalized decision boundaries across modalities. We also develop a novel multimodal FAS benchmark to evaluate generalizability under various deployment conditions. Extensive experiments across this benchmark show our method outperforms state-of-the-art FAS methods, demonstrating superior generalization capability. Xun Lin, Ajian Liu 0001, Zitong Yu, Rizhao Cai, Shuai Wang 0049, Yi Yu 0011, Jun Wan 0001, Zhen Lei 0001, Xiaochun Cao, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 10 |
| 2025 | Robust and Transferable Backdoor Attacks Against Deep Image Compression With Selective Frequency PriorabstractRecent advancements in deep learning-based compression techniques have demonstrated remarkable performance surpassing traditional methods. Nevertheless, deep neural networks have been observed to be vulnerable to backdoor attacks, where an added pre-defined trigger pattern can induce the malicious behavior of the models. In this paper, we propose a novel approach to launch a backdoor attack with multiple triggers against learned image compression models. Drawing inspiration from the widely used discrete cosine transform (DCT) in existing compression codecs and standards, we propose a frequency-based trigger injection model that adds triggers in the DCT domain. In particular, we design several attack objectives that are adapted for a series of diverse scenarios, including: 1) attacking compression quality in terms of bit-rate and reconstruction quality; 2) attacking task-driven measures, such as face recognition and semantic segmentation in downstream applications. To facilitate more efficient training, we develop a dynamic loss function that dynamically balances the impact of different loss terms with fewer hyper-parameters, which also results in more effective optimization of the attack objectives with improved performance. Furthermore, we consider several advanced scenarios. We evaluate the resistance of the proposed backdoor attack to the defensive pre-processing methods and then propose a two-stage training schedule along with the design of robust frequency selection to further improve resistance. To strengthen both the cross-model and cross-domain transferability on attacking downstream CV tasks, we propose to shift the classification boundary in the attack loss during training. Extensive experiments also demonstrate that by employing our trained trigger injection models and making slight modifications to the encoder parameters of the compression model, our proposed attack can successfully inject multiple backdoors accompanied by their corresponding triggers into a single image compression model. Yi Yu 0011, Yufei Wang 0006, Wenhan Yang, Lanqing Guo, Shijian Lu, Ling-Yu Duan, Yap-Peng Tan, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2025 | Clean-Label Attack on Face Authentication Systems Through Rolling Shutter MechanismabstractWe introduce a novel clean-label black-box face presentation attack on face authentication systems, i.e., face recognition and verification systems, under mild conditions. Different from other clean-label attacks which require inserting complicated or intensity patterns after the image-capturing phase, our designed pattern can be automatically inserted during the exposure by utilizing the rolling shutter mechanism and modulating environment LEDs in a specialized waveform. This method provides a potential way to conduct backdoor attacks in the physical domain. Additionally, we propose an optimization strategy based on evolutionary computing to optimize the parameters of the stripe patterns, enhancing the attack success rate. The experimental results on several face recognition models and face verification services provided by the leading technology companies demonstrate the effectiveness of our attack method. Our study reveals a new attack applicable in the physical world, highlighting significant security concerns for existing face recognition, verification, and face anti-spoofing techniques. Yufei Wang 0006, Haoliang Li, Liepiao Zhang, Yongjian Hu, Alex Chichung Kot |
IEEE Signal Process. Lett. | 5 |
| 2025 | CEAT: Continual Expansion and Absorption Transformer for Non-Exemplar Class-Incremental LearningabstractIn dynamic real-world scenarios, continuous learning without forgetting old knowledge is essential, particularly in environments with stricter privacy protection or resource-constrained edge devices where storing old exemplars is infeasible. Therefore, Non-Exemplar Class-Incremental Learning (NECIL) has garnered significant attention. Compared with normal settings, it faces a more severe plasticity-stability dilemma and classifier bias. To address those challenges, we propose a framework based on the vision transformer architecture, called the Continual Expansion and Absorption Transformer (CEAT), which consists of two core components. First, we propose the Continual Expansion and Absorption (CEA) method to alleviate the trade-off between new and old classes by parallelly expanding a set of parameters (i.e. EF layer) on the backbone to learn new tasks, while freezing the backbone to retain old task knowledge. The EF layers can be seamlessly absorbed into the ViT backbone through parameter recombination before inference, mitigating storage and computational burdens. Second, we propose a Dynamic Boundary-Aware (DBA) method to generate dynamic pseudo-features for classifier calibration to address the classifier bias. Extensive experiments demonstrate that our approach achieves state-of-the-art performance, particularly showcasing significant improvements of 4.82% and 5.92% on TinyImageNet and ImageNet-Subset, respectively. Songlin Dong, Xinyuan Gao, Yuhang He 0001, Zhengdong Zhou, Alex Chichung Kot, Yihong Gong |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Forgery-Aware Adaptive Learning With Vision Transformer for Generalized Face Forgery DetectionabstractWith the rapid progress of generative models, the current challenge in face forgery detection is how to effectively detect realistic manipulated faces from different unseen domains. Though previous studies show that pre-trained Vision Transformer (ViT) based models can achieve some promising results after fully fine-tuning on the Deepfake dataset, their generalization performances are still unsatisfactory. To this end, we present a Forgery-aware Adaptive Vision Transformer (FA-ViT) under the adaptive learning paradigm for generalized face forgery detection, where the parameters in the pre-trained ViT are kept fixed while the designed adaptive modules are optimized to capture forgery features. Specifically, a global adaptive module is designed to model long-range interactions among input tokens, which takes advantage of self-attention mechanism to mine global forgery clues. To further explore essential local forgery clues, a local adaptive module is proposed to expose local inconsistencies by enhancing the local contextual association. In addition, we introduce a fine-grained adaptive learning module that emphasizes the common compact representation of genuine faces through relationship learning in fine-grained pairs, driving these proposed adaptive modules to be aware of fine-grained forgery-aware information. Extensive experiments demonstrate that our FA-ViT achieves state-of-the-arts results in the cross-dataset evaluation, and enhances the robustness against unseen perturbations. Particularly, FA-ViT achieves 93.83% and 78.32% AUC scores on Celeb-DF and DFDC datasets in the cross-dataset evaluation. The code and trained model have been released at:https://github.com/LoveSiameseCat/FAViT. Anwei Luo, Rizhao Cai, Chenqi Kong, Yakun Ju, Xiangui Kang, Jiwu Huang, Alex Chichung Kot |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Facial Image Compression via Neural Image Manifold CompressionabstractAlthough the recent learning-based image and video coding techniques achieve rapid development, the signal fidelity-driven target in these methods leads to the divergence to a highly effective and efficient coding framework for both human and machine. In this paper, we aim to address the issue by making use of the power of generative models to bridge the gap between full fidelity (for human vision) and high discrimination (for machine vision). Therefore, relying on existing pretrained generative adversarial networks (GAN), we build a GAN inversion framework that projects the image into a low-dimensional natural image manifold. In this manifold, the feature is highly discriminative and also encodes the appearance information of the image, named aslatent code. Taking a variational bit-rate constraint with a hyperprior model to model/suppress the entropy of image manifold code, our method is capable of fulfilling the needs of both machine and human visions at very low bit-rates. To improve the visual quality of image reconstruction, we further proposemultiple latent codesandscalable inversion. The former gets several latent codes in the inversion, while the latter additionally compresses and transmits a shallow compact feature to support visual reconstruction. Experimental results demonstrate the superiority of our method in both human vision tasks,i.e. image reconstruction, and machine vision tasks, including semantic parsing and attribute prediction. Wenhan Yang, Haofeng Huang, Jiaying Liu 0001, Alex Chichung Kot |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Prototype-Driven Structure Synergy Network for Remote Sensing Images Segmentation
Jinjiang Li 0001, Yakun Ju, Alex Chichung Kot |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Toward Model Resistant to Transferable Adversarial Examples via Trigger ActivationabstractAdversarial examples, characterized by imperceptible perturbations, pose significant threats to deep neural networks by misleading their predictions. A critical aspect of these examples is their transferability, allowing them to deceive unseen models in closed-box scenarios. Despite the widespread exploration of defense methods, including those on transferability, they show limitations: inefficient deployment, ineffective defense, and degraded performance on clean images. In this work, we introduce a novel training paradigm aimed at enhancing robustness against transferable adversarial examples (TAEs) in a more efficient and effective way. We propose a model that exhibits random guessing behavior when presented with clean data$\boldsymbol {x}$as input, and generates accurate predictions when with triggered data$\boldsymbol {x}+\boldsymbol {\tau }$. Importantly, the trigger$\boldsymbol {\tau }$remains constant for all data instances. We refer to these models as models with trigger activation. We are surprised to find that these models exhibit certain robustness against TAEs. Through the consideration of first-order gradients, we provide a theoretical analysis of this robustness. Moreover, through the joint optimization of the learnable trigger and the model, we achieve improved robustness to transferable attacks. Extensive experiments conducted across diverse datasets, evaluating a variety of attacking methods, underscore the effectiveness and superiority of our approach. Yi Yu 0011, Song Xia, Xun Lin, Chenqi Kong, Wenhan Yang, Shijian Lu, Yap-Peng Tan, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 8 |
| 2025 | Cross-Frequency Attention and Color Contrast Constraint for Remote Sensing DehazingabstractCurrent deep learning-based methods for remote sensing image dehazing have developed rapidly, yet they still commonly struggle to simultaneously preserve fine texture details and restore accurate colors. The fundamental reason lies in the insufficient modeling of high-frequency information that captures structural details, as well as the lack of effective constraints for color restoration. To address the insufficient modeling of global high-frequency information, we first develop an omni-directional high-frequency feature in painting mechanism that leverages the wavelet transform to extract multi-directional high-frequency components. While maintaining the advantage of linear complexity, it models global long-range texture dependencies through cross-frequency perception. Then, to further strengthen local high-frequency representation, we design a high-frequency prompt attention module that dynamically injects wavelet-domain optimized high-frequency features as cross-level guidance signals, significantly enhancing the model's capability in edge sharpness restoration and texture detail reconstruction. Further, to alleviate the problem of inaccurate color restoration, we propose a color contrast loss function based on the HSV color space, which explicitly models the statistical distribution differences of brightness and saturation in hazy regions, guiding the model to generate dehazed images with consistent colors and natural visual appearance. Finally, extensive experiments on multiple benchmark datasets demonstrate that the proposed method outperforms existing approaches in both texture detail restoration and color consistency. Further results and code are available at: https://github.com/fyxnl/C4RSD. Jufeng Li, Yakun Ju, Chunxu Li, Weisheng Dong, Alex Chichung Kot |
IEEE Trans. Image Process. | 8 |
| 2025 | Digital Staining With Knowledge Distillation: A Unified Framework for Unpaired and Paired-but-Misaligned DataabstractStaining is essential in cell imaging and medical diagnostics but poses significant challenges, including high cost, time consumption, labor intensity, and irreversible tissue alterations. Recent advances in deep learning have enabled digital staining through supervised model training. However, collecting large-scale, perfectly aligned pairs of stained and unstained images remains difficult. In this work, we propose a novel unsupervised deep learning framework for digital cell staining that reduces the need for extensive paired data using knowledge distillation. We explore two training schemes: (1) unpaired and (2) paired-but-misaligned settings. For the unpaired case, we introduce a two-stage pipeline, comprising light enhancement followed by colorization, as a teacher model. Subsequently, we obtain a student staining generator through knowledge distillation with hybrid non-reference losses. To leverage the pixel-wise information between adjacent sections, we further extend to the paired-but-misaligned setting, adding the Learning to Align module to utilize pixel-level information. Experiment results on our dataset demonstrate that our proposed unsupervised deep staining method can generate stained images with more accurate positions and shapes of the cell targets in both settings. Compared with competing methods, our method achieves improved results both qualitatively and quantitatively (e.g., NIQE and PSNR). We applied our digital staining method to the White Blood Cell (WBC) dataset, investigating its potential for medical applications. Ziwang Xu, Lanqing Guo, Satoshi Tsutsui, Alex Chichung Kot, Bihan Wen |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Debiasing Medical Knowledge for Prompting Universal Model in CT Image SegmentationabstractWith the assistance of large language models, which offer universal medical prior knowledge via text prompts, state-of-the-art Universal Models (UM) have demonstrated considerable potential in the field of medical image segmentation. Semantically detailed text prompts, on the one hand, indicate comprehensive knowledge; on the other hand, they bring biases that may not be applicable to specific cases involving heterogeneous organs or rare cancers. To this end, we propose a Debiased Universal Model (DUM) to consider instance-level context information and remove knowledge biases in text prompts from the causal perspective. We are the first to discover and mitigate the bias introduced by universal knowledge. Specifically, we propose to extract organ-level text prompts via language models and instance-level context prompts from the visual features of each image. We aim to highlight more on factual instance-level information and mitigate organ-level's knowledge bias. This process can be derived and theoretically supported by a causal graph, and instantiated by designing a standard UM (SUM) and a biased UM. The debiased output is finally obtained by subtracting the likelihood distribution output by biased UM from that of the SUM. Experiments on three large-scale multi-center external datasets and MSD internal tumor datasets show that our method enhances the model's generalization ability in handling diverse medical scenarios and reducing the potential biases, even with an improvement of 4.16% compared with popular universal model on the AbdomenAtlas dataset, showing the strong generalizability. The code is publicly available at https://github.com/DeepMed-Lab-ECNU/DUM. Boxiang Yun, Shitian Zhao, Qingli Li, Alex Chichung Kot, Yan Wang 0033 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | CMoA: Contrastive Mixture of Adapters for Generalized Few-Shot Continual LearningabstractThe goal of Few-Shot Continual Learning (FSCL) is to incrementally learn novel tasks with limited labeled samples and preserve previous capabilities simultaneously. However, current FSCL works lack research on domain increment and domain generalization ability, which cannot cope with changes in the visual perception environment. In this paper, we set up a Generalized FSCL (GFSCL) protocol involving both class- and domain-incremental scenarios together with domain generalization assessment. Firstly, two benchmark datasets and protocols are newly arranged, and detailed baselines are provided for this unexplored configuration. Furthermore, we find that common continual learning methods have poor generalization ability on unseen domains and cannot better tackle catastrophic forgetting issue in cross-incremental tasks. Hence, we propose a rehearsal-free framework based on Vision Transformer (ViT) named Contrastive Mixture of Adapters (CMoA). It contains two non-conflicting parts: (1) By applying the fast-adaptation characteristic of adapter-embedded ViT, the mixture of Adapters (MoA) module is incorporated into ViT. For stability purpose, cosine similarity regularization and dynamic weighting are designed to make each adapter learn specific knowledge and concentrate on particular classes. (2) To further enhance domain generalization ability, we alleviate the intra-class variation by prototype-calibrated contrastive learning to improve domain-invariant representation learning. Finally, six evaluation indicators showing the overall performance and forgetting are compared by comprehensive experiments on two benchmark datasets to validate the efficacy of CMoA, and the results illustrate that CMoA can achieve comparative performance with rehearsal-based continual learning methods. Yawen Cui, Jian Zhao 0006, Zitong Yu, Rizhao Cai, Lei Jin 0003, Alex Chichung Kot, Li Liu 0002, Xuelong Li 0001 |
IEEE Trans. Multim. | 7 |
| 2025 | Analogical Augmentation and Significance Analysis for Online Task-Free Continual LearningabstractOnline task-free continual learning (OTFCL) is a more challenging variant of continual learning that emphasizes the gradual shift of task boundaries and learning in an online mode. Existing methods rely on a memory buffer of old samples to prevent forgetting. However, the use of memory buffers not only raises privacy concerns but also hinders the efficient learning of new samples. To address this problem, we propose a novel framework called I$^{2}$CANSAY that gets rid of the dependence on memory buffers and efficiently learns the knowledge of new data from one-shot samples. Concretely, our framework comprises two main modules. Firstly, theInter-Class Analogical Augmentation(ICAN) module generates diverse pseudo-features for old classes based on the inter-class analogy of feature distributions for different new classes, serving as a substitute for the memory buffer. Secondly, theIntra-Class Significance Analysis(ISAY) module analyzes the significance of attributes for each class via its distribution standard deviation, and generates an importance vector as a correction bias for the linear classifier, thereby enhancing the capability of learning from new samples. We run our experiments on four popular image classification datasets: CoRe50, CIFAR-10, CIFAR-100, and CUB-200, our approach outperforms the prior state-of-the-art by a large margin. Songlin Dong, Yuhang He 0001, Yuhan Jin, Alex Chichung Kot, Yihong Gong |
IEEE Trans. Multim. | 5 |
| 2025 | ContextualCoder: Adaptive In-Context Prompting for Programmatic Visual Question AnsweringabstractVisual Question Answering (VQA) presents a challenging task at the intersection of computer vision and natural language processing, aiming to bridge the semantic gap between visual perception and linguistic comprehension. Traditional VQA approaches do not distinguish between data processing and reasoning, limiting their interpretability and generalizability in complex and diverse scenarios. Conversely, Programmatic Visual Question Answering (PVQA) models leverage large language models (LLMs) to generate executable codes, providing answers with detailed and interpretable reasoning processes. However, existing PVQA models typically rely on simplistic input-output prompting, which struggles to elicit domain-specific knowledge from LLMs and often produces unclear or extraneous outputs. Furthermore, PVQA models typically rely on a basic in-context example (ICE) selection methodology that is heavily influenced by individual word similarity rather than the overall sentence context. This leads to suboptimal ICE selection and a reliance on dataset-specific ICE candidates. In this paper, we propose ContextualCoder, a novel prompting framework tailored for PVQA models. ContextualCoder leverages frozen LLMs for code generation and pre-trained visual models for code execution, eliminating the need for extensive training and enhancing model flexibility. By incorporating an innovative prompting methodology and a novel ICE selection strategy, ContextualCoder facilitates the use of diverse in-context information for code generation, thereby improving the performance of PVQA models. Our approach surpasses state-of-the-art models, as evidenced by comprehensive experiments across diverse VQA datasets, including multilingual scenarios. Ruoyue Shen, Nakamasa Inoue, Dayan Guan, Rizhao Cai, Alex Chichung Kot, Koichi Shinoda |
IEEE Trans. Multim. | 5 |
| 2024 | Omnipotent Distillation with LLMs for Weakly-Supervised Natural Language Video Localization: When Divergence Meets ConsistencyabstractNatural language video localization plays a pivotal role in video understanding, and leveraging weakly-labeled data is considered a promising approach to circumvent the laborintensive process of manual annotations. However, this approach encounters two significant challenges: 1) limited input distribution, namely that the limited writing styles of the language query, annotated by human annotators, hinder the model’s generalization to real-world scenarios with diverse vocabularies and sentence structures; 2) the incomplete ground truth, whose supervision guidance is insufficient. To overcome these challenges, we propose an omnipotent distillation algorithm with large language models (LLM). The distribution of the input sample is enriched to obtain diverse multi-view versions while a consistency then comes to regularize the consistency of their results for distillation. Specifically, we first train our teacher model with the proposed intra-model agreement, where multiple sub-models are supervised by each other. Then, we leverage the LLM to paraphrase the language query and distill the teacher model to a lightweight student model by enforcing the consistency between the localization results of the paraphrased sentence and the original one. In addition, to assess the generalization of the model across different dimensions of language variation, we create extensive datasets by building upon existing datasets. Our experiments demonstrate substantial performance improvements adaptively to diverse kinds of language queries. Peijun Bao, Zihao Shao, Wenhan Yang, Boon Poh Ng, Meng Hwa Er, Alex Chichung Kot |
AAAI | 6 |
| 2024 | Local-Global Multi-Modal Distillation for Weakly-Supervised Temporal Video GroundingabstractThis paper for the first time leverages multi-modal videos for weakly-supervised temporal video grounding. As labeling the video moment is labor-intensive and subjective, the weakly-supervised approaches have gained increasing attention in recent years. However, these approaches could inherently compromise performance due to inadequate supervision. Therefore, to tackle this challenge, we for the first time pay attention to exploiting complementary information extracted from multi-modal videos (e.g., RGB frames, optical flows), where richer supervision is naturally introduced in the weaklysupervised context. Our motivation is that by integrating different modalities of the videos, the model is learned from synergic supervision and thereby can attain superior generalization capability. However, addressing multiple modalities† would also inevitably introduce additional computational overhead, and might become inapplicable if a particular modality is inaccessible. To solve this issue, we adopt a novel route: building a multi-modal distillation algorithm to capitalize on the multi-modal knowledge as supervision for model training, while still being able to work with only the single modal input during inference. As such, we can utilize the benefits brought by the supplementary nature of multiple modalities, without compromising the applicability in practical scenarios. Specifically, we first propose a cross-modal mutual learning framework and train a sophisticated teacher model to learn collaboratively from the multi-modal videos. Then we identify two sorts of knowledge from the teacher model, i.e., temporal boundaries and semantic activation map. And we devise a local-global distillation algorithm to transfer this knowledge to a student model of single-modal input at both local and global levels. Extensive experiments on large-scale datasets demonstrate that our method achieves state-of-the-art performance with/without multi-modal inputs. Peijun Bao, Wenhan Yang, Boon Poh Ng, Meng Hwa Er, Alex Chichung Kot |
AAAI | 6 |
| 2024 | Cross-Domain Few-Shot Segmentation via Iterative Support-Query Correspondence MiningabstractCross-Domain Few-Shot Segmentation (CD-FSS) poses the challenge of segmenting novel categories from a distinct domain using only limited exemplars. In this paper, we undertake a comprehensive study of CD-FSS and uncover two crucial insights: (i) the necessity of a fine-tuning stage to effectively transfer the learned meta-knowledge across domains, and (ii) the overfitting risk during the naive fine-tuning due to the scarcity of novel category examples. With these insights, we propose a novel cross-domain fine-tuning strategy that addresses the challenging CD-FSS tasks. We first design Bi-directional Few-shot Prediction (BFP), which establishes support-query correspondence in bi-directional manner, crafting augmented supervision to reduce the overfitting risk. Then we further extend BFP into Iterative Few-shot Adaptor (IFA), which is a recursive framework to capture the support-query correspondence iteratively, targeting maximal exploitation of supervisory signals from the sparse novel category samples. Extensive empirical evaluations show that our method significantly outperforms the state-of-the-arts (+7.8%), which verifies that IFA tackles the cross-domain challenges and mitigates the overfitting simultaneously. Jiahao Nie 0002, Yun Xing 0001, Gongjie Zhang, Pei Yan, Aoran Xiao, Yap-Peng Tan, Alex Chichung Kot, Shijian Lu |
CVPR | 7 |
| 2024 | Suppress and Rebalance: Towards Generalized Multi-Modal Face Anti-SpoofingabstractFace Anti-Spoofing (FAS) is crucial for securing face recognition systems against presentation attacks. With ad-vancements in sensor manufacture and multi-modal learning techniques, many multi-modal FAS approaches have emerged. However, they face challenges in generalizing to unseen attacks and deployment conditions. These chal-lenges arise from (1) modality unreliability, where some modality sensors like depth and infrared undergo signifi-cant domain shifts in varying environments, leading to the spread of unreliable information during cross-modal feature fusion, and (2) modality imbalance, where training overly relies on a dominant modality hinders the conver-gence of others, reducing effectiveness against attack types that are indistinguishable by sorely using the dominant modality. To address modality unreliability, we propose the Uncertainty-Guided Cross-Adapter (U-Adapter) to recognize unreliably detected regions within each modality and suppress the impact of unreliable regions on other modal-ities. For modality imbalance, we propose a Rebalanced Modality Gradient Modulation (ReGrad) strategy to rebal-ance the convergence speed of all modalities by adaptively adjusting their gradients. Besides, we provide the first large-scale benchmark for evaluating multi-modal FAS per-formance under domain generalization scenarios. Exten-sive experiments demonstrate that our method outperforms state-of-the-art methods. Source codes and protocols are released on https://github.com/OMGGGGG/mmdg. Xun Lin, Shuai Wang 0049, Rizhao Cai, Yizhong Liu, Ying Fu 0001, Wenzhong Tang, Zitong Yu, Alex Chichung Kot |
CVPR | 8 |
| 2024 | SinSR: Diffusion-Based Image Super-Resolution in a Single StepabstractWhile super-resolution (SR) methods based on diffusion models exhibit promising results, their practical application is hindered by the substantial number of required inference steps. Recent methods utilize the degraded images in the initial state, thereby shortening the Markov chain. Nevertheless, these solutions either rely on a precise formulation of the degradation process or still necessitate a relatively lengthy generation path (e.g., 15 iterations). To enhance inference speed, we propose a simple yet effective method for achieving single-step SR generation, named SinSR. Specifically, we first derive a deterministic sampling process from the most recent state-of-the-art (SOTA) method for accelerating diffusion-based SR. This allows the mapping between the input random noise and the generated high-resolution image to be obtained in a reduced and acceptable number of inference steps during training. We show that this deterministic mapping can be distilled into a student model that performs SR within only one inference step. Additionally, we propose a novel consistency-preserving loss to simultaneously leverage the ground-truth image during the distillation process, ensuring that the performance of the student model is not solely bound by the feature manifold of the teacher model, resulting in further performance improvement. Extensive experiments conducted on synthetic and real-world datasets demonstrate that the proposed method can achieve comparable or even superior performance compared to both previous SOTA methods and the teacher model, in just one sampling step, resulting in a remarkable up to × 10 speedup for inference. Our code will be released at https://github.com/wyf0912/SinSR/. Yufei Wang 0006, Wenhan Yang, Yaohui Wang 0001, Lanqing Guo, Lap-Pui Chau, Ziwei Liu 0002, Yu Qiao 0001, Alex Chichung Kot, Bihan Wen |
CVPR | 9 |
| 2024 | E3M: Zero-Shot Spatio-Temporal Video Grounding with Expectation-Maximization Multimodal Modulation
Peijun Bao, Zihao Shao, Wenhan Yang, Boon Poh Ng, Alex Chichung Kot |
ECCV (83) | 5 |
| 2024 | BenchLMM: Benchmarking Cross-Style Visual Capability of Large Multimodal Models
Rizhao Cai, Zirui Song, Dayan Guan, Zhenhao Chen, Yaohang Li, Chenyu Yi, Alex Chichung Kot |
ECCV (50) | 8 |
| 2024 | STSP: Spatial-Temporal Subspace Projection for Video Class-Incremental Learning
Hao Cheng 0016, Siyuan Yang 0001, Chong Wang 0011, Joey Tianyi Zhou, Alex Chichung Kot, Bihan Wen |
ECCV (28) | 5 |
| 2024 | Towards Physical World Backdoor Attacks Against Skeleton Action Recognition
Qichen Zheng, Yi Yu 0011, Siyuan Yang 0001, Jun Liu 0036, Kwok-Yan Lam, Alex Chichung Kot |
ECCV (48) | 6 |
| 2024 | Unlearnable Examples Detection via Iterative Filtering
Yi Yu 0011, Qichen Zheng, Siyuan Yang 0001, Wenhan Yang, Jun Liu 0036, Shijian Lu, Yap-Peng Tan, Kwok-Yan Lam, Alex Chichung Kot |
ICANN (10) | 9 |
| 2024 | Flexible-Modal Deception Detection with Audio-Visual AdapterabstractDeception detection within audio-visual modalities is vital across diverse sectors, notably in customs security and multimedia anti-fraud. However, this notable efficacy is lost by the necessity to train and deploy separate models for each conceivable modality scenario, leading to redundancy and inefficiency. Moreover, real-world environments where multi-modal models are deployed often fail to meet these idealized conditions. To overcome these challenges and further elevate performance levels, we propose an advanced Transformer-based framework complemented by an Audio-Visual Adapter (AVA) integrating temporal features from both audio and visual modalities. In addition, we introduce an innovative multi-modal contrastive learning method that is designed to enhance the correlation between uni-modal features and their integrated counterparts within a consistent feature space. Our designed method can deal with the flexible-model scenario instead of deploying different models for various modalities. Empirical evaluations conducted on two benchmark datasets have validated the superiority of our proposed model over other multi-modal fusion techniques, particularly in scenarios characterized by varying and missing modalities. This strongly affirms the effectiveness of our approach in significantly boosting the accuracy of deception detection in complex, real-world multi-modal scenarios. The codes will be released soon. Zhaoxu Li, Zitong Yu, Xun Lin, Nithish Muthuchamy Selvaraj, Xiaobao Guo, Bingquan Shen, Adams Wai-Kin Kong, Alex Chichung Kot |
IJCB | 8 |
| 2024 | Color Space Learning for Cross-Color Person Re-IdentificationabstractThe primary color profile of the same identity is assumed to remain consistent in typical Person Re-identification (Person ReID) tasks. However, this assumption may be invalid in real-world situations and images hold variant color profiles, because of cross-modality cameras or identity with different clothing. To address this issue, we propose Color Space Learning (CSL) for those Cross-Color Person ReID problems. Specifically, CSL guides the model to be less color-sensitive with two modules: Image-level Color-Augmentation and Pixel-level Color-Transformation. The first module increases the color diversity of the inputs and guides the model to focus more on the non-color information. The second module projects every pixel of input images onto a new color space. In addition, we introduce a new Person ReID benchmark across RGB and Infrared modalities, NTU-Corridor, which is the first with privacy agreements from all participants. To evaluate the effectiveness and robustness of our proposed CSL, we evaluate it on several Cross-Color Person ReID benchmarks. Our method surpasses the state-of-the-art methods consistently. The code and benchmark are available at: https://github.com/niejiahao1998/CSL Jiahao Nie 0002, Alex Chichung Kot |
ICME | 3 |
| 2024 | Controllable and Gradual Facial Blemishes Retouching Via Physics-Based ModellingabstractFace retouching aims to remove facial blemishes, such as pigmentation and acne, and still retain fine-grain texture details. Nevertheless, existing methods just remove the blemishes but focus little on realism of the intermediate process, limiting their use more to beautifying facial images on social media rather than being effective tools for simulating changes in facial pigmentation and ance. Motivated by this limitation, we propose our Controllable and Gradual Face Retouching (CGFR). Our CGFR is based on physical modelling, adopting Sum-of-Gaussians to approximate skin subsurface scattering in a decomposed melanin and haemoglobin color space. Our CGFR offers a user-friendly control over the facial blemishes, achieving realistic and gradual blemishes retouching. Experimental results based on actual clinical data shows that CGFR can realistically simulate the blemishes’ gradual recovering process. Chenhao Shuai, Rizhao Cai, Bandara Dissanayake, Amanda Newman, Dayan Guan, Dennis Sng, Ling Li 0012, Alex Chichung Kot |
ICME | 8 |
| 2024 | Compress Clean Signal from Noisy Raw Image: A Self-Supervised ApproachabstractRaw images offer unique advantages in many low-level visual tasks due to their unprocessed nature. However, this unprocessed state accentuates noise, making raw images challenging to compress effectively. Current compression methods often overlook the ubiquitous noise in raw space, leading to increased bitrates and reduced quality. In this paper, we propose a novel raw image compression scheme that selectively compresses the noise-free component of the input, while discarding its real noise using a self-supervised approach. By excluding noise from the bitstream, both the coding efficiency and reconstruction quality are significantly enhanced. We curate an full-day dataset of raw images with calibrated noise parameters and reference images to evaluate the performance of models under a wide range of input signal-noise ratios. Experimental results demonstrate that our method surpasses existing compression techniques, achieving a more advantageous rate-distortion balance with improvements ranging from +2 to +10dB and yielding a bit saving of 2 to 50 times. The code will be released upon paper acceptance. Yufei Wang 0006, Alex Chichung Kot, Bihan Wen |
ICML | 3 |
| 2024 | Purify Unlearnable Examples via Rate-Constrained Variational AutoencodersabstractUnlearnable examples (UEs) seek to maximize testing error by making subtle modifications to training examples that are correctly labeled. Defenses against these poisoning attacks can be categorized based on whether specific interventions are adopted during training. The first approach is training-time defense, such as adversarial training, which can mitigate poisoning effects but is computationally intensive. The other approach is pre-training purification, e.g., image short squeezing, which consists of several simple compressions but often encounters challenges in dealing with various UEs. Our work provides a novel disentanglement mechanism to build an efficient pre-training purification method. Firstly, we uncover rate-constrained variational autoencoders (VAEs), demonstrating a clear tendency to suppress the perturbations in UEs. We subsequently conduct a theoretical analysis for this phenomenon. Building upon these insights, we introduce a disentangle variational autoencoder (D-VAE), capable of disentangling the perturbations with learnable class-wise embeddings. Based on this network, a two-stage purification approach is naturally developed. The first stage focuses on roughly eliminating perturbations, while the second stage produces refined, poison-free results, ensuring effectiveness and robustness across various scenarios. Extensive experiments demonstrate the remarkable performance of our method across CIFAR-10, CIFAR-100, and a 100-class ImageNet-subset. Code is available at https://github.com/yuyi-sd/D-VAE. Yi Yu 0011, Yufei Wang 0006, Song Xia, Wenhan Yang, Shijian Lu, Yap-Peng Tan, Alex Chichung Kot |
ICML | 7 |
| 2024 | HideMIA: Hidden Wavelet Mining for Privacy-Enhancing Medical Image Analysis
Xun Lin, Yi Yu 0011, Zitong Yu, Ruohan Meng, Jiale Zhou 0001, Ajian Liu 0001, Yizhong Liu, Shuai Wang 0049, Wenzhong Tang, Zhen Lei 0001, Alex Chichung Kot |
ACM Multimedia | 11 |
| 2024 | Evolving Storytelling: Benchmarks and Methods for New Character Customization with Diffusion ModelsabstractDiffusion-based models for story visualization have shown promise in generating content-coherent images for storytelling tasks. However, how to effectively integrate new characters into existing narratives while maintaining character consistency remains an open problem, particularly with limited data. Two major limitations hinder the progress: (1) the absence of a suitable benchmark due to potential character leakage and inconsistent text labeling, and (2) the challenge of distinguishing between new and old characters, leading to ambiguous results. To address these challenges, we introduce the NewEpisode benchmark, comprising refined datasets designed to evaluate generative models' adaptability in generating new stories with fresh characters using just a single example story. The refined dataset involves refined text prompts and eliminates character leakage. Additionally, to mitigate the character confusion of generated results, we propose EpicEvo, a method that customizes a diffusion-based visual story generation model with a single story featuring the new characters seamlessly integrating them into established character dynamics. EpicEvo introduces a novel adversarial character alignment module to align the generated images progressively in the diffusive process, with exemplar images of new characters, while applying knowledge distillation to prevent forgetting of characters and background details. Our evaluation quantitatively demonstrates that EpicEvo outperforms existing baselines on the NewEpisode benchmark, and qualitative studies confirm its superior customization of visual story generation in diffusion models. In summary, EpicEvo provides an effective way to incorporate new characters using only one example story, unlocking new possibilities for applications such as serialized cartoons. Yufei Wang 0006, Satoshi Tsutsui, Weisi Lin, Bihan Wen, Alex Chichung Kot |
ACM Multimedia | 6 |
| 2024 | From Chaos to Clarity: 3DGS in the DarkabstractNovel view synthesis from raw images provides superior high dynamic range (HDR) information compared to reconstructions from low dynamic range RGB images. However, the inherent noise in unprocessed raw images compromises the accuracy of 3D scene representation. Our study reveals that 3D Gaussian Splatting (3DGS) is particularly susceptible to this noise, leading to numerous elongated Gaussian shapes that overfit the noise, thereby significantly degrading reconstruction quality and reducing inference speed, especially in scenarios with limited views. To address these issues, we introduce a novel self-supervised learning framework designed to reconstruct HDR 3DGS from a limited number of noisy raw images. This framework enhances 3DGS by integrating a noise extractor and employing a noise-robust reconstruction loss that leverages a noise distribution prior. Experimental results show that our method outperforms LDR/HDR 3DGS and previous state-of-the-art (SOTA) self-supervised and supervised pre-trained models in both reconstruction quality and inference speed on the RawNeRF dataset across a broad range of training views. We will release the code upon paper acceptance. Yufei Wang 0006, Alex Chichung Kot, Bihan Wen |
NeurIPS | 3 |
| 2024 | ContextGS : Compact 3D Gaussian Splatting with Anchor Level Context ModelabstractRecently, 3D Gaussian Splatting (3DGS) has become a promising framework for novel view synthesis, offering fast rendering speeds and high fidelity. However, the large number of Gaussians and their associated attributes require effective compression techniques.
Existing methods primarily compress neural Gaussians individually and independently, i.e., coding all the neural Gaussians at the same time, with little design for their interactions and spatial dependence. Inspired by the effectiveness of the context model in image compression, we propose the first autoregressive model at the anchor level for 3DGS compression in this work. We divide anchors into different levels and the anchors that are not coded yet can be predicted based on the already coded ones in all the coarser levels, leading to more accurate modeling and higher coding efficiency. To further improve the efficiency of entropy coding, e.g., to code the coarsest level with no already coded anchors, we propose to introduce a low-dimensional quantized feature as the hyperprior for each anchor, which can be effectively compressed. Our work pioneers the context model in the anchor level for 3DGS representation, yielding an impressive size reduction of over 100 times compared to vanilla 3DGS and 15 times compared to the most recent state-of-the-art work Scaffold-GS, while achieving comparable or even higher rendering quality. Yufei Wang 0006, Lanqing Guo, Wenhan Yang, Alex Chichung Kot, Bihan Wen |
NeurIPS | 5 |
| 2024 | Beyond Learned Metadata-Based Raw Image Reconstruction
Yufei Wang 0006, Yi Yu 0011, Wenhan Yang, Lanqing Guo, Lap-Pui Chau, Alex Chichung Kot, Bihan Wen |
Int. J. Comput. Vis. | 6 |
| 2024 | Rethinking Vision Transformer and Masked Autoencoder in Multimodal Face Anti-SpoofingabstractAbstract Recently, vision transformer (ViT) based multimodal learning methods have been proposed to improve the robustness of face anti-spoofing (FAS) systems. However, there are still no works to explore the fundamental natures (e.g., modality-aware inputs, suitable multimodal pre-training, and efficient finetuning) in vanilla ViT for multimodal FAS. In this paper, we investigate three key factors (i.e., inputs, pre-training, and finetuning) in ViT for multimodal FAS with RGB, Infrared (IR), and Depth. First, in terms of the ViT inputs, we find that leveraging local feature descriptors (such as histograms of oriented gradients) benefits the ViT on IR modality but not RGB or Depth modalities. Second, in consideration of the task (FAS vs. generic object classification) and modality (multimodal vs. unimodal) gaps, ImageNet pre-trained models might be sub-optimal for the multimodal FAS task. Finally, in observation of the inefficiency on direct finetuning the whole or partial ViT, we design an adaptive multimodal adapter (AMA), which can efficiently aggregate local multimodal features while freezing majority of ViT parameters. To bridge these gaps, we propose the modality-asymmetric masked autoencoder (M $$^{2}$$ 2 A $$^{2}$$ 2 E) for multimodal FAS self-supervised pre-training without costly annotated labels. Compared with the previous modality-symmetric autoencoder, the proposed M $$^{2}$$ 2 A $$^{2}$$ 2 E is able to learn more intrinsic task-aware representation and compatible with modality-agnostic (e.g., unimodal, bimodal, and trimodal) downstream settings. Extensive experiments with both unimodal (RGB, Depth, IR) and multimodal (RGB+Depth, RGB+IR, Depth+IR, RGB+Depth+IR) settings conducted on multimodal FAS benchmarks demonstrate the superior performance of the proposed methods. One highlight is that the proposed method is robust under various missing-modality cases where previous multimodal FAS models suffer serious performance drops. We hope these findings and solutions can facilitate the future research for ViT-based multimodal FAS. Zitong Yu, Rizhao Cai, Yawen Cui, Xin Liu 0012, Yongjian Hu, Alex Chichung Kot |
Int. J. Comput. Vis. | 6 |
| 2024 | A unified deep semantic expansion framework for domain-generalized person re-identification
Eugene P. W. Ang, Alex Chichung Kot |
Neurocomputing | 3 |
| 2024 | GTADT: Gated tone-sensitive acne grading via augmented domain transfer
Min Tan 0005, Ruirui Wang, Ankur Purwar, Tao Jin 0004, Jun Yu 0002, Alex Chichung Kot |
Multim. Tools Appl. | 6 |
| 2024 | One-Shot Action Recognition via Multi-Scale Spatial-Temporal Skeleton MatchingabstractOne-shot skeleton action recognition, which aims to learn a skeleton action recognition model with a single training sample, has attracted increasing interest due to the challenge of collecting and annotating large-scale skeleton action data. However, most existing studies match skeleton sequences by comparing their feature vectors directly which neglects spatial structures and temporal orders of skeleton data. This paper presents a novel one-shot skeleton action recognition technique that handles skeleton action recognition via multi-scale spatial-temporal feature matching. We represent skeleton data at multiple spatial and temporal scales and achieve optimal feature matching from two perspectives. The first is multi-scale matching which captures the scale-wise semantic relevance of skeleton data at multiple spatial and temporal scales simultaneously. The second is cross-scale matching which handles different motion magnitudes and speeds by capturing sample-wise relevance across multiple scales. Extensive experiments over three large-scale datasets (NTU RGB+D, NTU RGB+D 120, and PKU-MMD) show that our method achieves superior one-shot skeleton action recognition, and outperforms SOTA consistently by large margins. Siyuan Yang 0001, Jun Liu 0036, Shijian Lu, Meng Hwa Er, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Self-Supervised 3D Action Representation Learning With Skeleton Cloud Colorizationabstract3D Skeleton-based human action recognition has attracted increasing attention in recent years. Most of the existing work focuses on supervised learning which requires a large number of labeled action sequences that are often expensive and time-consuming to annotate. In this paper, we address self-supervised 3D action representation learning for skeleton-based action recognition. We investigate self-supervised representation learning and design a novel skeleton cloud colorization technique that is capable of learning spatial and temporal skeleton representations from unlabeled skeleton sequence data. We represent a skeleton action sequence as a 3D skeleton cloud and colorize each point in the cloud according to its temporal and spatial orders in the original (unannotated) skeleton sequence. Leveraging the colorized skeleton point cloud, we design an auto-encoder framework that can learn spatial-temporal features from the artificial color labels of skeleton joints effectively. Specifically, we design a two-steam pretraining network that leverages fine-grained and coarse-grained colorization to learn multi-scale spatial-temporal features. In addition, we design a Masked Skeleton Cloud Repainting task that can pretrain the designed auto-encoder framework to learn informative representations. We evaluate our skeleton cloud colorization approach with linear classifiers trained under different configurations, including unsupervised, semi-supervised, fully-supervised, and transfer learning settings. Extensive experiments on NTU RGB+D, NTU RGB+D 120, PKU-MMD, NW-UCLA, and UWA3D datasets show that the proposed method outperforms existing unsupervised and semi-supervised 3D action recognition methods by large margins and achieves competitive performance in supervised 3D action recognition as well. Siyuan Yang 0001, Jun Liu 0036, Shijian Lu, Meng Hwa Er, Yongjian Hu, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Learning to Remove Rain in Video With Self-SupervisionabstractIn heavy rain video, rain streak and rain accumulation are the most common causes of degradation. They occlude background information and can significantly impair the visibility. Most existing methods rely heavily on the synthetic training data, and thus raise the domain gap problem that prevents the trained models from performing adequately in real testing cases. Unlike these methods, we introduce a self-learning method to remove both rain streaks and rain accumulation without using any ground-truth clean images in training our model, which consequently can alleviate the domain gap issue. The main idea is based on the assumptions that (1) adjacent clean frames can be aligned or warped from one frame to another frame, (2) rain streaks are distributed randomly in the temporal domain, (3) the rain streak/accumulation related variables/priors can be inferred reliably from the information within the images/sequences. Based on these assumptions, we construct an augmented Self-Learned Deraining Network (SLDNet+) to remove both rain streaks and rain accumulation by utilizing temporal correlation, consistency, and rain-related priors. For the temporal correlation, our SLDNet+ takes rain degraded adjacent frames as its input, aligns them, and learns to predict the clean version of the current frame. For the temporal consistency, a new loss is designed to build a robust mapping between the predicted clean frame and non-rain regions from the adjacent rain frames. For the rain-streak-related prior, the rain streak removal network is optimized jointly with motion estimation and rain region detection; while for the rain-accumulation-related prior, a novel non-local video rain accumulation removal method is developed to estimate the accumulation-lines from the whole input video and to offer better color constancy and temporal smoothness. Extensive experiments show the effectiveness of our approach, which provides superior results compared with the existing state of the art methods both quantitatively and qualitatively. The source code will be made publicly available at: https://github.com/flyywh/CVPR-2020-Self-Rain-Removal-Journal. Wenhan Yang, Robby T. Tan, Shiqi Wang 0001, Alex Chichung Kot, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Comm-Transformer: A Robust Deep Learning-Based Receiver for OFDM System Under TDL ChannelabstractIn this paper, we propose a deep learning (DL) based receiver named comm-transformer network (Comm-Trans Net), which is robust for different sub-types of tapped delay line (TDL) channels. The novel Comm-Trans Net considers the attention mechanism to compensate for the multi-path fading effect for different orthogonal frequency-division multiplexing (OFDM) subcarriers. We propose a novel positional encoding method for each OFDM subcarrier, which uses the attention mechanism to offset the deep fading effect. In particular, our proposed Comm-Trans Net serves as an integrated DL-based receiver for unknown channel conditions. Our results show that the proposed Comm-Trans Net can outperform the bit-error rate (BER) performance compared to minimum mean-square error (MMSE) channel estimation with generalized approximate message passing (GAMP) receiver, and also outperforms the state-of-the-art DL-based receivers. Moreover, our proposed Comm-Trans Net is robust for multiple communication scenarios ranging from the TDL-A channel to the TDL-E channel with different levels of non-line-of-sight (NLoS) and line-of-sight (LoS) components. Also, the computational complexity of the proposed Comm-Trans Net is studied, which shows that the attention mechanism for channel positional encoding is a cost-effective solution to improve the quality of service. Yihang Xie, Kah Chan Teh, Alex Chichung Kot |
IEEE Trans. Commun. | 3 |
| 2024 | Cross-Image Disentanglement for Low-Light Enhancement in Real WorldabstractImages captured in the low-light condition suffer from low visibility and various imaging artifacts, e.g., real noise. Existing supervised algorithms for low-light image enhancement require a large set of pixel-aligned training image pairs, which are hard to prepare in practice. Though some recent unsupervised methods can alleviate such data challenges, many real world artifacts inevitably get falsely amplified in the enhanced results due to the lack of corresponding supervision. In this paper, instead of using perfectly aligned images for training, we creatively employ the misaligned real world images as the guidance, which are considerably easier to collect. Specifically, we propose a Cross-Image Disentanglement Network (CIDN) with weakly supervised learning, to separately extract cross-image brightness and image-specific content features from low/normal-light images. Based on that, CIDN can simultaneously correct the brightness and suppress image artifacts in the feature domain, which largely increases the robustness of the pixel shifts between training pairs. By considering real world corruptions, we propose a new training dataset with misaligned and noisy image pairs and its corresponding evaluation dataset. Experimental results show that our model achieves state-of-the-art performances on both the newly proposed dataset and other popular low-light datasets. The code implementation is publicly available at:https://github.com/GuoLanqing/CIDN. Lanqing Guo, Renjie Wan, Wenhan Yang, Alex Chichung Kot, Bihan Wen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | PVASS-MDD: Predictive Visual-Audio Alignment Self-Supervision for Multimodal Deepfake DetectionabstractDeepfake techniques can forge the visual or audio signals in the video, which leads to inconsistencies between visual and audio (VA) signals. Therefore, multimodal detection methods expose deepfake videos by extracting VA inconsistencies. Recently, deepfake technology has started VA collaborative forgery to obtain more realistic deepfake videos, which poses new challenges for extracting VA inconsistencies. Recent multimodal detection methods propose to first extract natural VA correspondences in real videos in a self-supervised manner, and then use the learned real correspondences as targets to guide the extraction of VA inconsistencies in the subsequent deepfake detection stage. However, the inherent VA relations are difficult to extract due to the modality gap, which leads to the limited auxiliary performance of the aforementioned self-supervised methods. In this paper, we propose Predictive Visual-audio Alignment Self-supervision for Multimodal Deepfake Detection (PVASS-MDD), which consists of PVASS auxiliary and MDD stages. In the PVASS auxiliary stage in real videos, we first devise a three-stream network to associate two augmented visual views with corresponding audio clues, leading to explore common VA correspondences based on cross-view learning. Secondly, we introduce a novel cross-modal predictive align module for eliminating VA gaps to provide inherent VA correspondences. In the MDD stage, we propose to the auxiliary loss to utilize the frozen PVASS network to align VA features of real videos, to better assist multimodal deepfake detector for capturing subtle VA inconsistencies. We conduct extensive experiments on existing widely used and latest multimodal deepfake datasets. Our method obtains a significant performance improvement compared to state-of-the-art methods. Yang Yu 0039, Siyuan Yang 0001, Yao Zhao 0001, Alex Chichung Kot |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Benchmarking Joint Face Spoofing and Forgery Detection With Visual and Physiological CuesabstractFace anti-spoofing (FAS) and face forgery detection play vital roles in securing face biometric systems from presentation attacks (PAs) and vicious digital manipulation (e.g., deepfakes). Despite satisfactory performance upon large-scale data and powerful deep models, recent advances in face spoofing and forgery detection approaches usually focus on 1) unimodal visual appearance or physiological (i.e., remote photoplethysmography (rPPG)) cues; and 2) separated feature representation for FAS or face forgery detection. On one side, unimodal appearance and rPPG features are respectively vulnerable to high-fidelity face 3D mask and video replay attacks, inspiring us to design reliable multi-modal fusion mechanisms for generalized FAS. On the other side, there are rich common features across FAS and face forgery detection tasks (e.g., periodic rPPG rhythms and vanilla appearance for bonafides), providing solid evidence to design a joint FAS and face forgery detection system in a multi-task learning fashion. In this paper, we establish the first joint face spoofing and forgery detection benchmark using both visual appearance and physiological rPPG cues. To enhance the rPPG periodicity discrimination, we design a two-branch physiological network using both facial spatio-temporal rPPG signal map and its continuous wavelet transformed counterpart as inputs. To mitigate the modality bias and improve the fusion efficacy, we conduct a weighted batch and layer normalization for both appearance and rPPG features before multi-modal fusion. We also investigate prevalent deep models, feature fusion strategies and multi-task learning configurations for joint face spoofing and forgery detection. We find that the generalization capacities of both unimodal (appearance or rPPG) and multi-modal (appearance+rPPG) models can be obviously improved via joint training on these two tasks. We hope this new benchmark will facilitate the future research of both FAS and deepfake detection communities. The codes will be released athttps://github.com/ZitongYu/Benchmarking. Zitong Yu, Rizhao Cai, Zhi Li 0054, Wenhan Yang, Jingang Shi, Alex Chichung Kot |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2024 | S-Adapter: Generalizing Vision Transformer for Face Anti-Spoofing With Statistical TokensabstractFace Anti-Spoofing (FAS) aims to detect malicious attempts to invade a face recognition system by presenting spoofed faces. State-of-the-art FAS techniques predominantly rely on deep learning models but their cross-domain generalization capabilities are often hindered by the domain shift problem, which arises due to different distributions between training and testing data. In this study, we develop a generalized FAS method under the Efficient Parameter Transfer Learning (EPTL) paradigm, where we adapt the pre-trained Vision Transformer models for the FAS task. During training, the adapter modules are inserted into the pre-trained ViT model, and the adapters are updated while other pre-trained parameters remain fixed. We find the limitations of previous vanilla adapters in that they are based on linear layers, which lack a spoofing-aware inductive bias and thus restrict the cross-domain generalization. To address this limitation and achieve cross-domain generalized FAS, we propose a novel Statistical Adapter (S-Adapter) that gathers local discriminative and statistical information from localized token histograms. To further improve the generalization of the statistical tokens, we propose a novel Token Style Regularization (TSR), which aims to reduce domain style variance by regularizing Gram matrices extracted from tokens across different domains. Our experimental results demonstrate that our proposed S-Adapter and TSR provide significant benefits in both zero-shot and few-shot cross-domain testing, outperforming state-of-the-art methods on several benchmark tests. We will release the source code upon acceptance. Rizhao Cai, Zitong Yu, Chenqi Kong, Haoliang Li, Changsheng Chen 0001, Yongjian Hu, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2024 | Beyond the Prior Forgery Knowledge: Mining Critical Clues for General Face Forgery DetectionabstractFace forgery detection is essential in combating malicious digital face attacks. Previous methods mainly rely on prior expert knowledge to capture specific forgery clues, such as noise patterns, blending boundaries, and frequency artifacts. However, these methods tend to get trapped in local optima, resulting in limited robustness and generalization capability. To address these issues, we propose a novel Critical Forgery Mining (CFM) framework, which can be flexibly assembled with various backbones to boost their generalization and robustness performance. Specifically, we first build a fine-grained triplet and suppress specific forgery traces through prior knowledge-agnostic data augmentation. Subsequently, we propose a fine-grained relation learning prototype to mine critical information in forgeries through instance and local similarity-aware losses. Moreover, we design a novel progressive learning controller to guide the model to focus on principal feature components, enabling it to learn critical forgery features in a coarse-to-fine manner. The proposed method achieves state-of-the-art forgery detection performance under various challenging evaluation settings. The source code is available at:https://github.com/LoveSiameseCat/CFM. Anwei Luo, Chenqi Kong, Jiwu Huang, Yongjian Hu, Xiangui Kang, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2024 | Semantic Deep Hiding for Robust Unlearnable ExamplesabstractEnsuring data privacy and protection has become paramount in the era of deep learning. Unlearnable examples are proposed to mislead the deep learning models and prevent data from unauthorized exploration by adding small perturbations to data. However, such perturbations (e.g., noise, texture, color change) predominantly impact low-level features, making them vulnerable to common countermeasures. In contrast, semantic images with intricate shapes have a wealth of high-level features, making them more resilient to countermeasures and potential for producing robust unlearnable examples. In this paper, we propose a Deep Hiding (DH) scheme that adaptively hides semantic images enriched with high-level features. We employ an Invertible Neural Network (INN) to invisibly integrate predefined images, inherently hiding them with deceptive perturbations. To enhance data unlearnability, we introduce a Latent Feature Concentration module, designed to work with the INN, regularizing the intra-class variance of these perturbations. To further boost the robustness of unlearnable examples, we design a Semantic Images Generation module that produces hidden semantic images. By utilizing similar semantic information, this module generates similar semantic images for samples within the same classes, thereby enlarging the inter-class distance and narrowing the intra-class distance. Extensive experiments on CIFAR-10, CIFAR-100, and an ImageNet subset, against 18 countermeasures, reveal that our proposed method exhibits outstanding robustness for unlearnable examples, demonstrating its efficacy in preventing unauthorized data exploitation. Ruohan Meng, Chenyu Yi, Yi Yu 0011, Siyuan Yang 0001, Bingquan Shen, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2024 | Narrowing Domain Gaps With Bridging Samples for Generalized Face Forgery DetectionabstractFace forgery technology has developed rapidly, causing severe security issues in society. Recently, with the continuous emergence of forgery techniques and types, most forensics methods suffer from the generalization problem. In particular, it is difficult for existing generalized methods to detect fake faces with unseen fake types. The reason is that the distribution gaps among cross-forgery types are too large. In this article, we propose a novel generalized framework to narrow large gaps based on bridging cross-domain alignment to solve this problem. Specifically, our framework consists of three key steps: preventing, bridging and aligning distribution gaps. Firstly, in the feature mining stage, taking advantage of the ability of Instance Normalization (IN) to better tolerate domain gaps, we design Adaptive Batch and Instance Normalization (ABIN) to replace the commonly used BN to adaptively extract features to preliminarily prevent domain gaps. Secondly, we propose to generate bridging samples distributed among the inter-domains to fill large gaps based on progressive linear interpolation operation. Finally, with the help of bridging samples, the cross-domain alignment is performed to better narrow distribution gaps to refine data distribution, which helps to learn a more generalized framework. Extensive experiments show that our proposed framework achieves the state-of-the-art generalized performance. Yang Yu 0039, Siyuan Yang 0001, Yao Zhao 0001, Alex Chichung Kot |
IEEE Trans. Multim. | 5 |
| 2024 | Disentangled Feature Representation for Few-Shot Image ClassificationabstractLearning the generalizable feature representation is critical to few-shot image classification. While recent works exploited task-specific feature embedding using meta-tasks for few-shot learning, they are limited in many challenging tasks as being distracted by the excursive features such as the background, domain, and style of the image samples. In this work, we propose a novel disentangled feature representation (DFR) framework, dubbed DFR, for few-shot learning applications. DFR can adaptively decouple the discriminative features that are modeled by the classification branch, from the class-irrelevant component of the variation branch. In general, most of the popular deep few-shot learning methods can be plugged in as the classification branch, thus DFR can boost their performance on various few-shot tasks. Furthermore, we propose a novel FS-DomainNet dataset based on DomainNet, for benchmarking the few-shot domain generalization (DG) tasks. We conducted extensive experiments to evaluate the proposed DFR on general, fine-grained, and cross-domain few-shot classification, as well as few-shot DG, using the corresponding four benchmarks, i.e., mini-ImageNet, tiered-ImageNet, Caltech-UCSD Birds 200-2011 (CUB), and the proposed FS-DomainNet. Thanks to the effective feature disentangling, the DFR-based few-shot classifiers achieved state-of-the-art results on all datasets. Hao Cheng 0016, Yufei Wang 0006, Haoliang Li, Alex Chichung Kot, Bihan Wen |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Cross-Modal Label Contrastive Learning for Unsupervised Audio-Visual Event LocalizationabstractThis paper for the first time explores audio-visual event localization in an unsupervised manner. Previous methods tackle this problem in a supervised setting and require segment-level or video-level event category ground-truth to train the model. However, building large-scale multi-modality datasets with category annotations is human-intensive and thus not scalable to real-world applications. To this end, we propose cross-modal label contrastive learning to exploit multi-modal information among unlabeled audio and visual streams as self-supervision signals. At the feature representation level, multi-modal representations are collaboratively learned from audio and visual components by using self-supervised representation learning. At the label level, we propose a novel self-supervised pretext task i.e. label contrasting to self-annotate videos with pseudo-labels for localization model training. Note that irrelevant background would hinder the acquisition of high-quality pseudo-labels and thus lead to an inferior localization model. To address this issue, we then propose an expectation-maximization algorithm that optimizes the pseudo-label acquisition and localization model in a coarse-to-fine manner. Extensive experiments demonstrate that our unsupervised approach performs reasonably well compared to the state-of-the-art supervised methods. Peijun Bao, Wenhan Yang, Boon Poh Ng, Meng Hwa Er, Alex Chichung Kot |
AAAI | 5 |
| 2023 | Raw Image Reconstruction with Learned Compact MetadataabstractWhile raw images exhibit advantages over sRGB images (e.g., linearity and fine-grained quantization level), they are not widely used by common users due to the large storage requirements. Very recent works propose to compress raw images by designing the sampling masks in the raw image pixel space, leading to suboptimal image representations and redundant metadata. In this paper, we propose a novel framework to learn a compact representation in the latent space serving as the metadata in an end-to-end manner. Furthermore, we propose a novel sRGB-guided context model with the improved entropy estimation strategies, which leads to better reconstruction quality, smaller size of metadata, and faster speed. We illustrate how the proposed raw image compression scheme can adaptively allocate more bits to image regions that are important from a global perspective. The experimental results show that the proposed method can achieve superior raw image reconstruction results using a smaller size of the metadata on both uncompressed sRGB images and JPEG images. The code will be released at https://github.com/wyf0912/R2LCM Yufei Wang 0006, Yi Yu 0011, Wenhan Yang, Lanqing Guo, Lap-Pui Chau, Alex Chichung Kot, Bihan Wen |
CVPR | 6 |
| 2023 | Backdoor Attacks Against Deep Image Compression via Adaptive Frequency TriggerabstractRecent deep-learning-based compression methods have achieved superior performance compared with traditional approaches. However, deep learning models have proven to be vulnerable to backdoor attacks, where some specific trigger patterns added to the input can lead to malicious behavior of the models. In this paper, we present a novel backdoor attack with multiple triggers against learned image compression models. Motivated by the widely used discrete cosine transform (DCT) in existing compression systems and standards, we propose a frequency-based trigger injection model that adds triggers in the DCT domain. In particular, we design several attack objectives for various attacking scenarios, including: 1) attacking compression quality in terms of bit-rate and reconstruction quality; 2) attacking task-driven measures, such as downstream face recognition and semantic segmentation. Moreover, a novel simple dynamic loss is designed to balance the influence of different loss terms adaptively, which helps achieve more efficient training. Extensive experiments show that with our trained trigger injection models and simple modification of encoder parameters (of the compression model), the proposed attack can successfully inject several backdoors with corresponding triggers in a single image compression model. Yi Yu 0011, Yufei Wang 0006, Wenhan Yang, Shijian Lu, Yap-Peng Tan, Alex Chichung Kot |
CVPR | 6 |
| 2023 | Unsupervised Deep Digital Staining for Microscopic Cell Images via Knowledge DistillationabstractStaining is critical to cell imaging and medical diagnosis, which is expensive, time-consuming, labor-intensive, and causes irreversible changes to cell tissues. Recent advances in deep learning enabled digital staining via supervised model training. However, it is difficult to obtain large-scale stained/unstained cell image pairs in practice, which need to be perfectly aligned with the supervision. In this work, we propose a novel unsupervised deep learning framework for the digital staining of cell images using knowledge distillation and generative adversarial networks (GANs). A teacher model is first trained mainly for the colorization of bright-field images. After that, a student GAN for staining is obtained by knowledge distillation with hybrid non-reference losses. We show that the proposed unsupervised deep staining method can generate stained images with more accurate positions and shapes of the cell targets. Compared with other unsupervised deep generative models for staining, our method achieves much more promising results both qualitatively and quantitatively. Ziwang Xu, Lanqing Guo, Alex Chichung Kot, Bihan Wen |
ICASSP | 4 |
| 2023 | Towards Explainable Recommendation Via Bert-Guided Explanation GeneratorabstractExplainable recommender system has recently drawn increasing attention due to its capability of providing justification to recommendation. Rather than focusing on certain topics or specific item features, the explanation generated by existing works are too general without the guidance of aspects. However, such information is not given in the practical scenario. To address this issue, we propose a novel Explainable recommender system with BERT-guided explanation generator, named ExBERT to generate reliable explanation with finer granularity. More specifically, a multi-head self-attention based encoder is employed to incorporate pseudo user and item profiles into semantic representation. Moreover, we propose a novel matched explanation prediction task with discriminative ability to enable personalization of the generated sentence. Extensive experiments conducted on two real-world explainable recommendation datasets significantly outperform the state-of-the-art in generation. Huijing Zhan, Ling Li 0012, Weide Liu, Manas Gupta, Alex Chichung Kot |
ICASSP | 6 |
| 2023 | Rehearsal-Free Domain Continual Face Anti-Spoofing: Generalize More and Forget LessabstractFace Anti-Spoofing (FAS) is recently studied under the continual learning setting, where the FAS models are expected to evolve after encountering data from new domains. However, existing methods need extra replay buffers to store previous data for rehearsal, which becomes infeasible when previous data is unavailable because of privacy issues. In this paper, we propose the first rehearsal-free method for Domain Continual Learning (DCL) of FAS, which deals with catastrophic forgetting and unseen domain generalization problems simultaneously. For better generalization to unseen domains, we design the Dynamic Central Difference Convolutional Adapter (DCDCA) to adapt Vision Transformer (ViT) models during the continual learning sessions. To alleviate the forgetting of previous domains without using previous data, we propose the Proxy Prototype Contrastive Regularization (PPCR) to constrain the continual learning with previous domain knowledge from the proxy prototypes. Simulating practical DCL scenarios, we devise two new protocols which evaluate both generalization and anti-forgetting performance. Extensive experimental results show that our proposed method can improve the generalization performance in unseen domains and alleviate the catastrophic forgetting of previous knowledge. The code and protocol files are released on https://github.com/RizhaoCai/DCL-FAS-ICCV2023. Rizhao Cai, Yawen Cui, Zitong Yu, Haoliang Li, Yongjian Hu, Alex Chichung Kot |
ICCV | 7 |
| 2023 | Audio-Visual Deception Detection: DOLOS Dataset and Parameter-Efficient Crossmodal LearningabstractDeception detection in conversations is a challenging yet important task, having pivotal applications in many fields such as credibility assessment in business, multimedia anti-frauds, and custom security. Despite this, deception detection research is hindered by the lack of high-quality deception datasets, as well as the difficulties of learning multimodal features effectively. To address this issue, we introduce DOLOS1, the largest gameshow deception detection dataset with rich deceptive conversations. DOLOS includes 1, 675 video clips featuring 213 subjects, and it has been labeled with audio-visual feature annotations. We provide train-test, duration, and gender protocols to investigate the impact of different factors. We benchmark our dataset on previously proposed deception detection approaches. To further improve the performance by fine-tuning fewer parameters, we propose Parameter-Efficient Crossmodal Learning (PECL), where a Uniform Temporal Adapter (UT-Adapter) explores temporal attention in transformer-based architectures, and a crossmodal fusion module, Plug-in Audio-Visual Fusion (PAVF), combines crossmodal information from audio-visual features. Based on the rich fine-grained audio-visual annotations on DOLOS, we also exploit multi-task learning to enhance performance by concurrently predicting deception and audiovisual features. Experimental results demonstrate the desired quality of the DOLOS dataset and the effectiveness of the PECL. The DOLOS dataset and the source codes are available at here. Xiaobao Guo, Nithish Muthuchamy Selvaraj, Zitong Yu, Adams Wai-Kin Kong, Bingquan Shen, Alex Chichung Kot |
ICCV | 6 |
| 2023 | Virtual Try-On with Pose-Garment Keypoints Guided InpaintingabstractVirtual try-on is an important technology supporting on-line apparel shopping, which provides consumers with a virtual experience to fit garments without physically wearing them. Recently, the image-based virtual try-on has received growing research attention. However, the synthetic results of existing virtual try-on methods usually present distortions in garment shape and lose pattern details. In this paper, we propose a pose-garment keypoints guided inpainting method for the image-based virtual try-on task, which produces high-fidelity try-on images and well preserves the shapes and patterns of the garments. In our method, human pose and garment keypoints are extracted from source images and constructed as graphs to predict the garment keypoints at the target pose. After which, the predicted key-points are used as guide information to predict the target segmentation map and warp the garment image. The try-on image is finally generated with a semantic-conditioned inpainting scheme using the segmentation map and recomposed person image as conditions. To verify the effectiveness of our proposed method, we conduct extensive experiments on the VITON-HD dataset under both paired and unpaired experimental settings. The qualitative and quantitative results show that our method significantly outperforms prior methods at different image resolutions. The codes repository link is https://github.com/lizhi-ntu/KGI. Pengfei Wei 0001, Xiang Yin 0006, Zejun Ma 0001, Alex Chichung Kot |
ICCV | 5 |
| 2023 | ExposureDiffusion: Learning to Expose for Low-light Image EnhancementabstractPrevious raw image-based low-light image enhancement methods predominantly relied on feed-forward neural networks to learn deterministic mappings from low-light to normally-exposed images. However, they failed to capture critical distribution information, leading to visually undesirable results. This work addresses the issue by seamlessly integrating a diffusion model with a physics-based exposure model. Different from a vanilla diffusion model that has to perform Gaussian denoising, with the injected physics-based exposure model, our restoration process can directly start from a noisy image instead of pure noise. As such, our method obtains significantly improved performance and reduced inference time compared with vanilla diffusion models. To make full use of the advantages of different intermediate steps, we further propose an adaptive residual layer that effectively screens out the side-effect in the iterative refinement when the intermediate results have been already well-exposed. The proposed framework can work with both real-paired datasets, SOTA noise models, and different backbone networks. We evaluate the proposed method on various public benchmarks, achieving promising results with consistent improvements using different exposure models and backbones. Besides, the proposed method achieves better generalization capacity for unseen amplifying ratios and better performance than a larger feedforward neural model when few parameters are adopted. The code is released at https://github.com/wyf0912/ExposureDiffusion. Yufei Wang 0006, Yi Yu 0011, Wenhan Yang, Lanqing Guo, Lap-Pui Chau, Alex Chichung Kot, Bihan Wen |
ICCV | 6 |
| 2023 | Enhancing Low-Light Images Using Infrared Encoded ImagesabstractLow-light image enhancement task is essential yet challenging as it is ill-posed intrinsically. Previous arts mainly focus on the low-light images captured in the visible spectrum using pixel-wise loss, which limits the capacity of recovering the brightness, contrast, and texture details due to the small number of income photons. In this work, we propose a novel approach to increase the visibility of images captured under low-light environments by removing the in-camera infrared (IR) cut-off filter, which allows for the capture of more photons and results in improved signal-to-noise ratio due to the inclusion of information from the IR spectrum. To verify the proposed strategy, we collect a paired dataset of low-light images captured without the IR cut-off filter, with corresponding long-exposure reference images with an external filter. The experimental results on the proposed dataset demonstrate the effectiveness of the proposed method, showing better performance quantitatively and qualitatively. The dataset and code are publicly available at https://wyf0912.github.io/ELIEI/ Shulin Tian, Yufei Wang 0006, Renjie Wan, Wenhan Yang, Alex Chichung Kot, Bihan Wen |
ICIP | 5 |
| 2023 | Temporal Coherent Test Time Optimization for Robust Video Classification
Chenyu Yi, Siyuan Yang 0001, Yufei Wang 0006, Haoliang Li, Yap-Peng Tan, Alex Chichung Kot |
ICLR | 6 |
| 2023 | Adapter Incremental Continual Learning of Efficient Audio Spectrogram Transformers
Nithish Muthuchamy Selvaraj, Xiaobao Guo, Adams Wai-Kin Kong, Bingquan Shen, Alex Chichung Kot |
INTERSPEECH | 5 |
| 2023 | Removing Image Artifacts From Scratched Lens ProtectorsabstractA protector is placed in front of the camera lens for mobile devices to avoid damage, while the protector itself can be easily scratched accidentally, especially for plastic ones. The artifacts appear in a wide variety of patterns, making it difficult to see through them clearly. Removing image artifacts from the scratched lens protector is inherently challenging due to the occasional flare artifacts and the co-occurring interference within mixed artifacts. Though different methods have been proposed for some specific distortions, they seldom consider such inherent challenges. In our work, we consider the inherent challenges in a unified framework with two cooperative modules, which facilitate the performance boost of each other. We also collect a new dataset from the real world to facilitate training and evaluation purposes. The experimental results demonstrate that our method outperforms the baselines qualitatively and quantitatively. The code and datasets will be released at https://github.com/wyf0912/flare-removal Yufei Wang 0006, Renjie Wan, Wenhan Yang, Bihan Wen, Lap-Pui Chau, Alex Chichung Kot |
ISCAS | 6 |
| 2023 | PAR$^{2}$2Net: End-to-End Panoramic Image Reflection RemovalabstractIn this article, we investigate the problem of panoramic image reflection removal to relieve the content ambiguity between the reflection layer and the transmission scene. Although a partial view of the reflection scene is attainable in the panoramic image and provides additional information for reflection removal, it is not trivial to directly apply this for getting rid of undesired reflections due to its misalignment with the reflection-contaminated image. We propose an end-to-end framework to tackle this problem. By resolving misalignment issues with adaptive modules, the high-fidelity recovery of reflection layer and transmission scenes is accomplished. We further propose a new data generation approach that considers the physics-based formation model of mixture images and the in-camera dynamic range clipping to diminish the domain gap between synthetic and real data. Experimental results demonstrate the effectiveness of the proposed method and its applicability for mobile devices and industrial applications. Yuchen Hong, Lingran Zhao, Xudong Jiang 0001, Alex Chichung Kot, Boxin Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Benchmarking Single-Image Reflection Removal AlgorithmsabstractReflection removal has been discussed for more than decades. This paper aims to provide the analysis for different reflection properties and factors that influence image formation, an up-to-date taxonomy for existing methods, a benchmark dataset, and the unified benchmarking evaluations for state-of-the-art (especially learning-based) methods. Specifically, this paper presents a SIngle-image Reflection Removal Plus dataset “SIR$^{2+}$” with the new consideration for in-the-wild scenarios and glass with diverse color and unplanar shapes. We further perform quantitative and visual quality comparisons for state-of-the-art single-image reflection removal algorithms. Open problems for improving reflection removal algorithms are discussed at the end. Our dataset and follow-up update can be found athttps://reflectionremoval.github.io/sir2data/. Renjie Wan, Boxin Shi, Haoliang Li, Yuchen Hong, Ling-Yu Duan, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Deep Multimodal Sequence Fusion by Regularized Expressive Representation DistillationabstractMultimodal sequence learning aims to utilize information from different modalities to enhance overall performance. Mainstream works often follow an intermediate-fusion pipeline, which explores both modality-specific and modality-supplementary information for fusion. However, the unaligned and heterogeneously distributed multimodal sequences pose significant challenges to the fusion task: 1) to extract both effective unimodal and crossmodal representations and 2) to overcome the overfitting issue in joint multimodal sequence optimization. In this work, we propose regularized expressive representation distillation (RERD) that aims to seek effective multimodal representations and to enhance the generalization of fusion. First, to improve unimodal representation learning, unimodal representations are assigned to multi-head distillation encoders, where the unimodal representations are iteratively updated through distillation attention layers. Second, to alleviate the overfitting issue in joint crossmodal optimization, a multimodal sinkhorn distance regularizer is proposed to reinforce the expressive representation extraction and to reduce the modality gap before fusion adaptively. These representations produce a comprehensive view of the multimodal sequences, which are utilized for downstream fusion tasks. Experimental results on several popular benchmarks demonstrate that the proposed method achieves state-of-the-art performance, compared with widely used baselines for deep multimodal sequence fusion, as shown inhttps://github.com/Redaimao/RERD. Xiaobao Guo, Adams Wai-Kin Kong, Alex Chichung Kot |
IEEE Trans. Multim. | 3 |
| 2023 | Pace-Adaptive and Noise-Resistant Contrastive Learning for Multimodal Feature FusionabstractMultimodal feature fusion aims to draw complementary information from different modalities to achieve better performance. Contrastive learning is effective at discriminating coexisting semantic features (positive) from irrelative ones (negative) in multimodal signals. However, positive and negative pairs learn at separate rates, which undermines the overall performance of multimodal contrastive learning (MCL). Moreover, the learned representation model is not robust, as MCL utilizes supervision signals from potentially noisy modalities. To address these issues, a novel multimodal contrastive learning objective, Pace-adaptive and Noise-resistant Noise-Contrastive Estimation (PN-NCE), is proposed for multimodal fusion by directly using unimodal features. PN-NCE encourages the positive and negative pairs reaching to their optimal similarity scores adaptively and shows less susceptibility to noisy inputs during training. A theoretical analysis is performed on its robustness. Maximizing modality invariance information in the fused representation is expected to benefit the overall performance and therefore, an estimator that measures the difference between the fused representation and its unimodal representations is integrated into MCL to obtain a more modality-invariant fusion output. The proposed method is model-agnostic and can be adapted to various multimodal tasks. It also bears less performance degradation when reducing the number of training samples at the linear probing stage. With different networks and modality inputs from three multimodal datasets, experimental results show that PN-NCE achieves consistent enhancements compared with previous state-of-the-art approaches. Xiaobao Guo, Alex Chichung Kot, Adams Wai-Kin Kong |
IEEE Trans. Multim. | 2 |
| 2023 | Asymmetric Modality Translation for Face Presentation Attack DetectionabstractFace presentation attack detection (PAD) is an essentialmeasure to protect face recognition systems from being spoofed by malicious users and has attracted great attention from both academia and industry. Although most of the existing methods can achieve desired performance to some extent, the generalization issue of face presentation attack detection under cross-domain settings (e.g., the setting of unseen attacks and varying illumination) remains to be solved. In this paper, we propose a novel framework based on asymmetric modality translation for face presentation attack detection in bi-modality scenarios. Under the framework, we establish connections between two modality images of genuine faces. Specifically, a novel modality fusion scheme is presented that the image of one modality is translated to the other one through an asymmetric modality translator, then fused with its corresponding paired image. The fusion result is fed as the input to a discriminator for inference. The training of the translator is supervised by an asymmetric modality translation loss. Besides, an illumination normalization module based on Pattern of Local Gravitational Force (PLGF) representation is used to reduce the impact of illumination variation. We conduct extensive experiments on three public datasets, which validate that our method is effective in detecting various types of attacks and achieves state-of-the-art performance under different evaluation protocols. Zhi Li 0054, Haoliang Li, Yongjian Hu, Kwok-Yan Lam, Alex Chichung Kot |
IEEE Trans. Multim. | 6 |
| 2023 | Purifying Low-Light Images via Near-Infrared Enlightened ImageabstractCameras usually produce low-quality images under low-light conditions. Though many methods have been proposed to enhance the visibility of low-light images, they are mainly designed for illumination correction and less capable of sup-pressing the artifacts. In this paper, we propose to enhance the visibility and suppress artifacts by purifying low-light images under the guidance of the NIR enlightened image captured by using the near-infrared light as compensation. Specifically, we introduce a disentanglement framework to disentangle the structure and color components from the NIR enlightened and RGB images, respectively. Correspondingly, we introduce a new dataset with the RGB and NIR enlightened images for training and evaluation purposes. The experimental results show that our proposed method achieves promising results. Renjie Wan, Boxin Shi, Wenhan Yang, Bihan Wen, Ling-Yu Duan, Alex Chichung Kot |
IEEE Trans. Multim. | 6 |
| 2023 | Augmented Multi-Scale Spatiotemporal Inconsistency Magnifier for Generalized DeepFake DetectionabstractRecently, realistic DeepFake videos have raised severe security concerns in society. Existing video-based detection methods observe local spatial regions with the coarse temporal view, thus it is difficult to obtain subtle spatiotemporal information, resulting in limited generalization ability. In this paper, we propose a novel Augmented Multi-scale Spatiotemporal Inconsistency Magnifier (AMSIM) with a Global Inconsistency View (GIV) and a more meticulous Multi-timescale Local Inconsistency View (MLIV), focusing on mining comprehensive and more subtle spatiotemporal cues. Firstly, the GIV that includs the global spatial and long-term temporal views is established to ensure comprehensive spatiotemporal clues are captured. Then, the MLIV with the critical local spatial and multi-timescale local temporal views is designed for magnifying the indetectable spatiotemporal abnormality. Subsequently, GIV is utilized to guide MLIV to dynamically find local spatiotemporal anomalies that are highly relevant to the overall video. Finally, to further obtain a generalized framework, the adversarial data augmentation is specially designed to expand source domains and simulate unseen forgery domains. Extensive experiments on six large-scale datasets show that our AMSIM outperforms state-of-the-art detection methods and remains effective when applied to unseen forgery techniques and datasets. Yang Yu 0039, Siyuan Yang 0001, Yao Zhao 0001, Alex Chichung Kot |
IEEE Trans. Multim. | 6 |
| 2023 | Low-Rankness Guided Group Sparse Representation for Image RestorationabstractAs a spotlighted nonlocal image representation model, group sparse representation (GSR) has demonstrated a great potential in diverse image restoration tasks. Most of the existing GSR-based image restoration approaches exploit the nonlocal self-similarity (NSS) prior by clustering similar patches into groups and imposing sparsity to each group coefficient, which can effectively preserve image texture information. However, these methods have imposed only plain sparsity over each individual patch of the group, while neglecting other beneficial image properties, e.g., low-rankness (LR), leads to degraded image restoration results. In this article, we propose a novel low-rankness guided group sparse representation (LGSR) model for highly effective image restoration applications. The proposed LGSR jointly utilizes the sparsity and LR priors of each group of similar patches under a unified framework. The two priors serve as the complementary priors in LGSR for effectively preserving the texture and structure information of natural images. Moreover, we apply an alternating minimization algorithm with an adaptively adjusted parameter scheme to solve the proposed LGSR-based image restoration problem. Extensive experiments are conducted to demonstrate that the proposed LGSR achieves superior results compared with many popular or state-of-the-art algorithms in various image restoration tasks, including denoising, inpainting, and compressive sensing (CS). Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu, Alex Chichung Kot |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2022 | Low-Light Image Enhancement with Normalizing FlowabstractTo enhance low-light images to normally-exposed ones is highly ill-posed, namely that the mapping relationship between them is one-to-many. Previous works based on the pixel-wise reconstruction losses and deterministic processes fail to capture the complex conditional distribution of normally exposed images, which results in improper brightness, residual noise, and artifacts. In this paper, we investigate to model this one-to-many relationship via a proposed normalizing flow model. An invertible network that takes the low-light images/features as the condition and learns to map the distribution of normally exposed images into a Gaussian distribution. In this way, the conditional distribution of the normally exposed images can be well modeled, and the enhancement process, i.e., the other inference direction of the invertible network, is equivalent to being constrained by a loss function that better describes the manifold structure of natural images during the training. The experimental results on the existing benchmark datasets show our method achieves better quantitative and qualitative results, obtaining better-exposed illumination, less noise and artifact, and richer colors. Yufei Wang 0006, Renjie Wan, Wenhan Yang, Haoliang Li, Lap-Pui Chau, Alex Chichung Kot |
AAAI | 6 |
| 2022 | Towards Robust Rain Removal Against Adversarial Attacks: A Comprehensive Benchmark Analysis and BeyondabstractRain removal aims to remove rain streaks from images/videos and reduce the disruptive effects caused by rain. It not only enhances image/video visibility but also allows many computer vision algorithms to function properly. This paper makes the first attempt to conduct a comprehensive study on the robustness of deep learning-based rain removal methods against adversarial attacks. Our study shows that, when the image/video is highly degraded, rain removal methods are more vulnerable to the adversarial attacks as small distortions/perturbations become less noticeable or detectable. In this paper, we first present a comprehensive empirical evaluation of various methods at different levels of attacks and with various losses/targets to generate the perturbations from the perspective of human perception and machine analysis tasks. A systematic evaluation of key modules in existing methods is performed in terms of their robustness against adversarial attacks. From the insights of our analysis, we construct a more robust deraining method by integrating these effective modules. Finally, we examine various types of adversarial attacks that are specific to deraining problems and their effects on both human and machine vision tasks, including 1) rain region attacks, adding perturbations only in the rain regions to make the perturbations in the attacked rain images less visible; 2) object-sensitive attacks, adding perturbations only in regions near the given objects. Code is available at https://github.com/yuyi-sd/Robust_Rain_Removal. Yi Yu 0011, Wenhan Yang, Yap-Peng Tan, Alex Chichung Kot |
CVPR | 4 |
| 2022 | An Improved Two-Stage Based Multi-frame Track-Before-Detect Algorithm in Radar systems
Wujun Li, Kah Chan Teh, Xiujuan Lu, Wei Yi 0002, Alex Chichung Kot |
FUSION | 6 |
| 2022 | Adversarial Pairwise Reverse Attention for Camera Performance Imbalance in Person Re-Identification: New Dataset And MetricsabstractExisting evaluation metrics for Person Re-Identification (Person ReID) models focus on system-wide performance. However, our studies reveal weaknesses due to the uneven data distributions among cameras and different camera properties that expose the ReID system to exploitation. In this work, we raise the long-ignored ReID problem of camera performance imbalance and collect a real-world privacy-aware dataset from 38 cameras to assist the study of the imbalance issue. We propose new metrics to quantify camera performance imbalance and further propose the Adversarial Pairwise Reverse Attention (APRA) Module to guide the model towards learning camera invariant features with a novel pairwise attention inversion mechanism. Eugene P. W. Ang, Rahul Ahuja, Nemath Ahmed, Alex Chichung Kot |
ICIP | 5 |
| 2022 | Image Inpainting Detection via Enriched Attentive Pattern with Near Original Image AugmentationabstractAs deep learning-based inpainting methods have achieved increasingly better results, its malicious use, e.g. removing objects to report fake news or to provide fake evidence, is becoming threatening. Previous works have provided rich discussions on network architectures, e.g. even performing Neural Architecture Search to obtain the optimal model architecture. However, there are rooms in other aspects. In our work, we provide comprehensive efforts from data and feature aspects. From the data aspect, as harder samples in the training data usually lead to stronger detection models, we propose near original image augmentation that pushes the inpainted images closer to the original ones (without distortion and inpainting) as the input images, which is proved to improve the detection accuracy. From the feature aspect, we propose to extract the attentive pattern. With the designed attentive pattern, the knowledge of different inpainting methods can be better exploited during the training phase. Finally, extensive experiments are conducted. In our evaluation, we consider the scenarios where the inpainting masks, which are used to generate the testing set, have a distribution gap from those masks used to produce the training set. Thus, the comparisons are conducted on a newly proposed dataset, where testing masks are inconsistent with the training ones. The experimental results show the superiority of the proposed method and the effectiveness of each component. All our codes and data will be online available. Wenhan Yang, Rizhao Cai, Alex Chichung Kot |
ACM Multimedia | 3 |
| 2022 | On Generating Identifiable Virtual FacesabstractFace anonymization with generative models have become increasingly prevalent since they sanitize private information by generating virtual face images, ensuring both privacy and image utility. Such virtual face images are usually not identifiable after the removal or protection of the original identity. In this paper, we formalize and tackle the problem of generating identifiable virtual face images. Our virtual face images are visually different from the original ones for privacy protection. In addition, they are bound with new virtual identities, which can be directly used for face recognition. We propose an Identifiable Virtual Face Generator (IVFG) to generate the virtual face images. The IVFG projects the latent vectors of the original face images into virtual ones according to a user specific key, based on which the virtual face images are generated. To make the virtual face images identifiable, we propose a multi-task learning objective as well as a triplet styled training strategy to learn the IVFG. We evaluate the performance of our virtual face images using different face recognizers on diffident face image datasets, all of which demonstrate the effectiveness of the IVFG for generate identifiable virtual face images. Zhuowen Yuan, Zhengxin You, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001, Alex Chichung Kot |
ACM Multimedia | 6 |
| 2022 | Toward Efficiently Evaluating the Robustness of Deep Neural Networks in IoT Systems: A GAN-Based MethodabstractIntelligent Internet of Things (IoT) systems based on deep neural networks (DNNs) have been widely deployed in the real world. However, DNNs are found to be vulnerable to adversarial examples, which raises people’s concerns about intelligent IoT systems’ reliability and security. Testing and evaluating the robustness of IoT systems become necessary and essential. Recently, various attacks and strategies have been proposed, but the efficiency problem remains unsolved properly. Existing methods are either computationally extensive or time consuming, which is not applicable in practice. In this article, we propose a novel framework, called attack-inspired generative adversarial networks (AI-GAN) to generate adversarial examples conditionally. Once trained, it can generate adversarial perturbations efficiently given input images and target classes. We apply AI-GAN on different data sets in white-box settings, black-box settings, and targeted models protected by state-of-the-art defenses. Through extensive experiments, AI-GAN achieves high attack success rates, outperforming existing methods, and reduces generation time significantly. Moreover, for the first time, AI-GAN successfully scales to complex data sets, e.g., CIFAR-100 and ImageNet, with about 90% success rates among all classes. Jun Zhao 0007, Jinlin Zhu, Shoudong Han, Jiefeng Chen 0001, Bo Li 0026, Alex Chichung Kot |
IEEE Internet Things J. | 7 |
| 2022 | GMFAD: Towards Generalized Visual Recognition via Multilayer Feature Alignment and DisentanglementabstractThe deep learning based approaches which have been repeatedly proven to bring benefits to visual recognition tasks usually make a strong assumption that the training and test data are drawn from similar feature spaces and distributions. However, such an assumption may not always hold in various practical application scenarios on visual recognition tasks. Inspired by the hierarchical organization of deep feature representation that progressively leads to more abstract features at higher layers of representations, we propose to tackle this problem with a novel feature learning framework, which is called GMFAD, with better generalization capability in a multilayer perceptron manner. We first learn feature representations at the shallow layer where shareable underlying factors among domains (e.g., a subset of which could be relevant for each particular domain) can be explored. In particular, we propose to align the domain divergence between domain pair(s) by considering both inter-dimension and inter-sample correlations, which have been largely ignored by many cross-domain visual recognition methods. Subsequently, to learn more abstract information which could further benefit transferability, we propose to conduct feature disentanglement at the deep feature layer. Extensive experiments based on different visual recognition tasks demonstrate that our proposed framework can learn better transferable feature representation compared with state-of-the-art baselines. Haoliang Li, Shiqi Wang 0001, Renjie Wan, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Skeleton-based relational reasoning for group activity analysis
Mauricio Perez, Jun Liu 0036, Alex Chichung Kot |
Pattern Recognit. | 3 |
| 2022 | Learning Meta Pattern for Face Anti-SpoofingabstractFace Anti-Spoofing (FAS) is essential to secure face recognition systems and has been extensively studied in recent years. Although deep neural networks (DNNs) for the FAS task have achieved promising results in intra-dataset experiments with similar distributions of training and testing data, the DNNs’ generalization ability is limited under the cross-domain scenarios with different distributions of training and testing data. To improve the generalization ability, recent hybrid methods have been explored to extract task-aware handcrafted features (e.g., Local Binary Pattern) as discriminative information for the input of DNNs. However, the handcrafted feature extraction relies on experts’ domain knowledge, and how to choose appropriate handcrafted features is underexplored. To this end, we propose a learnable network to extract Meta Pattern (MP) in our learning-to-learn framework. By replacing handcrafted features with the MP, the discriminative information from MP is capable of learning a more generalized model. Moreover, we devise a two-stream network to hierarchically fuse the input RGB image and the extracted MP by using our proposed Hierarchical Fusion Module (HFM). We conduct comprehensive experiments and show that our MP outperforms the compared handcrafted features. Also, our proposed method with HFM and the MP can achieve state-of-the-art performance on two different domain generalization evaluation benchmarks. Rizhao Cai, Zhi Li 0054, Renjie Wan, Haoliang Li, Yongjian Hu, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2022 | One-Class Knowledge Distillation for Face Presentation Attack DetectionabstractFace presentation attack detection (PAD) has been extensively studied by research communities to enhance the security of face recognition systems. Although existing methods have achieved good performance on testing data with similar distribution as the training data, their performance degrades severely in application scenarios with data of unseen distributions. In situations where the training and testing data are drawn from different domains, a typical approach is to apply domain adaptation techniques to improve face PAD performance with the help of target domain data. However, it has always been a non-trivial challenge to collect sufficient data samples in the target domain, especially for attack samples. This paper introduces a teacher-student framework to improve the cross-domain performance of face PAD with one-class domain adaptation. In addition to the source domain data, the framework utilizes only a few genuine face samples of the target domain. Under this framework, a teacher network is trained with source domain samples to provide discriminative feature representations for face PAD. Student networks are trained to mimic the teacher network and learn similar representations for genuine face samples of the target domain. In the test phase, the similarity score between the representations of the teacher and student networks is used to distinguish attacks from genuine ones. To evaluate the proposed framework under one-class domain adaptation settings, we devised two new protocols and conducted extensive experiments. The experimental results show that our method outperforms baselines under one-class domain adaptation settings and even state-of-the-art methods with unsupervised domain adaptation. Zhi Li 0054, Rizhao Cai, Haoliang Li, Kwok-Yan Lam, Yongjian Hu, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2022 | Motion-Adaptive Detection of HEVC Double Compression With the Same Coding ParametersabstractHigh Efficiency Video Coding (HEVC) double compression detection is of prime significance in video forensics. However, double compression with the same parameters and video content with high motion displacement intensity have become two main factors that limit the performance of existing algorithms. To address these issues, a novel motion-adaptive algorithm is proposed in this paper. Firstly, the analysis of GOP structure in HEVC standard and the coding process of HEVC double compression are provided. Next, sub-features composed of fluctuation intensities of intra prediction modes and unstable Prediction Units (PUs) in normal Intra-Frames (I-frames) and optical flow in adaptive I-frames are exploited in our algorithm. Each sub-feature is extracted during the process of multiple decompression. We further combine these sub-features into a 27-dimensional detection feature, which is finally fed to the Support Vector Machine (SVM) classifier. By following a separation-fusion detection strategy, the experimental result shows that the proposed algorithm outperforms the existing state-of-the-art methods and demonstrates superior robustness to various motion displacement intensities and a wide variety of coding parameter settings. Qiang Xu 0007, Xinghao Jiang, Tanfeng Sun, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2022 | Interaction Relational Network for Mutual Action RecognitionabstractPerson-person mutual action recognition (also referred to as interaction recognition) is an important research branch of human activity analysis. Current solutions in the field – mainly dominated by CNNs, GCNs and LSTMs – often consist of complicated architectures and mechanisms to embed the relationships between the two persons on the architecture itself, to ensure the interaction patterns can be properly learned. Our main contribution with this work is by proposing a simpler yet very powerful architecture, named Interaction Relational Network, which utilizes minimal prior knowledge about the structure of the human body. We drive the network to identify by itself how to relate the body parts from the individuals interacting. In order to better represent the interaction, we define two different relationships, leading to specialized architectures and models for each. These multiple relationship models will then be fused into a single and special architecture, in order to leverage both streams of information for further enhancing the relational reasoning capability. Furthermore we define important structured pair-wise operations to extract meaningful extra information from each pair of joints – distance and motion. Ultimately, with the coupling of an LSTM, our IRN is capable of paramount sequential relational reasoning. These important extensions we made to our network can also be valuable to other problems that require sophisticated relational reasoning. Our solution is able to achieve state-of-the-art performance on the traditional interaction recognition datasets SBU and UT, and also on the mutual actions from the large-scale dataset NTU RGB+D. Furthermore, it obtains competitive performance in the NTU RGB+D 120 dataset interactions subset. Mauricio Perez, Jun Liu 0036, Alex Chichung Kot |
IEEE Trans. Multim. | 3 |
| 2022 | $A^3$-FKG: Attentive Attribute-Aware Fashion Knowledge Graph for Outfit Preference PredictionabstractWith the booming development of the online fashion industry, effective personalized recommender systems have become indispensable for the convenience they brought to the customers and the profits to the e-commercial platforms. Estimating the user’s preference towards the outfit is at the core of a personalized recommendation system. Existing works on fashion recommendation are largely centering on modelling the clothing compatibility without considering the user factor or characterizing the user’s preference over the single item. However, how to effectively model the outfits with either few or even none interactions, is yet under-explored. In this paper, we address the task of personalized outfit preference prediction via a novelAttentiveAttribute-AwareFashionKnowledgeGraph ($A^3$-FKG), which is incorporated to build the association between different outfits with both outfit- and item- level attributes. Additionally, a two-level attention mechanism is developed to capture the user’s preference: 1) User-specific relation-aware attention layer, which captures the user’s fine-grained preferences with different focus on relations for learning outfit representation; 2) Target-aware attention layer, which characterizes the user’s latent diverse interests from his/her behavior sequences for learning user representation. Extensive experiments conducted on a large-scale fashion outfit dataset demonstrate significant improvements over other methods, which verify the excellence of our proposed framework. Huijing Zhan, Jie Lin 0001, Kenan E. Ak, Boxin Shi, Ling-Yu Duan, Alex Chichung Kot |
IEEE Trans. Multim. | 6 |
| 2022 | A Hybrid Structural Sparsification Error Model for Image RestorationabstractRecent works on structural sparse representation (SSR), which exploit image nonlocal self-similarity (NSS) prior by grouping similar patches for processing, have demonstrated promising performance in various image restoration applications. However, conventional SSR-based image restoration methods directly fit the dictionaries or transforms to the internal (corrupted) image data. The trained internal models inevitably suffer from overfitting to data corruption, thus generating the degraded restoration results. In this article, we propose a novel hybrid structural sparsification error (HSSE) model for image restoration, which jointly exploits image NSS prior using both the internal and external image data that provide complementary information. Furthermore, we propose a general image restoration scheme based on the HSSE model, and an alternating minimization algorithm for a range of image restoration applications, including image inpainting, image compressive sensing and image deblocking. Extensive experiments are conducted to demonstrate that the proposed HSSE-based scheme outperforms many popular or state-of-the-art image restoration methods in terms of both objective metrics and visual perception. Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu, Alex Chichung Kot |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2021 | DEX: Domain Embedding Expansion for Generalized Person Re-identification
Eugene P. W. Ang, Alex Chichung Kot |
BMVC | 3 |
| 2021 | Panoramic Image Reflection RemovalabstractThis paper studies the problem of panoramic image reflection removal, aiming at reliving the content ambiguity between reflection and transmission scenes. Although a partial view of the reflection scene is included in the panoramic image, it cannot be utilized directly due to its misalignment with the reflection-contaminated image. We propose a two-step approach to solve this problem, by first accomplishing geometric and photometric alignment for the reflection scene via a coarse-to-fine strategy, and then restoring the transmission scene via a recovery network. The proposed method is trained with a synthetic dataset and verified quantitatively with a real panoramic image dataset. The effectiveness of the proposed method is validated by the significant performance advantage over single image-based reflection removal methods and generalization capacity to limited-FoV scenarios captured by conventional camera or mobile phone users. Yuchen Hong, Lingran Zhao, Xudong Jiang 0001, Alex Chichung Kot, Boxin Shi |
CVPR | 5 |
| 2021 | Single Image Reflection Removal With Absorption EffectabstractIn this paper, we consider the absorption effect for the problem of single image reflection removal. We show that the absorption effect can be numerically approximated by the average of refractive amplitude coefficient map. We then reformulate the image formation model and propose a two-step solution that explicitly takes the absorption effect into account. The first step estimates the absorption effect from a reflection-contaminated image, while the second step recovers the transmission image by taking a reflection-contaminated image and the estimated absorption effect as the input. Experimental results on four public datasets show that our two-step solution not only successfully removes reflection artifact, but also faithfully restores the intensity distortion caused by the absorption effect. Our ablation studies further demonstrate that our method achieves superior performance on the recovery of overall intensity and has good model generalization capacity. The code is available at https://github.com/q-zh/absorption. Boxin Shi, Jinnan Chen, Xudong Jiang 0001, Ling-Yu Duan, Alex Chichung Kot |
CVPR | 6 |
| 2021 | A Compact Joint Distillation Network for Visual Food RecognitionabstractVisual food recognition is emerging as an important application in dietary monitoring and management in recent years. Existing works use large backbone networks to achieve good performance. However, these networks are not able to be deployed on personal portable devices due to large size and computation cost. Some compact networks have been developed, however, their performance are usually lower than the large backbone networks. In view of this, this paper proposes a joint distillation framework that targets to achieve a high visual food recognition accuracy using a compact network. As opposed to the more traditional one-directional knowledge distillation methods, the proposed knowledge distillation framework trains both the large teacher network and the compact student network simultaneously. The framework introduces a new Multi-Layer Distillation (MLD) for simultaneous teacher-student learning at multiple layers of different abstraction. A novel Instance Activation Mapping (IAM) is proposed to jointly train the teacher and student networks using generated instance-level activation map that incorporates label information for each training image. Experimental results on the two benchmark datasets UECFood-256 and Food-101 show that the trained compact student network achieves state-of-the-art performance at 83.5% and 90.4%, respectively, while achieving more than 4 times deduction regarding network model size. Zhao Heng, Kim-Hui Yap, Alex Chichung Kot |
ICASSP | 3 |
| 2021 | Skeleton Cloud Colorization for Unsupervised 3D Action Representation LearningabstractSkeleton-based human action recognition has attracted increasing attention in recent years. However, most of the existing works focus on supervised learning which requiring a large number of annotated action sequences that are often expensive to collect. We investigate unsupervised representation learning for skeleton action recognition, and design a novel skeleton cloud colorization technique that is capable of learning skeleton representations from unlabeled skeleton sequence data. Specifically, we represent a skeleton action sequence as a 3D skeleton cloud and colorize each point in the cloud according to its temporal and spatial orders in the original (unannotated) skeleton sequence. Leveraging the colorized skeleton point cloud, we design an auto-encoder framework that can learn spatial-temporal features from the artificial color labels of skeleton joints effectively. We evaluate our skeleton cloud colorization approach with action classifiers trained under different configurations, including unsupervised, semi-supervised and fully-supervised settings. Extensive experiments on NTU RGB+D and NW-UCLA datasets show that the proposed method outperforms existing unsupervised and semi-supervised 3D action recognition methods by large margins, and it achieves competitive performance in supervised 3D action recognition as well. Siyuan Yang 0001, Jun Liu 0036, Shijian Lu, Meng Hwa Er, Alex Chichung Kot |
ICCV | 5 |
| 2021 | AI-GAN: Attack-Inspired Generation of Adversarial ExamplesabstractDeep neural networks (DNNs) are vulnerable to adversarial examples, which are crafted by adding imperceptible perturbations to inputs. Recently different attacks and strategies have been proposed, but how to generate adversarial examples perceptually realistic and more efficiently remains unsolved. This paper proposes a novel framework called Attack-Inspired GAN (AI-GAN), where a generator, a discriminator, and an attacker are trained jointly. Once trained, it can generate adversarial perturbations efficiently given input images and target classes. Through extensive experiments on several popular datasets e.g., MNIST and CFAR-10, AI-GAN achieves high attack success rates and reduces generation time significantly in various settings. Moreover, for the first time, AI-GAN successfully scales to complicated datasets e.g., CFAR-100 with around 90% success rates among all classes. Jun Zhao 0007, Jinlin Zhu, Shoudong Han, Jiefeng Chen 0001, Bo Li 0026, Alex Chichung Kot |
ICIP | 7 |
| 2021 | Multi-Scale Feature Guided Low-Light Image EnhancementabstractLow-light image enhancement aims at enlarging the intensity of image pixels to better match human perception and to improve the performance of subsequent vision tasks. While it is relatively easy to enlighten a globally low-light image, the lighting condition of realistic scenes is usually non-uniform and complex, e.g., some images may contain both bright and extremely dark regions, with or without rich features and information. Existing methods often generate abnormal light-enhancement results with over-exposure artifacts without proper guidance. To tackle this challenge, we propose a multi-scale feature guided attention mechanism in the deep generator, which can effectively perform a spatially-varying light enhancement. The attention map is fused by both the gray map and extracted feature map of the input image, to focus more on those dark and informative regions. Our baseline is an unsupervised generative adversarial network, which can be trained without any low/normal light image pair. Experimental results demonstrate the superiority in visual quality and performance of subsequent object detection over state-of-the-art alternatives. Lanqing Guo, Renjie Wan, Guan-Ming Su, Alex Chichung Kot, Bihan Wen |
ICIP | 4 |
| 2021 | Embracing the Dark Knowledge: Domain Generalization Using Regularized Knowledge DistillationabstractThough convolutional neural networks are widely used in different tasks, lack of generalization capability in the absence of sufficient and representative data is one of the challenges that hinders their practical application. In this paper, we propose a simple, effective, and plug-and-play training strategy named Knowledge Distillation for Domain Generalization (KDDG) which is built upon a knowledge distillation framework with the gradient filter as a novel regularization term. We find that both the "richer dark knowledge" from the teacher network, as well as the gradient filter we proposed, can reduce the difficulty of learning the mapping which further improves the generalization ability of the model. We also conduct experiments extensively to show that our framework can significantly improve the generalization capability of deep neural networks in different tasks including image classification, segmentation, reinforcement learning by comparing our method with existing state-of-the-art domain generalization techniques. Last but not the least, we propose to adopt two metrics to analyze our proposed method in order to better understand how our proposed method benefits the generalization capability of deep neural networks. Yufei Wang 0006, Haoliang Li, Lap-Pui Chau, Alex Chichung Kot |
ACM Multimedia | 4 |
| 2021 | Fusion Learning using Semantics and Graph Convolutional Network for Visual Food RecognitionabstractFood-related applications and services are essential for the health and well-being of people. With the rapid development of social networks and mobile devices, food images captured by people can offer rich knowledge about the food and also necessary dietary assistance for people that require special care. Known food recognition frameworks and approaches in computer vision have heavy reliance on many-shot training of a deep network on existing large-scale food datasets. However, it is common for many food categories that it is difficult to collect enough images for training. Traditional few-shot learning is unable to properly address the problem due to the complex characteristics and large variations of food images, and most few-shot frame-works cannot perform classification for many-shot and few-shot categories at the same time. In this paper, we propose a new fusion learning framework for food recognition. It unifies many-shot and few-shot under a single framework, by leveraging on extracted image representations and context sensitive semantic embeddings. Further, considering food categories are often correlated to each other for many commonalities such as same ingredients, cooking methods, the fusion learning framework utilizes a Graph Convolutional Network (GCN) to capture the inter-class relations between both image representations and semantic embeddings of different food categories. The final output fusion classifier will be more robust and discriminative. Comprehensive experimental results on two popular food benchmarks have shown the proposed framework achieves the state-of-the-art fusion performance. Kim-Hui Yap, Alex Chichung Kot |
WACV | 3 |
| 2021 | Unsupervised Domain Adaptation in the Wild via Disentangling Representation Learning
Haoliang Li, Renjie Wan, Shiqi Wang 0001, Alex Chichung Kot |
Int. J. Comput. Vis. | 4 |
| 2021 | Face Image Reflection Removal
Renjie Wan, Boxin Shi, Haoliang Li, Ling-Yu Duan, Alex Chichung Kot |
Int. J. Comput. Vis. | 5 |
| 2021 | Detection of HEVC double compression with non-aligned GOP structures via inter-frame quality degradation analysis
Qiang Xu 0007, Xinghao Jiang, Tanfeng Sun, Alex Chichung Kot |
Neurocomputing | 4 |
| 2021 | DeepImaging: A Ground Moving Target Imaging Based on CNN for SAR-GMTI SystemabstractImaging of ground multiple moving targets in a synthetic aperture radar (SAR) system is a challenging task due to the fact that targets are defocused owing to motions and contaminated by the strong background clutter. Motivated by recent advances in deep learning, a novel deep convolutional neural network (CNN)-based method, DeepImaging, is proposed for ground moving target imaging (GMTIm). Different from conventional imaging methods relying on the prior knowledge of imaging, the proposed DeepImaging is directly trained to learn an implicit imaging model of multiple moving targets. It is free of motion parameter estimation and iteration process. Then, the trained DeepImaging, as an imaging processor, can be applied to the SAR complex received data after clutter suppression to achieve the multiple moving target imaging and the residual clutter elimination simultaneously. Simulations and experiments on the Gotcha data show that the proposed method achieves significant improvements over existing state-of-the-art GMTIm methods in terms of imaging quality and efficiency. Huilin Mu, Yun Zhang 0023, Meng Hwa Er, Alex Chichung Kot |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2021 | Detection of transcoded HEVC videos based on in-loop filtering and PU partitioning analyses
Qiang Xu 0007, Xinghao Jiang, Tanfeng Sun, Alex Chichung Kot |
Signal Process. Image Commun. | 4 |
| 2021 | Detection of Spoofing Medium Contours for Face Anti-SpoofingabstractFace anti-spoofing is an important step for secure face recognition. In this paper, we target on building a general classifier to detect the face images with spoofing medium contours (termed as SMCs for simplicity). To this end, we consider the task of face anti-spoofing as the detection of SMCs from the image. We propose and train a Contour Enhanced Mask R-CNN (CEM-RCNN) model for the detection. This model detects the existence of the SMCs by incorporating the contour objectness which measures how likely an object contains the SMCs. The experimental results demonstrate the generality of the CEM-RCNN for identifying the face images with SMCs, which performs significantly better than the state-of-the-art on the cross-database scenario. Sheng Li 0006, Xinpeng Zhang 0001, Haoliang Li, Alex Chichung Kot |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | DRL-FAS: A Novel Framework Based on Deep Reinforcement Learning for Face Anti-SpoofingabstractInspired by the philosophy employed by human beings to determine whether a presented face example is genuine or not, i.e., to glance at the example globally first and then carefully observe the local regions to gain more discriminative information, for the face anti-spoofing problem, we propose a novel framework based on the Convolutional Neural Network (CNN) and the Recurrent Neural Network (RNN). In particular, we model the behavior of exploring face-spoofing-related information from image sub-patches by leveraging deep reinforcement learning. We further introduce a recurrent mechanism to learn representations of local information sequentially from the explored sub-patches with an RNN. Finally, for the classification purpose, we fuse the local information with the global one, which can be learned from the original input image through a CNN. Moreover, we conduct extensive experiments, including ablation study and visualization analysis, to evaluate our proposed framework on various public databases. The experiment results show that our method can generally achieve state-of-the-art performance among all scenarios, demonstrating its effectiveness. Rizhao Cai, Haoliang Li, Shiqi Wang 0001, Changsheng Chen 0001, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2021 | Multi-Domain Adversarial Feature Generalization for Person Re-IdentificationabstractWith the assistance of sophisticated training methods applied to single labeled datasets, the performance of fully-supervised person re-identification (Person Re-ID) has been improved significantly in recent years. However, these models trained on a single dataset usually suffer from considerable performance degradation when applied to videos of a different camera network. To make Person Re-ID systems more practical and scalable, several cross-dataset domain adaptation methods have been proposed, which achieve high performance without the labeled data from the target domain. However, these approaches still require the unlabeled data of the target domain during the training process, making them impractical. A practical Person Re-ID system pre-trained on other datasets should start running immediately after deployment on a new site without having to wait until sufficient images or videos are collected and the pre-trained model is tuned. To serve this purpose, in this paper, we reformulate person re-identification as a multi-dataset domain generalization problem. We propose a multi-dataset feature generalization network (MMFA-AAE), which is capable of learning a universal domain-invariant feature representation from multiple labeled datasets and generalizing it to 'unseen' camera systems. The network is based on an adversarial auto-encoder to learn a generalized domain-invariant latent feature representation with the Maximum Mean Discrepancy (MMD) measure to align the distributions across multiple domains. Extensive experiments demonstrate the effectiveness of the proposed method. Our MMFA-AAE approach not only outperforms most of the domain generalization Person Re-ID methods, but also surpasses many state-of-the-art supervised methods and unsupervised domain adaptation methods by a large margin. Chang-Tsun Li, Alex Chichung Kot |
IEEE Trans. Image Process. | 3 |
| 2021 | Pose-Normalized and Appearance-Preserved Street-to-Shop Clothing Image Generation and Feature LearningabstractWe tackle the task of street-to-shop clothing image synthesis. Given a daily person image with a particular clothing item captured in the street scenario, we aim to synthesize the frontal facing view of that item in the shop scenario. This problem has the following challenges: 1) the distinct visual discrepancy between the street and shop scenario; 2) the severe shape deformation of clothing in the presence of an arbitrary human pose; 3) the preservation of fine-grained details during the process of clothing image generation. In this paper, we jointly solve these difficulties by proposing a Pose-Normalized and Appearance-Preserved Generative Adversarial Network (PNAP-GAN). More specifically, conditioned on the clothing-agnostic representation (i.e., clothing landmarks and semantic parsing map), we disentangle the shape and appearance synthesis in a coarse-to-fine framework. Moreover, a semantic embedding loss is introduced to guide the domain transfer in the semantic level (i.e., keeping the clothing attributes). With the synthesized frontal shop image, a pose-normalized representation in complementary to the domain-invariant feature learnt from the original street image are integrated to facilitate the problem of street-to-shop clothing retrieval. Extensive experiments conducted demonstrate the effectiveness of the proposed PNAP-GAN on generating high quality frontal-view images and the excellence of the learnt pose-normalized features on the retrieval task than existing methods. In addition, we demonstrate that the pose-normalized retrieval feature benefits the cross-scenario (i.e., street-to-shop) clothing image generation in a semantic-preserved manner. Huijing Zhan, Chenyu Yi, Boxin Shi, Jie Lin 0001, Ling-Yu Duan, Alex Chichung Kot |
IEEE Trans. Multim. | 6 |
| 2020 | Reflection Scene Separation From a Single ImageabstractFor images taken through glass, existing methods focus on the restoration of the background scene by regarding the reflection components as noise. However, the scene reflected by glass surface also contains important information to be recovered, especially for the surveillance or criminal investigations. In this paper, instead of removing reflection components from the mixture image, we aim at recovering reflection scenes from the mixture image. We first propose a strategy to obtain such ground truth and its corresponding input images. Then, we propose a two-stage framework to obtain the visible reflection scene from the mixture image. Specifically, we train the network with a shift-invariant loss which is robust to misalignment between the input and output images. The experimental results show that our proposed method achieves promising results. Renjie Wan, Boxin Shi, Haoliang Li, Ling-Yu Duan, Alex Chichung Kot |
CVPR | 5 |
| 2020 | What Does Plate Glass Reveal About Camera Calibration?abstractThis paper aims to calibrate the orientation of glass and the field of view of the camera from a single reflection-contaminated image. We show how a reflective amplitude coefficient map can be used as a calibration cue. Different from existing methods, the proposed solution is free from image contents. To reduce the impact of a noisy calibration cue estimated from a reflection-contaminated image, we propose two strategies: an optimization-based method that imposes part of though reliable entries on the map and a learning-based method that fully exploits all entries. We collect a dataset containing 320 samples as well as their camera parameters for evaluation. We demonstrate that our method not only facilitates a general single image camera calibration method that leverages image contents but also contributes to improving the performance of single image reflection removal. Furthermore, we show our byproduct output helps alleviate the ill-posed problem of estimating the panorama from a single image. Jinnan Chen, Zhan Lu, Boxin Shi, Xudong Jiang 0001, Kim-Hui Yap, Ling-Yu Duan, Alex Chichung Kot |
CVPR | 8 |
| 2020 | Collaborative Learning of Gesture Recognition and 3D Hand Pose Estimation with Multi-order Feature Analysis
Siyuan Yang 0001, Jun Liu 0036, Shijian Lu, Meng Hwa Er, Alex Chichung Kot |
ECCV (3) | 5 |
| 2020 | Splitting Vs. Merging: Mining Object Regions with Discrepancy and Intersection Loss for Weakly Supervised Semantic Segmentation
Tianyi Zhang 0004, Guosheng Lin, Weide Liu, Jianfei Cai 0001, Alex Chichung Kot |
ECCV (22) | 5 |
| 2020 | A New Efficient Finger-Vein Verification Based on Lightweight Neural Network Using Multiple Schemes
Haocong Zheng, Yongjian Hu, Alex Chichung Kot |
ICANN (1) | 5 |
| 2020 | Unseen Face Presentation Attack Detection with Hypersphere LossabstractPresentation attack is one of the main threats to face verification systems and attracts great attention of research community. Recent methods achieve great success in intra-database test. However, the problem is more complex in practical scenario as the type of attack could be unseen to system designers. In this paper, we formulate the face presentation attack detection task under an open-set setting and address with our proposed deep anomaly detection based method. The training process is end-to-end supervised by a novel hypersphere loss function and the decision making is directly based on the learned feature representation. We conduct extensive experiments on multiple prevailing databases and evaluate our implemented models by using various metrics. The results show our proposed method is effective against unseen types of attacks and superior to latest state-of-the-art. Zhi Li 0054, Haoliang Li, Kwok-Yan Lam, Alex Chichung Kot |
ICASSP | 4 |
| 2020 | Heterogeneous Domain Generalization Via Domain MixupabstractOne of the main drawbacks of deep Convolutional Neural Networks (DCNN) is that they lack generalization capability. In this work, we focus on the problem of heterogeneous domain generalization which aims to improve the generalization capability across different tasks, which is, how to learn a DCNN model with multiple domain data such that the trained feature extractor can be generalized to supporting recognition of novel categories in a novel target domain. To solve this problem, we propose a novel heterogeneous domain generalization method by mixing up samples across multiple source domains with two different sampling strategies. Our experimental results based on the Visual Decathlon benchmark demonstrates the effectiveness of our proposed method. Yufei Wang 0006, Haoliang Li, Alex Chichung Kot |
ICASSP | 3 |
| 2020 | Data Representation in Hybrid Coding Framework for Feature Maps CompressionabstractRecently, a new paradigm of transmitting and compressing intermediate deep learning features (i.e., feature maps) for distributed visual analysis systems is emerging. As the fundamental infrastructure in such paradigm, research and standardization for feature maps coding has attracted more and more attention. In this paper, to improve the state-of-the-art hybrid coding framework which integrates the traditional video codecs to compress feature maps, we investigate the data representation procedure in such coding framework. Specifically, we proposed three modes in Repack module to help explore inter-channel redundancy, and we explore the fidelity maintenance ability of two modes in Pre-Quantization modules. It is worth mentioning that the proposed coding modes have been partially adopted in to the ongoing AVS (Audio Video Coding Standard Workgroup) - Visual Feature Coding Standard. Zhuo Chen 0006, Ling-Yu Duan, Shiqi Wang 0001, Weisi Lin, Alex Chichung Kot |
ICIP | 5 |
| 2020 | Attention Selective Network For Face Synthesis And Pose-Invariant Face RecognitionabstractFace recognition algorithms have improved significantly in recent years since the introduction of deep learning and the availability of large training datasets. However, their performance is still inadequate when the face pose varies as pose variation can dramatically increase intra-person variability. This work proposes a novel generative adversarial architecture called the Attention Selective Network (ASN) to address the problem of pose-invariant face recognition. The ASN introduces an efficient attention mechanism and a multi-part loss function to generate realistic-looking frontal face images from other face poses that can be used for recognizing faces under various poses. Thanks to the high quality of the synthesized images, the ASN achieves superior performance in terms of recognition rates compared to the state-of-the-art supervised methods. Jiashu Liao, Alex Chichung Kot, Tanaya Guha, Victor Sanchez |
ICIP | 2 |
| 2020 | Domain Generalization for Medical Imaging Classification with Linear-Dependency RegularizationabstractRecently, we have witnessed great progress in the field of medical imaging classification by adopting deep neural networks. However, the recent advanced models still require accessing sufficiently large and representative datasets for training, which is often unfeasible in clinically realistic environments. When trained on limited datasets, the deep neural network is lack of generalization capability, as the trained deep neural network on data within a certain distribution (e.g. the data captured by a certain device vendor or patient population) may not be able to generalize to the data with another distribution. In this paper, we introduce a simple but effective approach to improve the generalization capability of deep neural networks in the field of medical imaging classification. Motivated by the observation that the domain variability of the medical images is to some extent compact, we propose to learn a representative feature space through variational encoding with a novel linear-dependency regularization term to capture the shareable information among medical data collected from different domains. As a result, the trained neural network is expected to equip with better generalization capability to the ``unseen" medical data. Experimental results on two challenging medical imaging classification tasks indicate that our method can achieve better cross-domain generalization capability compared with state-of-the-art baselines. Haoliang Li, Yufei Wang 0006, Renjie Wan, Shiqi Wang 0001, Tie-Qiang Li, Alex Chichung Kot |
NeurIPS | 6 |
| 2020 | Improving Robustness of DNNs against Common Corruptions via Gaussian Adversarial TrainingabstractDeep neural networks have demonstrated tremendous success in image classification, but their performance sharply degrades when evaluated on slightly different test data (e.g., data with corruptions). To address these issues, we propose a minimax approach to improve common corruption robustness of deep neural networks via Gaussian Adversarial Training. To be specific, we propose to train neural networks with adversarial examples where the perturbations are Gaussian-distributed. Our experiments show that our proposed GAT can improve neural networks' robustness to noise corruptions more than other baseline methods. It also outperforms the state-of-the-art method in improving the overall robustness to common corruptions. Chenyu Yi, Haoliang Li, Renjie Wan, Alex Chichung Kot |
VCIP | 4 |
| 2020 | The Enhancement of Underexposed Images with Blurred ReflectanceabstractThe images captured in the low-light conditions always suffer from low visibility. Enhancing the visibility of the low-light image is of broad application to various computer vision tasks. Based on the classical Retinex model, previous methods assume the reflectance components as a well-exposed image. In this paper, we introduce the blurring distortion into the Retinex model to cover more general and challenging scenarios. We further propose a two-stage framework to extract the reflectance images and remove the blurring distortion separately. Specifically, we optimize the whole network by embedding a mechanism robust to the pixel misalignment in the training dataset. The experimental results show that our proposed method achieves promising results. Jinchao Zhou, Renjie Wan, Haoliang Li, Alex Chichung Kot |
VCIP | 4 |
| 2020 | Feature Boosting Network For 3D Pose EstimationabstractIn this paper, a feature boosting network is proposed for estimating 3D hand pose and 3D body pose from a single RGB image. In this method, the features learned by the convolutional layers are boosted with a new long short-term dependence-aware (LSTD) module, which enables the intermediate convolutional feature maps to perceive the graphical long short-term dependency among different hand (or body) parts using the designed Graphical ConvLSTM. Learning a set of features that are reliable and discriminatively representative of the pose of a hand (or body) part is difficult due to the ambiguities, texture and illumination variation, and self-occlusion in the real application of 3D pose estimation. To improve the reliability of the features for representing each body part and enhance the LSTD module, we further introduce a context consistency gate (CCG) in this paper, with which the convolutional feature maps are modulated according to their consistency with the context representations. We evaluate the proposed method on challenging benchmark datasets for 3D hand pose estimation and 3D full body pose estimation. Experimental results show the effectiveness of our method that achieves state-of-the-art performance on both of the tasks. Jun Liu 0036, Henghui Ding, Amir Shahroudy, Ling-Yu Duan, Xudong Jiang 0001, Gang Wang 0012, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2020 | NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity UnderstandingabstractResearch on depth-based human activity analysis achieved outstanding performance and demonstrated the effectiveness of 3D representation for action recognition. The existing depth-based and RGB+D-based action recognition benchmarks have a number of limitations, including the lack of large-scale training samples, realistic number of distinct class categories, diversity in camera views, varied environmental conditions, and variety of human subjects. In this work, we introduce a large-scale dataset for RGB+D human action recognition, which is collected from 106 distinct subjects and contains more than 114 thousand video samples and 8 million frames. This dataset contains 120 different action classes including daily, mutual, and health-related activities. We evaluate the performance of a series of existing 3D activity analysis methods on this dataset, and show the advantage of applying deep learning methods for 3D-based human action recognition. Furthermore, we investigate a novel one-shot 3D activity recognition problem on our dataset, and a simple yet effective Action-Part Semantic Relevance-aware (APSR) framework is proposed for this task, which yields promising results for recognition of the novel action classes. We believe the introduction of this large-scale dataset will enable the community to apply, adapt, and develop various data-hungry learning techniques for depth-based and RGB+D-based human activity understanding. Jun Liu 0036, Amir Shahroudy, Mauricio Perez, Gang Wang 0012, Ling-Yu Duan, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2020 | Skeleton-Based Online Action Prediction Using Scale Selection NetworkabstractAction prediction is to recognize the class label of an ongoing activity when only a part of it is observed. In this paper, we focus on online action prediction in streaming 3D skeleton sequences. A dilated convolutional network is introduced to model the motion dynamics in temporal dimension via a sliding window over the temporal axis. Since there are significant temporal scale variations in the observed part of the ongoing action at different time steps, a novel window scale selection method is proposed to make our network focus on the performed part of the ongoing action and try to suppress the possible incoming interference from the previous actions at each step. An activation sharing scheme is also proposed to handle the overlapping computations among the adjacent time steps, which enables our framework to run more efficiently. Moreover, to enhance the performance of our framework for action prediction with the skeletal input data, a hierarchy of dilated tree convolutions are also designed to learn the multi-level structured semantic representations over the skeleton joints at each frame. Our proposed approach is evaluated on four challenging datasets. The extensive experiments demonstrate the effectiveness of our method for skeleton-based online action prediction. Jun Liu 0036, Amir Shahroudy, Gang Wang 0012, Ling-Yu Duan, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2020 | CoRRN: Cooperative Reflection Removal NetworkabstractRemoving the undesired reflections from images taken through the glass is of broad application to various computer vision tasks. Non-learning based methods utilize different handcrafted priors such as the separable sparse gradients caused by different levels of blurs, which often fail due to their limited description capability to the properties of real-world reflections. In this paper, we propose a network with the feature-sharing strategy to tackle this problem in a cooperative and unified framework, by integrating image context information and the multi-scale gradient information. To remove the strong reflections existed in some local regions, we propose a statistic loss by considering the gradient level statistics between the background and reflections. Our network is trained on a new dataset with 3250 reflection images taken under diverse real-world scenes. Experiments on a public benchmark dataset show that the proposed method performs favorably against state-of-the-art methods. Renjie Wan, Boxin Shi, Haoliang Li, Ling-Yu Duan, Ah-Hwee Tan, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2020 | Face Spoofing Detection Based on Local Ternary Label Supervision in Fully Convolutional NetworksabstractFace verification systems are prone to spoofing attacks on photos, videos, and 3D masks. Face spoofing detection, i.e., face anti-spoofing, face liveness detection, or face presentation attack detection, is an important task for securing face verification systems in practice and presents many challenges. In this paper, a state-of-the-art face spoofing detection method based on a depth-based Fully Convolutional Network (FCN) is revisited. Different supervision schemes, including global and local label supervisions, are comprehensively investigated. A generic theoretical analysis and associated simulation are provided to demonstrate that local label supervision is more suitable than global label supervision for local tasks with insufficient training samples, such as the face spoofing detection task. Based on the analysis, the Spatial Aggregation of Pixel-level Local Classifiers (SAPLC), which is composed of an FCN part and an aggregation part, is proposed. The FCN part predicts the pixel-level ternary labels, which include the genuine foreground, the spoofed foreground, and the undetermined background. Then, these labels are aggregated together to yield an accurate image-level decision. Furthermore, to quantitatively evaluate the proposed SAPLC, experiments are carried out on the CASIA-FASD, Replay-Attack, OULU-NPU, and SiW datasets. The experiments show that the proposed SAPLC outperforms the representative deep networks, including two globally supervised CNNs, one depth-based FCN, two FCNs with binary labels, and two FCNs with ternary labels, and achieves competitive performances close to some state-of-the-art method performances under various common protocols. Overall, the results empirically verify the advantage of the proposed pixel-level local label supervision scheme. Wenyun Sun, Changsheng Chen 0001, Jiwu Huang, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2020 | Toward Intelligent Sensing: Intermediate Deep Feature CompressionabstractThe recent advances of hardware technology have made the intelligent analysis equipped at the front-end with deep learning more prevailing and practical. To better enable the intelligent sensing at the front-end, instead of compressing and transmitting visual signals or the ultimately utilized top-layer deep learning features, we propose to compactly represent and convey the intermediate-layer deep learning features with high generalization capability, to facilitate the collaborating approach between front and cloud ends. This strategy enables a good balance among the computational load, transmission load and the generalization ability for cloud servers when deploying the deep neural networks for large scale cloud based visual analysis. Moreover, the presented strategy also makes the standardization of deep feature coding more feasible and promising, as a series of tasks can simultaneously benefit from the transmitted intermediate layer features. We also present the results for evaluations of both lossless and lossy deep feature compression, which provide meaningful investigations and baselines for future research and standardization activities. Zhuo Chen 0006, Kui Fan, Shiqi Wang 0001, Ling-Yu Duan, Weisi Lin, Alex Chichung Kot |
IEEE Trans. Image Process. | 6 |
| 2020 | Heterogeneous Domain Adaptation via Nonlinear Matrix FactorizationabstractHeterogeneous domain adaptation (HDA) aims to solve the learning problems where the source- and the target-domain data are represented by heterogeneous types of features. The existing HDA approaches based on matrix completion or matrix factorization have proven to be effective to capture shareable information between heterogeneous domains. However, there are two limitations in the existing methods. First, a large number of corresponding data instances between the source domain and the target domain are required to bridge the gap between different domains for performing matrix completion. These corresponding data instances may be difficult to collect in real-world applications due to the limited size of data in the target domain. Second, most existing methods can only capture linear correlations between features and data instances while performing matrix completion for HDA. In this paper, we address these two issues by proposing a new matrix-factorization-based HDA method in a semisupervised manner, where only a few labeled data are required in the target domain without requiring any corresponding data instances between domains. Such labeled data are more practical to obtain compared with cross-domain corresponding data instances. Our proposed algorithm is based on matrix factorization in an approximated reproducing kernel Hilbert space (RKHS), where nonlinear correlations between features and data instances can be exploited to learn heterogeneous features for both the source and the target domains. Extensive experiments are conducted on cross-domain text classification and object recognition, and experimental results demonstrate the superiority of our proposed method compared with the state-of-the-art HDA approaches. Haoliang Li, Sinno Jialin Pan, Shiqi Wang 0001, Alex Chichung Kot |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2019 | Heterogeneous Transfer Learning via Deep Matrix Completion with Adversarial Kernel EmbeddingabstractHeterogeneous Transfer Learning (HTL) aims to solve transfer learning problems where a source domain and a target domain are of heterogeneous types of features. Most existing HTL approaches either explicitly learn feature mappings between the heterogeneous domains or implicitly reconstruct heterogeneous cross-domain features based on matrix completion techniques. In this paper, we propose a new HTL method based on a deep matrix completion framework, where kernel embedding of distributions is trained in an adversarial manner for learning heterogeneous features across domains. We conduct extensive experiments on two different vision tasks to demonstrate the effectiveness of our proposed method compared with a number of baseline methods. Haoliang Li, Sinno Jialin Pan, Renjie Wan, Alex Chichung Kot |
AAAI | 4 |
| 2019 | Detection of Real-world Fights in Surveillance VideosabstractCCTVs have since long been used to enforce security, e.g. to detect fights arising from many different situations. But their effectiveness is questionable, because they rely on continuous and specialized human supervision, demanding automated solutions. Previous work are either too superficial (classification of short-clips) or unrealistic (movies, sports, fake fights). None performed detection of actual fights on long duration CCTV recordings. In this work, we tackle this problem by firstly proposing CCTV-Fights1, a novel and challenging dataset containing 1,000 videos of real fights, with more than 8 hours of annotated CCTV footage. Then we propose a pipeline, on which we assess the impact of different feature extractors, through Two-stream CNN, 3D CNN and a local interest point descriptor, as well as different classifiers, such as end-to-end CNN, LSTM and SVM. Results confirm how challenging the problem is, and highlight the importance of explicit motion information to improve performance. Mauricio Perez, Alex Chichung Kot, Anderson Rocha 0001 |
ICASSP | 2 |
| 2019 | Learning to Jointly Generate and Separate ReflectionsabstractExisting learning-based single image reflection removal methods using paired training data have fundamental limitations about the generalization capability on real-world reflections due to the limited variations in training pairs. In this work, we propose to jointly generate and separate reflections within a weakly-supervised learning framework, aiming to model the reflection image formation more comprehensively with abundant unpaired supervision. By imposing the adversarial losses and combinable mapping mechanism in a multi-task structure, the proposed framework elegantly integrates the two separate stages of reflection generation and separation into a unified model. The gradient constraint is incorporated into the concurrent training process of the multi-task learning as well. In particular, we built up an unpaired reflection dataset with 4,027 images, which is useful for facilitating the weakly-supervised learning of reflection removal model. Extensive experiments on a public benchmark dataset show that our framework performs favorably against state-of-the-art methods and consistently produces visually appealing results. Daiqian Ma, Renjie Wan, Boxin Shi, Alex Chichung Kot, Ling-Yu Duan |
ICCV | 4 |
| 2019 | SPLINE-Net: Sparse Photometric Stereo Through Lighting Interpolation and Normal Estimation NetworksabstractThis paper solves the Sparse Photometric stereo through Lighting Interpolation and Normal Estimation using a generative Network (SPLINE-Net). SPLINE-Net contains a lighting interpolation network to generate dense lighting observations given a sparse set of lights as inputs followed by a normal estimation network to estimate surface normals. Both networks are jointly constrained by the proposed symmetric and asymmetric loss functions to enforce isotropic constrain and perform outlier rejection of global illumination effects. SPLINE-Net is verified to outperform existing methods for photometric stereo of general BRDFs by using only ten images of different lights instead of using nearly one hundred images. Yiming Jia, Boxin Shi, Xudong Jiang 0001, Ling-Yu Duan, Alex Chichung Kot |
ICCV | 6 |
| 2019 | Attention to Head Locations for Crowd Counting
Youmei Zhang, Chunluan Zhou, Faliang Chang, Alex Chichung Kot, Wei Zhang 0021 |
ICIG (2) | 4 |
| 2019 | Fashion Recommendation on Street ImagesabstractLearning the compatibility relationship is of vital importance to a fashion recommendation system, while existing works achieve this merely on product images but not on street images in the complex daily life scenario. In this paper, we propose a novel fashion recommendation system: Given a query item of interest in the street scenario, the system can return the compatible items. More specifically, a two-stage curriculum learning scheme is developed to transfer the semantics from the product to street outfit images. We also propose a domain-specific missing item imputation method based on style and color similarity to handle the incomplete outfits. To support the training of deep recommendation model, we collect a large dataset with street outfit images. The experiments on the dataset demonstrate the advantages of the proposed method over the state-of-the-art approaches on both the street images and the product images. Huijing Zhan, Boxin Shi, Ling-Yu Duan, Alex Chichung Kot |
ICIP | 6 |
| 2019 | Denoising Adversarial Networks for Rain Removal and Reflection RemovalabstractThis paper presents a novel adversarial scheme to perform image denoising for the tasks of rain streak removal and reflection removal. Similar to several previous works, the proposed method first estimates a prior image and then uses it to guide the inference of noise-free image. The novelty of our approach is to jointly learn the gradient and noise-free image based on an adversarial scheme. More specifically, we use the gradient map as the prior image. The inferred noise-free image guided by an estimated gradient is regarded as a negative sample, while the noise-free image guided by the ground truth of a gradient is taken as a positive sample. With the anchor defined by the ground truth of noise-free image, we play a min-max game to jointly train two optimizers for the estimation of the gradient and the inference of noise-free images. We show that both prior image and noise-free image can be accurately obtained under this adversarial scheme. Our state-of-the-art performance achieved on two public benchmark datasets validate the effectiveness of our approach. Boxin Shi, Xudong Jiang 0001, Ling-Yu Duan, Alex Chichung Kot |
ICIP | 5 |
| 2019 | Few-Shot and Many-Shot Fusion Learning in Mobile Visual Food RecognitionabstractMobile visual food recognition is emerging as an important application in food logging and dietary monitoring in recent years. Existing food recognition methods use conventional many-shot learning to train a large backbone network, which refers to the use of sufficient number of training data to train the network. However, these methods firstly do not consider the cases where certain food categories have limited training data. Therefore, they cannot use the conventional training using many-shot learning. Further, existing solutions focus on improving the food recognition performance by implementing state-of-the-art large full networks, and do not pay much attention to reduce the size and computational cost of the network. As a result, they are not amenable for deployment on mobile devices. In this paper, we address these issues by proposing a new few-shot and many-shot fusion learning for mobile visual food recognition, it has a compact framework and is able to learn from existing dataset categories, and also new food categories given only a few sample images. We construct a new Indian food dataset called NTU-IndianFood107 in order to evaluate the performance of the proposed method. The dataset has two parts: (i) a Base Dataset of 83 classes of Indian food images with over 600 images per class to perform many-shot learning, and (ii) a Food Diary of 24 classes captured in restaurants with limited number to simulate the few-shot learning on new food categories. The proposed fusion method achieves a Top-1 classification accuracy of 72.0% on the new dataset. Kim-Hui Yap, Alex Chichung Kot, Ling-Yu Duan, Ngai-Man Cheung |
ISCAS | 3 |
| 2019 | Lossy Intermediate Deep Learning Feature Compression and EvaluationabstractWith the unprecedented success of deep learning in computer vision tasks, many cloud-based visual analysis applications are powered by deep learning models. However, the deep learning models are also characterized with high computational complexity and are task-specific, which may hinder the large-scale implementation of the conventional data communication paradigms. To enable a better balance among bandwidth usage, computational load and the generalization capability for cloud-end servers, we propose to compress and transmit intermediate deep learning features instead of visual signals and ultimately utilized features. The proposed strategy also provides a promising way for the standardization of deep feature coding. As the first attempt to this problem, we present a lossy compression framework and evaluation metrics for intermediate deep feature compression. Comprehensive experimental results show the effectiveness of our proposed methods and the feasibility of the proposed data transmission strategy. It is worth mentioning that the proposed compression framework and evaluation metrics have been adopted into the ongoing AVS (Audio Video Coding Standard Workgroup) - Visual Feature Coding Standard. Zhuo Chen 0006, Kui Fan, Shiqi Wang 0001, Ling-Yu Duan, Weisi Lin, Alex Chichung Kot |
ACM Multimedia | 6 |
| 2019 | Semantic Segmentation via Domain Adaptation with Global Structure EmbeddingabstractIn this paper we focus on the problem of unsupervised domain adaptation for semantic segmentation. The previous works usually focus on adversarial learning either in pixel-level or feature-level. However, global structure knowledge is often neglected in the adversarial learning due to the possible reasons: First, the result of pixel-level adversarial learning does not necessarily preserve the semantic consistency of the input image. Second, global structure knowledge is not embedded to regularize the feature-level adversarial learning. In this work, we propose a framework for unsupervised domain adaptation in semantic segmentation which effectively incorporates pixel- level, feature-level adversarial learning and self-training strategy. Our framework embeds the global structure knowledge into the adversarial training step to tackle the problem of structure misalignment. Consequently, our proposed framework achieves the state-of-the-art semantic segmentation domain adaptation results on the task of transferring GTA5 to Cityscapes. Tianyi Zhang 0004, Guosheng Lin, Jianfei Cai 0001, Alex Chichung Kot |
VCIP | 4 |
| 2019 | Task-in-all Domain Adaptation for Semantic SegmentationabstractIn this work we tackle the problem of unsupervised domain adaptation for semantic segmentation. One pipeline is to sequentially train image-translation model and the final task segmentation model. In such pipeline, image translation is aimed to generate the translated source-domain images which are visually similar to the target-domain images and then the final task model is trained using the translated images and its corresponding groundtruth. However, the visually optimal translated-images are not necessarily optimal for the final task of segmenting the target-domain images. Thus we propose a Task-in-all pipeline for unsupervised domain adaptation on semantic segmentation, which incorporates image translation and final segmentation task into an end-to-end training pipeline. Our aim is to generate the translated images which better assists the final task, instead of just being visually similar to the target domain images. We show that in the task of adapting from GTA5 to Cityscapes dataset, the segmentation performance of our Task-in-all pipeline outperforms the sequentially training pipeline, with simpler model structure and less training complexity. Tianyi Zhang 0004, Chuanxia Zheng, Guosheng Lin, Jianfei Cai 0001, Alex Chichung Kot |
VCIP | 6 |
| 2019 | DeepShoe: An improved Multi-Task View-invariant CNN for street-to-shop shoe retrieval
Huijing Zhan, Boxin Shi, Ling-Yu Duan, Alex Chichung Kot |
Comput. Vis. Image Underst. | 4 |
| 2019 | Multi-resolution attention convolutional neural network for crowd countingabstractEstimating crowd counts remains a challenging task due to the problems of scale variations, non-uniform distribution and complex backgrounds. In this paper, we propose a multi-resolution attention convolutional neural network (MRA-CNN) to address this challenging task. Except for the counting task, we exploit an additional density-level classification task during training and combine features learned for the two tasks, thus forming multi-scale, multi-contextual features to cope with the scale variation and non-uniform distribution. Besides, we utilize a multi-resolution attention (MRA) model to generate score maps, where head locations are with higher scores to guide the network to focus on head regions and suppress non-head regions regardless of the complex backgrounds. During the generation of score maps, atrous convolution layers are used to expand the receptive field with fewer parameters, thus getting higher-level features and providing the MRA model more comprehensive information. Experiments on ShanghaiTech, WorldExpo’10 and UCF datasets demonstrate the effectiveness of our method. Youmei Zhang, Chunluan Zhou, Faliang Chang, Alex Chichung Kot |
Neurocomputing | 4 |
| 2019 | A scale adaptive network for crowd counting
Youmei Zhang, Chunluan Zhou, Faliang Chang, Alex Chichung Kot |
Neurocomputing | 4 |
| 2019 | Decoupled Spatial Neural Attention for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation receives much research attention since it alleviates the need to obtain a large amount of dense pixel-wise ground-truth annotations for the training images. Compared with other forms of weak supervision, image labels are quite efficient to obtain. In this paper, we focus on the weakly supervised semantic segmentation with image label annotations. Recent progress for this task has been largely dependent on the quality of generated pseudo-annotations. In this paper, inspired by spatial neural-attention for image captioning, we propose a decoupled spatial neural attention network for generating pseudo-annotations. Our decoupled attention structure could simultaneously identify the object regions and localize the discriminative parts, which generates high-quality pseudo-annotations in one forward path. The generated pseudo-annotations lead to the segmentation results that achieve the state of the art in weakly supervised semantic segmentation. Tianyi Zhang 0004, Guosheng Lin, Jianfei Cai 0001, Chunhua Shen, Alex Chichung Kot |
IEEE Trans. Multim. | 6 |
| 2018 | Multi-task Mid-level Feature Alignment Network for Unsupervised Cross-Dataset Person Re-Identification
Haoliang Li, Chang-Tsun Li, Alex Chichung Kot |
BMVC | 4 |
| 2018 | SSNet: Scale Selection Network for Online 3D Action PredictionabstractIn action prediction (early action recognition), the goal is to predict the class label of an ongoing action using its observed part so far. In this paper, we focus on online action prediction in streaming 3D skeleton sequences. A dilated convolutional network is introduced to model the motion dynamics in temporal dimension via a sliding window over the time axis. As there are significant temporal scale variations of the observed part of the ongoing action at different progress levels, we propose a novel window scale selection scheme to make our network focus on the performed part of the ongoing action and try to suppress the noise from the previous actions at each time step. Furthermore, an activation sharing scheme is proposed to deal with the overlapping computations among the adjacent steps, which allows our model to run more efficiently. The extensive experiments on two challenging datasets show the effectiveness of the proposed action prediction framework. Jun Liu 0036, Amir Shahroudy, Gang Wang 0012, Ling-Yu Duan, Alex Chichung Kot |
CVPR | 5 |
| 2018 | Domain Generalization With Adversarial Feature LearningabstractIn this paper, we tackle the problem of domain generalization: how to learn a generalized feature representation for an "unseen" target domain by taking the advantage of multiple seen source-domain data. We present a novel framework based on adversarial autoencoders to learn a generalized latent feature representation across domains for domain generalization. To be specific, we extend adversarial autoencoders by imposing the Maximum Mean Discrepancy (MMD) measure to align the distributions among different domains, and matching the aligned distribution to an arbitrary prior distribution via adversarial feature learning. In this way, the learned feature representation is supposed to be universal to the seen source domains because of the MMD regularization, and is expected to generalize well on the target domain because of the introduction of the prior distribution. We proposed an algorithm to jointly train different components of our proposed framework. Extensive experiments on various vision tasks demonstrate that our proposed framework can learn better generalized features for the unseen target domain compared with state-of-the-art domain generalization methods. Haoliang Li, Sinno Jialin Pan, Shiqi Wang 0001, Alex Chichung Kot |
CVPR | 4 |
| 2018 | Dual Attention Matching Network for Context-Aware Feature Sequence Based Person Re-IdentificationabstractTypical person re-identification (ReID) methods usually describe each pedestrian with a single feature vector and match them in a task-specific metric space. However, the methods based on a single feature vector are not sufficient enough to overcome visual ambiguity, which frequently occurs in real scenario. In this paper, we propose a novel end-to-end trainable framework, called Dual ATtention Matching network (DuATM), to learn context-aware feature sequences and perform attentive sequence comparison simultaneously. The core component of our DuATM framework is a dual attention mechanism, in which both intrasequence and inter-sequence attention strategies are used for feature refinement and feature-pair alignment, respectively. Thus, detailed visual cues contained in the intermediate feature sequences can be automatically exploited and properly compared. We train the proposed DuATM network as a siamese network via a triplet loss assisted with a decorrelation loss and a cross-entropy loss. We conduct extensive experiments on both image and video based ReID benchmark datasets. Experimental results demonstrate the significant advantages of our approach compared to the state-of-the-art methods. Jianlou Si, Honggang Zhang 0002, Chun-Guang Li, Jason Kuen, Xiangfei Kong, Alex Chichung Kot, Gang Wang 0012 |
CVPR | 6 |
| 2018 | CRRN: Multi-Scale Guided Concurrent Reflection Removal NetworkabstractRemoving the undesired reflections from images taken through the glass is of broad application to various computer vision tasks. Non-learning based methods utilize different handcrafted priors such as the separable sparse gradients caused by different levels of blurs, which often fail due to their limited description capability to the properties of real-world reflections. In this paper, we propose the Concurrent Reflection Removal Network (CRRN) to tackle this problem in a unified framework. Our proposed network integrates image appearance information and multi-scale gradient information with human perception inspired loss function, and is trained on a new dataset with 3250 reflection images taken under diverse real-world scenes. Extensive experiments on a public benchmark dataset show that the proposed method performs favorably against state-of-the-art methods. Renjie Wan, Boxin Shi, Ling-Yu Duan, Ah-Hwee Tan, Alex Chichung Kot |
CVPR | 5 |
| 2018 | Abandoned Object Detection Using Pixel-Based Finite State Machine and Single Shot Multibox DetectorabstractThis paper proposes a robust, scalable framework for automatic detection of abandoned, stationary objects in real time surveillance videos that can pose a security threat. We use the sViBe background modeling method to generate a long-term and a short-term background model to extract foreground objects. Subsequently, a pixel-based FSM detects stationary candidate objects based on the temporal transition of code patterns. In order to classify the stationary candidate objects, we use deep learning method (SSD: Single Shot MultiBox Detector) to detect person and some suspected type of objects which include backpack, handbag. In order to suppress any false alarm, we remove other stationary candidate objects other than the suspected stationary objects. After stationary object detection, we also check if there is no person near by the suspected detected objects for a particular time. We tested the system on four standard public datasets. The results show that our method outperforms the performance of existing results while also being robust to temporary occlusions and illumination changes. Devadeep Shyam, Alex Chichung Kot, Chinmayee Athalye |
ICME | 2 |
| 2018 | Skeleton-Based Action Recognition Using Spatio-Temporal LSTM Network with Trust GatesabstractSkeleton-based human action recognition has attracted a lot of research attention during the past few years. Recent works attempted to utilize recurrent neural networks to model the temporal dependencies between the 3D positional configurations of human body joints for better analysis of human activities in the skeletal data. The proposed work extends this idea to spatial domain as well as temporal domain to better analyze the hidden sources of action-related information within the human skeleton sequences in both of these domains simultaneously. Based on the pictorial structure of Kinect's skeletal data, an effective tree-structure based traversal framework is also proposed. In order to deal with the noise in the skeletal data, a new gating mechanism within LSTM module is introduced, with which the network can learn the reliability of the sequential data and accordingly adjust the effect of the input data on the updating procedure of the long-term context representation stored in the unit's memory cell. Moreover, we introduce a novel multi-modal feature fusion strategy within the LSTM unit in this paper. The comprehensive experimental results on seven challenging benchmark datasets for human action recognition demonstrate the effectiveness of the proposed method. Jun Liu 0036, Amir Shahroudy, Dong Xu 0001, Alex Chichung Kot, Gang Wang 0012 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2018 | Color space construction by optimizing luminance and chrominance components for face recognition
Ze Lu, Xudong Jiang 0001, Alex Chichung Kot |
Pattern Recognit. | 3 |
| 2018 | Feature fusion with covariance matrix regularization in face recognitionabstractThe fusion of multiple features is important for achieving state-of-the-art face recognition results. This has been proven in both traditional and deep learning approaches. Existing feature fusion methods either reduce the dimensionality of each feature first and then concatenate all low-dimensional feature vectors, named as DR-Cat, or the vice versa, named as Cat-DR. However, DR-Cat ignores the correlation information between different features which is useful for classification. In Cat-DR, on the other hand, the correlation information estimated from the training data may not be reliable especially when the number of training samples is limited. We propose a covariance matrix regularization (CMR) technique to solve problems of DR-Cat and Cat-DR. It works by assigning weights to cross-feature covariances in the covariance matrix of training data . Thus the feature correlation estimated from training data is regularized before being used to train the feature fusion model. The proposed CMR is applied to 4 feature fusion schemes: fusion of pixel values from 3 color channels, fusion of LBP features from 3 color channels, fusion of pixel values and LBP features from a single color channel, and fusion of CNN features extracted by 2 deep models. Extensive experiments of face recognition and verification are conducted on databases including MultiPIE, Georgia Tech, AR and LFW. Results show that the proposed CMR technique significantly and consistently outperforms the best single feature , DR-Cat and Cat-DR. Ze Lu, Xudong Jiang 0001, Alex Chichung Kot |
Signal Process. | 3 |
| 2018 | Deep Coupled ResNet for Low-Resolution Face RecognitionabstractFace images captured by surveillance cameras are often of low resolution (LR), which adversely affects the performance of their matching with high-resolution (HR) gallery images. Existing methods including super resolution, coupled mappings (CMs), multidimensional scaling, and convolutional neural network yield only modest performance. In this letter, we propose the deep coupled ResNet (DCR) model. It consists of one trunk network and two branch networks. The trunk network, trained by face images of three significantly different resolutions, is used to extract discriminative features robust to the resolution change. Two branch networks, trained by HR images and images of the targeted LR, work as resolution-specific CMs to transform HR and corresponding LR features to a space where their difference is minimized. Model parameters of branch networks are optimized using our proposed CM loss function, which considers not only the discriminability of HR and LR features, but also the similarity between them. In order to deal with various possible resolutions of probe images, we train multiple pairs of small branch networks while using the same trunk network. Thorough evaluation on LFW and SCface databases shows that the proposed DCR model achieves consistently and considerably better performance than the state of the arts. Ze Lu, Xudong Jiang 0001, Alex Chichung Kot |
IEEE Signal Process. Lett. | 3 |
| 2018 | Learning Generalized Deep Feature Representation for Face Anti-SpoofingabstractIn this paper, we propose a novel framework leveraging the advantages of the representational ability of deep learning and domain generalization for face spoofing detection. In particular, the generalized deep feature representation is achieved by taking both spatial and temporal information into consideration, and a 3D convolutional neural network architecture tailored for the spatial-temporal input is proposed. The network is first initialized by training with augmented facial samples based on cross-entropy loss and further enhanced with a specifically designed generalization loss, which coherently serves as the regularization term. The training samples from different domains can seamlessly work together for learning the generalized feature representation by manipulating their feature distribution distances. We evaluate the proposed framework with different experimental setups using various databases. Experimental results indicate that our method can learn more discriminative and generalized information compared with the state-of-the-art methods. Haoliang Li, Peisong He, Shiqi Wang 0001, Anderson Rocha 0001, Xinghao Jiang, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2018 | Unsupervised Domain Adaptation for Face Anti-SpoofingabstractFace anti-spoofing (a.k.a. presentation attack detection) has recently emerged as an active topic with great significance for both academia and industry due to the rapidly increasing demand in user authentication on mobile phones, PCs, tablets, and so on. Recently, numerous face spoofing detection schemes have been proposed based on the assumption that training and testing samples are in the same domain in terms of the feature space and marginal probability distribution. However, due to unlimited variations of the dominant conditions (illumination, facial appearance, camera quality, and so on) in face acquisition, such single domain methods lack generalization capability, which further prevents them from being applied in practical applications. In light of this, we introduce an unsupervised domain adaptation face anti-spoofing scheme to address the real-world scenario that learns the classifier for the target domain based on training samples in a different source domain. In particular, an embedding function is first imposed based on source and target domain data, which maps the data to a new space where the distribution similarity can be measured. Subsequently, the Maximum Mean Discrepancy between the latent features in source and target domains is minimized such that a more generalized classifier can be learned. State-of-the-art representations including both hand-crafted and deep neural network learned features are further adopted into the framework to quest the capability of them in domain adaptation. Moreover, we introduce a new database for face spoofing detection, which contains more than 4000 face samples with a large variety of spoofing types, capture devices, illuminations, and so on. Extensive experiments on existing benchmark databases and the new database verify that the proposed approach can gain significantly better generalization capability in cross-domain scenarios by providing consistently better anti-spoofing performance. Haoliang Li, Wen Li 0001, Shiqi Wang 0001, Feiyue Huang, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2018 | Fast MPEG-CDVS Encoder With GPU-CPU Hybrid ComputingabstractThe compact descriptors for visual search (CDVS) standard from ISO/IEC moving pictures experts group has succeeded in enabling the interoperability for efficient and effective image retrieval by standardizing the bitstream syntax of compact feature descriptors. However, the intensive computation of a CDVS encoder unfortunately hinders its widely deployment in industry for large-scale visual search. In this paper, we revisit the merits of low complexity design of CDVS core techniques and present a very fast CDVS encoder by leveraging the massive parallel execution resources of graphics processing unit (GPU). We elegantly shift the computation-intensive and parallel-friendly modules to the state-of-the-arts GPU platforms, in which the thread block allocation as well as the memory access mechanism are jointly optimized to eliminate performance loss. In addition, those operations with heavy data dependence are allocated to CPU for resolving the extra but non-necessary computation burden for GPU. Furthermore, we have demonstrated the proposed fast CDVS encoder can work well with those convolution neural network approaches which enables to leverage the advantages of GPU platforms harmoniously, and yield significant performance improvements. Comprehensive experimental results over benchmarks are evaluated, which has shown that the fast CDVS encoder using GPU-CPU hybrid computing is promising for scalable visual search. Ling-Yu Duan, Wei Sun 0029, Xinfeng Zhang 0001, Shiqi Wang 0001, Jie Chen 0006, Jianxiong Yin, Simon See, Tiejun Huang 0001, Alex Chichung Kot, Wen Gao 0001 |
IEEE Trans. Image Process. | 9 |
| 2018 | Skeleton-Based Human Action Recognition With Global Context-Aware Attention LSTM NetworksabstractHuman action recognition in 3D skeleton sequences has attracted a lot of research attention. Recently, long short-term memory (LSTM) networks have shown promising performance in this task due to their strengths in modeling the dependencies and dynamics in sequential data. As not all skeletal joints are informative for action recognition, and the irrelevant joints often bring noise which can degrade the performance, we need to pay more attention to the informative ones. However, the original LSTM network does not have explicit attention ability. In this paper, we propose a new class of LSTM network, global context-aware attention LSTM, for skeleton-based action recognition, which is capable of selectively focusing on the informative joints in each frame by using a global context memory cell. To further improve the attention capability, we also introduce a recurrent attention mechanism, with which the attention performance of our network can be enhanced progressively. Besides, a two-stream framework, which leverages coarse-grained attention and fine-grained attention, is also introduced. The proposed method achieves state-of-the-art performance on five challenging datasets for skeleton-based action recognition. Jun Liu 0036, Gang Wang 0012, Ling-Yu Duan, Kamila Abdiyeva, Alex Chichung Kot |
IEEE Trans. Image Process. | 5 |
| 2018 | Region-Aware Reflection Removal With Unified Content and Gradient PriorsabstractRemoving the undesired reflections in images taken through the glass is of broad application to various image processing and computer vision tasks. Existing single image based solutions heavily rely on scene priors such as separable sparse gradients caused by different levels of blur, and they are fragile when such priors are not observed. In this paper, we notice that strong reflections usually dominant a limited region in the whole image, and propose a Region-aware Reflection Removal (R3) approach by automatically detecting and heterogeneously processing regions with and without reflections. We integrate content and gradient priors to jointly achieve missing contents restoration as well as background and reflection separation in a unified optimization framework. Extensive validation using 50 sets of real data shows that the proposed method outperforms state-of-the-art on both quantitative metrics and visual qualities. Renjie Wan, Boxin Shi, Ling-Yu Duan, Ah-Hwee Tan, Wen Gao 0001, Alex Chichung Kot |
IEEE Trans. Image Process. | 6 |
| 2017 | DeepShoe: A Multi-Task View-Invariant CNN for Street-to-Shop Shoe Retrieval
Huijing Zhan, Boxin Shi, Alex Chichung Kot |
BMVC | 3 |
| 2017 | Global Context-Aware Attention LSTM Networks for 3D Action RecognitionabstractLong Short-Term Memory (LSTM) networks have shown superior performance in 3D human action recognition due to their power in modeling the dynamics and dependencies in sequential data. Since not all joints are informative for action analysis and the irrelevant joints often bring a lot of noise, we need to pay more attention to the informative ones. However, original LSTM does not have strong attention capability. Hence we propose a new class of LSTM network, Global Context-Aware Attention LSTM (GCA-LSTM), for 3D action recognition, which is able to selectively focus on the informative joints in the action sequence with the assistance of global contextual information. In order to achieve a reliable attention representation for the action sequence, we further propose a recurrent attention mechanism for our GCA-LSTM network, in which the attention performance is improved iteratively. Experiments show that our end-to-end network can reliably focus on the most informative joints in each frame of the skeleton sequence. Moreover, our network yields state-of-the-art performance on three challenging datasets for 3D action recognition. Jun Liu 0036, Gang Wang 0012, Ping Hu 0001, Ling-Yu Duan, Alex Chichung Kot |
CVPR | 5 |
| 2017 | Compact Deep Invariant Descriptors for Video RetrievalabstractWith emerging demand for large-scale video analysis, the Motion Picture Experts Group (MPEG) initiated the Compact Descriptor for Video Analysis (CDVA) standardization in 2014. In this work, we develop novel deep-learning features and incorporate them into the well-established CDVA evaluation framework to study its effectiveness in video analysis. In particular, we propose a Nested Invariance Pooling (NIP) method to obtain compact and robust Convolutional Neural Network (CNNs) descriptors. The CNNs descriptors are generated by applying three different pooling operations to the feature maps of CNNs in a nested way towards rotation and scale invariant feature representation. In particular, the rational, advantages and performance on the combination of CNNs and handcrafted descriptors are provided to better investigate the complementary effects of deep learnt and handcrafted features. Extensive experimental results show that the proposed CNNs descriptors outperform both state-of-the-art CNNs descriptors and canonical handcrafted descriptors adopted in CDVA Experimental Model (CXM) with significant mAP gains of 11.3% and 4.7%, respectively. Moreover, the combination of NIP derived deep invariant descriptors and handcrafted descriptors not only fulfills the lowest bitrate budget of CDVA, but also significantly advances the performance of CDVA core techniques. Yihang Lou, Jie Lin 0001, Shiqi Wang 0001, Jie Chen 0006, Vijay Chandrasekhar 0001, Ling-Yu Duan, Tiejun Huang 0001, Alex Chichung Kot, Wen Gao 0001 |
DCC | 9 |
| 2017 | A novel LBP-based Color descriptor for face recognitionabstractLBP-based color features have shown excellent performance for color face recognition tasks, such as Color LBP, CLBP and LCVBP. However, existing methods encode the inter-channel information on pairs of color channels by applying the same spatial structure as that used in the intra-channel encoding. This results in a very high dimensional feature vector yet ineffective in encoding inter-channel information. Moreover, the difference of pixel values across color channels may not be a proper measure if they are not quantitatively comparable. To tackle these problems of existing methods, we propose a novel LBP-based color feature, Ternary-Color LBP (TCLBP), to encode the inter-channel information more effectively and efficiently. Extensive experiments on 4 public face databases, Color FERET, Georgia Tech, FRGC and LFW, are conducted to verify the effectiveness of the proposed TCLBP color feature for face recognition. Results show that the proposed TCLBP leads to visibly better face recognition performance than Color LBP, CLBP and LCVBP consistently over the 4 databases. Ze Lu, Xudong Jiang 0001, Alex Chichung Kot |
ICASSP | 3 |
| 2017 | Benchmarking Single-Image Reflection Removal AlgorithmsabstractRemoving undesired reflections from a photo taken in front of a glass is of great importance for enhancing the efficiency of visual computing systems. Various approaches have been proposed and shown to be visually plausible on small datasets collected by their authors. A quantitative comparison of existing approaches using the same dataset has never been conducted due to the lack of suitable benchmark data with ground truth. This paper presents the first captured Single-image Reflection Removal dataset ‘SIR2’ with 40 controlled and 100 wild scenes, ground truth of background and reflection. For each controlled scene, we further provide ten sets of images under varying aperture settings and glass thicknesses. We perform quantitative and visual quality comparisons for four state-of-the-art single-image reflection removal algorithms using four error metrics. Open problems for improving reflection removal algorithms are discussed at the end. Renjie Wan, Boxin Shi, Ling-Yu Duan, Ah-Hwee Tan, Alex Chichung Kot |
ICCV | 5 |
| 2017 | Deep regional feature pooling for video matchingabstractIn this work, we study the problem of deep global descriptors for video matching with regional feature pooling. We aim to analyze the joint effect of ROI (Region of Interest) size and pooling moment on video matching performance. To this end, we propose to mathematically model the distribution of video matching function with a pooling function nested in. Matching performance can be estimated by the separability of these class-conditional distributions between matching and non-matching pairs. Empirical studies on the challenging MPEG CDVA dataset demonstrate that performance trends are consistent with the estimation and experimental results, though the theoretical model is largely simplified compared to video matching and retrieval in practice. Jie Lin 0001, Vijay Chandrasekhar 0001, Yihang Lou, Shiqi Wang 0001, Ling-Yu Duan, Tiejun Huang 0001, Alex Chichung Kot |
ICIP | 8 |
| 2017 | Street-to-shop shoe retrieval with multi-scale viewpoint invariant triplet networkabstractIn this paper we aim to find exactly the same shoes given a daily shoe photo (street scenario) that matches the online shop shoe photo (shop scenario). There are large visual differences between the street and shop scenario shoe images. To handle the discrepancy of different scenarios, we learn a feature embedding for shoes via a viewpoint-invariant triplet network, the feature activations of which reflect the inherent similarity between any two shoe images. Specifically, we propose a new loss function that minimizes the distances between images of the same shoes captured from different viewpoints. Moreover, we train the proposed triplet network at two different scales so that the representation of shoes incorporates different levels of invariance at different scales. To support training the multi-scale triplet networks, we collect a large dataset with shoe images from the daily life and online shopping websites. Experiments on the dataset show excellence over state-of-the-art approaches, which demonstrate the effectiveness of our proposed method. Huijing Zhan, Boxin Shi, Alex Chichung Kot |
ICIP | 3 |
| 2017 | Sparsity based reflection removal using external patch searchabstractReflection removal aims at separating the mixture of the desired background scenes and the undesired reflections, when the photos are taken through the glass. It has both aesthetic and practical applications which can largely improve the performance of many multimedia tasks. Existing reflection removal approaches heavily rely on scene priors such as separable sparse gradients brought by different levels of blur, and they easily fail when such priors are not observed in many real scenes. Sparse representation models and nonlocal image priors have shown their effectiveness in image restoration with self similarity. In this work, we propose a reflection removal method benefited from the sparsity and nonlocal image prior as a unified optimization framework. We leverage the retrieved image patch from an external database to overcome the limited prior information in the input mixture image and self similarity search. The experimental results show that our proposed model performs better than the existing state-of-the-art reflection removal method for both objective and subjective image qualities. Renjie Wan, Boxin Shi, Ah-Hwee Tan, Alex Chichung Kot |
ICME | 4 |
| 2017 | Fashion analysis with a subordinate attribute classification networkabstractIn this paper we deal with two image-based object search tasks in the fashion domain, clothing attribute prediction and cross-domain shoe retrieval. Clothing attribute prediction is about describing the appearances of clothes via semantic attributes and cross-domain shoe retrieval aims at retrieving the same shoe items from online stores given a daily life shoe photo. We jointly solve these two problems by a novel Subordinate Attribute Convolutional Neural Network (SA-CNN), with the newly designed loss function that systematically merges semantic attributes of closer visual appearance to prevent images with obvious visual differences being confused with each other. A three-level feature representation is further developed based on SA-CNN for shoes from different domains. The experimental results demonstrate that the clothing attribute prediction using the proposed SA-CNN achieves better performance than that using traditional features and fine-tuned conventional CNN. Moreover, for the task of cross-domain shoe retrieval, the top-20 retrieval accuracy with deep features extracted from SA-CNN has a significant improvement of 43% compared to that with the pretrained CNN features. Huijing Zhan, Boxin Shi, Alex Chichung Kot |
ICME | 3 |
| 2017 | Cross-domain shoe retrieval using a three-level deep feature representationabstractIn this paper, we address the problem of matching the shoes from the daily life photos to exactly the same shoes from online shops. The problem is extremely challenging because of the significant visual differences between street domain images (shoe images captured in the daily life scenario) and online domain photos (images from online shops taken in the controlled environment). This paper presents a semantic Shoe Attribute-Guided Convolutional Neural Network (SAG-CNN) to extract the deep features. Moreover, we develop a three-level feature representation based on SAG-CNN. The deep features extracted from the image, region and part levels effectively match the images across different domains. We collect a novel shoe dataset, which consists of 8021 street domain and 5821 online domain images. The experimental results on our dataset show that the top-20 retrieval accuracy of our approach improves over that using the pre-trained CNN features by about 40%. Huijing Zhan, Boxin Shi, Alex Chichung Kot |
ISCAS | 3 |
| 2017 | Action proposals using hierarchical clustering of super-trajectoriesabstractAction localization aims to determine the spatial and temporal location of certain action which appears in a video. To facilitate action localization, spatio-temporal proposals which are likely to contain the action of interest are extracted to reduce the search space of candidate locations in a video, inspired by the object proposals in images. In this paper, considering the effectiveness of spatio-temporal trajectories for video action recognition and action proposal generation, we build our unsupervised action proposal generation pipeline upon super-trajectories. Specifically, we first group trajectories into super-trajectories inspired by super-voxels, and then employ hierarchical clustering on super-trajectories by taking different aspect and temporal ratios into consideration. Comprehensive experiments on two benchmark datasets (i.e., UCF-sports and MSR-II) demonstrate that our action proposal generation pipeline not only achieves the state-of-the-art recall, but also achieves competitive results for the action localization task. Tianyi Zhang 0004, Li Niu 0002, Jianfei Cai 0001, Alex Chichung Kot |
VCIP | 4 |
| 2017 | Cross-Domain Shoe Retrieval With a Semantic Hierarchy of Attribute Classification NetworkabstractCross-domain shoe image retrieval is a challenging problem, because the query photo from the street domain (daily life scenario) and the reference photo in the online domain (online shop images) have significant visual differences due to the viewpoint and scale variation, self-occlusion, and cluttered background. This paper proposes the semantic hierarchy of attribute convolutional neural network (SHOE-CNN) with a three-level feature representation for discriminative shoe feature expression and efficient retrieval. The SHOE-CNN with its newly designed loss function systematically merges semantic attributes of closer visual appearances to prevent shoe images with the obvious visual differences being confused with each other; the features extracted from image, region, and part levels effectively match the shoe images across different domains. We collect a large-scale shoe data set composed of 14341 street domain and 12652 corresponding online domain images with fine-grained attributes to train our network and evaluate our system. The top-20 retrieval accuracy improves significantly over the solution with the pre-trained CNN features. Huijing Zhan, Boxin Shi, Alex Chichung Kot |
IEEE Trans. Image Process. | 3 |
| 2017 | HNIP: Compact Deep Invariant Representations for Video Matching, Localization, and RetrievalabstractWith emerging demand for large-scale video analysis, MPEG initiated the compact descriptor for video analysis (CDVA) standardization in 2014. Beyond handcrafted descriptors adopted by the current MPEG-CDVA reference model, we study the problem of deep learned global descriptors for video matching, localization, and retrieval. First, inspired by a recent invariance theory, we propose a nested invariance pooling (NIP) method to derive compact deep global descriptors from convolutional neural networks (CNNs), by progressively encoding translation, scale, and rotation invariances into the pooled descriptors. Second, our empirical studies have shown that a sequence of well designed pooling moments (e.g., max or average) may drastically impact video matching performance, which motivates us to design hybrid pooling operations via NIP (HNIP). HNIP has further improved the discriminability of deep global descriptors. Third, the technical merits and performance improvements by combining deep and handcrafted descriptors are provided to better investigate the complementary effects. We evaluate the effectiveness of HNIP within the well-established MPEG-CDVA evaluation framework. The extensive experiments have demonstrated that HNIP outperforms the state-of-the-art deep and canonical handcrafted descriptors with significant mAP gains of 5.5% and 4.7%, respectively. In particular the combination of HNIP incorporated and handcrafted global descriptors has significantly boosted the performance of CDVA core techniques with comparable descriptor size. Jie Lin 0001, Ling-Yu Duan, Shiqi Wang 0001, Yihang Lou, Vijay Chandrasekhar 0001, Tiejun Huang 0001, Alex Chichung Kot, Wen Gao 0001 |
IEEE Trans. Multim. | 8 |
| 2016 | An effective color space for face recognitionabstractThe three color components specifying a color can be defined in various ways leading to significantly different classification abilities. Several effective color spaces including RQCr, DCS and ZRG have been proposed to achieve better face recognition performance. However, their performance is not consistent on different databases. What's more, the framework of effective color spaces has not been thoroughly studied yet. In this paper, we propose an effective color space LC\C2 based on a framework of effective color spaces. LC\C2 consists of one discriminant luminance component L and two discriminant chrominance components C\C2. To find the discriminant luminance component, 4 luminance components from existing effective color models are compared. After that, the weighted color space normalization technique (WCSN) is applied on the DCS color space to generate two complementary and discriminative chrominance components. Experiments conducted on three databases (FRGC, AR and CMU Multi-PIE) show that the proposed color space LC\C2 achieves the best face recognition performance consistently. Ze Lu, Xudong Jiang 0001, Alex Chichung Kot |
ICASSP | 3 |
| 2016 | Efficient object feature selection for action recognitionabstractCurrently most action recognition or video classification tasks highly rely on the motion features such as state-of-the-art Improved Dense Trajectory (IDT) features. Despite the huge success, IDT features lack of rich static object-level information. In this paper, we make use of the object-level features for action recognition tasks. For efficiently and effectively processing large-scale video data, we propose a two-layer feature selection framework including local object feature selection (LS) and global feature selection (GS). Both of the selection methods can improve recognition accuracy while greatly reducing the feature dimension or feature processing complexity. Experimental results show that the selected object-level features contain complimentary information to IDT features and the combination with IDT features can further improve the recognition accuracy significantly. Tianyi Zhang 0004, Yu Zhang 0004, Jianfei Cai 0001, Alex Chichung Kot |
ICASSP | 4 |
| 2016 | Depth of field guided reflection removalabstractReflection removal aims at separating the mixture of the desired scene and the undesired reflections. Locating reflection and background edges is a key step for reflection removal. In this paper, we present a visual depth guided method to remove reflections. Our idea is to use Depth of Field (DoF) to label the background and reflection edges. We propose a DoF confidence map where pixels with higher DoF values are assumed to belong to the desired background components. Moreover, we observe that images with different resolutions show different properties in the DoF map. Thus, we introduce a multi-scale DoF computing strategy to classify edge pixels more efficiently. Based on the results of edge classification, the background and reflection layers can be separated. Experimental results validate the effectiveness of our method using real-world photos. Renjie Wan, Boxin Shi, Ah-Hwee Tan, Alex Chichung Kot |
ICIP | 4 |
| 2016 | Color space identification from single imagesabstractIn this paper, we focus on the problem of RGB color space identification from a single image. At the moment, RGB color spaces are widely adopted in photography for image producing. The problems with respect to color space identification, such as to get the consistent printing or displaying quality on screen devices and software applications and prevention of multimedia unauthorized usage(shown or printed by other device via gamut mapping), need to be concerned. Current techniques are all relying on EXchangeable Image File Format (EXIF) to extract color space information. In this paper, we use a two-dimensional non-causal regressive model to explore the image demosaicing properties in order to extract discriminative features and train them on SVM classifier for image color space detection without relying on EXIF. In our experiment, images in three different color spaces (sRGB, adobeRGB and pro PhotoRGB) are generated for color space identification task. The experimental results show that the proposed technique has an good performance on image color space identification. Haoliang Li, Alex Chichung Kot, Leida Li |
ISCAS | 2 |
| 2016 | No-Reference Image Blur Assessment Based on Discrete Orthogonal MomentsabstractBlur is a key determinant in the perception of image quality. Generally, blur causes spread of edges, which leads to shape changes in images. Discrete orthogonal moments have been widely studied as effective shape descriptors. Intuitively, blur can be represented using discrete moments since noticeable blur affects the magnitudes of moments of an image. With this consideration, this paper presents a blind image blur evaluation algorithm based on discrete Tchebichef moments. The gradient of a blurred image is first computed to account for the shape, which is more effective for blur representation. Then the gradient image is divided into equal-size blocks and the Tchebichef moments are calculated to characterize image shape. The energy of a block is computed as the sum of squared non-DC moment values. Finally, the proposed image blur score is defined as the variance-normalized moment energy, which is computed with the guidance of a visual saliency model to adapt to the characteristic of human visual system. The performance of the proposed method is evaluated on four public image quality databases. The experimental results demonstrate that our method can produce blur scores highly consistent with subjective evaluations. It also outperforms the state-of-the-art image blur metrics and several general-purpose no-reference quality metrics. Leida Li, Weisi Lin, Xuesong Wang 0001, Gaobo Yang, Khosro Bahrami, Alex Chichung Kot |
IEEE Trans. Cybern. | 6 |
| 2016 | Sparse Representation-Based Image Quality Index With Adaptive Sub-DictionariesabstractDistortions cause structural changes in digital images, leading to degraded visual quality. Dictionary-based sparse representation has been widely studied recently due to its ability to extract inherent image structures. Meantime, it can extract image features with slightly higher level semantics. Intuitively, sparse representation can be used for image quality assessment, because visible distortions can cause significant changes to the sparse features. In this paper, a new sparse representation-based image quality assessment model is proposed based on the construction of adaptive sub-dictionaries. An overcomplete dictionary trained from natural images is employed to capture the structure changes between the reference and distorted images by sparse feature extraction via adaptive sub-dictionary selection. Based on the observation that image sparse features are invariant to weak degradations and the perceived image quality is generally influenced by diverse issues, three auxiliary quality features are added, including gradient, color, and luminance information. The proposed method is not sensitive to training images, so a universal dictionary can be adopted for quality evaluation. Extensive experiments on five public image quality databases demonstrate that the proposed method produces the state-of-the-art results, and it delivers consistently well performances when tested in different image quality databases. Leida Li, Hao Cai 0004, Yabin Zhang 0002, Weisi Lin, Alex Chichung Kot, Xingming Sun |
IEEE Trans. Image Process. | 5 |
| 2016 | Efficient Image Sharpness Assessment Based on Content Aware Total VariationabstractState-of-the-art sharpness assessment methods are mostly based on edge width, gradient, high-frequency energy, or pixel intensity variation. Such methods consider very little the image content variation in conjunction with the sharpness assessment which causes the sharpness metric to be less effective for different content images. In this paper, we propose an efficient no-reference image sharpness assessment called content aware total variation (CATV) by considering the importance of image content variation in sharpness measurement. By parameterizing the image TV statistics using generalized Gaussian distribution, the sharpness measure is identified by the standard deviation, and the image content variation evaluator is indicated by the shape parameter. However, the standard deviation is content-dependent which is different for the regions with strong edges, high frequency textures, low frequency textures, and blank areas. By incorporating the shape-parameter in moderating of the standard deviation, we propose a content aware sharpness metric. The experimental results show that the proposed method is highly correlated with the human vision system and has better sharpness assessment results than the state-of-the-art techniques on the blurred subset images of LIVE, TID2008, CSIQ, and IVC databases. Also, our method has very low computational complexity which is suitable for online applications. The correlations with the subjective of the four databases and statistical significance analysis reveal that our method has superior results when compared with previous techniques. Khosro Bahrami, Alex Chichung Kot |
IEEE Trans. Multim. | 2 |
| 2016 | Image Sharpness Assessment by Sparse RepresentationabstractRecent advances in sparse representation show that overcomplete dictionaries learned from natural images can capture high-level features for image analysis. Since atoms in the dictionaries are typically edge patterns and image blur is characterized by the spread of edges, an overcomplete dictionary can be used to measure the extent of blur. Motivated by this, this paper presents a no-reference sparse representation-based image sharpness index. An overcomplete dictionary is first learned using natural images. The blurred image is then represented using the dictionary in a block manner, and block energy is computed using the sparse coefficients. The sharpness score is defined as the variance-normalized energy over a set of selected high-variance blocks, which is achieved by normalizing the total block energy using the sum of block variances. The proposed method is not sensitive to training images, so a universal dictionary can be used to evaluate the sharpness of images. Experiments on six public image quality databases demonstrate the advantages of the proposed method. Leida Li, Jinjian Wu, Haoliang Li, Weisi Lin, Alex Chichung Kot |
IEEE Trans. Multim. | 6 |
| 2016 | On Branded Handbag RecognitionabstractManufacturing branded handbags is a big business in the fashion world. Shoppers’ feedback showing photos of their purchased handbags in social networks or blogs is important for branding purposes. In this paper, we deal with handbag recognition. It is a challenging problem due to the inter-class style similarity and the intra-class color variation. We focus on developing discriminative representations of handbag style and color. For handbag style representation, two supervised mid-level patch selection procedures are proposed to select discriminative patches, regarding individual classes and pairwise classes. We also propose a low-level complementary feature, extracted from texture-enhanced mid-level patches, to capture the fine details of the mid-level patches. For handbag color representation, we propose to extract dominant color features to handle the illumination changes. The performance of our proposed method is evaluated on a newly built branded handbag dataset. The results show that our method performs favorably in recognizing handbags, with around$10\%$improvement in accuracy when compared with the existing fine-grained or generic object recognition methods. Yan Wang 0033, Sheng Li 0006, Alex Chichung Kot |
IEEE Trans. Multim. | 3 |
| 2015 | Dominant SIFT: A novel compact descriptorabstractDefinition and extraction of local features play a very important role in image retrieval (IR), pattern recognition and computer vision. Fast growth of technology today calls for local features to be as compact as possible toward real-time and limited bandwidth applications. In this paper, we study the problem of representing images in a compact way to achieve low bit-rate transmission while maintaining good performance. To be more specific, we propose a novel compact descriptor, dominant SIFT, which only uses 48 bits to describe local features. Importantly, our descriptor is training-free, vocabulary-free and suitable for real-time and mobile applications. We show the effectiveness of the proposed compact descriptor in image retrieval. Anh T. Tra, Weisi Lin, Alex Chichung Kot |
ICASSP | 3 |
| 2015 | Joint learning for image-based handbag recommendationabstractFashion recommendation helps shoppers to find desirable fashion items, which facilitates online interaction and product promotion. In this paper, we propose a method to recommend handbags to each shopper, based on the handbag images the shopper has clicked. This is performed by Joint learning of attribute Projection and One-class SVM classification (JPO) based on the images of the shopper's preferred handbags. More specifically, for the handbag images clicked by each shopper, we project the original image feature space into an attribute space which is more compact. The projection matrix is learned jointly with a one-class SVM to yield a shopper-specific one-class classifier. The results show that the proposed JPO handbag recommendation performs favorably based on initial subject testing. Yan Wang 0033, Sheng Li 0006, Alex Chichung Kot |
ICME | 3 |
| 2015 | Image splicing localization based on blur type inconsistencyabstractIn a spliced blurred image, the spliced region and the original image may have different blur types. Splicing localization in this image is challenging when a forger uses image resizing as anti-forensics to remove the splicing traces anomalies. In this paper, we overcome this problem by proposing a method for splicing localization based on partial blur type inconsistency. In this method, after the block-based image partitioning, a local blur type detection feature is extracted from the estimated local blur kernels. The image blocks are classified into out-of-focus or motion blur based on this feature to generate invariant blur type regions. Finally a fine splicing localization is applied to increase the precision of regions boundary. We can use the blur type differences of the regions to trace the inconsistency for the splicing localization. Our experimental results show the efficiency of the proposed method in the detection and the classification of the out-of-focus and motion blur types. Khosro Bahrami, Alex Chichung Kot |
ISCAS | 2 |
| 2015 | A Low Complexity Interest Point DetectorabstractInterest point detection is a fundamental approach to feature extraction in computer vision tasks. To handle the scale invariance, interest points usually work on the scale-space representation of an image. In this letter, we propose a novel block-wise scale-space representation to significantly reduce the computational complexity of an interest point detector. Laplacian of Gaussian (LoG) filtering is applied to implement the block-wise scale-space representation. Extensive comparison experiments have shown the block-wise scale-space representation enables the efficient and effective implementation of an interest point detector in terms of memory and time complexity reduction, as well as promising performance in visual search. Jie Chen 0006, Ling-Yu Duan, Feng Gao 0014, Jianfei Cai 0001, Alex Chichung Kot, Tiejun Huang 0001 |
IEEE Signal Process. Lett. | 5 |
| 2015 | A Color Channel Fusion Approach for Face RecognitionabstractDue to high dimensionality of images or generated color features, different color channels are usually processed separately and then concatenated together into a feature vector for classification. This makes channel fusion a crucial step in color face recognition (FR) systems. However, existing methods simply concatenate channel-wise color features without identifying the importance or reliability of features in different color channels. In this paper, we propose a color channel fusion (CCF) approach using jointly dimension reduction algorithms to select more features from reliable and discriminative channels. Experiments using two different dimension reduction approaches, two different types of features on three image datasets show that CCF achieves consistently better performance than color channel concatenation (CCC) method which deals with different color channels equally. Ze Lu, Xudong Jiang 0001, Alex Chichung Kot |
IEEE Signal Process. Lett. | 3 |
| 2015 | Blurred Image Splicing Localization by Exposing Blur Type InconsistencyabstractIn a tampered blurred image generated by splicing, the spliced region and the original image may have different blur types. Splicing localization in this image is a challenging problem when a forger uses some postprocessing operations as antiforensics to remove the splicing traces anomalies by resizing the tampered image or blurring the spliced region boundary. Such operations remove the artifacts that make detection of splicing difficult. In this paper, we overcome this problem by proposing a novel framework for blurred image splicing localization based on the partial blur type inconsistency. In this framework, after the block-based image partitioning, a local blur type detection feature is extracted from the estimated local blur kernels. The image blocks are classified into out-of-focus or motion blur based on this feature to generate invariant blur type regions. Finally, a fine splicing localization is applied to increase the precision of regions boundary. We can use the blur type differences of the regions to trace the inconsistency for the splicing localization. Our experimental results show the efficiency of the proposed method in the detection and the classification of the out-of-focus and motion blur types. For splicing localization, the result demonstrates that our method works well in detecting the inconsistency in the partial blur types of the tampered images. However, our method can be applied to blurred images only. Khosro Bahrami, Alex Chichung Kot, Leida Li, Haoliang Li |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2015 | DeepBag: Recognizing Handbag ModelsabstractIn this paper, we address the problem of branded handbag recognition. It is a challenging problem due to the non-rigid deformation, illumination changes, and inter-class similarity. We propose a novel framework based on deep convolutional neural network (CNN). Concretely, we propose a new CNN model, called feature selective joint classification - regression CNN (FSCR-CNN). Its advantages lie in two folds: 1) it alleviates the illumination changes by a feature selection strategy to focus on the color- nondiscriminative features in the network learning, and 2) rather than only targeting on the hard label (i.e., the handbag model), it also incorporates a soft label (i.e., a distribution measuring the similarity between the ground truth model and all the models to be trained) to construct the loss function for training CNN, which leads to a better classifier for handbags with large inter-class similarity. We evaluate the performance of our framework on a newly built branded handbag dataset. The results show that it performs favorably for recognizing handbags with 94.48% in accuracy. We also apply the proposed FSCR-CNN model in recognizing other fine-grained objects with state-of-the-art CNN architectures, which is able to achieve over 5% improvement in accuracy. Yan Wang 0033, Sheng Li 0006, Alex Chichung Kot |
IEEE Trans. Multim. | 3 |
| 2014 | Image tampering detection by exposing blur type inconsistencyabstractIn this paper, we propose a novel method for image tampering detection in multi-type blurred images. After block-based image partitioning, a space-variant prior for local blur kernels is proposed for local blur kernels estimation. Then, the image blocks are clustered using a k-means clustering based on the similarity of local blur kernels to generate blur type invariant regions. Finally, blur types of the regions are classified into out-of-focus or motion blur using a minimum distance classifier. The experimental results show that the proposed method successfully detects and classifies the regions blur types which outperforms the state-of-the-art techniques. Our proposed approach is used to detect inconsistency in the partial blur types of an image as an evidence of image tampering. Khosro Bahrami, Alex Chichung Kot |
ICASSP | 2 |
| 2014 | Complementary feature extraction for branded handbag recognitionabstractFine-grained object recognition aims at recognizing objects belonging to the same basic-level class such as dog, bird or fish, which is a challenging problem in computer vision. In this paper, we consider the problem of recognizing handbags that belong to a specific brand. In order to identify the subtle differences among handbags, we propose to enhance the handbag local structure pattern by using the Hölder exponent, and extract the feature from the enhanced handbag image to complement the feature extracted directly from the original handbag image. We term such two types of features as the complementary and original features. These features will then be fused by using Multiple Kernel Learning (MKL) for branded handbag recognition. We conduct the experiments on a newly built branded handbag dataset, the results of which demonstrate the effectiveness of the proposed complementary feature in recognizing the handbags. Yan Wang 0033, Sheng Li 0006, Alex Chichung Kot |
ICIP | 3 |
| 2014 | An Integrated Clustering-Based Approach to Filtering Unfair Multi-Nominal TestimoniesabstractReputation systems have contributed much to the success of electronic marketplaces. However, the problem of unfair testimonies has to be addressed effectively to improve the robustness of reputation systems. Until now, most of the existing approaches focus only on reputation systems using binary testimonies, and thus have limited applicability and effectiveness. In this paper, We propose an integrated CLUstering‐Based approach called iCLUB to filter unfair testimonies for reputation systems using multinominal testimonies, in an example application of multiagent‐based e‐commerce. It adopts clustering techniques and considers buyer agents’ local as well as global knowledge about seller agents. Experimental evaluation demonstrates the promising results of our approach in filtering various types of unfair testimonies, its robustness against collusion attacks, and better performance compared to competing models. Siyuan Liu 0003, Jie Zhang 0002, Chunyan Miao, Yin Leng Theng, Alex Chichung Kot |
Comput. Intell. | 5 |
| 2014 | A Fast Approach for No-Reference Image Sharpness Assessment Based on Maximum Local VariationabstractThis letter proposes a simple and fast approach for no-reference image sharpness quality assessment. In this proposal, we define the maximum local variation (MLV) of each pixel as the maximum intensity variation of the pixel with respect to its 8-neighbors. The MLV distribution of the pixels is an indicative of sharpness. We use standard deviation of the MLV distribution as a feature to measure sharpness. Since high variations in the pixel intensities is a better indicator of the sharpness than low variations, the MLV of the pixels are subjected to a weighting scheme in such a way that heavier weights are assigned to greater MLVs to make the tail end of MLV distribution thicker. The weighting leads to an improvement of the MLV distribution to be more discriminative for different blur degrees. Finally, the standard deviation of the weighted MLV distribution is used as a metric to measure sharpness. The proposed approach has a very low computational complexity and the performance analysis shows that our approach outperforms the state-of-the-art techniques in terms of correlation with human vision system on several commonly used databases. Khosro Bahrami, Alex Chichung Kot |
IEEE Signal Process. Lett. | 2 |
| 2013 | A novel approach for partial blur detection and segmentationabstractThis paper proposes a novel approach for partial blur detection and segmentation. The local blur kernels of image blocks are firstly estimated and then a reblurring technique is used to measure relative blur degrees of the local blur kernels. The output of reblurring is a metric to classify blurred and non-blurred image blocks. Furthermore, block-based and pixel-based techniques are incorporated for a fine segmentation of blurred and non-blurred regions. Our approach is evaluated for out-of-focus and motion blurred images. The experimental results show that the proposed approach detects and segments the blurred and non-blurred regions in partial blurred images with 88% accuracy for natural out-of-focus blur, 86% accuracy for artificial out-of-focus blur and 83% accuracy for artificial motion blur, which outperforms the state-of-the-art approaches of partial blur detection and segmentation. Khosro Bahrami, Alex Chichung Kot, Jiayuan Fan 0001 |
ICME | 2 |
| 2013 | Mobile media communication, processing, and analysis: A review of recent advancesabstractIn this paper, we review recent advances in mobile media communication, processing, and analysis. To identify the opportunities and challenges in fast growing mobile media computing, we discuss several emerging topics including mobile visual search, retargeting, mobile video streaming, and cloud based mobile media computing. According to the infrastructure of mobile devices vs. servers, we come up with essential concerns in mobile media computing such as wireless bandwidth consumption, mobile energy saving, media adaptation for better quality of services, the computational load shift from mobiles to servers, etc. With booming mobile Apps on diverse media consumption, it is envisioned that mobile media research and development is bringing about significant achievements in traditional topics of communication, processing, and analytics. Wen Gao 0001, Ling-Yu Duan, Jun Sun 0007, Junsong Yuan 0001, Yonggang Wen 0001, Yap-Peng Tan, Jianfei Cai 0001, Alex Chichung Kot |
ISCAS | 8 |
| 2013 | A two-stage quality measure for mobile phone captured 2D barcode images
Changsheng Chen 0001, Alex Chichung Kot, Huijuan Yang |
Pattern Recognit. | 2 |
| 2013 | On Establishing Edge Adaptive Grid for Bilevel Image Data HidingabstractWe propose, in this paper, a novel edge-adaptive data hiding method for authenticating binary host images. Through establishing a dense edge-adaptive grid (EAG) along the object contours, we use a simple binary image to show that EAG more efficiently selects good data carrying pixel locations (DCPL) associated with “ l-shaped” patterns than block-based methods. Our method employs a dynamic system structure with the redesigned fundamental content adaptive processes (CAP) switch to iteratively trace new contour segments and to search for new DCPLs. By maintaining and updating a location status map, a protective mechanism is proposed to preserve the context of each CAP and their corresponding outcomes. We prove that our method is robust against the interferences caused by close-by contours, image noises, and invariantly selects the same sequence of DCPLs for an arbitrary binary host image and its various marked versions. Comparison shows that our method achieves a good tradeoff between large payload and minimal visual distortion as compared with several classic prior arts for diverse types of binary host images. Moreover, our method well supports state-of-the-art hybrid authentication that integrates data hiding and modern cryptographic techniques. Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2013 | Estimating EXIF Parameters Based on Noise Features for Image Manipulation DetectionabstractWe propose in this paper a novel technique to correlate statistical image noise features with three EXchangeable Image File format (EXIF) header features for manipulation detection. By formulating each EXIF feature as a weighted sum of selected statistical image noise features using sequential floating forward selection, the weights are then solved as a least squares solution for modeling the correlation between the intact image and the corresponding EXIF header. Image manipulations like brightness and contrast adjustment can affect these noise features and lead to enlarged numerical difference between each actual and its estimated EXIF feature from the noise features. By using the numerical difference as a manipulation indicator, we achieve excellent performance in detecting common brightness and contrast adjustment. Based on cameras of different brands, our manipulation detection is also demonstrated to work well in a blind mode, where the camera brand/model source is unavailable. Several detection examples suggest that our model can be applied in detecting real-world forgeries. Jiayuan Fan 0001, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2013 | Fingerprint Combination for Privacy ProtectionabstractWe propose here a novel system for protecting fingerprint privacy by combining two different fingerprints into a new identity. In the enrollment, two fingerprints are captured from two different fingers. We extract the minutiae positions from one fingerprint, the orientation from the other fingerprint, and the reference points from both fingerprints. Based on this extracted information and our proposed coding strategies, a combined minutiae template is generated and stored in a database. In the authentication, the system requires two query fingerprints from the same two fingers which are used in the enrollment. A two-stage fingerprint matching process is proposed for matching the two query fingerprints against a combined minutiae template. By storing the combined minutiae template, the complete minutiae feature of a single fingerprint will not be compromised when the database is stolen. Furthermore, because of the similarity in topology, it is difficult for the attacker to distinguish a combined minutiae template from the original minutiae templates. With the help of an existing fingerprint reconstruction approach, we are able to convert the combined minutiae template into a real-look alike combined fingerprint. Thus, a new virtual identity is created for the two different fingerprints, which can be matched using minutiae-based fingerprint matching algorithms. The experimental results show that our system can achieve a very low error rate with FRR = 0.4% at FAR = 0.1%. Compared with the state-of-the-art technique, our work has the advantage in creating a better new virtual identity when the two different fingerprints are randomly chosen. Sheng Li 0006, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2013 | Securing Online Reputation Systems Through Trust Modeling and Temporal AnalysisabstractWith the rapid development of reputation systems in various online social networks, manipulations against such systems are evolving quickly. In this paper, we propose scheme TATA, the abbreviation of joint Temporal And Trust Analysis, which protects reputation systems from a new angle: the combination of time domain anomaly detection and Dempster–Shafer theory-based trust computation. Real user attack data collected from a cyber competition is used to construct the testing data set. Compared with two representative reputation schemes and our previous scheme, TATA achieves a significantly better performance in terms of identifying items under attack, detecting malicious users who insert dishonest ratings, and recovering reputation scores. Yuhong Liu 0003, Yan Lindsay Sun, Siyuan Liu 0003, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2012 | A Dempster-Shafer theory based witness trustworthiness model to cope with unfair ratings in e-marketplaceabstractReputation systems have contributed much to resisting against the threats from malicious seller agents in electronic marketplaces. Buyer agents can benefit from modeling the reputation of seller agents to make a decision on which seller to have transaction with. However, the existence of unfair ratings decreases the accuracy of the seller agents' reputation evaluation, which will lead to inappropriate sellers to be selected and hence harm buyers' profits. In this paper, to address the problem of unfair ratings, we propose a witness trustworthiness model based on Dempster-Shafer theory to evaluate the trustworthiness of a witness's ratings regarding a particular seller. The proposed approach uses local and global ratings together to model a witness's trustworthiness which is specific for different sellers. The experimental results demonstrate that the proposed approach can effectively address the problem of unfair ratings and outperform the comparative approaches, especially when collusion attacks exist. Siyuan Liu 0003, Alex Chichung Kot, Chunyan Miao, Yin Leng Theng |
ICEC | 2 |
| 2012 | Step-edge reconstruction using 2D finite rate of innovation principleabstractParametric signals that have a finite number of degrees of freedom per unit of time are defined as signals with Finite Rate of Innovation (FRI). Sampling and reconstruction schemes have been developed based on the 1D FRI principle and applied to reconstructing step edge images on a row by row basis. In this paper, we derive the 2D FRI principle by exploiting the separability of the B-spline sampling kernel. The proposed 2D FRI principle regards the sampling and reconstruction as block by block operations. The step-edge parameters can be retrieved in high accuracy with no post-processing. The performance on synthetic images shows that our proposed technique is more precise than the row by row approaches on Signal-to-Noise Ratio (SNR) levels larger than 4 dB. Experimental results on real images demonstrate that the proposed method can reconstruct the step-edge precisely under noisy and practical sampling conditions. Changsheng Chen 0001, Pina Marziliano, Alex Chichung Kot |
ICASSP | 3 |
| 2012 | Manipulation Detection on Image Patches Using FusionBoostabstractIn this paper, we propose a novel manipulation detection framework for image patches using a fusion procedure, called FusionBoost, in conjunction with accurately detected derivative correlation features. By first dividing all demosaiced samples of a color image into a number of categories, we estimate their underlying demosaicing formulas based on partial derivative correlation models and extract several types of derivative correlation features. The features are organized into small subsets according to both the demosaicing category and the feature type. For each subset, we train a lightweight manipulation detector using probabilistic support vector machines. FusionBoost is then proposed to learn the weights of an ensemble detector for achieving the minimum error rate. By applying the ensemble detector on cropped photo patches from different image sources, large-scale experiments show that our proposed method achieves low average detection error rates of 2.0% to 4.3% in simultaneously detecting a large variety of common manipulation attempts for image patches from several different source models. Our framework shows good learning efficiency for highly imbalanced tasks. In several patch-based detection examples, we demonstrate the efficacy of the proposed method in detecting image manipulations on local patches. Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2012 | An Improved Scheme for Full Fingerprint ReconstructionabstractDifferent fingerprint recognition systems store minutiae-based fingerprint templates differently. Some store them inside a small token; some can be found in a server database. As the minutiae template is very compact, many take it for granted that the template does not contain sufficient information for reconstructing the original fingerprint. This paper proposes a scheme to reconstruct a full fingerprint image from the minutiae points based on the amplitude and frequency modulated (AM-FM) fingerprint model. The scheme starts with generating a binary ridge pattern which has a similar ridge flow to that of the original fingerprint. The continuous phase is intuitively reconstructed by removing the spirals in the phase image estimated from the ridge pattern. To reduce the artifacts due to the discontinuity in the continuous phase, a refinement process is introduced for the reconstructed phase image, which is the combination of the continuous phase and the spiral phase (corresponding to the minutiae). Finally, the refined phase image is used to produce a thinned version of the fingerprint, from which a real-look alike gray-scale fingerprint image is reconstructed. The experimental results show that our proposed scheme performs better than the-state-of-the-art technique. Sheng Li 0006, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2012 | Binarization of Low-Quality Barcode Images Captured by Mobile Phones Using Local Window of Adaptive Location and SizeabstractIt is difficult to directly apply existing binarization approaches to the barcode images captured by mobile device due to their low quality. This paper proposes a novel scheme for the binarization of such images. The barcode and background regions are differentiated by the number of edge pixels in a search window. Unlike existing approaches that center the pixel to be binarized with a window of fixed size, we propose to shift the window center to the nearest edge pixel so that the balance of the number of object and background pixels can be achieved. The window size is adaptive either to the minimum distance to edges or minimum element width in the barcode. The threshold is calculated using the statistics in the window. Our proposed method has demonstrated its capability in handling the nonuniform illumination problem and the size variation of objects. Experimental results conducted on 350 images captured by five mobile phones achieve about 100% of recognition rate in good lighting conditions, and about 95% and 83% in bad lighting conditions. Comparisons made with nine existing binarization methods demonstrate the advancement of our proposed scheme. Huijuan Yang, Alex Chichung Kot, Xudong Jiang 0001 |
IEEE Trans. Image Process. | 2 |
| 2011 | A novel system for fingerprint privacy protection1abstractThis paper proposes a novel system for protecting the fingerprint privacy without using a token or key. In the enrollment, two fingerprints are captured from two of an user's fingers. We extract the minutiae positions from one fingerprint, the orientation from the other fingerprint and the primary cores from both fingerprints. Based on these extracted information, a combined minutiae template is generated and stored in a database. In the authentication, the user needs to provide two query fingerprints from the same two fingers which are used in the enrollment. By storing the combined minutiae template, the complete minutiae feature of a single fingerprint will not be compromised when the database is stolen. Furthermore, because of the similarity in topology, it is also difficult for the attacker to distinguish our template from the minutiae of an original fingerprint. We evaluate the performance of our system over the FVC2002 DB2_A database. The results show that the False Rejection Rate of our system is 3% when the False Acceptance Rate is 0.01%. Sheng Li 0006, Alex Chichung Kot |
IAS | 2 |
| 2011 | Modeling the EXIF-Image correlation for image manipulation detectionabstractEXchangeable Image File format (EXIF) is a metadata header containing shot-related camera settings such as aperture, exposure time, ISO speed etc. These settings can affect the photo content in many ways. In this paper, we investigate the underlying EXIF-Image correlation and propose a novel model, which correlates image statistical noise features with several commonly used EXIF features. By formulating each EXIF feature as a weighted combination of different image statistical noise features, we first select a compact image statistical noise feature set using sequential floating forward selection. The underlying correlation as a set of regression weights is then solved using a least squares solution. When applying our learned correlation to detect image manipulation, we achieve average test accuracies of 94.6%, 94.1% and 94.9% in three different cameras to detect the presence of common image brightness and contrast adjustment. Jiayuan Fan 0001, Alex Chichung Kot, Farook Sattar |
ICIP | 2 |
| 2011 | Unsharp Masking Sharpening Detection via Overshoot Artifacts AnalysisabstractIn this letter, we propose a new method in detecting unsharp masking (USM) sharpening operation in digital images. Overshoot artifacts are found to occur around side-planar edges in the sharpened images. Such artifacts, measured by a sharpening detector, can serve as a rather unique feature for identifying the previous performance of sharpening operation. Test results on photograph images with regard to various sharpening operators show the effectiveness of our proposed method. Gang Cao 0001, Yao Zhao 0001, Alex Chichung Kot |
IEEE Signal Process. Lett. | 4 |
| 2011 | Privacy Protection of Fingerprint DatabaseabstractA fingerprint authentication system for the privacy protection of the fingerprint template stored in a database is introduced here. The considered fingerprint data is a binary thinned fingerprint image, which will be embedded with some private user information without causing obvious abnormality in the enrollment phase. In the authentication phase, these hidden user data can be extracted from the stored template for verifying the authenticity of the person who provides the query fingerprint. A novel data hiding scheme is proposed for the thinned fingerprint template. This scheme does not produce any boundary pixel in the thinned fingerprint during data embedding. Thus, the abnormality caused by data hiding is visually imperceptible in the marked-thinned fingerprint. Compared with using existing binary image data hiding techniques, the proposed method causes the least abnormality for a thinned fingerprint without compromising the performance of the fingerprint identification. Sheng Li 0006, Alex Chichung Kot |
IEEE Signal Process. Lett. | 2 |
| 2010 | Identification of recaptured photographs on LCD screensabstractWith advances in image display technology, recapturing good-quality images from the high-fidelity artificial scenery on a LCD screen becomes possible. Such image recapturing posts a security threat, which allows the forgery images to bypass the current forensic systems. In this paper, we first recapture some good-quality photos on different LCD screens by properly setting up the recapturing environment and tuning the controllable settings. In a perceptional study, we find that such finely recaptured images can hardly be identified by human eyes. To prevent the image recapturing attack, we propose a set of statistical features, which capture the common anomalies introduced in the camera recapturing process on LCD screens. With a probabilistic support vector machine classifier, comparison results show that our proposed features work very well, which outperform the conventional image forensic features in identification of the finely recaptured images. Alex Chichung Kot |
ICASSP | 2 |
| 2010 | Knowledge guided adaptive binarization for 2D barcode images captured by mobile phonesabstractIn this paper, we investigate the problem of binarization of the 2D barcode images captured by mobile devices. The poor quality of the images due to noise, nonuniform illumination and inherent limitation and distortion of the camera makes the task of binarization more challenging. The global thresholding techniques employed by most existing approaches do not work well for the barcode images captured under uncontrolled conditions. Hence, we propose an adaptive binarization technique by taking the prior knowledge such as the edge structure and maximum element width of the 2D barcode into consideration. Incrementally updating the elements in a moving window is proposed to expediate the process for mobile phone applications. Comparisons with other binarization methods showthe superior performance of our method, which achieves about 96.6% recognition rate on a dataset of 787 images. Huijuan Yang, Alex Chichung Kot, Xudong Jiang 0001 |
ICASSP | 2 |
| 2010 | A quality measure of mobile phone captured 2D barcode imagesabstract2D barcode based mobile tagging technique has not fully matured yet. One of the difficulties is that mobile phone cameras induce inevitable distortions in captured barcode images. For a variety of reasons to be discussed herein, evaluation of the decodability of a barcode image is needed before the decoding process. We propose a novel blurriness measure to classify the samples according to their decodabilities. The proposed quality measure of 2D barcode image is computed by utilizing distinctive histogram features. Experimental results show that the classification accuracy is as high as 94% in our database which consists of samples captured with different phones and conditions. Comparisons with state-of-art quality measures are also conducted. The proposed method is not sensitive to noise, rotation and scaling. It is also applicable to many popular 2D barcode patterns. Changsheng Chen 0001, Alex Chichung Kot, Huijuan Yang |
ICIP | 2 |
| 2010 | Accurate localization of four extreme corners for barcode images captured by mobile phonesabstractIn this paper, we propose a novel method to locate the four extreme corners of barcodes in the images captured by mobile phones. To achieve this goal, the two nearly-parallel outer boundary lines are firstly localized by utilizing the prior knowledge of the relative distances and angles between the lines, which are subsequently employed to obtain the initially localized corners. A novel post-localization process based on edge tracing is also proposed to further validate the initially localized corners. This is achieved by setting the constraints in maximum direction change when tracing from the current to the next candidate corners. The distinctive feature of the proposed algorithm lies in the capability of handling those curved barcode images taken by the mobile phones. Experiments conducted on a data set of 1410 images show the accuracy of the corner localization is about 92.3% and 90.5% for indoor and outdoor barcode images, respectively. The processing time taken is about 0.8 second using Matlab. Huijuan Yang, Xudong Jiang 0001, Alex Chichung Kot |
ICIP | 3 |
| 2010 | Dynamic window construction for the binarization of barcode images captured by mobile phonesabstractIt is difficult to directly apply existing binarization methods to the barcode images captured by mobile devices under uncontrolled lighting conditions. Employing a fix-sized window to locally binarize the image cannot handle the situation when small and large objects co-exist in an image. This paper proposes a novel scheme to dynamically determine the size of the binarization window, which changes with the presence of the high gradient pixels in a search window centering the processing pixel. The proposed method has demonstrated its capability in handling objects of different sizes and the uneven illumination problem. Further, it is not constrained to certain barcode types. Experimental results conducted on 330 images captured using five mobile phones in indoor and outdoor environments achieve about 93% recognition rate evaluated using a barcode decoder. Comparisons with some existing approaches show its superior binarization performance. Huijuan Yang, Alex Chichung Kot, Xudong Jiang 0001 |
ICIP | 2 |
| 2010 | Privacy protection of fingerprint database using lossless data hidingabstractIn this paper, we introduce a fingerprint authentication system for protecting the privacy of the fingerprint template stored in a database. The template, which is a binary fingerprint image after thinning, will be embedded with private personal data in the user enrollment phase. In the user authentication phase, these hidden personal data can be extracted from the stored template for verifying the authenticity of the person who provides the query fingerprint. A novel lossless data hiding scheme is proposed for a thinned fingerprint. By adopting “embeddability criterion”, data is hidden into the template by just adding some boundary pixels in the template. These boundary pixels can be extracted and removed to reconstruct the original thinned fingerprint so that fingerprint matching accuracy is not affected. Compared with using existing binary image data hiding techniques, our scheme has a better performance for a thinned fingerprint. Sheng Li 0006, Alex Chichung Kot |
ICME | 2 |
| 2010 | Image watermarking using dual-tree complex wavelet by coefficients swapping and group of coefficients quantizationabstractIn this paper, we present two watermarking schemes for images using the low-pass frequency coefficients and high-pass complex frequency coefficients of the Dual-Tree Complex Wavelet Transform (DT-CWT) of the image, the distinctive characteristics of the two schemes can lead to different applications. We investigate the problem of embedding binary watermark sequence in DT-CWT domain, which is a challenging problem. The coefficients swapping is done by carefully selecting those 2×2 blocks with high-energy in the low frequency and swap the coefficients that lie in the median range in the block to embed the watermark. Whereas the group of coefficients quantization quantizes a group of high-pass complex frequency coefficients to make the quantized coefficients lie in the middle of the quantization range and distribute the changes among the coefficients. Experimental results conducted on 100 standard images achieve about 96% and 92% detection rate for the low-pass and high-pass frequency coefficients-based schemes, respectively. Robustness tests against some common signal processing attacks such as additive noise, median filtering and lossy JPEG compression confirm the superior performance of the proposed schemes. Huijuan Yang, Xudong Jiang 0001, Alex Chichung Kot |
ICME | 3 |
| 2010 | Mobile camera identification using demosaicing featuresabstractMobile cameras are typically low-end cameras equipped on handheld devices such as personal digital assistants and cellular phones. The fast proliferation of these mobile cameras has brought up concerns on the origin and integrity of their output images. In this paper, we identify blindly the source mobile cameras by combining 3 types of demosaicing features extracted from a test image. Through Eigenfeature regularization and feature reduction, comparison results show our Eigen demosaicing features perform significantly better than several conventional features in distinguishing 9 mobile cameras of dissimilar models based on cropped image blocks. By including cameras of the same and very similar models, in 15-cam identification, our Eigen demosaicing features achieves excellent classification accuracies in distinguishing cameras from dissimilar models and the classification accuracies expectedly tend to confuse among cameras of the same or very similar models. Alex Chichung Kot |
ISCAS | 2 |
| 2010 | Detection of Tampering Inconsistencies on Mobile Photos
Alex Chichung Kot |
IWDW | 2 |
| 2010 | Individuality of alphabet knowledge in online writer identification
Guo Xian Tan, Christian Viard-Gaudin, Alex Chichung Kot |
Int. J. Document Anal. Recognit. | 3 |
| 2010 | Two-Dimensional Polar Harmonic Transforms for Invariant Image RepresentationabstractThis paper introduces a set of 2D transforms, based on a set of orthogonal projection bases, to generate a set of features which are invariant to rotation. We call these transforms Polar Harmonic Transforms (PHTs). Unlike the well-known Zernike and pseudo-Zernike moments, the kernel computation of PHTs is extremely simple and has no numerical stability issue whatsoever. This implies that PHTs encompass the orthogonality and invariance advantages of Zernike and pseudo-Zernike moments, but are free from their inherent limitations. This also means that PHTs are well suited for application where maximal discriminant information is needed. Furthermore, PHTs make available a large set of features for further feature selection in the process of seeking for the best discriminative or representative features for a particular application. Pew-Thian Yap, Xudong Jiang 0001, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2010 | Prediction of eigenvalues and regularization of eigenfeatures for human face verification
Bappaditya Mandal, Xudong Jiang 0001, How-Lung Eng, Alex Chichung Kot |
Pattern Recognit. Lett. | 4 |
| 2010 | Lossless data embedding in electronic inksabstractThis paper presents a novel lossless data embedding algorithm for electronic inks. The proposed algorithm first computes the analytical ink-curve for each stroke as a set of smoothly concatenated cubic Bezier curves. During embedding, a set of data carrier points on the ink-curve are evaluated, perturbed, and then inserted into the original point array. Our proposed smooth ink-curve generation method iteratively refines the parameterization based on the local curve geometry. We demonstrate experimentally that the inserted points incur the least perceptional distortion on the marked ink-curves even after large mount of secret message is embedded. On extraction, the perturbed data carriers are identified and the original point array can be recovered. By subtracting these perturbed data carriers from their recomputed reference locations, we derive a sequence of perturbation vectors, from which the embedded secret message can be decoded. Based on the proposed embedding technique and the public-key cryptography, a hybrid secure ink authentication system is designed to secure electronic writings for tampering detection. Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2009 | Lossless data hiding for electronic inkabstractThis paper presents a novel lossless data hiding technique for electronic inks. The proposed algorithm first computes the analytical ink-curve function for each stroke as a set of smoothly concatenated cubic Bezier curves. During embedding, a set of data carrier points on the ink-curve are computed, perturbed and inserted back into the original point array together with some marker points. In extraction, data carrier points are identified and compared with their original positions in order to decode the secret message. Through experiments, we demonstrate that significant mount of secret message can be embedded and the ink-curve computed from the marked ink data closely resembles the original ink-curve. Alex Chichung Kot |
ICASSP | 2 |
| 2009 | Impact of Alphabet Knowledge on Online Writer IdentificationabstractCharacter prototype approaches for writer identification produces a consistent set of templates that are used to model the handwriting styles of writers, thereby allowing high accuracies to be attained. This paper extends such work on writer identification by investigating the usage of alphabet knowledge derived from the character prototypes. In addition, we demonstrate the concept of discriminative power of alphabets. It is not unconceivable that certain alphabets allow writers to express their individuality of handwriting with a more distinct and unique style compared with other alphabets. This paper establishes that such alphabets have higher discriminative powers in identifying writers. Experiments related to the reduction in dimensionality of the writer identification system are also reported. Our results show that the discriminative power of alphabet can be used to reduce the complexity while maintaining the same level of performance for the writer identification system. Guo Xian Tan, Christian Viard-Gaudin, Alex Chichung Kot |
ICDAR | 3 |
| 2009 | Information Retrieval Model for Online Handwritten Script IdentificationabstractScript identification has always been a topic of much research interest in the field of document analysis. The accurate determination of the identity of the script is paramount to many post-processing steps such as document sorting, translation and in determining the choice of linguistic resources to use for OCR or handwriting recognition. However, few works exist with regards to the identification of online handwritten scripts, partly due to the large variations and challenges innate in handwritten scripts. This paper proposes a novel approach for online handwritten script identification based on the information retrieval model. We attempt to identify among three script families; Arabic, Roman and Tamil scripts, which attained an average accuracy of 93.3% from our results. This signifies promising potential in utilizing information retrieval models for script identification. Guo Xian Tan, Christian Viard-Gaudin, Alex Chichung Kot |
ICDAR | 3 |
| 2009 | RAW tool identification through detected demosaicing regularityabstractRAW tools are PC software tools that develop the RAWs, i.e. the camera sensor data, into full-color photos. In this paper, we propose to study the internal processing characteristics of these RAW tools using 3 heterogeneous sets of demosaicing features. Through feature-level fusion, normalization and an Eigen-space regularization technique, we derive a compact set of discriminant features. Experimentally, we find that the compact feature set can be used to accurately distinguish 40 RAW-tool classes. A dissimilarity study also shows that the cropped image blocks from different RAW-tool or positional classes have a great deal of dissimilarity in our extracted demosaicing features. Alex Chichung Kot |
ICIP | 2 |
| 2009 | Accurate Detection of Demosaicing Regularity from Output ImagesabstractDemosaicing regularity is an important processing regularity associated with the internal camera processing and its detection from output photos is useful for non-intrusive forensic engineering. In this paper, we propose a reverse grouping technique to improve the detection accuracy of our earlier proposed detection model based on second-order image derivatives. Comparison results based on syntactic images shows that the proposed technique significantly reduces the reprediction errors for some commonly used demosaicing algorithms. When applied to a real application, i.e. camera model identification, our demosaicing features in conjunction with probabilistic support vector machine classifier achieve excellent classification performance. Alex Chichung Kot |
ISCAS | 2 |
| 2009 | Complete discriminant evaluation and feature extraction in kernel space for face recognition
Xudong Jiang 0001, Bappaditya Mandal, Alex Chichung Kot |
Mach. Vis. Appl. | 3 |
| 2009 | A multi-prototype clustering algorithm
Manhua Liu, Xudong Jiang 0001, Alex Chichung Kot |
Pattern Recognit. | 3 |
| 2009 | Automatic writer identification framework for online handwritten documents using character prototypes
Guo Xian Tan, Christian Viard-Gaudin, Alex Chichung Kot |
Pattern Recognit. | 3 |
| 2009 | Steganalysis of halftone image using inverse halftoning
Jun Cheng 0003, Alex Chichung Kot |
Signal Process. | 2 |
| 2009 | Accurate detection of demosaicing regularity for digital image forensicsabstractIn this paper, we propose a novel accurate detection framework of demosaicing regularity from different source images. The proposed framework first reversely classifies the demosaiced samples into several categories and then estimates the underlying demosaicing formulas for each category based on partial second-order derivative correlation models, which detect both the intrachannel and the cross-channel demosaicing correlation. An expectation-maximization reverse classification scheme is used to iteratively resolve the ambiguous demosaicing axes in order to best reveal the implicit grouping adopted by the underlying demosaicing algorithm. Comparison results based on syntactic images show that our proposed formulation significantly improves the accuracy of the regenerated demosaiced samples from the sensor samples for a large number of diversified demosaicing algorithms. By running sequential forward feature selection, our reduced feature sets used in conjunction with the probabilistic support vector machine classifier achieve superior performance in identifying 16 demosaicing algorithms in the presence of common camera post demosaicing processing. When applied to real applications, including camera model and RAW-tool identification, our selected features achieve nearly perfect classification performances based on large sets of cropped image blocks. Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2008 | A generalized model for detection of demosaicing characteristicsabstractDetection of demosaicing characteristics based on correlation at pixel level is useful for image forensic purposes. In this paper, we propose a new model to detect the demosaicing characteristics based on the second-order derivative correlation. Without the influences of the uneven DC components in the three color channels, the new model can estimate both intra-channel and cross-channel correlation. Experimental results show efficacy of the proposed model. Alex Chichung Kot |
ICME | 2 |
| 2008 | Verification of human faces using predicted eigenvaluesabstractTo alleviate the conventional problems of LDA and its variants, we propose a procedure of predicting eigenvalues using few reliable eigenvalues from the range space. Partitioning of entire eigenspace is performed using two control points, however, the effective low dimensional discriminative vectors are extracted from the whole eigenspace. This prediction strategy enables to perform discriminant evaluation in the full eigenspace. The proposed method is evaluated and compared with 8 popular subspace based methods for face verification task. Experimental results on popular face databases show that our method consistently outperforms others. Bappaditya Mandal, Xudong Jiang 0001, Alex Chichung Kot |
ICPR | 3 |
| 2008 | A stochastic nearest neighbor character prototype approach for online writer identificationabstractOne novel technique for identifying the writer of an online handwritten document is proposed. This technique makes use of a character prototype distribution to model the specific allographs used by a given writer. In this paper, we propose to extend and improve upon this newly established methodology [1] by making use of a stochastic nearest neighbor algorithm to estimate the character prototype distribution. The proposed method is text independent and relies on the automatic segmentation of the handwritten text at the character level. Our results show that this approach attained a writer identification rate of 99.2% when a reference database of 120 writers is used. Experiments related to the effect of the length of text of the document on the performance of the writer identification system are also reported. Guo Xian Tan, Christian Viard-Gaudin, Alex Chichung Kot |
ICPR | 3 |
| 2008 | Boosted complex moments for discriminant rotation invariant object recognitionabstractThis paper proposes a method for constructing a discriminative rotation invariant object recognition system from the set of complex moments by using a multi-class boosting algorithm. Experimental results show that a large of number images can be discriminated accurately with only a small number of features. This basically means economy of computational effort in feature acquisition and also possibility of higher speed in recognition task. Pew-Thian Yap, Xudong Jiang 0001, Alex Chichung Kot |
ICPR | 3 |
| 2008 | An interactive and secure user authentication scheme for mobile devicesabstractGraphical password (i.e., image based authentication) is considered as a promising alternative to traditional textual password for mobile devices, to achieve better tradeoff between usability and security. However, previous proposals of graphical password have the limitation of limited entropy. In this paper, we propose a new scheme incorporating user face based authentication into the association-based graphical password solution we proposed before, aiming at achieving higher security without compromising user-friendliness for mobile application scenarios. System performance analysis and comparisons with other schemes are presented to validate our scheme. Qibin Sun, Zhi Li 0001, Xudong Jiang 0001, Alex Chichung Kot |
ISCAS | 4 |
| 2008 | Backward-forward distortion minimization for binary images data hidingabstractIn our previous paper, we propose to use the morphological transform for binary images data hiding for authentication purpose. We view flipping an edge pixel in binary images as shifting the edge location one pixel horizontally and vertically. Based on this observation, we propose an interlaced morphological binary wavelet transform to track the shifted edges. The two processing cases that flipping the candidates of one does not affect the flippability conditions of another are employed such that a large capacity can be achieved. However, employing double processing cases sacrifices the visual quality of the watermarked image. In this paper, we further investigate and propose a novel effective Backward-Forward Minimization method to minimize the visual distortion for double processing cases. Experimental results demonstrate its superior performance. Huijuan Yang, Alex Chichung Kot |
ISCAS | 2 |
| 2008 | Eigenfeature Regularization and Extraction in Face RecognitionabstractThis work proposes a subspace approach that regularizes and extracts eigenfeatures from the face image. Eigenspace of the within-class scatter matrix is decomposed into three subspaces: a reliable subspace spanned mainly by the facial variation, an unstable subspace due to noise and finite number of training samples and a null subspace. Eigenfeatures are regularized differently in these three subspaces based on an eigenspectrum model to alleviate problems of instability, over-fitting or poor generalization. This also enables the discriminant evaluation performed in the whole space. Feature extraction or dimensionality reduction occurs only at the final stage after the discriminant assessment. These efforts facilitate a discriminative and stable low-dimensional feature representation of the face image. Experiments comparing the proposed approach with some other popular subspace methods on the FERET, ORL, AR and GT databases show that our method consistently outperforms others. Xudong Jiang 0001, Bappaditya Mandal, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2008 | Performance of Space-Time Codes: Gallager Bounds and Weight EnumerationabstractSince the standard union bound for space–time codes may diverge in quasi-static fading channels, the limit-before-average (LBA) technique has been exploited to derive tight performance bounds. However, it suffers from the computational burden arising from a multidimensional integral. In this paper, efficient bounding techniques for space–time codes are developed in the framework of Gallager bounds. Two closed-form upper bounds, the ellipsoidal bound and the spherical bound, are proposed that come close to simulation results within a few tenths of a decibel. In addition, two novel methods of weight enumeration operating on a further reduced state diagram are presented, which, in conjunction with the bounding techniques, give a thorough treatment of performance bounds for space–time codes. Cong Ling 0001, Kwok Hung Li, Alex Chichung Kot |
IEEE Trans. Inf. Theory | 3 |
| 2008 | Orthogonal Data Embedding for Binary Images in Morphological Transform Domain- A High-Capacity ApproachabstractThis paper proposes a data-hiding technique for binary images in morphological transform domain for authentication purpose. To achieve blind watermark extraction, it is difficult to use the detail coefficients directly as a location map to determine the data-hiding locations. Hence, we view flipping an edge pixel in binary images as shifting the edge location one pixel horizontally and vertically. Based on this observation, we propose an interlaced morphological binary wavelet transform to track the shifted edges, which thus facilitates blind watermark extraction and incorporation of cryptographic signature. Unlike existing block-based approach, in which the block size is constrained by 3times3 pixels or larger, we process an image in 2times2 pixel blocks. This allows flexibility in tracking the edges and also achieves low computational complexity. The two processing cases that flipping the candidates of one does not affect theflippabilityconditions of another are employed for orthogonal embedding, which renders more suitable candidates can be identified such that a larger capacity can be achieved. A novel effective Backward-Forward Minimization method is proposed, which considers both backwardly those neighboring processedembeddablecandidates and forwardly those unprocessedflippablecandidates that may be affected by flipping the current pixel. In this way, the total visual distortion can be minimized. Experimental results demonstrate the validity of our arguments. Huijuan Yang, Alex Chichung Kot, Susanto Rahardja |
IEEE Trans. Multim. | 2 |
| 2007 | Steganalysis of Binary Cartoon Image using Distortion MeasureabstractWe present a steganalysis technique for data hiding in binary cartoon images. Due to the perturbation from the embedding, the contours of a stego binary cartoon image are distorted. When calculating the distortion of the image based on a distortion measure, the distortion score between a stego image and its de-noised version should be different from that between an original image and its de-noised version. Different binary image distortion measures are used to calculate the distortion scores, which are used as the features for classification. The sequential floating forward search (SFFS) method is used to search for the combination of the features that yields the best classification results. Jun Cheng 0003, Alex Chichung Kot, Susanto Rahardja |
ICASSP (2) | 2 |
| 2007 | Face Recognition Based on Discriminant Evaluation in the Whole SpaceabstractThis paper proposes a face recognition approach that performs linear discriminant analysis in the whole eigenspace. It decomposes the eigenspace into two subspaces: a reliable subspace spanned mainly by the facial variation and an unstable subspace due to finite number of training samples. Eigenvalues in the unstable subspace are replaced by a constant. This alleviates the over-fitting problem and enables the discriminant evaluation in the whole space. Feature extraction or dimensionality reduction occurs only at the final stage after the discriminant assessment. These efforts facilitate a discriminative and stable low-dimensional feature representation of the face image. Experimental results comparing some popular subspace methods on FERET and ORL databases show that our approach consistently outperforms others. Xudong Jiang 0001, Bappaditya Mandal, Alex Chichung Kot |
ICASSP (2) | 3 |
| 2007 | Database Clustering Based on Multi-Prototype Representation of ClusterabstractClustering is a useful technique to provide the organization of multimedia database. Using single prototype to represent each cluster may not adequately model the different types of clusters and hence limits the clustering performance on the complex data structure. This paper proposes a clustering algorithm based on multi-prototype representation of cluster. The square-error clustering is used to produce a number of prototypes to locate the regions of high density. The prototypes are organized into a given number of clusters in agglomerative method based on a proposed separation measure. New prototypes are iteratively added to improve the poor cluster boundaries. As a result, the proposed algorithm can discover the clusters of complex structure. Experimental results demonstrate the effectiveness of the proposed clustering algorithm. Manhua Liu, Xudong Jiang 0001, Alex Chichung Kot |
ICME | 3 |
| 2007 | High Capacity Data Hiding for Binary Images in Morphological Wavelet Transform DomainabstractThis paper investigates the problem of data hiding for binary images authentication in morphological binary wavelet transform domain. Directly using the detail coefficients as a location map to determine the data hiding locations is difficult to achieve blind watermark extraction. Hence, we view flipping an edge pixel as shifting the edge location one pixel horizontally and vertically. Based on this observation, we propose an interlaced transform to track the shifted edges. Different from existing block-based approach, in which the block size is constrained by not less than 3 x 3, we process the image in 2 x 2 blocks. The two processing cases that the flippantly condition of one is not affected by flipping the candidates of another are combined such that an extremely large capacity can be achieved, which slightly sacrifices the visual quality. Huijuan Yang, Alex Chichung Kot, Susanto Rahardja, Xudong Jiang 0001 |
ICME | 2 |
| 2007 | Data Hiding For Binary Images Authentication By Considering A Larger NeighborhoodabstractThis paper proposes a novel data hiding method for binary images authentication aims at preserving the connectivity of pixels in a local 3times3 neighborhood and dynamically improving the "flipppability" decision by considering a larger neighborhood. The interlaced morphological binary wavelet transform (IMBWT) is employed as the basic processing tool. The primary objective of considering a larger neighborhood is to improve the visual quality of the watermarked image. The post-checking constraints ensure that the "embeddability" of a pixel is invariant in the data embedding process thus facilitates blind watermark extraction and incorporation of the cryptographic signature. "Single flipping" and "pair flipping" are discussed. The "connectivity-preserving" and "transition change" patterns are properly paired and classified so that each category is isolated from other categories, hence no ambiguity exists. The "end stroke" patterns that are orthogonal to the "connectivity-preserving" and "transition change" patterns are employed such that a double data hiding scheme can be designed. Huijuan Yang, Alex Chichung Kot |
ISCAS | 2 |
| 2007 | A General Data Hiding Framework and Multi-level Signature for Binary Images
Huijuan Yang, Alex Chichung Kot |
IWDW | 2 |
| 2007 | Efficient fingerprint search based on database clustering
Manhua Liu, Xudong Jiang 0001, Alex Chichung Kot |
Pattern Recognit. | 3 |
| 2007 | Objective Distortion Measure for Binary Text Image Based on Edge Line Segment SimilarityabstractThis paper proposes a new approach to measure the distortion introduced by changing individual edge pixels in binary text images. The approach considers not only how many pixels are changed but also where the pixels are changed and how the flipping affects the overall shape formed by the edge line. Similarities between the edge line segments in the original and distorted image are compared to measure the distortion. Subjective testing shows that the new distortion measure correlates well with human visual perception. Jun Cheng 0003, Alex Chichung Kot |
IEEE Trans. Image Process. | 2 |
| 2007 | Gallager Bounds for Noncoherent Decoders in Fading ChannelsabstractRecently, Gallager's bounding techniques have been used to derive tight performance bounds for coded systems in fading channels. Most works in this field have thus far dealt with coherent decoding. This paper develops Gallager bounds for noncoherent systems in fading channels. Unlike coherent decoding, the exact error probability of a noncoherent decoder/detector conditioned on the fading coefficients does not admit a closed-form expression. This difficulty is overcome in this paper by employing the Chernoff technique. Although it weakens the bounds to some extent, the Chernoff technique enables the derivations of the limit-before-average (LBA) bound and Gallager bounds in closed form for noncoherent fading channels. Numerical examples show that the proposed bounds are convergent and are tighter than the conventional union bound. Cong Ling 0001, Xiaofu Wu, Kwok Hung Li, Alex Chichung Kot |
IEEE Trans. Inf. Theory | 4 |
| 2007 | Pattern-Based Data Hiding for Binary Image Authentication by Connectivity-PreservingabstractIn this paper, a novel blind data hiding method for binary images authentication aims at preserving the connectivity of pixels in a local neighborhood is proposed. The "flippability" of a pixel is determined by imposing three transition criteria in a 3 times 3 moving window centered at the pixel. The "embeddability" of a block is invariant in the watermark embedding process, hence the watermark can be extracted without referring to the original image. The "uneven embeddability" of the host image is handled by embedding the watermark only in those "embeddable" blocks. The locations are chosen in such a way that the visual quality of the watermarked image is guaranteed. Different types of blocks are studied and their abilities to increase the capacity are compared. The problem of how to locate the "embeddable" pixels in a block for different block schemes is addressed which facilitates the incorporation of the cryptographic signature as the hard authenticator watermark to ensure integrity and authenticity of the image. Discussions on the security considerations, visual quality against capacity, counter measures against steganalysis and analysis of the computational load are provided. Comparisons with prior methods show superiority of the proposed scheme Huijuan Yang, Alex Chichung Kot |
IEEE Trans. Multim. | 2 |
| 2006 | Two-Layer Binary Image Authentication With Tampering LocalizationabstractIn this paper, a novel two-layer blind binary image authentication scheme is proposed, in which the first layer is targeted at the overall authentication and the second layer is targeted at identifying the tampering locations. The "flippability" of a pixel is determined by the three "Connectivity-Preserving" transition criterions described in [4]. The image is partitioned into multiple Micro-Blocks and the Micro-Blocks are classified into eight categories. The Block Identifier is defined adaptively for each class and embedded in those "Qualified" and "Self-Detecting" Micro-Blocks in order to identify the tampered locations. The Block Identifier is defined in such a way that any changes occurred either to the "Qualified" or its neighboring "Un-qualified' Micro-Blocks will render the retrieved Block Identifier different from the one embedded. Experimental results validate the arguments made. Discussions on the accuracy of localization of tamperings are provided. Huijuan Yang, Alex Chichung Kot |
ICASSP (2) | 2 |
| 2006 | Fault and intrusion tolerance of wireless sensor networksabstractThe following three questions should be answered in developing new topology with more powerful ability to tolerate node-failure in wireless sensor network. First, what is node-failure tolerance of topologies? Second, how to evaluate this tolerance ability? Third, which type of topologies is more efficient in tolerating node-failure? Without giving the answers, the existing work regards fault-tolerance topology as the multiply connected graph, and use the connectivity of the graph as the standard to evaluate tolerance ability. In this paper, we argue that fault tolerance of topologies is not equivalent to the connectivity of multiply connected graph by illustrating two concrete examples. Then the definition of node-failure tolerance is presented. According fault and intrusion, the two sources of failure nodes, we define fault tolerance and intrusion tolerance as the standards to evaluate the tolerance ability of topologies, and analyze the tolerance performance of hierarchical structure of wireless sensor network by using these standards. Finally, the function relation between hierarchical topology and its tolerance abilities of fault and intrusion is obtained, and an obvious corollary is that fault tolerance increase with the ratio of cluster head hierarchical structure, but with the intrusion tolerance decreasing. Liangmin Wang 0001, Jianfeng Ma 0001, Chao Wang 0085, Alex Chichung Kot |
IPDPS | 4 |
| 2006 | Binary Image Authentication With Tampering Localization by Embedding Cryptographic Signature and Block IdentifierabstractThis letter proposes a novel two-layer blind binary image authentication scheme, in which the first layer is targeted at the overall authentication and the second layer is targeted at identifying the tampering locations. The "flippability" of a pixel is determined by the "connectivity-preserving" transition criterion. The image is partitioned into multiple macro-blocks that are subsequently classified into eight categories. The block identifier is defined adaptively for each class and embedded in those "qualified" and "self-detecting" macro-blocks in order to identify the tampered locations. Experimental results validate the arguments made Huijuan Yang, Alex Chichung Kot |
IEEE Signal Process. Lett. | 2 |
| 2006 | Fingerprint Retrieval for IdentificationabstractThis paper presents a front-end filtering algorithm for fingerprint identification, which uses orientation field and dominant ridge distance as retrieval features. We propose a new distance measure that better quantifies the similarity evaluation between two orientation fields than the conventional Euclidean and Manhattan distance measures. Furthermore, fingerprints in the data base are clustered to facilitate a fast retrieval process that avoids exhaustive comparisons of an input fingerprint with all fingerprints in the data base. This makes the proposed approach applicable to large databases. Experimental results on the National Institute of Standards and Technology data base-4 show consistent better retrieval performance of the proposed approach compared to other continuous and exclusive fingerprint classification methods as well as minutia-based indexing schemes Xudong Jiang 0001, Manhua Liu, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2005 | Steganalysis of binary text imagesabstractWe present in this paper a technique for the steganalysis of electronic text documents. The proposed method averages the similar patterns together in a binary text image to estimate the original patterns. The proposed method can detect the existence of a secret message hidden by any boundary flipping techniques as well as estimate the message length and locate the flipped pixels. Jun Cheng 0003, Alex Chichung Kot, Jun Liu 0069 |
ICASSP (4) | 2 |
| 2005 | Data Hiding For Text Document Image Authentication by Connectivity-PreservingabstractA novel blind data hiding method for text document images is proposed; it aims to preserve the connectivity in a local neighborhood. The "flippability" of a pixel is determined by imposing three transition criteria in a 3/spl times/3 moving window which is centered at the pixel. The "embeddability" of a block is invariant in the watermark embedding process, while the "flipped" pixels can be located by imposing a constraint. The "uneven embeddability" of the host image is considered by embedding the watermark only in those "embeddable" blocks. The location is chosen in such a way that the visual quality of the watermarked image is guaranteed. Different types of blocks are employed and their abilities to increase the capacity are compared. A hard authenticator watermark is also generated to ensure the integrity and authenticity of the document. Huijuan Yang, Alex Chichung Kot |
ICASSP (2) | 2 |
| 2005 | Differential lattice decoding in noncoherent MIMOabstractWe present improved differential lattice decoding (DLD) for multi-antenna differential modulation by exploiting both basis reduction and sphere decoding. The extra complexity of DLD is shown to be worthwhile in terms of the obtained performance gain over Clarkson et al.'s decoding scheme. With roughly another fold of complexity of basis reduction, DLD augmented by a local search practically attains the maximum-likelihood decoding performance. Cong Ling 0001, Wai Ho Mow, Kwok Hung Li, Alex Chichung Kot |
ICC | 4 |
| 2005 | Detection of data hiding in binary text imagesabstractWe present in this paper a technique for the steganalysis of electronic binary text images. The proposed method utilizes the similarity between same characters or symbols. The proposed method can detect the existence of a secret message hidden by the embedding algorithms, which hide information by flipping centers of L-shape patterns (COL). Jun Cheng 0003, Alex Chichung Kot, Jun Liu 0069 |
ICIP (3) | 2 |
| 2005 | Relationships and unification of binary images data-hiding methodsabstractThere are several approaches proposed for embedding watermark data into binary images. In this paper, we discuss the relationship between these embedding schemes, in particular, the pixel-pattern method and key-weight method. We argue that a pixel-pattern method can be linked to key-weight method by selecting a set of (key, weight) matrices. A unification framework is given to show the advantages and disadvantages of these methods. Jun Liu 0069, Huijuan Yang, Alex Chichung Kot |
ICIP (1) | 3 |
| 2005 | Multiple-antenna differential lattice decodingabstractFrom a lattice viewpoint, Clarkson, Sweldens and Zheng significantly reduced the complexity of multiantenna differential decoding. Their approximate decoding algorithm, however, has not unleashed the full potential of lattice decoding. In this paper, we present several improved algorithms, generally referred to as differential lattice decoding (DLD), for multiantenna communication. We first analyze two distinct approximate DLD algorithms, and then develop an algorithm that exactly finds the closest lattice point in the Euclidean space. This exact DLD is subsequently augmented by local search to compensate for the remaining approximation. The small amount of extra complexity of the exact or augmented DLD is rewarded by a clear performance gain. We find that employing basis reduction is very effective to reduce the overall decoding complexity for high lattice dimensions. Moreover, the dimension of the lattice defined in this paper is independent of the number of receive antennas, which results in not only lower complexity, but also better performance for a multiantenna receiver. Cong Ling 0001, Wai Ho Mow, Kwok Hung Li, Alex Chichung Kot |
IEEE J. Sel. Areas Commun. | 4 |
| 2004 | Text document authentication by integrating inter character and word spaces watermarkingabstractA method for watermarking on text document images to authenticate the owner or authorized user is proposed. The proposed method makes use of the integrated inter character and word spaces for watermark embedding. An overlapping component which is of size three is utilized, whereby the relationship of the left and right spaces of the character is employed for the watermark embedding. The integrity of the document can be ensured by comparing the hash value of the character components of the document before and after watermark embedding, which can be applied to other line shifting and word-shifting methods as well. While the authenticity of the document can be ensured by generating the Gold-like sequence, which takes the secret key of the authorized user/owner as the seed value, and it is subsequently XORed with hash value of the character components of the document to generate the content-based watermark. The capacity of the watermark has increased compared with conventional line shifting and word shifting methods. Retrieval of the watermark does not require the presence of the original text document image. Discussions on the cons and pros of the method are also provided. Experimental results are given to validate the arguments. Huijuan Yang, Alex Chichung Kot |
ICME | 2 |
| 2004 | Parameter estimation of a real single tone from short data records
H. W. Fung, Alex Chichung Kot, Kwok Hung Li, Kah Chan Teh |
Signal Process. | 2 |
| 2004 | Distance-reciprocal distortion measure for binary document imagesabstractIn this letter, we present a novel objective distortion measure for binary document images. This measure is based on the reciprocal of distance that is straightforward to calculate. Our results show that the proposed distortion measure matches well to subjective evaluation by human visual perception. Haiping Lu, Alex Chichung Kot, Yun Q. Shi 0001 |
IEEE Signal Process. Lett. | 2 |
| 2004 | On decision-feedback detection of differential space-time modulation in continuous fadingabstractWe show that linear prediction (LP)-based decision-feedback detection (DFD) for nondiagonal differential space-time modulation (DSTM) may suffer from a severe performance degradation in continuously fading channels. DSTM constellations that incur no degradation in LP-DFD are identified as those with a diagonal generator. To cater to other constellations, we propose a low-complexity DFD scheme by inserting decision-feedback symbols into the metric of multiple-symbol differential detection. Cong Ling 0001, Kwok Hung Li, Alex Chichung Kot |
IEEE Trans. Commun. | 3 |
| 2004 | Errata to "Nonlinear Dynamic System Identification Using Chebyshev Functional Link Artificial Neural Networks"
Jagdish C. Patra, Alex Chichung Kot |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2003 | Decision-feedback multiple-symbol differential detection of differential space-time modulation in continuously fading channelsabstractLinear prediction (LP)-based decision-feedback differential detection (DFDD) only works for diagonal differential space-time modulation (DSTM) when fading is changing fast and continuously. For other constellations, it suffers bad performance. We propose DFDD based on multiple-symbol differential detection (MSDD) for DSTM to cope with continuous fading. A key observation is that the correlation matrix of the received signal can be expressed in terms of DSTM matrices corresponding to the sent information symbols. In this way, decision feedback can be inserted into the MSDD metric, yielding a DF-MSDD receiver while maintaining almost the same performance as MSDD. Cong Ling 0001, Kwok Hung Li, Alex Chichung Kot |
ICASSP (4) | 3 |
| 2003 | On decision-feedback detection of nondiagonal differential space-time modulation in temporally correlated fading channelsabstractExisting work on decision feedback detection (DFD) of differential space-time modulation (DTSM) were largely confined to diagonal constellations. In other works considering nondiagonal constellations, the temporal variation of fading in the DTSM supersymbol duration was ignored. In this paper, we take a close look at the DFD receiver structure for nondiagonal DSTM in temporally correlated fading. We identify the constellations of which the error performance is not influenced by ignorance of fading variation. In addition, by analyzing the error performance, we demonstrate that existing DFD, is however, fundamentally mismatched to other constellations in fast fading. A major conclusion is that diagonal design still plays an important role for the efficacy of DFD for nondiagonal constellation in fast fading. Cong Ling 0001, Kwok Hung Li, Alex Chichung Kot |
ICC | 3 |
| 2003 | Binary image watermarking through biased binarizationabstractThis paper presents a watermarking algorithm for binary images. The original binary image is blurred to a gray-level image and we embed the watermark by biasing the threshold in binarization. A loop is used to control the quality of watermarked images and robustness, and a key is generated for extraction. We employ error correction codes to reduce extraction error. This algorithm can be applied to general binary images except dithered images. Experiments show that the distortion in the watermarked image is not obtrusive and the algorithm provides some degree of robustness. Haiping Lu, Alex Chichung Kot, Susanto Rahardja |
ICME | 2 |
| 2003 | Multisampling decision-feedback linear prediction receivers for differential space-time modulation over Rayleigh fast-fading channelsabstractNovel decision-feedback (DF) linear prediction (LP) receivers, which process multiple samples per symbol interval in conjunction with optimal sample combining, are proposed for differential space-time modulation (DSTM) over Rayleigh fast-fading channels. Performance analysis demonstrates that multisampling DF-LP receivers outperform their symbol-rate sampling counterpart in fast fading substantially. In addition, an asymptotically tight upper bound on the pairwise error probability is derived. In view of this bound, the design criterion of DSTM for fast fading is the same as that for block-wise static fading. To avoid the estimation of the second-order statistics of the channel, a polynomial-model-based DF-LP receiver is proposed. It can approach the performance of the optimum DF-LP receiver at high signal-to noise ratios, provided fading is moderate. Cong Ling 0001, Kwok Hung Li, Alex Chichung Kot, Keith Q. T. Zhang |
IEEE Trans. Commun. | 3 |
| 2003 | Noncoherent sequence detection of differential space-time modulatioabstractApproximate maximum-likelihood noncoherent sequence detection (NSD) for differential space-time modulation (DSTM) in time-selective fading channels is proposed. The starting point is the optimum multiple-symbol differential detection for DSTM that is characterized by exponential complexity. By truncating the memory of the incremental metric, a finite-state trellis is obtained so that a Viterbi algorithm can be implemented to perform sequence detection. Compared to existing linear predictive receivers, a distinguished feature of NSD is that it can accommodate nondiagonal constellations in continuous fading. Error analysis demonstrates that significant improvement in performance is achievable over linear prediction receivers. By incorporating the reduced-state sequence detection techniques, performance and complexity tradeoffs can be controlled by the branch memory and trellis size. Numerical results show that most of the performance gain can be achieved by using an L-state trellis, where L is the size of the DSTM constellation. Cong Ling 0001, Kwok Hung Li, Alex Chichung Kot |
IEEE Trans. Inf. Theory | 3 |
| 2002 | Digital image-in-image watermarking for copyright protection of satellite images using the fast Hadamard transformabstractIn this paper, a robust and efficient digital image watermarking algorithm using the fast Hadamard transform (FFIT) is proposed for the copyright protection of satellite images. This algorithm can embed or hide an entire image or pattern as a watermark such as a company's logo or trademark directly into the original satellite image. The performance of the proposed algorithm is evaluated using a benchmarking tool called Stirmark. Results show that this algorithm is very robust and can survive up to 60% of all Stirmark attacks. These attacks were tested on a number of satellite test images of size 512/spl times/512/spl times/8 bit, embedded with a watermark image of size 64/spl times/64/spl times/8 bits. The simplicity of the fast Hadamard transform also offers a significant advantage in shorter processing time and ease of hardware implementation. Anthony Tung Shuen Ho, Jun Shen 0001, Soon Hie Tan, Alex Chichung Kot |
IGARSS | 4 |
| 2002 | Fast algorithm for multi-dimensional discrete Hartley transform with size ql1×ql2×...×qlr
Yonghong Zeng, Guoan Bi, Alex Chichung Kot |
Signal Process. | 3 |
| 2002 | Extending the sound impulse response of room using extrapolationabstractAn analytic method is used to extend the sound impulse response of a room. It is based on the extrapolation theory for band-limited signals. Only two matrix operations are needed for extending a discrete impulse response. Due to the approximation of the geometrical acoustic model of the observed data and the computation errors, the problem may be ill-conditioned. The optimum regularization method is used to solve the ill-posed problem. For discrete signals, the extrapolation is not unique, therefore additional constraints are used to obtain the admissible result. The method is evaluated with a set of psychoacoustic experiments. The evaluations are made on five psychological characteristics, and a general conclusion is given based on general fuzzy clustering. Zihou Meng, Kimihiro Sakagami, Masayuki Morimoto, Guoan Bi, Alex Chichung Kot |
IEEE Trans. Speech Audio Process. | 5 |
| 2002 | Nonlinear dynamic system identification using Chebyshev functional link artificial neural networksabstractA computationally efficient artificial neural network (ANN) for the purpose of dynamic nonlinear system identification is proposed. The major drawback of feedforward neural networks, such as multilayer perceptrons (MLPs) trained with the backpropagation (BP) algorithm, is that they require a large amount of computation for learning. We propose a single-layer functional-link ANN (FLANN) in which the need for a hidden layer is eliminated by expanding the input pattern by Chebyshev polynomials. The novelty of this network is that it requires much less computation than that of a MLP. We have shown its effectiveness in the problem of nonlinear dynamic system identification. In the presence of additive Gaussian noise, the performance of the proposed network is found to be similar or superior to that of a MLP. A performance comparison in terms of computational complexity has also been carried out. Jagdish C. Patra, Alex Chichung Kot |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2001 | Adaptive beamforming and power control for DS-CDMA mobile radio communicationsabstractJoint beamforming and power control is an efficient technique to increase the capacity of DS-CDMA systems. It is shown that the optimal solution of the uplink beamforming weights is independent of the combined path loss/shadowing effect. Based on this, a novel approach is proposed to implement adaptive beamforming and power control by converting the joint problem into a simple power control problem. It generates the optimal beamforming weights in each slot based on estimated channel responses. Computer simulations show that the new approach has much faster convergence for transmission power than the traditional LMS algorithm. Ying-Chang Liang, Francois P. S. Chin, Alex Chichung Kot |
ICC | 3 |
| 2000 | Blind code timing estimation for DS-CDMA signals in unknown colored noiseabstractIn previous years, the MUSIC timing estimator were proposed for direct-sequence code-division multiple access (DS-CDMA) signals. It has been generally assumed that the channel noise is temporally white. However it may be invalid in practice due to, for instance, the presence of some narrowband interference or the lumping of secondary users into with the noise. We consider blind code timing in unknown colored noise by using the matrix decomposition technique. A subspace-based code timing estimation algorithm is presented and its performance is evaluated by computer simulation and compared with the MUSIC method. It is proved that the proposed timing estimation algorithm provides better performance than the MUSIC method in colored noise. Yugang Ma, Kwok Hung Li, Alex Chichung Kot, G. T. Ye |
GLOBECOM | 3 |
| 2000 | DS-CDMA synchronization with dual-antenna in unknown correlated noiseabstractWe consider the propagation delay estimation of a direct-sequence code-division multiple-access (DS-CDMA) system in unknown correlated noise. The new propagation delay estimation method proposed here is based on generalized correlation decomposition (GCD). First, the received signal with a dual-antenna through a flat-fading channel is modeled. Then, two GCD-based propagation delay estimators: dual-antenna singular value decomposition (SVD) and dual-antenna canonical correlation decomposition (CCD) estimators are presented and compared with common MUSIC estimator. Simulation results show that the proposed estimators, especially the dual-antenna CCD estimator, are robust against the near-far problem and outperform the common MUSIC estimator in unknown correlated noise. Yugang Ma, G. T. Ye, Kwok Hung Li, Alex Chichung Kot |
GLOBECOM | 4 |
| 2000 | A novel approach to FM interference suppression in DS spread-spectrum communication systemsabstractA new recursive technique for the excision of FM interference in DS spread-spectrum communication system is introduced. The proposed algorithm tabulates the jammer-corrupted received samples as a matrix to model a time-varying process. Input vectors are sent into the matrix to produce upon multiplication output vectors from which the frequency information in the matrix is deduced. The scheme resembles the familiar closed-loop frequency response estimation procedures in classical control theories. The estimated frequency then tunes an adaptive notch filter for interference suppression. The method provides more accurate IF estimates than given by the Wigner-Ville distribution and a performance close to notch filtering at the exact IF at a significantly lower computational requirement. H. W. Fung, Alex Chichung Kot, Kwok Hung Li |
ICASSP | 2 |
| 2000 | A Floating Feature Detector for Handwritten Numeral RecognitionabstractA novel feature extraction method for handwritten numeral recognition is proposed based on character's geometric structures. A group of stable and reliable global features are defined and extracted. Furthermore, a floating feature detector is proposed to detect and extract tiny segments as fine features. A neural network is employed as the recognisor to conduct experiments on evaluating the feasibility of the new approach. This proposed method demonstrates that the combination of fine features with global features can greatly improve the handwritten character recognition rate compared to those using global features only. Lihui Chen 0001, Alex Chichung Kot |
ICPR | 3 |
| 2000 | Chebyschev functional link artificial neural networks for nonlinear dynamic system identificationabstractAn alternative novel artificial neural network (ANN) for the purpose of dynamic nonlinear system identification is proposed. The main drawback of feedforward neural networks such as a multi-layer perceptron (MLP) trained with backpropagation (BP) algorithm is that it requires a large amount of computation and the rate of error convergence is slow. The proposed Chebyschev functional link ANN (C-FLANN) is found to have much less computational requirement and its performance is found to be superior to that of a MLP for the complex task of nonlinear dynamic system identification, even in the case of additive input noise to the system. Jagdish C. Patra, Alex Chichung Kot, Yan Qiu Chen |
SMC | 2 |
| 2000 | An intelligent pressure sensor with self-calibration capability using artificial neural networksabstractThe nonlinear response characteristics of a capacitive pressure sensor (CPS) changes when the ambient temperature changes widely. In such conditions, the calibration becomes difficult, and to obtain an accurate pressure readout, appropriate compensation of the CPS characteristics is needed. We propose an intelligent CPS using artificial neural networks (ANNs) to provide self-calibration and compensation. The proposed ANN model can provide automatic nonlinear compensation and calibration of the CPS characteristics. A microcontroller unit (MCU) based implementation scheme for this model is also considered. Simulation results show that this model can estimate the pressure with a maximum full-scale error of /spl plusmn/1% over a variation of temperature from -50 to 150/spl deg/C. Santanu Kumar Rath, Jagdish C. Patra, Alex Chichung Kot |
SMC | 3 |
| 2000 | A novel hybrid classifier for recognition of handwritten numeralsabstractA hybrid neural network and tree classification system for handwritten numeral recognition is proposed. The recognition system consists of coarse and fine classification based on a variety of stable and reliable global features and local features. For the coarse classifier: a four-layer feedforward neural network with backpropagation learning algorithm is employed to distinguish six subsets {0}, {6}, {8}, {1,7}, {4,9}, {2,3,5} based on the similarity of character's geometrical features. Three character classes {0}, {6} and {8} are directly recognized from ANN. For each of the last three subsets, a decision tree classifier is built for fine classification as follows: firstly, the specific feature-class relationship is heuristically and empirically created between the feature primitives and corresponding semantic class. Then, an iterative growing and pruning algorithm is used to form a tree classifier. Experiments demonstrated that the proposed hybrid recognition system is robust and flexible, which can achieve a high recognition rate. Lihui Chen 0001, Alex Chichung Kot |
SMC | 3 |
| 2000 | Error probabilities of an FFH/BFSK self-normalizing receiver in a Rician fading channel with multitone jammingabstractWe derive the analytical bit-error rate (BER) expressions for a fast frequency-hopped binary frequency-shift keying self-normalizing receiver over a fading channel with the worst-case band multitone jamming (MTJ) and additive white Gaussian noise (AWGN). The desired signal and MTJ are assumed to undergo independent Rician fading and our analyses, validated with simulation results, show that the system performance is not sensitive to different types of MTJ fading conditions. The self-normalizing receiver is found to be superior to the linear-combining receiver when the signal amplitude does not experience severe fading, while the converse is true under Rayleigh fading signal conditions. Under a Rician fading channel and AWGN conditions, the worst-case MTJ and the worst-case partial-band noise jamming are shown to have similar effects on the BER performance of the self-normalizing receiver with diversity. Kah Chan Teh, Alex Chichung Kot, Kwok Hung Li |
IEEE Trans. Commun. | 2 |
| 1999 | Error probabilities and performance comparisons of FFH/BFSK receivers with multitone jamming and AWGNabstractThis paper studies the bit-error rate (BER) performance of a fast frequency-hopped (FFH) binary frequency-shift-keying (BFSK) clipper receiver in the presence of multitone jamming (MTJ) and additive white Gaussian noise (AWGN). By using the Taylor-series expansion and the quantization approach, the BER expressions for higher diversity levels can be obtained without much extra computational complexity. The analytical BER results, validated by simulations, show that there is an optimum diversity level for the clipper receiver. Performance comparisons among various receivers demonstrate that the BER performance of the clipper receiver is significantly better than that of the linear-combining receiver. In addition, the clipper receiver also outperforms the product-combining receiver and the self-normalizing receiver provided that the clipping threshold is set at the desired signal power level. Kah Chan Teh, Alex Chichung Kot, Kwok Hung Li |
ICASSP | 2 |
| 1999 | Performance of antenna diversity reception with correlated Rayleigh fading signalsabstractThis paper studies the performance of antenna diversity reception with selection combining (SC) and maximal ratio combining (MRC) in a correlated Rayleigh fading environment. Simple PDF expressions are derived for the output SNR of the SC and MRC diversity receivers. Based on the PDF, the average BERs for DPSK and NCFSK modulation are calculated to show the effects of the correlation. Moreover, the relationship between the average BER and the diversity antennas' separation is given. The presented numerical results show that SC and MRC diversity reception can achieve a good performance only if the distance between two diversity antennas is larger than 0.3/spl lambda/. Liquan Fang, Guoan Bi, Alex Chichung Kot |
ICC | 3 |
| 1999 | Performance study of an FFH/BFSK clipper receiver with multitone jamming over a Rician-fading channelabstractWe study the bit-error rate performance of a fast frequency-hopped binary frequency-shift-keying (FFH/BFSK) spread spectrum (SS) communication system over fading channels with the worst-case band multitone jamming (MTJ) and additive white Gaussian noise (AWGN). The FFH system employing a clipper receiver is investigated and the corresponding bit-error expressions are obtained using the quantization approach. The desired signal and MTJ are assumed to undergo independent fading and our analysis, validated with simulation results, shows that the performance of the system is not sensitive to a particular type of MTJ fading model. Kah Chan Teh, Alex Chichung Kot, Kwok Hung Li |
ICC | 2 |
| 1999 | Performance study of a maximum-likelihood receiver for FFH/BFSK systems with multitone jammingabstractWe derive the optimum structure of a maximum-likelihood (ML) receiver for a fast frequency-hopped binary frequency-shift-keying (FFH/BFSK) spread-spectrum (SS) communication system operating in the presence of multitone jamming (MTJ) and additive white Gaussian noise (AWGN). It is shown that the side information of noise variance, signal tone amplitude, and multiple interfering tone amplitude at each hop, as well as the computation of nonlinear modified Bessel function are required to implement the optimum ML receiver. We have also derived and analyzed two suboptimum receivers-namely, the ML-I and ML-II receivers-for large and small signal-to-noise ratio (SNR), respectively. Performance comparisons among various receivers show that the ML receiver gives the best performance, while the ML-I and ML-II receivers also outperform the other existing methods under both high and low SNR conditions. Kah Chan Teh, Alex Chichung Kot, Kwok Hung Li |
IEEE Trans. Commun. | 2 |
| 1998 | Multitone jamming rejection of FFH/BFSK spread-spectrum system over fading channelsabstractAnalytical expressions for bit-error probability are derived for a fast frequency-hopping binary frequency-shift keying (FFH/BFSK) spread-spectrum communication system over a fading channel with worst-case band multitone jamming (MTJ) and additive white Gaussian noise (AWGN). An FFH system employing either a linear-combining receiver or a clipper receiver is investigated. The desired signal and MTJ are assumed to undergo independent fading, and our analysis, validated with simulation results, shows that the performance of the system is slightly improved as the severity of the MTJ fading is increased. The clipper receiver is found to be superior to the linear-combining receiver when the jamming power is strong. The worst-case MTJ is shown to be more harmful than the corresponding worst-case partial-band noise jamming over a fading channel with AWGN. Kah Chan Teh, Alex Chichung Kot, Kwok Hung Li |
IEEE Trans. Commun. | 2 |
| 1997 | A new adaptive neural network multiuser detector in synchronous CDMA systemsabstractA new adaptive neural network multiuser detector is proposed and investigated for synchronous code-division multiple-access (CDMA) systems. The proposed multiuser detector includes two parts: a decorrelating detector as its auxiliary detector and an adaptive multiple layer perceptron (MLP) detector as its main detector. At the setup stage, the auxiliary detector detects the user's transmitted data and at the same time feeds these output data to the main MLP detector as its training data, and the main detector is trained by using the well-known backpropagation (BP) algorithm. After the training process, the auxiliary detector stops work and the main detector starts detecting the user's transmitted data self-adaptively. The proposed detector is blind and can provide near minimum bit-error-rate performance. Guiqing He, Alex Chichung Kot |
ICASSP | 2 |
| 1996 | Error-based constraints for efficient learning of pattern recognition problem
Haroon Atique Babri, Alex Chichung Kot, S. L. Tay, S. P. Ngian |
Pattern Recognit. Lett. | 2 |
| 1994 | A MUSIC approach for estimation of directions of arrival of multiple narrowband and broadband sources
Boon Poh Ng, Meng Hwa Er, Alex Chichung Kot |
Signal Process. | 3 |
| 1990 | Performance analysis of a high resolution time delay estimation algorithmabstractThe convergence properties of an iterative algorithm for high-resolution time delay estimation are studied for a specific example. It is shown that the functional relationship between successive iterates is a contraction mapping for high signal-to-noise ratio. The contraction mapping condition is shown by deriving the derivative of the functional relationship and then calculating the mean and variance of the derivative. The analytical computation is validated by computer simulation.> Ivars P. Kirsteins, Alex Chichung Kot |
ICASSP | 2 |
| 1989 | Analysis of estimation of signal parameters by linear-prediction at high SNR using matrix approximationabstractA greatly simplified calculation of the accuracy of estimates of the parameters of exponential signals is introduced. The authors assume that the estimates are obtained by first estimating the coefficients of a prediction-error filter (PEF) that arises in signal modeling by linear prediction. A simplified approach based on matrix approximation is used to calculate the statistics of the errors in the estimated signal parameters at high SNR. This approach provides insights for design and tractable formulas even for the case of multicomponent signals and Hankel or Toeplitz data matrices. The analyses are verified by simulation examples that include single- and multiple-component signals consisting of damped or undamped complex exponential signals in noise.> Donald W. Tufts, R. J. Vacarro, Alex Chichung Kot |
ICASSP | 3 |
| 1988 | The threshold analysis of SVD-based algorithmsabstractThe problem of analyzing the threshold effect of signal processing algorithms which use the singular-value decomposition (SVD) is addressed. The probability of obtaining an outlier is calculated and used to determine the threshold SNR at which the variance of parameter estimation errors depart from Cramer-Rao bound behavior. Simulation results using low rank approximation and linear prediction for frequency estimation verify the analysis. The same method of analysis can be applied to a broad class of parameter-estimation methods in which the principal-component technique or low rank approximations to matrices are used.> Donald W. Tufts, Alex Chichung Kot, Richard J. Vaccaro |
ICASSP | 2 |
| 1987 | The statistical performance of state-variable balancing and Prony's method in parameter estimationabstractThis paper presents a statistical analysis of the state-variable balancing and Prony methods for estimating the parameters of exponential signals in the presence of additive noise. The case of frequency estimation for a single sinusoid is carried out in detail. Analytical expressions for the variances of the frequency estimates at high signal-to-noise ratios are derived. The calculated variances are compared to the Cramer-Rao bound. The results are validated by simulations over a wide range of signal-to-noise ratios. The analysis and simulations show that the state-variable balancing method can provide slightly more accurate frequency estimates while avoiding the problem of selecting the signal zeros of the Prony polynomial. Alex Chichung Kot, Sarangarajan Parthasarathy, Donald W. Tufts, Richard J. Vaccaro |
ICASSP | 1 |
| 1987 | A perturbation theory for the analysis of SVD-based algorithmsabstractThe problem of statistically analyzing the performance of signal processing algorithms which use the singular value decomposition is adressed in this paper. Such decomposition, which is widely used in system identification and parameter estimation, is a non-linear operation. Consequently, when applied to random data, statistical results are extremely difficult to obtain. The first-order Taylor series expansion is generally used in computing the statistics but the derivative term makes the staistical analysis very difficult. In this paper, a power-like method which results in a simple expression is proposed. The singular vector perturbation using both approaches for the case of low rank approximation is examined. Richard J. Vaccaro, Alex Chichung Kot |
ICASSP | 2 |