Chun Pong Lau 0001

dblp:129/2661-1 · DBLP profile ↗
← Back
19ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0003-3748-4160ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 10 since 2021Security and privacy · 5 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Segment and inpaint: Defending object detectors with a unified segment and inpaint paradigm
Jiaming Guo, Chun Pong Lau 0001
Pattern Recognit.2
2026 Artwork protection against unauthorized neural style transfer and aesthetic color distance metric
Zhongliang Guo 0001, Yifei Qian, Shuai Zhao 0007, Junhao Dong 0001, Ognjen Arandjelovic, Lei Fang 0001, Chun Pong Lau 0001
Pattern Recognit.8
2026 DiffProtect: Generative adversarial examples using diffusion models for facial privacy protection
abstract
The increasingly pervasive facial recognition (FR) systems raise serious concerns about personal privacy, especially for billions of users who have publicly shared their photos on social media. To address this challenge, several adversarial attack methods have been proposed to protect individuals from being identified by unauthorized FR systems with perturbed facial images. However, these approaches suffer from poor visual quality or low attack success rates, which limit their practical utility. Recently, diffusion models have achieved tremendous success in image generation. In this work, we ask: can diffusion models be used to generate adversarial examples against FR systems to improve both visual quality and attack performance? We propose DiffProtect, a novel method leveraging a diffusion autoencoder to generate semantically meaningful perturbations on FR systems. Extensive experiments demonstrate that DiffProtect produces more natural-looking encrypted images than state-of-the-art methods while achieving significantly higher attack success rates, e.g. , 24.5 % and 25.1 % absolute improvements on the CelebA-HQ and FFHQ datasets. We further evaluate the effectiveness of DiffProtect in the real world using a commercial FR API and validate its usefulness in practice through a user study. Our code is available at https://github.com/joellliu/DiffProtect .
Jiang Liu 0014, Chun Pong Lau 0001, Zhongliang Guo 0001, Yuxiang Guo 0001, Zhao-Yang Wang, Rama Chellappa
Pattern Recognit.2
2025 Instant Adversarial Purification with Adversarial Consistency Distillation
abstract
Neural networks have revolutionized numerous fields with their exceptional performance, yet they remain susceptible to adversarial attacks through subtle perturbations. While diffusion-based purification methods like DiffPure offer promising defense mechanisms, their computational overhead presents a significant practical limitation. In this paper, we introduce One Step Control Purification (OSCP), a novel defense framework that achieves robust adversarial purification in a single Neural Function Evaluation (NFE) within diffusion models. We propose Gaussian Adversarial Noise Distillation (GAND) as the distillation objective and Controlled Adversarial Purification (CAP) as the inference pipeline, which makes OSCP demonstrate remarkable effi-ciency while maintaining defense efficacy. Our proposed GAND addresses a fundamental tension between consistency distillation and adversarial perturbation, bridging the gap between natural and adversarial manifolds in the latent space, while remaining computationally efficient through Parameter-Efficient Fine-Tuning (PEFT) methods such as LoRA, eliminating the high computational budget request from full parameter fine-tuning. The CAP guides the purifi-cation process through the unlearnable edge detection operator calculated by the input image as an extra prompt, effectively preventing the purified images from deviating from their original appearance when large purification steps are used. Our experimental results on ImageNet showcase OSCP’s superior performance, achieving a 74.19% defense success rate with merely 0.1s per purification — a 100-fold speedup compared to conventional approaches.
Chun Tong Lei, Hon Ming Yam, Zhongliang Guo 0001, Yifei Qian, Chun Pong Lau 0001
CVPR5
2025 T2ICount: Enhancing Cross-modal Understanding for Zero-Shot Counting
abstract
Zero-Shot object counting aims to count instances of arbitrary object categories specified by text descriptions. Existing methods typically rely on vision-language models like CLIP, but often exhibit limited sensitivity to text prompts. We present T21 Count, a diffusion-based framework that lever-ages rich prior knowledge and fine-grained visual understanding from pretrained diffusion models. While one-step demising ensures efficiency, it leads to weakened text sensitivity. To address this challenge, we propose a Hierarchical Semantic Correction Module that progressively refines text-image feature alignment, and a Representational Regional Coherence Loss that provides reliable supervision signals by leveraging the cross-attention maps extracted from the demising U-Net. Furthermore, we observe that current benchmarks mainly focus on majority objects in images, potentially masking models' text sensitivity. To address this, we contribute a challenging re-annotated subset of FSC147 for better evaluation of text-guided counting ability. Extensive experiments demonstrate that our method achieves superior performance across different benchmarks. Code is available at https://github.com/chal5yq/T2lCount.
Yifei Qian, Zhongliang Guo 0001, Bowen Deng 0006, Chun Tong Lei, Shuai Zhao 0007, Chun Pong Lau 0001, Xiaopeng Hong, Michael P. Pound
CVPR6
2025 A Gray-Box Attack Against Latent Diffusion Model-Based Image Editing by Posterior Collapse
abstract
Recent advancements in Latent Diffusion Models (LDMs) have revolutionized image synthesis and manipulation, raising significant concerns about data misappropriation and intellectual property infringement. While adversarial attacks have been extensively explored as a protective measure against such misuse of generative AI, current approaches are severely limited by their heavy reliance on model-specific knowledge and substantial computational costs. Drawing inspiration from the posterior collapse phenomenon observed in VAE training, we propose the Posterior Collapse Attack (PCA), a novel framework for protecting images from unauthorized manipulation. Through comprehensive theoretical analysis and empirical validation, we identify two distinct collapse phenomena during VAE inference: diffusion collapse and concentration collapse. Based on this discovery, we design a unified loss function that can flexibly achieve both types of collapse through parameter adjustment, each corresponding to different protection objectives in preventing image manipulation. Our method significantly reduces dependence on model-specific knowledge by requiring access to only the VAE encoder, which constitutes less than 4% of LDM parameters. Notably, PCA achieves prompt-invariant protection by operating on the VAE encoder before text conditioning occurs, eliminating the need for empty prompt optimization required by existing methods. This minimal requirement enables PCA to maintain adequate transferability across various VAE-based LDM architectures while effectively preventing unauthorized image editing. Extensive experiments show PCA outperforms existing techniques in protection effectiveness, computational efficiency (runtime and VRAM), and generalization across VAE-based LDM variants. Our code is available at https://github.com/ZhongliangGuo/PosteriorCollapseAttack.
Zhongliang Guo 0001, Chun Tong Lei, Lei Fang 0001, Shuai Zhao 0007, Yifei Qian, Zeyu Wang 0010, Cunjian Chen, Ognjen Arandjelovic, Chun Pong Lau 0001
IEEE Trans. Inf. Forensics Secur.10
2024 Identifying Attack-Specific Signatures in Adversarial Examples
abstract
The adversarial attack literature contains numerous algorithms for crafting perturbations which manipulate neural network predictions. Many of these adversarial attacks optimize inputs with the same constraints and have similar downstream impact on the models they attack. In this work, we first show how to reconstruct an adversarial perturbation, namely the difference between an adversarial example and the original natural image, from an adversarial example. Then, we classify reconstructed adversarial perturbations based on the algorithm that generated them. This pipeline, REDRL, can detect the attack algorithm used to generate a sample from only the sample itself. The ability to determine which algorithm generated an example implies that different attack algorithms actually produce unique signatures in their adversarial examples.
Hossein Souri, Pirazh Khorramshahi, Chun Pong Lau 0001, Micah Goldblum, Rama Chellappa
ICASSP3
2024 Distillation-guided Representation Learning for Unconstrained Gait Recognition
abstract
Gait recognition holds the promise of robustly identifying subjects based on walking patterns instead of appearance information. While previous approaches have performed well for curated indoor data, they tend to underperform in unconstrained situations, e.g. in outdoor, long distance scenes, etc. We propose a framework, termed GAit DEtection and Recognition (GADER), for human authentication in challenging outdoor scenarios. Specifically, GADER leverages a Double Helical Signature to detect segments that contain human movement and builds discriminative features through a novel gait recognition method, where only frames containing gait information are used. To further enhance robustness, GADER encodes viewpoint information in its architecture, and distills representation from an auxiliary RGB recognition model, which enables GADER to learn from silhouette and RGB data at training time. At test time, GADER only infers from the silhouette modality. We evaluate our method on multiple State-of-The-Arts(SoTA) gait baselines and demonstrate consistent improvements on indoor and outdoor datasets, especially with a significant 25.2% improvement on unconstrained, remote gait data.
Yuxiang Guo 0001, Siyuan Huang 0005, Ram Prabhakar, Chun Pong Lau 0001, Rama Chellappa, Cheng Peng 0008
IJCB4
2024 HyperGait: A Video-based Multitask Network for Gait Recognition and Human Attribute Estimation at Range and Altitude
abstract
Gait recognition is one of the mainstream approaches for identifying individuals when face information is not available. Most previous methods achieve good performance on structured indoor walking sequences with silhouettes provided. However, when these methods are applied to unconstrained outdoor sequences, a significant reduction in performance is inevitably observed due to factors such as turbulence, occlusion, view angle, and oversized clothing. To make gait recognition methods stable and effective for real-world settings, we extend gait-only-based approaches by introducing more useful biometric information such as gender, age, height, weight, and body mass index to cooperatively work with the gait recognition module. In this paper, we propose a video-based multitasking network for gait recognition and human attribute prediction at ranges of up to 1000 meters and high-pitch angles to mutually improve the robustness and accuracy of each task. Through a series of experiments on OU-MVLP and BRIAR datasets, we show that our multitasking network outperforms previous methods and provides more useful biometric information for human identification tasks.
Zhao-Yang Wang, Jiang Liu 0014, Ram Prabhakar Kathirvel, Chun Pong Lau 0001, Rama Chellappa
IJCB4
2024 Diffuse and Restore: A Region-Adaptive Diffusion Model for Identity-Preserving Blind Face Restoration
abstract
Blind face restoration (BFR) from severely degraded face images in the wild is a highly ill-posed problem. Due to the complex unknown degradation, existing generative works typically struggle to restore realistic details when the input is of poor quality. Recently, diffusion-based approaches were successfully used for high-quality image synthesis. But, for BFR, maintaining a balance between the fidelity of the restored image and the reconstructed identity information is important. Minor changes in certain facial regions may alter the identity or degrade the perceptual quality. With this observation, we present a conditional diffusion-based framework for BFR. We alleviate the drawbacks of existing diffusion-based approaches and design a region-adaptive strategy. Specifically, we use an identity preserving conditioner network to recover the identity information from the input image as much as possible and use that to guide the reverse diffusion process, specifically for important facial locations that contribute the most to the identity. This leads to a significant improvement in perceptual quality as well as face-recognition scores over existing GAN and diffusion-based restoration models. Our approach achieves superior results to prior art on a range of real and synthetic datasets, particularly for severely degraded face images.
Maitreya Suin, Nithin Gopalakrishnan Nair, Chun Pong Lau 0001, Vishal M. Patel, Rama Chellappa
WACV3
2023 Multi-Modal Human Authentication Using Silhouettes, Gait and RGB
abstract
Whole-body-based human authentication is a promising approach for remote biometrics scenarios. Current literature focuses on either body recognition based on RGB images or gait recognition based on body shapes and walking patterns; both have their advantages and drawbacks. In this work, we propose Dual-Modal Ensemble (DME), which combines both RGB and silhouette data to achieve more robust performances for indoor and outdoor whole-body based recognition. Within DME, we propose GaitPattern, which is inspired by the double helical gait pattern used in traditional gait analysis. The GaitPattern contributes to robust identification performance over a large range of viewing angles. Extensive experimental results on the CASIA-B dataset demonstrate that the proposed method outperforms state-of-the-art recognition systems. We also provide experimental results using the newly collected BRIAR dataset.
Yuxiang Guo 0001, Cheng Peng 0008, Chun Pong Lau 0001, Rama Chellappa
FG3
2023 ATDetect: Face Detection and Keypoint Extraction at Range and Altitude
abstract
Face detection and alignment are the crucial preprocessing steps in face recognition. While face detection works well in ideal situations, the performance deteriorates significantly when the image is degraded, due to factors such as blur, deformation, low resolution, and extreme headpose. However, there are very few works on face detection with realistic data captured from a long range (100m to 500m) and high altitude (30° to 50° pitch angle). We first evaluated several state-of-the-art methods on data collected at ranges of 100-500 meters and large pitch angles. One challenge is videos captured from long ranges usually lack bounding boxes and keypoint annotations, needed for training deep networks. This motivates us to develop a face detection and alignment algorithm that could perform effectively on videos captured from a long range and high altitude without groundtruth annotations. Moreover, meta information such as age, gender, and headpose of the subject could help face recognition. Therefore, we propose a single-stage face localization model ATDetect, which detects face bounding boxes, keypoints, and meta information simultaneously with realistic video captured at range and altitude.
Chun Pong Lau 0001, Maitreya Suin, Rama Chellappa
IJCB1
2023 Multi-View Action Recognition using Contrastive Learning
abstract
In this work, we present a method for RGB-based action recognition using multi-view videos. We present a supervised contrastive learning framework to learn a feature embedding robust to changes in viewpoint, by effectively leveraging multi-view data. We use an improved supervised contrastive loss and augment the positives with those coming from synchronized viewpoints. We also propose a new approach to use classifier probabilities to guide the selection of hard negatives in the contrastive loss, to learn a more discriminative representation. Negative samples from confusing classes based on posterior are weighted higher. We also show that our method leads to better domain generalization compared to the standard supervised training based on synthetic multi-view data. Extensive experiments on real (NTU-60, NTU-120, NUMA) and synthetic (RoCoG) data demonstrate the effectiveness of our approach.
Ketul Shah, Anshul Shah 0001, Chun Pong Lau 0001, Celso de Melo, Rama Chellappa
WACV3
2023 Interpolated Joint Space Adversarial Training for Robust and Generalizable Defenses
abstract
Adversarial training (AT) is considered to be one of the most reliable defenses against adversarial attacks. However, models trained with AT sacrifice standard accuracy and do not generalize well to unseen attacks. Recent works show generalization improvement with adversarial samples under unseen threat models such as on-manifold threat model or neural perceptual threat model. However, the former requires exact manifold information while the latter requires algorithm relaxation. Motivated by these considerations, we propose a novel threat model called Joint Space Threat Model (JSTM), which exploits the underlying manifold information with Normalizing Flow, ensuring that the exact manifold assumption holds. Under JSTM, we develop novel adversarial attacks and defenses. Specifically, we propose the Robust Mixup strategy in which we maximize the adversity of the interpolated images and gain robustness and prevent overfitting. Our experiments show that Interpolated Joint Space Adversarial Training (IJSAT) achieves good performance in standard accuracy, robustness, and generalization. IJSAT is also flexible and can be used as a data augmentation method to improve standard accuracy and combined with many existing AT approaches to improve robustness. We demonstrate the effectiveness of our approach on three benchmark datasets, CIFAR-10/100, OM-ImageNet and CIFAR-10-C.
Chun Pong Lau 0001, Jiang Liu 0014, Hossein Souri, Wei-An Lin, Soheil Feizi, Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Segment and Complete: Defending Object Detectors against Adversarial Patch Attacks with Robust Patch Detection
abstract
Object detection plays a key role in many security-critical systems. Adversarial patch attacks, which are easy to implement in the physical world, pose a serious threat to state-of-the-art object detectors. Developing reliable defenses for object detectors against patch attacks is critical but severely understudied. In this paper, we propose Segment and Complete defense (SAC), a general framework for defending object detectors against patch attacks through detection and removal of adversarial patches. We first train a patch segmenter that outputs patch masks which provide pixel-level localization of adversarial patches. We then propose a self adversarial training algorithm to robustify the patch segmenter. In addition, we design a robust shape completion algorithm, which is guaranteed to remove the entire patch from the images if the outputs of the patch segmenter are within a certain Hamming distance of the ground-truth patch masks. Our experiments on COCO and xView datasets demonstrate that SAC achieves superior robustness even under strong adaptive attacks with no reduction in performance on clean images, and generalizes well to unseen patch shapes, attack budgets, and unseen attack methods. Furthermore, we present the APRICOT-Mask dataset, which augments the APRICOT dataset with pixel-level annotations of adversarial patches. We show SAC can significantly reduce the targeted attack success rate of physical patch attacks. Our code is available at https://github.com/joellliu/SegmentAndComplete.
Jiang Liu 0014, Alexander Levine 0001, Chun Pong Lau 0001, Rama Chellappa, Soheil Feizi
CVPR3
2022 Mutual Adversarial Training: Learning Together is Better Than Going Alone
abstract
Recent studies have shown that robustness to adversarial attacks can be transferred across deep neural networks. In other words, we can make a weak model more robust with the help of a strong teacher model. In this paper, we ask if models can “learn together” and “teach each other” to achieve better robustness instead of learning from a static teacher.We study how interactions among models affect robustness via knowledge distillation. We propose mutual adversarial training (MAT), in which multiple models are trained together and share the knowledge of adversarial examples to achieve improved robustness. MAT allows robust models to explore a larger space of adversarial samples and find more robust feature spaces and decision boundaries. Through extensive experiments on the CIFAR-10, CIFAR-100, and mini-ImageNet datasets, we demonstrate that MAT can effectively improve model robustness and outperform state-of-the-art methods under white-box attacks. In addition, we show that MAT can also mitigate the robustness trade-off among different perturbation types. Specially, we train specialist models that learn to defend a specific perturbation type and a generalist model that learns to defend multiple perturbation types by learning from the specialists, which brings as much as 13.4% accuracy gain to AT baselines against the union ofl∞,l2, andl1 attacks. Our results show the superiority of the proposed method and demonstrate that collaborative learning is an effective strategy for designing robust models.
Jiang Liu 0014, Chun Pong Lau 0001, Hossein Souri, Soheil Feizi, Rama Chellappa
IEEE Trans. Inf. Forensics Secur.2
2020 ATFaceGAN: Single Face Image Restoration and Recognition from Atmospheric Turbulence
abstract
Image degradation due to atmospheric turbulence is common while capturing images at long ranges. To mitigate the degradation due to turbulence which includes deformation and blur, we propose a generative single frame restoration algorithm which disentangles the blur and deformation due to turbulence and reconstructs a restored image. The disentanglement is achieved by decomposing the distortion due to turbulence into blur and deformation components using deblur generator and deformation correction generator respectively. Two paths of restoration are implemented to regularize the disentanglement and generate two restored images from one degraded image. A fusion function combines the features of the restored images to reconstruct a sharp image with rich details. Adversarial and perceptual losses are added to reconstruct a sharp image and suppress the artifacts respectively. Extensive experiments demonstrate the effectiveness of the proposed restoration algorithm, which achieves satisfactory performance in face restoration and face recognition.
Chun Pong Lau 0001, Hossein Souri, Rama Chellappa
FG1
2020 Dual Manifold Adversarial Robustness: Defense against Lp and non-Lp Adversarial Attacks
abstract
Adversarial training is a popular defense strategy against attack threat models with bounded Lp norms. However, it often degrades the model performance on normal images and more importantly, the defense does not generalize well to novel attacks. Given the success of deep generative models such as GANs and VAEs in characterizing the underlying manifold of images, we investigate whether or not the aforementioned deficiencies of adversarial training can be remedied by exploiting the underlying manifold information. To partially answer this question, we consider the scenario when the manifold information of the underlying data is available. We use a subset of ImageNet natural images where an approximate underlying manifold is learned using StyleGAN. We also construct an ``On-Manifold ImageNet'' (OM-ImageNet) dataset by projecting the ImageNet samples onto the learned manifold. For OM-ImageNet, the underlying manifold information is exact. Using OM-ImageNet, we first show that on-manifold adversarial training improves both standard accuracy and robustness to on-manifold attacks. However, since no out-of-manifold perturbations are realized, the defense can be broken by Lp adversarial attacks. We further propose Dual Manifold Adversarial Training (DMAT) where adversarial perturbations in both latent and image spaces are used in robustifying the model. Our DMAT improves performance on normal images, and achieves comparable robustness to the standard adversarial training against Lp attacks. In addition, we observe that models defended by DMAT achieve improved robustness against novel attacks which manipulate images by global color shifts or various types of image filtering. Interestingly, similar improvements are also achieved when the defended models are tested on (out-of-manifold) natural images. These results demonstrate the potential benefits of using manifold information in enhancing robustness of deep learning models against various types of novel adversarial attacks.
Wei-An Lin, Chun Pong Lau 0001, Alexander Levine 0001, Rama Chellappa, Soheil Feizi
NeurIPS2
2018 Image Retargeting via Beltrami Representation
abstract
Image retargeting aims to resize an image to one with a prescribed aspect ratio. Simple scaling inevitably introduces unnatural geometric distortions on the important content of the image. In this paper, we propose a simple and yet effective method to resize an image, which preserves the geometry of the important content, using the Beltrami representation. Our algorithm allows users to interactively label content regions as well as line structures. Image resizing can then be achieved by warping the image by an orientation-preserving bijective warping map with controlled distortion. The warping map is represented by its Beltrami representation, which captures the local geometric distortion of the map. By carefully prescribing the values of the Beltrami representation, images with different complexity can be effectively resized. Our method does not require solving any optimization problems and tuning parameters throughout the process. This results in a simple and efficient algorithm to solve the image retargeting problem. Extensive experiments have been carried out, which demonstrate the efficacy of our proposed method.
Chun Pong Lau 0001, Chun Pang Yung, Lok Ming Lui
IEEE Trans. Image Process.1