VLDB 2026 Research / reviewers in the wild / expert
Adams Wai-Kin Kong
dblp:16/3792
· DBLP profile ↗
91ranked-venue papers
14as first author
37since 2021 · last 2026
0000-0002-9728-9511ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 55 · 10 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 45 · 4 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 5 since 2021Security and privacy · 9 · 1 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MACS: Multi-source Audio-to-image Generation with Contextual Significance and Semantic AlignmentabstractPropelled by the breakthrough in deep generative models, audio-to-image generation has emerged as a pivotal cross-modal task that converts complex auditory signals into rich visual representations. However, previous works only focus on single-source audio inputs for image generation, ignoring the multi-source characteristic in natural auditory scenes, thus limiting the performance in generating comprehensive visual content. To bridge this gap, we propose a method called MACS to conduct multi-source audio-to-image generation. To our best knowledge, this is the first work that explicitly separates multi-source audio to capture the rich audio components before image generation. MACS is a two-stage method. In the first stage, multi-source audio inputs are separated by a weakly supervised method, where the audio and text labels are semantically aligned by casting into a common space using the large pre-trained CLAP model. We introduce a ranking loss to consider the contextual significance of the separated audio signals. In the second stage, effective image generation is achieved by mapping the separated audio signals to the generation condition using only a trainable adapter and a MLP layer. We preprocess the LLP dataset as the first full multi-source audio-to-image generation benchmark. The experiments are conducted on multi-source, mixed-source, and single-source audio-to-image generation tasks. The proposed MACS outperforms the current state-of-the-art methods in 17 out of the 21 evaluation indexes on all tasks and delivers superior visual quality. Xiaobao Guo, Yuzhe Zhu, Adams Wai-Kin Kong |
AAAI | 4 |
| 2025 | Decreasing Word Error Rates in Paragraph Handwritten Text Recognition with Synthetic DataabstractHandwritten Text Recognition (HTR) faces a persistent challenge with the scarcity of data at the paragraph level, arising from the difficulty of acquiring diverse, cost-efficient, and cleanly labeled datasets for training. As such, works in HTR leverage segmentation, regularization techniques, and language modeling to excel in a low-data environment. While synthetic generation methods gain traction on the word and line-level recognition, this success has not translated to the paragraph level. Hence, our work seeks to mimic the nuances of paragraph-level text images with a custom synthetic data engine using Wikipedia texts. Experiments show that by using our synthetic dataset in tandem with a simple encoder-decoder Transformer, we can achieve the best Word Error Rate (WER) amongst the state-of-the-art methods for handwriting recognition on the IAM dataset. Additionally, we show the model pretrained on English texts can also recognize French and German texts with minimal finetuning. Ernest Yu Kai Chew, Adams Wai-Kin Kong, Joo-Hwee Lim |
ICASSP | 2 |
| 2025 | When and Where Do Data Poisons Attack Textual Inversion?abstractPoisoning attacks pose significant challenges to the robustness of diffusion models (DMs). In this paper, we systematically analyze when and where poisoning attacks textual inversion (TI), a widely used personalization technique for DMs. We first introduce Semantic Sensitivity Maps, a novel method for visualizing the influence of poisoning on text embeddings. Second, we identify and experimentally verify that DMs exhibit non-uniform learning behavior across timesteps, focusing on lower-noise samples. Poisoning attacks inherit this bias and inject adversarial signals predominantly at lower timesteps. Lastly, we observe that adversarial signals distract learning away from relevant concept regions within training data, corrupting the TI process. Based on these insights, we propose Safe-Zone Training (SZT), a novel defense mechanism comprised of 3 key components: (1) JPEG compression to weaken high-frequency poison signals, (2) restriction to high timesteps during TI training to avoid adversarial signals at lower timesteps, and (3) loss masking to constrain learning to relevant regions. Extensive experiments across multiple poisoning methods demonstrate that SZT greatly enhances the robustness of TI against all poisoning attacks, improving generative quality beyond prior published defenses. Code: www.github.com/JStyborski/Diff_Lab Data: www.github.com/JStyborski/NC10 Jeremy Styborski, Mingzhi Lyu, Jiayou Lu, Nupur Kapur, Adams Wai-Kin Kong |
ICCV | 5 |
| 2025 | A Unified Model for Paragraph and Line-Level Handwritten Text Recognition
Ernest Yu Kai Chew, Adams Wai-Kin Kong, Joo-Hwee Lim |
ICDAR (2) | 2 |
| 2025 | Dual Downsample Vision Transformer for Handwritten Text Recognition
Yew Lee Tan, Ernest Yu Kai Chew, Jung-Jae Kim 0001, Adams Wai-Kin Kong |
ICDAR (2) | 4 |
| 2025 | Robust Watermarking Using Generative Priors Against Image Editing: From Benchmarking to AdvancesabstractCurrent image watermarking methods are vulnerable to advanced image editing techniques enabled by large-scale text-to-image models. These models can distort embedded watermarks during editing, posing significant challenges to copyright protection. In this work, we introduce W-Bench, the first comprehensive benchmark designed to evaluate the robustness of watermarking methods against a wide range of image editing techniques, including image regeneration, global editing, local editing, and image-to-video generation. Through extensive evaluations of eleven representative watermarking methods against prevalent editing techniques, we demonstrate that most methods fail to detect watermarks after such edits. To address this limitation, we propose VINE, a watermarking method that significantly enhances robustness against various image editing techniques while maintaining high image quality. Our approach involves two key innovations: (1) we analyze the frequency characteristics of image editing and identify that blurring distortions exhibit similar frequency properties, which allows us to use them as surrogate attacks during training to bolster watermark robustness; (2) we leverage a large-scale pretrained diffusion model SDXL-Turbo, adapting it for the watermarking task to achieve more imperceptible and robust watermark embedding. Experimental results show that our method achieves outstanding watermarking performance under various image editing techniques, outperforming existing methods in both image quality and robustness. Code is available at https://github.com/Shilin-LU/VINE Shilin Lu, Jiayou Lu, Adams Wai-Kin Kong |
ICLR | 5 |
| 2025 | Easing Training Process of Rectified Flow Models Via Lengthening Inter-Path DistanceabstractRecent research pinpoints that different diffusion methods and architectures
trained on the same dataset produce similar results for the same input noise.
This property suggests that they have some preferable noises for a given sample.
By visualizing the noise-sample pairs of rectified flow models and stable diffusion models in two-dimensional spaces,
we observe that the preferable paths, connecting preferable noises to the corresponding samples,
are better organized with significant fewer crossings comparing with
the random paths, connecting random noises to training samples.
In high-dimensional space, paths rarely intersect.
The path crossings in two-dimensional spaces indicate the shorter inter-path distance
in the corresponding high-dimensional spaces.
Inspired by this observation, we propose the Distance-Aware Noise-Sample Matching (DANSM) method
to lengthen the inter-path distance for speeding up the model training.
DANSM is derived from rectified flow models, which allow using a closed-form formula to calculate the inter-path distance.
To further simplify the optimization, we derive the relationship between inter-path distance and path length,
and use the latter in the optimization surrogate.
DANSM is evaluated on both image and latent spaces by rectified flow models and diffusion models.
The experimental results show that DANSM can significantly improve the training speed by 30\% $\sim$ 40\%
without sacrificing the generation quality. Shifeng Xu, Yanzhu Liu, Adams Wai-Kin Kong |
ICLR | 3 |
| 2025 | Transferable Attack against Face Swapping in an Extended SpaceabstractAlthough deep Face Swapping (FS) models may benefit the entertainment industry, they pose severe threats to privacy and security. Existing protections, including deepfake detection and adversarial perturbation, are either passive responses or ineffective to unseen subject-agnostic FS models. In this paper, we propose a transferable attack against subject-agnostic FS models named Additive Identity attack based on a Relighting function (AIR). AIR leverages reillumination and additive perturbations to mislead the identity extraction modules in subject-agnostic FS models. By using these two types of perturbations simultaneously, the attack space is extended such that stronger but more visually natural adversarial examples can be identified. To further enhance the visual quality while preserving the effectiveness of the attack, an adaptive translation-invariant operation and an illumination control scheme are designed for AIR. Unlike other methods, AIR does not require a surrogate FS model to achieve high transferability. In addition, a mathematical proof is given for the extension of the attack space. Extensive experiments using 1000 image pairs across various state-of-the-art subject-agnostic FS models, including GAN and diffusion-based FS models, show that AIR surpasses all existing attacks in terms of both attack success rate and image quality. Mingzhi Lyu, Yi Huang 0013, Adams Wai-Kin Kong |
ICME | 6 |
| 2025 | Variance-Reduction Guidance: Sampling Trajectory Optimization for Diffusion ModelsabstractDiffusion models have become emerging generative models. Their sampling process involves multiple steps, and in each step the models predict the noise from a noisy sample. When the models make prediction, the output deviates from the ground truth, and we call such a deviation as prediction error. The prediction error accumulates over the sampling process and deteriorates generation quality. This paper introduces a novel technique for statistically measuring the prediction error and proposes the Variance-Reduction Guidance (VRG) method to mitigate this error. VRG does not require model fine-tuning or modification. Given a predefined sampling trajectory, it searches for a new trajectory which has the same number of sampling steps but produces higher quality results. VRG is applicable to both conditional and unconditional generation. Experiments on various datasets and baselines demonstrate that VRG can significantly improve the generation quality of diffusion models. Source code is available at https://github.com/shifengxu/VRG. Shifeng Xu, Yanzhu Liu, Adams Wai-Kin Kong |
ICME | 3 |
| 2025 | DDFD: Diffusion-Based Denoising Fusion for Object Detection in Infrared-Visible ImagesabstractInfrared-visible image fusion for object detection (IVIF-OD) aims to utilize complementary information in the two modalities to synthesize new images with richer information to serve object detection. Most existing works focus on how to better fuse pixel-level details while ignoring object-related information required for detection and introducing redundant and object-irrelevant information in the fused images. To address the limitations of previous studies, this paper proposes a diffusion-based denoising fusion for object detection in infrared-visible images, termed DDFD. Specifically, DDFD treats image fusion as a diffusion-based denoising process to generate fused images that are informative yet non-redundant. Since visible imaging is easily affected by adverse conditions, DDFD exploits an image-adaptive enhancement (IAE) module that adaptively improves visible images to achieve better fusion. To extract key fusion features and remove redundancy, DDFD uses an image-aware noise estimator (INE) to determine the noise in the input infrared-visible images for promoting the diffusion denoising network. To take advantage of both the fusion network and object detection network, DDFD jointly optimizes them such that the fusion network can receive object information to improve the fused images, and the improved images can provide high-quality features to enhance object detection performance. Extensive experiments on the M3FD, DroneVehicle, and VEDAI public datasets reveal the superior object detection performance of DDFD and confirm the effectiveness of IVIF-based object detection under challenging weather conditions. Min Dang, Gang Liu 0006, Jingqi Zhao, Adams Wai-Kin Kong, Nan Luo, Di Wang 0011 |
ACM Multimedia | 4 |
| 2025 | Set You Straight: Auto-Steering Denoising Trajectories to Sidestep Unwanted ConceptsabstractEnsuring the ethical deployment of text-to-image models requires effective techniques to prevent the generation of harmful or inappropriate content. While concept erasure methods offer a promising solution, existing finetuning-based approaches suffer from notable limitations. Anchor-free methods risk disrupting sampling trajectories, leading to visual artifacts, while anchor-based methods rely on the heuristic selection of anchor concepts. To overcome these shortcomings, we introduce a finetuning framework, dubbed ANT, which Automatically guides deNoising Trajectories to avoid unwanted concepts. ANT is built on a key insight: reversing the condition direction of classifier-free guidance during mid-to-late denoising stages enables precise content modification without sacrificing early-stage structural integrity. This inspires a trajectory-aware objective that preserves the integrity of the early-stage score function field-which steers samples toward the natural image manifold-without relying on heuristic anchor concept selection. For single-concept erasure, we propose an augmentation-enhanced weight saliency map to precisely identify the critical parameters that most significantly contribute to the unwanted concept, enabling more thorough and efficient erasure. For multi-concept erasure, our objective function offers a versatile plug-and-play solution that significantly boosts performance. Extensive experiments demonstrate that ANT achieves state-of-the-art results in both single and multi-concept erasure, delivering high-quality, safe outputs without compromising the generative fidelity. Code is available at https://github.com/lileyang1210/ANT Leyang Li, Shilin Lu, Yan Ren 0002, Adams Wai-Kin Kong |
ACM Multimedia | 4 |
| 2025 | Enhancing Bioactivity Prediction via Spatial Emptiness Representation of Protein-ligand Complex and Union of Multiple PocketsabstractPredicting the bioactivity of candidate ligands remains a central challenge in drug discovery. Ligands and endogenous substrates often compete for the same binding sites on target proteins, and the extent to which a ligand can modulate protein function depends not only on its binding but also on how effectively it occupies the relevant pocket. However, most existing methods focus narrowly on local interactions within protein–ligand complexes and neglect spatial emptiness—the unoccupied regions within the binding site that may permit endogenous molecules to engage or interfere. Such unfilled space can diminish the ligand’s functional impact, regardless of binding affinity. To overcome this key limitation in protein–ligand modeling, we propose LigoSpace, a novel method integrating three core components. LigoSpace introduces GeoREC (Geometric Representation of Spatial Emptiness in Complexes) to quantify atomic-level empty space and Union-Pocket to unify multiple protein pockets, providing a global view of binding sites. Additionally, LigoSpace employs a pairwise loss instead of commonly used MSE loss, to better capture relative relationships critical for drug discovery. Extensive experiments on multiple datasets with diverse bioactivity types demonstrate that LigoSpace significantly improves performance when integrated into state-of-the-art models, highlighting the effectiveness of its novel components. Zhiyuan Zhou 0001, Yueming Yin, Yuguang Mu, Hoi-Yeung Li, Adams Wai-Kin Kong |
NeurIPS | 6 |
| 2024 | MACE: Mass Concept Erasure in Diffusion ModelsabstractThe rapid expansion of large-scale text-to-image diffusion models has raised growing concerns regarding their potential misuse in creating harmful or misleading content. In this paper, we introduce MACE, a finetuning framework for the task of MAss Concept Erasure. This task aims to prevent models from generating images that embody unwanted concepts when prompted. Existing concept erasure methods are typically restricted to handling fewer than five concepts simultaneously and struggle to find a balance between erasing concept synonyms (generality) and maintaining unrelated concepts (specificity). In contrast, MACE differs by successfully scaling the erasure scope up to 100 concepts and by achieving an effective balance between generality and specificity. This is achieved by leveraging closed-form cross-attention refinement along with LoRA finetuning, collectively eliminating the information of undesirable concepts. Furthermore, MACE integrates multiple LoRAs without mutual interference. We conduct extensive evaluations of MACE against prior methods across four different tasks: object erasure, celebrity erasure, explicit content erasure, and artistic style erasure. Our results reveal that MACE surpasses prior methods in all evaluated tasks. Code is available at https://github.com/Shilin-LU/MACE. Shilin Lu, Zilan Wang, Leyang Li, Yanzhu Liu, Adams Wai-Kin Kong |
CVPR | 5 |
| 2024 | Leveraging Imperfect Restoration for Data Availability Attack
Yi Huang 0013, Jeremy Styborski, Mingzhi Lyu, Fan Wang 0018, Adams Wai-Kin Kong |
ECCV (73) | 5 |
| 2024 | Adaptive Multi-task Learning for Few-Shot Object Detection
Yan Ren 0002, Adams Wai-Kin Kong |
ECCV (7) | 3 |
| 2024 | Exploiting Supervised Poison Vulnerability to Strengthen Self-supervised Defense
Jeremy Styborski, Mingzhi Lyu, Yi Huang 0013, Adams Wai-Kin Kong |
ECCV (82) | 4 |
| 2024 | Flexible-Modal Deception Detection with Audio-Visual AdapterabstractDeception detection within audio-visual modalities is vital across diverse sectors, notably in customs security and multimedia anti-fraud. However, this notable efficacy is lost by the necessity to train and deploy separate models for each conceivable modality scenario, leading to redundancy and inefficiency. Moreover, real-world environments where multi-modal models are deployed often fail to meet these idealized conditions. To overcome these challenges and further elevate performance levels, we propose an advanced Transformer-based framework complemented by an Audio-Visual Adapter (AVA) integrating temporal features from both audio and visual modalities. In addition, we introduce an innovative multi-modal contrastive learning method that is designed to enhance the correlation between uni-modal features and their integrated counterparts within a consistent feature space. Our designed method can deal with the flexible-model scenario instead of deploying different models for various modalities. Empirical evaluations conducted on two benchmark datasets have validated the superiority of our proposed model over other multi-modal fusion techniques, particularly in scenarios characterized by varying and missing modalities. This strongly affirms the effectiveness of our approach in significantly boosting the accuracy of deception detection in complex, real-world multi-modal scenarios. The codes will be released soon. Zhaoxu Li, Zitong Yu, Xun Lin, Nithish Muthuchamy Selvaraj, Xiaobao Guo, Bingquan Shen, Adams Wai-Kin Kong, Alex Chichung Kot |
IJCB | 7 |
| 2024 | Finite Volume Features, Global Geometry Representations, and Residual Training for Deep Learning-based CFD SimulationabstractComputational fluid dynamics (CFD) simulation is an irreplaceable modelling step in many engineering designs, but it is often computationally expensive. Some graph neural network (GNN)-based CFD methods have been proposed. However, the current methods inherit the weakness of traditional numerical simulators, as well as ignore the cell characteristics in the mesh used in the finite volume method, a common method in practical CFD applications. Specifically, the input nodes in these GNN methods have very limited information about any object immersed in the simulation domain and its surrounding environment. Also, the cell characteristics of the mesh such as cell volume, face surface area, and face centroid are not included in the message-passing operations in the GNN methods. To address these weaknesses, this work proposes two novel geometric representations: Shortest Vector (SV) and Directional Integrated Distance (DID). Extracted from the mesh, the SV and DID provide global geometry perspective to each input node, thus removing the need to collect this information through message-passing. This work also introduces the use of Finite Volume Features (FVF) in the graph convolutions as node and edge attributes, enabling its message-passing operations to adjust to different nodes. Finally, this work is the first to demonstrate how residual training, with the availability of low-resolution data, can be adopted to improve the flow field prediction accuracy. Experimental results on two datasets with five different state-of-the-art GNN methods for CFD indicate that SV, DID, FVF and residual training can effectively reduce the predictive error of current GNN-based methods by as much as 41%. Our codes and datasets are available at https://github.com/toggled/FvFGeo. Loh Sher En Jessica, Naheed Anjum Arafat, Wei Xian Lim, Wai Lee Chan, Adams Wai-Kin Kong |
ICML | 5 |
| 2024 | OLB-AC: toward optimizing ligand bioactivities through deep graph learning and activity cliffsabstractMOTIVATION: Deep graph learning (DGL) has been widely employed in the realm of ligand-based virtual screening. Within this field, a key hurdle is the existence of activity cliffs (ACs), where minor chemical alterations can lead to significant changes in bioactivity. In response, several DGL models have been developed to enhance ligand bioactivity prediction in the presence of ACs. Yet, there remains a largely unexplored opportunity within ACs for optimizing ligand bioactivity, making it an area ripe for further investigation. RESULTS: We present a novel approach to simultaneously predict and optimize ligand bioactivities through DGL and ACs (OLB-AC). OLB-AC possesses the capability to optimize ligand molecules located near ACs, providing a direct reference for optimizing ligand bioactivities with the matching of original ligands. To accomplish this, a novel attentive graph reconstruction neural network and ligand optimization scheme are proposed. Attentive graph reconstruction neural network reconstructs original ligands and optimizes them through adversarial representations derived from their bioactivity prediction process. Experimental results on nine drug targets reveal that out of the 667 molecules generated through OLB-AC optimization on datasets comprising 974 low-activity, noninhibitor, or highly toxic ligands, 49 are recognized as known highly active, inhibitor, or nontoxic ligands beyond the datasets' scope. The 27 out of 49 matched molecular pairs generated by OLB-AC reveal novel transformations not present in their training sets. The adversarial representations employed for ligand optimization originate from the gradients of bioactivity predictions. Therefore, we also assess OLB-AC's prediction accuracy across 33 different bioactivity datasets. Results show that OLB-AC achieves the best Pearson correlation coefficient (r2) on 27/33 datasets, with an average improvement of 7.2%-22.9% against the state-of-the-art bioactivity prediction methods. AVAILABILITY AND IMPLEMENTATION: The code and dataset developed in this work are available at github.com/Yueming-Yin/OLB-AC. Yueming Yin, Haifeng Hu 0004, Jitao Yang, Chun Ye, Wilson Wen Bin Goh, Adams Wai-Kin Kong |
Bioinform. | 6 |
| 2024 | L3AM: Linear Adaptive Additive Angular Margin Loss for Video-Based Hand Gesture Authentication
Wenwei Song, Wenxiong Kang, Adams Wai-Kin Kong, Yitao Qiao |
Int. J. Comput. Vis. | 3 |
| 2024 | Bridging the Gap Between Vitiligo Segmentation and Clinical ScoresabstractQuantitative evaluation of vitiligo is crucial for assessing treatment response. Dermatologists evaluate vitiligo regularly to adjust their treatment plans, which requires extra work. Furthermore, the evaluations may not be objective due to inter- and intra-assessor variability. Though automatic vitiligo segmentation methods provide an objective evaluation, previous methods mainly focus on patch-wise images, and their results cannot be translated into clinical scores for treatment adjustment. Thus, full-body vitiligo segmentation needs to be developed for recording vitiligo changes in different body parts of a patient and for calculating the clinical scores. To bridge this gap, the first full-body vitiligo dataset with 1740 images, following the international vitiligo photo standard, was established. Compared with patch-wise images, full-body images have more complicated ambient light conditions and larger variances in lesion size and distribution. Additionally, in some hand and foot images, skin can be fully covered by either vitiligo or healthy skin. Previous patch-wise segmentation studies completely ignore these cases, as they assume that the contrast between vitiligo and healthy skin is available in each image for segmentation. To address the aforementioned challenges, the proposed algorithm in this study exploits a tailor-made contrast enhancement scheme and long-range comparison. Furthermore, a novel confidence score refinement module is proposed to manage images fully covered by vitiligo or healthy skin. Our results can be converted to clinical scores and used by clinicians. Compared to the state-of-the-art method, the proposed algorithm reduces the average per-image vitiligo involvement percentage error from 3.69% to 1.81%, and the top 10% per-image errors from 23.17% to 8.29%. Our algorithm achieves 1.17% and 3.11% for the mean and max error for the per-patient vitiligo involvement percentage, which is better than an experienced dermatologist's naked-eye evaluation. Steven Tien Guan Thng, Adams Wai-Kin Kong |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Advancing Bioactivity Prediction Through Molecular Docking and Self-AttentionabstractBioactivity refers to the ability of a substance to induce biological effects within living systems, often describing the influence of molecules, drugs, or chemicals on organisms. In drug discovery, predicting bioactivity streamlines early-stage candidate screening by swiftly identifying potential active molecules. The popular deep learning methods in bioactivity prediction primarily model the ligand structure-bioactivity relationship under the premise of Quantitative Structure-Activity Relationship (QSAR). However, bioactivity is determined by multiple factors, including not only the ligand structure but also drug-target interactions, signaling pathways, reaction environments, pharmacokinetic properties, and species differences. Our study first integrates drug-target interactions into bioactivity prediction using protein-ligand complex data from molecular docking. We devise a Drug-Target Interaction Graph Neural Network (DTIGN), infusing interatomic forces into intermolecular graphs. DTIGN employs multi-head self-attention to identify native-like binding pockets and poses within molecular docking results. To validate the fidelity of the self-attention mechanism, we gather ground truth data from crystal structure databases. Subsequently, we employ these limited native structures to refine bioactivity prediction via semi-supervised learning. For this study, we establish a unique benchmark dataset for evaluating bioactivity prediction models in the context of protein-ligand complexes, showcasing the superior performance of our method (with an average improvement of 27.03%) through comparison with 9 leading deep learning-based bioactivity prediction methods. Yueming Yin, Hilbert Yuen In Lam, Yuguang Mu, Hoi-Yeung Li, Adams Wai-Kin Kong |
IEEE J. Biomed. Health Informatics | 5 |
| 2023 | A Practical Upper Bound for the Worst-Case Attribution DeviationsabstractModel attribution is a critical component of deep neural networks (DNNs) for its interpretability to complex models. Recent studies bring up attention to the security of attribution methods as they are vulnerable to attribution attacks that generate similar images with dramatically different attributions. Existing works have been investigating empirically improving the robustness of DNNs against those attacks; however, none of them explicitly quantifies the actual deviations of attributions. In this work, for the first time, a constrained optimization problem is formulated to derive an upper bound that measures the largest dissimilarity of attributions after the samples are perturbed by any noises within a certain region while the classification results remain the same. Based on the formulation, different practical approaches are introduced to bound the attributions above using Euclidean distance and cosine similarity under both$\ell_{2}$and$\ell_{\infty}$-norm perturbations constraints. The bounds developed by our theoretical study are validated on various datasets and two different types of attacks (PGD attack and IFIA attribution attack). Over 10 million attacks in the experiments indicate that the proposed upper bounds effectively quantify the robustness of models based on the worst-case attribution dissimilarities. Fan Wang 0018, Adams Wai-Kin Kong |
CVPR | 2 |
| 2023 | Audio-Visual Deception Detection: DOLOS Dataset and Parameter-Efficient Crossmodal LearningabstractDeception detection in conversations is a challenging yet important task, having pivotal applications in many fields such as credibility assessment in business, multimedia anti-frauds, and custom security. Despite this, deception detection research is hindered by the lack of high-quality deception datasets, as well as the difficulties of learning multimodal features effectively. To address this issue, we introduce DOLOS1, the largest gameshow deception detection dataset with rich deceptive conversations. DOLOS includes 1, 675 video clips featuring 213 subjects, and it has been labeled with audio-visual feature annotations. We provide train-test, duration, and gender protocols to investigate the impact of different factors. We benchmark our dataset on previously proposed deception detection approaches. To further improve the performance by fine-tuning fewer parameters, we propose Parameter-Efficient Crossmodal Learning (PECL), where a Uniform Temporal Adapter (UT-Adapter) explores temporal attention in transformer-based architectures, and a crossmodal fusion module, Plug-in Audio-Visual Fusion (PAVF), combines crossmodal information from audio-visual features. Based on the rich fine-grained audio-visual annotations on DOLOS, we also exploit multi-task learning to enhance performance by concurrently predicting deception and audiovisual features. Experimental results demonstrate the desired quality of the DOLOS dataset and the effectiveness of the PECL. The DOLOS dataset and the source codes are available at here. Xiaobao Guo, Nithish Muthuchamy Selvaraj, Zitong Yu, Adams Wai-Kin Kong, Bingquan Shen, Alex Chichung Kot |
ICCV | 4 |
| 2023 | TF-ICON: Diffusion-Based Training-Free Cross-Domain Image CompositionabstractText-driven diffusion models have exhibited impressive generative capabilities, enabling various image editing tasks. In this paper, we propose TF-ICON, a novel Training-Free Image COmpositioN framework that harnesses the power of text-driven diffusion models for cross-domain image-guided composition. This task aims to seamlessly integrate user-provided objects into a specific visual context. Current diffusion-based methods often involve costly instance-based optimization or finetuning of pre-trained models on customized datasets, which can potentially undermine their rich prior. In contrast, TF-ICON can leverage off-the-shelf diffusion models to perform cross-domain image-guided composition without requiring additional training, finetuning, or optimization. Moreover, we introduce the exceptional prompt, which contains no information, to facilitate text-driven diffusion models in accurately inverting real images into latent representations, forming the basis for compositing. Our experiments show that equipping Stable Diffusion with the exceptional prompt outperforms state-of-the-art inversion methods on various datasets (CelebA-HQ, COCO, and ImageNet), and that TF-ICON surpasses prior baselines in versatile visual domains. Code is available at https://github.com/Shilin-LU/TF-ICON Shilin Lu, Yanzhu Liu, Adams Wai-Kin Kong |
ICCV | 3 |
| 2023 | Adapter Incremental Continual Learning of Efficient Audio Spectrogram Transformers
Nithish Muthuchamy Selvaraj, Xiaobao Guo, Adams Wai-Kin Kong, Bingquan Shen, Alex Chichung Kot |
INTERSPEECH | 3 |
| 2023 | Adversarial Attack for Robust Watermark Protection Against Inpainting-based and Blind Watermark RemoversabstractThe rise of social media platforms, especially those focusing on image sharing, has made visible watermarks increasingly important in protecting image copyrights. However, multiple studies have revealed that watermarks are vulnerable to both inpainting-based removers and blind watermark removers. Though two adversarial attack methods have been proposed to defend against watermark removers, they are tailored to a particular type of removers in a white-box setting, which significantly limits their practicality and applicability. To date, there is no adversarial attack method that can protect watermarks against the two types of watermark removers simultaneously. In this paper, we propose a novel method, named Adversarial Watermark Defender with Attribution-Guided Perturbation (AWD-AGP), that defends against both inpainting-based and blind watermark removers under a black-box setting. AWD-AGP is the first watermark protection method employing adversarial location. The adversarial location is generated by a Watermark Positioning Network, which predicts an optimal location for watermark placement, making watermark removal challenging for inpainting-based removers. Since inpainting-based removers and blind watermark removers exploit information in different regions of an image to perform removal, we propose an attribution-guided scheme, which automatically assigns attack strengths to different pixels against different removers. With this design, the generated perturbation can attack the two types of watermark removers concurrently. Experiments on seven models, including four inpainting-based removers and three blind watermark removers demonstrate the effectiveness of AWD-AGP. Mingzhi Lyu, Yi Huang 0013, Adams Wai-Kin Kong |
ACM Multimedia | 3 |
| 2023 | A fully differentiable ligand pose optimization framework guided by deep learning and a traditional scoring functionabstractThe recently reported machine learning- or deep learning-based scoring functions (SFs) have shown exciting performance in predicting protein-ligand binding affinities with fruitful application prospects. However, the differentiation between highly similar ligand conformations, including the native binding pose (the global energy minimum state), remains challenging that could greatly enhance the docking. In this work, we propose a fully differentiable, end-to-end framework for ligand pose optimization based on a hybrid SF called DeepRMSD+Vina combined with a multi-layer perceptron (DeepRMSD) and the traditional AutoDock Vina SF. The DeepRMSD+Vina, which combines (1) the root mean square deviation (RMSD) of the docking pose with respect to the native pose and (2) the AutoDock Vina score, is fully differentiable; thus is capable of optimizing the ligand binding pose to the energy-lowest conformation. Evaluated by the CASF-2016 docking power dataset, the DeepRMSD+Vina reaches a success rate of 94.4%, which outperforms most reported SFs to date. We evaluated the ligand conformation optimization framework in practical molecular docking scenarios (redocking and cross-docking tasks), revealing the high potentialities of this framework in drug design and discovery. Structural analysis shows that this framework has the ability to identify key physical interactions in protein-ligand binding, such as hydrogen-bonding. Our work provides a paradigm for optimizing ligand conformations based on deep learning algorithms. The DeepRMSD+Vina model and the optimization framework are available at GitHub repository https://github.com/zchwang/DeepRMSD-Vina_Optimization. Zechen Wang, Liangzhen Zheng, Sheng Wang 0001, Mingzhi Lin, Adams Wai-Kin Kong, Yuguang Mu, Yanjie Wei |
Briefings Bioinform. | 6 |
| 2023 | Deep Multimodal Sequence Fusion by Regularized Expressive Representation DistillationabstractMultimodal sequence learning aims to utilize information from different modalities to enhance overall performance. Mainstream works often follow an intermediate-fusion pipeline, which explores both modality-specific and modality-supplementary information for fusion. However, the unaligned and heterogeneously distributed multimodal sequences pose significant challenges to the fusion task: 1) to extract both effective unimodal and crossmodal representations and 2) to overcome the overfitting issue in joint multimodal sequence optimization. In this work, we propose regularized expressive representation distillation (RERD) that aims to seek effective multimodal representations and to enhance the generalization of fusion. First, to improve unimodal representation learning, unimodal representations are assigned to multi-head distillation encoders, where the unimodal representations are iteratively updated through distillation attention layers. Second, to alleviate the overfitting issue in joint crossmodal optimization, a multimodal sinkhorn distance regularizer is proposed to reinforce the expressive representation extraction and to reduce the modality gap before fusion adaptively. These representations produce a comprehensive view of the multimodal sequences, which are utilized for downstream fusion tasks. Experimental results on several popular benchmarks demonstrate that the proposed method achieves state-of-the-art performance, compared with widely used baselines for deep multimodal sequence fusion, as shown inhttps://github.com/Redaimao/RERD. Xiaobao Guo, Adams Wai-Kin Kong, Alex Chichung Kot |
IEEE Trans. Multim. | 2 |
| 2023 | Pace-Adaptive and Noise-Resistant Contrastive Learning for Multimodal Feature FusionabstractMultimodal feature fusion aims to draw complementary information from different modalities to achieve better performance. Contrastive learning is effective at discriminating coexisting semantic features (positive) from irrelative ones (negative) in multimodal signals. However, positive and negative pairs learn at separate rates, which undermines the overall performance of multimodal contrastive learning (MCL). Moreover, the learned representation model is not robust, as MCL utilizes supervision signals from potentially noisy modalities. To address these issues, a novel multimodal contrastive learning objective, Pace-adaptive and Noise-resistant Noise-Contrastive Estimation (PN-NCE), is proposed for multimodal fusion by directly using unimodal features. PN-NCE encourages the positive and negative pairs reaching to their optimal similarity scores adaptively and shows less susceptibility to noisy inputs during training. A theoretical analysis is performed on its robustness. Maximizing modality invariance information in the fused representation is expected to benefit the overall performance and therefore, an estimator that measures the difference between the fused representation and its unimodal representations is integrated into MCL to obtain a more modality-invariant fusion output. The proposed method is model-agnostic and can be adapted to various multimodal tasks. It also bears less performance degradation when reducing the number of training samples at the linear probing stage. With different networks and modality inputs from three multimodal datasets, experimental results show that PN-NCE achieves consistent enhancements compared with previous state-of-the-art approaches. Xiaobao Guo, Alex Chichung Kot, Adams Wai-Kin Kong |
IEEE Trans. Multim. | 3 |
| 2022 | Pure Transformer with Integrated Experts for Scene Text Recognition
Yew Lee Tan, Adams Wai-Kin Kong, Jung-Jae Kim 0001 |
ECCV (28) | 2 |
| 2022 | Transferable Adversarial Attack based on Integrated Gradients
Yi Huang 0013, Adams Wai-Kin Kong |
ICLR | 2 |
| 2022 | Portmanteauing Features for Scene Text RecognitionabstractScene text images have different shapes and are subjected to various distortions, e.g. perspective distortions. To handle these challenges, the state-of-the-art methods rely on a rectification network, which is connected to the text recognition network. They form a linear pipeline which uses text rectification on all input images, even for images that can be recognized without it. Undoubtedly, the rectification network improves the overall text recognition performance. However, in some cases, the rectification network generates unnecessary distortions on images, resulting in incorrect predictions in images that would have otherwise been correct without it. In order to alleviate the unnecessary distortions, the portmanteauing of features is proposed. The portmanteau feature, inspired by the portmanteau word, is a feature containing information from both the original text image and the rectified image. To generate the portmanteau feature, a non-linear input pipeline with a block matrix initialization is presented. In this work, the transformer is chosen as the recognition network due to its utilization of attention and inherent parallelism, which can effectively handle the portmanteau feature. The proposed method is examined on 6 benchmarks and compared with 13 state-of-the-art methods. The experimental results show that the proposed method outperforms the state-of-the-art methods on various of the benchmarks. Yew Lee Tan, Ernest Yu Kai Chew, Adams Wai-Kin Kong, Jung-Jae Kim 0001, Joo-Hwee Lim |
ICPR | 3 |
| 2022 | Exploiting the Relationship Between Kendall's Rank Correlation and Cosine Similarity for Attribution ProtectionabstractModel attributions are important in deep neural networks as they aid practitioners in understanding the models, but recent studies reveal that attributions can be easily perturbed by adding imperceptible noise to the input. The non-differentiable Kendall's rank correlation is a key performance index for attribution protection. In this paper, we first show that the expected Kendall's rank correlation is positively correlated to cosine similarity and then indicate that the direction of attribution is the key to attribution robustness. Based on these findings, we explore the vector space of attribution to explain the shortcomings of attribution defense methods using $\ell_p$ norm and propose integrated gradient regularizer (IGR), which maximizes the cosine similarity between natural and perturbed attributions. Our analysis further exposes that IGR encourages neurons with the same activation states for natural samples and the corresponding perturbed samples. Our experiments on different models and datasets confirm our analysis on attribution protection and demonstrate a decent improvement in adversarial robustness. Fan Wang 0018, Adams Wai-Kin Kong |
NeurIPS | 2 |
| 2021 | Unimodal and Crossmodal Refinement Network for Multimodal Sequence FusionabstractEffective unimodal representation and complementary crossmodal representation fusion are both important in multimodal representation learning.Prior works often modulate one modal feature to another straightforwardly and thus, underutilizing both unimodal and crossmodal representation refinements, which incurs a bottleneck of performance improvement.In this paper, Unimodal and Crossmodal Refinement Network (UCRN) is proposed to enhance both unimodal and crossmodal representations.Specifically, to improve unimodal representations, a unimodal refinement module is designed to refine modality-specific learning via iteratively updating the distribution with transformer-based attention layers.Self-quality improvement layers are followed to generate the desired weighted representations progressively.Subsequently, those unimodal representations are projected into a common latent space, regularized by a multimodal Jensen-Shannon divergence loss for better crossmodal refinement.Lastly, a crossmodal refinement module is employed to integrate all information.By hierarchical explorations on unimodal, bimodal, and trimodal interactions, UCRN is highly robust against missing modality and noisy data.Experimental results on MOSI and MOSEI datasets illustrated that the proposed UCRN outperforms recent state-of-the-art techniques and its robustness is highly preferred in real multimodal sequence fusion scenarios.Codes will be shared publicly 1 . Xiaobao Guo, Adams Wai-Kin Kong |
EMNLP (1) | 2 |
| 2021 | Pixel-wise ordinal classification for salient object grading
Yanzhu Liu, Adams Wai-Kin Kong |
Image Vis. Comput. | 3 |
| 2021 | Segmenting Vitiligo on Clinical Face Images Using CNN Trained on Synthetic and Internet ImagesabstractAccurately diagnosing and describing the severity of vitiligo is crucial for prognostication, treatment selection and comparison. Currently, disease severity scores require dermatologists to estimate percentage area of involvement, which is subjected to inter and intra-assessor variability. Previous studies focus on pure skin but vitiligo on the face, which has a more serious impact on patients' quality of life, was completely neglected. Convolutional neural networks (CNNs) have good performance on many segmentation tasks. However, due to data privacy, it is hard to have a large clinical vitiligo face image dataset to train a CNN. To address this challenge, images from two different sources, the Internet and the proposed vitiligo face synthesis algorithm, are employed in training. 843 vitiligo images taken from different viewpoints were collected from the Internet. These images are hugely different from the target clinical images collected according to a newly established international standard. To have more vitiligo face images similar to the target clinical images to enhance segmentation performance, an image synthesis algorithm is proposed. Both synthetic and Internet images are used to train a CNN which is modified from the fully convolutional network (FCN) to segment face vitiligo lesions. The results show that 1) the synthetic images effectively improve segmentation performance; 2) the proposed algorithm achieves 1.06 % error for the face vitiligo area estimation and 3) it is more accurate than two dermatologists and all the previous automated vitiligo segmentation methods, which were designed for segmentation vitiligo on pure skin. Adams Wai-Kin Kong, Steven Thng |
IEEE J. Biomed. Health Informatics | 2 |
| 2020 | New Threats Against Object Detector with Non-local Block
Yi Huang 0013, Fan Wang 0018, Adams Wai-Kin Kong, Kwok-Yan Lam |
ECCV (20) | 3 |
| 2020 | Gender and Ethnicity Classification based on Palmprint and Palmar Hand Images from Uncontrolled EnvironmentabstractSoft biometric attributes such as gender, ethnicity or age may provide useful information for biometrics and forensics applications. Researchers used, e.g., face, gait, iris, and hand, etc. to classify such attributes. Even though hand has been widely studied for biometric recognition, relatively less attention has been given to soft biometrics from hand. Previous studies of soft biometrics based on hand images focused on gender and well-controlled imaging environment. In this paper, the gender and ethnicity classification in uncontrolled environment are considered. Gender and ethnicity labels are collected and provided for subjects in a publicly available database, which contains hand images from the Internet. Five deep learning models are fine-tuned and evaluated in gender and ethnicity classification scenarios based on palmar 1) full hand, 2) segmented hand and 3) palmprint images. The experimental results indicate that for gender and ethnicity classification in uncontrolled environment, full and segmented hand images are more suitable than palmprint images. Wojciech Michal Matkowski, Adams Wai-Kin Kong |
IJCB | 2 |
| 2020 | A portrait photo-to-tattoo transform based on digital tattooing
Xingpeng Xu, Wojciech Michal Matkowski, Adams Wai-Kin Kong |
Multim. Tools Appl. | 3 |
| 2020 | A survey on image and video cosegmentation: Methods, challenges and analyses
Yan Ren 0002, Adams Wai-Kin Kong, Licheng Jiao |
Pattern Recognit. | 2 |
| 2020 | Palmprint Recognition in Uncontrolled and Uncooperative EnvironmentabstractOnline palmprint recognition and latent palmprint identification are two branches of palmprint studies. The former uses middle-resolution images collected by a digital camera in a well-controlled or contact-based environment with user cooperation for commercial applications and the latter uses high-resolution latent palmprints collected in crime scenes for forensic investigation. However, these two branches do not cover some palmprint images which have the potential for forensic investigation. Due to the prevalence of smartphone and consumer camera, more evidence is in the form of digital images taken in uncontrolled and uncooperative environment, e.g., child pornographic images and terrorist images, where the criminals commonly hide or cover their face. However, their palms can be observable. To study palmprint identification on images collected in uncontrolled and uncooperative environment, a new palmprint database is established and an end-to-end deep learning algorithm is proposed. The new database named NTU Palmprints from the Internet (NTU-PI-v1) contains 7881 images from 2035 palms collected from the Internet. The proposed algorithm consists of an alignment network and a feature extraction network and is end-to-end trainable. The proposed algorithm is compared with the state-of-the-art online palmprint recognition methods and evaluated on three public contactless palmprint databases, IITD, CASIA, and PolyU and two new databases, NTU-PI-v1 and NTU contactless palmprint database. The experimental results showed that the proposed algorithm outperforms the existing palmprint recognition methods. Wojciech Michal Matkowski, Tingting Chai, Adams Wai-Kin Kong |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2019 | Adversarial Signboard against Object Detector
Yi Huang 0013, Adams Wai-Kin Kong, Kwok-Yan Lam |
BMVC | 2 |
| 2019 | Probabilistic Deep Ordinal Regression Based on Gaussian ProcessesabstractWith excellent representation power for complex data, deep neural networks (DNNs) based approaches are state-of-the-art for ordinal regression problem which aims to classify instances into ordinal categories. However, DNNs are not able to capture uncertainties and produce probabilistic interpretations. As a probabilistic model, Gaussian Processes (GPs) on the other hand offers uncertainty information, which is nonetheless lack of scalability for large datasets. This paper adapts traditional GPs regression for ordinal regression problem by using both conjugate and non-conjugate ordinal likelihood. Based on that, it proposes a deep neural network with a GPs layer on the top, which is trained end-to-end by the stochastic gradient descent method for both neural network parameters and GPs parameters. The parameters in the ordinal likelihood function are learned as neural network parameters so that the proposed framework is able to produce fitted likelihood functions for training sets and make probabilistic predictions for test points. Experimental results on three real-world benchmarks - image aesthetics rating, historical image grading and age group estimation - demonstrate that in terms of mean absolute error, the proposed approach outperforms state-of-the-art ordinal regression approaches and provides the confidence for predictions. Yanzhu Liu, Fan Wang 0018, Adams Wai-Kin Kong |
ICCV | 3 |
| 2019 | Giant Panda Face Recognition Using Small DatasetabstractGiant panda (panda) is a highly endangered animal. Significant efforts and resources have been put on panda conservation. To measure effectiveness of conservation schemes, estimating its population size in wild is an important task. The current population estimation approaches, including capture-recapture, human visual identification and collection of DNA from hair or feces, are invasive, subjective, costly or even dangerous to the workers who perform these tasks in wild. Cameras have been widely installed in the regions where pandas live. It opens a new possibility for non-invasive image based panda recognition. Panda face recognition is naturally a small dataset problem, because of the number of pandas in the world and the number of qualified images captured by the cameras in each encounter. In this paper, a panda face recognition algorithm, which includes alignment, large feature set extraction and matching is proposed and evaluated on a dataset consisting of 163 images. The experimental results are encouraging. Wojciech Michal Matkowski, Adams Wai-Kin Kong, Han Su 0002, Rong Hou, Zhihe Zhang |
ICIP | 2 |
| 2019 | Towards Touch-to-Access Device Authentication Using Induced Body Electric PotentialsabstractThis paper presents TouchAuth, a new touch-to-access device authentication approach using induced body electric potentials (iBEPs) caused by the indoor ambient electric field that is mainly emitted from the building's electrical cabling. The design of TouchAuth is based on the electrostatics of iBEP generation and a resulting property, i.e., the iBEPs at two close locations on the same human body are similar, whereas those from different human bodies are distinct. Extensive experiments verify the above property and show that TouchAuth achieves high-profile receiver operating characteristics in implementing the touch-to-access policy. Our experiments also show that a range of possible interfering sources including appliances' electromagnetic emanations and noise injections into the power network do not affect the performance of TouchAuth. A key advantage of TouchAuth is that the iBEP sensing requires a simple analog-to-digital converter only, which is widely available on microcontrollers. Compared with existing approaches including intra-body communication and physiological sensing, TouchAuth is a low-cost, lightweight, and convenient approach for authorized users to access the smart objects found in indoor environments. Zhenyu Yan 0002, Qun Song 0001, Rui Tan 0001, Yang Li 0147, Adams Wai-Kin Kong |
MobiCom | 5 |
| 2019 | Attacking Object Detectors Without Changing the Target Object
Yi Huang 0013, Adams Wai-Kin Kong, Kwok-Yan Lam |
PRICAI (3) | 2 |
| 2019 | A study on wrist identification for forensic investigation
Wojciech Michal Matkowski, Frodo Kin-Sun Chan, Adams Wai-Kin Kong |
Image Vis. Comput. | 3 |
| 2018 | A Constrained Deep Neural Network for Ordinal RegressionabstractOrdinal regression is a supervised learning problem aiming to classify instances into ordinal categories. It is challenging to automatically extract high-level features for representing intraclass information and interclass ordinal relationship simultaneously. This paper proposes a constrained optimization formulation for the ordinal regression problem which minimizes the negative loglikelihood for multiple categories constrained by the order relationship between instances. Mathematically, it is equivalent to an unconstrained formulation with a pairwise regularizer. An implementation based on the CNN framework is proposed to solve the problem such that high-level features can be extracted automatically, and the optimal solution can be learned through the traditional back-propagation method. The proposed pairwise constraints make the algorithm work even on small datasets, and a proposed efficient implementation make it be scalable for large datasets. Experimental results on four real-world benchmarks demonstrate that the proposed algorithm outperforms the traditional deep learning approaches and other state-of-the-art approaches based on hand-crafted features. Yanzhu Liu, Adams Wai-Kin Kong, Chi Keong Goh |
CVPR | 2 |
| 2018 | Using Object Information for Spotting Text
Shitala Prasad, Adams Wai-Kin Kong |
ECCV (16) | 2 |
| 2018 | A further study of low resolution androgenic hair patterns as a soft biometric traitabstractSoft biometric traits such as skin color, tattoos, shoe size, height, and weight have been regularly used for forensic investigation , especially when hard biometric traits , e.g., faces and fingerprints are not available. Recently, a new soft biometric trait, androgenic hair also called body hair, was evaluated. The previous study showed that low resolution androgenic hair patterns have potential for forensic investigation . However, it was believed that they are not a distinctive biometric trait because of the reported accuracy. To explore discriminative information in androgenic hair patterns, in this paper, a new algorithm, which makes use of leg geometry to align lower leg images, large feature sets (about 60,000 features) extracted through multi-directional grid systems to increase discriminative power and robustness, and class-specific partial least squares (PLS) models to utilize the features effectively, is employed. To further enhance the performance of the class-specific PLS models trained on very limited positive samples, one to three images per model in the experiments, and further enhance robustness against viewpoint and pose variations, a scheme is designed to generate more positive samples from a single image. Experimental results on 1493 low resolution leg images with large viewpoint and pose variations from 412 legs demonstrate that low resolution androgenic hair patterns contain rich information and the impression of low discriminative power on androgenic hair is due to the method used in the previous study. Frodo Kin-Sun Chan, Adams Wai-Kin Kong |
Image Vis. Comput. | 2 |
| 2017 | Deep Ordinal Regression Based on Data Relationship for Small DatasetsabstractOrdinal regression aims to classify instances into ordinal categories. As with other supervised learning problems, learning an effective deep ordinal model from a small dataset is challenging. This paper proposes a new approach which transforms the ordinal regression problem to binary classification problems and uses triplets with instances from different categories to train deep neural networks such that high-level features describing their ordinal relationship can be extracted automatically. In the testing phase, triplets are formed by a testing instance and other instances with known ranks. A decoder is designed to estimate the rank of the testing instance based on the outputs of the network. Because of the data argumentation by permutation, deep learning can work for ordinal regression even on small datasets. Experimental results on the historical color image benchmark and MSRA image search datasets demonstrate that the proposed algorithm outperforms the traditional deep learning approach and is comparable with other state-of-the-art methods, which are highly based on prior knowledge to design effective features. Yanzhu Liu, Adams Wai-Kin Kong, Chi Keong Goh |
IJCAI | 2 |
| 2017 | A multi-model restoration algorithm for recovering blood vessels in skin images
Adams Wai-Kin Kong |
Image Vis. Comput. | 2 |
| 2017 | A Study of Distinctiveness of Skin Texture for Forensic Applications Through Comparison With Blood VesselsabstractSkin texture without obvious features is different from other hard biometrics on the skin, such as fingerprints and palmprints. Skin texture gives an impression that it is not distinctive like other soft biometric traits. It was proposed for personal identification a decade ago but did not draw attention from the biometric community, partially due to the success of other biometric technologies for commercial applications. However, in some forensic cases, e.g., identifying masked terrorists in images, skin texture may be the only option. Faces, tattoos, and skin marks are not always available for identification. To address these forensic needs, researchers have recently attempted to visualize blood vessels hidden in color images. Their performance is highly sensitive to image quality. Skin texture that is easily captured even in low-resolution images, such as that of the forearm skin, is suitable for these forensic applications. To study the distinctiveness of low-resolution skin texture, in this paper, an algorithm composed of a positive sample generation scheme, dynamic and directional grids, a large feature set generation scheme, and partial least squares regression has been proposed. More than 6300 inner forearm and thigh images collected from a laboratory environment and from the internet with large pose, viewpoint, and illumination variations were employed in this paper. The proposed algorithm was compared with the state-of-the-art texture recognition methods, and skin texture was compared with blood vessels, a hard biometric trait, extracted from color and infrared images. The results showed that the proposed algorithm performed significantly better than did the texture recognition methods, and skin texture outperformed blood vessels in all of the experiments, achieving encouraging performance. Frodo Kin-Sun Chan, Adams Wai-Kin Kong |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2016 | A geometric-based tattoo retrieval systemabstractVarious soft biometric traits have been used as hints in forensic investigation. Tattoo, as one of those soft biometric traits, has been used extensively because it is easy to be remembered and described by witnesses and appears very often among criminals and victims. Most of the tattoo retrieval systems currently used in police departments are still text-based systems. They depend on labels tagged on the tattoo images in databases. This manual labelling process is very time consuming and subject to loss of some detailed information, such as the accurate location and the shape of tattoos. These problems can make tattoo retrieval inefficiently and misleading. To address them, a geometric based tattoo retrieval system is developed. It allows witnesses to draw the boundary of a query tattoo and retrieve all the tattoos with similar shape around the particular location. The system comprises a tattoo detection algorithm, which detects tattoos from full body images, a full body coordinate algorithm, which defines locations of input tattoo boundaries and locations of tattoos in databases and a tattoo shape matching algorithm, which measures similarity between input boundaries and boundaries of tattoos in databases. The experimental results on 2188 images show the effectiveness of the proposed system. Xingpeng Xu, Adams Wai-Kin Kong |
ICPR | 2 |
| 2016 | A Fully Automatic Method for Gridding Bright Field Images of Bead-Based MicroarraysabstractIn this paper, a fully automatic method for gridding bright field images of bead-based microarrays is proposed. There have been numerous techniques developed for gridding fluorescence images of traditional spotted microarrays but to our best knowledge, no algorithm has yet been developed for gridding bright field images of bead-based microarrays. The proposed gridding method is designed for automatic quality control during fabrication and assembly of bead-based microarrays. The method begins by estimating the grid parameters using an evolutionary algorithm. This is followed by a grid-fitting step that rigidly aligns an ideal grid with the image. Finally, a grid refinement step deforms the ideal grid to better fit the image. The grid fitting and refinement are performed locally and the final grid is a nonlinear (piecewise affine) grid. To deal with extreme corruptions in the image, the initial grid parameter estimation and grid-fitting steps employ robust search techniques. The proposed method does not have any free parameters that need tuning. The method is capable of identifying the grid structure even in the presence of extreme amounts of artifacts and distortions. Evaluation results on a variety of images are presented. Abhik Datta, Adams Wai-Kin Kong, Kin Choong Yow 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2015 | A Statistical Analysis of IrisCode and Its Security ImplicationsabstractIrisCode has been used to gather iris data for 430 million people. Because of the huge impact of IrisCode, it is vital that it is completely understood. This paper first studies the relationship between bit probabilities and a mean of iris images (The mean of iris images is defined as the average of independent iris images.) and then uses the Chi-square statistic, the correlation coefficient and a resampling algorithm to detect statistical dependence between bits. The results show that the statistical dependence forms a graph with a sparse and structural adjacency matrix. A comparison of this graph with a graph whose edges are defined by the inner product of the Gabor filters that produce IrisCodes shows that partial statistical dependence is induced by the filters and propagates through the graph. Using this statistical information, the security risk associated with two patented template protection schemes that have been deployed in commercial systems for producing application-specific IrisCodes is analyzed. To retain high identification speed, they use the same key to lock all IrisCodes in a database. The belief has been that if the key is not compromised, the IrisCodes are secure. This study shows that even without the key, application-specific IrisCodes can be unlocked and that the key can be obtained through the statistical dependence detected. Adams Wai-Kin Kong |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | On Criminal Identification in Color Skin Images Using Skin Marks (RPPVSM) and Fusion With Inferred Vein PatternsabstractRelatively Permanent Pigmented or Vascular Skin Marks (RPPVSM) were recently introduced as a biometric trait for identification in the cases in which the evidence images show only the nonfacial body parts of the criminals or victims, such as in child sexual abuse and riots. As manual RPPVSM identification is tiring and time-consuming, an automated RPPVSM identification system is proposed in this paper. The system comprises skin segmentation, RPPVSM detection, and RPPVSM matching algorithms. The system was evaluated on 1,200 back images collected from 283 Asian and Caucasian subjects in varying pose and viewpoint conditions. The system achieved rank-1 and rank-10 identification accuracies of 76.79% and 88.97%, respectively, higher than the identification accuracies given by existing skin mark detection methods previously proposed for face recognition systems. To handle identification with limited numbers of RPPVSM, a fusion scheme with inferred vein patterns is also proposed. The fusion was evaluated on 2,360 images of chests, forearms, and thighs collected mostly from Asian subjects, who tend to have fewer RPPVSM than Caucasian subjects. The results show that the fusion improves vein identification in all body parts with improvement rates varying between 2% and 5% depending on the number of RPPVSM detected. To the best of our knowledge, this is the first work on automated identification in color skin images based on nonfacial skin marks and fusion with inferred vein patterns in forensic settings. Arfika Nurhudatiana, Adams Wai-Kin Kong |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2014 | Using Leg Geometry to Align Androgenic Hair Patterns in Low Resolution Images for Criminal and Victim IdentificationabstractIdentifying criminals and victims in digital media can be very challenging when neither their faces nor tattoos are observable. These criminals and victims can be masked gunmen, paedophiles, and victims in child pornographic images. Though skin marks and blood vessel patterns hidden in color images have been proposed to address this problem, they have different limitations. Blood vessel patterns are not suitable for subjects with high concentration of body fat or melanin and skin marks are not suitable for the cases where only low resolution images are available. A recent paper pointed out that androgenic hair, which is also called body hair, and its follicles can be used as biometric traits and demonstrated that androgenic hair patterns in low resolution images are effective for personal identification. In that study, viewpoint and pose variations were ignored, but these variations are unavoidable in real cases. To overcome the weaknesses of the previous method, this paper proposes an algorithm based on geometric information for aligning androgenic hair patterns on lower legs. The experimental results on 1,138 low resolution images from 283 different legs show that the proposed alignment algorithm offers more than 5% improvement. Frodo Kin-Sun Chan, Adams Wai-Kin Kong |
ICPR | 2 |
| 2014 | A Preliminary Study of Lower Leg Geometry as a Soft Biometric Trait for Forensic InvestigationabstractCriminal and victim identification is always vital in forensic investigation. Many biometric traits, such as DNA, fingerprint, face and palm print, have been regularly used by law enforcement agencies. However, they are not applicable to legal cases where only non-facial body sites of criminals or victims in evidence images are available for identification. These cases include but are not limited to violent protests, masked gunmen and child pornography. To address this challenging identification problem, skin marks, blood vessels hidden in color images, androgenic hair patterns and tattoos have been considered. Tattoos are not always available. Skin marks and blood vessels are suitable for high resolution images. Androgenic hair patterns provide useful identification information even in low resolution images, but their performance is still far from perfect. Thus, new biometric traits are still demanded especially for low resolution evidence images. This paper evaluates lower leg geometry as a soft biometric trait for criminal and victim identification. Lower legs are considered in this study because they are often observable in evidence images. The algorithm utilized in this evaluation first aligns two lower leg shapes from input images and extracts geometric features, including the partial sum of squared difference, the polynomial coefficients and the number of intersection points of the aligned leg shapes. Support vector machines, neural networks and decision trees are used to perform the classification. The algorithm is applied to 1,138 images from 283 subjects. The experimental results indicate that lower leg geometry is an effective soft biometric trait. This study provides a foundation for further research on criminal and victim identification based on body geometry. Md. Rabiul Islam 0002, Frodo Kin-Sun Chan, Adams Wai-Kin Kong |
ICPR | 3 |
| 2014 | Vein Pattern Visualization through Multiple Mapping Models and Local Parameter Estimation for Forensic InvestigationabstractForensic investigation methods based on some human traits, including fingerprint, face, and palm print, have been developed significantly, but some major aspects of particular crimes such as child pornography still lack of notable research efforts. Unlike common forensic identification methods, techniques for identifying criminals in child pornographic images should be developed based on partial non-facial skin observable in the images because criminals always hide their faces. Few methods published recently have shown the potential of vein patterns visualized from color images as a criminal and victim identification tool. However, these methods have two weaknesses: 1) they use single model to visualize vein patterns hidden in color images, which neglects the diversity of skin properties and 2) even though their parameters are determined automatically by an optimization, they do not adapt to fit local image characteristics. To address these weaknesses, this paper proposes an algorithm composed of a bank of mapping models which transform color images to near infrared (NIR) images for visualizing vein patterns and a local parameter estimation scheme for handling different image characteristics in different regions. Imbalanced data regression is also used to systematically construct the model bank. The proposed algorithm is examined and compared with the previous methods on a database of 920 thigh images from 230 subjects. It outperforms the previous methods. Hamid R. Sharifzadeh, Hengyi Zhang, Adams Wai-Kin Kong |
ICPR | 3 |
| 2014 | A Study on Low Resolution Androgenic Hair Patterns for Criminal and Victim IdentificationabstractIdentifying criminals and victims in images (e.g., child pornography and masked gunmen) can be a challenging task, especially when neither their faces nor tattoos are observable. Skin mark patterns and blood vessel patterns are recently proposed to address this problem. However, they are invisible in low-resolution images and dense androgenic hair can cover them completely. Medical research results have implied that androgenic hair patterns are a stable biometric trait and have potential to overcome the weaknesses of skin mark patterns and blood vessel patterns. To the best of our knowledge, no one has studied androgenic hair patterns for criminal and victim identification before. This paper aims to study matching performance of androgenic hair patterns in low-resolution images. An algorithm designed for this paper uses Gabor filters to compute orientation fields of androgenic hair patterns, histograms on a dynamic grid system to describe their local orientation fields, and the blockwise Chi-square distance to measure the dissimilarity between two patterns. The 4552 images from 283 different legs with resolutions of 25, 18.75, 12.5, and 6.25 dpi were examined. The experimental results indicate that androgenic hair patterns even in low-resolution images are an effective biometric trait and the proposed Gabor orientation histograms are comparable with other well-known texture recognition methods, including local binary patterns, local Gabor binary patterns, and histograms of oriented gradients. Han Su 0002, Adams Wai-Kin Kong |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2013 | A fast point pattern matching algorithm for robust spatially addressable bead encodingabstractBead encoding is a key problem central to all bead based microarrays. Recently a spatially addressable bead encoding technique has been developed ([1], [2]) that alleviates the need for costly hardware while still allowing high-throughput analysis. This paper proposes a pattern matching based scheme that extends this bead encoding technique's usability to uncontrolled environments. A novel affine invariant point pattern matching algorithm is developed to achieve this. The proposed algorithm uses local features to overcome the combinatorial explosion problem encountered in matching corrupted point patterns. The use of efficient data structures is emphasized to make the algorithm fast and scalable. The proposed scheme can decode bead identities in assays involving thousands of beads in a few seconds. Evaluation results using both real and simulated data are presented. Abhik Datta, Adams Wai-Kin Kong, Soumita Ghosh, Dieter Trau |
BIBE | 2 |
| 2013 | Automated identification of Relatively Permanent Pigmented or Vascular Skin Marks (RPPVSM)abstractIn cases of child pornography and child sexual abuse, criminals are usually careful to hide or cover their faces and tattoos, thus making identification difficult. However, naturally occurring skin marks can be observed in close-up views of their back, chest, or thighs, which are usually present in evidence images. Recently, a group of skin marks named Relatively Permanent Pigmented or Vascular Skin Marks (RPPVSM) was proposed as a biometric trait for identification. Manual RPPVSM identification can be tiring and time consuming. We propose in this paper an automated RPPVSM identification system, which is composed of RPPVSM detection and matching algorithms. Three learning-based detection algorithms were developed to automatically detect RPPVSMs in color images. To evaluate these algorithms, experiments were performed on a database containing 216 back torso images from 118 subjects. The results show that high identification accuracy can be achieved and that the proposed RPPVSM identification system has high potential for forensic investigation. Arfika Nurhudatiana, Adams Wai-Kin Kong, Lisa Altieri, Noah Craft |
ICASSP | 2 |
| 2013 | The Individuality of Relatively Permanent Pigmented or Vascular Skin Marks (RPPVSM) in Independently and Uniformly Distributed PatternsabstractWith recent advances in multimedia technology, the involvement of digital images/videos in crimes has been increasing significantly. Identification of individuals in these images/videos can be challenging. For example, in cases of child sexual abuse, child pornography, and masked gunmen, the faces of criminals or victims are often hidden or covered and only some body parts (e.g., back, thigh, and arm) can be observed from the digital evidence. Although tattoos and scars can be used for identification in some cases, they are neither universal nor unique. We propose a group of skin marks named Relatively Permanent Pigmented or Vascular Skin Marks (RPPVSM) as a biometric trait for forensic identification. To support the scientific underpinnings of using RPPVSM patterns as a novel biometric trait, the individuality was studied. RPPVSM on the backs of 269 male subjects were examined. We found that RPPVSM in middle to low density patterns tend to form an independent and uniform distribution, while RPPVSM in high density patterns tend to form clusters. We present in this paper an individuality model for the independently and uniformly distributed RPPVSM patterns. When compared to the empirical results, this model fits the empirical distribution very well. Finally, the predicted error rates for verification and identification are reported. Arfika Nurhudatiana, Adams Wai-Kin Kong, Keyan Matinpour, Deborah Chon, Lisa Altieri, Siu-Yeung Cho, Noah Craft |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2013 | Modeling IrisCode and Its Variants as Convex Polyhedral Cones and Its Security ImplicationsabstractIrisCode, developed by Daugman, in 1993, is the most influential iris recognition algorithm. A thorough understanding of IrisCode is essential, because over 100 million persons have been enrolled by this algorithm and many biometric personal identification and template protection methods have been developed based on IrisCode. This paper indicates that a template produced by IrisCode or its variants is a convex polyhedral cone in a hyperspace. Its central ray, being a rough representation of the original biometric signal, can be computed by a simple algorithm, which can often be implemented in one Matlab command line. The central ray is an expected ray and also an optimal ray of an objective function on a group of distributions. This algorithm is derived from geometric properties of a convex polyhedral cone but does not rely on any prior knowledge (e.g., iris images). The experimental results show that biometric templates, including iris and palmprint templates, produced by different recognition methods can be matched through the central rays in their convex polyhedral cones and that templates protected by a method extended from IrisCode can be broken into. These experimental results indicate that, without a thorough security analysis, convex polyhedral cone templates cannot be assumed secure. Additionally, the simplicity of the algorithm implies that even junior hackers without knowledge of advanced image processing and biometric databases can still break into protected templates and reveal relationships among templates produced by different recognition methods. Adams Wai-Kin Kong |
IEEE Trans. Image Process. | 1 |
| 2012 | Visualizing vein patterns from color skin images based on image mapping for forensics analysis
Chaoying Tang, Hengyi Zhang, Adams Wai-Kin Kong, Noah Craft |
ICPR | 3 |
| 2012 | IrisCode Decompression Based on the Dependence between Its Bit PairsabstractIrisCode is an iris recognition algorithm developed in 1993 and continuously improved by Daugman. Understanding IrisCode's properties is extremely important because over 60 million people have been mathematically enrolled by the algorithm. In this paper, IrisCode is proved to be a compression algorithm, which is to say its templates are compressed iris images. In our experiments, the compression ratio of these images is 1:655. An algorithm is designed to perform this decompression by exploiting a graph composed of the bit pairs in IrisCode, prior knowledge from iris image databases, and the theoretical results. To remove artifacts, two postprocessing techniques that carry out optimization in the Fourier domain are developed. Decompressed iris images obtained from two public iris image databases are evaluated by visual comparison, two objective image quality assessment metrics, and eight iris recognition methods. The experimental results show that the decompressed iris images retain iris texture that their quality is roughly equivalent to a JPEG quality factor of 10 and that the iris recognition methods can match the original images with the decompressed images. This paper also discusses the impacts of these theoretical and experimental findings on privacy and security. Adams Wai-Kin Kong |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | Uncovering vein patterns from color skin images for forensic analysisabstractRecent technological advances have allowed for a proliferation of digital evidence images. Using these images as evidence in legal cases (e.g. child sexual abuse, child pornography and masked gunmen) can be very challenging, because the faces of criminals or victims are not visible. Although large skin marks and tattoos have been used, they are ineffective in some legal cases, because the skin exposed in evidence images neither have unique tattoos nor enough skin marks for identification. The blood vessel between the skin and the muscle covering most parts of the human body is a powerful biometric trait, because of its universality, permanence and distinctiveness. Traditionally, it was impossible to use vein patterns for forensic identification, because they were not visible in color images. This paper proposes an algorithm to uncover vein patterns from the skin exposed in color images for personal identification. Based on the principles of optics and skin biophysics, we modeled the inverse process of skin color formation in an image and derived spatial distributions of biophysical parameters from color images, where vein patterns can be observed. Experimental results are very encouraging. The clarity of the vein patterns in resultant images is comparable to or even better than that in near infrared images. Chaoying Tang, Adams Wai-Kin Kong, Noah Craft |
CVPR | 2 |
| 2011 | A knowledge-based algorithm to remove blocking artifacts in skin images for forensic analysisabstractIdentifying criminals and victims in evidence images, where their faces are covered or obstructed, is a challenging task. In the legal case, United States v. Michael Joseph Pepe (2008), Craft and Kong, who served as expert witnesses, used nevi to identify a pedophile in evidence images. Their expert opinions were challenged, partially because the blocking artifacts generated by the standard JPEG algorithm adversely affected the visibility of the nevi. In addition to this case, a huge amount of JPEG compressed child pornography is posted on-line every day. Although many methods have been proposed to remove blocking artifacts, they are ineffective for our target application. In this paper, a knowledge-based algorithm, which simultaneously removes JPEG blocking artifacts, and recovers skin features, is proposed. Given a training dataset which contains original and compressed skin images, the relationship between original blocks and compressed blocks can be established. This prior information is used to infer original blocks of compressed evidence images. An indexing mechanism is also proposed to deal with large datasets efficiently. Extensive experiments are conducted on images with different characteristics and compression ratios. Both visual comparison and subjective evaluation demonstrate that the proposed algorithm is more effective than other methods. Chaoying Tang, Adams Wai-Kin Kong, Noah Craft |
ICASSP | 2 |
| 2011 | Fundamental statistics of relatively permanent pigmented or vascular skin marks for criminal and victim identificationabstractRecent technological advances have allowed for a proliferation of digital images that may be involved in crimes. Using these images as evidence in legal cases like child pornography and masked gunmen can be challenging because usually the faces of the suspects are not visible. To perform personal identification in these images, we propose a biometric trait composed of a group of skin marks including, but not limited to, nevi, lentigines, cherry hemangiomas, and seborrheic keratoses. Due to their biological characteristics, we have grouped these as "Relatively Permanent Pigmented or Vascular Skin Marks," abbreviated as RPPVSM. As statistical study of RPPVSM is essential before investigating their discriminative power, we present in this paper the fundamental statistics of RPPVSM. Back torso images were collected from 144 Caucasian, Asian, and Latino males, and a researcher trained in dermatology manually identified their RPPVSMs. The statistical results show that Caucasians tend to have more RPPVSMs than Asians and Latinos, and over 80 percent of middle to low density RPPVSM patterns are independently and uniformly distributed. Arfika Nurhudatiana, Adams Wai-Kin Kong, Keyan Matinpour, Siu-Yeung Cho, Noah Craft |
IJCB | 2 |
| 2011 | Using a Knowledge-Based Approach to Remove Blocking Artifacts in Skin Images for Forensic AnalysisabstractIdentifying individuals in evidence images, where their faces are covered or obstructed, is a challenging task. In the legal case, United States v. Michael Joseph Pepe (2008), Craft and Kong, who served as expert witnesses, used pigmented skin marks to identify a suspect in evidence images. Their expert opinions were challenged, partially because the blocking artifacts generated by the standard JPEG algorithm adversely affect the visibility of the small skin marks. In addition to this case, a huge amount of JPEG-compressed child pornography is posted online every day. Although many methods have been developed to remove blocking artifacts, they are ineffective for our target application. In this paper, a knowledge-based (KB) approach, which simultaneously removes JPEG blocking artifacts, and recovers skin features, is proposed. Given a training set containing both original and compressed skin images, the relationship between original blocks and compressed blocks can be established. This prior information is used to infer the original blocks of compressed evidence images. A Markov-model-based algorithm and a faster one-pass algorithm were developed to make inference, and a block synthesis algorithm was developed to handle the cases where compressed blocks are not contained in the training set. An indexing mechanism was also proposed to deal with large datasets efficiently. Extensive experiments were conducted on images with different characteristics and compression ratios. Both subjective and objective evaluations demonstrated that the KB approach is more effective than other methods. In summary, the KB approach is capable of removing blocking artifacts to recover useful skin features. Chaoying Tang, Adams Wai-Kin Kong, Noah Craft |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2010 | An alternative Gabor filtering schemeabstractAn elementary function that is now commonly referred to as Gabor filter was derived from uncertainty relation for information by Gabor to overcome the representation limit of Fourier analysis. For image-based applications (e.g. biometric recognition), researchers amend the weights of Gabor filters to produce zero DC (direct current) Gabor filters. This amendment can significantly change the shape of Gabor filters, when they use a special range of parameters. The aim of this paper is to develop a new Gabor filtering scheme to overcome this problem. Two different types of zero DC Gabor filters are compared with the proposed scheme on face recognition. FERET database is employed in the experiments. The Gabor phase from the proposed filtering scheme gives improvement in the range between 11% and 8.5%, and the performance of its magnitude is comparable with other schemes. Adams Wai-Kin Kong |
ICIP | 1 |
| 2010 | An Analysis of IrisCodeabstractIrisCode is an iris recognition algorithm developed in 1993 and continuously improved by Daugman. It has been extensively applied in commercial iris recognition systems. IrisCode representing an iris based on coarse phase has a number of properties including rapid matching, binomial impostor distribution and a predictable false acceptance rate. Because of its successful applications and these properties, many similar coding methods have been developed for iris and palmprint identification. However, we lack a detailed analysis of IrisCode. The aim of this paper is to provide such an analysis as a way of better understanding IrisCode, extending the coarse phase representation to a precise phase representation, and uncovering the relationship between IrisCode and other coding methods. Our analysis demonstrates that IrisCode is a clustering algorithm with four prototypes; the locus of a Gabor function is a 2-D ellipse with respect to a phase parameter and can be approximated by a circle in many cases; Gabor function can be considered as a phase-steerable filter and the bitwise hamming distance can be regarded as a bitwise phase distance. We also discuss the theoretical foundation of the impostor binomial distribution. We use this analysis to develop a precise phase representation which can enhance accuracy. Finally, we relate IrisCode and other coding methods. Adams Wai-Kin Kong, David Zhang 0001, Mohamed S. Kamel |
IEEE Trans. Image Process. | 1 |
| 2009 | A survey of palmprint recognition
Adams Wai-Kin Kong, David Zhang 0001, Mohamed S. Kamel |
Pattern Recognit. | 1 |
| 2008 | An evaluation of Gabor orientation as a feature for face recognitionabstractIdentifying a reliable feature is extremely important for all pattern recognition systems. The Gabor filter, which simultaneously captures spatial and frequency information, has been a vital component in numerous systems as a feature extractor. This filter produces three basic features - magnitude, phase, and orientation. Most face recognition methods based on Gabor filters use either the magnitude feature alone or a combination of the phase and magnitude features; very few are purely based on the phase feature, and the orientation feature is ignored. The aim of this paper is to evaluate these three basic features for face recognition using the FERET and AR face databases. The results show that the orientation feature is the most robust and distinctive feature, 20% and over 10% more accurate than the phase and magnitude features, respectively. Adams Wai-Kin Kong |
ICPR | 1 |
| 2008 | Three measures for secure palmprint identification
Adams Wai-Kin Kong, David Zhang 0001, Mohamed S. Kamel |
Pattern Recognit. | 1 |
| 2006 | An analysis of BioHashing and its variants
Adams Wai-Kin Kong, King Hong Cheung, David Zhang 0001, Mohamed S. Kamel, Jane You |
Pattern Recognit. | 1 |
| 2006 | Palmprint identification using feature-level fusion
Adams Wai-Kin Kong, David Zhang 0001, Mohamed S. Kamel |
Pattern Recognit. | 1 |
| 2006 | A study of identical twins' palmprints for personal verification
Adams Wai-Kin Kong, David Zhang 0001, Guangming Lu 0002 |
Pattern Recognit. | 1 |
| 2006 | Analysis of Brute-Force Break-Ins of a Palmprint Authentication SystemabstractBiometric authentication systems are widely applied because they offer inherent advantages over classical knowledge-based and token-based personal-identification approaches. This has led to the development of products using palmprints as biometric traits and their use in several real applications. However, as biometric systems are vulnerable to replay, database, and brute-force attacks, such potential attacks must be analyzed before biometric systems are massively deployed in security systems. This correspondence proposes a projected multinomial distribution for studying the probability of successfully using brute-force attacks to break into a palmprint system. To validate the proposed model, we have conducted a simulation. Its results demonstrate that the proposed model can accurately estimate the probability. The proposed model indicates that it is computationally infeasible to break into the palmprint system using brute-force attacks. Adams Wai-Kin Kong, David Zhang 0001, Mohamed S. Kamel |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2005 | An Analysis on Accuracy of Cancelable Biometrics Based on BioHashing
King Hong Cheung, Adams Wai-Kin Kong, David Zhang 0001, Mohamed S. Kamel, Jane You, Ho-Wang Lam |
KES (3) | 2 |
| 2005 | A new approach to appearance-based face recognitionabstractCurrent holistic appearance based face recognition methods require a high dimensional feature space to attain fruitful performance. In this paper, we have proposed a relatively low feature dimensional, template-matching scheme to cope with the transformed appearance-based face recognition problem. We use aggregated Gabor filter responses to represent face images. We investigated the effect of "duplicate" images (images from different sessions) and the effect of facial expressions. Our results indicate that the proposed method is more robust in recognizing "duplicate" images with variations in facial expression than the principal component analysis method. King Hong Cheung, Adams Wai-Kin Kong, Jane You, Qin Li 0001, David Zhang 0001, Prabir Bhattacharya |
SMC | 2 |
| 2005 | Online Palmprint Identification System for Civil Applications
David Zhang 0001, Guangming Lu 0002, Adams Wai-Kin Kong |
J. Comput. Sci. Technol. | 3 |
| 2004 | A study of aggregated 2D Gabor features on appearance-based face recognitionabstractExisting approaches to holistic appearance based face recognition require a high dimensional feature space to attain fruitful performance. We have proposed a relatively low feature dimensional scheme to deal with the face recognition problem. We use the aggregated responses of 2D Gabor filters to represent face images. We have investigated the effect of "duplicate" images and the effect of facial expressions. Our results show that the proposed method is more robust than the PCA-based method under varying facial expressions, especially in recognizing "duplicate" images. King Hong Cheung, Jane You, Adams Wai-Kin Kong, David Zhang 0001 |
ICIG | 3 |
| 2004 | On hierarchical palmprint coding with multiple features for personal identification in large databasesabstractAutomatic personal identification is a significant component of security systems with many challenges and practical applications. The advances in biometric technology have led to the very rapid growth in identity authentication. This paper presents a new approach to personal identification using palmprints. To tackle the key issues such as feature extraction, representation, indexing, similarity measurement, and fast search for the best match, we propose a hierarchical multifeature coding scheme to facilitate coarse-to-fine matching for efficient and effective palmprint verification and identification in a large database. In our approach, four-level features are defined: global geometry-based key point distance (Level-1 feature), global texture energy (Level-2 feature), fuzzy "interest" line (Level-3 feature), and local directional texture energy (Level-4 feature). In contrast to the existing systems that employ a fixed mechanism for feature extraction and similarity measurement, we extract multiple features and adopt different matching criteria at different levels to achieve high performance by a coarse-to-fine guided search. The proposed method has been tested in a database with 7752 palmprint images from 386 different palms. The use of Level-1, Level-2, and Level-3 features can remove candidates from the database by 9.6%, 7.8%, and 60.6%, respectively. For a system embedded with an Intel Pentium III processor (500 MHz), the execution time of the simulation of our hierarchical coding scheme for a large database with 10/sup 6/ palmprint samples is 2.8 s while the traditional sequential approach requires 6.7 s with 4.5% verification equal error rate. Our experimental results demonstrate the feasibility and effectiveness of the proposed method. Jane You, Adams Wai-Kin Kong, David Zhang 0001, King Hong Cheung |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2003 | An Integration of Principal Component Analysis and Self-Organizing Map for Effective Palmprint Retrieval
King Hong Cheung, Adams Wai-Kin Kong, Jane You, David Zhang 0001 |
CAINE | 2 |
| 2003 | On Hierarchical Palmprint Coding with Multi-Features for Personal Identification in Large Databases
Jane You, Adams Wai-Kin Kong, David Zhang 0001, King Hong Cheung |
CAINE | 2 |
| 2003 | Detecting Eyelash and Reflection for Accurate Iris SegmentationabstractAccurate iris segmentation is presented in this paper, which is composed of two parts, reflection detection and eyelash detection. Eyelashes are classified into two categories, separable and multiple. An edge detector is applied to detect separable eyelashes, and intensity variances are used to recognize multiple eyelashes. Reflection is also divided into two types, strong and weak. A threshold and statistical model is proposed to recognize the strong and weak reflection, respectively. We have developed an iris recognition approach for testing the effectiveness of the proposed segmentation method. The results show that the proposed method can reduce recognition error for the iris recognition approach. Adams Wai-Kin Kong, David Zhang 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2003 | Online Palmprint IdentificationabstractBiometrics-based personal identification is regarded as an effective method for automatically recognizing, with a high confidence, a person's identity. This paper presents a new biometric approach to online personal identification using palmprint technology. In contrast to the existing methods, our online palmprint identification system employs low-resolution palmprint images to achieve effective personal identification. The system consists of two parts: a novel device for online palmprint image acquisition and an efficient algorithm for fast palmprint recognition. A robust image coordinate system is defined to facilitate image alignment for feature extraction. In addition, a 2D Gabor phase encoding scheme is proposed for palmprint feature extraction and representation. The experimental results demonstrate the feasibility of the proposed system. David Zhang 0001, Adams Wai-Kin Kong, Jane You |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2003 | Palmprint feature extraction using 2-D Gabor filters
Adams Wai-Kin Kong, David Zhang 0001, Wenxin Li 0007 |
Pattern Recognit. | 1 |