Raphael C.-W. Phan

dblp:p/RCWPhan · also Raphael Chung-Wei Phan, Raphaël C.-W. Phan · DBLP profile ↗
← Back
137ranked-venue papers
27as first author
55since 2021 · last 2026
0000-0001-7448-4595ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 47 · 1 first-author · 30 since 2021Security and privacy · 42 · 14 first-author · 2 since 2021Artificial intelligence and machine learning · 21 · 1 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 3 first-author · 14 since 2021Computer networks · 8 · 4 first-author · 2 since 2021Systems, architecture and hardware · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 1 since 2021Theory of computation · 5 · 1 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 FLAG-4D: Flow-Guided Local-Global Dual-Deformation Model for 4D Reconstruction
abstract
We introduce FLAG-4D, a novel framework for generating novel views of dynamic scenes by reconstructing how 3D Gaussian primitives evolve through space and time. Existing methods typically rely on a single Multilayer Perceptron(MLP) to model temporal deformations, and they often struggle to capture complex point motions and fine-grained dynamic details consistently over time, especially from sparse input views. Our approach, FLAG-4D overcomes this by employing a dual-deformation network that dynamically warps a canonical set of 3D Gaussians over time into new positions and anisotropic shapes. This dual-deformation network consists of an Instantaneous Deformation Network (IDN) for modeling fine-grained, local deformations, and Global Motion Network (GMN) for capturing long-range dynamics, refined via mutual learning. To ensure these deformations are both accurate and temporally smooth, FLAG-4D incorporates dense motion features from a pretrained optical flow backbone. We fuse these motion cues from adjacent timeframes and use a deformation-guided attention mechanism to align this flow information with the current state of each evolving 3D Gaussian. Extensive experiments demonstrate that FLAG-4D achieves higher-fidelity and more temporally coherent reconstructions with finer detail preservation than state-of-the-art methods.
Guan Yuan Tan, Ngoc Tuan Vu, Arghya Pal, Sailaja Rajanala, Raphael C.-W. Phan, Mettu Srinivas, Chee-Ming Ting
AAAI5
2026 DepthPolyp: Pseudo-depth Guided Lightweight Segmentation for Real-Time Colonoscopy
Zhuoyu Wu, Wenhui Ou, Lexi Zhang, Pei-Sze Tan, Dongjun Wu, Junhe Zhao, Wenqi Fang, Raphael C.-W. Phan
ICPR (7)8
2026 A new time-decay radiomics integrated network (TRINet) for breast cancer risk prediction
abstract
To facilitate early detection of breast cancer, there is a need to develop risk prediction schemes that can prescribe personalized screening mammography regimens for women. In this study, we propose a new deep learning architecture called TRINet that implements time-decay attention to focus on recent mammographic screenings, as current models do not account for the relevance of newer images. We integrate radiomic features with an Attention-based Multiple Instance Learning (AMIL) framework to weigh and combine multiple views for better risk estimation. In addition, we introduce a continual learning approach with a new label assignment strategy based on bilateral asymmetry to make the model more adaptable to asymmetrical cancer indicators. Finally, we add a time-embedded additive hazard layer to perform dynamic, multi-year risk forecasting based on individualized screening intervals. We used two public datasets, namely 8528 patients from the American EMBED dataset and 8723 patients from the Swedish CSAW dataset in our experiments. Evaluation results on the EMBED test set show that our approach performs comparably with state-of-the-art models, achieving AUC scores of 0.851, 0.811, 0.796, 0.793, and 0.789 across 1-, 2-, to 5-year intervals, respectively. Our results underscore the importance of integrating temporal attention, radiomic features, time embeddings, bilateral asymmetry, and continual learning strategies, providing a more adaptive and precise tool for breast cancer risk prediction.
Hong Hui Yeoh, Fredrik Strand, Raphael C.-W. Phan, Kartini Rahmat, Maxine Tan
Medical Image Anal.3
2026 Causal-Ex: Causal graph-based micro and macro expression spotting
abstract
Detecting concealed emotions within apparently normal expressions is crucial for identifying potential mental health issues and facilitating timely support and intervention. The task of spotting macro- and micro-expressions involves predicting the emotional timeline within a video by identifying the onset (i.e., the beginning), apex (the peak of emotion), and offset (the end of emotion) frames of the displayed emotions. More particularly, closely monitoring the key emotion-conveying regions of the face; namely, the foundational muscle-movement cues known as facial action units (AUs)–greatly aids in the clear identification of micro-expressions. One major roadblock is the inadvertent introduction of biases into the training process, which degrades performance regardless of feature quality. Biases are spurious factors that falsely inflate or deflate performance metrics. For instance, the neural networks tend to falsely attribute certain AUs in specific facial regions to particular emotion classes, a phenomenon also termed as Inductive biases. To remove these false attributions, we must identify and mitigate biases that arise from mere correlation between some features and the output class labels. We hence introduce action-unit causal graphs. Unlike the traditional action-unit graph, which connects AUs based solely on spatial adjacency, the causal AU graph is derived from statistical tests and retains edges between AUs only when there is significant evidence that one AU causally influences another. Our model, named Causal-Ex ( Causal -based Ex pression spotting), employs a fast causal inference algorithm to construct a causal graph of facial region of interests (ROIs). This enables us to select causally relevant facial action units in the ROIs. Our work demonstrates improvement in overall F1-scores compared to state-of-the-art approaches with 0.388 on CAS(ME) 2 and 0.3701 on SAMM-Long Video datasets. Our code can be found at: https://github.com/noobasuna/causal_ex.git .
Pei-Sze Tan, Sailaja Rajanala, Arghya Pal, Raphael C.-W. Phan, Huey Fang Ong
Pattern Recognit. Lett.4
2026 A Unified Framework for Sparse Reconstruction via Preconditioning and Nonconvex Regularization
abstract
Compressed Sensing (CS) is an effective technique to recover sparse signals with fewer samples than what is required by the classical Shannon Nyquist sampling theorem. The sensing matrix, sparsifying transform, and sparse recovery algorithm are three key factors for accurate reconstruction in CS. Traditional CS uses a convex $l_{1}$-norm sparse regularizer which may lead to biased estimates and is suboptimal in promoting sparsity. Another challenge is the design of incoherent sensing matrices which is crucial for accurate sparse recovery. In this paper, we propose a novel CS framework combining a preconditioned sensing matrix and nonconvex regularization for improved sparse signal recovery. First, we formulate an optimization problem to find an incoherent sensing matrix via a preconditioner. It allows for a direct computation of the optimal preconditioner and preconditioned sensing matrix, simultaneously. Secondly, we consider a generalized CS model for signal recovery based on the incoherent sensing matrix and a nonconvex $\ell _{1/2}$-norm regularizer. We then derive an Alternating Direction Method of Multipliers (ADMM) algorithm to solve this nonconvex optimization problem. The proposed model is applied to sparse-view Computed Tomography (CT) reconstruction with highly-undersampled and noisy data. Qualitative and quantitative results show significantly better image reconstruction using the preconditioned sensing matrix and $\ell _{1/2}$ regularizer, compared to methods without preconditioning and using the $\ell _{1}$ regularizer.
Prasad Theeda, Fuad Noman, Arghya Pal, Raphael C.-W. Phan, Hernando C. Ombao, Chee-Ming Ting
IEEE J. Biomed. Health Informatics4
2025 GENIE: Socially Unbiased Generative Text-to-Image Editing
abstract
Generative diffusion models often exhibit societal biases in sensitive personal attributes such as age, gender, and race. In this work, we describe GENIE – a method to reduce such biases in a variety of classifier-free diffusion models used for image editing. Our method implicitly incorporates debiasing terms together with the user’s explicit edit instruction to reduce bias. This automatic method relieves the user from needing to modify edit instructions in order to avoid bias. Further, no additional training is needed. Experimental results are provided based on modifications to four diffusion models, namely InstructPix2Pix, Stable Diffusion 1.5, Stable Diffusion 2.1, and Stable Diffusion XL. We show that, on average, bias is reduced by 31% in gender, 15% in age, 39% in race.
Julia Kaiwen Lau, Raphael C.-W. Phan, Sailaja Rajanala, Ingemar J. Cox, Arghya Pal
ICASSP2
2025 Post-Hoc Adversarial Stickers Against Micro-Expression Leakage
abstract
Securing micro-expressions against leakage is crucial for privacy, as these subtle facial movements convey genuine emotions and are inherently personal. This study aims to protect micro-expression data from potential adversarial attacks, ensuring the preservation of individuals’ privacy and preventing unauthorized access or misuse of sensitive emotional information. Unlike traditional methods, which often require training and extensive access to models, this research introduces a novel post-hoc method that does not require additional training. We focus on physical adversarial attacks in micro-expression recognition, involving intentional manipulation of visual cues to deceive recognition systems and protect individual emotional privacy. Our approach leverages a causal discovery algorithm to identify causal relationships between facial parts, enabling rapid identification of the optimal locations for adversarial patches in frames with triggered micro-expressions. This method exhibits a more consistent attack success rate than randomly placed adversarial stickers, demonstrating effective generalization across different emotions, stickers, and models. Particularly relevant in scenarios with restricted access to the model, our technique requires only a single interaction during the attack process, highlighting its efficiency and minimal need for querying the target model. The proposed method effectively balances privacy protection with high generalization capability, setting a new standard for defending against adversarial threats in micro-expression recognition. The code is available at https://github.com/noobasuna/au-sticker.
Pei-Sze Tan, Sailaja Rajanala, Yee-Fan Tan, Arghya Pal, Chun-Ling Tan, Raphael C.-W. Phan, Huey Fang Ong
ICASSP6
2025 Guided Diffusion For Class-Conditioned Synthesis & Classification Of Microscopic Blood Cell Images
abstract
Microscopic visualization of diseased cells plays a vital role in the diagnosis and understanding of various medical conditions. Recent advances in deep learning generative models have shown remarkable potential as formidable tools for generating high-quality medical images. However, training these models generally requires large, annotated datasets, which are often costly and time-consuming to obtain. To overcome this challenge, we propose a fast-sampling guided score-based diffusion model with a classifier-free guidance strategy for class-conditioned generation of microscopic peripheral blood cell images. Our model achieves a Fréchet Inception Distance (FID) score of 10.24, demonstrating its ability to generate realistic synthetic blood cell images. Furthermore, our experimental results show that augmenting training data with these synthetic images significantly improves classification accuracy compared to relying solely on real data, highlighting the potential of synthetic data augmentation in hematology.
Kar-Ee Hoh, Junn Yong Loo, Yee-Fan Tan, Raphael C.-W. Phan, Chee-Ming Ting
ICIP4
2025 Polyfit generative model: can a group of lower-order polynomials generate high resolution diverse images?
abstract
Implicit neural representations (INRs) have recently gained popularity as a means to model images as continuous functions of spatial coordinates, synthesizing each pixel independently and yielding impressive results in tasks such as scene reconstruction and image generation. A notable advancement, PolyINR, utilizes element-wise multiplications between features and affine-transformed coordinates to achieve higher-order polynomial functions, eliminating the need for positional encodings. However, the finite encoding capacity of INRs, coupled with PolyINR’s recursive polynomial estimation, necessitates substantial training parameters, resulting in high computational costs and limiting applicability across diverse computer vision domains. In this work, we address these challenges by representing images as grids of smaller patches, within which we fit low-degree polynomials to capture local intensity variations. Our approach substantially reduces parameter requirements and computational demands. We evaluate our model qualitatively and quantitatively on large-scale datasets, ImageNet, CelebA, LSUN Bedroom, and Flower102; thus demonstrating competitive performance with state-of-the-art generative models, despite the absence of convolutional, normalization, or self-attention layers.
Arghya Pal, Ai-Fang Chai, Sailaja Rajanala, Raphael C.-W. Phan, Koksheik Wong, Chee-Ming Ting
IJCNN4
2025 Adaptive Graph Learning with Multi-graph Convolutions for Brain Disorder Classification
Fuad Noman, Raphael C.-W. Phan, Hernando C. Ombao, Chee-Ming Ting
MICCAI (12)2
2025 T2I-Diff: fMRI Signal Generation via Time-Frequency Image Transform and Classifier-Free Denoising Diffusion Models
Hwa Hui Tew, Junn Yong Loo, Yee-Fan Tan, Hernando C. Ombao, Fuad Noman, Raphael C.-W. Phan, Chee-Ming Ting
MICCAI (3)7
2025 Securing Face ID: Privacy Preservation for Non-Retentive Face Recognition System
abstract
Facial recognition technology is increasingly integrated into various applications. While face recognition systems have streamlined the authentication process and reduced the need for manual verification, the advancement in artificial intelligence (AI) generating realistic-looking media raises significant privacy concerns due to the potential misuse of biometric data. While current security protocols ensure data protection, storing biometric data in the system is a latent risk. In cases where the stored data is compromised, the users are susceptible to attacks such as face swapping, deepfake and identity theft. To address this, this paper presents a privacypreserving algorithm that omits the need to store human raw biometric data in the system. This is achieved by utilizing a new combination of Locality-Sensitive Hashing (LSH), salting, and RSA encryption for face recognition. The proposed method ensures data security by securely hashing and encrypting facial features while maintaining high recognition accuracy. The proposed framework is evaluated on the Labeled Faces in the Wild (LFW) and achieves a comparable performance with the state-of-the-art techniques.
Megan Chua, Chanelle Yue-Ting Yeow, Cassandra Xin-Yee Chwee, Shu-Min Leong, Raphael C.-W. Phan
TENCON5
2025 BitRelation: Exploring Bit-Level Dependencies in Neural Cryptanalysis
abstract
This paper applies Explainable Artificial Intelligence (XAI) to improve the interpretability of neural differential cryptanalysis on the SPECK cipher. We use Local Interpretable Model-agnostic Explanations (LIME) to analyse and visualise feature importance in neural distinguishers, giving signed contributions and absolute rankings. Signed contributions show whether, and how strongly, specific bit positions influence the model's decision, while absolute rankings reflect their importance regardless of sign. To study interactions beyond single bits, we introduce a Systematic Masking Approach to reveal relations among bits by testing if chosen combinations of masked bits alter classification accuracy. On Gohr's 8-round SPECK32/64 distinguisher, masking up to four-bit combinations shows that decisions involve multi-bit interactions rather than isolated single-bit effects. Although LIME highlights strong single-bit signals, masking reveals interaction patterns consistent with differential cryptanalysis. These findings clarify model behaviour in neural cryptanalysis and show XAI's value for exposing and visualising interaction structure in ciphertext features and decisions.
Yue-Tian Goi, Shu-Min Leong, Raphael C.-W. Phan, Ana Salagean, Shangqi Lai, Wei-Chuen Yau
TENCON3
2025 Res-SH: Unbiased Residual Learning for Self-Healing Interface Toughness Prediction with Limited Data
abstract
The development of self-healing materials is often hindered by the high costs and material waste associated with traditional characterization methods. Current approaches to toughness prediction, primarily based on convolutional neural networks (CNNs), are limited by their tendency to capture only surface-level features, which can lead to biased predictions. Moreover, working with small datasets, which is common in materials science, further increases the risk of biased training due to overfitting, posing a critical challenge to the reliability and generalizability of predictive models. This study introduces an unbiased residual learning framework designed explicitly for predicting self-healing interface toughness under limiteddata conditions. Our approach, ResNet-inspired approach for predicting self-healing material toughness, named Res-SH, used the power of residual networks to capture deeper, more complex patterns in the data, thereby addressing critical challenges in materials research. Res-SH minimises resource consumption and experimental overhead by focusing on unbiased learning, achieving accurate predictions with fewer training epochs and lower R2score and root mean square prediction errors compared to conventional CNN and lightweight model MobileNetv2. This novel framework provides a cost-effective and resource-efficient alternative to traditional material characterization methods, reducing material waste and accelerating the discovery and optimization of self-healing material systems.
Pei-Sze Tan, Karen Jia-Jun Koh, Sailaja Rajanala, Arghya Pal, Raphael C.-W. Phan, Nan Ze, Fuad Noman, Chee-Ming Ting, Norfadilah Dolmat, Nik Nur Wahidah Nik Hashim, Afidalina Tumian
TENCON5
2025 Disentangle Class Imbalance in Micro-Expression Recognition with Causal Structure Learning
abstract
Micro-expression recognition is a critical task in affective computing, yet it is often hindered by biases inherent in datasets, leading to skewed and unreliable outcomes. This work introduces a novel approach to disentangling bias in microexpression recognition using causal structure learning. By modeling causal relationships within the data, we identify and mitigate sources of bias that traditional machine learning models often overlook. Our framework integrates causal discovery techniques to uncover biased patterns in widely used micro-expression datasets. We employ debiasing strategies to enhance the fairness and accuracy of recognition models and generate counterfactual examples to address sensitive attributes such as gender and age, allowing us to observe the effects of imbalanced classes on classification results. Experiments conducted on baseline microexpression recognition models demonstrate comparable results after undersampling to create emotion class balance, revealing label bias in current training datasets including CASME2, SAMM, and SMIC. Further evaluation on balanced gender and age classes using generated counterfactual data as additional training instances showed performance improvements for the 4DME dataset.
Pei-Sze Tan, Sailaja Rajanala, Raphael C.-W. Phan
TENCON3
2025 Attack-SH: Adversarial Attacks on Self-Healing Material Properties Prediction Model
abstract
AI is now increasingly applied in diverse domains. The recent Nobel prizes for Physics and Chemistry awarded to computational scientists shows the significant impact that AI has on real-world scientific applications. Adversarial attacks pose a significant threat to the reliability of AI systems, particularly in high-stakes applications such as those in the materials sciences domain, which affect interactions with materials that exist in the real world. This paper examines the impact of two widely used adversarial attack approaches, notably the Fast Gradient Sign Method (FGSM) and Projected Gradient Descent (PGD), on a target AI model's performance for realworld materials. Experimental results demonstrate that both methods effectively degrade predictive accuracy, with PGD showing a more severe effect. Notably, high Structural Similarity Index (SSIM) scores across all perturbed samples suggest that the attacks introduce imperceptible changes, increasing their potential risk as such attacks are then undetectable. Further analysis using metrics such as Mean Squared Error (MSE), Adversarial MSE (A-MSE), and Relative Error Increase (REI) confirms substantial shifts in model output and reconstruction quality. These findings highlight critical vulnerabilities in current model architectures and emphasize the urgent need for more resilient defense strategies, such as adversarial training and input pre-processing.
Min-Xuan Tan, Pei-Sze Tan, Raphael C.-W. Phan, Shu-Min Leong
TENCON3
2025 PK-YOLO: Pretrained Knowledge Guided YOLO for Brain Tumor Detection in Multiplanar MRI Slices
abstract
Brain tumor detection in multiplane Magnetic Resonance Imaging (MRI) slices is a challenging task due to the various appearances and relationships in the structure of the multiplane images. In this paper, we propose a new You Only Look Once (YOLO)-based detection model that incorporates Pretrained Knowledge (PK), called PK-YOLO, to improve the performance for brain tumor detection in multiplane MRI slices. To our best knowledge, PK-YOLO is the first pretrained knowledge guided YOLO-based object detector. The main components of the new method are a pretrained pure lightweight convolutional neural network-based backbone via sparse masked modeling, a YOLO architecture with the pretrained backbone, and a regression loss function for improving small object detection. The pre-trained backbone allows for feature transferability of object queries on individual plane MRI slices into the model encoders, and the learned domain knowledge base can improve in-domain detection. The improved loss function can further boost detection performance on small-size brain tumors in multiplanar two-dimensional MRI slices. Experimental results show that the proposed PK-YOLO achieves competitive performance on the multiplanar MRI brain tumor detection datasets compared to state-of-the-art YOLO-like and DETR-like object detectors. The code is available at https://github.com/mkang315/PK-YOLO.
Ming Kang 0002, Fung Fung Ting, Raphael C.-W. Phan, Chee-Ming Ting
WACV3
2025 @LM DeceptionNet: A multimodal approach for efficient transfer learning-based deception detection
abstract
In terms of deception detection, traditional contact-based techniques often require collecting physiological signals, which can negatively impact device accuracy and participant comfort. While multimodal features extracted from audio and video modalities have been shown to outperform human observers on public datasets, the generalizability of existing audio and visual-based deception detection methods in different scenarios remains insufficiently explored. To narrow this gap, this work proposes a novel domain knowledge transfer learning method for deception detection in cross-scenario applications, which enhances its generalization and adaptability. Additionally, we designed a multimodal framework that filters out irrelevant information from other modalities when a particular modality yields reliable results, further improving overall system accuracy and robustness. We evaluate the proposed method on different public datasets, achieving promising generalizability results with consistent enhancements using four variations and networks. Apart from this, the proposed @LM DeceptionNet demonstrates better generalization capacity in computational efficiency, feature extraction, and adaptability compared to a larger model when employing fewer parameters.
Yuanya Zhuo, Vishnu Monn Baskaran, Lillian Yee Kiaw Wang, Raphael C.-W. Phan
Knowl. Based Syst.4
2024 BrainFC-CGAN: A Conditional Generative Adversarial Network for Brain Functional Connectivity Augmentation and Aging Synthesis
abstract
Brain functional connectivity (FC) changes are associated with neuropsychiatric disorders and other underlying factors, such as age and gender. Due to small training sample, data augmentation has been increasingly used for deep learning-based classification of brain FC. Although deep generative models could generate brain FCs to enhance downstream classification, most existing methods neglect the underlying factors involved in the generation process and fail to preserve the subject identity. We propose a novel brain FC conditional Generative Adversarial Network (GAN) called BrainFC-CGAN with specialized layers and filters to preserve the symmetry property and topological structure of brain FCs. We design a FC generator that captures the complex variations between brain FCs, ages, and health statuses to generate synthetic FCs that preserve the subject identity. We categorized true brain FCs into different age groups; an augmented age-specific dataset generated from BrainFC-CGAN is combined with the training set for classification. Experimental results on major depressive disorder (MDD) resting-state functional magnetic resonance imaging data show that the proposed method synthesizes realistic brain FCs of different target age groups, significantly improving downstream classification performance over baseline without augmentation, and also outperforming several state-of-the-art GANs.
Yee-Fan Tan, Junn Yong Loo, Chee-Ming Ting, Fuad Noman, Raphael C.-W. Phan, Hernando C. Ombao
ICASSP5
2024 Causally Uncovering Bias in Video Micro-Expression Recognition
abstract
Detecting microexpressions presents formidable challenges, primarily due to their fleeting nature and the limited diversity in existing datasets. Our studies find that these datasets exhibit a pronounced bias towards specific ethnicities and suffer from significant imbalances in terms of both class and gender representation among the samples. These disparities create fertile ground for various biases to permeate deep learning models, leading to skewed results and inadequate portrayal of specific demographic groups. Our research is driven by a compelling need to identify and rectify these biases within model architectures. To achieve this, we commence by constructing a causal graph that elucidates the intricate relationships between the model, input features, and training outcomes. This graphical representation forms the foundation for our analytical framework. Leveraging this causal framework, we conduct comprehensive case studies, employing counterfactuals as a diagnostic tool to unveil biases arising from dataset-induced class imbalances, gender inequalities, and variations in facial action units. Our final step involves a highly efficient counterfactual debiasing process, eliminating the necessity for additional data collection or model retraining. Our results showcase superior performance compared to state-of-the-art methods across the CASME II, SAMM, and SMIC datasets.
Pei-Sze Tan, Sailaja Rajanala, Arghya Pal, Shu-Min Leong, Raphael C.-W. Phan, Huey Fang Ong
ICASSP5
2024 Cafct-Net: A Cnn-Transformer Hybrid Network With Contextual And Attentional Feature Fusion For Liver Tumor Segmentation
abstract
Medical image semantic segmentation techniques can help identify tumors automatically from computed tomography (CT) scans. In this paper, we propose a Contextual and Attentional feature Fusions enhanced Convolutional Neural Network (CNN) and Transformer hybrid network (CAFCT-Net) for liver tumor segmentation. We incorporate three novel modules in the CAFCT-Net architecture: Attentional Feature Fusion (AFF), Atrous Spatial Pyramid Pooling (ASPP) of DeepLabv3, and Attention Gates (AGs) to improve contextual information related to tumor boundaries for accurate segmentation. Experimental results show that the proposed model achieves a mean Intersection over Union (IoU) of 76.54% and Dice coefficient of 84.29%, respectively, on the Liver Tumor Segmentation Benchmark (LiTS) dataset, outperforming pure CNN or Transformer methods, e.g., Attention U-Net and PVTFormer.
Ming Kang 0002, Chee-Ming Ting, Fung Fung Ting, Raphael C.-W. Phan
ICIP4
2024 CST-Yolo: A Novel Method For Blood Cell Detection Based On Improved Yolov7 And CNN-Swin Transformer
abstract
Blood cell detection is a typical small-scale object detection problem in computer vision. In this paper, we propose a CST-YOLO model for blood cell detection based on YOLOv7 architecture and enhance it with the CNN-Swin Transformer (CST), which is a new attempt at CNN-Transformer fusion. We also introduce three other useful modules: Weighted Efficient Layer Aggregation Networks (W-ELAN), Multiscale Channel Split (MCS), and Concatenate Convolutional Layers (CatConv) in our CST-YOLO to improve small-scale object detection precision. Experimental results show that the proposed CST-YOLO achieves 92.7%, 95.6%, and 91.1% mAP @ 0.5, respectively, on three blood cell datasets, outperforming state-of-the-art object detectors, e.g., RT-DETR, YOLOv5, and YOLOv7. Our code is available at https://github.com/mkang315/CST-YOLO.
Ming Kang 0002, Chee-Ming Ting, Fung Fung Ting, Raphael C.-W. Phan
ICIP4
2024 Deep Multi-Graph Embedded Clustering for Community Detection in FMRI Functional Brain Networks Across Individuals
abstract
Analyzing the community structure of brain networks provides new insights into human brain function. Existing studies broadly use conventional network clustering approaches. While graph neural networks have recently shown promise in modeling brain functional connectivity (FC) networks, their applications to brain community detection still need improvement and further refinement. Moreover, identifying common community structure while resolving the single-subject partitions across multiple individual networks remains underexplored. We propose a Deep Multi-Graph Embedded Clustering (DMGEC) framework to identify shared community partition in brain FC networks over a cohort of individuals. By incorporating the consensus information aggregated across network structures, DMGEC leverages a graph autoencoder to produce consensus-aware latent representations of individual networks, and applies deep embedded clustering on the multi-subject network representation to produce common community assignment of brain nodes. Simulations show superior community recovery by our method compared to conventional approaches, especially for networks with large number of communities. When applied to functional magnetic resonance imaging (fMRI) data, the DMGEC achieves outstanding alikeness over individual partitions, and uncovers group-level differences in brain community motifs between major depressive disorder patients and normal controls.
Kai-Jun See, Chee-Ming Ting, Fuad Noman, Junn Yong Loo, Yee-Fan Tan, Hernando C. Ombao, Raphael C.-W. Phan
ICIP7
2024 Dynamic MRI Reconstruction Using Low-Rank Plus Sparse Decomposition With Smoothness Regularization
abstract
The low-rank plus sparse (L+S) decomposition model has enabled better reconstruction of dynamic magnetic resonance imaging (dMRI) with separation into background (L) and dynamic (S) component. However, use of low-rank prior alone may not fully explain the slow variations or smoothness of the background part at the local scale. In this paper, we propose a smoothness-regularized L+S (SR-L+S) model for dMRI reconstruction from highly undersampled k-t-space data. We exploit joint low-rank and smooth priors on the background component of dMRI to better capture both its global and local temporal correlated structures. Extending the L+S formulation, the low-rank property is encoded by the nuclear norm, while the smoothness by a general $\ell_{p}$-norm penalty on the local differences of the columns of L. The additional smoothness regularizer can promote piecewise local consistency between neighboring frames. By smoothing out the noise and dynamic activities, it allows accurate recovery of the background part, and subsequently more robust dMRI reconstruction. Extensive experiments on multi-coil cardiac and synthetic data shows that the SR-L+S model outperforms several state-of-the-art methods in terms of recovery accuracy.
Chee-Ming Ting, Fuad Noman, Raphael C.-W. Phan, Hernando C. Ombao
ICIP3
2024 ActNetFormer: Transformer-ResNet Hybrid Method for Semi-supervised Action Recognition in Videos
Sharana Dharshikgan Suresh Dass, Hrishav Bakul Barua, Ganesh Krishnasamy, Raveendran Paramesran, Raphael C.-W. Phan
ICPR (15)5
2024 A Deep Probabilistic Spatiotemporal Framework for Dynamic Graph Representation Learning with Application to Brain Disorder Identification
Sin-Yee Yap, Junn Yong Loo, Chee-Ming Ting, Fuad Noman, Raphael C.-W. Phan, Adeel Razi, David L. Dowe
IJCAI5
2024 Video Deception Detection through the Fusion of Multimodal Feature Extraction and Neural Networks
abstract
Detecting deceptive behavior in videos is a complex task within several domains, including academic fraud assessment, commercial anti-fraud activities, judicial system evidence analysis, suspicious activity detection in security monitoring systems, and behavioral intent analysis in psychological research. In this study, we present a novel approach to video deception detection by integrating visual and audio models for deep feature fusion, primarily targeting advanced deception detection datasets. Our visual model leverages hierarchical image feature learning to enhance deceptive cue detection, complemented by an audio model that processes acoustic signals for precise speech pattern analysis. This multimodal method significantly boosts detection accuracy and lessens reliance on extensive training data. Notably, our visual model incorporates knowledge distillation technology, improving efficiency and reducing computational resource needs without compromising performance. We implement a transformer architecture using distillation tokens for effective learning and incorporate convolutional neural network insights to enrich our model’s interpretative capabilities. Experimental results demonstrate that our approach surpasses existing technologies in various standards and scenarios, offering enhanced deception recognition capabilities and addressing the challenge of limited training data.
Yuanya Zhuo, Vishnu Monn Baskaran, Lillian Yee Kiaw Wang, Raphael C.-W. Phan
IJCNN4
2024 BGF-YOLO: Enhanced YOLOv8 with Multiscale Attentional Feature Fusion for Brain Tumor Detection
Ming Kang 0002, Chee-Ming Ting, Fung Fung Ting, Raphael C.-W. Phan
MICCAI (8)4
2024 Unveiling the Black Box: Neural Cryptanalysis with XAI
abstract
At CRYPTO'19, Gohr[1] presented ResNet-based neural distinguishers (ND) for the round-reduced SPECK32/64 cipher. However, due to the black-box use of such deep learning models, it is hard for humans to understand why these distinguishers work, impeding advancements in cryptanalytic knowledge. In this work, we aim to effectively adapt eXplainable Artificial Intelligence (XAI) techniques, notably Local Interpretable Model-Agnostic Explanations (LIME) and Shapley Additive Explanations (SHAP), to gain a detailed understanding of the important features useful in Gohr's neural distinguishers.
Yue-Tian Goi, Shu-Min Leong, Raphael C.-W. Phan, Shangqi Lai, Ana Salagean
SMC3
2024 Distingusic: Distinguishing Synthesized Music from Human
abstract
In this paper we focus on a problem that is increasingly plaguing the music industry; to a large extent due to the proliferation of generative AI models that enable the generation of new realistic and indistinguishable content for diverse modalities: text, image, audio, video. We address this problem from the perspective of audio watermarking; to our best knowledge, this is the first-known watermarking based approach to solve the problem of distinguishing realistic songs synthesized from generative AI models from real songs sung by humans. In more detail, our approach specifically utilizes the SHA-256 hash function, Singular Value Decomposition (SVD) and Discrete Wavelet Transform (DWT) for robust audio watermarking of synthesized songs. Before embedding, the audio is subjected to an attack phase to pinpoint less vulnerable regions for QR watermark placement. During the embedding process, the audio chunks first undergo a 1-level Discrete Wavelet Transform (DWT), and then the resulting approximate coefficients go through Singular Value Decompo-sition (SVD). Additionally, the watermarked array is subjected to SHA-256 hashing for collision-resistant conciseness, which is subsequently embedded into the singular values of the audio. Experimental findings demonstrate the superiority of our method over existing audio watermarking approaches under various signal attack scenarios.
Zi Qian Yong, Shu-Min Leong, Sailaja Rajanala, Arghya Pal, Raphael C.-W. Phan
SMC5
2024 ASF-YOLO: A novel YOLO model with attentional scale sequence fusion for cell instance segmentation
abstract
We propose a novel Attentional Scale Sequence Fusion based You Only Look Once (YOLO) framework (ASF-YOLO) which combines spatial and scale features for accurate and fast cell instance segmentation. Built on the YOLO segmentation framework, we employ the Scale Sequence Feature Fusion (SSFF) module to enhance the multiscale information extraction capability of the network, and the Triple Feature Encoder (TFE) module to fuse feature maps of different scales to increase detailed information. We further introduce a Channel and Position Attention Mechanism (CPAM) to integrate both the SSFF and TFE modules, which focus on informative channels and spatial position-related small objects for improved detection and segmentation performance. Experimental validations on two cell datasets show remarkable segmentation accuracy and speed of the proposed ASF-YOLO model. It achieves a box mAP of 0.91, mask mAP of 0.887, and an inference speed of 47.3 FPS on the 2018 Data Science Bowl dataset, outperforming the state-of-the-art methods. The source code is available at https://github.com/mkang315/ASF-YOLO.
Ming Kang 0002, Chee-Ming Ting, Fung Fung Ting, Raphael C.-W. Phan
Image Vis. Comput.4
2024 Emotion-specific AUs for micro-expression recognition
abstract
Abstract The Facial Action Coding System (FACS) comprehensively describes facial expressions with facial action units (AUs). It is a well-used technique by researchers in emotions research to understand human emotions better. Most micro-expression datasets provide FACS-coded AU ground truths corresponding to micro-expressions classes. It is commonly accepted in computer vision-based emotions research that certain emotions are reliably revealed when specific combinations of AUs occur. However, the reliability of the ground truth AUs in the micro-expression datasets is lower than that of normal expressions, as they have lower AU intensities. Moreover, these micro-expression datasets only report the overall reliability of all AUs. It could not be identified which AUs had been accurately coded. This work aims to revisit the ground truth AUs of popular micro-expression datasets, namely CASME II, SAMM and CAS(ME) $$^2$$ 2 , and inspect whether any AUs crucial for micro-expression recognition may need to be reconsidered. This paper also provides a detailed AU analysis which yields new AU-based RoIs for each dataset. These new RoIs improve the micro-expression recognition performances compared to the baselines considered in this work. The proposed RoIs for CASME II, SAMM and CAS(ME) $$^2$$ 2 improve the recognition rates by $$2\%$$ 2 % , $$1\%$$ 1 % and $$4\%$$ 4 % , respectively, when compared with the existing RoIs.
Shu-Min Leong, Raphael C.-W. Phan, Vishnu Monn Baskaran
Multim. Tools Appl.2
2024 Graph Autoencoders for Embedding Learning in Brain Networks and Major Depressive Disorder Identification
abstract
Brain functional connectivity (FC) networks inferred from functional magnetic resonance imaging (fMRI) have shown altered or aberrant brain functional connectome in various neuropsychiatric disorders. Recent application of deep neural networks to connectome-based classification mostly relies on traditional convolutional neural networks (CNNs) using input FCs on a regular Euclidean grid to learn spatial maps of brain networks neglecting the topological information of the brain networks, leading to potentially sub-optimal performance in brain disorder identification. We propose a novel graph deep learning framework that leverages non-Euclidean information inherent in the graph structure for classifying brain networks in major depressive disorder (MDD). We introduce a novel graph autoencoder (GAE) architecture, built upon graph convolutional networks (GCNs), to embed the topological structure and node content of large fMRI networks into low-dimensional representations. For constructing the brain networks, we employ the Ledoit-Wolf (LDW) shrinkage method to efficiently estimate high-dimensional FC metrics from fMRI data. We explore both supervised and unsupervised techniques for graph embedding learning. The resulting embeddings serve as feature inputs for a deep fully-connected neural network (FCNN) to distinguish MDD from healthy controls (HCs). Evaluating our model on resting-state fMRI MDD dataset, we observe that the GAE-FCNN outperforms several state-of-the-art methods for brain connectome classification, achieving the highest accuracy when using LDW-FC edges as node features. The graph embeddings of fMRI FC networks also reveal significant group differences between MDD and HCs. Our framework demonstrates the feasibility of learning graph embeddings from brain networks, providing valuable discriminative information for diagnosing brain disorders.
Fuad Noman, Chee-Ming Ting, Hakmook Kang, Raphael C.-W. Phan, Hernando C. Ombao
IEEE J. Biomed. Health Informatics4
2023 $\mathrm{C}\eta\iota \text{DAE}$: Cryptographically Distinguishing Autoencoder for Cipher Cryptanalysis
abstract
We propose a new autoencoder (AE) construction$\mathrm{C}\eta\iota \text{DAE}$(Cryptographically Distinguishing AE) based on a novel loss formulation to solve the cipher cryptanalysis distinguishing problem in the domain of cryptology. Vanilla AE and variational AE are unable to address this problem as they are designed to draw new samples which are either similar to the input sample or are from the same distribution. Such generated samples do not facilitate the cryptanalysis task. We show that our AE construction enables the discovery of cipher distinguishers, which are the fundamental building blocks that make or break new cipher design proposals. This also answers an open question on the applicability of autoencoders for cipher cryptanalysis; as to date, only discriminative models have been applied for cryptanalysis problems. To the best of our knowledge,$\mathrm{C}\eta\iota \text{DAE}$is the first-known generative model designed to solve crypt-analysis problems. We apply our$\mathrm{C}\eta\iota \text{DAE}$model to discover distinguishing properties for up to 10 rounds of the NSA-designed Speck32/64 cipher that allows to distinguish it from a random permutation. This contrasts with the best-known machine learning-discovered neural distinguisher in the literature that covers up to 8 rounds of Speck32/64. Unlike these recent related work which leverage on white box analysis and human-guided differential or linear analysis in order for machine learning models to be applicable, our$\mathrm{C}\eta\iota \text{DAE}$distinguisher does not require prior human cryptanalytic knowledge. This motivates the new direction of human-unsupervised machine learning-based cryptanalysis techniques.
Raphael C.-W. Phan, Arghya Pal, Koksheik Wong, Sailaja Rajanala
GLOBECOM1
2023 Self Supervised Bert for Legal Text Classification
abstract
Critical BERT-based text classification tasks, such as legal text classification, require huge amounts of accurately labeled data. Legal text classification faces two trivial problems: labeling legal data is a sensitive process and can only be carried out by skilled professionals, and legal text is prone to privacy issues hence not all the data can be made available in the public domain. This means that we have limited diversity in the textual data, and to account for this data paucity, we propose a self-supervision approach to train Legal-BERT classifiers. We use the BERT text classifier’s knowledge of the class boundaries and perform gradient ascent w.r.t. class logits. Synthetic latent texts are generated through activation maximization. The main advantages over existing SOTAs are that our model: is easy to train, does not require much data but instead uses the synthesized data as fake samples; has less variance that helps to generate texts with good sample quality and diversity. We show the efficacy of the proposed method on the ECHR Violation (Multi-Label) Dataset and the Over-ruling Task Dataset.
Arghya Pal, Sailaja Rajanala, Raphael C.-W. Phan, Koksheik Wong
ICASSP3
2023 USURP: Universal Single-Source Adversarial Perturbations on Multimodal Emotion Recognition
abstract
The field of affective computing has progressed from traditional unimodal analysis to more complex multimodal analysis due to the proliferation of videos posted online. Multimodal learning has shown remarkable performance in emotion recognition tasks, but its robustness in an adversarial setting remains unknown. This paper investigates the robustness of multimodal emotion recognition models against worst-case adversarial perturbations on a single modality. We found that standard multimodal models are susceptible to single-source adversaries and can be easily fooled by perturbations on any single modality. We draw some key observations that serve as guidelines for designing universal adversarial attacks on multimodal emotion recognition models. Motivated by these findings, we propose a novel universal single-source adversarial perturbations framework on multimodal emotion recognition models: USURP. Through our analysis of adversarial robustness, we demonstrate the necessity of studying adversarial attacks on multimodal models. Our experimental results show that the proposed USURP method achieves high attack success rates and significantly improves adversarial transferability in multimodal settings. The observations and novel attack methods presented in this paper provide a new understanding of the adversarial robustness of multimodal models, contributing to their safe and reliable deployment in more real- world scenarios.
Yin Yin Low, Raphael C.-W. Phan, Arghya Pal, Xiaojun Chang
ICIP2
2023 A Unified Framework for Static and Dynamic Functional Connectivity Augmentation for Multi-Domain Brain Disorder Classification
abstract
Deep learning (DL) methods recently show promise on accurate brain disorder classification using functional connectivity (FC) estimated from functional magnetic resonance imaging (fMRI). However, DL model building can be hindered by small sample-size settings of fMRI. Moreover, most studies utilize either static (sFC) or dynamic FC (dFC) for classification. We propose a unified framework for data augmentation of both sFC and dFC for multi-domain joint classification of brain disorders. We exploit generative adversarial networks (GAN) to synthesize realistic FCs for data augmentation. Notably, we adopted the TimeGAN for dFC generation that can capture temporal dependencies in real dFC, and the GR-SPD-GAN for sFC generation that preserves the spatial connectivity structure. We further develop BrainFusionNet - a specialized DL model for multi-domain FC that simultaneously learns embedded features from both sFC and dFC to provide complementary spatio-temporal information for downstream classification. The synthetic FC data are augmented in training data to improve the BrainFusionNet performance and generalizability. Experimental results on major depressive disorder (MDD) identification using resting-state fMRI show substantial improvement in classification accuracy by our framework, outperforming competing models without FC augmentation and using sFC or dFC features alone.
Yee-Fan Tan, Chee-Ming Ting, Fuad Noman, Raphael C.-W. Phan, Hernando C. Ombao
ICIP4
2023 Scheduling Dependent Batching Tasks
abstract
We formulate and analyze the following problem of scheduling dependent batching tasks on a common resource. Each task is associated with an execution time window. Some can be executed simultaneously in batch, while others need exclusively use of the resource. A task may depend on a set of tasks such that it becomes executable only if all its ancestor tasks are completed. We look for a task scheduling policy maximizing the system reward. We investigate both the offline and online settings by focusing on a typical scenario where the task dependency forms a tree or a forest. In both settings, we formally establish the hardness of the scheduling problem by showing that the offline scheduling is NP-hard and the online counterpart admits no scheduling policy with finite competitive ratio. We then develop approximation scheduling algorithms for both cases with deterministic worst-case performance guarantee in terms of system utility. We further conduct numeric experiments to evaluate our algorithms under a variety of parameter settings to demonstrate the effectiveness of our scheduling algorithms.
Hehuan Shi, Lin Chen 0002, Raphael C.-W. Phan
ICPP4
2023 Cross-domain Transfer Learning and State Inference for Soft Robots via a Semi-supervised Sequential Variational Bayes Framework
abstract
Recently, data-driven models such as deep neural networks have shown to be promising tools for modelling and state inference in soft robots. However, voluminous amounts of data are necessary for deep models to perform effectively, which requires exhaustive and quality data collection, particularly of state labels. Consequently, obtaining labelled state data for soft robotic systems is challenged for various reasons, including difficulty in the sensorization of soft robots and the inconvenience of collecting data in unstructured environments. To address this challenge, in this paper, we propose a semi-supervised sequential variational Bayes (DSVB) framework for transfer learning and state inference in soft robots with missing state labels on certain robot configurations. Considering that soft robots may exhibit distinct dynamics under different robot configurations, a feature space transfer strategy is also incorporated to promote the adaptation of latent features across multiple configurations. Unlike existing transfer learning approaches, our proposed DSVB employs a recurrent neural network to model the nonlinear dynamics and temporal coherence in soft robot data. The proposed framework is validated on multiple setup configurations of a pneumatic-based soft robot finger. Experimental results on four transfer scenarios demonstrate that DSVB performs effective transfer learning and accurate state inference amidst missing state labels.
Shageenderan Sapai, Junn Yong Loo, Ze Yang Ding, Chee Pin Tan, Raphael C.-W. Phan, Vishnu Monn Baskaran, Surya Girinatha Nurzaman
ICRA5
2023 Schatten p-norm based Image-to-Video Adaptation for Video Action Recognition
abstract
Human action recognition has been receiving extensive interest among researchers from the computer vision community. Numerous successful action recognition techniques have demonstrated the effectiveness of learning action knowledge from still images or motion videos. The relevant action information learned for the same action via various media types, such as images or videos, may be correlated. Nonetheless, less attention has been paid to adapting the action knowledge from images to videos to enhance action recognition performance in videos. Furthermore, most existing video action recognition methods suffer from insufficient labeled training videos. Overfitting could be an issue in these circumstances; hence, action recognition performance can be inhibited. This paper proposes an adaptation framework to transfer knowledge from images to videos for action recognition. A multi-task learning framework is designed to optimize the image and video domain classifiers jointly. The general Schatten p-norm is applied to the classifiers to mine the shared knowledge between these two domains. In this way, our framework can learn the correlated action semantics by leveraging the shared components of labeled images and videos. Our proposed approach can fully use the action knowledge from images and performs better in the case of poor and limited video data compared with the existing state-of-the-art action recognition techniques.
Sharana Dharshikgan Suresh Dass, Ganesh Krishnasamy, Raveendran Paramesran, Raphael C.-W. Phan
IJCNN4
2023 RCS-YOLO: A Fast and High-Accuracy Object Detector for Brain Tumor Detection
Ming Kang 0002, Chee-Ming Ting, Fung Fung Ting, Raphael C.-W. Phan
MICCAI (4)4
2023 Attacking Mouse Dynamics Authentication Using Novel Wasserstein Conditional DCGAN
abstract
Behavioral biometrics is an emerging trend due to their cost-effectiveness and non-intrusive implementations that support remote access for user identification. This is the case especially in recent times of social distancing and working from home arrangements, where online attendance is the preferred option in contrast to physical presence. In this work, we explore the limitations of mouse dynamics authentication by impersonating legitimate user mouse action sequences. Specifically, towards that aim, we develop a novel generative WC-DCGAN model to generate highly accurate fake user action sequences. We apply our WC-DCGAN to this problem and show that it causes the target classifier can be tricked into identifying a fraudster as a legitimate user. WC-DCGAN has several benefits, including: achieving dominated convergence, hence implying the existence of solutions and optimal discriminator regardless of data and generator distributions; and acting as an unsupervised model for a fixed class label and generator. Experiments are conducted to verify these points. Subsequently, we analyzed the cause of misclassifications, and propose a novel mouse dynamics strategy that offers much tighter authentication with significant reductions in misclassification events.
Arunava Roy, Koksheik Wong, Raphael C.-W. Phan
IEEE Trans. Inf. Forensics Secur.3
2022 Post-Quantum Verifiable Random Function from Symmetric Primitives in PoS Blockchain
Maxime Buser, Rafael Dowsley, Muhammed F. Esgin, Shabnam Kasra Kermanshahi, Veronika Kuchta, Joseph K. Liu, Raphael C.-W. Phan, Zhenfei Zhang
ESORICS (1)7
2022 AdverFacial: Privacy-Preserving Universal Adversarial Perturbation Against Facial Micro-Expression Leakages
abstract
Privacy safeguards are crucial, notably now with increased virtual conferencing usage during the Covid pandemic. In contrast to conventional facial expressions that are visually obvious to humans, micro-expressions are involuntary and transient facial expressions, commonly manifested involuntarily when we aim to withhold our emotions. Advanced micro-expression recognition techniques exist that can reveal the genuine emotions that people attempt to conceal, thus threatening individual emotional privacy, as fundamental human rights would dictate that one should have a choice of what emotion is being shown or not shown. We propose the novel universal adversarial perturbation-based approach - AdverFacial - for privacy concealment against automated micro-expression analysis via deep learning techniques. We derive the optimal strategy to achieve micro-expression misclassification with a high success rate, low perceptibility and cross neural network transferability. We perform experiments on two popular datasets with state-of-the-art microexpression spotting and recognition models and demonstrate our approach’s effectiveness in emotional concealment.
Yin Yin Low, Angeline Tanvy, Raphael C.-W. Phan, Xiaojun Chang
ICASSP3
2022 GraphEx: Facial Action Unit Graph for Micro-Expression Classification
abstract
Facial micro-expressions are crucial cues for expressing human emotions. Existing works have shown substantial progress in detecting micro-expressions for various applications in the computer vision field. However, it is still onerous for existing methods to handle and interpret micro-expressions efficiently. This paper proposes a deep learning-based approach leveraging spatio-temporal and graph representation learning for micro-expression classification. We design a novel Spatial-Temporal Info Extraction Network (STIENet) for learning facial appearance and muscle motion from high dimensional video clip frames and summarizes them into more meaningful feature maps. We construct an action unit (AU) relation graph to further represent the AU co-occurrence in the same micro-expression video clip. A graph neural network (GNN) is used to learn AU-related graph embedding for the downstream classification task. Performance evaluation on two mainstream micro-expression datasets, i.e., CASME II and SAMM, show that the proposed framework outperforms other state-of-the-art methods for micro-expression classification.
Shu-Min Leong, Fuad Noman, Raphael C.-W. Phan, Vishnu Monn Baskaran, Chee-Ming Ting
ICIP3
2022 Graph Autoencoder-Based Embedded Learning in Dynamic Brain Networks for Autism Spectrum Disorder Identification
abstract
Recent applications of pattern recognition techniques to brain connectome-based classification focus on static functional connectivity (FC) neglecting the dynamics of FC over time, and use input connectivity matrices on a regular Euclidean grid. We exploit the graph convolutional networks (GCNs) to learn irregular structural patterns in brain FC networks and propose extensions to capture dynamic changes in network topology. We develop a dynamic graph autoencoder (DyGAE)-based framework to leverage the time-varying topological structures of dynamic brain networks for identification of autism spectrum disorder (ASD). The framework combines a GCN-based DyGAE to encode individual-level dynamic networks into time-varying low-dimensional network embeddings, and classifiers based on weighted fully-connected neural network (FCNN) and long short-term memory (LSTM) to facilitate dynamic graph classification via the learned spatial-temporal information. Evaluation on a large ABIDE resting-state functional magnetic resonance imaging (rs-fMRI) dataset shows that our method outperformed state-of-the-art methods in detecting altered FC in ASD. Dynamic FC analyses with DyGAE learned embeddings also reveal apparent group difference between ASD and healthy controls in network profiles and switching dynamics of brain states.
Fuad Noman, Sin-Yee Yap, Raphael C.-W. Phan, Hernando C. Ombao, Chee-Ming Ting
ICIP3
2022 Persistent Items Tracking in Large Data Streams Based on Adaptive Sampling
abstract
We address the problem of persistent item tracking in large-scale data streams. A persistent item refers to the one that persists to occur in the stream over a long timespan. Tracking persistent items is an important and pivotal functionality for many networking and computing applications as persistent items, though not necessarily contributing significantly to the data volume, may convey valuable information on the data pattern about the stream. The state-of-the-art solutions of tracking persistent items require to know the monitoring time horizon to set the sampling rate. This limitation is further accentuated when we need to track the persistent items in recent w slots where w can be any value between 0 and T to support different monitoring granularity. Motivated by this limitation, we develop a persistent item tracking algorithm that can function without knowing the monitoring time horizon beforehand, and can thus track persistent items up to the current time t or within a certain time window at any moment. Our central technicality is adaptively reducing the sampling rate such that the total memory overhead can be limited while still meeting the target tracking accuracy. Through both theoretical and empirical analysis, we fully characterize the performance of our proposition.
Lin Chen 0002, Raphael C.-W. Phan, Dan Huang 0001
INFOCOM2
2022 Guess-It-Generator: Generating in a Lewis Signaling Framework through Logical Reasoning
abstract
Human minds spontaneously integrate two inherited cognitive capabilities: perception and reasoning to accomplish cognitive tasks such as problem solving, imagination, and causation. It is observed in the primate brains that perception offers the assistance required for problem comprehension, whilst the reasoning elucidates upon the facts recovered during perception in order to make a decision. The field of artificial intelligence (AI) thus considers perception and reasoning as two complementary areas that are realized by machine learning and logic programming, respectively. In this work, we propose a generative model using a collaborative guessing game of the kind first introduced by David Lewis in his famous work called the Lewis signaling game that is synonymous with the "20 Questions'' game. Our proposed model, Guess-It-Generator (GIG) is a collaborative framework that engages two recurrent neural networks in a guessing game. GIG unifies perception and reasoning with a view to generating labeled images by capturing, (X, y), the underlying density of a data distribution, i.e. (X, y) - p(X, y). An encoder attends to a region of the input image and encodes that onto a latent variable that acts as a perception signal to a decoder. In contrast, the decoder leverages on the perception signals to guess the image and verifies the guess by reasoning with logical facts derived from the domain knowledge. Our experiments and comprehensive studies on seven datasets: PCAM, Chest-Xray-14, FIRE, HAM10000 from the medical domain, and CIFAR 10, LSUN, ImageNet, among standard benchmark datasets, show significant promise for the proposed method.
Arghya Pal, Sailaja Rajanala, Raphael C.-W. Phan, Koksheik Wong
ACM Multimedia3
2022 DeSCoVeR: Debiased Semantic Context Prior for Venue Recommendation
abstract
We present a novel semantic context prior-based venue recommendation system that uses only the title and the abstract of a paper. Based on the intuition that the text in the title and abstract have both semantic and syntactic components, we demonstrate that a joint training of a semantic feature extractor and syntactic feature extractor collaboratively leverages meaningful information that helps to provide venues for papers. The proposed methodology that we call DeSCoVeR at first elicits these semantic and syntactic features using a Neural Topic Model and text classifier respectively. The model then executes a transfer learning optimization procedure to perform a contextual transfer between the feature distributions of the Neural Topic Model and the text classifier during the training phase. DeSCoVeR also mitigates the document-level label bias using a Causal back-door path criterion and a sentence-level keyword bias removal technique. Experiments on the DBLP dataset show that DeSCoVeR outperforms the state-of-the-art methods.
Sailaja Rajanala, Arghya Pal, Manish Singh 0002, Raphael C.-W. Phan, Koksheik Wong
SIGIR4
2022 A comprehensive overview of Deepfake: Generation, detection, datasets, and opportunities
Jia Wen Seow, Mei Kuan Lim, Raphael C.-W. Phan, Joseph K. Liu
Neurocomputing3
2022 Invisible emotion magnification algorithm (IEMA) for real-time micro-expression recognition with graph-based features
Adamu Muhammad Buhari, Chee-Pun Ooi, Vishnu Monn Baskaran, Raphael C.-W. Phan, Koksheik Wong, Wooi-Haw Tan
Multim. Tools Appl.4
2021 Synthesize-It-Classifier: Learning a Generative Classifier Through Recurrent Self-Analysis
abstract
We show the generative capability of an image classifier network by synthesizing high-resolution, photo-realistic, and diverse images at scale. The overall methodology, called Synthesize-It-Classifier (STIC), does not require an explicit generator network to estimate the density of the data distribution and sample images from that, but instead uses the classifier’s knowledge of the boundary to perform gradient ascent w.r.t. class logits and then synthesizes images using the Gram Matrix Metropolis Adjusted Langevin Algorithm (GRMALA) by drawing on a blank canvas. During training, the classifier iteratively uses these synthesized images as fake samples and re-estimates the class boundary in a recurrent fashion to improve both the classification accuracy and quality of synthetic images. The STIC shows that mixing of the hard fake samples (i.e. those synthesized by the one-hot class conditioning), and the soft fake samples (which are synthesized as a convex combination of classes, i.e. a mixup of classes [36]) improves class interpolation. We demonstrate an Attentive-STIC network that shows iterative drawing of synthesized images on the ImageNet dataset that has thousands of classes. In addition, we introduce the synthesis using a class conditional score classifier (Score-STIC) instead of a normal image classifier and show improved results on several real world datasets, i.e. ImageNet, LSUN and CIFAR 10.
Arghya Pal, Raphael C.-W. Phan, Koksheik Wong
CVPR2
2021 Network Reconfiguration via Diversity: Theoretical Foundation and Algorithm Design
abstract
Moving Target Defense (MTD) is a powerful weapon to mitigate cyber attacks by increasing the attacker's efforts and complexity in fulfilling its goal. One effective technique of MTD is to deploy diverse implementations and configurations to provide equivalent functionality so as to increase the network resilience. In this paper, we investigate the algorithmic aspect of employing diversity to cause the network to be the most resilient possible. Specifically, we study two algorithmic optimization problems of both theoretical and practical importance: (1) given the network topology, how to assign different variants to different network nodes so as to maximize the network resilience; (2) when the variant assignment is fixed, but the network topology is configurable, what is the optimal topology maximizing the network resilience. We mathematically formulate the problems of variant assignment and network topology configuration and develop efficient algorithms, which can serve as design guidelines in the deployment of diversity-based MTD techniques to enhance the security of the network.
Lin Chen 0002, Raphael C.-W. Phan
VTC Fall2
2021 Faceless identification based on temporal strips
Shu-Min Leong, Raphael C.-W. Phan, Vishnu Monn Baskaran, Chee-Pun Ooi
Multim. Tools Appl.2
2021 Strengthening speech content authentication against tampering
Raphael C.-W. Phan, Yin Yin Low, Koksheik Wong, Kazuki Minemura
Speech Commun.1
2020 Optimized IoT Cryptoprocessor Based on QC-MPDC Key Encapsulation Mechanism
abstract
The key encapsulation mechanism (KEM) is an important cryptographic tool to protect communication in the Internet of Things (IoT). In the near future, classical algorithms used to construct KEMs, such as RSA and elliptic curve cryptography, will be vulnerable to attacks from quantum computers. Recently, Yamada et al. proposed the quasicyclic medium density parity check (QC-MDPC) KEM, which is considered one of the most advanced code-based cryptosystems to resist quantum attacks. In this article, an optimized implementation of QC-MDPC KEM for IoT applications is presented. Our main contributions are threefold: 1) the fastest QC-MDPC McEliece decryption in field-programmable gate array (FPGA); 2) the first QC-MDPC KEM implementation in FPGA; and 3) the first iteration count attack-resistant QC-MDPC decoder in FPGA. To improve the decryption speed, we introduce a novel customized rotation engine (CRE) and incorporated several recent techniques reported in the literature, including adaptive threshold and Hamming weight estimation. The best-achieved throughput in our implementation on Xilinx Virtex 7 FPGA is 12.7% faster than the state-of-the-art result reported by Heyse et al. The proposed CRE was then integrated with QC-MDPC KEM to produce a fast and secure KEM. Furthermore, to prevent timing attacks demonstrated recently, a constant-time implementation of the QC-MDPC McEliece decoder was presented.
Jun-Hoe Phoon, Wai-Kong Lee, Denis Chee-Keong Wong, Wun-She Yap, Bok-Min Goi, Raphael C.-W. Phan
IEEE Internet Things J.6
2020 Reduced contact lifting of latent fingerprints from curved surfaces
Mohammad Mogharen Askarin, Koksheik Wong, Raphael C.-W. Phan
J. Inf. Secur. Appl.3
2020 Cryptanalysis of genetic algorithm-based encryption scheme
Kuan-Wai Wong, Wun-She Yap, Denis Chee-Keong Wong, Raphael C.-W. Phan, Bok-Min Goi
Multim. Tools Appl.4
2019 Dual-stream Shallow Networks for Facial Micro-expression Recognition
abstract
Micro-expressions are spontaneous, brief and subtle facial muscle movements that exposes underlying emotions. Motivated by recent exploits into deep learning for micro-expression analysis, we propose a lightweight dual-stream shallow network in the form of a pair of truncated CNNs with heterogeneous input features. The merging of the convolutional features allows for discriminative learning of micro-expression classes stemming from both streams. Using activation heatmaps, we further demonstrate that salient facial areas are well emphasized, and correspond closely to relevant action units belonging to emotion classes. We empirically validate the proposed network on three benchmark databases, obtaining state-of-the-art performance on the CASME II and SAMM while remaining competitive on the SMIC. Further observations point towards the sufficiency of utilizing shallower deep networks for micro-expression recognition.
Huai-Qian Khor, John See, Sze-Teng Liong, Raphael C.-W. Phan, Weiyao Lin
ICIP4
2019 Terabit encryption in a second: Performance evaluation of block ciphers in GPU with Kepler, Maxwell, and Pascal architectures
abstract
Summary With the emergence of IoT and cloud computing technologies, massive data are generated from various applications everyday and communicated through the Internet. Secure communication is essential to protect these data from malicious attacks. Block ciphers are one mechanism to offer such protection but unfortunately involve intensive computations that can be performance bottlenecks to the servers, especially when the data center needs to handle thousands of concurrent transactions. In this paper, we investigate the feasibility of the GPU as an accelerator to perform high‐speed encryption in server environments. We present optimized implementations of a conventional block cipher (AES) and new lightweight block ciphers (LEA, Chaskey, SIMON, SPECK, and SIMECK) across three new GPU architectures (Kepler, Maxwell, and Pascal). For AES, we improve the fine‐grain implementation by utilizing the warp shuffle instruction available in these three new GPU architectures, which yield a 6%‐16% improvement over the previous implementations. For LEA, Chaskey, SIMON, SPECK, and SIMECK, we first analyze why they cannot have efficient fine‐grain implementations in the GPU and then present our optimization techniques, which are able to achieve impressive encryption speeds of 1.912, 637, 1.485, 2.291, and 1.478 Tb/s, respectively, in GTX1080.
Wai-Kong Lee, Bok-Min Goi, Raphael C.-W. Phan
Concurr. Comput. Pract. Exp.3
2019 Signature Gateway: Offloading Signature Generation to IoT Gateway Accelerated by GPU
abstract
The emergence of Internet of Things (IoT) brings us the possibility to form a well connected network for ubiquitous sensing, intelligent analysis, and timely actuation, which opens up many innovative applications in our daily life. To secure the communication between sensor nodes, gateway devices and cloud servers, cryptographic algorithms (e.g., digital signature, block cipher, and hash function) are widely used. Although cryptographic algorithms are effective in preventing malicious attacks, they involve heavy computation that may not be executed efficiently in resource constraint sensor nodes. In particular, the authentication of a sensor node is usually performed through a digital signature (e.g., RSA and elliptic curve cryptography), which can be slow when executed on a microcontroller. In this paper, an IoT architecture that offloads the digital signature generation to a nearby signature gateway equipped with graphic processing unit (GPU) accelerator are proposed. The communication process for signature offloading, together with optimized implementation techniques for RSA in signature gateway, are also presented in this paper. We have evaluated two different ways to implement modular exponentiation in RSA, namely residue number system and multiprecision montgomery multiplication (MPMM). The experimental results show that our RSA implementation using MPMM is 10.1% faster than the best RSA implementation in GPU. Our proposed IoT architecture with signature gateway can successfully reduce the burden of sensor nodes to generate signatures, at the same time preserve the ability to authenticate the sensor nodes.
Chin-Chen Chang 0001, Wai-Kong Lee, Yanjun Liu 0002, Bok-Min Goi, Raphael C.-W. Phan
IEEE Internet Things J.5
2019 Paradigm Shifts in Cryptographic Engineering
abstract
The papers in this special section identify paradigm shifts in cryptographic engineering. Modern cryptography is almost five decades old, and we have seen some interesting breakthroughs throughout its relatively young history. These include the engineering foundations of symmetric key cryptography and block ciphers in particular (like the DES and the AES), the discovery of public-key cryptography, protocols for secure computations for general and specific tasks using interactions among parties and tools like homomorphic encryption (this line of work has been enhancing the use of cryptography beyond secure messaging into secure computing).
Raphael C.-W. Phan, Moti Yung
IEEE Trans. Dependable Secur. Comput.1
2018 Enriched Long-Term Recurrent Convolutional Network for Facial Micro-Expression Recognition
abstract
Facial micro-expression (ME) recognition has posed a huge challenge to researchers for its subtlety in motion and limited databases. Recently, handcrafted techniques have achieved superior performance in micro-expression recognition but at the cost of domain specificity and cumbersome parametric tunings. In this paper, we propose an Enriched Long-term Recurrent Convolutional Network (ELRCN) that first encodes each micro-expression frame into a feature vector through CNN module(s), then predicts the micro-expression by passing the feature vector through a Long Short-term Memory (LSTM) module. The framework contains 2 different network variants: (1) Channel-wise stacking of input data for spatial enrichment, (2) Feature-wise stacking of features for temporal enrichment. We demonstrate that the proposed approach is able to achieve reasonably good performance, without data augmentation. In addition, we also present ablation studies conducted on the framework and visualizations of what CNN "sees" when predicting the micro-expression classes.
Huai-Qian Khor, John See, Raphael C.-W. Phan, Weiyao Lin
FG3
2018 Micro-Expression Motion Magnification: Global Lagrangian vs. Local Eulerian Approaches
abstract
Micro-expressions are difficult to spot but are utterly important for engaging in a conversation or negotiation. Through motion magnification, these expressions become much more distinguishable and easily recognized. This work proposes Global Lagrangian Motion Magnification (GLMM) for consistent exaggeration of facial expressions and dynamics across a whole video. As the proposal takes an opposite approach to a previous pivotal work, i.e. local Amplitude-based Eulerian Motion Magnification (AEMM). GLMM and AEMM are theoretically analyzed for potential advantages and disadvantages, especially with respect to how magnified noise and distortions are dealt with. Then, both GLMM and AEMM are empirically evaluated and compared using the CASME II micro-expression corpus.
Anh Cat Le Ngo, Alan Johnston, Raphael C.-W. Phan, John See
FG3
2018 Cryptography and Future Security
Jongsung Kim, Hongjun Wu 0001, Raphael C.-W. Phan
Discret. Appl. Math.3
2018 Separable authentication in encrypted HEVC video
Yiqi Tew, Koksheik Wong, Raphael C.-W. Phan, King Ngi Ngan
Multim. Tools Appl.3
2018 Less is more: Micro-expression recognition from video using apex frame
Sze-Teng Liong, John See, Koksheik Wong, Raphael C.-W. Phan
Signal Process. Image Commun.4
2017 Higher order differentiation over finite fields with applications to generalising the cube attack
abstract
Higher order differentiation was introduced in a cryptographic context by Lai. Several attacks can be viewed in the context of higher order differentiations, amongst them the cube attack of Dinur and Shamir and the AIDA attack of Vielhaber. All of the above have been developed for the binary case. We examine differentiation in larger fields, starting with the field $$\mathrm {GF}(p)$$ of integers modulo a prime p, and apply these techniques to generalising the cube attack to $$\mathrm {GF}(p)$$ . The crucial difference is that now the degree in each variable can be higher than one, and our proposed attack will differentiate several times with respect to each variable (unlike the classical cube attack and its larger field version described by Dinur and Shamir, both of which differentiate at most once with respect to each variable). Connections to the Moebius/Reed Muller Transform over $$\mathrm {GF}(p)$$ are also examined. Finally we describe differentiation over finite fields $$\mathrm {GF}(p^s)$$ with $$p^s$$ elements and show that it can be reduced to differentiation over $$\mathrm {GF}(p)$$ , so a cube attack over $$\mathrm {GF}(p^s)$$ would be equivalent to cube attacks over $$\mathrm {GF}(p)$$ .
Ana Salagean, Richard Winter, Matei Mandache-Salagean, Raphael C.-W. Phan
Des. Codes Cryptogr.4
2017 Effective recognition of facial micro-expressions with video motion magnification
Yandan Wang, John See, Yee-Hui Oh, Raphael C.-W. Phan, Yo Rahul, Huo-Chong Ling, Su-Wei Tan, Xujie Li 0002
Multim. Tools Appl.4
2017 Sparsity in Dynamics of Spontaneous Subtle Emotions: Analysis and Application
abstract
Subtle emotions are present in diverse real-life situations: in hostile environments, enemies and/or spies maliciouslyconceal their emotions as part of their deception; in life-threatening situations, victims under duress have no choice but to withhold theirreal feelings; in the medical scene, patients with psychological conditions such as depression could either be intentionally orsubconsciously suppressing their anguish from loved ones. Under such circumstances, it is often crucial that these subtle emotions arerecognized before it is too late. These spontaneous subtle emotions are typically expressed through micro-expressions, which are tiny,sudden and short-lived dynamics of facial muscles; thus, such micro-expressions pose a great challenge for visual recognition. Theabrupt but significant dynamics for the recognition task are temporally sparse while the rest, i.e. irrelevant dynamics, are temporallyredundant. In this work, we analyze and enforce sparsity constraints to learn significant temporal and spectral structures whileeliminating irrelevant facial dynamics of micro-expressions, which would ease the challenge in the visual recognition of spontaneoussubtle emotions. The hypothesis is confirmed through experimental results of automatic spontaneous subtle emotion recognition withseveral sparsity levels on CASME II and SMIC, the two well-established and publicly available spontaneous subtle emotion databases.The overall performances of the automatic subtle emotion recognition are boosted when only significant dynamics of the originalsequences are preserved.
Anh Cat Le Ngo, John See, Raphael C.-W. Phan
IEEE Trans. Affect. Comput.3
2017 A Novel Sketch Attack for H.264/AVC Format-Compliant Encrypted Video
abstract
In this paper, we propose a novel sketch attack for H.264 advanced video coding (H.264/AVC) format-compliant encrypted video. We briefly describe the notion of sketch attack, review the conventional sketch attacks designed for discrete cosine transform (DCT)-based compressed image, and identify their shortcomings when applied to attack compressed video. Specifically, the conventional DCT-based sketch attacks are incapable in sketching outlines for inter frame, which is deployed to significantly reduce temporal redundancy in video compression. To sketch directly from inter frame, we put forward a sketch attack by considering the partially decoded information of the H.264/AVC compressed video, namely, the number of bits spent on coding a macroblock. To evaluate the sketch image, we consider the Canny edge map as the ideal outline image. Experiments are conducted to verify the performance of the proposed sketch attack using ICADR2013, High Efficiency Video Coding dash, and Xiph video data sets. Results suggest that the proposed sketch attack can generate the outline image of the original frame for not only intra frame but also inter frame.
Kazuki Minemura, Koksheik Wong, Raphael C.-W. Phan, Kiyoshi Tanaka
IEEE Trans. Circuits Syst. Video Technol.3
2016 Eulerian emotion magnification for subtle expression recognition
abstract
Subtle emotions are expressed through tiny and brief movements of facial muscles, called micro-expressions; thus, recognition of these hidden expressions is as challenging as inspection of microscopic worlds without microscopes. In this paper, we show that through motion magnification, subtle expressions can be realistically exaggerated and become more easily recognisable. We magnify motions of facial expressions in the Eulerian perspective by manipulating their amplitudes or phases. To evaluate effects of exaggerating facial expressions, we use a common framework (LBP-TOP features and SVM classifiers) to perform 5-class subtle emotion recognition on the CASME II corpus, a spontaneous subtle emotion database. According to experimental results, significant improvements in recognition rates of magnified micro-expressions over normal ones are confirmed and measured. Furthermore, we estimate upper bounds of effective magnification factors and empirically corroborate these theoretical calculations with experimental data.
Anh Cat Le Ngo, Yee-Hui Oh, Raphael C.-W. Phan, John See
ICASSP3
2016 Intrinsic two-dimensional local structures for micro-expression recognition
abstract
An elapsed facial emotion involves changes of facial contour due to the motions (such as contraction or stretch) of facial muscles located at the eyes, nose, lips and etc. Thus, the important information such as corners of facial contours that are located in various regions of the face are crucial to the recognition of facial expressions, and even more apparent for micro-expressions. In this paper, we propose the first known notion of employing intrinsic two-dimensional (i2D) local structures to represent these features for micro-expression recognition. To retrieve i2D local structures such as phase and orientation, higher order Riesz transforms are employed by means of monogenic curvature tensors. Experiments performed on micro-expression datasets show the effectiveness of i2D local structures in recognizing micro-expressions.
Yee-Hui Oh, Anh Cat Le Ngo, Raphael C.-W. Phan, John See, Huo-Chong Ling
ICASSP3
2016 Smart, secure and seamless access control scheme for mobile devices
abstract
Smart devices capture users' activity such as unlock failures, application usage, location and proximity of devices in and around their surrounding environment. This activity information varies between users and can be used as digital fingerprints of the users' behaviour. Traditionally, users are authenticated to access restricted data using long term static attributes such as password and roles. In this paper, in order to allow secure and seamless data access in mobile environment, we combine both the user behaviour captured by the smart device and the static attributes to develop a novel access control technique. Security and performance analyses show that the proposed scheme substantially reduces the computational complexity while enhances the security compared to the conventional schemes.
Yo Rahul, Muttukrishnan Rajarajan, Raphael C.-W. Phan
ICC3
2016 Multi-layer authentication scheme for HEVC video based on embedded statistics
Yiqi Tew, Koksheik Wong, Raphael C.-W. Phan, King Ngi Ngan
J. Vis. Commun. Image Represent.3
2016 On the effective subkey space of some image encryption algorithms using external key
Wun-She Yap, Raphael C.-W. Phan, Bok-Min Goi, Wei-Chuen Yau, Swee-Huay Heng
J. Vis. Commun. Image Represent.2
2016 Spontaneous subtle expression detection and recognition based on facial strain
Sze-Teng Liong, John See, Raphael C.-W. Phan, Yee-Hui Oh, Anh Cat Le Ngo, Koksheik Wong, Su-Wei Tan
Signal Process. Image Commun.3
2015 HEVC video authentication using data embedding technique
abstract
A HEVC video authentication scheme by utilizing data embedding technique is proposed. The concept of authentication, layout and implementation are described under the HEVC standard. The authentication scheme includes weight generation, video feature extraction and two layers of authentication. Simulation results confirm that the overall perceptual video quality is maintained after the insertion of authentication code into commonly considered classes of video sequence. By analyzing the behavior of video tampering within and across video slices, the proposed authentication scheme is able to detect the tampered region and verify the integrity of the video in question.
Yiqi Tew, Koksheik Wong, Raphael C.-W. Phan
ICIP3
2015 Comparison of Cube Attacks Over Different Vector Spaces
Richard Winter, Ana Salagean, Raphael C.-W. Phan
IMACC3
2015 Cryptanalysis of a certificateless identification scheme
abstract
ABSTRACT In 2013, Dehkordi and Alimoradi proposed a certificateless identification scheme using supersingular elliptic curves. This proposal came independent of the parallel work of Chin et al. in proposing the first known security models for certificateless identification with provable security. In this paper, we show that there are some design flaws in the Dehkordi–Alimoradi scheme, which lead one to conclude that their scheme is insecure. Copyright © 2014 John Wiley & Sons, Ltd.
Ji-Jian Chin, Rouzbeh Behnia, Swee-Huay Heng, Raphael C.-W. Phan
Secur. Commun. Networks4
2014 Spontaneous Subtle Expression Recognition: Imbalanced Databases and Solutions
Anh Cat Le Ngo, Raphael C.-W. Phan, John See
ACCV (4)2
2014 LBP with Six Intersection Points: Reducing Redundant Information in LBP-TOP for Micro-expression Recognition
Yandan Wang, John See, Raphael C.-W. Phan, Yee-Hui Oh
ACCV (1)3
2014 Privacy-Preserving Multi-Class Support Vector Machine for Outsourcing the Data Classification in Cloud
abstract
Emerging cloud computing infrastructure replaces traditional outsourcing techniques and provides flexible services to clients at different locations via Internet. This leads to the requirement for data classification to be performed by potentially untrusted servers in the cloud. Within this context, classifier built by the server can be utilized by clients in order to classify their own data samples over the cloud. In this paper, we study a privacy-preserving (PP) data classification technique where the server is unable to learn any knowledge about clients’ input data samples while the server side classifier is also kept secret from the clients during the classification process. More specifically, to the best of our knowledge, we propose the first known client-server data classification protocol using support vector machine. The proposed protocol performs PP classification for both two-class and multi-class problems. The protocol exploits properties of Pailler homomorphic encryption and secure two-party computation. At the core of our protocol lies an efficient, novel protocol for securely obtaining the sign of Pailler encrypted numbers.
Yo Rahul, Raphael C.-W. Phan, Suresh Veluru 0001, K. Cumanan, Muttukrishnan Rajarajan
IEEE Trans. Dependable Secur. Comput.2
2014 Privacy-Preserving Clinical Decision Support System Using Gaussian Kernel-Based Classification
abstract
A clinical decision support system forms a critical capability to link health observations with health knowledge to influence choices by clinicians for improved healthcare. Recent trends toward remote outsourcing can be exploited to provide efficient and accurate clinical decision support in healthcare. In this scenario, clinicians can use the health knowledge located in remote servers via the Internet to diagnose their patients. However, the fact that these servers are third party and therefore potentially not fully trusted raises possible privacy concerns. In this paper, we propose a novel privacy-preserving protocol for a clinical decision support system where the patients' data always remain in an encrypted form during the diagnosis process. Hence, the server involved in the diagnosis process is not able to learn any extra knowledge about the patient's data and results. Our experimental results on popular medical datasets from UCI-database demonstrate that the accuracy of the proposed protocol is up to 97.21% and the privacy of patient data is not compromised.
Yo Rahul, Suresh Veluru 0001, Raphael C.-W. Phan, Jonathon A. Chambers, Muttukrishnan Rajarajan
IEEE J. Biomed. Health Informatics3
2013 Vickrey-Clarke-Groves for privacy-preserving collaborative classification
Anastasia Panoui, Sangarapillai Lambotharan, Raphael C.-W. Phan
FedCSIS3
2013 Efficient Generation of Elementary Sequences
David Gardner, Ana Salagean, Raphael C.-W. Phan
IMACC3
2013 An Efficient and Provably Secure Certificateless Identification Scheme
Ji-Jian Chin, Raphael C.-W. Phan, Rouzbeh Behnia, Swee-Huay Heng
SECRYPT2
2013 On the Security of the XOR Sandwiching Paradigm for Multiple Keyed Block Ciphers
Ruth Ng Ii-Yung, Khoongming Khoo, Raphael C.-W. Phan
SECRYPT3
2013 An authentication framework for peer-to-peer cloud
abstract
Cloud computing provides on demand computation and storage services delivered via applications, system software and hardware rendered as services. Due to its on demand nature, it has high variable workloads and requires real-time efficiency and availability. Most cloud computing systems use a centralised model to provision services, but reliance on a central entity to control scheduling decision and maintain all cloud hosts may constitute a computing bottleneck. A system failure will cause service outage, sometimes for a few hours as had happened before. In addition, the central entity needs to support heavy workloads in terms of service provisioning to all resource hosts. These issues can be addressed by distributing cloud resources using structured peer-to-peer (P2P) overlay networks as was recently proposed. However these proposals do not examine potential security issues of a P2P-based cloud, one of them being how peers verify the identities of one another over a decentralised setting. Therefore we propose an authentication framework for P2P cloud consisting of various approaches for authenticating entities and messages. The framework combines cryptographic primitives and security mechanisms proposed for existing structured P2P network.
Geong Sen Poh, Mohd Amril Nurman Mohd Nazir, Bok-Min Goi, Syh-Yuan Tan, Raphael C.-W. Phan, Maryam Safiyah Shamsudin
SIN5
2013 On the security of a modified Beth identity-based identification scheme
Ji-Jian Chin, Syh-Yuan Tan, Swee-Huay Heng, Raphael C.-W. Phan
Inf. Process. Lett.4
2013 Facial Expression Recognition in the Encrypted Domain Based on Local Fisher Discriminant Analysis
abstract
Facial expression recognition forms a critical capability desired by human-interacting systems that aim to be responsive to variations in the human's emotional state. Recent trends toward cloud computing and outsourcing has led to the requirement for facial expression recognition to be performed remotely by potentially untrusted servers. This paper presents a system that addresses the challenge of performing facial expression recognition when the test image is in the encrypted domain. More specifically, to the best of our knowledge, this is the first known result that performs facial expression recognition in the encrypted domain. Such a system solves the problem of needing to trust servers since the test image for facial expression recognition can remain in encrypted form at all times without needing any decryption, even during the expression recognition process. Our experimental results on popular JAFFE and MUG facial expression databases demonstrate that recognition rate of up to 95.24 percent can be achieved even in the encrypted domain.
Yo Rahul, Raphael C.-W. Phan, Jonathon A. Chambers, David J. Parish
IEEE Trans. Affect. Comput.2
2012 Index Tables of Finite Fields and Modular Golomb Rulers
Ana Salagean, David Gardner, Raphael C.-W. Phan
SETA3
2012 A Formally Verified Device Authentication Protocol Using Casper/FDR
abstract
For communication in Next Generation Networks, highly-developed mobile devices will enable users to store and manage a lot of credentials on their terminals. Furthermore, these terminals will represent and act on behalf of users when accessing different networks and connecting to a wide variety of services. In this situation, it is essential for users to trust their terminals and for all transactions using them to be secure. This paper analyses a number of the Authentication and Key Agreement protocols between the users and mobile terminals, then proposes a novel device authentication protocol. The proposed protocol is analysed and verified using a formal methods approach based on Casper/FDR compiler.
Mahdi Aiash, Glenford E. Mapp, Raphael C.-W. Phan, Aboubaker Lasebae, Jonathan Loo
TrustCom3
2012 Efficient encryption with keyword search in mobile networks
abstract
ABSTRACT On these days, users tend to access to online content via mobile devices, for example, e‐mails. Because these devices have constrained resources, users may wish to instruct e‐mail gateways to search through new e‐mails and only download those corresponding to particular keywords, such as “urgent.” Yet, this searching should not compromise the user's privacy. A public key encryption with keyword search (PEKS) scheme achieves both these requirements. Most PEKS schemes are constructed on the basis of bilinear pairings. Recently, Khader proposed the first PEKS scheme that does not require bilinear pairings and is provably indistinguishable chosen‐keyword attack (IND‐CKA) secure in the standard model. Such a scheme is more efficient than pairing‐based ones. In this paper, we show a drawback of Khader's scheme in that it depends on an unnecessary security assumption: Its IND‐CKA security requires its underlying identity‐based encryption building block to be indistinguishable chosen‐ciphertext attack secure. We construct a more efficient PEKS scheme that achieves the same level of PEKS security as Khader's but that only requires the underlying identity‐based encryption to be indistinguishable chosen‐plaintext attack secure. We give a direct proof that the proposed scheme is IND‐CKA secure. Our scheme outperforms other recent PEKS schemes in literature. Copyright © 2012 John Wiley & Sons, Ltd.
Wei-Chuen Yau, Swee-Huay Heng, Syh-Yuan Tan, Bok-Min Goi, Raphael C.-W. Phan
Secur. Commun. Networks5
2011 Cryptanalysis of a Provably Secure Cross-Realm Client-to-Client Password-Authenticated Key Agreement Protocol of CANS '09
Wei-Chuen Yau, Raphael C.-W. Phan, Bok-Min Goi, Swee-Huay Heng
CANS2
2011 On the Stability of m-Sequences
Alex J. Burrage, Ana Salagean, Raphael C.-W. Phan
IMACC3
2011 Linear complexity for sequences with characteristic polynomial ƒv
abstract
We present several generalisations of the Games-Chan algorithm. For a fixed monic irreducible polynomial f we consider the sequences s that have as characteristic polynomial a power of f. We propose an algorithm for computing the linear complexity of s given a full (not necessarily minimal) period of s. We give versions of the algorithm for fields of characteristic 2 and for arbitrary finite characteristic p, the latter generalising an algorithm of Kaida et al. We also propose an algorithm which computes the linear complexity given only a finite portion of s (of length greater than or equal to the linear complexity), generalising an algorithm of Meidl. All our algorithms have linear computational complexity. The algorithms for computing the linear complexity when a full period is known can be further generalised to sequences for which it is known a priori that the irreducible factors of the minimal polynomial belong to a given small set of polynomials.
Alex J. Burrage, Ana Salagean, Raphael C.-W. Phan
ISIT3
2011 On the Security of a Hybrid SVD-DCT Watermarking Method Based on LPSNR
Huo-Chong Ling, Raphael C.-W. Phan, Swee-Huay Heng
PSIVT (1)2
2011 Notions and relations for RKA-secure permutation and function families
Jongsung Kim, Jaechul Sung, Ermaliza Razali, Raphael C.-W. Phan, Marc Joye
Des. Codes Cryptogr.4
2011 On the cryptanalysis of the hash function Fugue: Partitioning and inside-out distinguishers
Jean-Philippe Aumasson, Raphael C.-W. Phan
Inf. Process. Lett.2
2011 VLSI Characterization of the Cryptographic Hash Function BLAKE
abstract
Cryptographic hash functions are used to protect information integrity and authenticity in a wide range of applications. After the discovery of weaknesses in the current deployed standards, the U.S. Institute of Standards and Technology started a public competition to develop the future standard SHA-3, which will be implemented in a multitude of environments, after its selection in 2012. In this paper, we investigate high-speed and low-area hardware architectures of one of the 14 “second-round” candidates in this competition: BLAKE. VLSI performance results of the proposed high-speed designs indicate a throughput improvement between 16% and 36% compared to the current standard SHA-2. Additionally, we propose a compact implementation of BLAKE with memory optimization that fits in 0.127 mm2of a 0.18 μ m CMOS. Measurements reveal a minimal power dissipation of 9.59 μW/MHz at 0.65 V, which suggests that BLAKE is suitable for resource-limited systems.
Luca Henzen, Jean-Philippe Aumasson, Willi Meier, Raphael C.-W. Phan
IEEE Trans. Very Large Scale Integr. Syst.4
2011 Non-repudiable authentication and billing architecture for wireless mesh networks
Raphael C.-W. Phan
Wirel. Networks1
2010 Integral Distinguishers of Some SHA-3 Candidates
Marine Minier, Raphael C.-W. Phan, Benjamin Pousse
CANS2
2010 Quasi-Linear Cryptanalysis of a Secure RFID Ultralightweight Authentication Protocol
Pedro Peris-Lopez, Julio César Hernández Castro, Raphael C.-W. Phan, Juan Tapiador, Tieyan Li
Inscrypt3
2010 Security Models for Heterogeneous Networking
Glenford E. Mapp, Mahdi Aiash, Aboubaker Lasebae, Raphael C.-W. Phan
SECRYPT4
2009 Improved Cryptanalysis of Skein
Jean-Philippe Aumasson, Çagdas Çalik, Willi Meier, Onur Özen, Raphael C.-W. Phan, Kerem Varici
ASIACRYPT5
2009 On Hashing with Tweakable Ciphers
abstract
Cryptographic hash functions are often built on block ciphers in order to reduce the security analysis of the hash to that of the cipher, and to minimize the hardware size. Well known hash constructs are used in international standards like MD5 and SHA-1. Recently, researchers proposed new modes of operations for hash functions to protect against generic attacks, and it remains open how to base such functions on block ciphers. An attracting and intuitive choice is to combine previous constructions with tweakable block ciphers. We investigate such constructions, and show the surprising result that combining a provably secure mode of operation with a provably secure tweakable cipher does not guarantee the security of the constructed hash function. In fact, simple attacks can be possible when the interaction between secure components leaves some additional "freedom" to an adversary. Our techniques are derived from the principle of slide attacks, which were introduced for attacking block ciphers.
Raphael C.-W. Phan, Jean-Philippe Aumasson
ICC1
2009 Analysis of Two Pairing-Based Three-Party Password Authenticated Key Exchange Protocols
abstract
Password-Authenticated Key Exchange (PAKE) protocols allow parties to share secret keys in an authentic manner based on an easily memorizable password. Recently, Nam et al. showed that a provably secure three-party password-based authenticated key exchange protocol using Weil pairing by Wen et al. is vulnerable to a man-in-the-middle attack. In doing so, Nam et al. showed the flaws in the proof of Wen et al. and described how to fix the problem so that their attack no longer works. In this paper, we show that both Wen et al. and Nam et al. variants fall to key compromise impersonation by any adversary. Our results underline the fact that although the provable security approach is necessary to designing PAKEs, gaps still exist between what can be proven and what are really secure in practice.
Raphael C.-W. Phan, Wei-Chuen Yau, Bok-Min Goi
NSS1
2009 Cryptanalysis of a New Ultralightweight RFID Authentication Protocol—SASI
abstract
Since RFID tags are ubiquitous and at times even oblivious to the human user, all modern RFID protocols are designed to resist tracking so that the location privacy of the human RFID user is not violated. Another design criterion for RFIDs is the low computational effort required for tags, in view that most tags are passive devices that derive power from an RFID reader's signals. Along this vein, a class of ultralightweight RFID authentication protocols has been designed, which uses only the most basic bitwise and arithmetic operations like exclusive-OR, OR, addition, rotation, and so forth. In this paper, we analyze the security of the SASI protocol, a recently proposed ultralightweight RFID protocol with better claimed security than earlier protocols. We show that SASI does not achieve resistance to tracking, which is one of its design objectives.
Raphael C.-W. Phan
IEEE Trans. Dependable Secur. Comput.1
2008 Traceable Privacy of Recent Provably-Secure RFID Protocols
Khaled Ouafi, Raphael C.-W. Phan
ACNS2
2008 The Hash Function Family LAKE
Jean-Philippe Aumasson, Willi Meier, Raphael C.-W. Phan
FSE3
2008 Privacy of Recent RFID Authentication Protocols
Khaled Ouafi, Raphael C.-W. Phan
ISPEC2
2008 Proxy Re-signatures in the Standard Model
Sherman S. M. Chow, Raphael C.-W. Phan
ISC2
2008 Cryptanalysis of simple three-party key exchange protocol (S-3PAKE)
Raphael C.-W. Phan, Wei-Chuen Yau, Bok-Min Goi
Inf. Sci.1
2008 Tampering with a watermarking-based image authentication scheme
Raphael C.-W. Phan
Pattern Recognit.1
2007 Cryptanalysis of Two Non-anonymous Buyer-Seller Watermarking Protocols for Content Protection
Bok-Min Goi, Raphael C.-W. Phan, Hean-Teik Chuah
ICCSA (1)2
2007 (In)Security of an Efficient Fingerprinting Scheme with Symmetric and Commutative Encryption of IWDW 2005
Raphael C.-W. Phan, Bok-Min Goi
IWDW1
2007 Security of a Leakage-Resilient Protocol for Key Establishment and Mutual Authentication
Raphael C.-W. Phan, Kim-Kwang Raymond Choo, Swee-Huay Heng
ProvSec1
2007 On the Notions of PRP - RKA , KR and KR - RKA for Block Ciphers
Ermaliza Razali, Raphael C.-W. Phan, Marc Joye
ProvSec2
2007 On the Analysis and Design of a Family Tree of Smart Card Based User Authentication Schemes
Raphael C.-W. Phan, Bok-Min Goi
UIC1
2006 Cryptanalysis of the N-Party Encrypted Diffie-Hellman Key Exchange Using Different Passwords
Raphael C.-W. Phan, Bok-Min Goi
ACNS1
2006 Amplifying Side-Channel Attacks with Techniques from Block Cipher Cryptanalysis
Raphael C.-W. Phan, Sung-Ming Yen
CARDIS1
2006 Cryptanalysis of some improved password-authenticated key exchange schemes
Raphael C.-W. Phan, Bok-Min Goi, Kah-Hoong Wong
Comput. Commun.1
2006 Cryptanalysis of two password-based authentication schemes using smart cards
Raphael C.-W. Phan
Comput. Secur.1
2006 Security considerations for incremental hash functions based on pair block chaining
Raphael C.-W. Phan, David A. Wagner 0001
Comput. Secur.1
2006 A Framework for Describing Block Cipher Cryptanalysis
abstract
Block ciphers provide confidentiality by encrypting confidential messages into unintelligible form, which are irreversible without knowledge of the secret key used. During the design of a block cipher, its security against cryptanalysis must be considered. History has shown that a cipher designed without an adequate treatment of this will often lead to flaws and attacks by other researchers, sometimes devastatingly so. The problem for an aspiring cipher designer is that there are no standard texts on block cipher cryptanalysis because it is a fast changing field. The commonly available references are academic journals and conference proceedings, which may not be easy to grasp for researchers new to cryptanalysis. This paper presents the Xi framework, which is designed to compactly describe the block cipher cryptanalysis techniques regardless of their individual differences. This provides the cryptanalyst with a general framework to describe attacks on block ciphers, with the additional capabilities of allowing specification of the technical details of each different type of attack and of comparison of their respective strengths. Comparing different distinguishers in this framework also allows us to see natural generalizations and trigger nice open problems. We then show how to apply this Xi framework to the description of various attacks on popular and recent block ciphers.
Raphael C.-W. Phan, Mohammad Umar Siddiqi
IEEE Trans. Computers1
2005 Cryptanalysis of an Improved Client-to-Client Password-Authenticated Key Exchange (C2C-PAKE) Scheme
Raphael C.-W. Phan, Bok-Min Goi
ACNS1
2005 On the Rila-Mitchell Security Protocols for Biometrics-Based Cardholder Authentication in Smartcards
Raphael C.-W. Phan, Bok-Min Goi
ICCSA (1)1
2005 On the Rila-Mitchell Security Protocols for Biometrics-Based Cardholder Authentication in Smartcards
Raphael C.-W. Phan, Bok-Min Goi
ICCSA (4)1
2005 Related-Mode Attacks on Block Cipher Modes of Operation
Raphael C.-W. Phan, Mohammad Umar Siddiqi
ICCSA (3)1
2005 On the Security Bounds of CMC, EME, EME+ and EME* Modes of Operation
Raphael C.-W. Phan, Bok-Min Goi
ICICS1
2005 On the Security of the WinRAR Encryption Method
Gary S.-W. Yeo, Raphael C.-W. Phan
ISC2
2004 Cryptanalysis of Two Anonymous Buyer-Seller Watermarking Protocols and an Improvement for True Anonymity
Bok-Min Goi, Raphael C.-W. Phan, Yanjiang Yang, Feng Bao 0001, Robert H. Deng, Mohammad Umar Siddiqi
ACNS2
2004 Related-Key Attacks on Triple-DES and DESX Variants
Raphael C.-W. Phan
CT-RSA1
2004 On Related-Key and Collision Attacks: The Case for the IBM 4758 Cryptoprocessor
Raphael C.-W. Phan, Helena Handschuh
ISC1
2004 Flaws in Generic Watermarking Protocols Based on Zero-Knowledge Proofs
Raphael C.-W. Phan, Huo-Chong Ling
IWDW1
2004 Impossible differential cryptanalysis of 7-round Advanced Encryption Standard (AES)
Raphael C.-W. Phan
Inf. Process. Lett.1