VLDB 2026 Research / reviewers in the wild / expert
Raphael C.-W. Phan
dblp:p/RCWPhan · also Raphael Chung-Wei Phan, Raphaël C.-W. Phan
· DBLP profile ↗
137ranked-venue papers
27as first author
55since 2021 · last 2026
0000-0001-7448-4595ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 47 · 1 first-author · 30 since 2021Security and privacy · 42 · 14 first-author · 2 since 2021Artificial intelligence and machine learning · 21 · 1 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 3 first-author · 14 since 2021Computer networks · 8 · 4 first-author · 2 since 2021Systems, architecture and hardware · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 1 since 2021Theory of computation · 5 · 1 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FLAG-4D: Flow-Guided Local-Global Dual-Deformation Model for 4D ReconstructionabstractWe introduce FLAG-4D, a novel framework for generating novel views of dynamic scenes by reconstructing how 3D Gaussian primitives evolve through space and time. Existing methods typically rely on a single Multilayer Perceptron(MLP) to model temporal deformations, and they often struggle to capture complex point motions and fine-grained dynamic details consistently over time, especially from sparse input views. Our approach, FLAG-4D overcomes this by employing a dual-deformation network that dynamically warps a canonical set of 3D Gaussians over time into new positions and anisotropic shapes. This dual-deformation network consists of an Instantaneous Deformation Network (IDN) for modeling fine-grained, local deformations, and Global Motion Network (GMN) for capturing long-range dynamics, refined via mutual learning. To ensure these deformations are both accurate and temporally smooth, FLAG-4D incorporates dense motion features from a pretrained optical flow backbone. We fuse these motion cues from adjacent timeframes and use a deformation-guided attention mechanism to align this flow information with the current state of each evolving 3D Gaussian. Extensive experiments demonstrate that FLAG-4D achieves higher-fidelity and more temporally coherent reconstructions with finer detail preservation than state-of-the-art methods. Guan Yuan Tan, Ngoc Tuan Vu, Arghya Pal, Sailaja Rajanala, Raphael C.-W. Phan, Mettu Srinivas, Chee-Ming Ting |
AAAI | 5 |
| 2026 | DepthPolyp: Pseudo-depth Guided Lightweight Segmentation for Real-Time Colonoscopy
Zhuoyu Wu, Wenhui Ou, Lexi Zhang, Pei-Sze Tan, Dongjun Wu, Junhe Zhao, Wenqi Fang, Raphael C.-W. Phan |
ICPR (7) | 8 |
| 2026 | A new time-decay radiomics integrated network (TRINet) for breast cancer risk predictionabstractTo facilitate early detection of breast cancer, there is a need to develop risk prediction schemes that can prescribe personalized screening mammography regimens for women. In this study, we propose a new deep learning architecture called TRINet that implements time-decay attention to focus on recent mammographic screenings, as current models do not account for the relevance of newer images. We integrate radiomic features with an Attention-based Multiple Instance Learning (AMIL) framework to weigh and combine multiple views for better risk estimation. In addition, we introduce a continual learning approach with a new label assignment strategy based on bilateral asymmetry to make the model more adaptable to asymmetrical cancer indicators. Finally, we add a time-embedded additive hazard layer to perform dynamic, multi-year risk forecasting based on individualized screening intervals. We used two public datasets, namely 8528 patients from the American EMBED dataset and 8723 patients from the Swedish CSAW dataset in our experiments. Evaluation results on the EMBED test set show that our approach performs comparably with state-of-the-art models, achieving AUC scores of 0.851, 0.811, 0.796, 0.793, and 0.789 across 1-, 2-, to 5-year intervals, respectively. Our results underscore the importance of integrating temporal attention, radiomic features, time embeddings, bilateral asymmetry, and continual learning strategies, providing a more adaptive and precise tool for breast cancer risk prediction. Hong Hui Yeoh, Fredrik Strand, Raphael C.-W. Phan, Kartini Rahmat, Maxine Tan |
Medical Image Anal. | 3 |
| 2026 | Causal-Ex: Causal graph-based micro and macro expression spottingabstractDetecting concealed emotions within apparently normal expressions is crucial for identifying potential mental health issues and facilitating timely support and intervention. The task of spotting macro- and micro-expressions involves predicting the emotional timeline within a video by identifying the onset (i.e., the beginning), apex (the peak of emotion), and offset (the end of emotion) frames of the displayed emotions. More particularly, closely monitoring the key emotion-conveying regions of the face; namely, the foundational muscle-movement cues known as facial action units (AUs)–greatly aids in the clear identification of micro-expressions. One major roadblock is the inadvertent introduction of biases into the training process, which degrades performance regardless of feature quality. Biases are spurious factors that falsely inflate or deflate performance metrics. For instance, the neural networks tend to falsely attribute certain AUs in specific facial regions to particular emotion classes, a phenomenon also termed as Inductive biases. To remove these false attributions, we must identify and mitigate biases that arise from mere correlation between some features and the output class labels. We hence introduce action-unit causal graphs. Unlike the traditional action-unit graph, which connects AUs based solely on spatial adjacency, the causal AU graph is derived from statistical tests and retains edges between AUs only when there is significant evidence that one AU causally influences another. Our model, named Causal-Ex ( Causal -based Ex pression spotting), employs a fast causal inference algorithm to construct a causal graph of facial region of interests (ROIs). This enables us to select causally relevant facial action units in the ROIs. Our work demonstrates improvement in overall F1-scores compared to state-of-the-art approaches with 0.388 on CAS(ME) 2 and 0.3701 on SAMM-Long Video datasets. Our code can be found at: https://github.com/noobasuna/causal_ex.git . Pei-Sze Tan, Sailaja Rajanala, Arghya Pal, Raphael C.-W. Phan, Huey Fang Ong |
Pattern Recognit. Lett. | 4 |
| 2026 | A Unified Framework for Sparse Reconstruction via Preconditioning and Nonconvex RegularizationabstractCompressed Sensing (CS) is an effective technique to recover sparse signals with fewer samples than what is required by the classical Shannon Nyquist sampling theorem. The sensing matrix, sparsifying transform, and sparse recovery algorithm are three key factors for accurate reconstruction in CS. Traditional CS uses a convex $l_{1}$-norm sparse regularizer which may lead to biased estimates and is suboptimal in promoting sparsity. Another challenge is the design of incoherent sensing matrices which is crucial for accurate sparse recovery. In this paper, we propose a novel CS framework combining a preconditioned sensing matrix and nonconvex regularization for improved sparse signal recovery. First, we formulate an optimization problem to find an incoherent sensing matrix via a preconditioner. It allows for a direct computation of the optimal preconditioner and preconditioned sensing matrix, simultaneously. Secondly, we consider a generalized CS model for signal recovery based on the incoherent sensing matrix and a nonconvex $\ell _{1/2}$-norm regularizer. We then derive an Alternating Direction Method of Multipliers (ADMM) algorithm to solve this nonconvex optimization problem. The proposed model is applied to sparse-view Computed Tomography (CT) reconstruction with highly-undersampled and noisy data. Qualitative and quantitative results show significantly better image reconstruction using the preconditioned sensing matrix and $\ell _{1/2}$ regularizer, compared to methods without preconditioning and using the $\ell _{1}$ regularizer. Prasad Theeda, Fuad Noman, Arghya Pal, Raphael C.-W. Phan, Hernando C. Ombao, Chee-Ming Ting |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | GENIE: Socially Unbiased Generative Text-to-Image EditingabstractGenerative diffusion models often exhibit societal biases in sensitive personal attributes such as age, gender, and race. In this work, we describe GENIE – a method to reduce such biases in a variety of classifier-free diffusion models used for image editing. Our method implicitly incorporates debiasing terms together with the user’s explicit edit instruction to reduce bias. This automatic method relieves the user from needing to modify edit instructions in order to avoid bias. Further, no additional training is needed. Experimental results are provided based on modifications to four diffusion models, namely InstructPix2Pix, Stable Diffusion 1.5, Stable Diffusion 2.1, and Stable Diffusion XL. We show that, on average, bias is reduced by 31% in gender, 15% in age, 39% in race. Julia Kaiwen Lau, Raphael C.-W. Phan, Sailaja Rajanala, Ingemar J. Cox, Arghya Pal |
ICASSP | 2 |
| 2025 | Post-Hoc Adversarial Stickers Against Micro-Expression LeakageabstractSecuring micro-expressions against leakage is crucial for privacy, as these subtle facial movements convey genuine emotions and are inherently personal. This study aims to protect micro-expression data from potential adversarial attacks, ensuring the preservation of individuals’ privacy and preventing unauthorized access or misuse of sensitive emotional information. Unlike traditional methods, which often require training and extensive access to models, this research introduces a novel post-hoc method that does not require additional training. We focus on physical adversarial attacks in micro-expression recognition, involving intentional manipulation of visual cues to deceive recognition systems and protect individual emotional privacy. Our approach leverages a causal discovery algorithm to identify causal relationships between facial parts, enabling rapid identification of the optimal locations for adversarial patches in frames with triggered micro-expressions. This method exhibits a more consistent attack success rate than randomly placed adversarial stickers, demonstrating effective generalization across different emotions, stickers, and models. Particularly relevant in scenarios with restricted access to the model, our technique requires only a single interaction during the attack process, highlighting its efficiency and minimal need for querying the target model. The proposed method effectively balances privacy protection with high generalization capability, setting a new standard for defending against adversarial threats in micro-expression recognition. The code is available at https://github.com/noobasuna/au-sticker. Pei-Sze Tan, Sailaja Rajanala, Yee-Fan Tan, Arghya Pal, Chun-Ling Tan, Raphael C.-W. Phan, Huey Fang Ong |
ICASSP | 6 |
| 2025 | Guided Diffusion For Class-Conditioned Synthesis & Classification Of Microscopic Blood Cell ImagesabstractMicroscopic visualization of diseased cells plays a vital role in the diagnosis and understanding of various medical conditions. Recent advances in deep learning generative models have shown remarkable potential as formidable tools for generating high-quality medical images. However, training these models generally requires large, annotated datasets, which are often costly and time-consuming to obtain. To overcome this challenge, we propose a fast-sampling guided score-based diffusion model with a classifier-free guidance strategy for class-conditioned generation of microscopic peripheral blood cell images. Our model achieves a Fréchet Inception Distance (FID) score of 10.24, demonstrating its ability to generate realistic synthetic blood cell images. Furthermore, our experimental results show that augmenting training data with these synthetic images significantly improves classification accuracy compared to relying solely on real data, highlighting the potential of synthetic data augmentation in hematology. Kar-Ee Hoh, Junn Yong Loo, Yee-Fan Tan, Raphael C.-W. Phan, Chee-Ming Ting |
ICIP | 4 |
| 2025 | Polyfit generative model: can a group of lower-order polynomials generate high resolution diverse images?abstractImplicit neural representations (INRs) have recently gained popularity as a means to model images as continuous functions of spatial coordinates, synthesizing each pixel independently and yielding impressive results in tasks such as scene reconstruction and image generation. A notable advancement, PolyINR, utilizes element-wise multiplications between features and affine-transformed coordinates to achieve higher-order polynomial functions, eliminating the need for positional encodings. However, the finite encoding capacity of INRs, coupled with PolyINR’s recursive polynomial estimation, necessitates substantial training parameters, resulting in high computational costs and limiting applicability across diverse computer vision domains. In this work, we address these challenges by representing images as grids of smaller patches, within which we fit low-degree polynomials to capture local intensity variations. Our approach substantially reduces parameter requirements and computational demands. We evaluate our model qualitatively and quantitatively on large-scale datasets, ImageNet, CelebA, LSUN Bedroom, and Flower102; thus demonstrating competitive performance with state-of-the-art generative models, despite the absence of convolutional, normalization, or self-attention layers. Arghya Pal, Ai-Fang Chai, Sailaja Rajanala, Raphael C.-W. Phan, Koksheik Wong, Chee-Ming Ting |
IJCNN | 4 |
| 2025 | Adaptive Graph Learning with Multi-graph Convolutions for Brain Disorder Classification
Fuad Noman, Raphael C.-W. Phan, Hernando C. Ombao, Chee-Ming Ting |
MICCAI (12) | 2 |
| 2025 | T2I-Diff: fMRI Signal Generation via Time-Frequency Image Transform and Classifier-Free Denoising Diffusion Models
Hwa Hui Tew, Junn Yong Loo, Yee-Fan Tan, Hernando C. Ombao, Fuad Noman, Raphael C.-W. Phan, Chee-Ming Ting |
MICCAI (3) | 7 |
| 2025 | Securing Face ID: Privacy Preservation for Non-Retentive Face Recognition SystemabstractFacial recognition technology is increasingly integrated into various applications. While face recognition systems have streamlined the authentication process and reduced the need for manual verification, the advancement in artificial intelligence (AI) generating realistic-looking media raises significant privacy concerns due to the potential misuse of biometric data. While current security protocols ensure data protection, storing biometric data in the system is a latent risk. In cases where the stored data is compromised, the users are susceptible to attacks such as face swapping, deepfake and identity theft. To address this, this paper presents a privacypreserving algorithm that omits the need to store human raw biometric data in the system. This is achieved by utilizing a new combination of Locality-Sensitive Hashing (LSH), salting, and RSA encryption for face recognition. The proposed method ensures data security by securely hashing and encrypting facial features while maintaining high recognition accuracy. The proposed framework is evaluated on the Labeled Faces in the Wild (LFW) and achieves a comparable performance with the state-of-the-art techniques. Megan Chua, Chanelle Yue-Ting Yeow, Cassandra Xin-Yee Chwee, Shu-Min Leong, Raphael C.-W. Phan |
TENCON | 5 |
| 2025 | BitRelation: Exploring Bit-Level Dependencies in Neural CryptanalysisabstractThis paper applies Explainable Artificial Intelligence (XAI) to improve the interpretability of neural differential cryptanalysis on the SPECK cipher. We use Local Interpretable Model-agnostic Explanations (LIME) to analyse and visualise feature importance in neural distinguishers, giving signed contributions and absolute rankings. Signed contributions show whether, and how strongly, specific bit positions influence the model's decision, while absolute rankings reflect their importance regardless of sign. To study interactions beyond single bits, we introduce a Systematic Masking Approach to reveal relations among bits by testing if chosen combinations of masked bits alter classification accuracy. On Gohr's 8-round SPECK32/64 distinguisher, masking up to four-bit combinations shows that decisions involve multi-bit interactions rather than isolated single-bit effects. Although LIME highlights strong single-bit signals, masking reveals interaction patterns consistent with differential cryptanalysis. These findings clarify model behaviour in neural cryptanalysis and show XAI's value for exposing and visualising interaction structure in ciphertext features and decisions. Yue-Tian Goi, Shu-Min Leong, Raphael C.-W. Phan, Ana Salagean, Shangqi Lai, Wei-Chuen Yau |
TENCON | 3 |
| 2025 | Res-SH: Unbiased Residual Learning for Self-Healing Interface Toughness Prediction with Limited DataabstractThe development of self-healing materials is often hindered by the high costs and material waste associated with traditional characterization methods. Current approaches to toughness prediction, primarily based on convolutional neural networks (CNNs), are limited by their tendency to capture only surface-level features, which can lead to biased predictions. Moreover, working with small datasets, which is common in materials science, further increases the risk of biased training due to overfitting, posing a critical challenge to the reliability and generalizability of predictive models. This study introduces an unbiased residual learning framework designed explicitly for predicting self-healing interface toughness under limiteddata conditions. Our approach, ResNet-inspired approach for predicting self-healing material toughness, named Res-SH, used the power of residual networks to capture deeper, more complex patterns in the data, thereby addressing critical challenges in materials research. Res-SH minimises resource consumption and experimental overhead by focusing on unbiased learning, achieving accurate predictions with fewer training epochs and lower R2score and root mean square prediction errors compared to conventional CNN and lightweight model MobileNetv2. This novel framework provides a cost-effective and resource-efficient alternative to traditional material characterization methods, reducing material waste and accelerating the discovery and optimization of self-healing material systems. Pei-Sze Tan, Karen Jia-Jun Koh, Sailaja Rajanala, Arghya Pal, Raphael C.-W. Phan, Nan Ze, Fuad Noman, Chee-Ming Ting, Norfadilah Dolmat, Nik Nur Wahidah Nik Hashim, Afidalina Tumian |
TENCON | 5 |
| 2025 | Disentangle Class Imbalance in Micro-Expression Recognition with Causal Structure LearningabstractMicro-expression recognition is a critical task in affective computing, yet it is often hindered by biases inherent in datasets, leading to skewed and unreliable outcomes. This work introduces a novel approach to disentangling bias in microexpression recognition using causal structure learning. By modeling causal relationships within the data, we identify and mitigate sources of bias that traditional machine learning models often overlook. Our framework integrates causal discovery techniques to uncover biased patterns in widely used micro-expression datasets. We employ debiasing strategies to enhance the fairness and accuracy of recognition models and generate counterfactual examples to address sensitive attributes such as gender and age, allowing us to observe the effects of imbalanced classes on classification results. Experiments conducted on baseline microexpression recognition models demonstrate comparable results after undersampling to create emotion class balance, revealing label bias in current training datasets including CASME2, SAMM, and SMIC. Further evaluation on balanced gender and age classes using generated counterfactual data as additional training instances showed performance improvements for the 4DME dataset. Pei-Sze Tan, Sailaja Rajanala, Raphael C.-W. Phan |
TENCON | 3 |
| 2025 | Attack-SH: Adversarial Attacks on Self-Healing Material Properties Prediction ModelabstractAI is now increasingly applied in diverse domains. The recent Nobel prizes for Physics and Chemistry awarded to computational scientists shows the significant impact that AI has on real-world scientific applications. Adversarial attacks pose a significant threat to the reliability of AI systems, particularly in high-stakes applications such as those in the materials sciences domain, which affect interactions with materials that exist in the real world. This paper examines the impact of two widely used adversarial attack approaches, notably the Fast Gradient Sign Method (FGSM) and Projected Gradient Descent (PGD), on a target AI model's performance for realworld materials. Experimental results demonstrate that both methods effectively degrade predictive accuracy, with PGD showing a more severe effect. Notably, high Structural Similarity Index (SSIM) scores across all perturbed samples suggest that the attacks introduce imperceptible changes, increasing their potential risk as such attacks are then undetectable. Further analysis using metrics such as Mean Squared Error (MSE), Adversarial MSE (A-MSE), and Relative Error Increase (REI) confirms substantial shifts in model output and reconstruction quality. These findings highlight critical vulnerabilities in current model architectures and emphasize the urgent need for more resilient defense strategies, such as adversarial training and input pre-processing. Min-Xuan Tan, Pei-Sze Tan, Raphael C.-W. Phan, Shu-Min Leong |
TENCON | 3 |
| 2025 | PK-YOLO: Pretrained Knowledge Guided YOLO for Brain Tumor Detection in Multiplanar MRI SlicesabstractBrain tumor detection in multiplane Magnetic Resonance Imaging (MRI) slices is a challenging task due to the various appearances and relationships in the structure of the multiplane images. In this paper, we propose a new You Only Look Once (YOLO)-based detection model that incorporates Pretrained Knowledge (PK), called PK-YOLO, to improve the performance for brain tumor detection in multiplane MRI slices. To our best knowledge, PK-YOLO is the first pretrained knowledge guided YOLO-based object detector. The main components of the new method are a pretrained pure lightweight convolutional neural network-based backbone via sparse masked modeling, a YOLO architecture with the pretrained backbone, and a regression loss function for improving small object detection. The pre-trained backbone allows for feature transferability of object queries on individual plane MRI slices into the model encoders, and the learned domain knowledge base can improve in-domain detection. The improved loss function can further boost detection performance on small-size brain tumors in multiplanar two-dimensional MRI slices. Experimental results show that the proposed PK-YOLO achieves competitive performance on the multiplanar MRI brain tumor detection datasets compared to state-of-the-art YOLO-like and DETR-like object detectors. The code is available at https://github.com/mkang315/PK-YOLO. Ming Kang 0002, Fung Fung Ting, Raphael C.-W. Phan, Chee-Ming Ting |
WACV | 3 |
| 2025 | @LM DeceptionNet: A multimodal approach for efficient transfer learning-based deception detectionabstractIn terms of deception detection, traditional contact-based techniques often require collecting physiological signals, which can negatively impact device accuracy and participant comfort. While multimodal features extracted from audio and video modalities have been shown to outperform human observers on public datasets, the generalizability of existing audio and visual-based deception detection methods in different scenarios remains insufficiently explored. To narrow this gap, this work proposes a novel domain knowledge transfer learning method for deception detection in cross-scenario applications, which enhances its generalization and adaptability. Additionally, we designed a multimodal framework that filters out irrelevant information from other modalities when a particular modality yields reliable results, further improving overall system accuracy and robustness. We evaluate the proposed method on different public datasets, achieving promising generalizability results with consistent enhancements using four variations and networks. Apart from this, the proposed @LM DeceptionNet demonstrates better generalization capacity in computational efficiency, feature extraction, and adaptability compared to a larger model when employing fewer parameters. Yuanya Zhuo, Vishnu Monn Baskaran, Lillian Yee Kiaw Wang, Raphael C.-W. Phan |
Knowl. Based Syst. | 4 |
| 2024 | BrainFC-CGAN: A Conditional Generative Adversarial Network for Brain Functional Connectivity Augmentation and Aging SynthesisabstractBrain functional connectivity (FC) changes are associated with neuropsychiatric disorders and other underlying factors, such as age and gender. Due to small training sample, data augmentation has been increasingly used for deep learning-based classification of brain FC. Although deep generative models could generate brain FCs to enhance downstream classification, most existing methods neglect the underlying factors involved in the generation process and fail to preserve the subject identity. We propose a novel brain FC conditional Generative Adversarial Network (GAN) called BrainFC-CGAN with specialized layers and filters to preserve the symmetry property and topological structure of brain FCs. We design a FC generator that captures the complex variations between brain FCs, ages, and health statuses to generate synthetic FCs that preserve the subject identity. We categorized true brain FCs into different age groups; an augmented age-specific dataset generated from BrainFC-CGAN is combined with the training set for classification. Experimental results on major depressive disorder (MDD) resting-state functional magnetic resonance imaging data show that the proposed method synthesizes realistic brain FCs of different target age groups, significantly improving downstream classification performance over baseline without augmentation, and also outperforming several state-of-the-art GANs. Yee-Fan Tan, Junn Yong Loo, Chee-Ming Ting, Fuad Noman, Raphael C.-W. Phan, Hernando C. Ombao |
ICASSP | 5 |
| 2024 | Causally Uncovering Bias in Video Micro-Expression RecognitionabstractDetecting microexpressions presents formidable challenges, primarily due to their fleeting nature and the limited diversity in existing datasets. Our studies find that these datasets exhibit a pronounced bias towards specific ethnicities and suffer from significant imbalances in terms of both class and gender representation among the samples. These disparities create fertile ground for various biases to permeate deep learning models, leading to skewed results and inadequate portrayal of specific demographic groups. Our research is driven by a compelling need to identify and rectify these biases within model architectures. To achieve this, we commence by constructing a causal graph that elucidates the intricate relationships between the model, input features, and training outcomes. This graphical representation forms the foundation for our analytical framework. Leveraging this causal framework, we conduct comprehensive case studies, employing counterfactuals as a diagnostic tool to unveil biases arising from dataset-induced class imbalances, gender inequalities, and variations in facial action units. Our final step involves a highly efficient counterfactual debiasing process, eliminating the necessity for additional data collection or model retraining. Our results showcase superior performance compared to state-of-the-art methods across the CASME II, SAMM, and SMIC datasets. Pei-Sze Tan, Sailaja Rajanala, Arghya Pal, Shu-Min Leong, Raphael C.-W. Phan, Huey Fang Ong |
ICASSP | 5 |
| 2024 | Cafct-Net: A Cnn-Transformer Hybrid Network With Contextual And Attentional Feature Fusion For Liver Tumor SegmentationabstractMedical image semantic segmentation techniques can help identify tumors automatically from computed tomography (CT) scans. In this paper, we propose a Contextual and Attentional feature Fusions enhanced Convolutional Neural Network (CNN) and Transformer hybrid network (CAFCT-Net) for liver tumor segmentation. We incorporate three novel modules in the CAFCT-Net architecture: Attentional Feature Fusion (AFF), Atrous Spatial Pyramid Pooling (ASPP) of DeepLabv3, and Attention Gates (AGs) to improve contextual information related to tumor boundaries for accurate segmentation. Experimental results show that the proposed model achieves a mean Intersection over Union (IoU) of 76.54% and Dice coefficient of 84.29%, respectively, on the Liver Tumor Segmentation Benchmark (LiTS) dataset, outperforming pure CNN or Transformer methods, e.g., Attention U-Net and PVTFormer. Ming Kang 0002, Chee-Ming Ting, Fung Fung Ting, Raphael C.-W. Phan |
ICIP | 4 |
| 2024 | CST-Yolo: A Novel Method For Blood Cell Detection Based On Improved Yolov7 And CNN-Swin TransformerabstractBlood cell detection is a typical small-scale object detection problem in computer vision. In this paper, we propose a CST-YOLO model for blood cell detection based on YOLOv7 architecture and enhance it with the CNN-Swin Transformer (CST), which is a new attempt at CNN-Transformer fusion. We also introduce three other useful modules: Weighted Efficient Layer Aggregation Networks (W-ELAN), Multiscale Channel Split (MCS), and Concatenate Convolutional Layers (CatConv) in our CST-YOLO to improve small-scale object detection precision. Experimental results show that the proposed CST-YOLO achieves 92.7%, 95.6%, and 91.1% mAP @ 0.5, respectively, on three blood cell datasets, outperforming state-of-the-art object detectors, e.g., RT-DETR, YOLOv5, and YOLOv7. Our code is available at https://github.com/mkang315/CST-YOLO. Ming Kang 0002, Chee-Ming Ting, Fung Fung Ting, Raphael C.-W. Phan |
ICIP | 4 |
| 2024 | Deep Multi-Graph Embedded Clustering for Community Detection in FMRI Functional Brain Networks Across IndividualsabstractAnalyzing the community structure of brain networks provides new insights into human brain function. Existing studies broadly use conventional network clustering approaches. While graph neural networks have recently shown promise in modeling brain functional connectivity (FC) networks, their applications to brain community detection still need improvement and further refinement. Moreover, identifying common community structure while resolving the single-subject partitions across multiple individual networks remains underexplored. We propose a Deep Multi-Graph Embedded Clustering (DMGEC) framework to identify shared community partition in brain FC networks over a cohort of individuals. By incorporating the consensus information aggregated across network structures, DMGEC leverages a graph autoencoder to produce consensus-aware latent representations of individual networks, and applies deep embedded clustering on the multi-subject network representation to produce common community assignment of brain nodes. Simulations show superior community recovery by our method compared to conventional approaches, especially for networks with large number of communities. When applied to functional magnetic resonance imaging (fMRI) data, the DMGEC achieves outstanding alikeness over individual partitions, and uncovers group-level differences in brain community motifs between major depressive disorder patients and normal controls. Kai-Jun See, Chee-Ming Ting, Fuad Noman, Junn Yong Loo, Yee-Fan Tan, Hernando C. Ombao, Raphael C.-W. Phan |
ICIP | 7 |
| 2024 | Dynamic MRI Reconstruction Using Low-Rank Plus Sparse Decomposition With Smoothness RegularizationabstractThe low-rank plus sparse (L+S) decomposition model has enabled better reconstruction of dynamic magnetic resonance imaging (dMRI) with separation into background (L) and dynamic (S) component. However, use of low-rank prior alone may not fully explain the slow variations or smoothness of the background part at the local scale. In this paper, we propose a smoothness-regularized L+S (SR-L+S) model for dMRI reconstruction from highly undersampled k-t-space data. We exploit joint low-rank and smooth priors on the background component of dMRI to better capture both its global and local temporal correlated structures. Extending the L+S formulation, the low-rank property is encoded by the nuclear norm, while the smoothness by a general $\ell_{p}$-norm penalty on the local differences of the columns of L. The additional smoothness regularizer can promote piecewise local consistency between neighboring frames. By smoothing out the noise and dynamic activities, it allows accurate recovery of the background part, and subsequently more robust dMRI reconstruction. Extensive experiments on multi-coil cardiac and synthetic data shows that the SR-L+S model outperforms several state-of-the-art methods in terms of recovery accuracy. Chee-Ming Ting, Fuad Noman, Raphael C.-W. Phan, Hernando C. Ombao |
ICIP | 3 |
| 2024 | ActNetFormer: Transformer-ResNet Hybrid Method for Semi-supervised Action Recognition in Videos
Sharana Dharshikgan Suresh Dass, Hrishav Bakul Barua, Ganesh Krishnasamy, Raveendran Paramesran, Raphael C.-W. Phan |
ICPR (15) | 5 |
| 2024 | A Deep Probabilistic Spatiotemporal Framework for Dynamic Graph Representation Learning with Application to Brain Disorder Identification
Sin-Yee Yap, Junn Yong Loo, Chee-Ming Ting, Fuad Noman, Raphael C.-W. Phan, Adeel Razi, David L. Dowe |
IJCAI | 5 |
| 2024 | Video Deception Detection through the Fusion of Multimodal Feature Extraction and Neural NetworksabstractDetecting deceptive behavior in videos is a complex task within several domains, including academic fraud assessment, commercial anti-fraud activities, judicial system evidence analysis, suspicious activity detection in security monitoring systems, and behavioral intent analysis in psychological research. In this study, we present a novel approach to video deception detection by integrating visual and audio models for deep feature fusion, primarily targeting advanced deception detection datasets. Our visual model leverages hierarchical image feature learning to enhance deceptive cue detection, complemented by an audio model that processes acoustic signals for precise speech pattern analysis. This multimodal method significantly boosts detection accuracy and lessens reliance on extensive training data. Notably, our visual model incorporates knowledge distillation technology, improving efficiency and reducing computational resource needs without compromising performance. We implement a transformer architecture using distillation tokens for effective learning and incorporate convolutional neural network insights to enrich our model’s interpretative capabilities. Experimental results demonstrate that our approach surpasses existing technologies in various standards and scenarios, offering enhanced deception recognition capabilities and addressing the challenge of limited training data. Yuanya Zhuo, Vishnu Monn Baskaran, Lillian Yee Kiaw Wang, Raphael C.-W. Phan |
IJCNN | 4 |
| 2024 | BGF-YOLO: Enhanced YOLOv8 with Multiscale Attentional Feature Fusion for Brain Tumor Detection
Ming Kang 0002, Chee-Ming Ting, Fung Fung Ting, Raphael C.-W. Phan |
MICCAI (8) | 4 |
| 2024 | Unveiling the Black Box: Neural Cryptanalysis with XAIabstractAt CRYPTO'19, Gohr[1] presented ResNet-based neural distinguishers (ND) for the round-reduced SPECK32/64 cipher. However, due to the black-box use of such deep learning models, it is hard for humans to understand why these distinguishers work, impeding advancements in cryptanalytic knowledge. In this work, we aim to effectively adapt eXplainable Artificial Intelligence (XAI) techniques, notably Local Interpretable Model-Agnostic Explanations (LIME) and Shapley Additive Explanations (SHAP), to gain a detailed understanding of the important features useful in Gohr's neural distinguishers. Yue-Tian Goi, Shu-Min Leong, Raphael C.-W. Phan, Shangqi Lai, Ana Salagean |
SMC | 3 |
| 2024 | Distingusic: Distinguishing Synthesized Music from HumanabstractIn this paper we focus on a problem that is increasingly plaguing the music industry; to a large extent due to the proliferation of generative AI models that enable the generation of new realistic and indistinguishable content for diverse modalities: text, image, audio, video. We address this problem from the perspective of audio watermarking; to our best knowledge, this is the first-known watermarking based approach to solve the problem of distinguishing realistic songs synthesized from generative AI models from real songs sung by humans. In more detail, our approach specifically utilizes the SHA-256 hash function, Singular Value Decomposition (SVD) and Discrete Wavelet Transform (DWT) for robust audio watermarking of synthesized songs. Before embedding, the audio is subjected to an attack phase to pinpoint less vulnerable regions for QR watermark placement. During the embedding process, the audio chunks first undergo a 1-level Discrete Wavelet Transform (DWT), and then the resulting approximate coefficients go through Singular Value Decompo-sition (SVD). Additionally, the watermarked array is subjected to SHA-256 hashing for collision-resistant conciseness, which is subsequently embedded into the singular values of the audio. Experimental findings demonstrate the superiority of our method over existing audio watermarking approaches under various signal attack scenarios. Zi Qian Yong, Shu-Min Leong, Sailaja Rajanala, Arghya Pal, Raphael C.-W. Phan |
SMC | 5 |
| 2024 | ASF-YOLO: A novel YOLO model with attentional scale sequence fusion for cell instance segmentationabstractWe propose a novel Attentional Scale Sequence Fusion based You Only Look Once (YOLO) framework (ASF-YOLO) which combines spatial and scale features for accurate and fast cell instance segmentation. Built on the YOLO segmentation framework, we employ the Scale Sequence Feature Fusion (SSFF) module to enhance the multiscale information extraction capability of the network, and the Triple Feature Encoder (TFE) module to fuse feature maps of different scales to increase detailed information. We further introduce a Channel and Position Attention Mechanism (CPAM) to integrate both the SSFF and TFE modules, which focus on informative channels and spatial position-related small objects for improved detection and segmentation performance. Experimental validations on two cell datasets show remarkable segmentation accuracy and speed of the proposed ASF-YOLO model. It achieves a box mAP of 0.91, mask mAP of 0.887, and an inference speed of 47.3 FPS on the 2018 Data Science Bowl dataset, outperforming the state-of-the-art methods. The source code is available at https://github.com/mkang315/ASF-YOLO. Ming Kang 0002, Chee-Ming Ting, Fung Fung Ting, Raphael C.-W. Phan |
Image Vis. Comput. | 4 |
| 2024 | Emotion-specific AUs for micro-expression recognitionabstractAbstract The Facial Action Coding System (FACS) comprehensively describes facial expressions with facial action units (AUs). It is a well-used technique by researchers in emotions research to understand human emotions better. Most micro-expression datasets provide FACS-coded AU ground truths corresponding to micro-expressions classes. It is commonly accepted in computer vision-based emotions research that certain emotions are reliably revealed when specific combinations of AUs occur. However, the reliability of the ground truth AUs in the micro-expression datasets is lower than that of normal expressions, as they have lower AU intensities. Moreover, these micro-expression datasets only report the overall reliability of all AUs. It could not be identified which AUs had been accurately coded. This work aims to revisit the ground truth AUs of popular micro-expression datasets, namely CASME II, SAMM and CAS(ME) $$^2$$ 2 , and inspect whether any AUs crucial for micro-expression recognition may need to be reconsidered. This paper also provides a detailed AU analysis which yields new AU-based RoIs for each dataset. These new RoIs improve the micro-expression recognition performances compared to the baselines considered in this work. The proposed RoIs for CASME II, SAMM and CAS(ME) $$^2$$ 2 improve the recognition rates by $$2\%$$ 2 % , $$1\%$$ 1 % and $$4\%$$ 4 % , respectively, when compared with the existing RoIs. Shu-Min Leong, Raphael C.-W. Phan, Vishnu Monn Baskaran |
Multim. Tools Appl. | 2 |
| 2024 | Graph Autoencoders for Embedding Learning in Brain Networks and Major Depressive Disorder IdentificationabstractBrain functional connectivity (FC) networks inferred from functional magnetic resonance imaging (fMRI) have shown altered or aberrant brain functional connectome in various neuropsychiatric disorders. Recent application of deep neural networks to connectome-based classification mostly relies on traditional convolutional neural networks (CNNs) using input FCs on a regular Euclidean grid to learn spatial maps of brain networks neglecting the topological information of the brain networks, leading to potentially sub-optimal performance in brain disorder identification. We propose a novel graph deep learning framework that leverages non-Euclidean information inherent in the graph structure for classifying brain networks in major depressive disorder (MDD). We introduce a novel graph autoencoder (GAE) architecture, built upon graph convolutional networks (GCNs), to embed the topological structure and node content of large fMRI networks into low-dimensional representations. For constructing the brain networks, we employ the Ledoit-Wolf (LDW) shrinkage method to efficiently estimate high-dimensional FC metrics from fMRI data. We explore both supervised and unsupervised techniques for graph embedding learning. The resulting embeddings serve as feature inputs for a deep fully-connected neural network (FCNN) to distinguish MDD from healthy controls (HCs). Evaluating our model on resting-state fMRI MDD dataset, we observe that the GAE-FCNN outperforms several state-of-the-art methods for brain connectome classification, achieving the highest accuracy when using LDW-FC edges as node features. The graph embeddings of fMRI FC networks also reveal significant group differences between MDD and HCs. Our framework demonstrates the feasibility of learning graph embeddings from brain networks, providing valuable discriminative information for diagnosing brain disorders. Fuad Noman, Chee-Ming Ting, Hakmook Kang, Raphael C.-W. Phan, Hernando C. Ombao |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | $\mathrm{C}\eta\iota \text{DAE}$: Cryptographically Distinguishing Autoencoder for Cipher CryptanalysisabstractWe propose a new autoencoder (AE) construction$\mathrm{C}\eta\iota \text{DAE}$(Cryptographically Distinguishing AE) based on a novel loss formulation to solve the cipher cryptanalysis distinguishing problem in the domain of cryptology. Vanilla AE and variational AE are unable to address this problem as they are designed to draw new samples which are either similar to the input sample or are from the same distribution. Such generated samples do not facilitate the cryptanalysis task. We show that our AE construction enables the discovery of cipher distinguishers, which are the fundamental building blocks that make or break new cipher design proposals. This also answers an open question on the applicability of autoencoders for cipher cryptanalysis; as to date, only discriminative models have been applied for cryptanalysis problems. To the best of our knowledge,$\mathrm{C}\eta\iota \text{DAE}$is the first-known generative model designed to solve crypt-analysis problems. We apply our$\mathrm{C}\eta\iota \text{DAE}$model to discover distinguishing properties for up to 10 rounds of the NSA-designed Speck32/64 cipher that allows to distinguish it from a random permutation. This contrasts with the best-known machine learning-discovered neural distinguisher in the literature that covers up to 8 rounds of Speck32/64. Unlike these recent related work which leverage on white box analysis and human-guided differential or linear analysis in order for machine learning models to be applicable, our$\mathrm{C}\eta\iota \text{DAE}$distinguisher does not require prior human cryptanalytic knowledge. This motivates the new direction of human-unsupervised machine learning-based cryptanalysis techniques. Raphael C.-W. Phan, Arghya Pal, Koksheik Wong, Sailaja Rajanala |
GLOBECOM | 1 |
| 2023 | Self Supervised Bert for Legal Text ClassificationabstractCritical BERT-based text classification tasks, such as legal text classification, require huge amounts of accurately labeled data. Legal text classification faces two trivial problems: labeling legal data is a sensitive process and can only be carried out by skilled professionals, and legal text is prone to privacy issues hence not all the data can be made available in the public domain. This means that we have limited diversity in the textual data, and to account for this data paucity, we propose a self-supervision approach to train Legal-BERT classifiers. We use the BERT text classifier’s knowledge of the class boundaries and perform gradient ascent w.r.t. class logits. Synthetic latent texts are generated through activation maximization. The main advantages over existing SOTAs are that our model: is easy to train, does not require much data but instead uses the synthesized data as fake samples; has less variance that helps to generate texts with good sample quality and diversity. We show the efficacy of the proposed method on the ECHR Violation (Multi-Label) Dataset and the Over-ruling Task Dataset. Arghya Pal, Sailaja Rajanala, Raphael C.-W. Phan, Koksheik Wong |
ICASSP | 3 |
| 2023 | USURP: Universal Single-Source Adversarial Perturbations on Multimodal Emotion RecognitionabstractThe field of affective computing has progressed from traditional unimodal analysis to more complex multimodal analysis due to the proliferation of videos posted online. Multimodal learning has shown remarkable performance in emotion recognition tasks, but its robustness in an adversarial setting remains unknown. This paper investigates the robustness of multimodal emotion recognition models against worst-case adversarial perturbations on a single modality. We found that standard multimodal models are susceptible to single-source adversaries and can be easily fooled by perturbations on any single modality. We draw some key observations that serve as guidelines for designing universal adversarial attacks on multimodal emotion recognition models. Motivated by these findings, we propose a novel universal single-source adversarial perturbations framework on multimodal emotion recognition models: USURP. Through our analysis of adversarial robustness, we demonstrate the necessity of studying adversarial attacks on multimodal models. Our experimental results show that the proposed USURP method achieves high attack success rates and significantly improves adversarial transferability in multimodal settings. The observations and novel attack methods presented in this paper provide a new understanding of the adversarial robustness of multimodal models, contributing to their safe and reliable deployment in more real- world scenarios. Yin Yin Low, Raphael C.-W. Phan, Arghya Pal, Xiaojun Chang |
ICIP | 2 |
| 2023 | A Unified Framework for Static and Dynamic Functional Connectivity Augmentation for Multi-Domain Brain Disorder ClassificationabstractDeep learning (DL) methods recently show promise on accurate brain disorder classification using functional connectivity (FC) estimated from functional magnetic resonance imaging (fMRI). However, DL model building can be hindered by small sample-size settings of fMRI. Moreover, most studies utilize either static (sFC) or dynamic FC (dFC) for classification. We propose a unified framework for data augmentation of both sFC and dFC for multi-domain joint classification of brain disorders. We exploit generative adversarial networks (GAN) to synthesize realistic FCs for data augmentation. Notably, we adopted the TimeGAN for dFC generation that can capture temporal dependencies in real dFC, and the GR-SPD-GAN for sFC generation that preserves the spatial connectivity structure. We further develop BrainFusionNet - a specialized DL model for multi-domain FC that simultaneously learns embedded features from both sFC and dFC to provide complementary spatio-temporal information for downstream classification. The synthetic FC data are augmented in training data to improve the BrainFusionNet performance and generalizability. Experimental results on major depressive disorder (MDD) identification using resting-state fMRI show substantial improvement in classification accuracy by our framework, outperforming competing models without FC augmentation and using sFC or dFC features alone. Yee-Fan Tan, Chee-Ming Ting, Fuad Noman, Raphael C.-W. Phan, Hernando C. Ombao |
ICIP | 4 |
| 2023 | Scheduling Dependent Batching TasksabstractWe formulate and analyze the following problem of scheduling dependent batching tasks on a common resource. Each task is associated with an execution time window. Some can be executed simultaneously in batch, while others need exclusively use of the resource. A task may depend on a set of tasks such that it becomes executable only if all its ancestor tasks are completed. We look for a task scheduling policy maximizing the system reward. We investigate both the offline and online settings by focusing on a typical scenario where the task dependency forms a tree or a forest. In both settings, we formally establish the hardness of the scheduling problem by showing that the offline scheduling is NP-hard and the online counterpart admits no scheduling policy with finite competitive ratio. We then develop approximation scheduling algorithms for both cases with deterministic worst-case performance guarantee in terms of system utility. We further conduct numeric experiments to evaluate our algorithms under a variety of parameter settings to demonstrate the effectiveness of our scheduling algorithms. Hehuan Shi, Lin Chen 0002, Raphael C.-W. Phan |
ICPP | 4 |
| 2023 | Cross-domain Transfer Learning and State Inference for Soft Robots via a Semi-supervised Sequential Variational Bayes FrameworkabstractRecently, data-driven models such as deep neural networks have shown to be promising tools for modelling and state inference in soft robots. However, voluminous amounts of data are necessary for deep models to perform effectively, which requires exhaustive and quality data collection, particularly of state labels. Consequently, obtaining labelled state data for soft robotic systems is challenged for various reasons, including difficulty in the sensorization of soft robots and the inconvenience of collecting data in unstructured environments. To address this challenge, in this paper, we propose a semi-supervised sequential variational Bayes (DSVB) framework for transfer learning and state inference in soft robots with missing state labels on certain robot configurations. Considering that soft robots may exhibit distinct dynamics under different robot configurations, a feature space transfer strategy is also incorporated to promote the adaptation of latent features across multiple configurations. Unlike existing transfer learning approaches, our proposed DSVB employs a recurrent neural network to model the nonlinear dynamics and temporal coherence in soft robot data. The proposed framework is validated on multiple setup configurations of a pneumatic-based soft robot finger. Experimental results on four transfer scenarios demonstrate that DSVB performs effective transfer learning and accurate state inference amidst missing state labels. Shageenderan Sapai, Junn Yong Loo, Ze Yang Ding, Chee Pin Tan, Raphael C.-W. Phan, Vishnu Monn Baskaran, Surya Girinatha Nurzaman |
ICRA | 5 |
| 2023 | Schatten p-norm based Image-to-Video Adaptation for Video Action RecognitionabstractHuman action recognition has been receiving extensive interest among researchers from the computer vision community. Numerous successful action recognition techniques have demonstrated the effectiveness of learning action knowledge from still images or motion videos. The relevant action information learned for the same action via various media types, such as images or videos, may be correlated. Nonetheless, less attention has been paid to adapting the action knowledge from images to videos to enhance action recognition performance in videos. Furthermore, most existing video action recognition methods suffer from insufficient labeled training videos. Overfitting could be an issue in these circumstances; hence, action recognition performance can be inhibited. This paper proposes an adaptation framework to transfer knowledge from images to videos for action recognition. A multi-task learning framework is designed to optimize the image and video domain classifiers jointly. The general Schatten p-norm is applied to the classifiers to mine the shared knowledge between these two domains. In this way, our framework can learn the correlated action semantics by leveraging the shared components of labeled images and videos. Our proposed approach can fully use the action knowledge from images and performs better in the case of poor and limited video data compared with the existing state-of-the-art action recognition techniques. Sharana Dharshikgan Suresh Dass, Ganesh Krishnasamy, Raveendran Paramesran, Raphael C.-W. Phan |
IJCNN | 4 |
| 2023 | RCS-YOLO: A Fast and High-Accuracy Object Detector for Brain Tumor Detection
Ming Kang 0002, Chee-Ming Ting, Fung Fung Ting, Raphael C.-W. Phan |
MICCAI (4) | 4 |
| 2023 | Attacking Mouse Dynamics Authentication Using Novel Wasserstein Conditional DCGANabstractBehavioral biometrics is an emerging trend due to their cost-effectiveness and non-intrusive implementations that support remote access for user identification. This is the case especially in recent times of social distancing and working from home arrangements, where online attendance is the preferred option in contrast to physical presence. In this work, we explore the limitations of mouse dynamics authentication by impersonating legitimate user mouse action sequences. Specifically, towards that aim, we develop a novel generative WC-DCGAN model to generate highly accurate fake user action sequences. We apply our WC-DCGAN to this problem and show that it causes the target classifier can be tricked into identifying a fraudster as a legitimate user. WC-DCGAN has several benefits, including: achieving dominated convergence, hence implying the existence of solutions and optimal discriminator regardless of data and generator distributions; and acting as an unsupervised model for a fixed class label and generator. Experiments are conducted to verify these points. Subsequently, we analyzed the cause of misclassifications, and propose a novel mouse dynamics strategy that offers much tighter authentication with significant reductions in misclassification events. Arunava Roy, Koksheik Wong, Raphael C.-W. Phan |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | Post-Quantum Verifiable Random Function from Symmetric Primitives in PoS Blockchain
Maxime Buser, Rafael Dowsley, Muhammed F. Esgin, Shabnam Kasra Kermanshahi, Veronika Kuchta, Joseph K. Liu, Raphael C.-W. Phan, Zhenfei Zhang |
ESORICS (1) | 7 |
| 2022 | AdverFacial: Privacy-Preserving Universal Adversarial Perturbation Against Facial Micro-Expression LeakagesabstractPrivacy safeguards are crucial, notably now with increased virtual conferencing usage during the Covid pandemic. In contrast to conventional facial expressions that are visually obvious to humans, micro-expressions are involuntary and transient facial expressions, commonly manifested involuntarily when we aim to withhold our emotions. Advanced micro-expression recognition techniques exist that can reveal the genuine emotions that people attempt to conceal, thus threatening individual emotional privacy, as fundamental human rights would dictate that one should have a choice of what emotion is being shown or not shown. We propose the novel universal adversarial perturbation-based approach - AdverFacial - for privacy concealment against automated micro-expression analysis via deep learning techniques. We derive the optimal strategy to achieve micro-expression misclassification with a high success rate, low perceptibility and cross neural network transferability. We perform experiments on two popular datasets with state-of-the-art microexpression spotting and recognition models and demonstrate our approach’s effectiveness in emotional concealment. Yin Yin Low, Angeline Tanvy, Raphael C.-W. Phan, Xiaojun Chang |
ICASSP | 3 |
| 2022 | GraphEx: Facial Action Unit Graph for Micro-Expression ClassificationabstractFacial micro-expressions are crucial cues for expressing human emotions. Existing works have shown substantial progress in detecting micro-expressions for various applications in the computer vision field. However, it is still onerous for existing methods to handle and interpret micro-expressions efficiently. This paper proposes a deep learning-based approach leveraging spatio-temporal and graph representation learning for micro-expression classification. We design a novel Spatial-Temporal Info Extraction Network (STIENet) for learning facial appearance and muscle motion from high dimensional video clip frames and summarizes them into more meaningful feature maps. We construct an action unit (AU) relation graph to further represent the AU co-occurrence in the same micro-expression video clip. A graph neural network (GNN) is used to learn AU-related graph embedding for the downstream classification task. Performance evaluation on two mainstream micro-expression datasets, i.e., CASME II and SAMM, show that the proposed framework outperforms other state-of-the-art methods for micro-expression classification. Shu-Min Leong, Fuad Noman, Raphael C.-W. Phan, Vishnu Monn Baskaran, Chee-Ming Ting |
ICIP | 3 |
| 2022 | Graph Autoencoder-Based Embedded Learning in Dynamic Brain Networks for Autism Spectrum Disorder IdentificationabstractRecent applications of pattern recognition techniques to brain connectome-based classification focus on static functional connectivity (FC) neglecting the dynamics of FC over time, and use input connectivity matrices on a regular Euclidean grid. We exploit the graph convolutional networks (GCNs) to learn irregular structural patterns in brain FC networks and propose extensions to capture dynamic changes in network topology. We develop a dynamic graph autoencoder (DyGAE)-based framework to leverage the time-varying topological structures of dynamic brain networks for identification of autism spectrum disorder (ASD). The framework combines a GCN-based DyGAE to encode individual-level dynamic networks into time-varying low-dimensional network embeddings, and classifiers based on weighted fully-connected neural network (FCNN) and long short-term memory (LSTM) to facilitate dynamic graph classification via the learned spatial-temporal information. Evaluation on a large ABIDE resting-state functional magnetic resonance imaging (rs-fMRI) dataset shows that our method outperformed state-of-the-art methods in detecting altered FC in ASD. Dynamic FC analyses with DyGAE learned embeddings also reveal apparent group difference between ASD and healthy controls in network profiles and switching dynamics of brain states. Fuad Noman, Sin-Yee Yap, Raphael C.-W. Phan, Hernando C. Ombao, Chee-Ming Ting |
ICIP | 3 |
| 2022 | Persistent Items Tracking in Large Data Streams Based on Adaptive SamplingabstractWe address the problem of persistent item tracking in large-scale data streams. A persistent item refers to the one that persists to occur in the stream over a long timespan. Tracking persistent items is an important and pivotal functionality for many networking and computing applications as persistent items, though not necessarily contributing significantly to the data volume, may convey valuable information on the data pattern about the stream. The state-of-the-art solutions of tracking persistent items require to know the monitoring time horizon to set the sampling rate. This limitation is further accentuated when we need to track the persistent items in recent w slots where w can be any value between 0 and T to support different monitoring granularity. Motivated by this limitation, we develop a persistent item tracking algorithm that can function without knowing the monitoring time horizon beforehand, and can thus track persistent items up to the current time t or within a certain time window at any moment. Our central technicality is adaptively reducing the sampling rate such that the total memory overhead can be limited while still meeting the target tracking accuracy. Through both theoretical and empirical analysis, we fully characterize the performance of our proposition. Lin Chen 0002, Raphael C.-W. Phan, Dan Huang 0001 |
INFOCOM | 2 |
| 2022 | Guess-It-Generator: Generating in a Lewis Signaling Framework through Logical ReasoningabstractHuman minds spontaneously integrate two inherited cognitive capabilities: perception and reasoning to accomplish cognitive tasks such as problem solving, imagination, and causation. It is observed in the primate brains that perception offers the assistance required for problem comprehension, whilst the reasoning elucidates upon the facts recovered during perception in order to make a decision. The field of artificial intelligence (AI) thus considers perception and reasoning as two complementary areas that are realized by machine learning and logic programming, respectively. In this work, we propose a generative model using a collaborative guessing game of the kind first introduced by David Lewis in his famous work called the Lewis signaling game that is synonymous with the "20 Questions'' game. Our proposed model, Guess-It-Generator (GIG) is a collaborative framework that engages two recurrent neural networks in a guessing game. GIG unifies perception and reasoning with a view to generating labeled images by capturing, (X, y), the underlying density of a data distribution, i.e. (X, y) - p(X, y). An encoder attends to a region of the input image and encodes that onto a latent variable that acts as a perception signal to a decoder. In contrast, the decoder leverages on the perception signals to guess the image and verifies the guess by reasoning with logical facts derived from the domain knowledge. Our experiments and comprehensive studies on seven datasets: PCAM, Chest-Xray-14, FIRE, HAM10000 from the medical domain, and CIFAR 10, LSUN, ImageNet, among standard benchmark datasets, show significant promise for the proposed method. Arghya Pal, Sailaja Rajanala, Raphael C.-W. Phan, Koksheik Wong |
ACM Multimedia | 3 |
| 2022 | DeSCoVeR: Debiased Semantic Context Prior for Venue RecommendationabstractWe present a novel semantic context prior-based venue recommendation system that uses only the title and the abstract of a paper. Based on the intuition that the text in the title and abstract have both semantic and syntactic components, we demonstrate that a joint training of a semantic feature extractor and syntactic feature extractor collaboratively leverages meaningful information that helps to provide venues for papers. The proposed methodology that we call DeSCoVeR at first elicits these semantic and syntactic features using a Neural Topic Model and text classifier respectively. The model then executes a transfer learning optimization procedure to perform a contextual transfer between the feature distributions of the Neural Topic Model and the text classifier during the training phase. DeSCoVeR also mitigates the document-level label bias using a Causal back-door path criterion and a sentence-level keyword bias removal technique. Experiments on the DBLP dataset show that DeSCoVeR outperforms the state-of-the-art methods. Sailaja Rajanala, Arghya Pal, Manish Singh 0002, Raphael C.-W. Phan, Koksheik Wong |
SIGIR | 4 |
| 2022 | A comprehensive overview of Deepfake: Generation, detection, datasets, and opportunities
Jia Wen Seow, Mei Kuan Lim, Raphael C.-W. Phan, Joseph K. Liu |
Neurocomputing | 3 |
| 2022 | Invisible emotion magnification algorithm (IEMA) for real-time micro-expression recognition with graph-based features
Adamu Muhammad Buhari, Chee-Pun Ooi, Vishnu Monn Baskaran, Raphael C.-W. Phan, Koksheik Wong, Wooi-Haw Tan |
Multim. Tools Appl. | 4 |
| 2021 | Synthesize-It-Classifier: Learning a Generative Classifier Through Recurrent Self-AnalysisabstractWe show the generative capability of an image classifier network by synthesizing high-resolution, photo-realistic, and diverse images at scale. The overall methodology, called Synthesize-It-Classifier (STIC), does not require an explicit generator network to estimate the density of the data distribution and sample images from that, but instead uses the classifier’s knowledge of the boundary to perform gradient ascent w.r.t. class logits and then synthesizes images using the Gram Matrix Metropolis Adjusted Langevin Algorithm (GRMALA) by drawing on a blank canvas. During training, the classifier iteratively uses these synthesized images as fake samples and re-estimates the class boundary in a recurrent fashion to improve both the classification accuracy and quality of synthetic images. The STIC shows that mixing of the hard fake samples (i.e. those synthesized by the one-hot class conditioning), and the soft fake samples (which are synthesized as a convex combination of classes, i.e. a mixup of classes [36]) improves class interpolation. We demonstrate an Attentive-STIC network that shows iterative drawing of synthesized images on the ImageNet dataset that has thousands of classes. In addition, we introduce the synthesis using a class conditional score classifier (Score-STIC) instead of a normal image classifier and show improved results on several real world datasets, i.e. ImageNet, LSUN and CIFAR 10. Arghya Pal, Raphael C.-W. Phan, Koksheik Wong |
CVPR | 2 |
| 2021 | Network Reconfiguration via Diversity: Theoretical Foundation and Algorithm DesignabstractMoving Target Defense (MTD) is a powerful weapon to mitigate cyber attacks by increasing the attacker's efforts and complexity in fulfilling its goal. One effective technique of MTD is to deploy diverse implementations and configurations to provide equivalent functionality so as to increase the network resilience. In this paper, we investigate the algorithmic aspect of employing diversity to cause the network to be the most resilient possible. Specifically, we study two algorithmic optimization problems of both theoretical and practical importance: (1) given the network topology, how to assign different variants to different network nodes so as to maximize the network resilience; (2) when the variant assignment is fixed, but the network topology is configurable, what is the optimal topology maximizing the network resilience. We mathematically formulate the problems of variant assignment and network topology configuration and develop efficient algorithms, which can serve as design guidelines in the deployment of diversity-based MTD techniques to enhance the security of the network. Lin Chen 0002, Raphael C.-W. Phan |
VTC Fall | 2 |
| 2021 | Faceless identification based on temporal strips
Shu-Min Leong, Raphael C.-W. Phan, Vishnu Monn Baskaran, Chee-Pun Ooi |
Multim. Tools Appl. | 2 |
| 2021 | Strengthening speech content authentication against tampering
Raphael C.-W. Phan, Yin Yin Low, Koksheik Wong, Kazuki Minemura |
Speech Commun. | 1 |
| 2020 | Optimized IoT Cryptoprocessor Based on QC-MPDC Key Encapsulation MechanismabstractThe key encapsulation mechanism (KEM) is an important cryptographic tool to protect communication in the Internet of Things (IoT). In the near future, classical algorithms used to construct KEMs, such as RSA and elliptic curve cryptography, will be vulnerable to attacks from quantum computers. Recently, Yamada et al. proposed the quasicyclic medium density parity check (QC-MDPC) KEM, which is considered one of the most advanced code-based cryptosystems to resist quantum attacks. In this article, an optimized implementation of QC-MDPC KEM for IoT applications is presented. Our main contributions are threefold: 1) the fastest QC-MDPC McEliece decryption in field-programmable gate array (FPGA); 2) the first QC-MDPC KEM implementation in FPGA; and 3) the first iteration count attack-resistant QC-MDPC decoder in FPGA. To improve the decryption speed, we introduce a novel customized rotation engine (CRE) and incorporated several recent techniques reported in the literature, including adaptive threshold and Hamming weight estimation. The best-achieved throughput in our implementation on Xilinx Virtex 7 FPGA is 12.7% faster than the state-of-the-art result reported by Heyse et al. The proposed CRE was then integrated with QC-MDPC KEM to produce a fast and secure KEM. Furthermore, to prevent timing attacks demonstrated recently, a constant-time implementation of the QC-MDPC McEliece decoder was presented. Jun-Hoe Phoon, Wai-Kong Lee, Denis Chee-Keong Wong, Wun-She Yap, Bok-Min Goi, Raphael C.-W. Phan |
IEEE Internet Things J. | 6 |
| 2020 | Reduced contact lifting of latent fingerprints from curved surfaces
Mohammad Mogharen Askarin, Koksheik Wong, Raphael C.-W. Phan |
J. Inf. Secur. Appl. | 3 |
| 2020 | Cryptanalysis of genetic algorithm-based encryption scheme
Kuan-Wai Wong, Wun-She Yap, Denis Chee-Keong Wong, Raphael C.-W. Phan, Bok-Min Goi |
Multim. Tools Appl. | 4 |
| 2019 | Dual-stream Shallow Networks for Facial Micro-expression RecognitionabstractMicro-expressions are spontaneous, brief and subtle facial muscle movements that exposes underlying emotions. Motivated by recent exploits into deep learning for micro-expression analysis, we propose a lightweight dual-stream shallow network in the form of a pair of truncated CNNs with heterogeneous input features. The merging of the convolutional features allows for discriminative learning of micro-expression classes stemming from both streams. Using activation heatmaps, we further demonstrate that salient facial areas are well emphasized, and correspond closely to relevant action units belonging to emotion classes. We empirically validate the proposed network on three benchmark databases, obtaining state-of-the-art performance on the CASME II and SAMM while remaining competitive on the SMIC. Further observations point towards the sufficiency of utilizing shallower deep networks for micro-expression recognition. Huai-Qian Khor, John See, Sze-Teng Liong, Raphael C.-W. Phan, Weiyao Lin |
ICIP | 4 |
| 2019 | Terabit encryption in a second: Performance evaluation of block ciphers in GPU with Kepler, Maxwell, and Pascal architecturesabstractSummary With the emergence of IoT and cloud computing technologies, massive data are generated from various applications everyday and communicated through the Internet. Secure communication is essential to protect these data from malicious attacks. Block ciphers are one mechanism to offer such protection but unfortunately involve intensive computations that can be performance bottlenecks to the servers, especially when the data center needs to handle thousands of concurrent transactions. In this paper, we investigate the feasibility of the GPU as an accelerator to perform high‐speed encryption in server environments. We present optimized implementations of a conventional block cipher (AES) and new lightweight block ciphers (LEA, Chaskey, SIMON, SPECK, and SIMECK) across three new GPU architectures (Kepler, Maxwell, and Pascal). For AES, we improve the fine‐grain implementation by utilizing the warp shuffle instruction available in these three new GPU architectures, which yield a 6%‐16% improvement over the previous implementations. For LEA, Chaskey, SIMON, SPECK, and SIMECK, we first analyze why they cannot have efficient fine‐grain implementations in the GPU and then present our optimization techniques, which are able to achieve impressive encryption speeds of 1.912, 637, 1.485, 2.291, and 1.478 Tb/s, respectively, in GTX1080. Wai-Kong Lee, Bok-Min Goi, Raphael C.-W. Phan |
Concurr. Comput. Pract. Exp. | 3 |
| 2019 | Signature Gateway: Offloading Signature Generation to IoT Gateway Accelerated by GPUabstractThe emergence of Internet of Things (IoT) brings us the possibility to form a well connected network for ubiquitous sensing, intelligent analysis, and timely actuation, which opens up many innovative applications in our daily life. To secure the communication between sensor nodes, gateway devices and cloud servers, cryptographic algorithms (e.g., digital signature, block cipher, and hash function) are widely used. Although cryptographic algorithms are effective in preventing malicious attacks, they involve heavy computation that may not be executed efficiently in resource constraint sensor nodes. In particular, the authentication of a sensor node is usually performed through a digital signature (e.g., RSA and elliptic curve cryptography), which can be slow when executed on a microcontroller. In this paper, an IoT architecture that offloads the digital signature generation to a nearby signature gateway equipped with graphic processing unit (GPU) accelerator are proposed. The communication process for signature offloading, together with optimized implementation techniques for RSA in signature gateway, are also presented in this paper. We have evaluated two different ways to implement modular exponentiation in RSA, namely residue number system and multiprecision montgomery multiplication (MPMM). The experimental results show that our RSA implementation using MPMM is 10.1% faster than the best RSA implementation in GPU. Our proposed IoT architecture with signature gateway can successfully reduce the burden of sensor nodes to generate signatures, at the same time preserve the ability to authenticate the sensor nodes. Chin-Chen Chang 0001, Wai-Kong Lee, Yanjun Liu 0002, Bok-Min Goi, Raphael C.-W. Phan |
IEEE Internet Things J. | 5 |
| 2019 | Paradigm Shifts in Cryptographic EngineeringabstractThe papers in this special section identify paradigm shifts in cryptographic engineering. Modern cryptography is almost five decades old, and we have seen some interesting breakthroughs throughout its relatively young history. These include the engineering foundations of symmetric key cryptography and block ciphers in particular (like the DES and the AES), the discovery of public-key cryptography, protocols for secure computations for general and specific tasks using interactions among parties and tools like homomorphic encryption (this line of work has been enhancing the use of cryptography beyond secure messaging into secure computing). Raphael C.-W. Phan, Moti Yung |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2018 | Enriched Long-Term Recurrent Convolutional Network for Facial Micro-Expression RecognitionabstractFacial micro-expression (ME) recognition has posed a huge challenge to researchers for its subtlety in motion and limited databases. Recently, handcrafted techniques have achieved superior performance in micro-expression recognition but at the cost of domain specificity and cumbersome parametric tunings. In this paper, we propose an Enriched Long-term Recurrent Convolutional Network (ELRCN) that first encodes each micro-expression frame into a feature vector through CNN module(s), then predicts the micro-expression by passing the feature vector through a Long Short-term Memory (LSTM) module. The framework contains 2 different network variants: (1) Channel-wise stacking of input data for spatial enrichment, (2) Feature-wise stacking of features for temporal enrichment. We demonstrate that the proposed approach is able to achieve reasonably good performance, without data augmentation. In addition, we also present ablation studies conducted on the framework and visualizations of what CNN "sees" when predicting the micro-expression classes. Huai-Qian Khor, John See, Raphael C.-W. Phan, Weiyao Lin |
FG | 3 |
| 2018 | Micro-Expression Motion Magnification: Global Lagrangian vs. Local Eulerian ApproachesabstractMicro-expressions are difficult to spot but are utterly important for engaging in a conversation or negotiation. Through motion magnification, these expressions become much more distinguishable and easily recognized. This work proposes Global Lagrangian Motion Magnification (GLMM) for consistent exaggeration of facial expressions and dynamics across a whole video. As the proposal takes an opposite approach to a previous pivotal work, i.e. local Amplitude-based Eulerian Motion Magnification (AEMM). GLMM and AEMM are theoretically analyzed for potential advantages and disadvantages, especially with respect to how magnified noise and distortions are dealt with. Then, both GLMM and AEMM are empirically evaluated and compared using the CASME II micro-expression corpus. Anh Cat Le Ngo, Alan Johnston, Raphael C.-W. Phan, John See |
FG | 3 |
| 2018 | Cryptography and Future Security
Jongsung Kim, Hongjun Wu 0001, Raphael C.-W. Phan |
Discret. Appl. Math. | 3 |
| 2018 | Separable authentication in encrypted HEVC video
Yiqi Tew, Koksheik Wong, Raphael C.-W. Phan, King Ngi Ngan |
Multim. Tools Appl. | 3 |
| 2018 | Less is more: Micro-expression recognition from video using apex frame
Sze-Teng Liong, John See, Koksheik Wong, Raphael C.-W. Phan |
Signal Process. Image Commun. | 4 |
| 2017 | Higher order differentiation over finite fields with applications to generalising the cube attackabstractHigher order differentiation was introduced in a cryptographic context by Lai. Several attacks can be viewed in the context of higher order differentiations, amongst them the cube attack of Dinur and Shamir and the AIDA attack of Vielhaber. All of the above have been developed for the binary case. We examine differentiation in larger fields, starting with the field $$\mathrm {GF}(p)$$ of integers modulo a prime p, and apply these techniques to generalising the cube attack to $$\mathrm {GF}(p)$$ . The crucial difference is that now the degree in each variable can be higher than one, and our proposed attack will differentiate several times with respect to each variable (unlike the classical cube attack and its larger field version described by Dinur and Shamir, both of which differentiate at most once with respect to each variable). Connections to the Moebius/Reed Muller Transform over $$\mathrm {GF}(p)$$ are also examined. Finally we describe differentiation over finite fields $$\mathrm {GF}(p^s)$$ with $$p^s$$ elements and show that it can be reduced to differentiation over $$\mathrm {GF}(p)$$ , so a cube attack over $$\mathrm {GF}(p^s)$$ would be equivalent to cube attacks over $$\mathrm {GF}(p)$$ . Ana Salagean, Richard Winter, Matei Mandache-Salagean, Raphael C.-W. Phan |
Des. Codes Cryptogr. | 4 |
| 2017 | Effective recognition of facial micro-expressions with video motion magnification
Yandan Wang, John See, Yee-Hui Oh, Raphael C.-W. Phan, Yo Rahul, Huo-Chong Ling, Su-Wei Tan, Xujie Li 0002 |
Multim. Tools Appl. | 4 |
| 2017 | Sparsity in Dynamics of Spontaneous Subtle Emotions: Analysis and ApplicationabstractSubtle emotions are present in diverse real-life situations: in hostile environments, enemies and/or spies maliciouslyconceal their emotions as part of their deception; in life-threatening situations, victims under duress have no choice but to withhold theirreal feelings; in the medical scene, patients with psychological conditions such as depression could either be intentionally orsubconsciously suppressing their anguish from loved ones. Under such circumstances, it is often crucial that these subtle emotions arerecognized before it is too late. These spontaneous subtle emotions are typically expressed through micro-expressions, which are tiny,sudden and short-lived dynamics of facial muscles; thus, such micro-expressions pose a great challenge for visual recognition. Theabrupt but significant dynamics for the recognition task are temporally sparse while the rest, i.e. irrelevant dynamics, are temporallyredundant. In this work, we analyze and enforce sparsity constraints to learn significant temporal and spectral structures whileeliminating irrelevant facial dynamics of micro-expressions, which would ease the challenge in the visual recognition of spontaneoussubtle emotions. The hypothesis is confirmed through experimental results of automatic spontaneous subtle emotion recognition withseveral sparsity levels on CASME II and SMIC, the two well-established and publicly available spontaneous subtle emotion databases.The overall performances of the automatic subtle emotion recognition are boosted when only significant dynamics of the originalsequences are preserved. Anh Cat Le Ngo, John See, Raphael C.-W. Phan |
IEEE Trans. Affect. Comput. | 3 |
| 2017 | A Novel Sketch Attack for H.264/AVC Format-Compliant Encrypted VideoabstractIn this paper, we propose a novel sketch attack for H.264 advanced video coding (H.264/AVC) format-compliant encrypted video. We briefly describe the notion of sketch attack, review the conventional sketch attacks designed for discrete cosine transform (DCT)-based compressed image, and identify their shortcomings when applied to attack compressed video. Specifically, the conventional DCT-based sketch attacks are incapable in sketching outlines for inter frame, which is deployed to significantly reduce temporal redundancy in video compression. To sketch directly from inter frame, we put forward a sketch attack by considering the partially decoded information of the H.264/AVC compressed video, namely, the number of bits spent on coding a macroblock. To evaluate the sketch image, we consider the Canny edge map as the ideal outline image. Experiments are conducted to verify the performance of the proposed sketch attack using ICADR2013, High Efficiency Video Coding dash, and Xiph video data sets. Results suggest that the proposed sketch attack can generate the outline image of the original frame for not only intra frame but also inter frame. Kazuki Minemura, Koksheik Wong, Raphael C.-W. Phan, Kiyoshi Tanaka |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | Eulerian emotion magnification for subtle expression recognitionabstractSubtle emotions are expressed through tiny and brief movements of facial muscles, called micro-expressions; thus, recognition of these hidden expressions is as challenging as inspection of microscopic worlds without microscopes. In this paper, we show that through motion magnification, subtle expressions can be realistically exaggerated and become more easily recognisable. We magnify motions of facial expressions in the Eulerian perspective by manipulating their amplitudes or phases. To evaluate effects of exaggerating facial expressions, we use a common framework (LBP-TOP features and SVM classifiers) to perform 5-class subtle emotion recognition on the CASME II corpus, a spontaneous subtle emotion database. According to experimental results, significant improvements in recognition rates of magnified micro-expressions over normal ones are confirmed and measured. Furthermore, we estimate upper bounds of effective magnification factors and empirically corroborate these theoretical calculations with experimental data. Anh Cat Le Ngo, Yee-Hui Oh, Raphael C.-W. Phan, John See |
ICASSP | 3 |
| 2016 | Intrinsic two-dimensional local structures for micro-expression recognitionabstractAn elapsed facial emotion involves changes of facial contour due to the motions (such as contraction or stretch) of facial muscles located at the eyes, nose, lips and etc. Thus, the important information such as corners of facial contours that are located in various regions of the face are crucial to the recognition of facial expressions, and even more apparent for micro-expressions. In this paper, we propose the first known notion of employing intrinsic two-dimensional (i2D) local structures to represent these features for micro-expression recognition. To retrieve i2D local structures such as phase and orientation, higher order Riesz transforms are employed by means of monogenic curvature tensors. Experiments performed on micro-expression datasets show the effectiveness of i2D local structures in recognizing micro-expressions. Yee-Hui Oh, Anh Cat Le Ngo, Raphael C.-W. Phan, John See, Huo-Chong Ling |
ICASSP | 3 |
| 2016 | Smart, secure and seamless access control scheme for mobile devicesabstractSmart devices capture users' activity such as unlock failures, application usage, location and proximity of devices in and around their surrounding environment. This activity information varies between users and can be used as digital fingerprints of the users' behaviour. Traditionally, users are authenticated to access restricted data using long term static attributes such as password and roles. In this paper, in order to allow secure and seamless data access in mobile environment, we combine both the user behaviour captured by the smart device and the static attributes to develop a novel access control technique. Security and performance analyses show that the proposed scheme substantially reduces the computational complexity while enhances the security compared to the conventional schemes. Yo Rahul, Muttukrishnan Rajarajan, Raphael C.-W. Phan |
ICC | 3 |
| 2016 | Multi-layer authentication scheme for HEVC video based on embedded statistics
Yiqi Tew, Koksheik Wong, Raphael C.-W. Phan, King Ngi Ngan |
J. Vis. Commun. Image Represent. | 3 |
| 2016 | On the effective subkey space of some image encryption algorithms using external key
Wun-She Yap, Raphael C.-W. Phan, Bok-Min Goi, Wei-Chuen Yau, Swee-Huay Heng |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | Spontaneous subtle expression detection and recognition based on facial strain
Sze-Teng Liong, John See, Raphael C.-W. Phan, Yee-Hui Oh, Anh Cat Le Ngo, Koksheik Wong, Su-Wei Tan |
Signal Process. Image Commun. | 3 |
| 2015 | HEVC video authentication using data embedding techniqueabstractA HEVC video authentication scheme by utilizing data embedding technique is proposed. The concept of authentication, layout and implementation are described under the HEVC standard. The authentication scheme includes weight generation, video feature extraction and two layers of authentication. Simulation results confirm that the overall perceptual video quality is maintained after the insertion of authentication code into commonly considered classes of video sequence. By analyzing the behavior of video tampering within and across video slices, the proposed authentication scheme is able to detect the tampered region and verify the integrity of the video in question. Yiqi Tew, Koksheik Wong, Raphael C.-W. Phan |
ICIP | 3 |
| 2015 | Comparison of Cube Attacks Over Different Vector Spaces
Richard Winter, Ana Salagean, Raphael C.-W. Phan |
IMACC | 3 |
| 2015 | Cryptanalysis of a certificateless identification schemeabstractABSTRACT In 2013, Dehkordi and Alimoradi proposed a certificateless identification scheme using supersingular elliptic curves. This proposal came independent of the parallel work of Chin et al. in proposing the first known security models for certificateless identification with provable security. In this paper, we show that there are some design flaws in the Dehkordi–Alimoradi scheme, which lead one to conclude that their scheme is insecure. Copyright © 2014 John Wiley & Sons, Ltd. Ji-Jian Chin, Rouzbeh Behnia, Swee-Huay Heng, Raphael C.-W. Phan |
Secur. Commun. Networks | 4 |
| 2014 | Spontaneous Subtle Expression Recognition: Imbalanced Databases and Solutions
Anh Cat Le Ngo, Raphael C.-W. Phan, John See |
ACCV (4) | 2 |
| 2014 | LBP with Six Intersection Points: Reducing Redundant Information in LBP-TOP for Micro-expression Recognition
Yandan Wang, John See, Raphael C.-W. Phan, Yee-Hui Oh |
ACCV (1) | 3 |
| 2014 | Privacy-Preserving Multi-Class Support Vector Machine for Outsourcing the Data Classification in CloudabstractEmerging cloud computing infrastructure replaces traditional outsourcing techniques and provides flexible services to clients at different locations via Internet. This leads to the requirement for data classification to be performed by potentially untrusted servers in the cloud. Within this context, classifier built by the server can be utilized by clients in order to classify their own data samples over the cloud. In this paper, we study a privacy-preserving (PP) data classification technique where the server is unable to learn any knowledge about clients’ input data samples while the server side classifier is also kept secret from the clients during the classification process. More specifically, to the best of our knowledge, we propose the first known client-server data classification protocol using support vector machine. The proposed protocol performs PP classification for both two-class and multi-class problems. The protocol exploits properties of Pailler homomorphic encryption and secure two-party computation. At the core of our protocol lies an efficient, novel protocol for securely obtaining the sign of Pailler encrypted numbers. Yo Rahul, Raphael C.-W. Phan, Suresh Veluru 0001, K. Cumanan, Muttukrishnan Rajarajan |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2014 | Privacy-Preserving Clinical Decision Support System Using Gaussian Kernel-Based ClassificationabstractA clinical decision support system forms a critical capability to link health observations with health knowledge to influence choices by clinicians for improved healthcare. Recent trends toward remote outsourcing can be exploited to provide efficient and accurate clinical decision support in healthcare. In this scenario, clinicians can use the health knowledge located in remote servers via the Internet to diagnose their patients. However, the fact that these servers are third party and therefore potentially not fully trusted raises possible privacy concerns. In this paper, we propose a novel privacy-preserving protocol for a clinical decision support system where the patients' data always remain in an encrypted form during the diagnosis process. Hence, the server involved in the diagnosis process is not able to learn any extra knowledge about the patient's data and results. Our experimental results on popular medical datasets from UCI-database demonstrate that the accuracy of the proposed protocol is up to 97.21% and the privacy of patient data is not compromised. Yo Rahul, Suresh Veluru 0001, Raphael C.-W. Phan, Jonathon A. Chambers, Muttukrishnan Rajarajan |
IEEE J. Biomed. Health Informatics | 3 |
| 2013 | Vickrey-Clarke-Groves for privacy-preserving collaborative classification
Anastasia Panoui, Sangarapillai Lambotharan, Raphael C.-W. Phan |
FedCSIS | 3 |
| 2013 | Efficient Generation of Elementary Sequences
David Gardner, Ana Salagean, Raphael C.-W. Phan |
IMACC | 3 |
| 2013 | An Efficient and Provably Secure Certificateless Identification Scheme
Ji-Jian Chin, Raphael C.-W. Phan, Rouzbeh Behnia, Swee-Huay Heng |
SECRYPT | 2 |
| 2013 | On the Security of the XOR Sandwiching Paradigm for Multiple Keyed Block Ciphers
Ruth Ng Ii-Yung, Khoongming Khoo, Raphael C.-W. Phan |
SECRYPT | 3 |
| 2013 | An authentication framework for peer-to-peer cloudabstractCloud computing provides on demand computation and storage services delivered via applications, system software and hardware rendered as services. Due to its on demand nature, it has high variable workloads and requires real-time efficiency and availability. Most cloud computing systems use a centralised model to provision services, but reliance on a central entity to control scheduling decision and maintain all cloud hosts may constitute a computing bottleneck. A system failure will cause service outage, sometimes for a few hours as had happened before. In addition, the central entity needs to support heavy workloads in terms of service provisioning to all resource hosts. These issues can be addressed by distributing cloud resources using structured peer-to-peer (P2P) overlay networks as was recently proposed. However these proposals do not examine potential security issues of a P2P-based cloud, one of them being how peers verify the identities of one another over a decentralised setting. Therefore we propose an authentication framework for P2P cloud consisting of various approaches for authenticating entities and messages. The framework combines cryptographic primitives and security mechanisms proposed for existing structured P2P network. Geong Sen Poh, Mohd Amril Nurman Mohd Nazir, Bok-Min Goi, Syh-Yuan Tan, Raphael C.-W. Phan, Maryam Safiyah Shamsudin |
SIN | 5 |
| 2013 | On the security of a modified Beth identity-based identification scheme
Ji-Jian Chin, Syh-Yuan Tan, Swee-Huay Heng, Raphael C.-W. Phan |
Inf. Process. Lett. | 4 |
| 2013 | Facial Expression Recognition in the Encrypted Domain Based on Local Fisher Discriminant AnalysisabstractFacial expression recognition forms a critical capability desired by human-interacting systems that aim to be responsive to variations in the human's emotional state. Recent trends toward cloud computing and outsourcing has led to the requirement for facial expression recognition to be performed remotely by potentially untrusted servers. This paper presents a system that addresses the challenge of performing facial expression recognition when the test image is in the encrypted domain. More specifically, to the best of our knowledge, this is the first known result that performs facial expression recognition in the encrypted domain. Such a system solves the problem of needing to trust servers since the test image for facial expression recognition can remain in encrypted form at all times without needing any decryption, even during the expression recognition process. Our experimental results on popular JAFFE and MUG facial expression databases demonstrate that recognition rate of up to 95.24 percent can be achieved even in the encrypted domain. Yo Rahul, Raphael C.-W. Phan, Jonathon A. Chambers, David J. Parish |
IEEE Trans. Affect. Comput. | 2 |
| 2012 | Index Tables of Finite Fields and Modular Golomb Rulers
Ana Salagean, David Gardner, Raphael C.-W. Phan |
SETA | 3 |
| 2012 | A Formally Verified Device Authentication Protocol Using Casper/FDRabstractFor communication in Next Generation Networks, highly-developed mobile devices will enable users to store and manage a lot of credentials on their terminals. Furthermore, these terminals will represent and act on behalf of users when accessing different networks and connecting to a wide variety of services. In this situation, it is essential for users to trust their terminals and for all transactions using them to be secure. This paper analyses a number of the Authentication and Key Agreement protocols between the users and mobile terminals, then proposes a novel device authentication protocol. The proposed protocol is analysed and verified using a formal methods approach based on Casper/FDR compiler. Mahdi Aiash, Glenford E. Mapp, Raphael C.-W. Phan, Aboubaker Lasebae, Jonathan Loo |
TrustCom | 3 |
| 2012 | Efficient encryption with keyword search in mobile networksabstractABSTRACT On these days, users tend to access to online content via mobile devices, for example, e‐mails. Because these devices have constrained resources, users may wish to instruct e‐mail gateways to search through new e‐mails and only download those corresponding to particular keywords, such as “urgent.” Yet, this searching should not compromise the user's privacy. A public key encryption with keyword search (PEKS) scheme achieves both these requirements. Most PEKS schemes are constructed on the basis of bilinear pairings. Recently, Khader proposed the first PEKS scheme that does not require bilinear pairings and is provably indistinguishable chosen‐keyword attack (IND‐CKA) secure in the standard model. Such a scheme is more efficient than pairing‐based ones. In this paper, we show a drawback of Khader's scheme in that it depends on an unnecessary security assumption: Its IND‐CKA security requires its underlying identity‐based encryption building block to be indistinguishable chosen‐ciphertext attack secure. We construct a more efficient PEKS scheme that achieves the same level of PEKS security as Khader's but that only requires the underlying identity‐based encryption to be indistinguishable chosen‐plaintext attack secure. We give a direct proof that the proposed scheme is IND‐CKA secure. Our scheme outperforms other recent PEKS schemes in literature. Copyright © 2012 John Wiley & Sons, Ltd. Wei-Chuen Yau, Swee-Huay Heng, Syh-Yuan Tan, Bok-Min Goi, Raphael C.-W. Phan |
Secur. Commun. Networks | 5 |
| 2011 | Cryptanalysis of a Provably Secure Cross-Realm Client-to-Client Password-Authenticated Key Agreement Protocol of CANS '09
Wei-Chuen Yau, Raphael C.-W. Phan, Bok-Min Goi, Swee-Huay Heng |
CANS | 2 |
| 2011 | On the Stability of m-Sequences
Alex J. Burrage, Ana Salagean, Raphael C.-W. Phan |
IMACC | 3 |
| 2011 | Linear complexity for sequences with characteristic polynomial ƒvabstractWe present several generalisations of the Games-Chan algorithm. For a fixed monic irreducible polynomial f we consider the sequences s that have as characteristic polynomial a power of f. We propose an algorithm for computing the linear complexity of s given a full (not necessarily minimal) period of s. We give versions of the algorithm for fields of characteristic 2 and for arbitrary finite characteristic p, the latter generalising an algorithm of Kaida et al. We also propose an algorithm which computes the linear complexity given only a finite portion of s (of length greater than or equal to the linear complexity), generalising an algorithm of Meidl. All our algorithms have linear computational complexity. The algorithms for computing the linear complexity when a full period is known can be further generalised to sequences for which it is known a priori that the irreducible factors of the minimal polynomial belong to a given small set of polynomials. Alex J. Burrage, Ana Salagean, Raphael C.-W. Phan |
ISIT | 3 |
| 2011 | On the Security of a Hybrid SVD-DCT Watermarking Method Based on LPSNR
Huo-Chong Ling, Raphael C.-W. Phan, Swee-Huay Heng |
PSIVT (1) | 2 |
| 2011 | Notions and relations for RKA-secure permutation and function families
Jongsung Kim, Jaechul Sung, Ermaliza Razali, Raphael C.-W. Phan, Marc Joye |
Des. Codes Cryptogr. | 4 |
| 2011 | On the cryptanalysis of the hash function Fugue: Partitioning and inside-out distinguishers
Jean-Philippe Aumasson, Raphael C.-W. Phan |
Inf. Process. Lett. | 2 |
| 2011 | VLSI Characterization of the Cryptographic Hash Function BLAKEabstractCryptographic hash functions are used to protect information integrity and authenticity in a wide range of applications. After the discovery of weaknesses in the current deployed standards, the U.S. Institute of Standards and Technology started a public competition to develop the future standard SHA-3, which will be implemented in a multitude of environments, after its selection in 2012. In this paper, we investigate high-speed and low-area hardware architectures of one of the 14 “second-round” candidates in this competition: BLAKE. VLSI performance results of the proposed high-speed designs indicate a throughput improvement between 16% and 36% compared to the current standard SHA-2. Additionally, we propose a compact implementation of BLAKE with memory optimization that fits in 0.127 mm2of a 0.18 μ m CMOS. Measurements reveal a minimal power dissipation of 9.59 μW/MHz at 0.65 V, which suggests that BLAKE is suitable for resource-limited systems. Luca Henzen, Jean-Philippe Aumasson, Willi Meier, Raphael C.-W. Phan |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2011 | Non-repudiable authentication and billing architecture for wireless mesh networks
Raphael C.-W. Phan |
Wirel. Networks | 1 |
| 2010 | Integral Distinguishers of Some SHA-3 Candidates
Marine Minier, Raphael C.-W. Phan, Benjamin Pousse |
CANS | 2 |
| 2010 | Quasi-Linear Cryptanalysis of a Secure RFID Ultralightweight Authentication Protocol
Pedro Peris-Lopez, Julio César Hernández Castro, Raphael C.-W. Phan, Juan Tapiador, Tieyan Li |
Inscrypt | 3 |
| 2010 | Security Models for Heterogeneous Networking
Glenford E. Mapp, Mahdi Aiash, Aboubaker Lasebae, Raphael C.-W. Phan |
SECRYPT | 4 |
| 2009 | Improved Cryptanalysis of Skein
Jean-Philippe Aumasson, Çagdas Çalik, Willi Meier, Onur Özen, Raphael C.-W. Phan, Kerem Varici |
ASIACRYPT | 5 |
| 2009 | On Hashing with Tweakable CiphersabstractCryptographic hash functions are often built on block ciphers in order to reduce the security analysis of the hash to that of the cipher, and to minimize the hardware size. Well known hash constructs are used in international standards like MD5 and SHA-1. Recently, researchers proposed new modes of operations for hash functions to protect against generic attacks, and it remains open how to base such functions on block ciphers. An attracting and intuitive choice is to combine previous constructions with tweakable block ciphers. We investigate such constructions, and show the surprising result that combining a provably secure mode of operation with a provably secure tweakable cipher does not guarantee the security of the constructed hash function. In fact, simple attacks can be possible when the interaction between secure components leaves some additional "freedom" to an adversary. Our techniques are derived from the principle of slide attacks, which were introduced for attacking block ciphers. Raphael C.-W. Phan, Jean-Philippe Aumasson |
ICC | 1 |
| 2009 | Analysis of Two Pairing-Based Three-Party Password Authenticated Key Exchange ProtocolsabstractPassword-Authenticated Key Exchange (PAKE) protocols allow parties to share secret keys in an authentic manner based on an easily memorizable password. Recently, Nam et al. showed that a provably secure three-party password-based authenticated key exchange protocol using Weil pairing by Wen et al. is vulnerable to a man-in-the-middle attack. In doing so, Nam et al. showed the flaws in the proof of Wen et al. and described how to fix the problem so that their attack no longer works. In this paper, we show that both Wen et al. and Nam et al. variants fall to key compromise impersonation by any adversary. Our results underline the fact that although the provable security approach is necessary to designing PAKEs, gaps still exist between what can be proven and what are really secure in practice. Raphael C.-W. Phan, Wei-Chuen Yau, Bok-Min Goi |
NSS | 1 |
| 2009 | Cryptanalysis of a New Ultralightweight RFID Authentication Protocol—SASIabstractSince RFID tags are ubiquitous and at times even oblivious to the human user, all modern RFID protocols are designed to resist tracking so that the location privacy of the human RFID user is not violated. Another design criterion for RFIDs is the low computational effort required for tags, in view that most tags are passive devices that derive power from an RFID reader's signals. Along this vein, a class of ultralightweight RFID authentication protocols has been designed, which uses only the most basic bitwise and arithmetic operations like exclusive-OR, OR, addition, rotation, and so forth. In this paper, we analyze the security of the SASI protocol, a recently proposed ultralightweight RFID protocol with better claimed security than earlier protocols. We show that SASI does not achieve resistance to tracking, which is one of its design objectives. Raphael C.-W. Phan |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2008 | Traceable Privacy of Recent Provably-Secure RFID Protocols
Khaled Ouafi, Raphael C.-W. Phan |
ACNS | 2 |
| 2008 | The Hash Function Family LAKE
Jean-Philippe Aumasson, Willi Meier, Raphael C.-W. Phan |
FSE | 3 |
| 2008 | Privacy of Recent RFID Authentication Protocols
Khaled Ouafi, Raphael C.-W. Phan |
ISPEC | 2 |
| 2008 | Proxy Re-signatures in the Standard Model
Sherman S. M. Chow, Raphael C.-W. Phan |
ISC | 2 |
| 2008 | Cryptanalysis of simple three-party key exchange protocol (S-3PAKE)
Raphael C.-W. Phan, Wei-Chuen Yau, Bok-Min Goi |
Inf. Sci. | 1 |
| 2008 | Tampering with a watermarking-based image authentication scheme
Raphael C.-W. Phan |
Pattern Recognit. | 1 |
| 2007 | Cryptanalysis of Two Non-anonymous Buyer-Seller Watermarking Protocols for Content Protection
Bok-Min Goi, Raphael C.-W. Phan, Hean-Teik Chuah |
ICCSA (1) | 2 |
| 2007 | (In)Security of an Efficient Fingerprinting Scheme with Symmetric and Commutative Encryption of IWDW 2005
Raphael C.-W. Phan, Bok-Min Goi |
IWDW | 1 |
| 2007 | Security of a Leakage-Resilient Protocol for Key Establishment and Mutual Authentication
Raphael C.-W. Phan, Kim-Kwang Raymond Choo, Swee-Huay Heng |
ProvSec | 1 |
| 2007 | On the Notions of PRP - RKA , KR and KR - RKA for Block Ciphers
Ermaliza Razali, Raphael C.-W. Phan, Marc Joye |
ProvSec | 2 |
| 2007 | On the Analysis and Design of a Family Tree of Smart Card Based User Authentication Schemes
Raphael C.-W. Phan, Bok-Min Goi |
UIC | 1 |
| 2006 | Cryptanalysis of the N-Party Encrypted Diffie-Hellman Key Exchange Using Different Passwords
Raphael C.-W. Phan, Bok-Min Goi |
ACNS | 1 |
| 2006 | Amplifying Side-Channel Attacks with Techniques from Block Cipher Cryptanalysis
Raphael C.-W. Phan, Sung-Ming Yen |
CARDIS | 1 |
| 2006 | Cryptanalysis of some improved password-authenticated key exchange schemes
Raphael C.-W. Phan, Bok-Min Goi, Kah-Hoong Wong |
Comput. Commun. | 1 |
| 2006 | Cryptanalysis of two password-based authentication schemes using smart cards
Raphael C.-W. Phan |
Comput. Secur. | 1 |
| 2006 | Security considerations for incremental hash functions based on pair block chaining
Raphael C.-W. Phan, David A. Wagner 0001 |
Comput. Secur. | 1 |
| 2006 | A Framework for Describing Block Cipher CryptanalysisabstractBlock ciphers provide confidentiality by encrypting confidential messages into unintelligible form, which are irreversible without knowledge of the secret key used. During the design of a block cipher, its security against cryptanalysis must be considered. History has shown that a cipher designed without an adequate treatment of this will often lead to flaws and attacks by other researchers, sometimes devastatingly so. The problem for an aspiring cipher designer is that there are no standard texts on block cipher cryptanalysis because it is a fast changing field. The commonly available references are academic journals and conference proceedings, which may not be easy to grasp for researchers new to cryptanalysis. This paper presents the Xi framework, which is designed to compactly describe the block cipher cryptanalysis techniques regardless of their individual differences. This provides the cryptanalyst with a general framework to describe attacks on block ciphers, with the additional capabilities of allowing specification of the technical details of each different type of attack and of comparison of their respective strengths. Comparing different distinguishers in this framework also allows us to see natural generalizations and trigger nice open problems. We then show how to apply this Xi framework to the description of various attacks on popular and recent block ciphers. Raphael C.-W. Phan, Mohammad Umar Siddiqi |
IEEE Trans. Computers | 1 |
| 2005 | Cryptanalysis of an Improved Client-to-Client Password-Authenticated Key Exchange (C2C-PAKE) Scheme
Raphael C.-W. Phan, Bok-Min Goi |
ACNS | 1 |
| 2005 | On the Rila-Mitchell Security Protocols for Biometrics-Based Cardholder Authentication in Smartcards
Raphael C.-W. Phan, Bok-Min Goi |
ICCSA (1) | 1 |
| 2005 | On the Rila-Mitchell Security Protocols for Biometrics-Based Cardholder Authentication in Smartcards
Raphael C.-W. Phan, Bok-Min Goi |
ICCSA (4) | 1 |
| 2005 | Related-Mode Attacks on Block Cipher Modes of Operation
Raphael C.-W. Phan, Mohammad Umar Siddiqi |
ICCSA (3) | 1 |
| 2005 | On the Security Bounds of CMC, EME, EME+ and EME* Modes of Operation
Raphael C.-W. Phan, Bok-Min Goi |
ICICS | 1 |
| 2005 | On the Security of the WinRAR Encryption Method
Gary S.-W. Yeo, Raphael C.-W. Phan |
ISC | 2 |
| 2004 | Cryptanalysis of Two Anonymous Buyer-Seller Watermarking Protocols and an Improvement for True Anonymity
Bok-Min Goi, Raphael C.-W. Phan, Yanjiang Yang, Feng Bao 0001, Robert H. Deng, Mohammad Umar Siddiqi |
ACNS | 2 |
| 2004 | Related-Key Attacks on Triple-DES and DESX Variants
Raphael C.-W. Phan |
CT-RSA | 1 |
| 2004 | On Related-Key and Collision Attacks: The Case for the IBM 4758 Cryptoprocessor
Raphael C.-W. Phan, Helena Handschuh |
ISC | 1 |
| 2004 | Flaws in Generic Watermarking Protocols Based on Zero-Knowledge Proofs
Raphael C.-W. Phan, Huo-Chong Ling |
IWDW | 1 |
| 2004 | Impossible differential cryptanalysis of 7-round Advanced Encryption Standard (AES)
Raphael C.-W. Phan |
Inf. Process. Lett. | 1 |