VLDB 2026 Research / reviewers in the wild / expert
Shu-Min Leong
dblp:278/2022
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0003-1041-4170ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Securing Face ID: Privacy Preservation for Non-Retentive Face Recognition SystemabstractFacial recognition technology is increasingly integrated into various applications. While face recognition systems have streamlined the authentication process and reduced the need for manual verification, the advancement in artificial intelligence (AI) generating realistic-looking media raises significant privacy concerns due to the potential misuse of biometric data. While current security protocols ensure data protection, storing biometric data in the system is a latent risk. In cases where the stored data is compromised, the users are susceptible to attacks such as face swapping, deepfake and identity theft. To address this, this paper presents a privacypreserving algorithm that omits the need to store human raw biometric data in the system. This is achieved by utilizing a new combination of Locality-Sensitive Hashing (LSH), salting, and RSA encryption for face recognition. The proposed method ensures data security by securely hashing and encrypting facial features while maintaining high recognition accuracy. The proposed framework is evaluated on the Labeled Faces in the Wild (LFW) and achieves a comparable performance with the state-of-the-art techniques. Megan Chua, Chanelle Yue-Ting Yeow, Cassandra Xin-Yee Chwee, Shu-Min Leong, Raphael C.-W. Phan |
TENCON | 4 |
| 2025 | BitRelation: Exploring Bit-Level Dependencies in Neural CryptanalysisabstractThis paper applies Explainable Artificial Intelligence (XAI) to improve the interpretability of neural differential cryptanalysis on the SPECK cipher. We use Local Interpretable Model-agnostic Explanations (LIME) to analyse and visualise feature importance in neural distinguishers, giving signed contributions and absolute rankings. Signed contributions show whether, and how strongly, specific bit positions influence the model's decision, while absolute rankings reflect their importance regardless of sign. To study interactions beyond single bits, we introduce a Systematic Masking Approach to reveal relations among bits by testing if chosen combinations of masked bits alter classification accuracy. On Gohr's 8-round SPECK32/64 distinguisher, masking up to four-bit combinations shows that decisions involve multi-bit interactions rather than isolated single-bit effects. Although LIME highlights strong single-bit signals, masking reveals interaction patterns consistent with differential cryptanalysis. These findings clarify model behaviour in neural cryptanalysis and show XAI's value for exposing and visualising interaction structure in ciphertext features and decisions. Yue-Tian Goi, Shu-Min Leong, Raphael C.-W. Phan, Ana Salagean, Shangqi Lai, Wei-Chuen Yau |
TENCON | 2 |
| 2025 | Attack-SH: Adversarial Attacks on Self-Healing Material Properties Prediction ModelabstractAI is now increasingly applied in diverse domains. The recent Nobel prizes for Physics and Chemistry awarded to computational scientists shows the significant impact that AI has on real-world scientific applications. Adversarial attacks pose a significant threat to the reliability of AI systems, particularly in high-stakes applications such as those in the materials sciences domain, which affect interactions with materials that exist in the real world. This paper examines the impact of two widely used adversarial attack approaches, notably the Fast Gradient Sign Method (FGSM) and Projected Gradient Descent (PGD), on a target AI model's performance for realworld materials. Experimental results demonstrate that both methods effectively degrade predictive accuracy, with PGD showing a more severe effect. Notably, high Structural Similarity Index (SSIM) scores across all perturbed samples suggest that the attacks introduce imperceptible changes, increasing their potential risk as such attacks are then undetectable. Further analysis using metrics such as Mean Squared Error (MSE), Adversarial MSE (A-MSE), and Relative Error Increase (REI) confirms substantial shifts in model output and reconstruction quality. These findings highlight critical vulnerabilities in current model architectures and emphasize the urgent need for more resilient defense strategies, such as adversarial training and input pre-processing. Min-Xuan Tan, Pei-Sze Tan, Raphael C.-W. Phan, Shu-Min Leong |
TENCON | 4 |
| 2024 | Causally Uncovering Bias in Video Micro-Expression RecognitionabstractDetecting microexpressions presents formidable challenges, primarily due to their fleeting nature and the limited diversity in existing datasets. Our studies find that these datasets exhibit a pronounced bias towards specific ethnicities and suffer from significant imbalances in terms of both class and gender representation among the samples. These disparities create fertile ground for various biases to permeate deep learning models, leading to skewed results and inadequate portrayal of specific demographic groups. Our research is driven by a compelling need to identify and rectify these biases within model architectures. To achieve this, we commence by constructing a causal graph that elucidates the intricate relationships between the model, input features, and training outcomes. This graphical representation forms the foundation for our analytical framework. Leveraging this causal framework, we conduct comprehensive case studies, employing counterfactuals as a diagnostic tool to unveil biases arising from dataset-induced class imbalances, gender inequalities, and variations in facial action units. Our final step involves a highly efficient counterfactual debiasing process, eliminating the necessity for additional data collection or model retraining. Our results showcase superior performance compared to state-of-the-art methods across the CASME II, SAMM, and SMIC datasets. Pei-Sze Tan, Sailaja Rajanala, Arghya Pal, Shu-Min Leong, Raphael C.-W. Phan, Huey Fang Ong |
ICASSP | 4 |
| 2024 | Unveiling the Black Box: Neural Cryptanalysis with XAIabstractAt CRYPTO'19, Gohr[1] presented ResNet-based neural distinguishers (ND) for the round-reduced SPECK32/64 cipher. However, due to the black-box use of such deep learning models, it is hard for humans to understand why these distinguishers work, impeding advancements in cryptanalytic knowledge. In this work, we aim to effectively adapt eXplainable Artificial Intelligence (XAI) techniques, notably Local Interpretable Model-Agnostic Explanations (LIME) and Shapley Additive Explanations (SHAP), to gain a detailed understanding of the important features useful in Gohr's neural distinguishers. Yue-Tian Goi, Shu-Min Leong, Raphael C.-W. Phan, Shangqi Lai, Ana Salagean |
SMC | 2 |
| 2024 | Distingusic: Distinguishing Synthesized Music from HumanabstractIn this paper we focus on a problem that is increasingly plaguing the music industry; to a large extent due to the proliferation of generative AI models that enable the generation of new realistic and indistinguishable content for diverse modalities: text, image, audio, video. We address this problem from the perspective of audio watermarking; to our best knowledge, this is the first-known watermarking based approach to solve the problem of distinguishing realistic songs synthesized from generative AI models from real songs sung by humans. In more detail, our approach specifically utilizes the SHA-256 hash function, Singular Value Decomposition (SVD) and Discrete Wavelet Transform (DWT) for robust audio watermarking of synthesized songs. Before embedding, the audio is subjected to an attack phase to pinpoint less vulnerable regions for QR watermark placement. During the embedding process, the audio chunks first undergo a 1-level Discrete Wavelet Transform (DWT), and then the resulting approximate coefficients go through Singular Value Decompo-sition (SVD). Additionally, the watermarked array is subjected to SHA-256 hashing for collision-resistant conciseness, which is subsequently embedded into the singular values of the audio. Experimental findings demonstrate the superiority of our method over existing audio watermarking approaches under various signal attack scenarios. Zi Qian Yong, Shu-Min Leong, Sailaja Rajanala, Arghya Pal, Raphael C.-W. Phan |
SMC | 2 |
| 2024 | Emotion-specific AUs for micro-expression recognitionabstractAbstract The Facial Action Coding System (FACS) comprehensively describes facial expressions with facial action units (AUs). It is a well-used technique by researchers in emotions research to understand human emotions better. Most micro-expression datasets provide FACS-coded AU ground truths corresponding to micro-expressions classes. It is commonly accepted in computer vision-based emotions research that certain emotions are reliably revealed when specific combinations of AUs occur. However, the reliability of the ground truth AUs in the micro-expression datasets is lower than that of normal expressions, as they have lower AU intensities. Moreover, these micro-expression datasets only report the overall reliability of all AUs. It could not be identified which AUs had been accurately coded. This work aims to revisit the ground truth AUs of popular micro-expression datasets, namely CASME II, SAMM and CAS(ME) $$^2$$ 2 , and inspect whether any AUs crucial for micro-expression recognition may need to be reconsidered. This paper also provides a detailed AU analysis which yields new AU-based RoIs for each dataset. These new RoIs improve the micro-expression recognition performances compared to the baselines considered in this work. The proposed RoIs for CASME II, SAMM and CAS(ME) $$^2$$ 2 improve the recognition rates by $$2\%$$ 2 % , $$1\%$$ 1 % and $$4\%$$ 4 % , respectively, when compared with the existing RoIs. Shu-Min Leong, Raphael C.-W. Phan, Vishnu Monn Baskaran |
Multim. Tools Appl. | 1 |
| 2022 | GraphEx: Facial Action Unit Graph for Micro-Expression ClassificationabstractFacial micro-expressions are crucial cues for expressing human emotions. Existing works have shown substantial progress in detecting micro-expressions for various applications in the computer vision field. However, it is still onerous for existing methods to handle and interpret micro-expressions efficiently. This paper proposes a deep learning-based approach leveraging spatio-temporal and graph representation learning for micro-expression classification. We design a novel Spatial-Temporal Info Extraction Network (STIENet) for learning facial appearance and muscle motion from high dimensional video clip frames and summarizes them into more meaningful feature maps. We construct an action unit (AU) relation graph to further represent the AU co-occurrence in the same micro-expression video clip. A graph neural network (GNN) is used to learn AU-related graph embedding for the downstream classification task. Performance evaluation on two mainstream micro-expression datasets, i.e., CASME II and SAMM, show that the proposed framework outperforms other state-of-the-art methods for micro-expression classification. Shu-Min Leong, Fuad Noman, Raphael C.-W. Phan, Vishnu Monn Baskaran, Chee-Ming Ting |
ICIP | 1 |
| 2021 | Faceless identification based on temporal strips
Shu-Min Leong, Raphael C.-W. Phan, Vishnu Monn Baskaran, Chee-Pun Ooi |
Multim. Tools Appl. | 1 |