EDBT 2026 Demo / reviewers in the wild / expert
Koksheik Wong
dblp:19/6841 · also KokSheik Wong
· DBLP profile ↗
90ranked-venue papers
5as first author
37since 2021 · last 2026
0000-0002-4893-2291ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 63 · 4 first-author · 20 since 2021Artificial intelligence and machine learning · 8 · 7 since 2021Security and privacy · 8 · 1 first-author · 5 since 2021Computer networks · 5 · 3 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reversible data hiding in PDF files by overlapping charactersabstractMost reported PDF-based data hiding methods make subtle changes to existing PDF elements to conceal data, where the changes are predominantly made to the existing ‘TJ’ operator in a typical PDF document. However, this choice limits the volume of data that can be hidden, and most methods inevitably introduce visual distortion to the PDF document. Therefore, in this work, we propose a novel method called OPAC to hide data in PDF by means of o verla p ping char ac ters , which is the first of its kind for data hiding purposes. Specifically, overlapping employs a series of technical steps to strategically stack characters. First, all 26 characters in the English alphabet are arranged into a 4 × 7 table based on their frequency of occurrences, where each character can be referred to by specifying the column and row in which it appears. Subsequently, the encrypted message is hidden, one character at a time, by superimposing a special symbol called anchor onto a specific character to indicate the column number, and specific combination of the ‘Tf’ and ‘Tz’ operators is added to the PDF stream to indicate the row number. Meanwhile, the newly added ‘Tf’ and ‘Tz’ operators are further manipulated to reduce the size and width of anchor to allow for complete character overlapping. We evaluate the performance of the proposed method OPAC using 220 PDF documents generated from various famous texts. On average, OPAC can hide 3877 characters at the expense of 0.224 MB increase in PDF file size. Furthermore, experiment results confirmed that OPAC achieves distortion-free data hiding and reversibility. Moreover, OPAC is empirically verified to be resistant to several common PDF processing and alterations. Mohamed Mazeen Mujthaba, Koksheik Wong |
J. Inf. Secur. Appl. | 3 |
| 2025 | Polyfit generative model: can a group of lower-order polynomials generate high resolution diverse images?abstractImplicit neural representations (INRs) have recently gained popularity as a means to model images as continuous functions of spatial coordinates, synthesizing each pixel independently and yielding impressive results in tasks such as scene reconstruction and image generation. A notable advancement, PolyINR, utilizes element-wise multiplications between features and affine-transformed coordinates to achieve higher-order polynomial functions, eliminating the need for positional encodings. However, the finite encoding capacity of INRs, coupled with PolyINR’s recursive polynomial estimation, necessitates substantial training parameters, resulting in high computational costs and limiting applicability across diverse computer vision domains. In this work, we address these challenges by representing images as grids of smaller patches, within which we fit low-degree polynomials to capture local intensity variations. Our approach substantially reduces parameter requirements and computational demands. We evaluate our model qualitatively and quantitatively on large-scale datasets, ImageNet, CelebA, LSUN Bedroom, and Flower102; thus demonstrating competitive performance with state-of-the-art generative models, despite the absence of convolutional, normalization, or self-attention layers. Arghya Pal, Ai-Fang Chai, Sailaja Rajanala, Raphael C.-W. Phan, Koksheik Wong, Chee-Ming Ting |
IJCNN | 5 |
| 2025 | Conceptualizing Multi-scale Wavelet Attention and Ray-based Encoding for Human-Object Interaction DetectionabstractHuman-object interaction (HOI) detection is essential for accurately localizing and characterizing interactions between humans and objects, providing a comprehensive understanding of complex visual scenes across various domains. However, existing HOI detectors often struggle to deliver reliable predictions efficiently, relying on resource-intensive training methods and inefficient architectures. To address these challenges, we conceptualize a wavelet attention-like backbone and a novel ray-based encoder architecture tailored for HOI detection. Our wavelet backbone addresses the limitations of expressing middle-order interactions by aggregating discriminative features from the low- and high-order interactions extracted from diverse convolutional filters. Concurrently, the ray-based encoder facilitates multi-scale attention by optimizing the focus of the decoder on relevant regions of interest and mitigating computational overhead. As a result of harnessing the attenuated intensity of learnable ray origins, our decoder aligns query embeddings with emphasized regions of interest for accurate predictions. Experimental results on benchmark datasets, including ImageNet and HICO-DET, showcase the potential of our proposed architecture. The code is publicly available at [https://github.com/henrypay/RayEncoder]. Quan Bi Pay, Vishnu Monn Baskaran, Junn Yong Loo, Koksheik Wong, Simon See |
IJCNN | 4 |
| 2025 | SpaRTAN: Spatial Reinforcement Token-based Aggregation Network for Visual RecognitionabstractThe resurgence of convolutional neural networks (CNNs) in visual recognition tasks, exemplified by ConvNeXt, has demonstrated their capability to rival transformer-based architectures through advanced training methodologies and ViTinspired design principles. However, both CNNs and transformers exhibit a simplicity bias, favoring straightforward features over complex structural representations. Furthermore, modern CNNs often integrate MLP-like blocks akin to those in transformers, but these blocks suffer from significant information redundancies, necessitating high expansion ratios to sustain competitive performance. To address these limitations, we propose SpaRTAN, a lightweight architectural design that enhances spatial and channel-wise information processing. SpaRTAN employs kernels with varying receptive fields, controlled by kernel size and dilation factor, to capture discriminative multi-order spatial features effectively. A wave-based channel aggregation module further modulates and reinforces pixel interactions, mitigating channel-wise redundancies. Combining the two modules, the proposed network can efficiently gather and dynamically contextualize discriminative features. Experimental results in ImageNet and COCO demonstrate that SpaRTAN achieves remarkable parameter efficiency while maintaining competitive performance. In particular, on the ImageNet-1k benchmark, SpaRTAN achieves 77. 7% accuracy with only 3.8M parameters and approximately 1.0 GFLOPs, demonstrating its ability to deliver strong performance through an efficient design. On the COCO benchmark, it achieves 50.0% AP, surpassing the previous benchmark by 1.2% with only 21.5M parameters. The code is publicly available at [https://github.com/henry-pay/SpaRTAN]. Quan Bi Pay, Vishnu Monn Baskaran, Junn Yong Loo, Koksheik Wong, Simon See |
IJCNN | 4 |
| 2025 | GTA-HDR: A Large-Scale Synthetic Dataset for HDR Image ReconstructionabstractHigh Dynamic Range (HDR) content (i.e., images and videos) has a broad range of applications. However, capturing HDR content from real-world scenes is expensive and time-consuming. Therefore, the challenging task of reconstructing visually accurate HDR images from their Low Dynamic Range (LDR) counterparts is gaining attention in the vision research community. A major challenge is the lack of datasets, which capture diverse scene conditions (e.g., lighting, weather, locations) and various image features (e.g., color, contrast, saturation). To address this gap, we introduce GTA-HDR, a large-scale synthetic dataset of photo-realistic HDR images sampled from the GTA-V video game. We perform thorough evaluation of the proposed dataset, which enables significant qualitative and quantitative improvements of the state-of-the-art HDR image reconstruction methods. Furthermore, we demonstrate the effectiveness of the proposed dataset and its impact on additional computer vision tasks including 3D human pose estimation, human body part segmentation, and holistic scene segmentation. The dataset, data collection pipeline, and evaluation code are available at: https://github.com/HrishavBakulBarua/GTA-HDR. Hrishav Bakul Barua, Kalin Stefanov, Koksheik Wong, Abhinav Dhall, Ganesh Krishnasamy |
WACV | 3 |
| 2025 | Amogel: a multi-omics classification framework using associative graph neural networks with prior knowledge for biomarker identificationabstractThe advent of high-throughput sequencing technologies, such as DNA microarray and DNA sequencing, has enabled effective analysis of cancer subtypes and targeted treatment. Furthermore, numerous studies have highlighted the capability of graph neural networks (GNN) to model complex biological systems and capture non-linear interactions in high-throughput data. GNN has proven to be useful in leveraging multiple types of omics data, including prior biological knowledge from various sources, such as transcriptomics, genomics, proteomics, and metabolomics, to improve cancer classification. However, current works do not fully utilize the non-linear learning potential of GNN and lack of the integration ability to analyse high-throughput multi-omics data simultaneously with prior biological knowledge. Nevertheless, relying on limited prior knowledge in generating gene graphs might lead to less accurate classification due to undiscovered significant gene-gene interactions, which may require expert intervention and can be time-consuming. Hence, this study proposes a graph classification model called associative multi-omics graph embedding learning (AMOGEL) to effectively integrate multi-omics datasets and prior knowledge through GNN coupled with association rule mining (ARM). AMOGEL employs an early fusion technique using ARM to mine intra-omics and inter-omics relationships, forming a multi-omics synthetic information graph before the model training. Moreover, AMOGEL introduces multi-dimensional edges, with multi-omics gene associations or edges as the main contributors and prior knowledge edges as auxiliary contributors. Additionally, it uses a gene ranking technique based on attention scores, considering the relationships between neighbouring genes. Several experiments were performed on BRCA and KIPAN cancer subtypes to demonstrate the integration of multi-omics datasets (miRNA, mRNA, and DNA methylation) with prior biological knowledge of protein-protein interactions, KEGG pathways and Gene Ontology. The experimental results showed that the AMOGEL outperformed the current state-of-the-art models in terms of classification accuracy, F1 score and AUC score. The findings of this study represent a crucial step forward in advancing the effective integration of multi-omics data and prior knowledge to improve cancer subtype classification. Chia Yan Tan, Huey Fang Ong, Chern Hong Lim, Mei Sze Tan, Ean Hin Ooi, Koksheik Wong |
BMC Bioinform. | 6 |
| 2025 | Trade-off independent image watermarking using enhanced structured matrix decompositionabstractAbstract Image watermarking plays a vital role in providing protection from copyright violation. However, conventional watermarking techniques typically exhibit trade-offs in terms of image quality, robustness and capacity constrains. More often than not, these techniques optimize on one constrain while settling with the two other constraints. Therefore, in this paper, an enhanced saliency detection based watermarking method is proposed to simultaneously improve quality, capacity, and robustness. First, the enhanced structured matrix decomposition (E-SMD) is proposed to extract salient regions in the host image for producing a saliency mask. This mask is then applied to partition the foreground and background of the host and watermark images. Subsequently, the watermark (with the same dimension of host image) is shuffled using multiple Arnold and Logistic chaotic maps, and the resulting shuffled-watermark is embedded into the wavelet domain of the host image. Furthermore, a filtering operation is put forward to estimate the original host image so that the proposed watermarking method can also operate in blind mode. In the best case scenario, we could embed a 24-bit image as the watermark into another 24-bit image while maintaining an average SSIM of 0.9999 and achieving high robustness against commonly applied watermark attacks. Furthermore, as per our best knowledge, with high payload embedding, the significant improvement in these features (in terms of saliency, PSNR, SSIM, and NC) has not been achieved by the state-of-the-art methods. Thus, the outcomes of this research realizes a trade-off independent image watermarking method, which is a first of its kind in this domain. Koksheik Wong, Vishnu Monn Baskaran |
Multim. Tools Appl. | 2 |
| 2025 | High-fidelity reversible data hiding using novel comprehensive rhombus predictor
Rajeev Kumar 0007, Roberto Caldelli, Koksheik Wong, Aruna Malik, Ki-Hyun Jung |
Multim. Tools Appl. | 3 |
| 2024 | Reimagining Violent Action Detection with Human-Object InteractionabstractThe rising urban crime rates globally underscore the need for advanced video surveillance systems capable of autonomously detecting violent actions. Current deep learning models face limitations, struggling with subtle motions and lacking real-time capabilities. In response, we advocate for a paradigm shift in surveillance oriented violent action detection, emphasizing the pivotal role of human-object interaction (HOI) detection as opposed to conventional action recognition methodologies. Our contributions include unveiling Violence-HOI (V-HOI), a dataset capturing HOI interactions in static surveillance images. Additionally, we introduce Violence-Net (V Net), a novel convolutional-transformer network architecture, which outperforms existing HOI approaches by 5.25 percentage points in mean average precision. Moreover, when trained on V-HOI, V-Net achieves near real-time processing at 10.43 frames per second, demonstrating its practicality in dynamic surveillance scenarios. The code and dataset is available at https://github.com/MarcusLimJunYi/vhoi. Vishnu Monn Baskaran, Ricky Sutopo, JunYi Lim, Joanne Mun-Yee Lim, Koksheik Wong |
AVSS | 5 |
| 2024 | Histohdr-Net: Histogram Equalization for Single LDR to HDR Image TranslationabstractHigh Dynamic Range (HDR) imaging aims to replicate the high visual quality and clarity of real-world scenes. Due to the high costs associated with HDR imaging, the literature offers various data-driven methods for HDR image reconstruction from Low Dynamic Range (LDR) counterparts. A common limitation of these approaches is missing details in regions of the reconstructed HDR images, which are overor under-exposed in the input LDR images. To this end, we propose a simple and effective method, HistoHDR-Net, to recover the fine details (e.g., color, contrast, saturation, and brightness) of HDR images via a fusion-based approach utilizing histogram-equalized LDR images along with self-attention guidance. Our experiments demonstrate the efficacy of the proposed approach over the state-of-art methods. Hrishav Bakul Barua, Ganesh Krishnasamy, Koksheik Wong, Abhinav Dhall, Kalin Stefanov |
ICIP | 3 |
| 2024 | Music Form Analysis: A Case Study of The Theme and Variations FormabstractThe theme and variation music form is a hierarchical structure in music. It has a theme segment at the beginning, followed by a series of variation segments imitating the theme segment. Hence, the primary features of the theme and variation form are repetition and variation at different levels. However, due to the lack of available datasets, the theme and variation form analysis method has not been explored much. Therefore, in this work, we curate and contribute a dataset named Performance of Theme and Variation Form (PTV) and propose a theme and variation form segmentation framework to analyze the theme and variation form. Experiment results show that our method achieves an F1 score of 92.8% on our dataset with some constraints. In addition, we conduct analysis to support and encourage future studies of the theme and variation form. Jing Zhao 0033, Koksheik Wong, Vishnu Monn Baskaran, Kiki Maulana, David Taniar |
ICME | 2 |
| 2024 | A message verification scheme based on physical layer-enabled data hiding for flying ad hoc networkabstractAbstract The use of Unmanned Aerial Vehicles (UAVs) for military as well as civilian applications such as search and rescue operations, disaster management, parcel delivery and agriculture, is becoming increasingly popular, mainly due to their versatile nature and relatively inexpensive operating costs. However, one of the most significant barriers in the widespread adoption of UAVs for the aforementioned applications is the increasing concern for wireless communication security within the network. Unfortunately, the nodes that comprise these networks are extremely constrained in terms of storage space, processing power and battery life, which require light-weight security protocols that align with these limitations. The current standards for message authentication and verification predominantly rely on lightweight cryptographic techniques. However, there is continual scope for improvement, particularly in the realm of computational efficiency, to better align with the constraints inherent to UAVs. This paper explores data hiding in a multi-UAV environment, to form a light-weight message verification scheme to confirm the identity of the transmitting node. First, five venues having the potential to hide data are analyzed. Based on the analysis of bit error rate obtained from the simulated multi-UAV environment, both cyclic prefix and padding bits are selected. Next, a codeword is generated for each transmission and it is embedded along with the unique ID key, $$\kappa $$ κ of the transmitting node into each venue respectively, and the output is transmitted as Hello packets. Subsequently, the receiving node extracts and decodes the codeword in order to verify the authenticity of the transmitting node. An evaluation of the proposed scheme has been conducted by using a simulated UAV environment. In the best case scenario, the proposed authentication scheme can achieve a successful verification rate of at least 80% at a maximum hop-count of 10 hops. Through computational cost analysis, in comparison to the conventional methods, the proposed scheme was found to have a significantly lower total execution time of $$1.7\mu s$$ 1.7 μ s . In addition, the security analysis shows that when the proposed scheme is paired with physical unclonable functions (PUFs), it is able to resist common security attacks, including man-in-the-middle, replay and cloning attacks. Dilshani Mallikarachchi, Koksheik Wong, Joanne Mun-Yee Lim |
Multim. Tools Appl. | 2 |
| 2023 | $\mathrm{C}\eta\iota \text{DAE}$: Cryptographically Distinguishing Autoencoder for Cipher CryptanalysisabstractWe propose a new autoencoder (AE) construction$\mathrm{C}\eta\iota \text{DAE}$(Cryptographically Distinguishing AE) based on a novel loss formulation to solve the cipher cryptanalysis distinguishing problem in the domain of cryptology. Vanilla AE and variational AE are unable to address this problem as they are designed to draw new samples which are either similar to the input sample or are from the same distribution. Such generated samples do not facilitate the cryptanalysis task. We show that our AE construction enables the discovery of cipher distinguishers, which are the fundamental building blocks that make or break new cipher design proposals. This also answers an open question on the applicability of autoencoders for cipher cryptanalysis; as to date, only discriminative models have been applied for cryptanalysis problems. To the best of our knowledge,$\mathrm{C}\eta\iota \text{DAE}$is the first-known generative model designed to solve crypt-analysis problems. We apply our$\mathrm{C}\eta\iota \text{DAE}$model to discover distinguishing properties for up to 10 rounds of the NSA-designed Speck32/64 cipher that allows to distinguish it from a random permutation. This contrasts with the best-known machine learning-discovered neural distinguisher in the literature that covers up to 8 rounds of Speck32/64. Unlike these recent related work which leverage on white box analysis and human-guided differential or linear analysis in order for machine learning models to be applicable, our$\mathrm{C}\eta\iota \text{DAE}$distinguisher does not require prior human cryptanalytic knowledge. This motivates the new direction of human-unsupervised machine learning-based cryptanalysis techniques. Raphael C.-W. Phan, Arghya Pal, Koksheik Wong, Sailaja Rajanala |
GLOBECOM | 3 |
| 2023 | Self Supervised Bert for Legal Text ClassificationabstractCritical BERT-based text classification tasks, such as legal text classification, require huge amounts of accurately labeled data. Legal text classification faces two trivial problems: labeling legal data is a sensitive process and can only be carried out by skilled professionals, and legal text is prone to privacy issues hence not all the data can be made available in the public domain. This means that we have limited diversity in the textual data, and to account for this data paucity, we propose a self-supervision approach to train Legal-BERT classifiers. We use the BERT text classifier’s knowledge of the class boundaries and perform gradient ascent w.r.t. class logits. Synthetic latent texts are generated through activation maximization. The main advantages over existing SOTAs are that our model: is easy to train, does not require much data but instead uses the synthesized data as fake samples; has less variance that helps to generate texts with good sample quality and diversity. We show the efficacy of the proposed method on the ECHR Violation (Multi-Label) Dataset and the Over-ruling Task Dataset. Arghya Pal, Sailaja Rajanala, Raphael C.-W. Phan, Koksheik Wong |
ICASSP | 4 |
| 2023 | Computational Music: Analysis of Music Forms
Jing Zhao 0033, Koksheik Wong, Vishnu Monn Baskaran, Kiki Maulana, David Taniar |
ICCSA (1) | 2 |
| 2023 | ScratchHOI: Training Human-Object Interaction Detectors from ScratchabstractTransformer-based approaches have exhibited outstanding performances in the field of human-object interaction (HOI) detection. However, these approaches rely on underlying object detectors that have undergone large-scale pre-trainings on the ImageNet and MS-COCO dataset. This limits the potential of unique architectural designs and induces a learning bias, causing ineffective HOI representation learning. In this paper, we propose ScratchHOI, a transformer-based method for human-object interaction detection that can be trained from scratch, eliminating the need for pre-trained object detectors. ScratchHOI employs dynamic and static affinity-based feature aggregation for processing local and long-range visual information. Additional techniques are also employed to improve detection performance, such as dynamic and interactive anchor refinement for objects and interactions. Experiments on the HICO-Det dataset show that ScratchHOI achieves competitive performance against other state-of-the-art approaches over a variety of different evaluation measures. JunYi Lim, Vishnu Monn Baskaran, Joanne Mun-Yee Lim, Ricky Sutopo, Koksheik Wong, Massimo Tistarelli |
ICIP | 5 |
| 2023 | An authentication scheme for FANET packet payload using data hidingabstractFlying ad hoc networks (FANETs) are gaining much traction in recent years due to their wide range of applications, resulting in the emergence of a number of studies on the facilitation of their widespread adoption. One of the key areas being investigated is the establishment of security strategies to combat the large amount and variety of security threats that they inevitably face when they are operating in the field. Since a majority of FANET’s operations involve the transmission of network packets containing both instructions and sensitive data, the protection of these packets is critical. This paper proposes an authentication scheme to protect the payload of the packets being transmitted within the FANET. Since FANETs are highly resource-constrained, we opted for a lightweight and energy-efficient approach. The proposed scheme essentially encodes the payload of a physical layer data frame into a compact bit stream, which is subsequently embedded into the cyclic prefix (CP) at the physical layer. The encoding process includes the XOR-operation and JBIG2 lossless compression. At the receiver’s end, the hidden authentication data is extracted and decoded for comparison against the payload of the received data frame. Two enhancements are introduced to further enhance robustness against masquerade attack. The proposed scheme is evaluated in terms of bit error rate (BER) under different signal-to-noise ratio (SNR) settings. In all considered SNR settings, the proposed authentication scheme achieves BERs of <0.7∗10−4 when 7-bit Hamming code is implemented. Experiment results also confirm that proposed scheme is able to localize the tampered data frames. Dilshani Mallikarachchi, Koksheik Wong, Joanne Mun-Yee Lim |
J. Inf. Secur. Appl. | 2 |
| 2023 | High payload watermarking based on enhanced image saliency detectionabstractAbstract Nowadays, images are circulated rapidly over the internet and they are subject to some risk of misuses. To address this issue, various watermarking methods are proposed in the literature. However, most conventional methods achieve a certain trade-off among imperceptibility and high capacity payload, and they are not able to improve these criteria simultaneously. Therefore, in this paper, a robust saliency-based image watermarking method is proposed to achieve high payload and high quality watermarked image. First, an enhanced salient object model is proposed to produce a saliency map, followed by a binary mask to segments the foreground/background region of a host image. The same mask is then consulted to decompose the watermark image. Next, the RGB channels of the watermark are encrypted by using Arnold, 3-DES and multi-flipping permutation encoding (MFPE). Furthermore, the principal key used for encryption is embedded in the singular matrix of the blue channel. Moreover, the blue channel is encrypted by using the Okamoto-Uchiyama homomorphic encryption (OUHE) method. Finally, these encrypted watermark channels are diffused and embedded into the host channels. When the need arises, more watermarks can be embedded into the host at the expense of the quality of the embedded watermarks. Our method can embed watermark of the same dimension as the host image, which is the first of its kind. Experimental results suggest that the proposed method maintains robustness while achieving high image quality and high payload. It also outperforms the state-of-the-art (SOTA) methods. Koksheik Wong |
Multim. Tools Appl. | 2 |
| 2023 | Multi-mmlg: a novel framework of extracting multiple main melodies from MIDI filesabstractAbstract As an essential part of music, main melody is the cornerstone of music information retrieval. In the MIR’s sub-field of main melody extraction, the mainstream methods assume that the main melody is unique. However, the assumption cannot be established, especially for music with multiple main melodies such as symphony or music with many harmonies. Hence, the conventional methods ignore some main melodies in the music. To solve this problem, we propose a deep learning-based Multiple Main Melodies Generator (Multi-MMLG) framework that can automatically predict potential main melodies from a MIDI file. This framework consists of two stages: (1) main melody classification using a proposed MIDIXLNet model and (2) conditional prediction using a modified MuseBERT model. Experiment results suggest that the proposed MIDIXLNet model increases the accuracy of main melody classification from 89.62 to 97.37%. In addition, this model requires fewer parameters (71.8 million) than the previous state-of-art approaches. We also conduct ablation experiments on the Multi-MMLG framework. In the best-case scenario, predicting meaningful multiple main melodies for the music are achieved. Jing Zhao 0033, David Taniar, Kiki Maulana, Vishnu Monn Baskaran, Koksheik Wong |
Neural Comput. Appl. | 5 |
| 2023 | Attacking Mouse Dynamics Authentication Using Novel Wasserstein Conditional DCGANabstractBehavioral biometrics is an emerging trend due to their cost-effectiveness and non-intrusive implementations that support remote access for user identification. This is the case especially in recent times of social distancing and working from home arrangements, where online attendance is the preferred option in contrast to physical presence. In this work, we explore the limitations of mouse dynamics authentication by impersonating legitimate user mouse action sequences. Specifically, towards that aim, we develop a novel generative WC-DCGAN model to generate highly accurate fake user action sequences. We apply our WC-DCGAN to this problem and show that it causes the target classifier can be tricked into identifying a fraudster as a legitimate user. WC-DCGAN has several benefits, including: achieving dominated convergence, hence implying the existence of solutions and optimal discriminator regardless of data and generator distributions; and acting as an unsupervised model for a fixed class label and generator. Experiments are conducted to verify these points. Subsequently, we analyzed the cause of misclassifications, and propose a novel mouse dynamics strategy that offers much tighter authentication with significant reductions in misclassification events. Arunava Roy, Koksheik Wong, Raphael C.-W. Phan |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | ERNet: An Efficient and Reliable Human-Object Interaction Detection NetworkabstractHuman-Object Interaction (HOI) detection recognizes how persons interact with objects, which is advantageous in autonomous systems such as self-driving vehicles and collaborative robots. However, current HOI detectors are often plagued by model inefficiency and unreliability when making a prediction, which consequently limits its potential for real-world scenarios. In this paper, we address these challenges by proposing ERNet, an end-to-end trainable convolutional-transformer network for HOI detection. The proposed model employs an efficient multi-scale deformable attention to effectively capture vital HOI features. We also put forward a novel detection attention module to adaptively generate semantically rich instance and interaction tokens. These tokens undergo pre-emptive detections to produce initial region and vector proposals that also serve as queries which enhances the feature refinement process in the transformer decoders. Several impactful enhancements are also applied to improve the HOI representation learning. Additionally, we utilize a predictive uncertainty estimation framework in the instance and interaction classification heads to quantify the uncertainty behind each prediction. By doing so, we can accurately and reliably predict HOIs even under challenging scenarios. Experiment results on the HICO-Det, V-COCO, and HOI-A datasets demonstrate that the proposed model achieves state-of-the-art performance in detection accuracy and training efficiency. Codes are publicly available at https://github.com/Monash-CyPhi-AI-Research-Lab/ernet. JunYi Lim, Vishnu Monn Baskaran, Joanne Mun-Yee Lim, Koksheik Wong, John See, Massimo Tistarelli |
IEEE Trans. Image Process. | 4 |
| 2023 | Is it Violin or Viola? Classifying the Instruments' Music Pieces using Descriptive StatisticsabstractClassifying music pieces based on their instrument sounds is pivotal for analysis and application purposes. Given its importance, techniques using machine learning have been proposed to classify violin and viola music pieces. The violin and viola are two different instruments with three overlapping strings of the same notes, and it is challenging for ordinary people or even musicians to distinguish the sound produced by these instruments. However, the classification of musical instrument pieces was barely performed by prior research. To solve this problem, we propose a technique using descriptive statistics to reliably distinguish between violin and viola music pieces. Likewise, a similar technique on the basis of histogram is introduced alongside the main descriptive statistics approach. These approaches are derived based on the nature of the instruments’ strings and the range of their pieces. We also solve the problem in the current literature which divide the audio into segments for processing instead of managing the whole song. Thereby, we compile a dataset of recordings that comprises of violin and viola solo pieces from the Baroque, Classical, Romantic, and Modern eras. Experiment results suggest that our approach achieves high accuracy on solo pieces as compared to other methods with 0.97 accuracy on Baroque pieces. Chong Hong Tan, Koksheik Wong, Vishnu Monn Baskaran, Kiki Maulana, David Taniar |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Guess-It-Generator: Generating in a Lewis Signaling Framework through Logical ReasoningabstractHuman minds spontaneously integrate two inherited cognitive capabilities: perception and reasoning to accomplish cognitive tasks such as problem solving, imagination, and causation. It is observed in the primate brains that perception offers the assistance required for problem comprehension, whilst the reasoning elucidates upon the facts recovered during perception in order to make a decision. The field of artificial intelligence (AI) thus considers perception and reasoning as two complementary areas that are realized by machine learning and logic programming, respectively. In this work, we propose a generative model using a collaborative guessing game of the kind first introduced by David Lewis in his famous work called the Lewis signaling game that is synonymous with the "20 Questions'' game. Our proposed model, Guess-It-Generator (GIG) is a collaborative framework that engages two recurrent neural networks in a guessing game. GIG unifies perception and reasoning with a view to generating labeled images by capturing, (X, y), the underlying density of a data distribution, i.e. (X, y) - p(X, y). An encoder attends to a region of the input image and encodes that onto a latent variable that acts as a perception signal to a decoder. In contrast, the decoder leverages on the perception signals to guess the image and verifies the guess by reasoning with logical facts derived from the domain knowledge. Our experiments and comprehensive studies on seven datasets: PCAM, Chest-Xray-14, FIRE, HAM10000 from the medical domain, and CIFAR 10, LSUN, ImageNet, among standard benchmark datasets, show significant promise for the proposed method. Arghya Pal, Sailaja Rajanala, Raphael C.-W. Phan, Koksheik Wong |
ACM Multimedia | 4 |
| 2022 | DeSCoVeR: Debiased Semantic Context Prior for Venue RecommendationabstractWe present a novel semantic context prior-based venue recommendation system that uses only the title and the abstract of a paper. Based on the intuition that the text in the title and abstract have both semantic and syntactic components, we demonstrate that a joint training of a semantic feature extractor and syntactic feature extractor collaboratively leverages meaningful information that helps to provide venues for papers. The proposed methodology that we call DeSCoVeR at first elicits these semantic and syntactic features using a Neural Topic Model and text classifier respectively. The model then executes a transfer learning optimization procedure to perform a contextual transfer between the feature distributions of the Neural Topic Model and the text classifier during the training phase. DeSCoVeR also mitigates the document-level label bias using a Causal back-door path criterion and a sentence-level keyword bias removal technique. Experiments on the DBLP dataset show that DeSCoVeR outperforms the state-of-the-art methods. Sailaja Rajanala, Arghya Pal, Manish Singh 0002, Raphael C.-W. Phan, Koksheik Wong |
SIGIR | 5 |
| 2022 | Covert communication in multi-hop UAV network
Dilshani Mallikarachchi, Koksheik Wong, Joanne Mun-Yee Lim |
Ad Hoc Networks | 2 |
| 2022 | Risk score prediction model based on single nucleotide polymorphism for predicting malaria: a machine learning approachabstractBACKGROUND: The malaria risk prediction is currently limited to using advanced statistical methods, such as time series and cluster analysis on epidemiological data. Nevertheless, machine learning models have been explored to study the complexity of malaria through blood smear images and environmental data. However, to the best of our knowledge, no study analyses the contribution of Single Nucleotide Polymorphisms (SNPs) to malaria using a machine learning model. More specifically, this study aims to quantify an individual's susceptibility to the development of malaria by using risk scores obtained from the cumulative effects of SNPs, known as weighted genetic risk scores (wGRS). RESULTS: We proposed an SNP-based feature extraction algorithm that incorporates the susceptibility information of an individual to malaria to generate the feature set. However, it can become computationally expensive for a machine learning model to learn from many SNPs. Therefore, we reduced the feature set by employing the Logistic Regression and Recursive Feature Elimination (LR-RFE) method to select SNPs that improve the efficacy of our model. Next, we calculated the wGRS of the selected feature set, which is used as the model's target variables. Moreover, to compare the performance of the wGRS-only model, we calculated and evaluated the combination of wGRS with genotype frequency (wGRS + GF). Finally, Light Gradient Boosting Machine (LightGBM), eXtreme Gradient Boosting (XGBoost), and Ridge regression algorithms are utilized to establish the machine learning models for malaria risk prediction. CONCLUSIONS: Our proposed approach identified SNP rs334 as the most contributing feature with an importance score of 6.224 compared to the baseline, with an importance score of 1.1314. This is an important result as prior studies have proven that rs334 is a major genetic risk factor for malaria. The analysis and comparison of the three machine learning models demonstrated that LightGBM achieves the highest model performance with a Mean Absolute Error (MAE) score of 0.0373. Furthermore, based on wGRS + GF, all models performed significantly better than wGRS alone, in which LightGBM obtained the best performance (0.0033 MAE score). Kah Yee Tai, Jasbir Dhaliwal, Koksheik Wong |
BMC Bioinform. | 3 |
| 2022 | Invisible emotion magnification algorithm (IEMA) for real-time micro-expression recognition with graph-based features
Adamu Muhammad Buhari, Chee-Pun Ooi, Vishnu Monn Baskaran, Raphael C.-W. Phan, Koksheik Wong, Wooi-Haw Tan |
Multim. Tools Appl. | 5 |
| 2021 | Synthesize-It-Classifier: Learning a Generative Classifier Through Recurrent Self-AnalysisabstractWe show the generative capability of an image classifier network by synthesizing high-resolution, photo-realistic, and diverse images at scale. The overall methodology, called Synthesize-It-Classifier (STIC), does not require an explicit generator network to estimate the density of the data distribution and sample images from that, but instead uses the classifier’s knowledge of the boundary to perform gradient ascent w.r.t. class logits and then synthesizes images using the Gram Matrix Metropolis Adjusted Langevin Algorithm (GRMALA) by drawing on a blank canvas. During training, the classifier iteratively uses these synthesized images as fake samples and re-estimates the class boundary in a recurrent fashion to improve both the classification accuracy and quality of synthetic images. The STIC shows that mixing of the hard fake samples (i.e. those synthesized by the one-hot class conditioning), and the soft fake samples (which are synthesized as a convex combination of classes, i.e. a mixup of classes [36]) improves class interpolation. We demonstrate an Attentive-STIC network that shows iterative drawing of synthesized images on the ImageNet dataset that has thousands of classes. In addition, we introduce the synthesis using a class conditional score classifier (Score-STIC) instead of a normal image classifier and show improved results on several real world datasets, i.e. ImageNet, LSUN and CIFAR 10. Arghya Pal, Raphael C.-W. Phan, Koksheik Wong |
CVPR | 3 |
| 2021 | Baitradar: A Multi-Model Clickbait Detection Algorithm Using Deep LearningabstractFollowing the rising popularity of YouTube, there is an emerging problem on this platform called clickbait, which provokes users to click on videos using attractive titles and thumbnails. As a result, users ended up watching a video that does not have the content as publicized in the title. This issue is addressed in this study by proposing an algorithm called BaitRadar, which uses a deep learning technique where six inference models are jointly consulted to make the final classification decision. These models focus on different attributes of the video, including title, comments, thumbnail, tags, video statistics and audio transcript. The final classification is attained by computing the average of multiple models to provide a robust and accurate output even in situation where there is missing data. The proposed method is tested on 1,400 YouTube videos. On average, a test accuracy of 98% is achieved with an inference time of ≤ 2s. Bhanuka Gamage, Adnan Labib, Aisha Joomun, Chern Hong Lim, Koksheik Wong |
ICASSP | 5 |
| 2021 | Deep multi-level feature pyramids: Application for non-canonical firearm detection in video surveillance
JunYi Lim, Md Istiaque Al Jobayer, Vishnu Monn Baskaran, Joanne Mun-Yee Lim, John See, Koksheik Wong |
Eng. Appl. Artif. Intell. | 6 |
| 2021 | StealthPDF: Data hiding method for PDF file with no visual degradation
Minoru Kuribayashi, Koksheik Wong |
J. Inf. Secur. Appl. | 2 |
| 2021 | A survey of reversible data hiding in encrypted images - The first 12 years
Pauline Puteaux, Simying Ong, Koksheik Wong, William Puech |
J. Vis. Commun. Image Represent. | 3 |
| 2021 | CUPSEED - A combined use of prediction syntax elements to embed data in SHVC video
LieLin Pang, Yiqi Tew, Koksheik Wong, Mohamad Nizam Ayub |
Multim. Tools Appl. | 3 |
| 2021 | Innovative lane detection method to increase the accuracy of lane departure warning system
Ting Yau Teo, Ricky Sutopo, Joanne Mun-Yee Lim, Koksheik Wong |
Multim. Tools Appl. | 4 |
| 2021 | Appearance-based passenger counting in cluttered scenes with lateral movement compensation
Ricky Sutopo, Joanne Mun-Yee Lim, Vishnu Monn Baskaran, Koksheik Wong, Massimo Tistarelli, Heng Fui Liau |
Neural Comput. Appl. | 4 |
| 2021 | Strengthening speech content authentication against tampering
Raphael C.-W. Phan, Yin Yin Low, Koksheik Wong, Kazuki Minemura |
Speech Commun. | 3 |
| 2021 | Efficient Known-Sample Attack for Distance-Preserving Hashing Biometric Template Protection SchemesabstractThe rapid deployment of biometric authentication systems raises concern over user privacy and security. A biometric template protection scheme emerges as a solution to protect individual biometric templates stored in a database. Among all available protection schemes, a template protection scheme that relies on distance-preserving hashing has received much attention due to its simplicity and efficiency in offering privacy protection while archiving decent authentication performance. In this work, we introduce an efficient attack called known sample attack and demonstrate that most state-of-art template protection schemes that utilize distance-preserving hashing can be compromised in practice (within few seconds), especially when the output is significantly smaller than the original input sample size. These findings further motivated our subsequent work in proposing a secure authentication mechanism to resist such an attack with proper study over the distribution of the input samples. Furthermore, we conducted revocability, unlinkability analysis to demonstrate the satisfactory of general biometric template protection requirements; and showed the resistance of various security and privacy attacks, i.e., false acceptance attack, and attack via record multiplicity. Yen-Lung Lai, Zhe Jin 0001, Koksheik Wong, Massimo Tistarelli |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2020 | Information Hiding In Image EnhancementabstractThis paper proposes an information hiding method to embed data while executing image enhancement steps. The 2D Median Filter is adapted and re-engineered to demonstrate the feasibility of this concept. In particular, the filteringembedding steps are performed for each pixel in a sliding window manner. Pixels enclosed within the predefined window (neighborhood) are gathered, linearized and sorted. Then, the linearized pixels are divided into partitions, in which each partition is assigned to represent a certain sequence of bits. The performance of the proposed method is evaluated by using the BSD300 dataset for various settings. The embedding capacity, image quality, data extraction error rate are reported and analyzed. Besides, the robustness of the proposed method against brute force attack is also discussed. In the best case scenario, when the window size is 7 x 7, ~ 0.97 bpp is achieved with acceptable image quality while having ~ 3.5% data extraction error rate. Simying Ong, Koksheik Wong |
ICIP | 2 |
| 2020 | Complete Quality Preserving Data Hiding in Animated GIF with Reversibility and Scalable Capacity Functionalities
Koksheik Wong, Mohamed N. M. Nazeeb, Jean-Luc Dugelay |
IWDW | 1 |
| 2020 | Reduced contact lifting of latent fingerprints from curved surfaces
Mohammad Mogharen Askarin, Koksheik Wong, Raphael C.-W. Phan |
J. Inf. Secur. Appl. | 2 |
| 2019 | A Data Embedding Technique for Spatial Scalable Coded Video Using Motion Vector PredictorabstractThis work aims to embed data into scalable coded video by manipulating the syntax elements related to motion vectors. The payload is improved by forcing the prediction blocks to be coded by using smaller block sizes, which directly results in having more manipulable syntax elements to embed data. A threshold is introduced to guide the rate-distortion optimization and data embedding processes. Experimental results suggest that higher payload is achieved when smaller prediction blocks are coded at slightly higher bit rate, with negligible degradation in perceptual video quality. In the best case scenario, 475, 691 bits are embedded into the PartyScene test video sequence while PSNR drops by 0.06dB with a bit rate overhead of 4.72%. LieLin Pang, Koksheik Wong |
ICIP | 2 |
| 2019 | Format-Compliant Perceptual Encryption Method for JPEG XTabstractHigh dynamic range (HDR) image allows more details to be seen in a digital image, where objects in the darker (e.g., shadow) and brighter (e.g., sky) regions are made more visible. Therefore, there is an increasing number of devices that include the HDR imaging capturing capability. In this work, a method is put forward to encrypt high dynamic range image encoded in the JPEG XT standard. To achieve diffusion, DC coefficients in the base layer are shuffled in groups of variable sizes. Similarly, AC coefficients in the base and residual layers are shuffled in the form of (zero-run-size, value) pair. Our encryption method is able to achieve: (a) suppression of bit stream size expansion, (b) sufficiently mask the visual appearance of an HDR image, and (c) format-compliant. Experiments are carried out to verify the basic performance of the proposed method. In the best case scenario, an HDR image can be completely distorted to mask its perceptual semantics with slightly reduced file size. Jeffrey Ting, Koksheik Wong, Simying Ong |
ICIP | 2 |
| 2019 | Improved DM-QIM Watermarking Scheme for PDF Document
Minoru Kuribayashi, Koksheik Wong |
IWDW | 2 |
| 2019 | Data Hiding in Perceptually Masked OpenEXR ImageabstractHigh dynamic range (HDR) imaging is able to capture and display significantly more colors when compared to the legacy imaging technology. Recently, HDR environment is gaining popularity and becoming common in our daily life, which resulted in the growing number of HDR images. In this work, a technique is put forward to first perceptually mask a HDR image, and then hide data into the masked HDR image. Specifically, pixels in the OpenEXR file, which are stored in half floating point precision format, are manipulated. First, a predictor is utilized to predict pixel values, where well predicted pixel locations are flagged and utilized as the venues to hide data. On the other hand, the ill predicted pixel locations are flagged as unusable. Next, each pixel is divided into segments of 5 bits, and the segments are XOR-ed and permuted to mask the perceptual semantic of the HDR image. Data hiding then takes place at locations flagged as usable, which further distorts the quality of the image. The proposed method is both reversible and separable. The basic performance of the proposed joint technique is evaluated by using 6 HDR images. In the best case scenario, 91% of pixels in the image can be utilized for data hiding purpose while the perceptual semantic of the original image is completely masked. KaiLin Chia, Koksheik Wong, Jean-Luc Dugelay |
MMSP | 2 |
| 2019 | A Secure Visual-thermal Fused Face Recognition System Based on Non-Linear HashingabstractIn this paper, we propose a secure visual-thermal fused face recognition system using non-linear hashing. To extract features from both thermal and visible facial images, a deep neural network model pre-trained by visible images, namely InsightFace, is utilized in extracting deep features from both thermal and visible images. Next, we investigate into the effectiveness of using nonlinear hashing in protecting deep features extracted from both thermal and visible face images. To further boost the accuracy performance of the facial recognition system under unfavorable environment, feature- and score-level fusion of thermal and visible images for face matching are studied. The performance of different application scenarios are tested on the EURECOM VIS-TH face dataset. Experiment results suggest that: 1) feature- and score-level fusion techniques are effective in achieving higher accuracy under unfavorable situation; 2) non-linear hashing offers additional layer of protection, namely, privacy preservation, to face image. We also found that the deep model trained by using visible images is applicable to thermal images for feature extraction, which is particularly useful because there is no large thermal dataset available to train deep neural network. Xingbo Dong, Koksheik Wong, Zhe Jin 0001, Jean-Luc Dugelay |
MMSP | 2 |
| 2019 | Rewritable Data Embedding in JPEG XT Image using Coefficient Count Across LayersabstractA JPEG XT format-compliant data embedding method is put forward by exploiting the natural relationship of coefficient count in the base and residual layers. Specifically, the number of coefficients for each MCU in the base layer is often larger than that of the respective MCU in the residual layer. This happens when the quality factor of the base layer is ≥ 75 and the quality factor of the residual layer is lower than that of the base layer. The proposed method then swaps the MCUs within these two layers to embed data. In case when a MCU is indistinguishable, a marker is injected to explicitly mark this bad MCU. When only AC coefficients are swapped, the output image quality can be maintained with some expansion of bit stream size. On the other hand, when both DC and AC coefficients are swapped, the proposed method is able to generate a fully distorted output image, hence masking the perceptual semantic of the image. Experiments are carried out to evaluate the performance of the proposed data embedding method. Jeffrey Ting, Koksheik Wong, Simying Ong |
MMSP | 2 |
| 2019 | A new hybrid ensemble feature selection framework for machine learning-based phishing detection systemabstractThis paper proposes a new feature selection framework for machine learning-based phishing detection system, called the Hybrid Ensemble Feature Selection (HEFS). In the first phase of HEFS, a novel Cumulative Distribution Function gradient (CDF-g) algorithm is exploited to produce primary feature subsets, which are then fed into a data perturbation ensemble to yield secondary feature subsets. The second phase derives a set of baseline features from the secondary feature subsets by using a function perturbation ensemble. The overall experimental results suggest that HEFS performs best when it is integrated with Random Forest classifier , where the baseline features correctly distinguish 94.6% of phishing and legitimate websites using only 20.8% of the original features. In another experiment, the baseline features (10 in total) utilised on Random Forest outperforms the set of all features (48 in total) used on SVM, Naive Bayes, C4.5, JRip, and PART classifiers. HEFS also shows promising results when benchmarked using another well-known phishing dataset from the University of California Irvine (UCI) repository. Hence, the HEFS is a highly desirable and practical feature selection technique for machine learning-based phishing detection systems. Kang-Leng Chiew, Colin Choon Lin Tan, Koksheik Wong, Kelvin Sheng Chek Yong, Wei King Tiong |
Inf. Sci. | 3 |
| 2018 | Dominant speaker detection in multipoint video communication using Markov chain with non-linear weights and dynamic transition window
Vishnu Monn Baskaran, Yoong Choon Chang, Jonathan Loo, Koksheik Wong, Ming-Tao Gan |
Inf. Sci. | 4 |
| 2018 | Separable authentication in encrypted HEVC video
Yiqi Tew, Koksheik Wong, Raphael C.-W. Phan, King Ngi Ngan |
Multim. Tools Appl. | 2 |
| 2018 | Less is more: Micro-expression recognition from video using apex frame
Sze-Teng Liong, John See, Koksheik Wong, Raphael C.-W. Phan |
Signal Process. Image Commun. | 3 |
| 2017 | A scrambling framework for block transform compressed image
Kazuki Minemura, Koksheik Wong, Xiaojun Qi 0001, Kiyoshi Tanaka |
Multim. Tools Appl. | 2 |
| 2017 | Image encryption method based on chaotic fuzzy cellular neural networks
Kuru Ratnavelu, M. Kalpana, P. Balasubramaniam 0001, Koksheik Wong, Raveendran Paramesran |
Signal Process. | 4 |
| 2017 | Fast recovery of unknown coefficients in DCT-transformed images
Simying Ong, Shujun Li 0001, Koksheik Wong, KuanYew Tan |
Signal Process. Image Commun. | 3 |
| 2017 | A Novel Sketch Attack for H.264/AVC Format-Compliant Encrypted VideoabstractIn this paper, we propose a novel sketch attack for H.264 advanced video coding (H.264/AVC) format-compliant encrypted video. We briefly describe the notion of sketch attack, review the conventional sketch attacks designed for discrete cosine transform (DCT)-based compressed image, and identify their shortcomings when applied to attack compressed video. Specifically, the conventional DCT-based sketch attacks are incapable in sketching outlines for inter frame, which is deployed to significantly reduce temporal redundancy in video compression. To sketch directly from inter frame, we put forward a sketch attack by considering the partially decoded information of the H.264/AVC compressed video, namely, the number of bits spent on coding a macroblock. To evaluate the sketch image, we consider the Canny edge map as the ideal outline image. Experiments are conducted to verify the performance of the proposed sketch attack using ICADR2013, High Efficiency Video Coding dash, and Xiph video data sets. Results suggest that the proposed sketch attack can generate the outline image of the original frame for not only intra frame but also inter frame. Kazuki Minemura, Koksheik Wong, Raphael C.-W. Phan, Kiyoshi Tanaka |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Ocular Recognition for Blinking EyesabstractOcular recognition is expected to provide a higher flexibility in handling practical applications as oppose to the iris recognition, which only works for the ideal open-eye case. However, the accuracy of the recent efforts is still far from satisfactory at uncontrollable conditions, such as eye blinking which implies any poses of eyes. To address these issues, the skin texture, eyelids, and additional geometrical features are employed. In addition, to achieve higher accuracy, sequential forward floating selection is utilized to select the best feature combinations. Finally, the non-linear support vector machine is applied for identification purpose. Experimental results demonstrate that the proposed algorithm achieves the best accuracy for both open eye and blinking eye scenarios. As a result, it offers greater flexibility for the prospective subjects during recognition as well as higher reliability for security. Peizhong Liu, Jing-Ming Guo, Szu-Han Tseng, Koksheik Wong, Jiann-Der Lee, Chen-Chieh Yao, Daxin Zhu |
IEEE Trans. Image Process. | 4 |
| 2016 | PhishWHO: Phishing webpage detection via identity keywords extraction and target domain name finderabstractThis paper proposes a phishing detection technique based on the difference between the target and actual identities of a webpage. The proposed phishing detection approach, called PhishWHO, can be divided into three phases. The first phase extracts identity keywords from the textual contents of the website, where a novel weighted URL tokens system based on the N-gram model is proposed. The second phase finds the target domain name by using a search engine, and the target domain name is selected based on identity-relevant features. In the final phase, a 3-tier identity matching system is proposed to determine the legitimacy of the query webpage. The overall experimental results suggest that the proposed system outperforms the conventional phishing detection methods considered. Colin Choon Lin Tan, Kang-Leng Chiew, Koksheik Wong, San-Nah Sze |
Decis. Support Syst. | 3 |
| 2016 | AIPISteg: An active IP identification based steganographic method
Osamah Ibrahiem Abdullaziz, Vik Tor Goh, Huo-Chong Ling, Koksheik Wong |
J. Netw. Comput. Appl. | 4 |
| 2016 | Halftoning-based Block Truncation Coding image restoration
Jing-Ming Guo, Heri Prasetyo, Koksheik Wong |
J. Vis. Commun. Image Represent. | 3 |
| 2016 | Multi-layer authentication scheme for HEVC video based on embedded statistics
Yiqi Tew, Koksheik Wong, Raphael C.-W. Phan, King Ngi Ngan |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | Reversible data hiding by adaptive group modification on histogram of prediction errors
Reza Moradi Rad, Koksheik Wong, Jing-Ming Guo |
Signal Process. | 2 |
| 2016 | Fast watermarking scheme for real-time spatial scalable video coding
Adamu Muhammad Buhari, Huo-Chong Ling, Vishnu Monn Baskaran, Koksheik Wong |
Signal Process. Image Commun. | 4 |
| 2016 | Spontaneous subtle expression detection and recognition based on facial strain
Sze-Teng Liong, John See, Raphael C.-W. Phan, Yee-Hui Oh, Anh Cat Le Ngo, Koksheik Wong, Su-Wei Tan |
Signal Process. Image Commun. | 6 |
| 2016 | Accurate Facial Landmark ExtractionabstractFacial landmark extraction system is crucial in various applications, including face recognition, expression analysis, face tracking, and face animation. This letter aims to improve the performance of an existing landmark extraction method proposed by Ren et al. in terms of error rate. Specifically, the Gaussian blur filter is applied on the input image to reduce noise interference and the theta-based split rule is deployed to strengthen the performance of the random forests. Then, global linear regression is applied instead of treating each landmark independently. Experimental results demonstrate that the proposed modified facial landmark extraction algorithm outperforms the conventional methods for both the LFPW and Helen databases. Jing-Ming Guo, Szu-Han Tseng, Koksheik Wong |
IEEE Signal Process. Lett. | 3 |
| 2015 | HEVC video authentication using data embedding techniqueabstractA HEVC video authentication scheme by utilizing data embedding technique is proposed. The concept of authentication, layout and implementation are described under the HEVC standard. The authentication scheme includes weight generation, video feature extraction and two layers of authentication. Simulation results confirm that the overall perceptual video quality is maintained after the insertion of authentication code into commonly considered classes of video sequence. By analyzing the behavior of video tampering within and across video slices, the proposed authentication scheme is able to detect the tampered region and verify the integrity of the video in question. Yiqi Tew, Koksheik Wong, Raphael C.-W. Phan |
ICIP | 2 |
| 2015 | Design and implementation of parallel video combiner architecture for multi-user video conferencing at ultra-high definition resolution
Vishnu Monn Baskaran, Yoong Choon Chang, Jonathan Loo, Koksheik Wong |
Multim. Tools Appl. | 4 |
| 2015 | Data embedding in random domain
Mustafa S. Abdul Karim, Koksheik Wong |
Signal Process. | 2 |
| 2015 | Scrambling-embedding for JPEG compressed image
Simying Ong, Koksheik Wong, Kiyoshi Tanaka |
Signal Process. | 2 |
| 2015 | Beyond format-compliant encryption for JPEG image
Simying Ong, Koksheik Wong, Xiaojun Qi 0001, Kiyoshi Tanaka |
Signal Process. Image Commun. | 2 |
| 2014 | Optical flow based dynamic curved video text detectionabstractText detection in video is a challenging problem as it is useful in several real time applications in the field of video indexing and retrieval. Unlike existing methods that generally focus on horizontal caption or graphics text, the proposed method focuses on detecting dynamic curved text in video. The method explores the characteristics of the optical flow of text, namely, constant velocity, uniform magnitude distribution and unique angle distribution, to identify text candidates with the help of k-means clustering algorithm. We propose an iterative procedure which finds the standard deviation of text candidates between the first and its successive frames, and it terminates when there is a sudden decrease in the standard deviation values. The proposed method eliminates false text candidates based on the characteristics of optical flow at component level while retaining the potential text candidates. Then, direction guided boundary growing is proposed to traverse curved text lines in video. Furthermore, the characteristics of optical flow of text are utilized at block level to eliminate false positives. Experiments are conducted with various videos, including video with static text, static and dynamic text, and dynamic text only, to evaluate the proposed method. The results are benchmarked with the existing methods to verify the superiority of our method over the existing methods in terms of recall, precision, F-measure and average processing time. Palaiahnakote Shivakumara, Mohamed Lubani, Koksheik Wong, Tong Lu 0002 |
ICIP | 3 |
| 2014 | Information hiding in HEVC standard using adaptive coding block size decisionabstractIn this work, an information hiding techniques is proposed using the coding block size decision in HEVC. This approach manipulates the CB (coding block) size decision on every coding tree unit to embed information based on the predefined mapping rules. Each CB is forced to assume certain size to encode the external information without significantly compromising perceptual quality. To improve payload, the odd-even based information hiding technique is further deployed by manipulating the nonzero DCT coefficients in certain ranges, in which case each range depends on the CB size. Results suggest that by combining both approaches, improvement is achieved in the terms of payload for the higher bitrate scenario and insignificant degradation in perceptual video quality for the low bitrate scenario. Yiqi Tew, Koksheik Wong |
ICIP | 2 |
| 2014 | Data fusion in universal domain using dual semantic code
Mustafa S. Abdul Karim, Koksheik Wong |
Inf. Sci. | 2 |
| 2014 | Universal data embedding in encrypted domain
Mustafa S. Abdul Karim, Koksheik Wong |
Signal Process. | 2 |
| 2014 | A Scalable Reversible Data Embedding Method with progressive quality degradation functionality
Simying Ong, Koksheik Wong, Kiyoshi Tanaka |
Signal Process. Image Commun. | 2 |
| 2014 | Vehicle Verification Using Gabor Filter Magnitude with Gamma Distribution ModelingabstractThis letter presents a new method to derive the image feature descriptor for vehicle verification. The effectiveness of the proposed feature descriptor is based on the nature of the Gabor filter magnitude that tends to obey the Gamma distribution. The statistical parameters of the Gabor magnitude are computed using the Maximum Likelihood Estimation (MLE), which is later utilized to construct the feature descriptor. Conventionally, the Gabor magnitude is simply modeled by using Gaussian distribution, and thus the image descriptor consists of mean, standard deviation, and skewness values of the Gabor filter magnitude. However, recent investigations found that the skewness parameter is not contributing towards class separation. Based on our observation, the Gamma distribution provides a better statistical fitting to represent the Gabor filter magnitude when compared to the Gaussian distribution. As documented in the experimental results, the proposed feature descriptor yields higher accuracy for vehicle verification when compared to the conventional schemes. Jing-Ming Guo, Heri Prasetyo, Koksheik Wong |
IEEE Signal Process. Lett. | 3 |
| 2014 | An Overview of Information Hiding in H.264/AVC Compressed VideoabstractInformation hiding refers to the process of inserting information into a host to serve specific purpose(s). In this paper, information hiding methods in the H.264/AVC compressed video domain are surveyed. First, the general framework of information hiding is conceptualized by relating the state of an entity to a meaning (i.e., sequences of bits). This concept is illustrated by using various data representation schemes such as bit plane replacement, spread spectrum, histogram manipulation, divisibility, mapping rules, and matrix encoding. Venues at which information hiding takes place are then identified, including prediction process, transformation, quantization, and entropy coding. Related information hiding methods at each venue are briefly reviewed, along with the presentation of the targeted applications, appropriate diagrams, and references. A timeline diagram is constructed to chronologically summarize the invention of information hiding methods in the compressed still image and video domains since 1992. A comparison among the considered information hiding methods is also conducted in terms of venue, payload, bitstream size overhead, video quality, computational complexity, and video criteria. Further perspectives and recommendations are presented to provide a better understanding of the current trend of information hiding and to identify new opportunities for information hiding in compressed video. Yiqi Tew, Koksheik Wong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | A Unified Data Embedding and Scrambling MethodabstractConventionally, data embedding techniques aim at maintaining high-output image quality so that the difference between the original and the embedded images is imperceptible to the naked eye. Recently, as a new trend, some researchers exploited reversible data embedding techniques to deliberately degrade image quality to a desirable level of distortion. In this paper, a unified data embedding-scrambling technique called UES is proposed to achieve two objectives simultaneously, namely, high payload and adaptive scalable quality degradation. First, a pixel intensity value prediction method called checkerboard-based prediction is proposed to accurately predict 75% of the pixels in the image based on the information obtained from 25% of the image. Then, the locations of the predicted pixels are vacated to embed information while degrading the image quality. Given a desirable quality (quantified in SSIM) for the output image, UES guides the embedding-scrambling algorithm to handle the exact number of pixels, i.e., the perceptual quality of the embedded-scrambled image can be controlled. In addition, the prediction errors are stored at a predetermined precision using the structure side information to perfectly reconstruct or approximate the original image. In particular, given a desirable SSIM value, the precision of the stored prediction errors can be adjusted to control the perceptual quality of the reconstructed image. Experimental results confirmed that UES is able to perfectly reconstruct or approximate the original image with SSIM value > 0.99 after completely degrading its perceptual quality while embedding at 7.001 bpp on average. Reza Moradi Rad, Koksheik Wong, Jing-Ming Guo |
IEEE Trans. Image Process. | 2 |
| 2013 | Progressive quality degradation in JPEG compressed image using DC block orientation with rewritable data embedding functionalityabstractThis paper proposes a novel block rotational method to degrade quality and embed external data in JPEG compressed image. The orientation of each non-overlapping DC coefficients block is exploited to embed information while introducing distortion. To achieve progressive quality degradation, size of DC coefficients block is manipulated and the proposed embedding process is applied recursively by shrinking block size in each iteration. Markers are added into the blocks as pre-processing steps to ensure that the original orientation always yields the smallest difference. A post-processing is also proposed to erase the marker introduced for recovering image at higher quality, making the proposed method a rewritable method but not complete reversible. Experiments are conducted to verify the basic performance of the proposed method and comparisons with the conventional methods are also carried out. Simying Ong, Kazuki Minemura, Koksheik Wong |
ICIP | 3 |
| 2013 | An efficient sign prediction method for DCT coefficients and its application to reversible data embedding in scrambled JPEG imageabstractIn this paper, an efficient DCT sign prediction method is proposed. Unlike the conventional methods that depend on information from both spatial and frequency domains, the proposed method operates solely in the frequency domain by exploiting the pixel value patterns represented by the corresponding DCT basis vectors. In particular, each block is classified into five categories, namely, complex-pattern, complex-nonpattern, simple-smooth, simple-pattern and simple-texture, and each is treated differently using the proposed predictor. The proposed sign prediction method is then applied to realize reversible data embedding using sign information in a scrambled JPEG compressed image. This work is the first of its kind in using sign information for data embedding purposes. Basic performance of the proposed sign prediction and the proposed reversible data embedding method in scrambled JPEG image are verified using standard test images. Reza Moradi Rad, Koksheik Wong |
ICIP | 2 |
| 2013 | Rotational based rewritable data hiding in JPEGabstractThis paper proposes a novel rotational method on AC coefficient pairs to embed data into a JPEG compressed image. The purpose is to improve the carrier capacity while maintaining its original feature in controlling quality degradation. The proposed method exploits two properties in the quantized AC coefficients, namely, large magnitude and short run of zeros for low frequency subbands, and vice versa. The coefficients are first grouped into pairs and then rotated to the left or right directions to create distinctive states, where each can be utilized to represent external data. The AC coefficients are not modified and no additional AC coefficients are introduced for data embedding. However, preprocessing is needed so that all blocks satisfy the properties assumed to ensure correct data extraction and image recovery. The proposed method is rewritable because the host image can be re-utilized without causing further distortion. Experiments were conducted to verify the basic performance of the proposed method. On average, the proposed method is able to embed up to ~9318 bits in the test images of quality factor 80. Simying Ong, Koksheik Wong |
VCIP | 2 |
| 2013 | Quality degradative reversible data embedding using pixel replacementabstractConventionally, reversible data embedding methods aim at maintaining high output image quality while sacrificing carrier capacity. Recently, as a new trend, some researchers exploited reversible data embedding techniques to severely degrade image quality. In this paper, a novel high carrier capacity data embedding technique is proposed to achieve quality degradation. An efficient pixel value estimation method called checkerboard based prediction is proposed and exploited to realize data embedding while achieving scrambling effect. Here, locations of the predicted pixels are vacated to embed information while degrading the image quality. Basic performance of the proposed method is verified through experiments using various standard test images. In the best case scenario, carrier capacity of 7.31 bpp is achieved while the image is severely degraded. Reza Moradi Rad, Koksheik Wong |
VCIP | 2 |
| 2013 | Software-based serverless endpoint video combiner architecture for high-definition multiparty video conferencing
Vishnu Monn Baskaran, Yoong Choon Chang, Jonathan Loo, Koksheik Wong |
J. Netw. Comput. Appl. | 4 |
| 2012 | Dynamic semantic feature-based long-term cross-session learning approach to content-based image retrievalabstractThis paper proposes a novel content-based image retrieval technique, which facilitates short-term (intra-query) and long-term (inter-query) learning processes by integrating accumulated users' historical relevance feedback-based semantic knowledge. The history is efficiently represented as a dynamic semantic feature of the images. As such, the high-level semantic similarity measure can be dynamically adapted based on the semantic relevance derived from the dynamic semantic features. The short-term relevance feedback technique can benefit from long-term learning. Our extensive experiments show that the proposed system outperforms three peer systems in the context of both correct and erroneous relevance feedback. Zhongmiao Xiao, Matthew J. Clark, Koksheik Wong, Xiaojun Qi 0001 |
ICASSP | 3 |
| 2012 | Learning a weighted semantic manifold for content-based image retrievalabstractWe propose a novel weighted semantic manifold ranking system for content-based image retrieval. This manifold builds a more accurate intrinsic structure for the proper image space by combining visual and semantic relevance relations. Specifically, we apply the learning mechanism to capture users' semantic concepts in clusters and extract high-level semantic features for each database image. We then incorporate the reliability score, the fuzzy membership, and the composite low-level and high-level relation into the traditional affinity matrix to construct a weighted semantic manifold structure. We finally create an asymmetric relevance vector to propagate positive and negative labels via the proposed manifold structure to images with high similarities. Extensive experiments demonstrate our system outperforms other manifold systems and learning systems in the context of both correct and erroneous feedback. Ran Chang, Zhongmiao Xiao, Koksheik Wong, Xiaojun Qi 0001 |
ICIP | 3 |
| 2012 | JPEG image scrambling without expansion in bitstream sizeabstractIn this work, an algorithm is proposed to scramble an JPEG compressed image without causing bitstream size expansion. The causes of bitstream size expansion in the existing scrambling methods are first identified. Three recommendations on AC coefficients in the scrambled image are proposed to combat unauthorized viewing. As the first step of the scrambling algorithm, edges are identified directly in the frequency domain using solely AC coefficients without relying on any traditional methods. These edges then form a low resolution image of its original counterpart and the information is utilized to identify regions. The DC coefficients are encoded in region-basis to suppress bitstream size expansion while achieving scrambling effect. Experiments were carried out to verify the basic performance of the proposed scrambling method. For the parameter settings considered, most of the scrambled images are of smaller bitstream size than their original counter parts. Kazuki Minemura, Zahra Moayed, Koksheik Wong, Xiaojun Qi 0001, Kiyoshi Tanaka |
ICIP | 3 |
| 2012 | UniSpaCh: A text-based data hiding method using Unicode space characters
Lip Yee Por, Koksheik Wong, Kok Onn Chee |
J. Syst. Softw. | 2 |
| 2010 | Scalable image scrambling method using unified constructive permutation function on diagonal blocksabstractIn this paper, an extension of ScaScra [1] is proposed to scal-ably scramble an image in the diagonal direction for achieving distorted scanline-like effect. The non-overlapping diagonal blocks are first defined and the unified constructive permutation function is applied to scramble pixels in each diagonal block. Scalability in scrambling is achieved by varying the block size. Experiments were carried out to objectively and subjectively verify the basic performance of the proposed extension and compare them to the results of ScaScra by using standard test images. Evaluations on pixel correlation and entropy are also carried out to verify the performance of both ScaScra and the proposed extension. Koksheik Wong, Kiyoshi Tanaka |
PCS | 1 |
| 2009 | Complete Video Quality-Preserving Data HidingabstractAlthough many data hiding methods are proposed in the literature, all of them distort the quality of the host content during data embedding. In this paper, we propose a novel data hiding method in the compressed video domain that completely preserves the image quality of the host video while embedding information into it. Information is embedded into a compressed video by simultaneously manipulating Mquant and quantized discrete cosine transform coefficients, which are the significant parts of MPEG and H.26x-based compression standards. To the best of our knowledge, this data hiding method is the first attempt of its kind. When fed into an ordinary video decoder, the modified video completely reconstructs the original video even compared at the bit-to-bit level. Our method is also reversible, where the embedded information could be removed to obtain the original video. A new data representation scheme called reverse zerorun length (RZL) is proposed to exploit the statistics of macroblock for achieving high embedding efficiency while trading off with payload. It is theoretically and experimentally verified that RZL outperforms matrix encoding in terms of payload and embedding efficiency for this particular data hiding method. The problem of video bitstream size increment caused by data embedding is also addressed, and two independent solutions are proposed to suppress this increment. Basic performance of this data hiding method is verified through experiments on various existing MPEG-1 encoded videos. In the best case scenario, an average increase of four bits in the video bitstream size is observed for every message bit embedded. Koksheik Wong, Kiyoshi Tanaka, Koichi Takagi, Yasuyuki Nakajima |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2008 | An efficient data representation scheme for complete video quality preserving data hidingabstractThis paper proposes an efficient data representation scheme to improve the performance of a data hiding method [1] in MPEG compressed domain. Even though [1] completely preserves the quality of the modified video to that of the original (compressed) video and [1] is reversible, [1] suffers from consistent filesize increase caused by data embedding. To suppress filesize increase, reverse zerorun length (RZL) is proposed to efficiently encode the message. RZL utilizes the statistics of the macroblocks with respect to [1], and the distance between two excited macroblocks is considered to encode a message segment. RZL simultaneously achieves high payload and high embedding efficiency, thus RZL is able to suppress the filesize increase caused by data embedding. We theoretically analyzed that RZL outperformsmatrix encoding for both payload and embedding efficiency for this particular data hiding method. Experiments are also carried out to verify the theoretically deduced results, and the observed results agree with the expected outcomes. Koksheik Wong, Kiyoshi Tanaka, Koichi Takagi, Yasuyuki Nakajima |
ICME | 1 |
| 2007 | A DCT-based Mod4 steganographic method
Koksheik Wong, Xiaojun Qi 0001, Kiyoshi Tanaka |
Signal Process. | 1 |
| 2005 | An adaptive DCT-based Mod-4 steganographic methodabstractThis paper presents a novel Mod-4 steganographic method in discrete cosine transform (DCT) domain. A group of 2 /spl times/ 2 quantized DCT coefficients (GQC) is selected as the valid embedding area if more than two DCT coefficients are outside the interval of [-1,1]. The modulo 4 arithmetic operation is further applied to all the valid GQCs to embed a pair of binary bits using the shortest-route modification scheme. Each secret message is also encrypted to provide the system with more security. The proposed system has been extensively tested on a variety of images with different textures. Experimental results demonstrate that our system successfully preserves the quality of the images and stays undetected by the well-known steganalysis methods. Xiaojun Qi 0001, Koksheik Wong |
ICIP (2) | 2 |